{"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original resolution policy and the disputed September 14 renewal, current-contact correction request, account, and timing bindings. The two focus-evidence spans are complete factual sentences: one defines what field E measures and the other reports its base value. The counterfactual changes only that value from 6 to 9 days, without leaving a conflicting measurement elsewhere. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "full_context_fact_states": {"base": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "counterfactual": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "remove_left": {"request_within_seven_days": "unknown"}, "remove_right": {"request_within_seven_days": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"request_within_seven_days": "unknown"}, "negative_pair": {"request_within_seven_days": "refuted"}, "negative_sentence": {"request_within_seven_days": "unknown"}, "positive_pair": {"request_within_seven_days": "supported"}, "right": {"request_within_seven_days": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single relevant factual relationship; none is a bundled final resolution classification. The focus is the factual timing relationship between the correction request and renewal. Base and counter assignments can be realized by changing only the request date from within the seven-day window to outside it while holding the remaining facts fixed. Empty policy_evidence is correct because all substantive governing rules are already preserved in the unchanged questions object; the original state supplies case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a recorded and valid offer, pre-renewal conditional acceptance tied to applying the credit, omission of the credit, and a correction request within seven calendar days. These facts satisfy the credit-and-annual-resolution rubric. Refutation of two settled charges also excludes duplicate-payment routing.", "rule_index": 0, "sound": true}, {"reason": "With the other eligibility facts fixed, a correction request explicitly outside the seven-calendar-day limit makes the credit ineligible. The conjunction also establishes each expressly requested fallback component—annual cancellation, movement to monthly, and the unused-term refund—and refutes the charge multiplicity required for duplicate-payment routing.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "offer_recorded", "statement": "A $24 retention offer was recorded on the subscriber's account for the disputed September 14 annual renewal."}, {"id": "offer_valid", "statement": "The $24 retention offer recorded on the subscriber's account was valid for the disputed September 14 annual renewal."}, {"id": "acceptance_before_renewal", "statement": "The subscriber accepted the recorded $24 retention offer before the disputed September 14 annual renewal posted."}, {"id": "acceptance_condition_credit", "statement": "The condition in the subscriber's acceptance of the recorded offer was application of the eligible $24 retention credit."}, {"id": "credit_omitted", "statement": "The $24 retention credit was omitted from invoice INV-8841 for the disputed September 14 annual renewal."}, {"id": "request_within_seven_days", "statement": "The subscriber's correction request in the current contact was made no later than seven calendar days after the disputed September 14 annual renewal."}, {"id": "fallback_cancel", "statement": "The subscriber expressly requested cancellation of the annual subscription if the $24 retention credit was ineligible."}, {"id": "fallback_monthly", "statement": "The subscriber expressly requested a move to a monthly subscription if the $24 retention credit was ineligible."}, {"id": "fallback_refund", "statement": "The subscriber expressly requested the permitted refund of the unused annual term if the $24 retention credit was ineligible."}, {"id": "two_settled_charges", "statement": "The subscriber's account shows at least two settled subscription charges for the disputed September 14 annual renewal."}], "base_state_json": "[{\"speaker\":\"Account reviewer\",\"text\":\"The subscriber’s account records a $24 retention offer for the disputed September 14 annual renewal, and the offer was valid for that renewal.\"},{\"speaker\":\"Communications reviewer\",\"text\":\"Before the renewal posted, the subscriber accepted the recorded offer on the condition that the eligible $24 retention credit be applied.\"},{\"speaker\":\"Invoice reviewer\",\"text\":\"Invoice INV-8841 omitted the $24 retention credit from the disputed September 14 annual renewal.\"},{\"speaker\":\"Subscriber request record\",\"text\":\"In the current contact, the subscriber expressly requested that, if the credit was ineligible, the annual subscription be canceled, the account be moved to a monthly subscription, and the permitted refund for the unused annual term be issued.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact.\"},{\"speaker\":\"Audit record\",\"text\":\"Field E in audit record AR-7319 contains the value 6.\"},{\"speaker\":\"Payment reviewer\",\"text\":\"Payment reconciliation shows exactly one settled subscription charge for the disputed September 14 annual renewal, and no second settled subscription charge appears.\"}]", "base_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "counter_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "focus_atom": "request_within_seven_days", "focus_evidence": [{"path": ["4", "text"], "text": "Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact."}, {"path": ["5", "text"], "text": "Field E in audit record AR-7319 contains the value 6."}], "policy_evidence": [], "rules": [{"justification": "The omitted $24 credit is eligible because the valid recorded offer was accepted before renewal, the acceptance condition was application of that eligible credit, and correction was requested within seven calendar days. Refutation of two settled charges excludes duplicate-payment routing, and the actionable primary request excludes clarification and fallback.", "target": "apply_credit_keep_annual", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}, {"justification": "An explicitly late correction request makes the omitted credit ineligible under the seven-calendar-day requirement. The subscriber expressly supplied every component of the authorized fallback. Refutation of two settled charges excludes duplicate-payment routing, and the actionable fallback excludes clarification and none of the above.", "target": "execute_fallback_cancellation", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}]}, "verified_pair": {"left": "Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact.", "negative_left": "Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact.", "negative_right": "Field E in audit record AR-7319 contains the value 9.", "right": "Field E in audit record AR-7319 contains the value 6."}, "verifier_independent_model": false}, "family": "scale-diverse-013-002", "id": "scale-diverse-013-002-base", "input": {"questions": {"decision": {"criteria": {"apply_credit_keep_annual": "Add the $24 credit and preserve the annual subscription when the offer and pre-renewal acceptance are verified and the correction request is within seven calendar days.", "execute_fallback_cancellation": "Cancel annual, move to monthly, and process the permitted unused-term refund only when the credit is ineligible and the subscriber expressly supplied this fallback.", "none_of_above": "Use only when the evidence satisfies none of the four resolution rubrics above.", "route_duplicate_payment_review": "Route for duplicate-payment investigation only when the account shows at least two settled subscription charges for the disputed renewal.", "seek_intent_clarification": "Ask the subscriber to clarify only when neither the primary conditional request nor any fallback identifies an actionable outcome."}, "instructions": "Choose the authorized resolution. Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges.", "type": "choice"}}, "state": [{"speaker": "Account reviewer", "text": "The subscriber’s account records a $24 retention offer for the disputed September 14 annual renewal, and the offer was valid for that renewal."}, {"speaker": "Communications reviewer", "text": "Before the renewal posted, the subscriber accepted the recorded offer on the condition that the eligible $24 retention credit be applied."}, {"speaker": "Invoice reviewer", "text": "Invoice INV-8841 omitted the $24 retention credit from the disputed September 14 annual renewal."}, {"speaker": "Subscriber request record", "text": "In the current contact, the subscriber expressly requested that, if the credit was ineligible, the annual subscription be canceled, the account be moved to a monthly subscription, and the permitted refund for the unused annual term be issued."}, {"speaker": "Audit reviewer", "text": "Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact."}, {"speaker": "Audit record", "text": "Field E in audit record AR-7319 contains the value 6."}, {"speaker": "Payment reviewer", "text": "Payment reconciliation shows exactly one settled subscription charge for the disputed September 14 annual renewal, and no second settled subscription charge appears."}]}, "method": "c2d", "provenance": {"source_id": "diverse-013", "source_is_synthetic": true, "source_sha256": "640e873611af9a131e8dfe8de5027d31b2e53121b8e09c71e14040fbe16c6d8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "apply_credit_keep_annual"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original resolution policy and the disputed September 14 renewal, current-contact correction request, account, and timing bindings. The two focus-evidence spans are complete factual sentences: one defines what field E measures and the other reports its base value. The counterfactual changes only that value from 6 to 9 days, without leaving a conflicting measurement elsewhere. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "full_context_fact_states": {"base": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "counterfactual": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "remove_left": {"request_within_seven_days": "unknown"}, "remove_right": {"request_within_seven_days": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"request_within_seven_days": "unknown"}, "negative_pair": {"request_within_seven_days": "refuted"}, "negative_sentence": {"request_within_seven_days": "unknown"}, "positive_pair": {"request_within_seven_days": "supported"}, "right": {"request_within_seven_days": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single relevant factual relationship; none is a bundled final resolution classification. The focus is the factual timing relationship between the correction request and renewal. Base and counter assignments can be realized by changing only the request date from within the seven-day window to outside it while holding the remaining facts fixed. Empty policy_evidence is correct because all substantive governing rules are already preserved in the unchanged questions object; the original state supplies case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a recorded and valid offer, pre-renewal conditional acceptance tied to applying the credit, omission of the credit, and a correction request within seven calendar days. These facts satisfy the credit-and-annual-resolution rubric. Refutation of two settled charges also excludes duplicate-payment routing.", "rule_index": 0, "sound": true}, {"reason": "With the other eligibility facts fixed, a correction request explicitly outside the seven-calendar-day limit makes the credit ineligible. The conjunction also establishes each expressly requested fallback component—annual cancellation, movement to monthly, and the unused-term refund—and refutes the charge multiplicity required for duplicate-payment routing.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "offer_recorded", "statement": "A $24 retention offer was recorded on the subscriber's account for the disputed September 14 annual renewal."}, {"id": "offer_valid", "statement": "The $24 retention offer recorded on the subscriber's account was valid for the disputed September 14 annual renewal."}, {"id": "acceptance_before_renewal", "statement": "The subscriber accepted the recorded $24 retention offer before the disputed September 14 annual renewal posted."}, {"id": "acceptance_condition_credit", "statement": "The condition in the subscriber's acceptance of the recorded offer was application of the eligible $24 retention credit."}, {"id": "credit_omitted", "statement": "The $24 retention credit was omitted from invoice INV-8841 for the disputed September 14 annual renewal."}, {"id": "request_within_seven_days", "statement": "The subscriber's correction request in the current contact was made no later than seven calendar days after the disputed September 14 annual renewal."}, {"id": "fallback_cancel", "statement": "The subscriber expressly requested cancellation of the annual subscription if the $24 retention credit was ineligible."}, {"id": "fallback_monthly", "statement": "The subscriber expressly requested a move to a monthly subscription if the $24 retention credit was ineligible."}, {"id": "fallback_refund", "statement": "The subscriber expressly requested the permitted refund of the unused annual term if the $24 retention credit was ineligible."}, {"id": "two_settled_charges", "statement": "The subscriber's account shows at least two settled subscription charges for the disputed September 14 annual renewal."}], "base_state_json": "[{\"speaker\":\"Account reviewer\",\"text\":\"The subscriber’s account records a $24 retention offer for the disputed September 14 annual renewal, and the offer was valid for that renewal.\"},{\"speaker\":\"Communications reviewer\",\"text\":\"Before the renewal posted, the subscriber accepted the recorded offer on the condition that the eligible $24 retention credit be applied.\"},{\"speaker\":\"Invoice reviewer\",\"text\":\"Invoice INV-8841 omitted the $24 retention credit from the disputed September 14 annual renewal.\"},{\"speaker\":\"Subscriber request record\",\"text\":\"In the current contact, the subscriber expressly requested that, if the credit was ineligible, the annual subscription be canceled, the account be moved to a monthly subscription, and the permitted refund for the unused annual term be issued.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact.\"},{\"speaker\":\"Audit record\",\"text\":\"Field E in audit record AR-7319 contains the value 6.\"},{\"speaker\":\"Payment reviewer\",\"text\":\"Payment reconciliation shows exactly one settled subscription charge for the disputed September 14 annual renewal, and no second settled subscription charge appears.\"}]", "base_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "counter_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "focus_atom": "request_within_seven_days", "focus_evidence": [{"path": ["4", "text"], "text": "Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact."}, {"path": ["5", "text"], "text": "Field E in audit record AR-7319 contains the value 6."}], "policy_evidence": [], "rules": [{"justification": "The omitted $24 credit is eligible because the valid recorded offer was accepted before renewal, the acceptance condition was application of that eligible credit, and correction was requested within seven calendar days. Refutation of two settled charges excludes duplicate-payment routing, and the actionable primary request excludes clarification and fallback.", "target": "apply_credit_keep_annual", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}, {"justification": "An explicitly late correction request makes the omitted credit ineligible under the seven-calendar-day requirement. The subscriber expressly supplied every component of the authorized fallback. Refutation of two settled charges excludes duplicate-payment routing, and the actionable fallback excludes clarification and none of the above.", "target": "execute_fallback_cancellation", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}]}, "verified_pair": {"left": "Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact.", "negative_left": "Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact.", "negative_right": "Field E in audit record AR-7319 contains the value 9.", "right": "Field E in audit record AR-7319 contains the value 6."}, "verifier_independent_model": false}, "family": "scale-diverse-013-002", "id": "scale-diverse-013-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"apply_credit_keep_annual": "Add the $24 credit and preserve the annual subscription when the offer and pre-renewal acceptance are verified and the correction request is within seven calendar days.", "execute_fallback_cancellation": "Cancel annual, move to monthly, and process the permitted unused-term refund only when the credit is ineligible and the subscriber expressly supplied this fallback.", "none_of_above": "Use only when the evidence satisfies none of the four resolution rubrics above.", "route_duplicate_payment_review": "Route for duplicate-payment investigation only when the account shows at least two settled subscription charges for the disputed renewal.", "seek_intent_clarification": "Ask the subscriber to clarify only when neither the primary conditional request nor any fallback identifies an actionable outcome."}, "instructions": "Choose the authorized resolution. Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges.", "type": "choice"}}, "state": [{"speaker": "Account reviewer", "text": "The subscriber’s account records a $24 retention offer for the disputed September 14 annual renewal, and the offer was valid for that renewal."}, {"speaker": "Communications reviewer", "text": "Before the renewal posted, the subscriber accepted the recorded offer on the condition that the eligible $24 retention credit be applied."}, {"speaker": "Invoice reviewer", "text": "Invoice INV-8841 omitted the $24 retention credit from the disputed September 14 annual renewal."}, {"speaker": "Subscriber request record", "text": "In the current contact, the subscriber expressly requested that, if the credit was ineligible, the annual subscription be canceled, the account be moved to a monthly subscription, and the permitted refund for the unused annual term be issued."}, {"speaker": "Audit reviewer", "text": "Audit record AR-7319 states that its field E records the number of calendar days between the subscriber's disputed September 14 annual renewal and the correction request in the current contact."}, {"speaker": "Audit record", "text": "Field E in audit record AR-7319 contains the value 9."}, {"speaker": "Payment reviewer", "text": "Payment reconciliation shows exactly one settled subscription charge for the disputed September 14 annual renewal, and no second settled subscription charge appears."}]}, "method": "c2d", "provenance": {"source_id": "diverse-013", "source_is_synthetic": true, "source_sha256": "640e873611af9a131e8dfe8de5027d31b2e53121b8e09c71e14040fbe16c6d8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "execute_fallback_cancellation"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual changes only T’s latest-update time, an allowed case observation, while preserving the routing policy, Mira Chen, the verified-unresolved issue set, and the latest-update decision path. The two evidence spans are complete factual sentences. Moving T’s update from 14:32 to 06:41 makes R later than T and remains consistent with the unchanged assertion that the update times are unequal. Neither context embeds an answer choice, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"On 18 August 2026, an audit of subscriber Mira Chen’s billing file identified issues T and R as verified and still unresolved. Every verified unresolved billing issue for Mira was either T or R. Issue T concerned only correction of sales tax on an invoice. Issue R concerned the base renewal charge, which had been billed at the wrong standard price. R did not concern tax, a promotion, or a plan change, and its disputed invoice was not later voided. Exactly one settled payment covered the disputed subscription period, a count below two. The merchant had not issued a refund for R. The file recorded unequal latest timeline-update times. The latest timeline update for billing issue T was recorded at 14:32 UTC on 18 August 2026. The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The latest timeline update for billing issue T was recorded at 14:32 UTC on 18 August 2026."}, {"path": [], "text": "The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The latest timeline update for billing issue T was recorded at 14:32 UTC on 18 August 2026.", "negative_left": "The latest timeline update for billing issue T was recorded at 06:41 UTC on 18 August 2026.", "negative_right": "The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026.", "right": "The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-014-001", "id": "scale-diverse-014-001-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "On 18 August 2026, an audit of subscriber Mira Chen’s billing file identified issues T and R as verified and still unresolved. Every verified unresolved billing issue for Mira was either T or R. Issue T concerned only correction of sales tax on an invoice. Issue R concerned the base renewal charge, which had been billed at the wrong standard price. R did not concern tax, a promotion, or a plan change, and its disputed invoice was not later voided. Exactly one settled payment covered the disputed subscription period, a count below two. The merchant had not issued a refund for R. The file recorded unequal latest timeline-update times. The latest timeline update for billing issue T was recorded at 14:32 UTC on 18 August 2026. The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual changes only T’s latest-update time, an allowed case observation, while preserving the routing policy, Mira Chen, the verified-unresolved issue set, and the latest-update decision path. The two evidence spans are complete factual sentences. Moving T’s update from 14:32 to 06:41 makes R later than T and remains consistent with the unchanged assertion that the update times are unequal. Neither context embeds an answer choice, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"On 18 August 2026, an audit of subscriber Mira Chen’s billing file identified issues T and R as verified and still unresolved. Every verified unresolved billing issue for Mira was either T or R. Issue T concerned only correction of sales tax on an invoice. Issue R concerned the base renewal charge, which had been billed at the wrong standard price. R did not concern tax, a promotion, or a plan change, and its disputed invoice was not later voided. Exactly one settled payment covered the disputed subscription period, a count below two. The merchant had not issued a refund for R. The file recorded unequal latest timeline-update times. The latest timeline update for billing issue T was recorded at 14:32 UTC on 18 August 2026. The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The latest timeline update for billing issue T was recorded at 14:32 UTC on 18 August 2026."}, {"path": [], "text": "The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The latest timeline update for billing issue T was recorded at 14:32 UTC on 18 August 2026.", "negative_left": "The latest timeline update for billing issue T was recorded at 06:41 UTC on 18 August 2026.", "negative_right": "The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026.", "right": "The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-014-001", "id": "scale-diverse-014-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "On 18 August 2026, an audit of subscriber Mira Chen’s billing file identified issues T and R as verified and still unresolved. Every verified unresolved billing issue for Mira was either T or R. Issue T concerned only correction of sales tax on an invoice. Issue R concerned the base renewal charge, which had been billed at the wrong standard price. R did not concern tax, a promotion, or a plan change, and its disputed invoice was not later voided. Exactly one settled payment covered the disputed subscription period, a count below two. The merchant had not issued a refund for R. The file recorded unequal latest timeline-update times. The latest timeline update for billing issue T was recorded at 06:41 UTC on 18 August 2026. The latest timeline update for billing issue R was recorded at 09:17 UTC on 18 August 2026."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy and Mira Chen decision scope without adding exceptions, priorities, or missing-evidence defaults. The two focus-evidence spans are complete factual sentences. The counterfactual changes only T’s latest-update timestamp, coherently making R later than T while leaving all other facts unchanged and noncontradictory. Neither context includes an answer label, code, rule table, proposition identifier, label rationale, or output instruction; its route-related language states case facts using permitted natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"An evidence reconciliation for subscriber Mira Chen identified the following records. Billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Billing issue R is likewise verified and unresolved; it concerns the base renewal charge, which was billed at the wrong standard price. R does not concern tax, a promotion, or a plan change. The invoice disputed in R was not later voided. Fewer than two settled payments cover the subscription period disputed in R, and the merchant has not issued a refund for that issue. Every verified unresolved billing issue for Mira Chen is either T or R. The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-14 16:42:11 UTC. The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-14 16:42:11 UTC."}, {"path": [], "text": "The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-14 16:42:11 UTC.", "negative_left": "The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-10 18:53:24 UTC.", "negative_right": "The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC.", "right": "The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-014-002", "id": "scale-diverse-014-002-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "An evidence reconciliation for subscriber Mira Chen identified the following records. Billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Billing issue R is likewise verified and unresolved; it concerns the base renewal charge, which was billed at the wrong standard price. R does not concern tax, a promotion, or a plan change. The invoice disputed in R was not later voided. Fewer than two settled payments cover the subscription period disputed in R, and the merchant has not issued a refund for that issue. Every verified unresolved billing issue for Mira Chen is either T or R. The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-14 16:42:11 UTC. The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy and Mira Chen decision scope without adding exceptions, priorities, or missing-evidence defaults. The two focus-evidence spans are complete factual sentences. The counterfactual changes only T’s latest-update timestamp, coherently making R later than T while leaving all other facts unchanged and noncontradictory. Neither context includes an answer label, code, rule table, proposition identifier, label rationale, or output instruction; its route-related language states case facts using permitted natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"An evidence reconciliation for subscriber Mira Chen identified the following records. Billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Billing issue R is likewise verified and unresolved; it concerns the base renewal charge, which was billed at the wrong standard price. R does not concern tax, a promotion, or a plan change. The invoice disputed in R was not later voided. Fewer than two settled payments cover the subscription period disputed in R, and the merchant has not issued a refund for that issue. Every verified unresolved billing issue for Mira Chen is either T or R. The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-14 16:42:11 UTC. The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-14 16:42:11 UTC."}, {"path": [], "text": "The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-14 16:42:11 UTC.", "negative_left": "The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-10 18:53:24 UTC.", "negative_right": "The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC.", "right": "The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-014-002", "id": "scale-diverse-014-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "An evidence reconciliation for subscriber Mira Chen identified the following records. Billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Billing issue R is likewise verified and unresolved; it concerns the base renewal charge, which was billed at the wrong standard price. R does not concern tax, a promotion, or a plan change. The invoice disputed in R was not later voided. Fewer than two settled payments cover the subscription period disputed in R, and the merchant has not issued a refund for that issue. Every verified unresolved billing issue for Mira Chen is either T or R. The reconciled billing timeline for Mira Chen records billing issue T's latest update at 2026-08-10 18:53:24 UTC. The reconciled billing timeline for Mira Chen records billing issue R's latest update at 2026-08-12 09:17:36 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the routing policy, criteria, and scope, while both contexts retain the subscriber, billing-issue entity path, unresolved-status framing, and relevant timeline bindings. The two evidence spans are complete factual sentences. The counterfactual changes only issue T’s latest timeline-update timestamp, making it earlier than issue R without creating duplicate or contradictory measurements. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Operational handoff for subscriber Mira Chen: billing issues T and R are both verified and remain unresolved. The queue audit confirms they are the complete set of Mira’s verified unresolved billing issues. Issue T concerns only correction of sales tax on an invoice. The billing handoff log records 2026-08-14 16:42:18 UTC as the latest timeline-update time for billing issue T. Issue R concerns the base renewal charge, which was billed at the wrong standard price. It does not concern tax, a promotion, or a plan change, and the disputed invoice was not later voided. Exactly one settled payment covers the subscription period disputed in R, and the merchant has not issued a refund for that issue. The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The billing handoff log records 2026-08-14 16:42:18 UTC as the latest timeline-update time for billing issue T."}, {"path": [], "text": "The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The billing handoff log records 2026-08-14 16:42:18 UTC as the latest timeline-update time for billing issue T.", "negative_left": "The billing handoff log records 2026-08-14 14:31:26 UTC as the latest timeline-update time for billing issue T.", "negative_right": "The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R.", "right": "The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R."}, "verifier_independent_model": false}, "family": "scale-diverse-014-003", "id": "scale-diverse-014-003-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Operational handoff for subscriber Mira Chen: billing issues T and R are both verified and remain unresolved. The queue audit confirms they are the complete set of Mira’s verified unresolved billing issues. Issue T concerns only correction of sales tax on an invoice. The billing handoff log records 2026-08-14 16:42:18 UTC as the latest timeline-update time for billing issue T. Issue R concerns the base renewal charge, which was billed at the wrong standard price. It does not concern tax, a promotion, or a plan change, and the disputed invoice was not later voided. Exactly one settled payment covers the subscription period disputed in R, and the merchant has not issued a refund for that issue. The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the routing policy, criteria, and scope, while both contexts retain the subscriber, billing-issue entity path, unresolved-status framing, and relevant timeline bindings. The two evidence spans are complete factual sentences. The counterfactual changes only issue T’s latest timeline-update timestamp, making it earlier than issue R without creating duplicate or contradictory measurements. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Operational handoff for subscriber Mira Chen: billing issues T and R are both verified and remain unresolved. The queue audit confirms they are the complete set of Mira’s verified unresolved billing issues. Issue T concerns only correction of sales tax on an invoice. The billing handoff log records 2026-08-14 16:42:18 UTC as the latest timeline-update time for billing issue T. Issue R concerns the base renewal charge, which was billed at the wrong standard price. It does not concern tax, a promotion, or a plan change, and the disputed invoice was not later voided. Exactly one settled payment covers the subscription period disputed in R, and the merchant has not issued a refund for that issue. The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The billing handoff log records 2026-08-14 16:42:18 UTC as the latest timeline-update time for billing issue T."}, {"path": [], "text": "The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The billing handoff log records 2026-08-14 16:42:18 UTC as the latest timeline-update time for billing issue T.", "negative_left": "The billing handoff log records 2026-08-14 14:31:26 UTC as the latest timeline-update time for billing issue T.", "negative_right": "The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R.", "right": "The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R."}, "verifier_independent_model": false}, "family": "scale-diverse-014-003", "id": "scale-diverse-014-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Operational handoff for subscriber Mira Chen: billing issues T and R are both verified and remain unresolved. The queue audit confirms they are the complete set of Mira’s verified unresolved billing issues. Issue T concerns only correction of sales tax on an invoice. The billing handoff log records 2026-08-14 14:31:26 UTC as the latest timeline-update time for billing issue T. Issue R concerns the base renewal charge, which was billed at the wrong standard price. It does not concern tax, a promotion, or a plan change, and the disputed invoice was not later voided. Exactly one settled payment covers the subscription period disputed in R, and the merchant has not issued a refund for that issue. The billing handoff log records 2026-08-14 15:07:53 UTC as the latest timeline-update time for billing issue R."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the routing policy, criteria, and latest-issue priority, while neither context adds an exception, priority, or missing-evidence default. Both contexts retain Mira Chen, the billing-issue scope, and the latest-update comparison; the counterfactual changes only issue R’s timestamp, coherently making it later than T without conflicting with the stated inequality or any duplicate measurement. The two evidence spans are complete factual sentences. Neither context contains an answer code, explicit gold label, rule table, proposition identifier, label rationale, or output instruction; its route-related wording is ordinary factual policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Field note: Subscriber Mira Chen’s billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Her billing issue R is also verified and unresolved. No verified unresolved billing issue for Mira exists other than T or R.\\n\\nBilling issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC. Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 11:05 UTC. The two timestamps are unequal.\\n\\nIssue R concerns the base renewal charge, which was billed at the wrong standard price. It does not concern tax, a promotion, or a plan change. The disputed invoice was not later voided. The disputed subscription period is covered by exactly one settled payment; its settled-payment count is below two. The merchant has not issued a refund for issue R.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC."}, {"path": [], "text": "Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 11:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC.", "negative_left": "Billing issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC.", "negative_right": "Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 17:40 UTC.", "right": "Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 11:05 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-014-004", "id": "scale-diverse-014-004-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Field note: Subscriber Mira Chen’s billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Her billing issue R is also verified and unresolved. No verified unresolved billing issue for Mira exists other than T or R.\n\nBilling issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC. Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 11:05 UTC. The two timestamps are unequal.\n\nIssue R concerns the base renewal charge, which was billed at the wrong standard price. It does not concern tax, a promotion, or a plan change. The disputed invoice was not later voided. The disputed subscription period is covered by exactly one settled payment; its settled-payment count is below two. The merchant has not issued a refund for issue R."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the routing policy, criteria, and latest-issue priority, while neither context adds an exception, priority, or missing-evidence default. Both contexts retain Mira Chen, the billing-issue scope, and the latest-update comparison; the counterfactual changes only issue R’s timestamp, coherently making it later than T without conflicting with the stated inequality or any duplicate measurement. The two evidence spans are complete factual sentences. Neither context contains an answer code, explicit gold label, rule table, proposition identifier, label rationale, or output instruction; its route-related wording is ordinary factual policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Field note: Subscriber Mira Chen’s billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Her billing issue R is also verified and unresolved. No verified unresolved billing issue for Mira exists other than T or R.\\n\\nBilling issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC. Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 11:05 UTC. The two timestamps are unequal.\\n\\nIssue R concerns the base renewal charge, which was billed at the wrong standard price. It does not concern tax, a promotion, or a plan change. The disputed invoice was not later voided. The disputed subscription period is covered by exactly one settled payment; its settled-payment count is below two. The merchant has not issued a refund for issue R.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC."}, {"path": [], "text": "Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 11:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC.", "negative_left": "Billing issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC.", "negative_right": "Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 17:40 UTC.", "right": "Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 11:05 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-014-004", "id": "scale-diverse-014-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Field note: Subscriber Mira Chen’s billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Her billing issue R is also verified and unresolved. No verified unresolved billing issue for Mira exists other than T or R.\n\nBilling issue T for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 14:20 UTC. Billing issue R for subscriber Mira Chen has its latest timeline update timestamped 2031-07-18 17:40 UTC. The two timestamps are unequal.\n\nIssue R concerns the base renewal charge, which was billed at the wrong standard price. It does not concern tax, a promotion, or a plan change. The disputed invoice was not later voided. The disputed subscription period is covered by exactly one settled payment; its settled-payment count is below two. The merchant has not issued a refund for issue R."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all routing instructions and criteria, and neither context alters or adds governing policy. Both contexts retain the same subscriber, issue set, routing path, and latest-update comparison; the counterfactual coherently changes only R’s update time, with no duplicate or contradictory timestamp. The two evidence spans are complete factual sentences. Neither context contains an answer code, explicit route selection, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Billing records for subscriber Mira Chen identify exactly two verified, unresolved issues: T and R; no other verified unresolved billing issue exists. Issue T concerns only correction of sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price. R does not concern tax, a promotion, or a plan change, and its disputed invoice was not later voided. Exactly one settled payment covers R's disputed subscription period, and the merchant has not issued a refund for R. The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC. The audit log records the latest timeline-update time for billing issue R as 2026-04-18 09:12:45 UTC. The two recorded update times are different.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC."}, {"path": [], "text": "The audit log records the latest timeline-update time for billing issue R as 2026-04-18 09:12:45 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC.", "negative_left": "The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC.", "negative_right": "The audit log records the latest timeline-update time for billing issue R as 2026-04-18 16:08:11 UTC.", "right": "The audit log records the latest timeline-update time for billing issue R as 2026-04-18 09:12:45 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-014-005", "id": "scale-diverse-014-005-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Billing records for subscriber Mira Chen identify exactly two verified, unresolved issues: T and R; no other verified unresolved billing issue exists. Issue T concerns only correction of sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price. R does not concern tax, a promotion, or a plan change, and its disputed invoice was not later voided. Exactly one settled payment covers R's disputed subscription period, and the merchant has not issued a refund for R. The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC. The audit log records the latest timeline-update time for billing issue R as 2026-04-18 09:12:45 UTC. The two recorded update times are different."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all routing instructions and criteria, and neither context alters or adds governing policy. Both contexts retain the same subscriber, issue set, routing path, and latest-update comparison; the counterfactual coherently changes only R’s update time, with no duplicate or contradictory timestamp. The two evidence spans are complete factual sentences. Neither context contains an answer code, explicit route selection, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Billing records for subscriber Mira Chen identify exactly two verified, unresolved issues: T and R; no other verified unresolved billing issue exists. Issue T concerns only correction of sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price. R does not concern tax, a promotion, or a plan change, and its disputed invoice was not later voided. Exactly one settled payment covers R's disputed subscription period, and the merchant has not issued a refund for R. The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC. The audit log records the latest timeline-update time for billing issue R as 2026-04-18 09:12:45 UTC. The two recorded update times are different.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC."}, {"path": [], "text": "The audit log records the latest timeline-update time for billing issue R as 2026-04-18 09:12:45 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC.", "negative_left": "The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC.", "negative_right": "The audit log records the latest timeline-update time for billing issue R as 2026-04-18 16:08:11 UTC.", "right": "The audit log records the latest timeline-update time for billing issue R as 2026-04-18 09:12:45 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-014-005", "id": "scale-diverse-014-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Billing records for subscriber Mira Chen identify exactly two verified, unresolved issues: T and R; no other verified unresolved billing issue exists. Issue T concerns only correction of sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price. R does not concern tax, a promotion, or a plan change, and its disputed invoice was not later voided. Exactly one settled payment covers R's disputed subscription period, and the merchant has not issued a refund for R. The audit log records the latest timeline-update time for billing issue T as 2026-04-18 14:37:20 UTC. The audit log records the latest timeline-update time for billing issue R as 2026-04-18 16:08:11 UTC. The two recorded update times are different."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the adjustment and conflict-resolution policies, as well as Mara Venn, the July 3 Pro renewal, the verification scope, and the relevant ledger path. The two focus-evidence spans are complete factual sentences; the counterfactual coherently changes only L932’s final disposition without creating an internal duplicate-measurement conflict, and neither context embeds a gold answer, output instruction, code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Subscriber Mara Venn disputes her July 3 Pro plan renewal, for which a billing support agent opened a duplicate-payment case. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.\",\"evidence\":[\"Before reviewing the ledger extract, the subscription operations analyst recorded that the renewal’s complete ledger scope consisted of exactly the two entries named in the extract below. Each scoped entry carried one payment transaction reference, and the export recorded no alias references.\",\"In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively.\",\"In that ledger, entries L817 and L932 each have the final disposition “settled capture.”\"],\"request\":\"Verify whether Mara Venn’s case qualifies for a duplicate-payment adjustment under the stated policy.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "1"], "text": "In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively."}, {"path": ["evidence", "2"], "text": "In that ledger, entries L817 and L932 each have the final disposition “settled capture.”"}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively.", "negative_left": "In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively.", "negative_right": "In that ledger, entry L817 has the final disposition “settled capture,” while entry L932 has the final disposition “capture reversed,” not “settled capture.”", "right": "In that ledger, entries L817 and L932 each have the final disposition “settled capture.”"}, "verifier_independent_model": false}, "family": "scale-diverse-015-001", "id": "scale-diverse-015-001-base", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Subscriber Mara Venn disputes her July 3 Pro plan renewal, for which a billing support agent opened a duplicate-payment case. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.", "evidence": ["Before reviewing the ledger extract, the subscription operations analyst recorded that the renewal’s complete ledger scope consisted of exactly the two entries named in the extract below. Each scoped entry carried one payment transaction reference, and the export recorded no alias references.", "In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively.", "In that ledger, entries L817 and L932 each have the final disposition “settled capture.”"], "request": "Verify whether Mara Venn’s case qualifies for a duplicate-payment adjustment under the stated policy."}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the adjustment and conflict-resolution policies, as well as Mara Venn, the July 3 Pro renewal, the verification scope, and the relevant ledger path. The two focus-evidence spans are complete factual sentences; the counterfactual coherently changes only L932’s final disposition without creating an internal duplicate-measurement conflict, and neither context embeds a gold answer, output instruction, code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Subscriber Mara Venn disputes her July 3 Pro plan renewal, for which a billing support agent opened a duplicate-payment case. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.\",\"evidence\":[\"Before reviewing the ledger extract, the subscription operations analyst recorded that the renewal’s complete ledger scope consisted of exactly the two entries named in the extract below. Each scoped entry carried one payment transaction reference, and the export recorded no alias references.\",\"In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively.\",\"In that ledger, entries L817 and L932 each have the final disposition “settled capture.”\"],\"request\":\"Verify whether Mara Venn’s case qualifies for a duplicate-payment adjustment under the stated policy.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "1"], "text": "In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively."}, {"path": ["evidence", "2"], "text": "In that ledger, entries L817 and L932 each have the final disposition “settled capture.”"}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively.", "negative_left": "In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively.", "negative_right": "In that ledger, entry L817 has the final disposition “settled capture,” while entry L932 has the final disposition “capture reversed,” not “settled capture.”", "right": "In that ledger, entries L817 and L932 each have the final disposition “settled capture.”"}, "verifier_independent_model": false}, "family": "scale-diverse-015-001", "id": "scale-diverse-015-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Subscriber Mara Venn disputes her July 3 Pro plan renewal, for which a billing support agent opened a duplicate-payment case. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.", "evidence": ["Before reviewing the ledger extract, the subscription operations analyst recorded that the renewal’s complete ledger scope consisted of exactly the two entries named in the extract below. Each scoped entry carried one payment transaction reference, and the export recorded no alias references.", "In the final settlement ledger at 2026-07-06T14:20:00Z, payment transaction references P441 and P442 identify entries L817 and L932, respectively.", "In that ledger, entry L817 has the final disposition “settled capture,” while entry L932 has the final disposition “capture reversed,” not “settled capture.”"], "request": "Verify whether Mara Venn’s case qualifies for a duplicate-payment adjustment under the stated policy."}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts retain the original eligibility rule and ledger-priority rule, as well as Mara Venn, the July 3 Pro renewal, the duplicate-payment-adjustment request, and the relevant ledger time/path. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only P442’s final disposition from settled to reversed without conflicting with the unchanged transaction count or sole-record assertions. Neither constructed context embeds an answer, code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Subscriber Mara Venn disputes her July 3 Pro plan renewal, and billing support opened a duplicate-payment adjustment review. Operations reconciled the renewal record against the gateway index and found exactly two associated payment transaction references: P441 and P442; no additional reference was linked to that renewal. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.\",\"evidence\":[\"The reconciliation was completed using the renewal identifier attached to Mara Venn’s July 3 Pro plan charge rather than the account’s other billing activity.\",\"In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442.\",\"In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a settled capture.\"],\"request\":\"Verify whether this case qualifies for a duplicate-payment adjustment under the stated policy.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "1"], "text": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442."}, {"path": ["evidence", "2"], "text": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a settled capture."}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442.", "negative_left": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442.", "negative_right": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a reversed capture.", "right": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a settled capture."}, "verifier_independent_model": false}, "family": "scale-diverse-015-002", "id": "scale-diverse-015-002-base", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Subscriber Mara Venn disputes her July 3 Pro plan renewal, and billing support opened a duplicate-payment adjustment review. Operations reconciled the renewal record against the gateway index and found exactly two associated payment transaction references: P441 and P442; no additional reference was linked to that renewal. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.", "evidence": ["The reconciliation was completed using the renewal identifier attached to Mara Venn’s July 3 Pro plan charge rather than the account’s other billing activity.", "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442.", "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a settled capture."], "request": "Verify whether this case qualifies for a duplicate-payment adjustment under the stated policy."}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts retain the original eligibility rule and ledger-priority rule, as well as Mara Venn, the July 3 Pro renewal, the duplicate-payment-adjustment request, and the relevant ledger time/path. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only P442’s final disposition from settled to reversed without conflicting with the unchanged transaction count or sole-record assertions. Neither constructed context embeds an answer, code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Subscriber Mara Venn disputes her July 3 Pro plan renewal, and billing support opened a duplicate-payment adjustment review. Operations reconciled the renewal record against the gateway index and found exactly two associated payment transaction references: P441 and P442; no additional reference was linked to that renewal. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.\",\"evidence\":[\"The reconciliation was completed using the renewal identifier attached to Mara Venn’s July 3 Pro plan charge rather than the account’s other billing activity.\",\"In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442.\",\"In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a settled capture.\"],\"request\":\"Verify whether this case qualifies for a duplicate-payment adjustment under the stated policy.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "1"], "text": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442."}, {"path": ["evidence", "2"], "text": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a settled capture."}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442.", "negative_left": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442.", "negative_right": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a reversed capture.", "right": "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a settled capture."}, "verifier_independent_model": false}, "family": "scale-diverse-015-002", "id": "scale-diverse-015-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Subscriber Mara Venn disputes her July 3 Pro plan renewal, and billing support opened a duplicate-payment adjustment review. Operations reconciled the renewal record against the gateway index and found exactly two associated payment transaction references: P441 and P442; no additional reference was linked to that renewal. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.", "evidence": ["The reconciliation was completed using the renewal identifier attached to Mara Venn’s July 3 Pro plan charge rather than the account’s other billing activity.", "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 is the sole final-disposition record for P441, and entry L864 is the sole final-disposition record for P442.", "In the final settlement ledger certified at 18:00 UTC on July 8 for Mara Venn’s July 3 Pro plan renewal, entry L731 records a settled capture, and entry L864 records a reversed capture."], "request": "Verify whether this case qualifies for a duplicate-payment adjustment under the stated policy."}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original eligibility and conflict-resolution policies, while the unchanged questions preserve the decision criteria and instructions. The subscriber, July 3 Pro renewal, duplicate-payment issue, ledger-based evidence path, and verification request remain bound consistently. Each context contains exactly two complete factual evidence sentences. The counterfactual coherently changes only L917’s final disposition from settled to reversed, without conflicting duplicate measurements or assertions. Neither context states the required output, an answer code, proposition ID, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Subscriber Mara Venn’s duplicate-payment review concerns her July 3 Pro plan renewal. Billing support transferred the dispute to subscription operations for an operational handoff after the ledger close. The handoff records the roster and corresponding final dispositions used for review; invoice labels and a customer-supplied bank image remain in the case file as background materials. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.\",\"evidence\":[\"The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917.\",\"At the 2026-07-05 18:00 UTC ledger close, entries L903 and L917 each carry the final disposition “settled capture.”\"],\"request\":\"Verify whether the dispute qualifies for a duplicate-payment adjustment under the stated policy.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917."}, {"path": ["evidence", "1"], "text": "At the 2026-07-05 18:00 UTC ledger close, entries L903 and L917 each carry the final disposition “settled capture.”"}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917.", "negative_left": "The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917.", "negative_right": "At the 2026-07-05 18:00 UTC ledger close, entry L903 carries the final disposition “settled capture,” while entry L917 carries the final disposition “capture reversed.”", "right": "At the 2026-07-05 18:00 UTC ledger close, entries L903 and L917 each carry the final disposition “settled capture.”"}, "verifier_independent_model": false}, "family": "scale-diverse-015-003", "id": "scale-diverse-015-003-base", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Subscriber Mara Venn’s duplicate-payment review concerns her July 3 Pro plan renewal. Billing support transferred the dispute to subscription operations for an operational handoff after the ledger close. The handoff records the roster and corresponding final dispositions used for review; invoice labels and a customer-supplied bank image remain in the case file as background materials. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.", "evidence": ["The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917.", "At the 2026-07-05 18:00 UTC ledger close, entries L903 and L917 each carry the final disposition “settled capture.”"], "request": "Verify whether the dispute qualifies for a duplicate-payment adjustment under the stated policy."}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original eligibility and conflict-resolution policies, while the unchanged questions preserve the decision criteria and instructions. The subscriber, July 3 Pro renewal, duplicate-payment issue, ledger-based evidence path, and verification request remain bound consistently. Each context contains exactly two complete factual evidence sentences. The counterfactual coherently changes only L917’s final disposition from settled to reversed, without conflicting duplicate measurements or assertions. Neither context states the required output, an answer code, proposition ID, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Subscriber Mara Venn’s duplicate-payment review concerns her July 3 Pro plan renewal. Billing support transferred the dispute to subscription operations for an operational handoff after the ledger close. The handoff records the roster and corresponding final dispositions used for review; invoice labels and a customer-supplied bank image remain in the case file as background materials. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.\",\"evidence\":[\"The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917.\",\"At the 2026-07-05 18:00 UTC ledger close, entries L903 and L917 each carry the final disposition “settled capture.”\"],\"request\":\"Verify whether the dispute qualifies for a duplicate-payment adjustment under the stated policy.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917."}, {"path": ["evidence", "1"], "text": "At the 2026-07-05 18:00 UTC ledger close, entries L903 and L917 each carry the final disposition “settled capture.”"}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917.", "negative_left": "The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917.", "negative_right": "At the 2026-07-05 18:00 UTC ledger close, entry L903 carries the final disposition “settled capture,” while entry L917 carries the final disposition “capture reversed.”", "right": "At the 2026-07-05 18:00 UTC ledger close, entries L903 and L917 each carry the final disposition “settled capture.”"}, "verifier_independent_model": false}, "family": "scale-diverse-015-003", "id": "scale-diverse-015-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Subscriber Mara Venn’s duplicate-payment review concerns her July 3 Pro plan renewal. Billing support transferred the dispute to subscription operations for an operational handoff after the ledger close. The handoff records the roster and corresponding final dispositions used for review; invoice labels and a customer-supplied bank image remain in the case file as background materials. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries.", "evidence": ["The complete payment-reference roster for Mara Venn’s July 3 Pro plan renewal in the 2026-07-05 18:00 UTC operational handoff consists of P441 mapped to final-settlement-ledger entry L903 and P442 mapped to final-settlement-ledger entry L917.", "At the 2026-07-05 18:00 UTC ledger close, entry L903 carries the final disposition “settled capture,” while entry L917 carries the final disposition “capture reversed.”"], "request": "Verify whether the dispute qualifies for a duplicate-payment adjustment under the stated policy."}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original duplicate-payment completeness policy and preserve the account, amount, currency, date, disputed-charge identities, and settlement-status inquiry. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Q7L-844’s status from posted to pending, which is coherent with the unchanged statements and creates no duplicate or contradictory status assertion within that context. Neither context includes an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Billing support agent\",\"text\":\"The invoice and processor export identify both entries as attempts to collect the same annual renewal. The subscriber disputes the extra collection.\"},{\"speaker\":\"Subscription operations analyst\",\"text\":\"On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing.\"},{\"speaker\":\"Records specialist\",\"text\":\"I compared the processor export with the settlement ledger. The transaction identifiers are consistent across both sources, and neither record is a reversal or refund.\"},{\"speaker\":\"Subscription operations analyst\",\"text\":\"In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status posted rather than pending or processing.\"},{\"speaker\":\"Billing support agent\",\"text\":\"The submitted materials are legible, and the processor export is authenticated. No additional disputed entry appears in the reviewed records.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing."}, {"path": ["3", "text"], "text": "In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status posted rather than pending or processing."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing.", "negative_left": "On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing.", "negative_right": "In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status pending rather than posted or processing.", "right": "In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status posted rather than pending or processing."}, "verifier_independent_model": false}, "family": "scale-diverse-016-001", "id": "scale-diverse-016-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Billing support agent", "text": "The invoice and processor export identify both entries as attempts to collect the same annual renewal. The subscriber disputes the extra collection."}, {"speaker": "Subscription operations analyst", "text": "On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing."}, {"speaker": "Records specialist", "text": "I compared the processor export with the settlement ledger. The transaction identifiers are consistent across both sources, and neither record is a reversal or refund."}, {"speaker": "Subscription operations analyst", "text": "In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status posted rather than pending or processing."}, {"speaker": "Billing support agent", "text": "The submitted materials are legible, and the processor export is authenticated. No additional disputed entry appears in the reviewed records."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original duplicate-payment completeness policy and preserve the account, amount, currency, date, disputed-charge identities, and settlement-status inquiry. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Q7L-844’s status from posted to pending, which is coherent with the unchanged statements and creates no duplicate or contradictory status assertion within that context. Neither context includes an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Billing support agent\",\"text\":\"The invoice and processor export identify both entries as attempts to collect the same annual renewal. The subscriber disputes the extra collection.\"},{\"speaker\":\"Subscription operations analyst\",\"text\":\"On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing.\"},{\"speaker\":\"Records specialist\",\"text\":\"I compared the processor export with the settlement ledger. The transaction identifiers are consistent across both sources, and neither record is a reversal or refund.\"},{\"speaker\":\"Subscription operations analyst\",\"text\":\"In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status posted rather than pending or processing.\"},{\"speaker\":\"Billing support agent\",\"text\":\"The submitted materials are legible, and the processor export is authenticated. No additional disputed entry appears in the reviewed records.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing."}, {"path": ["3", "text"], "text": "In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status posted rather than pending or processing."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing.", "negative_left": "On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing.", "negative_right": "In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status pending rather than posted or processing.", "right": "In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status posted rather than pending or processing."}, "verifier_independent_model": false}, "family": "scale-diverse-016-001", "id": "scale-diverse-016-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Billing support agent", "text": "The invoice and processor export identify both entries as attempts to collect the same annual renewal. The subscriber disputes the extra collection."}, {"speaker": "Subscription operations analyst", "text": "On account SF-48219, the two disputed USD 84.00 charges dated September 12 for the StreamForge Plus renewal were recorded chronologically as transaction Q7L-391 first and transaction Q7L-844 second, and Q7L-391 had settlement status posted rather than pending or processing."}, {"speaker": "Records specialist", "text": "I compared the processor export with the settlement ledger. The transaction identifiers are consistent across both sources, and neither record is a reversal or refund."}, {"speaker": "Subscription operations analyst", "text": "In the settlement-ledger snapshot taken at 14:20 UTC on September 13, transaction Q7L-844 had settlement status pending rather than posted or processing."}, {"speaker": "Billing support agent", "text": "The submitted materials are legible, and the processor export is authenticated. No additional disputed entry appears in the reviewed records."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original completeness policy without adding exceptions or altered defaults, and they remain bound to the same duplicate-renewal case, account, amount, currency, transaction date, and settlement-status inquiry. The two focus-evidence spans are complete factual sentences. The counterfactual changes only SF9-7316's status from posted to pending at the same recorded time, which is coherent with the unchanged facts and creates no duplicate contradictory status assertion within that context. Neither context includes an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Reconciliation specialist\",\"text\":\"The review packet contains exactly two disputed charge rows for the StreamForge Plus renewal. Both rows are assigned to account SF-48219, list an amount of $84.00 USD, and carry a transaction date of September 12.\"},{\"speaker\":\"Records custodian\",\"text\":\"The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal.\"},{\"speaker\":\"Settlement reviewer\",\"text\":\"At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a posted status rather than pending or processing.\"},{\"speaker\":\"Reconciliation specialist\",\"text\":\"The first disputed charge, transaction SF9-7315, is marked posted in the processor extract. The invoice and processor records use the same renewal descriptor, and the review packet preserves the two charge rows separately.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal."}, {"path": ["2", "text"], "text": "At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a posted status rather than pending or processing."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal.", "negative_left": "The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal.", "negative_right": "At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a pending status rather than posted or processing.", "right": "At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a posted status rather than pending or processing."}, "verifier_independent_model": false}, "family": "scale-diverse-016-002", "id": "scale-diverse-016-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Reconciliation specialist", "text": "The review packet contains exactly two disputed charge rows for the StreamForge Plus renewal. Both rows are assigned to account SF-48219, list an amount of $84.00 USD, and carry a transaction date of September 12."}, {"speaker": "Records custodian", "text": "The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal."}, {"speaker": "Settlement reviewer", "text": "At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a posted status rather than pending or processing."}, {"speaker": "Reconciliation specialist", "text": "The first disputed charge, transaction SF9-7315, is marked posted in the processor extract. The invoice and processor records use the same renewal descriptor, and the review packet preserves the two charge rows separately."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original completeness policy without adding exceptions or altered defaults, and they remain bound to the same duplicate-renewal case, account, amount, currency, transaction date, and settlement-status inquiry. The two focus-evidence spans are complete factual sentences. The counterfactual changes only SF9-7316's status from posted to pending at the same recorded time, which is coherent with the unchanged facts and creates no duplicate contradictory status assertion within that context. Neither context includes an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Reconciliation specialist\",\"text\":\"The review packet contains exactly two disputed charge rows for the StreamForge Plus renewal. Both rows are assigned to account SF-48219, list an amount of $84.00 USD, and carry a transaction date of September 12.\"},{\"speaker\":\"Records custodian\",\"text\":\"The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal.\"},{\"speaker\":\"Settlement reviewer\",\"text\":\"At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a posted status rather than pending or processing.\"},{\"speaker\":\"Reconciliation specialist\",\"text\":\"The first disputed charge, transaction SF9-7315, is marked posted in the processor extract. The invoice and processor records use the same renewal descriptor, and the review packet preserves the two charge rows separately.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal."}, {"path": ["2", "text"], "text": "At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a posted status rather than pending or processing."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal.", "negative_left": "The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal.", "negative_right": "At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a pending status rather than posted or processing.", "right": "At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a posted status rather than pending or processing."}, "verifier_independent_model": false}, "family": "scale-diverse-016-002", "id": "scale-diverse-016-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Reconciliation specialist", "text": "The review packet contains exactly two disputed charge rows for the StreamForge Plus renewal. Both rows are assigned to account SF-48219, list an amount of $84.00 USD, and carry a transaction date of September 12."}, {"speaker": "Records custodian", "text": "The September 12 reconciliation record for StreamForge Plus account SF-48219 identifies transaction SF9-7316 as the second of the two disputed charges for that renewal."}, {"speaker": "Settlement reviewer", "text": "At 18:40 UTC on September 13, the settlement record for transaction SF9-7316 showed a pending status rather than posted or processing."}, {"speaker": "Reconciliation specialist", "text": "The first disputed charge, transaction SF9-7315, is marked posted in the processor extract. The invoice and processor records use the same renewal descriptor, and the review packet preserves the two charge rows separately."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged question preserves the governing completeness policy, while both contexts retain the same account, renewal, disputed-charge scope, amounts, currency, dates, transaction identities, and review snapshot. The two focus-evidence spans are each complete factual sentences. The counterfactual coherently changes only TX-731619 from posted to pending without conflicting with the remaining context, and neither context contains a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Handoff coordinator\",\"text\":\"The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219.\"},{\"speaker\":\"Ledger reviewer\",\"text\":\"At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status posted.\"},{\"speaker\":\"Case custodian\",\"text\":\"The subscriber says the dispute concerns an apparent duplicate renewal debit. Processor exports were attached directly to the case and checked against the billing invoice; the transaction identifiers are legible, and no third charge is alleged. The evidence snapshot was preserved at review time. The case is awaiting a completeness decision before any duplicate-payment adjustment is considered.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["0", "text"], "text": "The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219."}, {"path": ["1", "text"], "text": "At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status posted."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219.", "negative_left": "The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219.", "negative_right": "At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status pending.", "right": "At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status posted."}, "verifier_independent_model": false}, "family": "scale-diverse-016-003", "id": "scale-diverse-016-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Handoff coordinator", "text": "The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219."}, {"speaker": "Ledger reviewer", "text": "At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status posted."}, {"speaker": "Case custodian", "text": "The subscriber says the dispute concerns an apparent duplicate renewal debit. Processor exports were attached directly to the case and checked against the billing invoice; the transaction identifiers are legible, and no third charge is alleged. The evidence snapshot was preserved at review time. The case is awaiting a completeness decision before any duplicate-payment adjustment is considered."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged question preserves the governing completeness policy, while both contexts retain the same account, renewal, disputed-charge scope, amounts, currency, dates, transaction identities, and review snapshot. The two focus-evidence spans are each complete factual sentences. The counterfactual coherently changes only TX-731619 from posted to pending without conflicting with the remaining context, and neither context contains a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Handoff coordinator\",\"text\":\"The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219.\"},{\"speaker\":\"Ledger reviewer\",\"text\":\"At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status posted.\"},{\"speaker\":\"Case custodian\",\"text\":\"The subscriber says the dispute concerns an apparent duplicate renewal debit. Processor exports were attached directly to the case and checked against the billing invoice; the transaction identifiers are legible, and no third charge is alleged. The evidence snapshot was preserved at review time. The case is awaiting a completeness decision before any duplicate-payment adjustment is considered.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["0", "text"], "text": "The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219."}, {"path": ["1", "text"], "text": "At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status posted."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219.", "negative_left": "The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219.", "negative_right": "At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status pending.", "right": "At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status posted."}, "verifier_independent_model": false}, "family": "scale-diverse-016-003", "id": "scale-diverse-016-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Handoff coordinator", "text": "The operational handoff record identifies the two disputed charges for the September 12, 2026 StreamForge Plus renewal as first transaction TX-731604 and second transaction TX-731619, each dated September 12, 2026 for USD 84.00 on account SF-48219."}, {"speaker": "Ledger reviewer", "text": "At 14:35 UTC on September 14, 2026, the settlement ledger lists transaction TX-731604 with status posted and transaction TX-731619 with status pending."}, {"speaker": "Case custodian", "text": "The subscriber says the dispute concerns an apparent duplicate renewal debit. Processor exports were attached directly to the case and checked against the billing invoice; the transaction identifiers are legible, and no third charge is alleged. The evidence snapshot was preserved at review time. The case is awaiting a completeness decision before any duplicate-payment adjustment is considered."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the renewal-refund policy, exception conditions, prorated-remedy limitation, request scope, invoice/subscription identities, and relevant times. The two evidence spans are complete factual sentences; the counterfactual coherently changes the sole audited event from Premium use to a non-Premium profile view without conflicting with the exhaustive event count, and neither context embeds an answer level, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"On September 14, 2026, at 09:00 UTC, invoice R-841 charged $120.00 to renew subscription S-204 on the Premium annual plan. An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17. Event U-17 was an invocation of a Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026. At the stated September 16 time, Maya requested that S-204 be downgraded to the Basic annual plan, whose applicable annual price was $72.00. The request came 48 hours after the charge, with 363 of 365 service days remaining. The requested correction concerns exactly one paid invoice, R-841. renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17."}, {"path": [], "text": "Event U-17 was an invocation of a Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17.", "negative_left": "An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17.", "negative_right": "Event U-17 was an account-profile view that invoked no Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026.", "right": "Event U-17 was an invocation of a Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-017-001", "id": "scale-diverse-017-001-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "On September 14, 2026, at 09:00 UTC, invoice R-841 charged $120.00 to renew subscription S-204 on the Premium annual plan. An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17. Event U-17 was an invocation of a Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026. At the stated September 16 time, Maya requested that S-204 be downgraded to the Basic annual plan, whose applicable annual price was $72.00. The request came 48 hours after the charge, with 363 of 365 service days remaining. The requested correction concerns exactly one paid invoice, R-841. renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the renewal-refund policy, exception conditions, prorated-remedy limitation, request scope, invoice/subscription identities, and relevant times. The two evidence spans are complete factual sentences; the counterfactual coherently changes the sole audited event from Premium use to a non-Premium profile view without conflicting with the exhaustive event count, and neither context embeds an answer level, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"On September 14, 2026, at 09:00 UTC, invoice R-841 charged $120.00 to renew subscription S-204 on the Premium annual plan. An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17. Event U-17 was an invocation of a Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026. At the stated September 16 time, Maya requested that S-204 be downgraded to the Basic annual plan, whose applicable annual price was $72.00. The request came 48 hours after the charge, with 363 of 365 service days remaining. The requested correction concerns exactly one paid invoice, R-841. renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17."}, {"path": [], "text": "Event U-17 was an invocation of a Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17.", "negative_left": "An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17.", "negative_right": "Event U-17 was an account-profile view that invoked no Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026.", "right": "Event U-17 was an invocation of a Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-017-001", "id": "scale-diverse-017-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "On September 14, 2026, at 09:00 UTC, invoice R-841 charged $120.00 to renew subscription S-204 on the Premium annual plan. An exhaustive audit captured every event on subscription S-204 after its September 14, 2026 renewal charge at 09:00 UTC and before Maya Chen's September 16, 2026 downgrade request at 09:00 UTC, and it captured exactly one event, event U-17. Event U-17 was an account-profile view that invoked no Premium feature on subscription S-204 at 14:22 UTC on September 15, 2026. At the stated September 16 time, Maya requested that S-204 be downgraded to the Basic annual plan, whose applicable annual price was $72.00. The request came 48 hours after the charge, with 363 of 365 service days remaining. The requested correction concerns exactly one paid invoice, R-841. renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing renewal-and-downgrade policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the relevant subscriber, invoice, subscription, plan, and time bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes FV-731 from Premium use to Basic-only use while retaining the ledger assertion that it was the sole feature-use event. Neither context contains an answer level, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing records identify invoice R-841 as a $120.00 renewal charge posted September 14, renewing subscription S-204 on the Premium annual plan. On September 16, no more than 72 hours later, Maya Chen submitted a downgrade request for S-204 targeting the Basic annual plan, priced at $72.00. At that time, 363 of 365 service days remained. Event FV-731 is classified in the service audit as Premium feature use. The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731. Her requested correction covers exactly one paid invoice, R-841. The governing policy states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "Event FV-731 is classified in the service audit as Premium feature use."}, {"path": [], "text": "The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "Event FV-731 is classified in the service audit as Premium feature use.", "negative_left": "Event FV-731 is classified in the service audit as Basic-only feature use rather than Premium feature use.", "negative_right": "The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731.", "right": "The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731."}, "verifier_independent_model": false}, "family": "scale-diverse-017-002", "id": "scale-diverse-017-002-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing records identify invoice R-841 as a $120.00 renewal charge posted September 14, renewing subscription S-204 on the Premium annual plan. On September 16, no more than 72 hours later, Maya Chen submitted a downgrade request for S-204 targeting the Basic annual plan, priced at $72.00. At that time, 363 of 365 service days remained. Event FV-731 is classified in the service audit as Premium feature use. The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731. Her requested correction covers exactly one paid invoice, R-841. The governing policy states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing renewal-and-downgrade policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the relevant subscriber, invoice, subscription, plan, and time bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes FV-731 from Premium use to Basic-only use while retaining the ledger assertion that it was the sole feature-use event. Neither context contains an answer level, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing records identify invoice R-841 as a $120.00 renewal charge posted September 14, renewing subscription S-204 on the Premium annual plan. On September 16, no more than 72 hours later, Maya Chen submitted a downgrade request for S-204 targeting the Basic annual plan, priced at $72.00. At that time, 363 of 365 service days remained. Event FV-731 is classified in the service audit as Premium feature use. The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731. Her requested correction covers exactly one paid invoice, R-841. The governing policy states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "Event FV-731 is classified in the service audit as Premium feature use."}, {"path": [], "text": "The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "Event FV-731 is classified in the service audit as Premium feature use.", "negative_left": "Event FV-731 is classified in the service audit as Basic-only feature use rather than Premium feature use.", "negative_right": "The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731.", "right": "The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731."}, "verifier_independent_model": false}, "family": "scale-diverse-017-002", "id": "scale-diverse-017-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing records identify invoice R-841 as a $120.00 renewal charge posted September 14, renewing subscription S-204 on the Premium annual plan. On September 16, no more than 72 hours later, Maya Chen submitted a downgrade request for S-204 targeting the Basic annual plan, priced at $72.00. At that time, 363 of 365 service days remained. Event FV-731 is classified in the service audit as Basic-only feature use rather than Premium feature use. The complete activity ledger for subscription S-204 records exactly one feature-use event after the September 14 renewal charge and before Maya Chen's September 16 downgrade request: event FV-731. Her requested correction covers exactly one paid invoice, R-841. The governing policy states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original renewal policy, exception conditions, prorating rule, request scope, subscriber, subscription, invoice, prices, timing, and remaining-service-day bindings. The two evidence spans are complete factual sentences. The counterfactual changes only EV-7319 from Premium feature use to non-Premium account-profile activity, which remains coherent with the one-entry exhaustive ledger. Neither context states an impact level, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Operational handoff: Invoice R-841 is a renewal charge that renewed Maya Chen’s subscription S-204 on the Premium annual plan. The $120.00 charge posted September 14 at 10:15 UTC. On September 16 at 08:40 UTC, Maya submitted a downgrade request for S-204 targeting the Basic annual plan at the applicable $72.00 price. The request followed the charge by 46 hours and 25 minutes, no more than 72 hours. At that point, 363 of 365 service days remained. The requested correction covers exactly one paid invoice, R-841.\\n\\nThe gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319. Event EV-7319 records activation of the advanced analytics dashboard, classified as Premium feature use.\\n\\nPolicy excerpts: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319."}, {"path": [], "text": "Event EV-7319 records activation of the advanced analytics dashboard, classified as Premium feature use."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319.", "negative_left": "The gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319.", "negative_right": "Event EV-7319 records an account-profile view, classified as not Premium feature use.", "right": "Event EV-7319 records activation of the advanced analytics dashboard, classified as Premium feature use."}, "verifier_independent_model": false}, "family": "scale-diverse-017-003", "id": "scale-diverse-017-003-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Operational handoff: Invoice R-841 is a renewal charge that renewed Maya Chen’s subscription S-204 on the Premium annual plan. The $120.00 charge posted September 14 at 10:15 UTC. On September 16 at 08:40 UTC, Maya submitted a downgrade request for S-204 targeting the Basic annual plan at the applicable $72.00 price. The request followed the charge by 46 hours and 25 minutes, no more than 72 hours. At that point, 363 of 365 service days remained. The requested correction covers exactly one paid invoice, R-841.\n\nThe gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319. Event EV-7319 records activation of the advanced analytics dashboard, classified as Premium feature use.\n\nPolicy excerpts: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original renewal policy, exception conditions, prorating rule, request scope, subscriber, subscription, invoice, prices, timing, and remaining-service-day bindings. The two evidence spans are complete factual sentences. The counterfactual changes only EV-7319 from Premium feature use to non-Premium account-profile activity, which remains coherent with the one-entry exhaustive ledger. Neither context states an impact level, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Operational handoff: Invoice R-841 is a renewal charge that renewed Maya Chen’s subscription S-204 on the Premium annual plan. The $120.00 charge posted September 14 at 10:15 UTC. On September 16 at 08:40 UTC, Maya submitted a downgrade request for S-204 targeting the Basic annual plan at the applicable $72.00 price. The request followed the charge by 46 hours and 25 minutes, no more than 72 hours. At that point, 363 of 365 service days remained. The requested correction covers exactly one paid invoice, R-841.\\n\\nThe gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319. Event EV-7319 records activation of the advanced analytics dashboard, classified as Premium feature use.\\n\\nPolicy excerpts: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319."}, {"path": [], "text": "Event EV-7319 records activation of the advanced analytics dashboard, classified as Premium feature use."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319.", "negative_left": "The gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319.", "negative_right": "Event EV-7319 records an account-profile view, classified as not Premium feature use.", "right": "Event EV-7319 records activation of the advanced analytics dashboard, classified as Premium feature use."}, "verifier_independent_model": false}, "family": "scale-diverse-017-003", "id": "scale-diverse-017-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Operational handoff: Invoice R-841 is a renewal charge that renewed Maya Chen’s subscription S-204 on the Premium annual plan. The $120.00 charge posted September 14 at 10:15 UTC. On September 16 at 08:40 UTC, Maya submitted a downgrade request for S-204 targeting the Basic annual plan at the applicable $72.00 price. The request followed the charge by 46 hours and 25 minutes, no more than 72 hours. At that point, 363 of 365 service days remained. The requested correction covers exactly one paid invoice, R-841.\n\nThe gap-free, exhaustive activity ledger for subscription S-204 between the September 14 renewal charge at 10:15 UTC and Maya Chen's September 16 downgrade request at 08:40 UTC contains exactly one event entry, EV-7319. Event EV-7319 records an account-profile view, classified as not Premium feature use.\n\nPolicy excerpts: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing renewal-downgrade exception and prorated-difference rule, while the unchanged questions preserve the scoring criteria and instructions. Invoice R-841, subscription S-204, Maya Chen, the relevant plans, prices, request timing, one-invoice scope, and remaining-service period remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual changes only PX-73’s date from September 15 to September 13, placing it before the September 14 renewal without contradicting the complete log or any unchanged fact. Neither context includes a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing field note: Invoice R-841 is a renewal charge that renewed subscription S-204 on the Premium annual plan at $120.00. Maya Chen’s September 16 request is a downgrade of S-204 to the Basic annual plan, whose applicable price is $72.00. The service period remaining at the request was 363 of 365 days. The correction sought covers exactly one paid invoice, R-841. The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73. The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 15, 2026. renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73."}, {"path": [], "text": "The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 15, 2026."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73.", "negative_left": "The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73.", "negative_right": "The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 13, 2026.", "right": "The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 15, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-017-004", "id": "scale-diverse-017-004-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing field note: Invoice R-841 is a renewal charge that renewed subscription S-204 on the Premium annual plan at $120.00. Maya Chen’s September 16 request is a downgrade of S-204 to the Basic annual plan, whose applicable price is $72.00. The service period remaining at the request was 363 of 365 days. The correction sought covers exactly one paid invoice, R-841. The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73. The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 15, 2026. renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing renewal-downgrade exception and prorated-difference rule, while the unchanged questions preserve the scoring criteria and instructions. Invoice R-841, subscription S-204, Maya Chen, the relevant plans, prices, request timing, one-invoice scope, and remaining-service period remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual changes only PX-73’s date from September 15 to September 13, placing it before the September 14 renewal without contradicting the complete log or any unchanged fact. Neither context includes a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing field note: Invoice R-841 is a renewal charge that renewed subscription S-204 on the Premium annual plan at $120.00. Maya Chen’s September 16 request is a downgrade of S-204 to the Basic annual plan, whose applicable price is $72.00. The service period remaining at the request was 363 of 365 days. The correction sought covers exactly one paid invoice, R-841. The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73. The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 15, 2026. renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73."}, {"path": [], "text": "The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 15, 2026."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73.", "negative_left": "The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73.", "negative_right": "The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 13, 2026.", "right": "The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 15, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-017-004", "id": "scale-diverse-017-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing field note: Invoice R-841 is a renewal charge that renewed subscription S-204 on the Premium annual plan at $120.00. Maya Chen’s September 16 request is a downgrade of S-204 to the Basic annual plan, whose applicable price is $72.00. The service period remaining at the request was 363 of 365 days. The correction sought covers exactly one paid invoice, R-841. The complete and accurate September 2026 Premium feature-use log for subscription S-204 records exactly one event, identified as event PX-73. The September 14 renewal charge occurred at 09:20 on September 14, 2026, Maya Chen's September 16 downgrade request occurred at 16:40 on September 16, 2026, and event PX-73 occurred at 11:05 on September 13, 2026. renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the criteria and routing instructions verbatim, and the request remains bound to Mara’s U-889 charge referenced by “it.” The two focus spans are complete factual sentences. The counterfactual changes only U-889’s service interval, making it consecutive to rather than duplicative of U-882’s interval, without contradicting the remaining evidence. Neither context states a selected queue, answer code, label rationale, proposition ID, or classifier-output instruction beyond the preserved governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified a13 remains atomic over an explicit transaction set. The focus, equality of the two service-coverage periods, is factual rather than policy-based. Base and counter assignments differ only on that equality and are jointly realizable: equal periods produce a duplicate, while different periods produce a single disputed prorated upgrade charge. Empty policy_evidence is correct because the governing routing criteria and priority are already preserved verbatim in original_input.questions; the original state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies U-889 as the sole referenced unresolved charge and establishes that it is a later, distinct, settled payment matching U-882 for the same subscription and service period. This is sufficient for duplicate-payment review. Renewal and promotional-credit alternatives are explicitly excluded.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an invoice-supported, settled, prorated upgrade charge whose amount is disputed. U-882 covers a different period, and a13 excludes every other settled transaction from U-889's service period, so there is no repeated settled payment for that period. Renewal and promotional-credit alternatives are also excluded, making plan-change proration review sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pronoun “it” in the billing support agent’s note refers to the charge documented by invoice U-889."}, {"id": "a2", "statement": "Invoice U-889 documents the only unresolved charge disputed by Mara."}, {"id": "a3", "statement": "Invoice U-889 is invoice evidence for its charge."}, {"id": "a4", "statement": "The payment for invoice U-889 settled."}, {"id": "a5", "statement": "The payment for invoice U-882 settled."}, {"id": "a6", "statement": "The payment for invoice U-882 occurred before the payment for invoice U-889."}, {"id": "a7", "statement": "The payments for invoices U-882 and U-889 are distinct payment transactions."}, {"id": "a8", "statement": "The charges on invoices U-882 and U-889 have equal dollar amounts."}, {"id": "a9", "statement": "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription."}, {"id": "a10", "statement": "The service-coverage period on invoice U-882 equals the service-coverage period on invoice U-889."}, {"id": "a11", "statement": "The charge on invoice U-889 arose from Mara’s upgrade from Basic to Pro."}, {"id": "a12", "statement": "Mara disputes the amount charged on invoice U-889."}, {"id": "a13", "statement": "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889."}, {"id": "a14", "statement": "The unresolved issue concerning invoice U-889 is an omitted or misapplied promotional credit."}, {"id": "a15", "statement": "The charge on invoice U-889 is a scheduled renewal charge."}, {"id": "a16", "statement": "The amount on invoice U-889 is prorated for its recorded service-coverage period."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.\",\"1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.\",\"2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.\",\"3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.\",\"4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence.\"],\"instructions\":\"Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.\",\"type\":\"score\"}},\"state\":{\"context\":\"Mara asked billing support to review two invoice-backed charges on her StreamForge Pro subscription following her upgrade from Basic.\",\"evidence\":[\"On April 3, invoice U-882 recorded a prorated $36 upgrade charge. Mara paid it in a distinct transaction, and the payment settled.\",\"Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.\",\"Invoice U-889 later recorded a separate $36 charge for the same StreamForge Pro subscription. Its distinct payment settled after the U-882 payment.\",\"Invoice U-889 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.\",\"The U-889 charge arose from the Basic-to-Pro upgrade and is prorated for its recorded coverage; it is not a scheduled renewal charge.\",\"Account review found no omitted or misapplied promotional credit. Every other settled payment transaction, apart from those for U-882 and U-889, covers a service period different from U-889’s.\",\"The agent wrote, “Invoice U-889 documents Mara’s only unresolved charge. She disputes the amount and wants it refunded.”\"],\"request\":\"Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025."}, {"path": ["state", "evidence", "3"], "text": "Invoice U-889 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025."}], "policy_evidence": [], "rules": [{"justification": "The referenced unresolved U-889 charge is a later, distinct, settled payment matching an earlier settled payment in amount, subscription, and service period. It therefore duplicates payment for the same subscription service period.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}]}, {"justification": "The sole unresolved invoice-supported dispute concerns the amount of a settled prorated upgrade charge. U-882 covers a different period, and every other settled transaction also covers a different period, so no repeated settled payment exists for U-889’s service period.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.", "negative_left": "Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.", "negative_right": "Invoice U-889 records its service coverage as the interval from 00:00 UTC on May 3, 2025, through 23:59 UTC on June 1, 2025.", "right": "Invoice U-889 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025."}, "verifier_independent_model": false}, "family": "scale-diverse-018-001", "id": "scale-diverse-018-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"context": "Mara asked billing support to review two invoice-backed charges on her StreamForge Pro subscription following her upgrade from Basic.", "evidence": ["On April 3, invoice U-882 recorded a prorated $36 upgrade charge. Mara paid it in a distinct transaction, and the payment settled.", "Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.", "Invoice U-889 later recorded a separate $36 charge for the same StreamForge Pro subscription. Its distinct payment settled after the U-882 payment.", "Invoice U-889 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.", "The U-889 charge arose from the Basic-to-Pro upgrade and is prorated for its recorded coverage; it is not a scheduled renewal charge.", "Account review found no omitted or misapplied promotional credit. Every other settled payment transaction, apart from those for U-882 and U-889, covers a service period different from U-889’s.", "The agent wrote, “Invoice U-889 documents Mara’s only unresolved charge. She disputes the amount and wants it refunded.”"], "request": "Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-018", "source_is_synthetic": true, "source_sha256": "9f3816a01540b36b15d4decda36965966708aa93a8de1d0b982b71aa44023c53", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the criteria and routing instructions verbatim, and the request remains bound to Mara’s U-889 charge referenced by “it.” The two focus spans are complete factual sentences. The counterfactual changes only U-889’s service interval, making it consecutive to rather than duplicative of U-882’s interval, without contradicting the remaining evidence. Neither context states a selected queue, answer code, label rationale, proposition ID, or classifier-output instruction beyond the preserved governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified a13 remains atomic over an explicit transaction set. The focus, equality of the two service-coverage periods, is factual rather than policy-based. Base and counter assignments differ only on that equality and are jointly realizable: equal periods produce a duplicate, while different periods produce a single disputed prorated upgrade charge. Empty policy_evidence is correct because the governing routing criteria and priority are already preserved verbatim in original_input.questions; the original state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies U-889 as the sole referenced unresolved charge and establishes that it is a later, distinct, settled payment matching U-882 for the same subscription and service period. This is sufficient for duplicate-payment review. Renewal and promotional-credit alternatives are explicitly excluded.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an invoice-supported, settled, prorated upgrade charge whose amount is disputed. U-882 covers a different period, and a13 excludes every other settled transaction from U-889's service period, so there is no repeated settled payment for that period. Renewal and promotional-credit alternatives are also excluded, making plan-change proration review sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pronoun “it” in the billing support agent’s note refers to the charge documented by invoice U-889."}, {"id": "a2", "statement": "Invoice U-889 documents the only unresolved charge disputed by Mara."}, {"id": "a3", "statement": "Invoice U-889 is invoice evidence for its charge."}, {"id": "a4", "statement": "The payment for invoice U-889 settled."}, {"id": "a5", "statement": "The payment for invoice U-882 settled."}, {"id": "a6", "statement": "The payment for invoice U-882 occurred before the payment for invoice U-889."}, {"id": "a7", "statement": "The payments for invoices U-882 and U-889 are distinct payment transactions."}, {"id": "a8", "statement": "The charges on invoices U-882 and U-889 have equal dollar amounts."}, {"id": "a9", "statement": "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription."}, {"id": "a10", "statement": "The service-coverage period on invoice U-882 equals the service-coverage period on invoice U-889."}, {"id": "a11", "statement": "The charge on invoice U-889 arose from Mara’s upgrade from Basic to Pro."}, {"id": "a12", "statement": "Mara disputes the amount charged on invoice U-889."}, {"id": "a13", "statement": "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889."}, {"id": "a14", "statement": "The unresolved issue concerning invoice U-889 is an omitted or misapplied promotional credit."}, {"id": "a15", "statement": "The charge on invoice U-889 is a scheduled renewal charge."}, {"id": "a16", "statement": "The amount on invoice U-889 is prorated for its recorded service-coverage period."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.\",\"1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.\",\"2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.\",\"3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.\",\"4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence.\"],\"instructions\":\"Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.\",\"type\":\"score\"}},\"state\":{\"context\":\"Mara asked billing support to review two invoice-backed charges on her StreamForge Pro subscription following her upgrade from Basic.\",\"evidence\":[\"On April 3, invoice U-882 recorded a prorated $36 upgrade charge. Mara paid it in a distinct transaction, and the payment settled.\",\"Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.\",\"Invoice U-889 later recorded a separate $36 charge for the same StreamForge Pro subscription. Its distinct payment settled after the U-882 payment.\",\"Invoice U-889 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.\",\"The U-889 charge arose from the Basic-to-Pro upgrade and is prorated for its recorded coverage; it is not a scheduled renewal charge.\",\"Account review found no omitted or misapplied promotional credit. Every other settled payment transaction, apart from those for U-882 and U-889, covers a service period different from U-889’s.\",\"The agent wrote, “Invoice U-889 documents Mara’s only unresolved charge. She disputes the amount and wants it refunded.”\"],\"request\":\"Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025."}, {"path": ["state", "evidence", "3"], "text": "Invoice U-889 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025."}], "policy_evidence": [], "rules": [{"justification": "The referenced unresolved U-889 charge is a later, distinct, settled payment matching an earlier settled payment in amount, subscription, and service period. It therefore duplicates payment for the same subscription service period.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}]}, {"justification": "The sole unresolved invoice-supported dispute concerns the amount of a settled prorated upgrade charge. U-882 covers a different period, and every other settled transaction also covers a different period, so no repeated settled payment exists for U-889’s service period.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.", "negative_left": "Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.", "negative_right": "Invoice U-889 records its service coverage as the interval from 00:00 UTC on May 3, 2025, through 23:59 UTC on June 1, 2025.", "right": "Invoice U-889 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025."}, "verifier_independent_model": false}, "family": "scale-diverse-018-001", "id": "scale-diverse-018-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"context": "Mara asked billing support to review two invoice-backed charges on her StreamForge Pro subscription following her upgrade from Basic.", "evidence": ["On April 3, invoice U-882 recorded a prorated $36 upgrade charge. Mara paid it in a distinct transaction, and the payment settled.", "Invoice U-882 records its service coverage as the interval from 00:00 UTC on April 3, 2025, through 23:59 UTC on May 2, 2025.", "Invoice U-889 later recorded a separate $36 charge for the same StreamForge Pro subscription. Its distinct payment settled after the U-882 payment.", "Invoice U-889 records its service coverage as the interval from 00:00 UTC on May 3, 2025, through 23:59 UTC on June 1, 2025.", "The U-889 charge arose from the Basic-to-Pro upgrade and is prorated for its recorded coverage; it is not a scheduled renewal charge.", "Account review found no omitted or misapplied promotional credit. Every other settled payment transaction, apart from those for U-882 and U-889, covers a service period different from U-889’s.", "The agent wrote, “Invoice U-889 documents Mara’s only unresolved charge. She disputes the amount and wants it refunded.”"], "request": "Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-018", "source_is_synthetic": true, "source_sha256": "9f3816a01540b36b15d4decda36965966708aa93a8de1d0b982b71aa44023c53", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing criteria and instructions verbatim, and the request continues to target Mara’s U-889 charge referenced by “it.” The two focus spans are complete factual sentences. The counterfactual changes only U-889’s coverage period to 2025-04-07 through 2025-05-06, which is coherent with the distinct invoice and payment facts and removes the same-service-period duplication without creating contradictory measurements. Neither context states a queue selection, answer code, proposition ID, rule table, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified a13 remains atomic over an explicit transaction set. The focus, equality of the two service-coverage periods, is factual rather than policy-based. Base and counter assignments differ only on that equality and are jointly realizable: equal periods produce a duplicate, while different periods produce a single disputed prorated upgrade charge. Empty policy_evidence is correct because the governing routing criteria and priority are already preserved verbatim in original_input.questions; the original state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies U-889 as the sole referenced unresolved charge and establishes that it is a later, distinct, settled payment matching U-882 for the same subscription and service period. This is sufficient for duplicate-payment review. Renewal and promotional-credit alternatives are explicitly excluded.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an invoice-supported, settled, prorated upgrade charge whose amount is disputed. U-882 covers a different period, and a13 excludes every other settled transaction from U-889's service period, so there is no repeated settled payment for that period. Renewal and promotional-credit alternatives are also excluded, making plan-change proration review sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pronoun “it” in the billing support agent’s note refers to the charge documented by invoice U-889."}, {"id": "a2", "statement": "Invoice U-889 documents the only unresolved charge disputed by Mara."}, {"id": "a3", "statement": "Invoice U-889 is invoice evidence for its charge."}, {"id": "a4", "statement": "The payment for invoice U-889 settled."}, {"id": "a5", "statement": "The payment for invoice U-882 settled."}, {"id": "a6", "statement": "The payment for invoice U-882 occurred before the payment for invoice U-889."}, {"id": "a7", "statement": "The payments for invoices U-882 and U-889 are distinct payment transactions."}, {"id": "a8", "statement": "The charges on invoices U-882 and U-889 have equal dollar amounts."}, {"id": "a9", "statement": "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription."}, {"id": "a10", "statement": "The service-coverage period on invoice U-882 equals the service-coverage period on invoice U-889."}, {"id": "a11", "statement": "The charge on invoice U-889 arose from Mara’s upgrade from Basic to Pro."}, {"id": "a12", "statement": "Mara disputes the amount charged on invoice U-889."}, {"id": "a13", "statement": "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889."}, {"id": "a14", "statement": "The unresolved issue concerning invoice U-889 is an omitted or misapplied promotional credit."}, {"id": "a15", "statement": "The charge on invoice U-889 is a scheduled renewal charge."}, {"id": "a16", "statement": "The amount on invoice U-889 is prorated for its recorded service-coverage period."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.\",\"1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.\",\"2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.\",\"3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.\",\"4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence.\"],\"instructions\":\"Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.\",\"type\":\"score\"}},\"state\":{\"context\":\"An analyst reconciled Mara’s invoice records, payment ledger, and support notes for her StreamForge account.\",\"evidence\":[\"Account reconciliation identified two issued invoice line items for Mara’s StreamForge Pro subscription, each for $36.\",\"Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.\",\"U-882 arose from her Basic-to-Pro upgrade; its prorated amount was paid by transaction P-31, which settled on March 8.\",\"Invoice U-889 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.\",\"U-889 also arose from that upgrade and labels its $36 as prorated for its recorded coverage; distinct transaction P-32 settled on March 9.\",\"It is a one-time plan-change charge, not a scheduled renewal.\",\"Mara disputes U-889’s amount, and it is her only unresolved disputed charge.\",\"The agent noted beneath the U-889 charge entry, “Mara wants it refunded.”\",\"No discount or promotional credit was advertised, promised, documented, omitted, or misapplied.\",\"The complete ledger’s only other settled payment is R-410, covering 2025-02-07 through 2025-03-07; no additional settled transactions exist.\"],\"request\":\"Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive."}, {"path": ["state", "evidence", "3"], "text": "Invoice U-889 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive."}], "policy_evidence": [], "rules": [{"justification": "The referenced unresolved U-889 charge is a later, distinct, settled payment matching an earlier settled payment in amount, subscription, and service period. It therefore duplicates payment for the same subscription service period.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}]}, {"justification": "The sole unresolved invoice-supported dispute concerns the amount of a settled prorated upgrade charge. U-882 covers a different period, and every other settled transaction also covers a different period, so no repeated settled payment exists for U-889’s service period.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.", "negative_left": "Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.", "negative_right": "Invoice U-889 records a service-coverage period from 2025-04-07 through 2025-05-06, inclusive.", "right": "Invoice U-889 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive."}, "verifier_independent_model": false}, "family": "scale-diverse-018-002", "id": "scale-diverse-018-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"context": "An analyst reconciled Mara’s invoice records, payment ledger, and support notes for her StreamForge account.", "evidence": ["Account reconciliation identified two issued invoice line items for Mara’s StreamForge Pro subscription, each for $36.", "Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.", "U-882 arose from her Basic-to-Pro upgrade; its prorated amount was paid by transaction P-31, which settled on March 8.", "Invoice U-889 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.", "U-889 also arose from that upgrade and labels its $36 as prorated for its recorded coverage; distinct transaction P-32 settled on March 9.", "It is a one-time plan-change charge, not a scheduled renewal.", "Mara disputes U-889’s amount, and it is her only unresolved disputed charge.", "The agent noted beneath the U-889 charge entry, “Mara wants it refunded.”", "No discount or promotional credit was advertised, promised, documented, omitted, or misapplied.", "The complete ledger’s only other settled payment is R-410, covering 2025-02-07 through 2025-03-07; no additional settled transactions exist."], "request": "Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-018", "source_is_synthetic": true, "source_sha256": "9f3816a01540b36b15d4decda36965966708aa93a8de1d0b982b71aa44023c53", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing criteria and instructions verbatim, and the request continues to target Mara’s U-889 charge referenced by “it.” The two focus spans are complete factual sentences. The counterfactual changes only U-889’s coverage period to 2025-04-07 through 2025-05-06, which is coherent with the distinct invoice and payment facts and removes the same-service-period duplication without creating contradictory measurements. Neither context states a queue selection, answer code, proposition ID, rule table, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified a13 remains atomic over an explicit transaction set. The focus, equality of the two service-coverage periods, is factual rather than policy-based. Base and counter assignments differ only on that equality and are jointly realizable: equal periods produce a duplicate, while different periods produce a single disputed prorated upgrade charge. Empty policy_evidence is correct because the governing routing criteria and priority are already preserved verbatim in original_input.questions; the original state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies U-889 as the sole referenced unresolved charge and establishes that it is a later, distinct, settled payment matching U-882 for the same subscription and service period. This is sufficient for duplicate-payment review. Renewal and promotional-credit alternatives are explicitly excluded.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an invoice-supported, settled, prorated upgrade charge whose amount is disputed. U-882 covers a different period, and a13 excludes every other settled transaction from U-889's service period, so there is no repeated settled payment for that period. Renewal and promotional-credit alternatives are also excluded, making plan-change proration review sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pronoun “it” in the billing support agent’s note refers to the charge documented by invoice U-889."}, {"id": "a2", "statement": "Invoice U-889 documents the only unresolved charge disputed by Mara."}, {"id": "a3", "statement": "Invoice U-889 is invoice evidence for its charge."}, {"id": "a4", "statement": "The payment for invoice U-889 settled."}, {"id": "a5", "statement": "The payment for invoice U-882 settled."}, {"id": "a6", "statement": "The payment for invoice U-882 occurred before the payment for invoice U-889."}, {"id": "a7", "statement": "The payments for invoices U-882 and U-889 are distinct payment transactions."}, {"id": "a8", "statement": "The charges on invoices U-882 and U-889 have equal dollar amounts."}, {"id": "a9", "statement": "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription."}, {"id": "a10", "statement": "The service-coverage period on invoice U-882 equals the service-coverage period on invoice U-889."}, {"id": "a11", "statement": "The charge on invoice U-889 arose from Mara’s upgrade from Basic to Pro."}, {"id": "a12", "statement": "Mara disputes the amount charged on invoice U-889."}, {"id": "a13", "statement": "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889."}, {"id": "a14", "statement": "The unresolved issue concerning invoice U-889 is an omitted or misapplied promotional credit."}, {"id": "a15", "statement": "The charge on invoice U-889 is a scheduled renewal charge."}, {"id": "a16", "statement": "The amount on invoice U-889 is prorated for its recorded service-coverage period."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.\",\"1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.\",\"2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.\",\"3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.\",\"4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence.\"],\"instructions\":\"Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.\",\"type\":\"score\"}},\"state\":{\"context\":\"An analyst reconciled Mara’s invoice records, payment ledger, and support notes for her StreamForge account.\",\"evidence\":[\"Account reconciliation identified two issued invoice line items for Mara’s StreamForge Pro subscription, each for $36.\",\"Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.\",\"U-882 arose from her Basic-to-Pro upgrade; its prorated amount was paid by transaction P-31, which settled on March 8.\",\"Invoice U-889 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.\",\"U-889 also arose from that upgrade and labels its $36 as prorated for its recorded coverage; distinct transaction P-32 settled on March 9.\",\"It is a one-time plan-change charge, not a scheduled renewal.\",\"Mara disputes U-889’s amount, and it is her only unresolved disputed charge.\",\"The agent noted beneath the U-889 charge entry, “Mara wants it refunded.”\",\"No discount or promotional credit was advertised, promised, documented, omitted, or misapplied.\",\"The complete ledger’s only other settled payment is R-410, covering 2025-02-07 through 2025-03-07; no additional settled transactions exist.\"],\"request\":\"Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive."}, {"path": ["state", "evidence", "3"], "text": "Invoice U-889 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive."}], "policy_evidence": [], "rules": [{"justification": "The referenced unresolved U-889 charge is a later, distinct, settled payment matching an earlier settled payment in amount, subscription, and service period. It therefore duplicates payment for the same subscription service period.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}]}, {"justification": "The sole unresolved invoice-supported dispute concerns the amount of a settled prorated upgrade charge. U-882 covers a different period, and every other settled transaction also covers a different period, so no repeated settled payment exists for U-889’s service period.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.", "negative_left": "Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.", "negative_right": "Invoice U-889 records a service-coverage period from 2025-04-07 through 2025-05-06, inclusive.", "right": "Invoice U-889 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive."}, "verifier_independent_model": false}, "family": "scale-diverse-018-002", "id": "scale-diverse-018-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"context": "An analyst reconciled Mara’s invoice records, payment ledger, and support notes for her StreamForge account.", "evidence": ["Account reconciliation identified two issued invoice line items for Mara’s StreamForge Pro subscription, each for $36.", "Invoice U-882 records a service-coverage period from 2025-03-08 through 2025-04-06, inclusive.", "U-882 arose from her Basic-to-Pro upgrade; its prorated amount was paid by transaction P-31, which settled on March 8.", "Invoice U-889 records a service-coverage period from 2025-04-07 through 2025-05-06, inclusive.", "U-889 also arose from that upgrade and labels its $36 as prorated for its recorded coverage; distinct transaction P-32 settled on March 9.", "It is a one-time plan-change charge, not a scheduled renewal.", "Mara disputes U-889’s amount, and it is her only unresolved disputed charge.", "The agent noted beneath the U-889 charge entry, “Mara wants it refunded.”", "No discount or promotional credit was advertised, promised, documented, omitted, or misapplied.", "The complete ledger’s only other settled payment is R-410, covering 2025-02-07 through 2025-03-07; no additional settled transactions exist."], "request": "Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-018", "source_is_synthetic": true, "source_sha256": "9f3816a01540b36b15d4decda36965966708aa93a8de1d0b982b71aa44023c53", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both inputs preserve the questions object verbatim, including all routing criteria and instructions. The request remains bound to Mara’s account, the charge referenced by “it,” and invoice U-889; changing U-889’s service period is an allowed factual change. The two focus spans are complete factual sentences. In the counterfactual, U-882 and U-889 cover different periods, which is coherent with the remaining transaction, invoice, and dispute facts and creates no duplicate measurement or assertion. Neither context states a routing score, queue answer, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified a13 remains atomic over an explicit transaction set. The focus, equality of the two service-coverage periods, is factual rather than policy-based. Base and counter assignments differ only on that equality and are jointly realizable: equal periods produce a duplicate, while different periods produce a single disputed prorated upgrade charge. Empty policy_evidence is correct because the governing routing criteria and priority are already preserved verbatim in original_input.questions; the original state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies U-889 as the sole referenced unresolved charge and establishes that it is a later, distinct, settled payment matching U-882 for the same subscription and service period. This is sufficient for duplicate-payment review. Renewal and promotional-credit alternatives are explicitly excluded.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an invoice-supported, settled, prorated upgrade charge whose amount is disputed. U-882 covers a different period, and a13 excludes every other settled transaction from U-889's service period, so there is no repeated settled payment for that period. Renewal and promotional-credit alternatives are also excluded, making plan-change proration review sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pronoun “it” in the billing support agent’s note refers to the charge documented by invoice U-889."}, {"id": "a2", "statement": "Invoice U-889 documents the only unresolved charge disputed by Mara."}, {"id": "a3", "statement": "Invoice U-889 is invoice evidence for its charge."}, {"id": "a4", "statement": "The payment for invoice U-889 settled."}, {"id": "a5", "statement": "The payment for invoice U-882 settled."}, {"id": "a6", "statement": "The payment for invoice U-882 occurred before the payment for invoice U-889."}, {"id": "a7", "statement": "The payments for invoices U-882 and U-889 are distinct payment transactions."}, {"id": "a8", "statement": "The charges on invoices U-882 and U-889 have equal dollar amounts."}, {"id": "a9", "statement": "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription."}, {"id": "a10", "statement": "The service-coverage period on invoice U-882 equals the service-coverage period on invoice U-889."}, {"id": "a11", "statement": "The charge on invoice U-889 arose from Mara’s upgrade from Basic to Pro."}, {"id": "a12", "statement": "Mara disputes the amount charged on invoice U-889."}, {"id": "a13", "statement": "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889."}, {"id": "a14", "statement": "The unresolved issue concerning invoice U-889 is an omitted or misapplied promotional credit."}, {"id": "a15", "statement": "The charge on invoice U-889 is a scheduled renewal charge."}, {"id": "a16", "statement": "The amount on invoice U-889 is prorated for its recorded service-coverage period."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.\",\"1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.\",\"2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.\",\"3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.\",\"4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence.\"],\"instructions\":\"Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.\",\"type\":\"score\"}},\"state\":{\"context\":\"Concise billing field note for Mara’s StreamForge account.\",\"evidence\":[\"Each invoice charges $36, paid through separate card transactions that settled on March 2 and March 4, 2026, for U-882 and U-889 respectively.\",\"Both invoices charge for Mara’s same StreamForge Pro subscription.\",\"U-889 arose from Mara’s Basic-to-Pro upgrade, is not a scheduled renewal, and records a prorated amount for its service-coverage period.\",\"Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.\",\"Invoice U-889 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.\",\"The only settled transaction besides U-882 and U-889 is R-410, whose service period differs from U-889’s recorded period.\",\"Mara accepted R-410 and U-882 but disputes U-889’s amount; U-889 is her only unresolved disputed charge.\",\"No advertised, promised, or documented promotional credit is involved, and no credit was omitted or misapplied.\",\"The agent clarified that ‘it’ means the charge documented by U-889, which is retained as invoice evidence for that charge.\"]},\"request\":\"Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["state", "evidence", "3"], "text": "Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive."}, {"path": ["state", "evidence", "4"], "text": "Invoice U-889 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive."}], "policy_evidence": [], "rules": [{"justification": "The referenced unresolved U-889 charge is a later, distinct, settled payment matching an earlier settled payment in amount, subscription, and service period. It therefore duplicates payment for the same subscription service period.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}]}, {"justification": "The sole unresolved invoice-supported dispute concerns the amount of a settled prorated upgrade charge. U-882 covers a different period, and every other settled transaction also covers a different period, so no repeated settled payment exists for U-889’s service period.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.", "negative_left": "Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.", "negative_right": "Invoice U-889 records a service-coverage period from April 1, 2026, through April 30, 2026, inclusive.", "right": "Invoice U-889 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive."}, "verifier_independent_model": false}, "family": "scale-diverse-018-004", "id": "scale-diverse-018-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "request": "Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?", "state": {"context": "Concise billing field note for Mara’s StreamForge account.", "evidence": ["Each invoice charges $36, paid through separate card transactions that settled on March 2 and March 4, 2026, for U-882 and U-889 respectively.", "Both invoices charge for Mara’s same StreamForge Pro subscription.", "U-889 arose from Mara’s Basic-to-Pro upgrade, is not a scheduled renewal, and records a prorated amount for its service-coverage period.", "Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.", "Invoice U-889 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.", "The only settled transaction besides U-882 and U-889 is R-410, whose service period differs from U-889’s recorded period.", "Mara accepted R-410 and U-882 but disputes U-889’s amount; U-889 is her only unresolved disputed charge.", "No advertised, promised, or documented promotional credit is involved, and no credit was omitted or misapplied.", "The agent clarified that ‘it’ means the charge documented by U-889, which is retained as invoice evidence for that charge."]}}}, "method": "c2d", "provenance": {"source_id": "diverse-018", "source_is_synthetic": true, "source_sha256": "9f3816a01540b36b15d4decda36965966708aa93a8de1d0b982b71aa44023c53", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both inputs preserve the questions object verbatim, including all routing criteria and instructions. The request remains bound to Mara’s account, the charge referenced by “it,” and invoice U-889; changing U-889’s service period is an allowed factual change. The two focus spans are complete factual sentences. In the counterfactual, U-882 and U-889 cover different periods, which is coherent with the remaining transaction, invoice, and dispute facts and creates no duplicate measurement or assertion. Neither context states a routing score, queue answer, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified a13 remains atomic over an explicit transaction set. The focus, equality of the two service-coverage periods, is factual rather than policy-based. Base and counter assignments differ only on that equality and are jointly realizable: equal periods produce a duplicate, while different periods produce a single disputed prorated upgrade charge. Empty policy_evidence is correct because the governing routing criteria and priority are already preserved verbatim in original_input.questions; the original state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies U-889 as the sole referenced unresolved charge and establishes that it is a later, distinct, settled payment matching U-882 for the same subscription and service period. This is sufficient for duplicate-payment review. Renewal and promotional-credit alternatives are explicitly excluded.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an invoice-supported, settled, prorated upgrade charge whose amount is disputed. U-882 covers a different period, and a13 excludes every other settled transaction from U-889's service period, so there is no repeated settled payment for that period. Renewal and promotional-credit alternatives are also excluded, making plan-change proration review sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pronoun “it” in the billing support agent’s note refers to the charge documented by invoice U-889."}, {"id": "a2", "statement": "Invoice U-889 documents the only unresolved charge disputed by Mara."}, {"id": "a3", "statement": "Invoice U-889 is invoice evidence for its charge."}, {"id": "a4", "statement": "The payment for invoice U-889 settled."}, {"id": "a5", "statement": "The payment for invoice U-882 settled."}, {"id": "a6", "statement": "The payment for invoice U-882 occurred before the payment for invoice U-889."}, {"id": "a7", "statement": "The payments for invoices U-882 and U-889 are distinct payment transactions."}, {"id": "a8", "statement": "The charges on invoices U-882 and U-889 have equal dollar amounts."}, {"id": "a9", "statement": "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription."}, {"id": "a10", "statement": "The service-coverage period on invoice U-882 equals the service-coverage period on invoice U-889."}, {"id": "a11", "statement": "The charge on invoice U-889 arose from Mara’s upgrade from Basic to Pro."}, {"id": "a12", "statement": "Mara disputes the amount charged on invoice U-889."}, {"id": "a13", "statement": "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889."}, {"id": "a14", "statement": "The unresolved issue concerning invoice U-889 is an omitted or misapplied promotional credit."}, {"id": "a15", "statement": "The charge on invoice U-889 is a scheduled renewal charge."}, {"id": "a16", "statement": "The amount on invoice U-889 is prorated for its recorded service-coverage period."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.\",\"1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.\",\"2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.\",\"3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.\",\"4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence.\"],\"instructions\":\"Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.\",\"type\":\"score\"}},\"state\":{\"context\":\"Concise billing field note for Mara’s StreamForge account.\",\"evidence\":[\"Each invoice charges $36, paid through separate card transactions that settled on March 2 and March 4, 2026, for U-882 and U-889 respectively.\",\"Both invoices charge for Mara’s same StreamForge Pro subscription.\",\"U-889 arose from Mara’s Basic-to-Pro upgrade, is not a scheduled renewal, and records a prorated amount for its service-coverage period.\",\"Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.\",\"Invoice U-889 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.\",\"The only settled transaction besides U-882 and U-889 is R-410, whose service period differs from U-889’s recorded period.\",\"Mara accepted R-410 and U-882 but disputes U-889’s amount; U-889 is her only unresolved disputed charge.\",\"No advertised, promised, or documented promotional credit is involved, and no credit was omitted or misapplied.\",\"The agent clarified that ‘it’ means the charge documented by U-889, which is retained as invoice evidence for that charge.\"]},\"request\":\"Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["state", "evidence", "3"], "text": "Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive."}, {"path": ["state", "evidence", "4"], "text": "Invoice U-889 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive."}], "policy_evidence": [], "rules": [{"justification": "The referenced unresolved U-889 charge is a later, distinct, settled payment matching an earlier settled payment in amount, subscription, and service period. It therefore duplicates payment for the same subscription service period.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}]}, {"justification": "The sole unresolved invoice-supported dispute concerns the amount of a settled prorated upgrade charge. U-882 covers a different period, and every other settled transaction also covers a different period, so no repeated settled payment exists for U-889’s service period.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.", "negative_left": "Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.", "negative_right": "Invoice U-889 records a service-coverage period from April 1, 2026, through April 30, 2026, inclusive.", "right": "Invoice U-889 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive."}, "verifier_independent_model": false}, "family": "scale-diverse-018-004", "id": "scale-diverse-018-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "request": "Which billing queue should the subscription operations analyst select for the charge referenced by “it” in the agent’s note?", "state": {"context": "Concise billing field note for Mara’s StreamForge account.", "evidence": ["Each invoice charges $36, paid through separate card transactions that settled on March 2 and March 4, 2026, for U-882 and U-889 respectively.", "Both invoices charge for Mara’s same StreamForge Pro subscription.", "U-889 arose from Mara’s Basic-to-Pro upgrade, is not a scheduled renewal, and records a prorated amount for its service-coverage period.", "Invoice U-882 records a service-coverage period from March 1, 2026, through March 31, 2026, inclusive.", "Invoice U-889 records a service-coverage period from April 1, 2026, through April 30, 2026, inclusive.", "The only settled transaction besides U-882 and U-889 is R-410, whose service period differs from U-889’s recorded period.", "Mara accepted R-410 and U-882 but disputes U-889’s amount; U-889 is her only unresolved disputed charge.", "No advertised, promised, or documented promotional credit is involved, and no credit was omitted or misapplied.", "The agent clarified that ‘it’ means the charge documented by U-889, which is retained as invoice evidence for that charge."]}}}, "method": "c2d", "provenance": {"source_id": "diverse-018", "source_is_synthetic": true, "source_sha256": "9f3816a01540b36b15d4decda36965966708aa93a8de1d0b982b71aa44023c53", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policies from the original state, while the unchanged questions preserve all decision criteria and instructions. The pickup-order decision scope and 9:15 time binding remain fixed. The two evidence spans are complete factual sentences rather than policy text. Changing the scan result from SKU ZX-3148 to ZX-3196 coherently changes whether the located item is the requested SKU without contradicting the shared facts. Neither context contains an answer code, proposition ID, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization. At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3148. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization."}, {"path": [], "text": "At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3148."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization.", "negative_left": "For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization.", "negative_right": "At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3196.", "right": "At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3148."}, "verifier_independent_model": false}, "family": "scale-diverse-019-004", "id": "scale-diverse-019-004-base", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization. At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3148. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_exact_item"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policies from the original state, while the unchanged questions preserve all decision criteria and instructions. The pickup-order decision scope and 9:15 time binding remain fixed. The two evidence spans are complete factual sentences rather than policy text. Changing the scan result from SKU ZX-3148 to ZX-3196 coherently changes whether the located item is the requested SKU without contradicting the shared facts. Neither context contains an answer code, proposition ID, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization. At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3148. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization."}, {"path": [], "text": "At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3148."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization.", "negative_left": "For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization.", "negative_right": "At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3196.", "right": "At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3148."}, "verifier_independent_model": false}, "family": "scale-diverse-019-004", "id": "scale-diverse-019-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "For the 9:15 fulfillment decision on pickup order PO-6814, the order line records requested SKU ZX-3148, the picker proposed the located item, the latest customer note effective before 9:15 permits that item and states only that it must be navy and two liters, both of which match the item, and the decision record documents no safety, regulatory, or price-limit exception requiring supervisor authorization. At 9:12, the picker scan log for pickup order PO-6814 records the located item's SKU as ZX-3196. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_proposed_substitute"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing latest-note and return-to-picking policies from the original state, while the unchanged questions preserve all decision criteria. The pickup order, 9:15 decision, and item-identification path remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the scanner identification from SKU PX-4827 to SKU PX-6194, without conflicting with the unchanged attribute inspection or other assertions. Neither context includes an answer code, gold label, rule table, proposition identifier, label rationale, or classifier-output instruction; the action language is preserved governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827. At 9:10, the store picker proposed the item then located for fulfillment of that order at the 9:15 decision. At 9:12, the customer’s final timestamped note effective before the decision said, “The picker-located item is permitted if it is Green Vale, 32-ounce, and unsweetened,” and stated no other item attributes. At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-4827. At 9:14, inspection confirmed that the picker-located item was Green Vale, 32-ounce, and unsweetened. The order record documented that no safety exception concerning the item required supervisor authorization, no regulatory exception concerning it required such authorization, and no price-limit exception concerning it required such authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827."}, {"path": [], "text": "At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-4827."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827.", "negative_left": "At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827.", "negative_right": "At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-6194.", "right": "At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-4827."}, "verifier_independent_model": false}, "family": "scale-diverse-019-005", "id": "scale-diverse-019-005-base", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827. At 9:10, the store picker proposed the item then located for fulfillment of that order at the 9:15 decision. At 9:12, the customer’s final timestamped note effective before the decision said, “The picker-located item is permitted if it is Green Vale, 32-ounce, and unsweetened,” and stated no other item attributes. At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-4827. At 9:14, inspection confirmed that the picker-located item was Green Vale, 32-ounce, and unsweetened. The order record documented that no safety exception concerning the item required supervisor authorization, no regulatory exception concerning it required such authorization, and no price-limit exception concerning it required such authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_exact_item"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing latest-note and return-to-picking policies from the original state, while the unchanged questions preserve all decision criteria. The pickup order, 9:15 decision, and item-identification path remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the scanner identification from SKU PX-4827 to SKU PX-6194, without conflicting with the unchanged attribute inspection or other assertions. Neither context includes an answer code, gold label, rule table, proposition identifier, label rationale, or classifier-output instruction; the action language is preserved governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827. At 9:10, the store picker proposed the item then located for fulfillment of that order at the 9:15 decision. At 9:12, the customer’s final timestamped note effective before the decision said, “The picker-located item is permitted if it is Green Vale, 32-ounce, and unsweetened,” and stated no other item attributes. At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-4827. At 9:14, inspection confirmed that the picker-located item was Green Vale, 32-ounce, and unsweetened. The order record documented that no safety exception concerning the item required supervisor authorization, no regulatory exception concerning it required such authorization, and no price-limit exception concerning it required such authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827."}, {"path": [], "text": "At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-4827."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827.", "negative_left": "At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827.", "negative_right": "At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-6194.", "right": "At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-4827."}, "verifier_independent_model": false}, "family": "scale-diverse-019-005", "id": "scale-diverse-019-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At 9:03, the pickup order considered in the 9:15 fulfillment decision recorded the originally requested item under SKU PX-4827. At 9:10, the store picker proposed the item then located for fulfillment of that order at the 9:15 decision. At 9:12, the customer’s final timestamped note effective before the decision said, “The picker-located item is permitted if it is Green Vale, 32-ounce, and unsweetened,” and stated no other item attributes. At 9:13, the store picker's scanner identified the item located for the 9:15 fulfillment decision as SKU PX-6194. At 9:14, inspection confirmed that the picker-located item was Green Vale, 32-ounce, and unsweetened. The order record documented that no safety exception concerning the item required supervisor authorization, no regulatory exception concerning it required such authorization, and no price-limit exception concerning it required such authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_proposed_substitute"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same customer constraints and do not introduce waivers, exceptions, priorities, or altered missing-evidence rules. They concern the same pickup order, proposed sealed package, approval decision, and relevant handoff/inspection timeframe. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the inspection result from finding the required label claim to finding no such claim, without conflicting with the unchanged brand, size, assignment, or waiver facts. Neither context contains an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”\"},{\"speaker\":\"Handoff record\",\"text\":\"At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter.\"},{\"speaker\":\"Inspection record\",\"text\":\"At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found the exact claim “no added sugar” on its front label.\"},{\"speaker\":\"Contact log\",\"text\":\"The requested item remains unavailable. The customer has not withdrawn or waived any condition in the order note, and the inspected package remained assigned to this pickup order while the supervisor awaited an approval decision.\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["1", "text"], "text": "At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter."}, {"path": ["2", "text"], "text": "At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found the exact claim “no added sugar” on its front label."}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter.", "negative_left": "At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter.", "negative_right": "At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found no claim stating “no added sugar” or otherwise stating that the product has no added sugar.", "right": "At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found the exact claim “no added sugar” on its front label."}, "verifier_independent_model": false}, "family": "scale-diverse-021-003", "id": "scale-diverse-021-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Handoff record", "text": "At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter."}, {"speaker": "Inspection record", "text": "At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found the exact claim “no added sugar” on its front label."}, {"speaker": "Contact log", "text": "The requested item remains unavailable. The customer has not withdrawn or waived any condition in the order note, and the inspected package remained assigned to this pickup order while the supervisor awaited an approval decision."}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same customer constraints and do not introduce waivers, exceptions, priorities, or altered missing-evidence rules. They concern the same pickup order, proposed sealed package, approval decision, and relevant handoff/inspection timeframe. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the inspection result from finding the required label claim to finding no such claim, without conflicting with the unchanged brand, size, assignment, or waiver facts. Neither context contains an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”\"},{\"speaker\":\"Handoff record\",\"text\":\"At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter.\"},{\"speaker\":\"Inspection record\",\"text\":\"At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found the exact claim “no added sugar” on its front label.\"},{\"speaker\":\"Contact log\",\"text\":\"The requested item remains unavailable. The customer has not withdrawn or waived any condition in the order note, and the inspected package remained assigned to this pickup order while the supervisor awaited an approval decision.\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["1", "text"], "text": "At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter."}, {"path": ["2", "text"], "text": "At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found the exact claim “no added sugar” on its front label."}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter.", "negative_left": "At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter.", "negative_right": "At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found no claim stating “no added sugar” or otherwise stating that the product has no added sugar.", "right": "At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found the exact claim “no added sugar” on its front label."}, "verifier_independent_model": false}, "family": "scale-diverse-021-003", "id": "scale-diverse-021-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Handoff record", "text": "At the 4:18 p.m. fulfillment handoff for this pickup order, the clerk identified the package bearing seal H-731 as the proposed North Mill smooth peanut butter substitution, recorded its net quantity as exactly 16 ounces, and confirmed that its brand is the same North Mill brand as the requested North Mill creamy peanut butter."}, {"speaker": "Inspection record", "text": "At 4:21 p.m., a complete inspection of every printed surface of the package bearing seal H-731 found no claim stating “no added sugar” or otherwise stating that the product has no added sugar."}, {"speaker": "Contact log", "text": "The requested item remains unavailable. The customer has not withdrawn or waived any condition in the order note, and the inspected package remained assigned to this pickup order while the supervisor awaited an approval decision."}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve Lena, the pickup substitution decision, the infant-formula brand/stage facts, and the governing price, exception, and routing policies. The two evidence spans are complete factual sentences. The counterfactual changes only the proposed replacement price from $35.10 to $35.30; with an original-item price of $32.00, this is coherent and creates no duplicate contradictory measurement. Neither context supplies an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00. At 10:04, the picker’s product scan identified the unavailable item as BrightStart Stage 2 infant formula. The proposed replacement was identified as BrightStart Stage 2 infant formula in a different package size. At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.10.\\n\\nLena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”\\n\\nFulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00."}, {"path": [], "text": "At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.10."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00.", "negative_left": "At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00.", "negative_right": "At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.30.", "right": "At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.10."}, "verifier_independent_model": false}, "family": "scale-diverse-022-001", "id": "scale-diverse-022-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00. At 10:04, the picker’s product scan identified the unavailable item as BrightStart Stage 2 infant formula. The proposed replacement was identified as BrightStart Stage 2 infant formula in a different package size. At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.10.\n\nLena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”\n\nFulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve Lena, the pickup substitution decision, the infant-formula brand/stage facts, and the governing price, exception, and routing policies. The two evidence spans are complete factual sentences. The counterfactual changes only the proposed replacement price from $35.10 to $35.30; with an original-item price of $32.00, this is coherent and creates no duplicate contradictory measurement. Neither context supplies an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00. At 10:04, the picker’s product scan identified the unavailable item as BrightStart Stage 2 infant formula. The proposed replacement was identified as BrightStart Stage 2 infant formula in a different package size. At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.10.\\n\\nLena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”\\n\\nFulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00."}, {"path": [], "text": "At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.10."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00.", "negative_left": "At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00.", "negative_right": "At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.30.", "right": "At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.10."}, "verifier_independent_model": false}, "family": "scale-diverse-022-001", "id": "scale-diverse-022-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "At 10:02 on 17 September 2026, Lena’s pickup-order record listed the final pre-tax price of the single unavailable ordered item as $32.00. At 10:04, the picker’s product scan identified the unavailable item as BrightStart Stage 2 infant formula. The proposed replacement was identified as BrightStart Stage 2 infant formula in a different package size. At 10:06 on 17 September 2026, the same pickup-order record listed the final pre-tax price of the single proposed replacement as $35.30.\n\nLena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”\n\nFulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same substitution policy, request scope, customer, order, product, and routing decision. The two evidence spans are complete factual ledger sentences; the counterfactual coherently changes only the proposed replacement price from $27.10 to $27.40 without creating duplicate conflicting measurements, and neither context states or encodes the required yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"During evidence reconciliation for Lena’s pickup order, the inventory and packaging records were compared. The unavailable line is infant formula, and both the ordered item and the proposed replacement are identified as BrightStart Stage 2. The different package sizes do not alter those recorded brand and stage designations. At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80. At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.10. Lena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80."}, {"path": [], "text": "At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.10."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80.", "negative_left": "At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80.", "negative_right": "At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.40.", "right": "At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.10."}, "verifier_independent_model": false}, "family": "scale-diverse-022-002", "id": "scale-diverse-022-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "During evidence reconciliation for Lena’s pickup order, the inventory and packaging records were compared. The unavailable line is infant formula, and both the ordered item and the proposed replacement are identified as BrightStart Stage 2. The different package sizes do not alter those recorded brand and stage designations. At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80. At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.10. Lena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same substitution policy, request scope, customer, order, product, and routing decision. The two evidence spans are complete factual ledger sentences; the counterfactual coherently changes only the proposed replacement price from $27.10 to $27.40 without creating duplicate conflicting measurements, and neither context states or encodes the required yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"During evidence reconciliation for Lena’s pickup order, the inventory and packaging records were compared. The unavailable line is infant formula, and both the ordered item and the proposed replacement are identified as BrightStart Stage 2. The different package sizes do not alter those recorded brand and stage designations. At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80. At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.10. Lena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80."}, {"path": [], "text": "At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.10."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80.", "negative_left": "At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80.", "negative_right": "At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.40.", "right": "At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.10."}, "verifier_independent_model": false}, "family": "scale-diverse-022-002", "id": "scale-diverse-022-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "During evidence reconciliation for Lena’s pickup order, the inventory and packaging records were compared. The unavailable line is infant formula, and both the ordered item and the proposed replacement are identified as BrightStart Stage 2. The different package sizes do not alter those recorded brand and stage designations. At 14:05 UTC on 17 September 2026, Lena’s pickup-order ledger recorded the unavailable ordered infant formula at $24.80. At 14:05 UTC on 17 September 2026, the same pickup-order ledger recorded the proposed replacement infant formula at $27.40. Lena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same substitution policy, request scope, Lena pickup-order entity, and decision pathway. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes only the proposed replacement price from $50.50 to $51.00 without creating an internal duplicate or contradiction. Neither context states the case-specific answer, an answer code, proposition ID, rule table, label rationale, or an instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Operational handoff for Lena’s pickup order: The unavailable ordered item is infant formula. Package-label verification confirms that the unavailable can and proposed replacement both carry the BrightStart brand and Stage 2 designation. At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax. At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $50.50 per can before tax. The replacement remains pending the substitution decision before release to picking. Lena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax."}, {"path": [], "text": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $50.50 per can before tax."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax.", "negative_left": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax.", "negative_right": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $51.00 per can before tax.", "right": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $50.50 per can before tax."}, "verifier_independent_model": false}, "family": "scale-diverse-022-003", "id": "scale-diverse-022-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Operational handoff for Lena’s pickup order: The unavailable ordered item is infant formula. Package-label verification confirms that the unavailable can and proposed replacement both carry the BrightStart brand and Stage 2 designation. At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax. At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $50.50 per can before tax. The replacement remains pending the substitution decision before release to picking. Lena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same substitution policy, request scope, Lena pickup-order entity, and decision pathway. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes only the proposed replacement price from $50.50 to $51.00 without creating an internal duplicate or contradiction. Neither context states the case-specific answer, an answer code, proposition ID, rule table, label rationale, or an instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Operational handoff for Lena’s pickup order: The unavailable ordered item is infant formula. Package-label verification confirms that the unavailable can and proposed replacement both carry the BrightStart brand and Stage 2 designation. At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax. At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $50.50 per can before tax. The replacement remains pending the substitution decision before release to picking. Lena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax."}, {"path": [], "text": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $50.50 per can before tax."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax.", "negative_left": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax.", "negative_right": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $51.00 per can before tax.", "right": "At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $50.50 per can before tax."}, "verifier_independent_model": false}, "family": "scale-diverse-022-003", "id": "scale-diverse-022-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Operational handoff for Lena’s pickup order: The unavailable ordered item is infant formula. Package-label verification confirms that the unavailable can and proposed replacement both carry the BrightStart brand and Stage 2 designation. At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the unavailable ordered item at $46.00 per can before tax. At 14:20 UTC on 17 September 2026, Lena’s pickup-order handoff record listed the proposed replacement for the unavailable ordered item at $51.00 per can before tax. The replacement remains pending the substitution decision before release to picking. Lena’s order note states: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the current-order requirements, governing suitability scope, proposed 28 oz Mill Lane item, current-order binding, and relevant date/path details. The two evidence spans are complete factual sentences linking specimen Q7-41 to the proposed carton and reporting its assay result. The counterfactual changes only the assay from 0 to 3 grams of added sugar; this does not contradict the unchanged front-panel description that the product was labeled unsweetened. Neither context embeds a decision level, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"At 08:50 on 17 September 2026, the clerk logged the proposed replacement for the current pickup order. Its front panel identified it as Mill Lane oat milk, labeled unsweetened, with a net quantity of 28 oz.\"},{\"speaker\":\"Laboratory technician\",\"text\":\"At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order.\"},{\"speaker\":\"Laboratory record\",\"text\":\"At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 0 grams of added sugar for specimen Q7-41.\"},{\"speaker\":\"Order review\",\"text\":\"The requested Harvest Oat carton has a net quantity of 32 oz, whereas the proposed Mill Lane carton has a net quantity of 28 oz.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order."}, {"path": ["3", "text"], "text": "At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 0 grams of added sugar for specimen Q7-41."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order.", "negative_left": "At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order.", "negative_right": "At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 3 grams of added sugar for specimen Q7-41.", "right": "At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 0 grams of added sugar for specimen Q7-41."}, "verifier_independent_model": false}, "family": "scale-diverse-024-001", "id": "scale-diverse-024-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Receiving clerk", "text": "At 08:50 on 17 September 2026, the clerk logged the proposed replacement for the current pickup order. Its front panel identified it as Mill Lane oat milk, labeled unsweetened, with a net quantity of 28 oz."}, {"speaker": "Laboratory technician", "text": "At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order."}, {"speaker": "Laboratory record", "text": "At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 0 grams of added sugar for specimen Q7-41."}, {"speaker": "Order review", "text": "The requested Harvest Oat carton has a net quantity of 32 oz, whereas the proposed Mill Lane carton has a net quantity of 28 oz."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the current-order requirements, governing suitability scope, proposed 28 oz Mill Lane item, current-order binding, and relevant date/path details. The two evidence spans are complete factual sentences linking specimen Q7-41 to the proposed carton and reporting its assay result. The counterfactual changes only the assay from 0 to 3 grams of added sugar; this does not contradict the unchanged front-panel description that the product was labeled unsweetened. Neither context embeds a decision level, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"At 08:50 on 17 September 2026, the clerk logged the proposed replacement for the current pickup order. Its front panel identified it as Mill Lane oat milk, labeled unsweetened, with a net quantity of 28 oz.\"},{\"speaker\":\"Laboratory technician\",\"text\":\"At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order.\"},{\"speaker\":\"Laboratory record\",\"text\":\"At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 0 grams of added sugar for specimen Q7-41.\"},{\"speaker\":\"Order review\",\"text\":\"The requested Harvest Oat carton has a net quantity of 32 oz, whereas the proposed Mill Lane carton has a net quantity of 28 oz.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order."}, {"path": ["3", "text"], "text": "At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 0 grams of added sugar for specimen Q7-41."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order.", "negative_left": "At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order.", "negative_right": "At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 3 grams of added sugar for specimen Q7-41.", "right": "At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 0 grams of added sugar for specimen Q7-41."}, "verifier_independent_model": false}, "family": "scale-diverse-024-001", "id": "scale-diverse-024-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Receiving clerk", "text": "At 08:50 on 17 September 2026, the clerk logged the proposed replacement for the current pickup order. Its front panel identified it as Mill Lane oat milk, labeled unsweetened, with a net quantity of 28 oz."}, {"speaker": "Laboratory technician", "text": "At 09:14 on 17 September 2026, technician Mara Ives assigned specimen code Q7-41 exclusively to a sealed sample taken from the proposed 28 oz Mill Lane oat milk for the current pickup order."}, {"speaker": "Laboratory record", "text": "At 10:02 on 17 September 2026, the completed added-sugar assay recorded exactly 3 grams of added sugar for specimen Q7-41."}, {"speaker": "Order review", "text": "The requested Harvest Oat carton has a net quantity of 32 oz, whereas the proposed Mill Lane carton has a net quantity of 28 oz."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the current-order requirements, relevance policy, proposed 28 oz Mill Lane item, current-order scope, and timestamps. The two focus spans are complete factual sentences. The counterfactual changes only the laboratory measurement from zero to four grams of added sugar; this does not duplicate or directly contradict another measurement, because the separate statement that the carton is marked “unsweetened” describes its labeling. Neither context contains an answer level, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Inventory examiner\",\"text\":\"The sealed proposed carton for the current pickup order is a Mill Lane product identified as oat milk, marked unsweetened, and printed with a net quantity of 28 oz. The unavailable requested Harvest Oat carton is printed with a net quantity of 32 oz.\"},{\"speaker\":\"Evidence custodian\",\"text\":\"At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731.\"},{\"speaker\":\"Reporting laboratory\",\"text\":\"The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly zero grams of added sugar in the complete source contents represented by sample MX-731.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731."}, {"path": ["3", "text"], "text": "The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly zero grams of added sugar in the complete source contents represented by sample MX-731."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731.", "negative_left": "At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731.", "negative_right": "The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly four grams of added sugar in the complete source contents represented by sample MX-731.", "right": "The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly zero grams of added sugar in the complete source contents represented by sample MX-731."}, "verifier_independent_model": false}, "family": "scale-diverse-024-002", "id": "scale-diverse-024-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Inventory examiner", "text": "The sealed proposed carton for the current pickup order is a Mill Lane product identified as oat milk, marked unsweetened, and printed with a net quantity of 28 oz. The unavailable requested Harvest Oat carton is printed with a net quantity of 32 oz."}, {"speaker": "Evidence custodian", "text": "At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731."}, {"speaker": "Reporting laboratory", "text": "The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly zero grams of added sugar in the complete source contents represented by sample MX-731."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the current-order requirements, relevance policy, proposed 28 oz Mill Lane item, current-order scope, and timestamps. The two focus spans are complete factual sentences. The counterfactual changes only the laboratory measurement from zero to four grams of added sugar; this does not duplicate or directly contradict another measurement, because the separate statement that the carton is marked “unsweetened” describes its labeling. Neither context contains an answer level, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Inventory examiner\",\"text\":\"The sealed proposed carton for the current pickup order is a Mill Lane product identified as oat milk, marked unsweetened, and printed with a net quantity of 28 oz. The unavailable requested Harvest Oat carton is printed with a net quantity of 32 oz.\"},{\"speaker\":\"Evidence custodian\",\"text\":\"At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731.\"},{\"speaker\":\"Reporting laboratory\",\"text\":\"The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly zero grams of added sugar in the complete source contents represented by sample MX-731.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731."}, {"path": ["3", "text"], "text": "The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly zero grams of added sugar in the complete source contents represented by sample MX-731."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731.", "negative_left": "At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731.", "negative_right": "The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly four grams of added sugar in the complete source contents represented by sample MX-731.", "right": "The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly zero grams of added sugar in the complete source contents represented by sample MX-731."}, "verifier_independent_model": false}, "family": "scale-diverse-024-002", "id": "scale-diverse-024-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Inventory examiner", "text": "The sealed proposed carton for the current pickup order is a Mill Lane product identified as oat milk, marked unsweetened, and printed with a net quantity of 28 oz. The unavailable requested Harvest Oat carton is printed with a net quantity of 32 oz."}, {"speaker": "Evidence custodian", "text": "At 10:42 UTC on 17 September 2026, chain-of-custody record MX-731 identified the proposed 28 oz Mill Lane oat milk for the current pickup order as the complete source contents represented by laboratory sample MX-731."}, {"speaker": "Reporting laboratory", "text": "The final laboratory report issued at 14:18 UTC on 17 September 2026 recorded exactly four grams of added sugar in the complete source contents represented by sample MX-731."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the customer’s current-order constraints and the supervisor’s governing relevance policy, while the unchanged questions preserve all scoring criteria and instructions. The proposed item, order, quantity, and relevant chronology remain bound to the same entities and path. The two focus-evidence spans are complete factual sentences. Changing the sealed-vessel assay from exactly 0 grams to exactly 4 grams of added sugar is coherent with the unchanged transfer and chain-of-custody facts; the carton’s unsweetened label is merely a package claim and does not create a duplicate contradictory measurement. Neither context embeds a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Transfer log\",\"text\":\"At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material.\"},{\"speaker\":\"Laboratory log\",\"text\":\"At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 0 grams of added sugar in all material inside sealed vessel Q17.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"},{\"speaker\":\"Field note\",\"text\":\"Package inspection confirmed that the proposed Mill Lane carton is oat milk, bears an unsweetened label, and has a manufacturer-declared net quantity of 28 oz. The order record lists the requested Harvest Oat carton at 32 oz, so the proposed carton is smaller rather than equal in net quantity. Chain-of-custody records show that Q17 remained sealed and intact between transfer and assay.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material."}, {"path": ["2", "text"], "text": "At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 0 grams of added sugar in all material inside sealed vessel Q17."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material.", "negative_left": "At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material.", "negative_right": "At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 4 grams of added sugar in all material inside sealed vessel Q17.", "right": "At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 0 grams of added sugar in all material inside sealed vessel Q17."}, "verifier_independent_model": false}, "family": "scale-diverse-024-004", "id": "scale-diverse-024-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Transfer log", "text": "At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material."}, {"speaker": "Laboratory log", "text": "At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 0 grams of added sugar in all material inside sealed vessel Q17."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}, {"speaker": "Field note", "text": "Package inspection confirmed that the proposed Mill Lane carton is oat milk, bears an unsweetened label, and has a manufacturer-declared net quantity of 28 oz. The order record lists the requested Harvest Oat carton at 32 oz, so the proposed carton is smaller rather than equal in net quantity. Chain-of-custody records show that Q17 remained sealed and intact between transfer and assay."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the customer’s current-order constraints and the supervisor’s governing relevance policy, while the unchanged questions preserve all scoring criteria and instructions. The proposed item, order, quantity, and relevant chronology remain bound to the same entities and path. The two focus-evidence spans are complete factual sentences. Changing the sealed-vessel assay from exactly 0 grams to exactly 4 grams of added sugar is coherent with the unchanged transfer and chain-of-custody facts; the carton’s unsweetened label is merely a package claim and does not create a duplicate contradictory measurement. Neither context embeds a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Transfer log\",\"text\":\"At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material.\"},{\"speaker\":\"Laboratory log\",\"text\":\"At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 0 grams of added sugar in all material inside sealed vessel Q17.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"},{\"speaker\":\"Field note\",\"text\":\"Package inspection confirmed that the proposed Mill Lane carton is oat milk, bears an unsweetened label, and has a manufacturer-declared net quantity of 28 oz. The order record lists the requested Harvest Oat carton at 32 oz, so the proposed carton is smaller rather than equal in net quantity. Chain-of-custody records show that Q17 remained sealed and intact between transfer and assay.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material."}, {"path": ["2", "text"], "text": "At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 0 grams of added sugar in all material inside sealed vessel Q17."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material.", "negative_left": "At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material.", "negative_right": "At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 4 grams of added sugar in all material inside sealed vessel Q17.", "right": "At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 0 grams of added sugar in all material inside sealed vessel Q17."}, "verifier_independent_model": false}, "family": "scale-diverse-024-004", "id": "scale-diverse-024-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Transfer log", "text": "At 09:14 UTC on 17 September 2026, technician Mara Venn transferred all contents of the proposed 28 oz Mill Lane oat milk for the current pickup order into sealed vessel Q17 without adding or losing any material."}, {"speaker": "Laboratory log", "text": "At 09:26 UTC on 17 September 2026, a calibrated quantitative assay measured exactly 4 grams of added sugar in all material inside sealed vessel Q17."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}, {"speaker": "Field note", "text": "Package inspection confirmed that the proposed Mill Lane carton is oat milk, bears an unsweetened label, and has a manufacturer-declared net quantity of 28 oz. The order record lists the requested Harvest Oat carton at 32 oz, so the proposed carton is smaller rather than equal in net quantity. Chain-of-custody records show that Q17 remained sealed and intact between transfer and assay."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the customer’s current-order constraints and the supervisor’s governing scope rule, while the unchanged questions retain the full scoring policy. The proposed item, order, size, entity, and September 17, 2026 inspection path remain bound consistently. The focus evidence contains exactly two complete factual sentences. Changing the calibrated-meter result from zero to six grams is coherent with the unchanged accounting statement that every dispensed gram from lot Q7 is added sugar; the unsweetened statement concerns the carton’s label and is not a duplicate measurement. Neither context includes a gold score, answer code, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Receiving inspector\",\"text\":\"At 08:00 on September 17, 2026, inspection of the proposed Mill Lane carton identified the product as oat milk, with an unsweetened label and a net quantity of 28 oz.\"},{\"speaker\":\"Ingredient records clerk\",\"text\":\"At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar.\"},{\"speaker\":\"Meter records clerk\",\"text\":\"At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that zero grams from ingredient lot Q7 had been dispensed into that carton before it was sealed.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar."}, {"path": ["3", "text"], "text": "At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that zero grams from ingredient lot Q7 had been dispensed into that carton before it was sealed."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar.", "negative_left": "At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar.", "negative_right": "At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that six grams from ingredient lot Q7 had been dispensed into that carton before it was sealed.", "right": "At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that zero grams from ingredient lot Q7 had been dispensed into that carton before it was sealed."}, "verifier_independent_model": false}, "family": "scale-diverse-024-005", "id": "scale-diverse-024-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Receiving inspector", "text": "At 08:00 on September 17, 2026, inspection of the proposed Mill Lane carton identified the product as oat milk, with an unsweetened label and a net quantity of 28 oz."}, {"speaker": "Ingredient records clerk", "text": "At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar."}, {"speaker": "Meter records clerk", "text": "At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that zero grams from ingredient lot Q7 had been dispensed into that carton before it was sealed."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the customer’s current-order constraints and the supervisor’s governing scope rule, while the unchanged questions retain the full scoring policy. The proposed item, order, size, entity, and September 17, 2026 inspection path remain bound consistently. The focus evidence contains exactly two complete factual sentences. Changing the calibrated-meter result from zero to six grams is coherent with the unchanged accounting statement that every dispensed gram from lot Q7 is added sugar; the unsweetened statement concerns the carton’s label and is not a duplicate measurement. Neither context includes a gold score, answer code, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Receiving inspector\",\"text\":\"At 08:00 on September 17, 2026, inspection of the proposed Mill Lane carton identified the product as oat milk, with an unsweetened label and a net quantity of 28 oz.\"},{\"speaker\":\"Ingredient records clerk\",\"text\":\"At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar.\"},{\"speaker\":\"Meter records clerk\",\"text\":\"At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that zero grams from ingredient lot Q7 had been dispensed into that carton before it was sealed.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar."}, {"path": ["3", "text"], "text": "At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that zero grams from ingredient lot Q7 had been dispensed into that carton before it was sealed."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar.", "negative_left": "At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar.", "negative_right": "At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that six grams from ingredient lot Q7 had been dispensed into that carton before it was sealed.", "right": "At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that zero grams from ingredient lot Q7 had been dispensed into that carton before it was sealed."}, "verifier_independent_model": false}, "family": "scale-diverse-024-005", "id": "scale-diverse-024-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Receiving inspector", "text": "At 08:00 on September 17, 2026, inspection of the proposed Mill Lane carton identified the product as oat milk, with an unsweetened label and a net quantity of 28 oz."}, {"speaker": "Ingredient records clerk", "text": "At 08:10 on September 17, 2026, the finalized ingredient-accounting record for the proposed 28 oz Mill Lane oat milk carton in the current pickup order classified every gram dispensed from ingredient lot Q7 as added sugar and classified no other input as added sugar."}, {"speaker": "Meter records clerk", "text": "At 08:25 on September 17, 2026, the complete calibrated-meter log recorded that six grams from ingredient lot Q7 had been dispensed into that carton before it was sealed."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing points policy and the same purchase-points scope, order K-184, and September 17, 2026 assessment binding. The two focus-evidence spans are complete factual sentences, and the counterfactual coherently changes only T-62’s referenced delivery date from September 8 to September 13 without creating conflicting measurements or assertions. Neither context embeds a classification answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom a5 is factual rather than policy. The base and counter assignments differ only on a5 and are realizable using different system-recorded delivery dates. The policy evidence correctly preserves the substantive earning, waiting-period, field-validation, and member-recollection rules originating in the original state; criteria and instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a purchase-points issue, both required system fields, a $200 eligible subtotal yielding 200 due points, expiration of the seven-full-day waiting period, and only 0 posted points. It also excludes redemption and tier-status issues, so it is sufficient for ready_valid_earning_discrepancy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a purchase-points issue and both required system fields, while refutation of a5 entails that seven full calendar days have not elapsed. Under the cited earning policy, the 200 points are therefore not currently due, and 0 posted points proves no current shortfall. Redemption and tier-status issues are also excluded, making the condition sufficient for ready_invalid_earning_claim.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported issue for order K-184 concerns points earned from the purchase."}, {"id": "a2", "statement": "The order system record for K-184 contains a delivered-at date."}, {"id": "a3", "statement": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"id": "a4", "statement": "As of September 17, the loyalty account has 0 posted points for order K-184."}, {"id": "a5", "statement": "For order K-184, as of September 17, at least seven full calendar days have elapsed since the system-recorded delivered-at date."}, {"id": "a6", "statement": "The reported issue for order K-184 concerns reward redemption."}, {"id": "a7", "statement": "The reported issue for order K-184 concerns tier status."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My report for order K-184 is only about points earned from that purchase. I am not seeking reward redemption or asking about tier status.\"},{\"speaker\":\"Order specialist\",\"text\":\"The system-recorded eligible-merchandise subtotal is $200. At the assessment, the loyalty account shows 0 posted points for K-184.\"},{\"speaker\":\"Records custodian\",\"text\":\"In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62.\"},{\"speaker\":\"Audit clerk\",\"text\":\"For the September 17 assessment of order K-184, entry code T-62 references September 8, 2026, and the assessment date is September 17, 2026.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62."}, {"path": ["3", "text"], "text": "For the September 17 assessment of order K-184, entry code T-62 references September 8, 2026, and the assessment date is September 17, 2026."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}], "rules": [{"justification": "Both required system fields are present. The $200 eligible-merchandise subtotal makes 200 points due once seven full calendar days have elapsed, while only 0 points are posted. The issue is purchase earning, not redemption or tier status.", "target": "ready_valid_earning_discrepancy", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Both required system fields are present, but seven full calendar days have not elapsed since the system-recorded delivery date. Therefore the purchase points are not currently due, so 0 posted points does not establish a current earning discrepancy.", "target": "ready_invalid_earning_claim", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62.", "negative_left": "In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62.", "negative_right": "For the September 17 assessment of order K-184, entry code T-62 references September 13, 2026, and the assessment date is September 17, 2026.", "right": "For the September 17 assessment of order K-184, entry code T-62 references September 8, 2026, and the assessment date is September 17, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-026-001", "id": "scale-diverse-026-001-base", "input": {"questions": {"decision": {"criteria": {"incomplete_request_required_records": "Select when the issue is a purchase-points claim but at least one required system field—the delivered-at date or eligible-merchandise subtotal—is missing.", "ready_invalid_earning_claim": "Select only if both required system fields are present and prove that no points are currently missing or due.", "ready_valid_earning_discrepancy": "Select only if both required system fields are present and prove that fewer points posted than were due after the waiting period.", "wrong_loyalty_workflow": "Select only if the reported issue concerns reward redemption or tier status rather than points earned from a purchase."}, "instructions": "Classify the claim’s workflow readiness. Apply the option rubrics exactly; the asserted delivery date being at the seven-day boundary does not count as a system record.", "type": "choice"}}, "state": [{"speaker": "Loyalty member", "text": "My report for order K-184 is only about points earned from that purchase. I am not seeking reward redemption or asking about tier status."}, {"speaker": "Order specialist", "text": "The system-recorded eligible-merchandise subtotal is $200. At the assessment, the loyalty account shows 0 posted points for K-184."}, {"speaker": "Records custodian", "text": "In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62."}, {"speaker": "Audit clerk", "text": "For the September 17 assessment of order K-184, entry code T-62 references September 8, 2026, and the assessment date is September 17, 2026."}, {"speaker": "Rewards program analyst", "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-026", "source_is_synthetic": true, "source_sha256": "0b515cddd9ebafd477dd65c9518b5ab7e0ccbe04906abd9ba0b4e5dae77b43a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_valid_earning_discrepancy"}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing points policy and the same purchase-points scope, order K-184, and September 17, 2026 assessment binding. The two focus-evidence spans are complete factual sentences, and the counterfactual coherently changes only T-62’s referenced delivery date from September 8 to September 13 without creating conflicting measurements or assertions. Neither context embeds a classification answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom a5 is factual rather than policy. The base and counter assignments differ only on a5 and are realizable using different system-recorded delivery dates. The policy evidence correctly preserves the substantive earning, waiting-period, field-validation, and member-recollection rules originating in the original state; criteria and instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a purchase-points issue, both required system fields, a $200 eligible subtotal yielding 200 due points, expiration of the seven-full-day waiting period, and only 0 posted points. It also excludes redemption and tier-status issues, so it is sufficient for ready_valid_earning_discrepancy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a purchase-points issue and both required system fields, while refutation of a5 entails that seven full calendar days have not elapsed. Under the cited earning policy, the 200 points are therefore not currently due, and 0 posted points proves no current shortfall. Redemption and tier-status issues are also excluded, making the condition sufficient for ready_invalid_earning_claim.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported issue for order K-184 concerns points earned from the purchase."}, {"id": "a2", "statement": "The order system record for K-184 contains a delivered-at date."}, {"id": "a3", "statement": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"id": "a4", "statement": "As of September 17, the loyalty account has 0 posted points for order K-184."}, {"id": "a5", "statement": "For order K-184, as of September 17, at least seven full calendar days have elapsed since the system-recorded delivered-at date."}, {"id": "a6", "statement": "The reported issue for order K-184 concerns reward redemption."}, {"id": "a7", "statement": "The reported issue for order K-184 concerns tier status."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My report for order K-184 is only about points earned from that purchase. I am not seeking reward redemption or asking about tier status.\"},{\"speaker\":\"Order specialist\",\"text\":\"The system-recorded eligible-merchandise subtotal is $200. At the assessment, the loyalty account shows 0 posted points for K-184.\"},{\"speaker\":\"Records custodian\",\"text\":\"In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62.\"},{\"speaker\":\"Audit clerk\",\"text\":\"For the September 17 assessment of order K-184, entry code T-62 references September 8, 2026, and the assessment date is September 17, 2026.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62."}, {"path": ["3", "text"], "text": "For the September 17 assessment of order K-184, entry code T-62 references September 8, 2026, and the assessment date is September 17, 2026."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}], "rules": [{"justification": "Both required system fields are present. The $200 eligible-merchandise subtotal makes 200 points due once seven full calendar days have elapsed, while only 0 points are posted. The issue is purchase earning, not redemption or tier status.", "target": "ready_valid_earning_discrepancy", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Both required system fields are present, but seven full calendar days have not elapsed since the system-recorded delivery date. Therefore the purchase points are not currently due, so 0 posted points does not establish a current earning discrepancy.", "target": "ready_invalid_earning_claim", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62.", "negative_left": "In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62.", "negative_right": "For the September 17 assessment of order K-184, entry code T-62 references September 13, 2026, and the assessment date is September 17, 2026.", "right": "For the September 17 assessment of order K-184, entry code T-62 references September 8, 2026, and the assessment date is September 17, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-026-001", "id": "scale-diverse-026-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"incomplete_request_required_records": "Select when the issue is a purchase-points claim but at least one required system field—the delivered-at date or eligible-merchandise subtotal—is missing.", "ready_invalid_earning_claim": "Select only if both required system fields are present and prove that no points are currently missing or due.", "ready_valid_earning_discrepancy": "Select only if both required system fields are present and prove that fewer points posted than were due after the waiting period.", "wrong_loyalty_workflow": "Select only if the reported issue concerns reward redemption or tier status rather than points earned from a purchase."}, "instructions": "Classify the claim’s workflow readiness. Apply the option rubrics exactly; the asserted delivery date being at the seven-day boundary does not count as a system record.", "type": "choice"}}, "state": [{"speaker": "Loyalty member", "text": "My report for order K-184 is only about points earned from that purchase. I am not seeking reward redemption or asking about tier status."}, {"speaker": "Order specialist", "text": "The system-recorded eligible-merchandise subtotal is $200. At the assessment, the loyalty account shows 0 posted points for K-184."}, {"speaker": "Records custodian", "text": "In the order system record for K-184, the delivered-at field contains the date referenced by entry code T-62."}, {"speaker": "Audit clerk", "text": "For the September 17 assessment of order K-184, entry code T-62 references September 13, 2026, and the assessment date is September 17, 2026."}, {"speaker": "Rewards program analyst", "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-026", "source_is_synthetic": true, "source_sha256": "0b515cddd9ebafd477dd65c9518b5ab7e0ccbe04906abd9ba0b4e5dae77b43a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_invalid_earning_claim"}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original purchase-points policy, required-record rules, order K-184, and the September 17, 2026 assessment framing. The two evidence spans are complete factual sentences. The counterfactual changes only event E-73’s recorded date from September 9 to September 12, which coherently changes the system-derived delivery date without conflicting with unchanged facts. Neither context contains an answer label, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom a5 is factual rather than policy. The base and counter assignments differ only on a5 and are realizable using different system-recorded delivery dates. The policy evidence correctly preserves the substantive earning, waiting-period, field-validation, and member-recollection rules originating in the original state; criteria and instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a purchase-points issue, both required system fields, a $200 eligible subtotal yielding 200 due points, expiration of the seven-full-day waiting period, and only 0 posted points. It also excludes redemption and tier-status issues, so it is sufficient for ready_valid_earning_discrepancy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a purchase-points issue and both required system fields, while refutation of a5 entails that seven full calendar days have not elapsed. Under the cited earning policy, the 200 points are therefore not currently due, and 0 posted points proves no current shortfall. Redemption and tier-status issues are also excluded, making the condition sufficient for ready_invalid_earning_claim.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported issue for order K-184 concerns points earned from the purchase."}, {"id": "a2", "statement": "The order system record for K-184 contains a delivered-at date."}, {"id": "a3", "statement": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"id": "a4", "statement": "As of September 17, the loyalty account has 0 posted points for order K-184."}, {"id": "a5", "statement": "For order K-184, as of September 17, at least seven full calendar days have elapsed since the system-recorded delivered-at date."}, {"id": "a6", "statement": "The reported issue for order K-184 concerns reward redemption."}, {"id": "a7", "statement": "The reported issue for order K-184 concerns tier status."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My report for order K-184 is about points earned from this purchase. It is not a reward-redemption request or a question about tier status.\"},{\"speaker\":\"Order records specialist\",\"text\":\"The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred.\"},{\"speaker\":\"Ledger specialist\",\"text\":\"In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 9, 2026.\"},{\"speaker\":\"Order records specialist\",\"text\":\"The system-recorded eligible-merchandise subtotal for K-184 is $200; tax and other ineligible amounts are recorded separately.\"},{\"speaker\":\"Loyalty account specialist\",\"text\":\"As of September 17, 2026, the loyalty account shows 0 posted points for K-184.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred."}, {"path": ["2", "text"], "text": "In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 9, 2026."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}], "rules": [{"justification": "Both required system fields are present. The $200 eligible-merchandise subtotal makes 200 points due once seven full calendar days have elapsed, while only 0 points are posted. The issue is purchase earning, not redemption or tier status.", "target": "ready_valid_earning_discrepancy", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Both required system fields are present, but seven full calendar days have not elapsed since the system-recorded delivery date. Therefore the purchase points are not currently due, so 0 posted points does not establish a current earning discrepancy.", "target": "ready_invalid_earning_claim", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred.", "negative_left": "The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred.", "negative_right": "In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 12, 2026.", "right": "In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 9, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-026-002", "id": "scale-diverse-026-002-base", "input": {"questions": {"decision": {"criteria": {"incomplete_request_required_records": "Select when the issue is a purchase-points claim but at least one required system field—the delivered-at date or eligible-merchandise subtotal—is missing.", "ready_invalid_earning_claim": "Select only if both required system fields are present and prove that no points are currently missing or due.", "ready_valid_earning_discrepancy": "Select only if both required system fields are present and prove that fewer points posted than were due after the waiting period.", "wrong_loyalty_workflow": "Select only if the reported issue concerns reward redemption or tier status rather than points earned from a purchase."}, "instructions": "Classify the claim’s workflow readiness. Apply the option rubrics exactly; the asserted delivery date being at the seven-day boundary does not count as a system record.", "type": "choice"}}, "state": [{"speaker": "Loyalty member", "text": "My report for order K-184 is about points earned from this purchase. It is not a reward-redemption request or a question about tier status."}, {"speaker": "Order records specialist", "text": "The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred."}, {"speaker": "Ledger specialist", "text": "In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 9, 2026."}, {"speaker": "Order records specialist", "text": "The system-recorded eligible-merchandise subtotal for K-184 is $200; tax and other ineligible amounts are recorded separately."}, {"speaker": "Loyalty account specialist", "text": "As of September 17, 2026, the loyalty account shows 0 posted points for K-184."}, {"speaker": "Rewards program analyst", "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-026", "source_is_synthetic": true, "source_sha256": "0b515cddd9ebafd477dd65c9518b5ab7e0ccbe04906abd9ba0b4e5dae77b43a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_valid_earning_discrepancy"}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original purchase-points policy, required-record rules, order K-184, and the September 17, 2026 assessment framing. The two evidence spans are complete factual sentences. The counterfactual changes only event E-73’s recorded date from September 9 to September 12, which coherently changes the system-derived delivery date without conflicting with unchanged facts. Neither context contains an answer label, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom a5 is factual rather than policy. The base and counter assignments differ only on a5 and are realizable using different system-recorded delivery dates. The policy evidence correctly preserves the substantive earning, waiting-period, field-validation, and member-recollection rules originating in the original state; criteria and instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a purchase-points issue, both required system fields, a $200 eligible subtotal yielding 200 due points, expiration of the seven-full-day waiting period, and only 0 posted points. It also excludes redemption and tier-status issues, so it is sufficient for ready_valid_earning_discrepancy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a purchase-points issue and both required system fields, while refutation of a5 entails that seven full calendar days have not elapsed. Under the cited earning policy, the 200 points are therefore not currently due, and 0 posted points proves no current shortfall. Redemption and tier-status issues are also excluded, making the condition sufficient for ready_invalid_earning_claim.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported issue for order K-184 concerns points earned from the purchase."}, {"id": "a2", "statement": "The order system record for K-184 contains a delivered-at date."}, {"id": "a3", "statement": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"id": "a4", "statement": "As of September 17, the loyalty account has 0 posted points for order K-184."}, {"id": "a5", "statement": "For order K-184, as of September 17, at least seven full calendar days have elapsed since the system-recorded delivered-at date."}, {"id": "a6", "statement": "The reported issue for order K-184 concerns reward redemption."}, {"id": "a7", "statement": "The reported issue for order K-184 concerns tier status."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My report for order K-184 is about points earned from this purchase. It is not a reward-redemption request or a question about tier status.\"},{\"speaker\":\"Order records specialist\",\"text\":\"The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred.\"},{\"speaker\":\"Ledger specialist\",\"text\":\"In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 9, 2026.\"},{\"speaker\":\"Order records specialist\",\"text\":\"The system-recorded eligible-merchandise subtotal for K-184 is $200; tax and other ineligible amounts are recorded separately.\"},{\"speaker\":\"Loyalty account specialist\",\"text\":\"As of September 17, 2026, the loyalty account shows 0 posted points for K-184.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred."}, {"path": ["2", "text"], "text": "In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 9, 2026."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}], "rules": [{"justification": "Both required system fields are present. The $200 eligible-merchandise subtotal makes 200 points due once seven full calendar days have elapsed, while only 0 points are posted. The issue is purchase earning, not redemption or tier status.", "target": "ready_valid_earning_discrepancy", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Both required system fields are present, but seven full calendar days have not elapsed since the system-recorded delivery date. Therefore the purchase points are not currently due, so 0 posted points does not establish a current earning discrepancy.", "target": "ready_invalid_earning_claim", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred.", "negative_left": "The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred.", "negative_right": "In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 12, 2026.", "right": "In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 9, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-026-002", "id": "scale-diverse-026-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"incomplete_request_required_records": "Select when the issue is a purchase-points claim but at least one required system field—the delivered-at date or eligible-merchandise subtotal—is missing.", "ready_invalid_earning_claim": "Select only if both required system fields are present and prove that no points are currently missing or due.", "ready_valid_earning_discrepancy": "Select only if both required system fields are present and prove that fewer points posted than were due after the waiting period.", "wrong_loyalty_workflow": "Select only if the reported issue concerns reward redemption or tier status rather than points earned from a purchase."}, "instructions": "Classify the claim’s workflow readiness. Apply the option rubrics exactly; the asserted delivery date being at the seven-day boundary does not count as a system record.", "type": "choice"}}, "state": [{"speaker": "Loyalty member", "text": "My report for order K-184 is about points earned from this purchase. It is not a reward-redemption request or a question about tier status."}, {"speaker": "Order records specialist", "text": "The order system records order K-184's delivered-at date as the UTC calendar date on which logistics event E-73 occurred."}, {"speaker": "Ledger specialist", "text": "In the UTC ledger used for the September 17, 2026 assessment, logistics event E-73 occurred on September 12, 2026."}, {"speaker": "Order records specialist", "text": "The system-recorded eligible-merchandise subtotal for K-184 is $200; tax and other ineligible amounts are recorded separately."}, {"speaker": "Loyalty account specialist", "text": "As of September 17, 2026, the loyalty account shows 0 posted points for K-184."}, {"speaker": "Rewards program analyst", "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-026", "source_is_synthetic": true, "source_sha256": "0b515cddd9ebafd477dd65c9518b5ab7e0ccbe04906abd9ba0b4e5dae77b43a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_invalid_earning_claim"}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing rewards policy and retain the same claim scope, order K-184, and September 17, 2026 review binding. The two evidence spans are complete factual sentences, and the counterfactual coherently changes audit entry Q-62’s delivery date from September 8 to September 14 without conflicting with unchanged facts. Neither context contains an answer label, code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom a5 is factual rather than policy. The base and counter assignments differ only on a5 and are realizable using different system-recorded delivery dates. The policy evidence correctly preserves the substantive earning, waiting-period, field-validation, and member-recollection rules originating in the original state; criteria and instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a purchase-points issue, both required system fields, a $200 eligible subtotal yielding 200 due points, expiration of the seven-full-day waiting period, and only 0 posted points. It also excludes redemption and tier-status issues, so it is sufficient for ready_valid_earning_discrepancy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a purchase-points issue and both required system fields, while refutation of a5 entails that seven full calendar days have not elapsed. Under the cited earning policy, the 200 points are therefore not currently due, and 0 posted points proves no current shortfall. Redemption and tier-status issues are also excluded, making the condition sufficient for ready_invalid_earning_claim.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported issue for order K-184 concerns points earned from the purchase."}, {"id": "a2", "statement": "The order system record for K-184 contains a delivered-at date."}, {"id": "a3", "statement": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"id": "a4", "statement": "As of September 17, the loyalty account has 0 posted points for order K-184."}, {"id": "a5", "statement": "For order K-184, as of September 17, at least seven full calendar days have elapsed since the system-recorded delivered-at date."}, {"id": "a6", "statement": "The reported issue for order K-184 concerns reward redemption."}, {"id": "a7", "statement": "The reported issue for order K-184 concerns tier status."}], "base_state_json": "[{\"speaker\":\"Case intake\",\"text\":\"The member’s report for order K-184 concerns only points earned from that purchase; it does not concern reward redemption or tier status.\"},{\"speaker\":\"Order audit\",\"text\":\"For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date.\"},{\"speaker\":\"Audit registry\",\"text\":\"Audit entry Q-62 shows September 8, 2026.\"},{\"speaker\":\"Order record\",\"text\":\"The system-recorded eligible-merchandise subtotal for order K-184 is $200.\"},{\"speaker\":\"Loyalty ledger\",\"text\":\"As of the review, the loyalty account shows 0 posted points for order K-184.\"},{\"speaker\":\"Document check\",\"text\":\"The submitted receipt lists a $218 combined total without separately identifying merchandise and tax amounts.\"},{\"speaker\":\"Account review\",\"text\":\"The analyst matched the loyalty account and purchase record using the K-184 order reference.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date."}, {"path": ["2", "text"], "text": "Audit entry Q-62 shows September 8, 2026."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}], "rules": [{"justification": "Both required system fields are present. The $200 eligible-merchandise subtotal makes 200 points due once seven full calendar days have elapsed, while only 0 points are posted. The issue is purchase earning, not redemption or tier status.", "target": "ready_valid_earning_discrepancy", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Both required system fields are present, but seven full calendar days have not elapsed since the system-recorded delivery date. Therefore the purchase points are not currently due, so 0 posted points does not establish a current earning discrepancy.", "target": "ready_invalid_earning_claim", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date.", "negative_left": "For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date.", "negative_right": "Audit entry Q-62 shows September 14, 2026.", "right": "Audit entry Q-62 shows September 8, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-026-004", "id": "scale-diverse-026-004-base", "input": {"questions": {"decision": {"criteria": {"incomplete_request_required_records": "Select when the issue is a purchase-points claim but at least one required system field—the delivered-at date or eligible-merchandise subtotal—is missing.", "ready_invalid_earning_claim": "Select only if both required system fields are present and prove that no points are currently missing or due.", "ready_valid_earning_discrepancy": "Select only if both required system fields are present and prove that fewer points posted than were due after the waiting period.", "wrong_loyalty_workflow": "Select only if the reported issue concerns reward redemption or tier status rather than points earned from a purchase."}, "instructions": "Classify the claim’s workflow readiness. Apply the option rubrics exactly; the asserted delivery date being at the seven-day boundary does not count as a system record.", "type": "choice"}}, "state": [{"speaker": "Case intake", "text": "The member’s report for order K-184 concerns only points earned from that purchase; it does not concern reward redemption or tier status."}, {"speaker": "Order audit", "text": "For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date."}, {"speaker": "Audit registry", "text": "Audit entry Q-62 shows September 8, 2026."}, {"speaker": "Order record", "text": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"speaker": "Loyalty ledger", "text": "As of the review, the loyalty account shows 0 posted points for order K-184."}, {"speaker": "Document check", "text": "The submitted receipt lists a $218 combined total without separately identifying merchandise and tax amounts."}, {"speaker": "Account review", "text": "The analyst matched the loyalty account and purchase record using the K-184 order reference."}, {"speaker": "Rewards program analyst", "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-026", "source_is_synthetic": true, "source_sha256": "0b515cddd9ebafd477dd65c9518b5ab7e0ccbe04906abd9ba0b4e5dae77b43a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_valid_earning_discrepancy"}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing rewards policy and retain the same claim scope, order K-184, and September 17, 2026 review binding. The two evidence spans are complete factual sentences, and the counterfactual coherently changes audit entry Q-62’s delivery date from September 8 to September 14 without conflicting with unchanged facts. Neither context contains an answer label, code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom a5 is factual rather than policy. The base and counter assignments differ only on a5 and are realizable using different system-recorded delivery dates. The policy evidence correctly preserves the substantive earning, waiting-period, field-validation, and member-recollection rules originating in the original state; criteria and instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a purchase-points issue, both required system fields, a $200 eligible subtotal yielding 200 due points, expiration of the seven-full-day waiting period, and only 0 posted points. It also excludes redemption and tier-status issues, so it is sufficient for ready_valid_earning_discrepancy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a purchase-points issue and both required system fields, while refutation of a5 entails that seven full calendar days have not elapsed. Under the cited earning policy, the 200 points are therefore not currently due, and 0 posted points proves no current shortfall. Redemption and tier-status issues are also excluded, making the condition sufficient for ready_invalid_earning_claim.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported issue for order K-184 concerns points earned from the purchase."}, {"id": "a2", "statement": "The order system record for K-184 contains a delivered-at date."}, {"id": "a3", "statement": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"id": "a4", "statement": "As of September 17, the loyalty account has 0 posted points for order K-184."}, {"id": "a5", "statement": "For order K-184, as of September 17, at least seven full calendar days have elapsed since the system-recorded delivered-at date."}, {"id": "a6", "statement": "The reported issue for order K-184 concerns reward redemption."}, {"id": "a7", "statement": "The reported issue for order K-184 concerns tier status."}], "base_state_json": "[{\"speaker\":\"Case intake\",\"text\":\"The member’s report for order K-184 concerns only points earned from that purchase; it does not concern reward redemption or tier status.\"},{\"speaker\":\"Order audit\",\"text\":\"For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date.\"},{\"speaker\":\"Audit registry\",\"text\":\"Audit entry Q-62 shows September 8, 2026.\"},{\"speaker\":\"Order record\",\"text\":\"The system-recorded eligible-merchandise subtotal for order K-184 is $200.\"},{\"speaker\":\"Loyalty ledger\",\"text\":\"As of the review, the loyalty account shows 0 posted points for order K-184.\"},{\"speaker\":\"Document check\",\"text\":\"The submitted receipt lists a $218 combined total without separately identifying merchandise and tax amounts.\"},{\"speaker\":\"Account review\",\"text\":\"The analyst matched the loyalty account and purchase record using the K-184 order reference.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date."}, {"path": ["2", "text"], "text": "Audit entry Q-62 shows September 8, 2026."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}], "rules": [{"justification": "Both required system fields are present. The $200 eligible-merchandise subtotal makes 200 points due once seven full calendar days have elapsed, while only 0 points are posted. The issue is purchase earning, not redemption or tier status.", "target": "ready_valid_earning_discrepancy", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Both required system fields are present, but seven full calendar days have not elapsed since the system-recorded delivery date. Therefore the purchase points are not currently due, so 0 posted points does not establish a current earning discrepancy.", "target": "ready_invalid_earning_claim", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date.", "negative_left": "For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date.", "negative_right": "Audit entry Q-62 shows September 14, 2026.", "right": "Audit entry Q-62 shows September 8, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-026-004", "id": "scale-diverse-026-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"incomplete_request_required_records": "Select when the issue is a purchase-points claim but at least one required system field—the delivered-at date or eligible-merchandise subtotal—is missing.", "ready_invalid_earning_claim": "Select only if both required system fields are present and prove that no points are currently missing or due.", "ready_valid_earning_discrepancy": "Select only if both required system fields are present and prove that fewer points posted than were due after the waiting period.", "wrong_loyalty_workflow": "Select only if the reported issue concerns reward redemption or tier status rather than points earned from a purchase."}, "instructions": "Classify the claim’s workflow readiness. Apply the option rubrics exactly; the asserted delivery date being at the seven-day boundary does not count as a system record.", "type": "choice"}}, "state": [{"speaker": "Case intake", "text": "The member’s report for order K-184 concerns only points earned from that purchase; it does not concern reward redemption or tier status."}, {"speaker": "Order audit", "text": "For the September 17, 2026 review, the order system identifies the date shown on audit entry Q-62 as order K-184's delivered-at date."}, {"speaker": "Audit registry", "text": "Audit entry Q-62 shows September 14, 2026."}, {"speaker": "Order record", "text": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"speaker": "Loyalty ledger", "text": "As of the review, the loyalty account shows 0 posted points for order K-184."}, {"speaker": "Document check", "text": "The submitted receipt lists a $218 combined total without separately identifying merchandise and tax amounts."}, {"speaker": "Account review", "text": "The analyst matched the loyalty account and purchase record using the K-184 order reference."}, {"speaker": "Rewards program analyst", "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-026", "source_is_synthetic": true, "source_sha256": "0b515cddd9ebafd477dd65c9518b5ab7e0ccbe04906abd9ba0b4e5dae77b43a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_invalid_earning_claim"}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing rewards policy from the original state, while the unchanged questions preserve the decision criteria and instructions. The request remains a purchase-points claim for order K-184 evaluated at the September 17 cutoff using the system-recorded delivery event and eligible subtotal. The two evidence spans are complete factual sentences. Changing the elapsed time from eight days to six days is coherent within the counterfactual and creates no duplicate conflicting measurement inside that context. Neither context includes an answer label, code, proposition ID, rule table, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom a5 is factual rather than policy. The base and counter assignments differ only on a5 and are realizable using different system-recorded delivery dates. The policy evidence correctly preserves the substantive earning, waiting-period, field-validation, and member-recollection rules originating in the original state; criteria and instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a purchase-points issue, both required system fields, a $200 eligible subtotal yielding 200 due points, expiration of the seven-full-day waiting period, and only 0 posted points. It also excludes redemption and tier-status issues, so it is sufficient for ready_valid_earning_discrepancy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a purchase-points issue and both required system fields, while refutation of a5 entails that seven full calendar days have not elapsed. Under the cited earning policy, the 200 points are therefore not currently due, and 0 posted points proves no current shortfall. Redemption and tier-status issues are also excluded, making the condition sufficient for ready_invalid_earning_claim.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported issue for order K-184 concerns points earned from the purchase."}, {"id": "a2", "statement": "The order system record for K-184 contains a delivered-at date."}, {"id": "a3", "statement": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"id": "a4", "statement": "As of September 17, the loyalty account has 0 posted points for order K-184."}, {"id": "a5", "statement": "For order K-184, as of September 17, at least seven full calendar days have elapsed since the system-recorded delivered-at date."}, {"id": "a6", "statement": "The reported issue for order K-184 concerns reward redemption."}, {"id": "a7", "statement": "The reported issue for order K-184 concerns tier status."}], "base_state_json": "[{\"speaker\":\"Member\",\"text\":\"I am reporting missing points earned from the purchase on order K-184. This is not a request to redeem rewards or to review my tier status.\"},{\"speaker\":\"Account specialist\",\"text\":\"At the September 17 review cutoff, the loyalty ledger shows 0 posted points for K-184. The order system lists an eligible-merchandise subtotal of $200.\"},{\"speaker\":\"Operations reviewer\",\"text\":\"The chronology for order K-184 records exactly eight full calendar days as elapsed from event E-72 to the September 17 cutoff.\"},{\"speaker\":\"Records analyst\",\"text\":\"Event E-72 is the system-recorded delivered-at event for order K-184.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "The chronology for order K-184 records exactly eight full calendar days as elapsed from event E-72 to the September 17 cutoff."}, {"path": ["3", "text"], "text": "Event E-72 is the system-recorded delivered-at event for order K-184."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}], "rules": [{"justification": "Both required system fields are present. The $200 eligible-merchandise subtotal makes 200 points due once seven full calendar days have elapsed, while only 0 points are posted. The issue is purchase earning, not redemption or tier status.", "target": "ready_valid_earning_discrepancy", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Both required system fields are present, but seven full calendar days have not elapsed since the system-recorded delivery date. Therefore the purchase points are not currently due, so 0 posted points does not establish a current earning discrepancy.", "target": "ready_invalid_earning_claim", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The chronology for order K-184 records exactly eight full calendar days as elapsed from event E-72 to the September 17 cutoff.", "negative_left": "The chronology for order K-184 records exactly six full calendar days as elapsed from event E-72 to the September 17 cutoff.", "negative_right": "Event E-72 is the system-recorded delivered-at event for order K-184.", "right": "Event E-72 is the system-recorded delivered-at event for order K-184."}, "verifier_independent_model": false}, "family": "scale-diverse-026-005", "id": "scale-diverse-026-005-base", "input": {"questions": {"decision": {"criteria": {"incomplete_request_required_records": "Select when the issue is a purchase-points claim but at least one required system field—the delivered-at date or eligible-merchandise subtotal—is missing.", "ready_invalid_earning_claim": "Select only if both required system fields are present and prove that no points are currently missing or due.", "ready_valid_earning_discrepancy": "Select only if both required system fields are present and prove that fewer points posted than were due after the waiting period.", "wrong_loyalty_workflow": "Select only if the reported issue concerns reward redemption or tier status rather than points earned from a purchase."}, "instructions": "Classify the claim’s workflow readiness. Apply the option rubrics exactly; the asserted delivery date being at the seven-day boundary does not count as a system record.", "type": "choice"}}, "state": [{"speaker": "Member", "text": "I am reporting missing points earned from the purchase on order K-184. This is not a request to redeem rewards or to review my tier status."}, {"speaker": "Account specialist", "text": "At the September 17 review cutoff, the loyalty ledger shows 0 posted points for K-184. The order system lists an eligible-merchandise subtotal of $200."}, {"speaker": "Operations reviewer", "text": "The chronology for order K-184 records exactly eight full calendar days as elapsed from event E-72 to the September 17 cutoff."}, {"speaker": "Records analyst", "text": "Event E-72 is the system-recorded delivered-at event for order K-184."}, {"speaker": "Rewards program analyst", "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-026", "source_is_synthetic": true, "source_sha256": "0b515cddd9ebafd477dd65c9518b5ab7e0ccbe04906abd9ba0b4e5dae77b43a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_valid_earning_discrepancy"}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing rewards policy from the original state, while the unchanged questions preserve the decision criteria and instructions. The request remains a purchase-points claim for order K-184 evaluated at the September 17 cutoff using the system-recorded delivery event and eligible subtotal. The two evidence spans are complete factual sentences. Changing the elapsed time from eight days to six days is coherent within the counterfactual and creates no duplicate conflicting measurement inside that context. Neither context includes an answer label, code, proposition ID, rule table, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom a5 is factual rather than policy. The base and counter assignments differ only on a5 and are realizable using different system-recorded delivery dates. The policy evidence correctly preserves the substantive earning, waiting-period, field-validation, and member-recollection rules originating in the original state; criteria and instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a purchase-points issue, both required system fields, a $200 eligible subtotal yielding 200 due points, expiration of the seven-full-day waiting period, and only 0 posted points. It also excludes redemption and tier-status issues, so it is sufficient for ready_valid_earning_discrepancy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a purchase-points issue and both required system fields, while refutation of a5 entails that seven full calendar days have not elapsed. Under the cited earning policy, the 200 points are therefore not currently due, and 0 posted points proves no current shortfall. Redemption and tier-status issues are also excluded, making the condition sufficient for ready_invalid_earning_claim.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported issue for order K-184 concerns points earned from the purchase."}, {"id": "a2", "statement": "The order system record for K-184 contains a delivered-at date."}, {"id": "a3", "statement": "The system-recorded eligible-merchandise subtotal for order K-184 is $200."}, {"id": "a4", "statement": "As of September 17, the loyalty account has 0 posted points for order K-184."}, {"id": "a5", "statement": "For order K-184, as of September 17, at least seven full calendar days have elapsed since the system-recorded delivered-at date."}, {"id": "a6", "statement": "The reported issue for order K-184 concerns reward redemption."}, {"id": "a7", "statement": "The reported issue for order K-184 concerns tier status."}], "base_state_json": "[{\"speaker\":\"Member\",\"text\":\"I am reporting missing points earned from the purchase on order K-184. This is not a request to redeem rewards or to review my tier status.\"},{\"speaker\":\"Account specialist\",\"text\":\"At the September 17 review cutoff, the loyalty ledger shows 0 posted points for K-184. The order system lists an eligible-merchandise subtotal of $200.\"},{\"speaker\":\"Operations reviewer\",\"text\":\"The chronology for order K-184 records exactly eight full calendar days as elapsed from event E-72 to the September 17 cutoff.\"},{\"speaker\":\"Records analyst\",\"text\":\"Event E-72 is the system-recorded delivered-at event for order K-184.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "The chronology for order K-184 records exactly eight full calendar days as elapsed from event E-72 to the September 17 cutoff."}, {"path": ["3", "text"], "text": "Event E-72 is the system-recorded delivered-at event for order K-184."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}], "rules": [{"justification": "Both required system fields are present. The $200 eligible-merchandise subtotal makes 200 points due once seven full calendar days have elapsed, while only 0 points are posted. The issue is purchase earning, not redemption or tier status.", "target": "ready_valid_earning_discrepancy", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Both required system fields are present, but seven full calendar days have not elapsed since the system-recorded delivery date. Therefore the purchase points are not currently due, so 0 posted points does not establish a current earning discrepancy.", "target": "ready_invalid_earning_claim", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The chronology for order K-184 records exactly eight full calendar days as elapsed from event E-72 to the September 17 cutoff.", "negative_left": "The chronology for order K-184 records exactly six full calendar days as elapsed from event E-72 to the September 17 cutoff.", "negative_right": "Event E-72 is the system-recorded delivered-at event for order K-184.", "right": "Event E-72 is the system-recorded delivered-at event for order K-184."}, "verifier_independent_model": false}, "family": "scale-diverse-026-005", "id": "scale-diverse-026-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"incomplete_request_required_records": "Select when the issue is a purchase-points claim but at least one required system field—the delivered-at date or eligible-merchandise subtotal—is missing.", "ready_invalid_earning_claim": "Select only if both required system fields are present and prove that no points are currently missing or due.", "ready_valid_earning_discrepancy": "Select only if both required system fields are present and prove that fewer points posted than were due after the waiting period.", "wrong_loyalty_workflow": "Select only if the reported issue concerns reward redemption or tier status rather than points earned from a purchase."}, "instructions": "Classify the claim’s workflow readiness. Apply the option rubrics exactly; the asserted delivery date being at the seven-day boundary does not count as a system record.", "type": "choice"}}, "state": [{"speaker": "Member", "text": "I am reporting missing points earned from the purchase on order K-184. This is not a request to redeem rewards or to review my tier status."}, {"speaker": "Account specialist", "text": "At the September 17 review cutoff, the loyalty ledger shows 0 posted points for K-184. The order system lists an eligible-merchandise subtotal of $200."}, {"speaker": "Operations reviewer", "text": "The chronology for order K-184 records exactly six full calendar days as elapsed from event E-72 to the September 17 cutoff."}, {"speaker": "Records analyst", "text": "Event E-72 is the system-recorded delivered-at event for order K-184."}, {"speaker": "Rewards program analyst", "text": "Policy awards one point per eligible merchandise dollar, excluding tax, seven full calendar days after delivery. Validation requires the system-recorded delivery date and eligible subtotal; a member’s recollection cannot replace either field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-026", "source_is_synthetic": true, "source_sha256": "0b515cddd9ebafd477dd65c9518b5ab7e0ccbe04906abd9ba0b4e5dae77b43a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_invalid_earning_claim"}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same base-points rule, gift-card exclusion, and promotion scope without adding exceptions or defaults. Priya, transaction T1, receipt R1, the 200-point claim, and the relevant timestamps and record paths remain fixed. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the complete registry’s membership so that CCG-4817 is no longer identified as a gift-card SKU; this does not conflict with the unchanged statement that the item otherwise qualifies apart from any applicable gift-card exclusion. Neither constructed context embeds an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the two universal registry relations. The focus atom is the factual relation of SKU membership, not a policy conclusion. The base and counter assignments are jointly realizable and differ only on that focus: registry membership makes the item a gift card in the base case, while nonmembership plus registry exhaustiveness makes it a non-gift-card item in the countercase. The policy evidence correctly preserves the substantive program and promotion rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The sole item’s SKU is in the registry, and every registered SKU identifies a gift card. Therefore the transaction is a gift-card purchase, the stated exclusion applies, and the promotion cannot create points because it only doubles transactions that first qualify for base points. Priya is not entitled to the claimed 200 points.", "rule_index": 0, "sound": true}, {"reason": "Refutation of registry membership means the sole item’s SKU is not registered. Because every gift-card SKU is stated to appear in the registry, the item cannot be a gift card. Atom a5 establishes eligibility apart from that exclusion; the sole $100 item therefore earns 100 base points and 200 during the double-points weekend. With zero points awarded and a claim for 200, the claim is valid.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya Shah's missing-points claim requests 200 Cedar Circle points for transaction T1."}, {"id": "a2", "statement": "Receipt R1 documents transaction T1."}, {"id": "a3", "statement": "The total price shown on receipt R1 is $100."}, {"id": "a4", "statement": "Receipt R1 contains exactly one line item."}, {"id": "a5", "statement": "The sole line item on receipt R1 qualifies for base points apart from the gift-card exclusion."}, {"id": "a6", "statement": "The SKU of the sole line item on receipt R1 is a member of the program's gift-card SKU registry."}, {"id": "a7", "statement": "Every SKU in the program's gift-card SKU registry identifies a gift-card product."}, {"id": "a8", "statement": "Every gift-card product SKU appears in the program's gift-card SKU registry."}, {"id": "a9", "statement": "Transaction T1 occurred during the double-points weekend."}, {"id": "a10", "statement": "Priya Shah's account received zero Cedar Circle points for transaction T1."}], "base_state_json": "\"Priya Shah filed a missing-points claim requesting 200 Cedar Circle points for transaction T1. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. The item otherwise qualifies for base points apart from the gift-card exclusion. T1 occurred during the double-points weekend, but her account received zero Cedar Circle points. At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817. At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-4817, CCG-6392, and CCG-8504. Program administration confirms that registry entries identify gift-card products and that every gift-card product SKU is included in that complete registry. Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817."}, {"path": [], "text": "At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-4817, CCG-6392, and CCG-8504."}], "policy_evidence": [{"path": [], "text": "Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases."}, {"path": [], "text": "The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}], "rules": [{"justification": "The sole item on the claimed $100 transaction has a SKU in the gift-card registry, and every registry SKU identifies a gift-card product. The gift-card exclusion therefore applies, and the promotion cannot override it.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "The exhaustive registry contains every gift-card SKU, while the sole item's SKU is not in that registry, so the gift-card exclusion does not apply. The otherwise base-eligible $100 transaction earns 100 base points and, because it occurred during the double-points weekend, 200 total points; zero were awarded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817.", "negative_left": "At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817.", "negative_right": "At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-2751, CCG-6392, and CCG-8504.", "right": "At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-4817, CCG-6392, and CCG-8504."}, "verifier_independent_model": false}, "family": "scale-diverse-027-001", "id": "scale-diverse-027-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the gift-card exclusion applies, so Priya was not entitled to the claimed points.", "true": "Yes — the records and applicable rules show that Priya should have received the claimed 200 points."}, "instructions": "Decide whether Priya's claim for 200 missing points is valid under the supplied program and promotion rules. Answer yes or no.", "type": "noul"}}, "state": "Priya Shah filed a missing-points claim requesting 200 Cedar Circle points for transaction T1. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. The item otherwise qualifies for base points apart from the gift-card exclusion. T1 occurred during the double-points weekend, but her account received zero Cedar Circle points. At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817. At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-4817, CCG-6392, and CCG-8504. Program administration confirms that registry entries identify gift-card products and that every gift-card product SKU is included in that complete registry. Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}, "method": "c2d", "provenance": {"source_id": "diverse-027", "source_is_synthetic": true, "source_sha256": "24a57454b60c3c253b20a328c3963efba4da8dd26a29c30742ff6568152211d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same base-points rule, gift-card exclusion, and promotion scope without adding exceptions or defaults. Priya, transaction T1, receipt R1, the 200-point claim, and the relevant timestamps and record paths remain fixed. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the complete registry’s membership so that CCG-4817 is no longer identified as a gift-card SKU; this does not conflict with the unchanged statement that the item otherwise qualifies apart from any applicable gift-card exclusion. Neither constructed context embeds an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the two universal registry relations. The focus atom is the factual relation of SKU membership, not a policy conclusion. The base and counter assignments are jointly realizable and differ only on that focus: registry membership makes the item a gift card in the base case, while nonmembership plus registry exhaustiveness makes it a non-gift-card item in the countercase. The policy evidence correctly preserves the substantive program and promotion rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The sole item’s SKU is in the registry, and every registered SKU identifies a gift card. Therefore the transaction is a gift-card purchase, the stated exclusion applies, and the promotion cannot create points because it only doubles transactions that first qualify for base points. Priya is not entitled to the claimed 200 points.", "rule_index": 0, "sound": true}, {"reason": "Refutation of registry membership means the sole item’s SKU is not registered. Because every gift-card SKU is stated to appear in the registry, the item cannot be a gift card. Atom a5 establishes eligibility apart from that exclusion; the sole $100 item therefore earns 100 base points and 200 during the double-points weekend. With zero points awarded and a claim for 200, the claim is valid.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya Shah's missing-points claim requests 200 Cedar Circle points for transaction T1."}, {"id": "a2", "statement": "Receipt R1 documents transaction T1."}, {"id": "a3", "statement": "The total price shown on receipt R1 is $100."}, {"id": "a4", "statement": "Receipt R1 contains exactly one line item."}, {"id": "a5", "statement": "The sole line item on receipt R1 qualifies for base points apart from the gift-card exclusion."}, {"id": "a6", "statement": "The SKU of the sole line item on receipt R1 is a member of the program's gift-card SKU registry."}, {"id": "a7", "statement": "Every SKU in the program's gift-card SKU registry identifies a gift-card product."}, {"id": "a8", "statement": "Every gift-card product SKU appears in the program's gift-card SKU registry."}, {"id": "a9", "statement": "Transaction T1 occurred during the double-points weekend."}, {"id": "a10", "statement": "Priya Shah's account received zero Cedar Circle points for transaction T1."}], "base_state_json": "\"Priya Shah filed a missing-points claim requesting 200 Cedar Circle points for transaction T1. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. The item otherwise qualifies for base points apart from the gift-card exclusion. T1 occurred during the double-points weekend, but her account received zero Cedar Circle points. At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817. At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-4817, CCG-6392, and CCG-8504. Program administration confirms that registry entries identify gift-card products and that every gift-card product SKU is included in that complete registry. Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817."}, {"path": [], "text": "At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-4817, CCG-6392, and CCG-8504."}], "policy_evidence": [{"path": [], "text": "Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases."}, {"path": [], "text": "The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}], "rules": [{"justification": "The sole item on the claimed $100 transaction has a SKU in the gift-card registry, and every registry SKU identifies a gift-card product. The gift-card exclusion therefore applies, and the promotion cannot override it.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "The exhaustive registry contains every gift-card SKU, while the sole item's SKU is not in that registry, so the gift-card exclusion does not apply. The otherwise base-eligible $100 transaction earns 100 base points and, because it occurred during the double-points weekend, 200 total points; zero were awarded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817.", "negative_left": "At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817.", "negative_right": "At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-2751, CCG-6392, and CCG-8504.", "right": "At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-4817, CCG-6392, and CCG-8504."}, "verifier_independent_model": false}, "family": "scale-diverse-027-001", "id": "scale-diverse-027-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the gift-card exclusion applies, so Priya was not entitled to the claimed points.", "true": "Yes — the records and applicable rules show that Priya should have received the claimed 200 points."}, "instructions": "Decide whether Priya's claim for 200 missing points is valid under the supplied program and promotion rules. Answer yes or no.", "type": "noul"}}, "state": "Priya Shah filed a missing-points claim requesting 200 Cedar Circle points for transaction T1. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. The item otherwise qualifies for base points apart from the gift-card exclusion. T1 occurred during the double-points weekend, but her account received zero Cedar Circle points. At 10:12:04 on 8 August 2026, a complete line-by-line audit of receipt R1 recorded its sole line item's SKU as CCG-4817. At 10:13:29 on 8 August 2026, the program's complete gift-card SKU registry applicable to receipt R1 contained exactly CCG-2751, CCG-6392, and CCG-8504. Program administration confirms that registry entries identify gift-card products and that every gift-card product SKU is included in that complete registry. Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}, "method": "c2d", "provenance": {"source_id": "diverse-027", "source_is_synthetic": true, "source_sha256": "24a57454b60c3c253b20a328c3963efba4da8dd26a29c30742ff6568152211d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same base-points rule, gift-card exclusion, and promotion scope, while retaining Priya Shah, transaction T1, the 200-point claim, and the relevant transaction and registry timing. The two evidence spans are complete factual sentences. The counterfactual coherently changes the complete registry so that QV-4827 is absent; together with the stated registry completeness, this does not conflict with any unchanged factual assertion. Neither context contains an answer code, proposition ID, rule table, output instruction, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the two universal registry relations. The focus atom is the factual relation of SKU membership, not a policy conclusion. The base and counter assignments are jointly realizable and differ only on that focus: registry membership makes the item a gift card in the base case, while nonmembership plus registry exhaustiveness makes it a non-gift-card item in the countercase. The policy evidence correctly preserves the substantive program and promotion rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The sole item’s SKU is in the registry, and every registered SKU identifies a gift card. Therefore the transaction is a gift-card purchase, the stated exclusion applies, and the promotion cannot create points because it only doubles transactions that first qualify for base points. Priya is not entitled to the claimed 200 points.", "rule_index": 0, "sound": true}, {"reason": "Refutation of registry membership means the sole item’s SKU is not registered. Because every gift-card SKU is stated to appear in the registry, the item cannot be a gift card. Atom a5 establishes eligibility apart from that exclusion; the sole $100 item therefore earns 100 base points and 200 during the double-points weekend. With zero points awarded and a claim for 200, the claim is valid.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya Shah's missing-points claim requests 200 Cedar Circle points for transaction T1."}, {"id": "a2", "statement": "Receipt R1 documents transaction T1."}, {"id": "a3", "statement": "The total price shown on receipt R1 is $100."}, {"id": "a4", "statement": "Receipt R1 contains exactly one line item."}, {"id": "a5", "statement": "The sole line item on receipt R1 qualifies for base points apart from the gift-card exclusion."}, {"id": "a6", "statement": "The SKU of the sole line item on receipt R1 is a member of the program's gift-card SKU registry."}, {"id": "a7", "statement": "Every SKU in the program's gift-card SKU registry identifies a gift-card product."}, {"id": "a8", "statement": "Every gift-card product SKU appears in the program's gift-card SKU registry."}, {"id": "a9", "statement": "Transaction T1 occurred during the double-points weekend."}, {"id": "a10", "statement": "Priya Shah's account received zero Cedar Circle points for transaction T1."}], "base_state_json": "\"Reconciliation note: Priya Shah’s missing-points claim requests 200 Cedar Circle points for transaction T1. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. Apart from the gift-card exclusion, that item meets the requirements for base points. T1 occurred during the double-points weekend, but Priya Shah’s account ledger shows zero Cedar Circle points credited for it.\\n\\nAt 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827. The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly QV-4827, LM-6031, and TX-9146. Registry governance records confirm that every SKU in the registry identifies a gift-card product and that every gift-card product SKU appears in the registry.\\n\\nProgram rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "At 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827."}, {"path": [], "text": "The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly QV-4827, LM-6031, and TX-9146."}], "policy_evidence": [{"path": [], "text": "Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases."}, {"path": [], "text": "The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}], "rules": [{"justification": "The sole item on the claimed $100 transaction has a SKU in the gift-card registry, and every registry SKU identifies a gift-card product. The gift-card exclusion therefore applies, and the promotion cannot override it.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "The exhaustive registry contains every gift-card SKU, while the sole item's SKU is not in that registry, so the gift-card exclusion does not apply. The otherwise base-eligible $100 transaction earns 100 base points and, because it occurred during the double-points weekend, 200 total points; zero were awarded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827.", "negative_left": "At 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827.", "negative_right": "The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly BN-2754, LM-6031, and TX-9146.", "right": "The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly QV-4827, LM-6031, and TX-9146."}, "verifier_independent_model": false}, "family": "scale-diverse-027-002", "id": "scale-diverse-027-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the gift-card exclusion applies, so Priya was not entitled to the claimed points.", "true": "Yes — the records and applicable rules show that Priya should have received the claimed 200 points."}, "instructions": "Decide whether Priya's claim for 200 missing points is valid under the supplied program and promotion rules. Answer yes or no.", "type": "noul"}}, "state": "Reconciliation note: Priya Shah’s missing-points claim requests 200 Cedar Circle points for transaction T1. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. Apart from the gift-card exclusion, that item meets the requirements for base points. T1 occurred during the double-points weekend, but Priya Shah’s account ledger shows zero Cedar Circle points credited for it.\n\nAt 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827. The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly QV-4827, LM-6031, and TX-9146. Registry governance records confirm that every SKU in the registry identifies a gift-card product and that every gift-card product SKU appears in the registry.\n\nProgram rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}, "method": "c2d", "provenance": {"source_id": "diverse-027", "source_is_synthetic": true, "source_sha256": "24a57454b60c3c253b20a328c3963efba4da8dd26a29c30742ff6568152211d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same base-points rule, gift-card exclusion, and promotion scope, while retaining Priya Shah, transaction T1, the 200-point claim, and the relevant transaction and registry timing. The two evidence spans are complete factual sentences. The counterfactual coherently changes the complete registry so that QV-4827 is absent; together with the stated registry completeness, this does not conflict with any unchanged factual assertion. Neither context contains an answer code, proposition ID, rule table, output instruction, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the two universal registry relations. The focus atom is the factual relation of SKU membership, not a policy conclusion. The base and counter assignments are jointly realizable and differ only on that focus: registry membership makes the item a gift card in the base case, while nonmembership plus registry exhaustiveness makes it a non-gift-card item in the countercase. The policy evidence correctly preserves the substantive program and promotion rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The sole item’s SKU is in the registry, and every registered SKU identifies a gift card. Therefore the transaction is a gift-card purchase, the stated exclusion applies, and the promotion cannot create points because it only doubles transactions that first qualify for base points. Priya is not entitled to the claimed 200 points.", "rule_index": 0, "sound": true}, {"reason": "Refutation of registry membership means the sole item’s SKU is not registered. Because every gift-card SKU is stated to appear in the registry, the item cannot be a gift card. Atom a5 establishes eligibility apart from that exclusion; the sole $100 item therefore earns 100 base points and 200 during the double-points weekend. With zero points awarded and a claim for 200, the claim is valid.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya Shah's missing-points claim requests 200 Cedar Circle points for transaction T1."}, {"id": "a2", "statement": "Receipt R1 documents transaction T1."}, {"id": "a3", "statement": "The total price shown on receipt R1 is $100."}, {"id": "a4", "statement": "Receipt R1 contains exactly one line item."}, {"id": "a5", "statement": "The sole line item on receipt R1 qualifies for base points apart from the gift-card exclusion."}, {"id": "a6", "statement": "The SKU of the sole line item on receipt R1 is a member of the program's gift-card SKU registry."}, {"id": "a7", "statement": "Every SKU in the program's gift-card SKU registry identifies a gift-card product."}, {"id": "a8", "statement": "Every gift-card product SKU appears in the program's gift-card SKU registry."}, {"id": "a9", "statement": "Transaction T1 occurred during the double-points weekend."}, {"id": "a10", "statement": "Priya Shah's account received zero Cedar Circle points for transaction T1."}], "base_state_json": "\"Reconciliation note: Priya Shah’s missing-points claim requests 200 Cedar Circle points for transaction T1. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. Apart from the gift-card exclusion, that item meets the requirements for base points. T1 occurred during the double-points weekend, but Priya Shah’s account ledger shows zero Cedar Circle points credited for it.\\n\\nAt 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827. The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly QV-4827, LM-6031, and TX-9146. Registry governance records confirm that every SKU in the registry identifies a gift-card product and that every gift-card product SKU appears in the registry.\\n\\nProgram rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "At 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827."}, {"path": [], "text": "The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly QV-4827, LM-6031, and TX-9146."}], "policy_evidence": [{"path": [], "text": "Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases."}, {"path": [], "text": "The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}], "rules": [{"justification": "The sole item on the claimed $100 transaction has a SKU in the gift-card registry, and every registry SKU identifies a gift-card product. The gift-card exclusion therefore applies, and the promotion cannot override it.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "The exhaustive registry contains every gift-card SKU, while the sole item's SKU is not in that registry, so the gift-card exclusion does not apply. The otherwise base-eligible $100 transaction earns 100 base points and, because it occurred during the double-points weekend, 200 total points; zero were awarded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827.", "negative_left": "At 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827.", "negative_right": "The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly BN-2754, LM-6031, and TX-9146.", "right": "The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly QV-4827, LM-6031, and TX-9146."}, "verifier_independent_model": false}, "family": "scale-diverse-027-002", "id": "scale-diverse-027-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the gift-card exclusion applies, so Priya was not entitled to the claimed points.", "true": "Yes — the records and applicable rules show that Priya should have received the claimed 200 points."}, "instructions": "Decide whether Priya's claim for 200 missing points is valid under the supplied program and promotion rules. Answer yes or no.", "type": "noul"}}, "state": "Reconciliation note: Priya Shah’s missing-points claim requests 200 Cedar Circle points for transaction T1. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. Apart from the gift-card exclusion, that item meets the requirements for base points. T1 occurred during the double-points weekend, but Priya Shah’s account ledger shows zero Cedar Circle points credited for it.\n\nAt 09:42 UTC on 18 August 2026, the finalized receipt audit identified the SKU of the sole line item on receipt R1 as QV-4827. The complete program gift-card SKU registry snapshot effective at 09:42 UTC on 18 August 2026 listed exactly BN-2754, LM-6031, and TX-9146. Registry governance records confirm that every SKU in the registry identifies a gift-card product and that every gift-card product SKU appears in the registry.\n\nProgram rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}, "method": "c2d", "provenance": {"source_id": "diverse-027", "source_is_synthetic": true, "source_sha256": "24a57454b60c3c253b20a328c3963efba4da8dd26a29c30742ff6568152211d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original eligibility, gift-card exclusion, and promotion-scope rules while preserving Priya Shah’s 200-point claim, transaction, and double-points-weekend bindings. The two evidence spans are complete factual sentences. Changing only whether SKU GC-5842 appears in the complete gift-card registry is coherent with the unchanged registry specification and creates no duplicate contradictory measurement or assertion within either context. Neither context states the claim decision, supplies an answer code, or instructs the classifier what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the two universal registry relations. The focus atom is the factual relation of SKU membership, not a policy conclusion. The base and counter assignments are jointly realizable and differ only on that focus: registry membership makes the item a gift card in the base case, while nonmembership plus registry exhaustiveness makes it a non-gift-card item in the countercase. The policy evidence correctly preserves the substantive program and promotion rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The sole item’s SKU is in the registry, and every registered SKU identifies a gift card. Therefore the transaction is a gift-card purchase, the stated exclusion applies, and the promotion cannot create points because it only doubles transactions that first qualify for base points. Priya is not entitled to the claimed 200 points.", "rule_index": 0, "sound": true}, {"reason": "Refutation of registry membership means the sole item’s SKU is not registered. Because every gift-card SKU is stated to appear in the registry, the item cannot be a gift card. Atom a5 establishes eligibility apart from that exclusion; the sole $100 item therefore earns 100 base points and 200 during the double-points weekend. With zero points awarded and a claim for 200, the claim is valid.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya Shah's missing-points claim requests 200 Cedar Circle points for transaction T1."}, {"id": "a2", "statement": "Receipt R1 documents transaction T1."}, {"id": "a3", "statement": "The total price shown on receipt R1 is $100."}, {"id": "a4", "statement": "Receipt R1 contains exactly one line item."}, {"id": "a5", "statement": "The sole line item on receipt R1 qualifies for base points apart from the gift-card exclusion."}, {"id": "a6", "statement": "The SKU of the sole line item on receipt R1 is a member of the program's gift-card SKU registry."}, {"id": "a7", "statement": "Every SKU in the program's gift-card SKU registry identifies a gift-card product."}, {"id": "a8", "statement": "Every gift-card product SKU appears in the program's gift-card SKU registry."}, {"id": "a9", "statement": "Transaction T1 occurred during the double-points weekend."}, {"id": "a10", "statement": "Priya Shah's account received zero Cedar Circle points for transaction T1."}], "base_state_json": "\"Operations transferred Priya Shah’s missing-points claim requesting 200 Cedar Circle points for transaction T1 for review. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. That sole item qualifies for base points apart from the gift-card exclusion. During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item. The complete program gift-card SKU registry applicable to receipt R1 contained SKU GC-5842. The registry specification states that every SKU in the program’s gift-card SKU registry identifies a gift-card product and that every gift-card product SKU appears in the registry. T1 occurred during the double-points weekend, while Priya Shah’s account received zero Cedar Circle points for the transaction. Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item."}, {"path": [], "text": "The complete program gift-card SKU registry applicable to receipt R1 contained SKU GC-5842."}], "policy_evidence": [{"path": [], "text": "Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases."}, {"path": [], "text": "The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}], "rules": [{"justification": "The sole item on the claimed $100 transaction has a SKU in the gift-card registry, and every registry SKU identifies a gift-card product. The gift-card exclusion therefore applies, and the promotion cannot override it.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "The exhaustive registry contains every gift-card SKU, while the sole item's SKU is not in that registry, so the gift-card exclusion does not apply. The otherwise base-eligible $100 transaction earns 100 base points and, because it occurred during the double-points weekend, 200 total points; zero were awarded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item.", "negative_left": "During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item.", "negative_right": "The complete program gift-card SKU registry applicable to receipt R1 did not contain SKU GC-5842.", "right": "The complete program gift-card SKU registry applicable to receipt R1 contained SKU GC-5842."}, "verifier_independent_model": false}, "family": "scale-diverse-027-003", "id": "scale-diverse-027-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the gift-card exclusion applies, so Priya was not entitled to the claimed points.", "true": "Yes — the records and applicable rules show that Priya should have received the claimed 200 points."}, "instructions": "Decide whether Priya's claim for 200 missing points is valid under the supplied program and promotion rules. Answer yes or no.", "type": "noul"}}, "state": "Operations transferred Priya Shah’s missing-points claim requesting 200 Cedar Circle points for transaction T1 for review. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. That sole item qualifies for base points apart from the gift-card exclusion. During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item. The complete program gift-card SKU registry applicable to receipt R1 contained SKU GC-5842. The registry specification states that every SKU in the program’s gift-card SKU registry identifies a gift-card product and that every gift-card product SKU appears in the registry. T1 occurred during the double-points weekend, while Priya Shah’s account received zero Cedar Circle points for the transaction. Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}, "method": "c2d", "provenance": {"source_id": "diverse-027", "source_is_synthetic": true, "source_sha256": "24a57454b60c3c253b20a328c3963efba4da8dd26a29c30742ff6568152211d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original eligibility, gift-card exclusion, and promotion-scope rules while preserving Priya Shah’s 200-point claim, transaction, and double-points-weekend bindings. The two evidence spans are complete factual sentences. Changing only whether SKU GC-5842 appears in the complete gift-card registry is coherent with the unchanged registry specification and creates no duplicate contradictory measurement or assertion within either context. Neither context states the claim decision, supplies an answer code, or instructs the classifier what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the two universal registry relations. The focus atom is the factual relation of SKU membership, not a policy conclusion. The base and counter assignments are jointly realizable and differ only on that focus: registry membership makes the item a gift card in the base case, while nonmembership plus registry exhaustiveness makes it a non-gift-card item in the countercase. The policy evidence correctly preserves the substantive program and promotion rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The sole item’s SKU is in the registry, and every registered SKU identifies a gift card. Therefore the transaction is a gift-card purchase, the stated exclusion applies, and the promotion cannot create points because it only doubles transactions that first qualify for base points. Priya is not entitled to the claimed 200 points.", "rule_index": 0, "sound": true}, {"reason": "Refutation of registry membership means the sole item’s SKU is not registered. Because every gift-card SKU is stated to appear in the registry, the item cannot be a gift card. Atom a5 establishes eligibility apart from that exclusion; the sole $100 item therefore earns 100 base points and 200 during the double-points weekend. With zero points awarded and a claim for 200, the claim is valid.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya Shah's missing-points claim requests 200 Cedar Circle points for transaction T1."}, {"id": "a2", "statement": "Receipt R1 documents transaction T1."}, {"id": "a3", "statement": "The total price shown on receipt R1 is $100."}, {"id": "a4", "statement": "Receipt R1 contains exactly one line item."}, {"id": "a5", "statement": "The sole line item on receipt R1 qualifies for base points apart from the gift-card exclusion."}, {"id": "a6", "statement": "The SKU of the sole line item on receipt R1 is a member of the program's gift-card SKU registry."}, {"id": "a7", "statement": "Every SKU in the program's gift-card SKU registry identifies a gift-card product."}, {"id": "a8", "statement": "Every gift-card product SKU appears in the program's gift-card SKU registry."}, {"id": "a9", "statement": "Transaction T1 occurred during the double-points weekend."}, {"id": "a10", "statement": "Priya Shah's account received zero Cedar Circle points for transaction T1."}], "base_state_json": "\"Operations transferred Priya Shah’s missing-points claim requesting 200 Cedar Circle points for transaction T1 for review. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. That sole item qualifies for base points apart from the gift-card exclusion. During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item. The complete program gift-card SKU registry applicable to receipt R1 contained SKU GC-5842. The registry specification states that every SKU in the program’s gift-card SKU registry identifies a gift-card product and that every gift-card product SKU appears in the registry. T1 occurred during the double-points weekend, while Priya Shah’s account received zero Cedar Circle points for the transaction. Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item."}, {"path": [], "text": "The complete program gift-card SKU registry applicable to receipt R1 contained SKU GC-5842."}], "policy_evidence": [{"path": [], "text": "Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases."}, {"path": [], "text": "The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}], "rules": [{"justification": "The sole item on the claimed $100 transaction has a SKU in the gift-card registry, and every registry SKU identifies a gift-card product. The gift-card exclusion therefore applies, and the promotion cannot override it.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "The exhaustive registry contains every gift-card SKU, while the sole item's SKU is not in that registry, so the gift-card exclusion does not apply. The otherwise base-eligible $100 transaction earns 100 base points and, because it occurred during the double-points weekend, 200 total points; zero were awarded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item.", "negative_left": "During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item.", "negative_right": "The complete program gift-card SKU registry applicable to receipt R1 did not contain SKU GC-5842.", "right": "The complete program gift-card SKU registry applicable to receipt R1 contained SKU GC-5842."}, "verifier_independent_model": false}, "family": "scale-diverse-027-003", "id": "scale-diverse-027-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the gift-card exclusion applies, so Priya was not entitled to the claimed points.", "true": "Yes — the records and applicable rules show that Priya should have received the claimed 200 points."}, "instructions": "Decide whether Priya's claim for 200 missing points is valid under the supplied program and promotion rules. Answer yes or no.", "type": "noul"}}, "state": "Operations transferred Priya Shah’s missing-points claim requesting 200 Cedar Circle points for transaction T1 for review. Receipt R1 documents T1, shows a total price of $100, and contains exactly one line item. That sole item qualifies for base points apart from the gift-card exclusion. During the operational handoff for receipt R1, the verified item record identified SKU GC-5842 as the SKU of R1's sole line item. The complete program gift-card SKU registry applicable to receipt R1 did not contain SKU GC-5842. The registry specification states that every SKU in the program’s gift-card SKU registry identifies a gift-card product and that every gift-card product SKU appears in the registry. T1 occurred during the double-points weekend, while Priya Shah’s account received zero Cedar Circle points for the transaction. Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}, "method": "c2d", "provenance": {"source_id": "diverse-027", "source_is_synthetic": true, "source_sha256": "24a57454b60c3c253b20a328c3963efba4da8dd26a29c30742ff6568152211d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rewards-policy rules and preserve the order, account, transaction, session, and timestamp bindings. The two focus spans are complete factual sentences, and the counterfactual coherently changes the authenticated user from the member to Dana Voss without conflicting with the unchanged statement that the session user was the sole approver. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Account custodian\",\"text\":\"At the start of the claim period for order K-184, the member's loyalty account had an opening balance of exactly 2,000 points.\"},{\"speaker\":\"Ledger analyst\",\"text\":\"During the claim period for order K-184, exactly 450 eligible points were posted to the member's loyalty account.\"},{\"speaker\":\"Rewards policy\",\"text\":\"Policy awards one point per eligible merchandise dollar; gift cards earn none.\"},{\"speaker\":\"Transaction auditor\",\"text\":\"During the claim period for order K-184, exactly one reward-redemption transaction was posted to the account: R-700, for exactly 700 points.\"},{\"speaker\":\"Approval log\",\"text\":\"At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41.\"},{\"speaker\":\"Authentication log\",\"text\":\"The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was the member who owned the loyalty account charged by reward-redemption transaction R-700.\"},{\"speaker\":\"Account custodian\",\"text\":\"At the end of the claim period for order K-184, the actual account balance was exactly 1,750 points.\"},{\"speaker\":\"Rewards policy\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["4", "text"], "text": "At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41."}, {"path": ["5", "text"], "text": "The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was the member who owned the loyalty account charged by reward-redemption transaction R-700."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41.", "negative_left": "At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41.", "negative_right": "The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was account-services employee Dana Voss, who was not the member who owned the loyalty account charged by reward-redemption transaction R-700.", "right": "The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was the member who owned the loyalty account charged by reward-redemption transaction R-700."}, "verifier_independent_model": false}, "family": "scale-diverse-029-001", "id": "scale-diverse-029-001-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Account custodian", "text": "At the start of the claim period for order K-184, the member's loyalty account had an opening balance of exactly 2,000 points."}, {"speaker": "Ledger analyst", "text": "During the claim period for order K-184, exactly 450 eligible points were posted to the member's loyalty account."}, {"speaker": "Rewards policy", "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"speaker": "Transaction auditor", "text": "During the claim period for order K-184, exactly one reward-redemption transaction was posted to the account: R-700, for exactly 700 points."}, {"speaker": "Approval log", "text": "At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41."}, {"speaker": "Authentication log", "text": "The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was the member who owned the loyalty account charged by reward-redemption transaction R-700."}, {"speaker": "Account custodian", "text": "At the end of the claim period for order K-184, the actual account balance was exactly 1,750 points."}, {"speaker": "Rewards policy", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rewards-policy rules and preserve the order, account, transaction, session, and timestamp bindings. The two focus spans are complete factual sentences, and the counterfactual coherently changes the authenticated user from the member to Dana Voss without conflicting with the unchanged statement that the session user was the sole approver. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Account custodian\",\"text\":\"At the start of the claim period for order K-184, the member's loyalty account had an opening balance of exactly 2,000 points.\"},{\"speaker\":\"Ledger analyst\",\"text\":\"During the claim period for order K-184, exactly 450 eligible points were posted to the member's loyalty account.\"},{\"speaker\":\"Rewards policy\",\"text\":\"Policy awards one point per eligible merchandise dollar; gift cards earn none.\"},{\"speaker\":\"Transaction auditor\",\"text\":\"During the claim period for order K-184, exactly one reward-redemption transaction was posted to the account: R-700, for exactly 700 points.\"},{\"speaker\":\"Approval log\",\"text\":\"At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41.\"},{\"speaker\":\"Authentication log\",\"text\":\"The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was the member who owned the loyalty account charged by reward-redemption transaction R-700.\"},{\"speaker\":\"Account custodian\",\"text\":\"At the end of the claim period for order K-184, the actual account balance was exactly 1,750 points.\"},{\"speaker\":\"Rewards policy\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["4", "text"], "text": "At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41."}, {"path": ["5", "text"], "text": "The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was the member who owned the loyalty account charged by reward-redemption transaction R-700."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41.", "negative_left": "At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41.", "negative_right": "The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was account-services employee Dana Voss, who was not the member who owned the loyalty account charged by reward-redemption transaction R-700.", "right": "The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was the member who owned the loyalty account charged by reward-redemption transaction R-700."}, "verifier_independent_model": false}, "family": "scale-diverse-029-001", "id": "scale-diverse-029-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Account custodian", "text": "At the start of the claim period for order K-184, the member's loyalty account had an opening balance of exactly 2,000 points."}, {"speaker": "Ledger analyst", "text": "During the claim period for order K-184, exactly 450 eligible points were posted to the member's loyalty account."}, {"speaker": "Rewards policy", "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"speaker": "Transaction auditor", "text": "During the claim period for order K-184, exactly one reward-redemption transaction was posted to the account: R-700, for exactly 700 points."}, {"speaker": "Approval log", "text": "At 14:26:08 UTC on 9 May 2026, the sole person who knowingly approved the posting of reward-redemption transaction R-700 against the member's loyalty account was the user of terminal session QX-41."}, {"speaker": "Authentication log", "text": "The authenticated user of terminal session QX-41 at 14:26:08 UTC on 9 May 2026 was account-services employee Dana Voss, who was not the member who owned the loyalty account charged by reward-redemption transaction R-700."}, {"speaker": "Account custodian", "text": "At the end of the claim period for order K-184, the actual account balance was exactly 1,750 points."}, {"speaker": "Rewards policy", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same balance-calculation policy, earning exclusions, order K-184, claim period, transaction R-700, credential, and approval timestamp. The two focus spans are complete factual sentences. The counterfactual changes only the factual identity of Q-391's holder at the relevant time and does not conflict with another assertion in that context. Neither context states a magnitude level, computed discrepancy, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Ledger auditor\",\"text\":\"For the claim period associated with order K-184, the loyalty account opened with exactly 2,000 points and ended with an actual balance of exactly 1,750 points.\"},{\"speaker\":\"Reconciliation analyst\",\"text\":\"Eligible earnings posted during the period total exactly 450 points. The ledger contains exactly one posted reward-redemption transaction for the period: R-700, with an amount of exactly 700 points.\"},{\"speaker\":\"Approval-record auditor\",\"text\":\"The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026.\"},{\"speaker\":\"Credential custodian\",\"text\":\"At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by the member.\"},{\"speaker\":\"Rewards policy specialist\",\"text\":\"Policy awards one point per eligible merchandise dollar; gift cards earn none.\"},{\"speaker\":\"Rewards policy specialist\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["2", "text"], "text": "The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026."}, {"path": ["3", "text"], "text": "At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by the member."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026.", "negative_left": "The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026.", "negative_right": "At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by Nia Voss, who was not the member.", "right": "At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by the member."}, "verifier_independent_model": false}, "family": "scale-diverse-029-002", "id": "scale-diverse-029-002-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Ledger auditor", "text": "For the claim period associated with order K-184, the loyalty account opened with exactly 2,000 points and ended with an actual balance of exactly 1,750 points."}, {"speaker": "Reconciliation analyst", "text": "Eligible earnings posted during the period total exactly 450 points. The ledger contains exactly one posted reward-redemption transaction for the period: R-700, with an amount of exactly 700 points."}, {"speaker": "Approval-record auditor", "text": "The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026."}, {"speaker": "Credential custodian", "text": "At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by the member."}, {"speaker": "Rewards policy specialist", "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"speaker": "Rewards policy specialist", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same balance-calculation policy, earning exclusions, order K-184, claim period, transaction R-700, credential, and approval timestamp. The two focus spans are complete factual sentences. The counterfactual changes only the factual identity of Q-391's holder at the relevant time and does not conflict with another assertion in that context. Neither context states a magnitude level, computed discrepancy, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Ledger auditor\",\"text\":\"For the claim period associated with order K-184, the loyalty account opened with exactly 2,000 points and ended with an actual balance of exactly 1,750 points.\"},{\"speaker\":\"Reconciliation analyst\",\"text\":\"Eligible earnings posted during the period total exactly 450 points. The ledger contains exactly one posted reward-redemption transaction for the period: R-700, with an amount of exactly 700 points.\"},{\"speaker\":\"Approval-record auditor\",\"text\":\"The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026.\"},{\"speaker\":\"Credential custodian\",\"text\":\"At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by the member.\"},{\"speaker\":\"Rewards policy specialist\",\"text\":\"Policy awards one point per eligible merchandise dollar; gift cards earn none.\"},{\"speaker\":\"Rewards policy specialist\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["2", "text"], "text": "The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026."}, {"path": ["3", "text"], "text": "At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by the member."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026.", "negative_left": "The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026.", "negative_right": "At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by Nia Voss, who was not the member.", "right": "At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by the member."}, "verifier_independent_model": false}, "family": "scale-diverse-029-002", "id": "scale-diverse-029-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Ledger auditor", "text": "For the claim period associated with order K-184, the loyalty account opened with exactly 2,000 points and ended with an actual balance of exactly 1,750 points."}, {"speaker": "Reconciliation analyst", "text": "Eligible earnings posted during the period total exactly 450 points. The ledger contains exactly one posted reward-redemption transaction for the period: R-700, with an amount of exactly 700 points."}, {"speaker": "Approval-record auditor", "text": "The complete approval record for reward-redemption transaction R-700 shows that its sole approving person was the holder of credential Q-391 at 14:32:18 UTC on 8 April 2026."}, {"speaker": "Credential custodian", "text": "At 14:32:18 UTC on 8 April 2026, credential Q-391 was held exclusively by Nia Voss, who was not the member."}, {"speaker": "Rewards policy specialist", "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"speaker": "Rewards policy specialist", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the computation rule and scoring criteria, while both contexts retain the original governing earning, gift-card, tier-status, and pending-promotion policies. The order, account, claim-period, transaction, token, and as-of-time bindings remain consistent. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the authorization declaration and remains coherent with the posted redemption and balances; an unauthorized posting can still affect the actual balance. Neither context contains a score, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Account auditor\",\"text\":\"For order K-184, the loyalty account opened the claim period with exactly 2,000 points and ended it with an actual balance of exactly 1,750 points. The ledger records exactly 450 points as the total eligible earnings posted during that period.\"},{\"speaker\":\"Ledger custodian\",\"text\":\"During the same claim period, exactly one reward-redemption transaction was posted to the account: R-700, with a point amount of exactly 700 points.\"},{\"speaker\":\"Decision-record custodian\",\"text\":\"As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I authorize the reward redemption bearing this token.”\"},{\"speaker\":\"Posting-record custodian\",\"text\":\"The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token.\"},{\"speaker\":\"Program policy\",\"text\":\"Policy awards one point per eligible merchandise dollar; gift cards earn none. Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["2", "text"], "text": "As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I authorize the reward redemption bearing this token.”"}, {"path": ["3", "text"], "text": "The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I authorize the reward redemption bearing this token.”", "negative_left": "As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I do not authorize the reward redemption bearing this token.”", "negative_right": "The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token.", "right": "The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token."}, "verifier_independent_model": false}, "family": "scale-diverse-029-004", "id": "scale-diverse-029-004-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Account auditor", "text": "For order K-184, the loyalty account opened the claim period with exactly 2,000 points and ended it with an actual balance of exactly 1,750 points. The ledger records exactly 450 points as the total eligible earnings posted during that period."}, {"speaker": "Ledger custodian", "text": "During the same claim period, exactly one reward-redemption transaction was posted to the account: R-700, with a point amount of exactly 700 points."}, {"speaker": "Decision-record custodian", "text": "As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I authorize the reward redemption bearing this token.”"}, {"speaker": "Posting-record custodian", "text": "The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token."}, {"speaker": "Program policy", "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none. Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the computation rule and scoring criteria, while both contexts retain the original governing earning, gift-card, tier-status, and pending-promotion policies. The order, account, claim-period, transaction, token, and as-of-time bindings remain consistent. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the authorization declaration and remains coherent with the posted redemption and balances; an unauthorized posting can still affect the actual balance. Neither context contains a score, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Account auditor\",\"text\":\"For order K-184, the loyalty account opened the claim period with exactly 2,000 points and ended it with an actual balance of exactly 1,750 points. The ledger records exactly 450 points as the total eligible earnings posted during that period.\"},{\"speaker\":\"Ledger custodian\",\"text\":\"During the same claim period, exactly one reward-redemption transaction was posted to the account: R-700, with a point amount of exactly 700 points.\"},{\"speaker\":\"Decision-record custodian\",\"text\":\"As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I authorize the reward redemption bearing this token.”\"},{\"speaker\":\"Posting-record custodian\",\"text\":\"The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token.\"},{\"speaker\":\"Program policy\",\"text\":\"Policy awards one point per eligible merchandise dollar; gift cards earn none. Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["2", "text"], "text": "As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I authorize the reward redemption bearing this token.”"}, {"path": ["3", "text"], "text": "The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I authorize the reward redemption bearing this token.”", "negative_left": "As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I do not authorize the reward redemption bearing this token.”", "negative_right": "The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token.", "right": "The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token."}, "verifier_independent_model": false}, "family": "scale-diverse-029-004", "id": "scale-diverse-029-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Account auditor", "text": "For order K-184, the loyalty account opened the claim period with exactly 2,000 points and ended it with an actual balance of exactly 1,750 points. The ledger records exactly 450 points as the total eligible earnings posted during that period."}, {"speaker": "Ledger custodian", "text": "During the same claim period, exactly one reward-redemption transaction was posted to the account: R-700, with a point amount of exactly 700 points."}, {"speaker": "Decision-record custodian", "text": "As of 2026-09-17T18:00:00Z, the member's complete decision record for confirmation token C-9416 contained the sole signed declaration, “I do not authorize the reward redemption bearing this token.”"}, {"speaker": "Posting-record custodian", "text": "The posting record for reward-redemption transaction R-700 identifies C-9416 as its unique confirmation token."}, {"speaker": "Program policy", "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none. Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 200-points-per-$1 redemption policy and keep the inquiry bound to the loyalty member’s June 8 order and its redemption deduction. The focus evidence consists of exactly two complete factual sentences. Changing the ledger deduction from 8,037 to 9,037 points is coherent: against the unchanged 7,400-point requirement, both deductions remain above the requirement by at least one and no more than 5,000 points, so neither context conflicts with its integrity-check assertion. Neither context includes an answer code, category label, rule table, proposition identifier, output instruction, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards analyst reconciled the final worksheet against the final redemption ledger for the loyalty member’s June 8 order. The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order. The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 8,037 points for the reward applied to the loyalty member’s June 8 order. A separate integrity check found that the recorded deduction exceeded the policy-required amount by at least one point and remained no more than 5,000 points above it. No unrelated adjustments, reversals, or duplicate reward entries were posted to the account during the reviewed period. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order."}, {"path": [], "text": "The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 8,037 points for the reward applied to the loyalty member’s June 8 order."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order.", "negative_left": "The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order.", "negative_right": "The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 9,037 points for the reward applied to the loyalty member’s June 8 order.", "right": "The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 8,037 points for the reward applied to the loyalty member’s June 8 order."}, "verifier_independent_model": false}, "family": "scale-diverse-030-002", "id": "scale-diverse-030-002-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards analyst reconciled the final worksheet against the final redemption ledger for the loyalty member’s June 8 order. The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order. The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 8,037 points for the reward applied to the loyalty member’s June 8 order. A separate integrity check found that the recorded deduction exceeded the policy-required amount by at least one point and remained no more than 5,000 points above it. No unrelated adjustments, reversals, or duplicate reward entries were posted to the account during the reviewed period. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 200-points-per-$1 redemption policy and keep the inquiry bound to the loyalty member’s June 8 order and its redemption deduction. The focus evidence consists of exactly two complete factual sentences. Changing the ledger deduction from 8,037 to 9,037 points is coherent: against the unchanged 7,400-point requirement, both deductions remain above the requirement by at least one and no more than 5,000 points, so neither context conflicts with its integrity-check assertion. Neither context includes an answer code, category label, rule table, proposition identifier, output instruction, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards analyst reconciled the final worksheet against the final redemption ledger for the loyalty member’s June 8 order. The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order. The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 8,037 points for the reward applied to the loyalty member’s June 8 order. A separate integrity check found that the recorded deduction exceeded the policy-required amount by at least one point and remained no more than 5,000 points above it. No unrelated adjustments, reversals, or duplicate reward entries were posted to the account during the reviewed period. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order."}, {"path": [], "text": "The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 8,037 points for the reward applied to the loyalty member’s June 8 order."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order.", "negative_left": "The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order.", "negative_right": "The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 9,037 points for the reward applied to the loyalty member’s June 8 order.", "right": "The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 8,037 points for the reward applied to the loyalty member’s June 8 order."}, "verifier_independent_model": false}, "family": "scale-diverse-030-002", "id": "scale-diverse-030-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards analyst reconciled the final worksheet against the final redemption ledger for the loyalty member’s June 8 order. The reconciliation worksheet finalized at 09:14 UTC on June 9, 2026 records 7,400 points as the amount required by the supplied program policy for the reward applied to the loyalty member’s June 8 order. The redemption ledger finalized at 09:16 UTC on June 9, 2026 records a deduction of 9,037 points for the reward applied to the loyalty member’s June 8 order. A separate integrity check found that the recorded deduction exceeded the policy-required amount by at least one point and remained no more than 5,000 points above it. No unrelated adjustments, reversals, or duplicate reward entries were posted to the account during the reviewed period. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same completeness-and-routing scope and do not alter or supplement the governing readiness policy. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only M4’s timestamp from absent to present, without conflicting with unchanged facts. Neither context includes a readiness level, prescribed output, rule table, proposition identifier, or label rationale; terms such as “appropriate owner” and “completeness check” are permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the aggregate missing-field-count propositions. A3 is factual rather than a policy classification. The base and counter assignments are realizable: the base can have exactly one missing non-owner field, while the counter can have none; A1 and A2 remain supported in both. Empty policy_evidence is correct because all governing readiness rules are already retained in the questions object, and no substantive rule from the original state is needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 establishes appropriate ownership for every unresolved item, and A3 establishes exactly one missing required field instance. These conditions are sufficient for Level 3. A2 is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "A2 entails that the nonnegative integer count of missing required field instances is zero or one, while refuted A3 entails that it is not exactly one. Therefore the count is zero. Together with A1, this is sufficient for Level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every item marked unresolved in the 18:00 handoff has an appropriate owner."}, {"id": "A2", "statement": "Across all items marked unresolved in the 18:00 handoff, the total number of missing required field instances—owner, current status, update timestamp, or record link—is at most one."}, {"id": "A3", "statement": "Across all items marked unresolved in the 18:00 handoff, the total number of missing required field instances—owner, current status, update timestamp, or record link—is exactly one."}], "base_state_json": "[{\"speaker\":\"Duty roster clerk\",\"text\":\"At 17:35, the duty roster assigned Inez Cole to validation work for Q7 and Dario Venn to vendor follow-up for M4; both assignments remained active through the handoff review.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \\\"awaiting validation,\\\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \\\"vendor response pending,\\\" and the record link https://records.example/M4, but contained no update timestamp.\"},{\"speaker\":\"Review observer\",\"text\":\"At 18:06, the reviewer used the saved handoff snapshot for the completeness check and recorded that the roster assignments matched the work represented by the two entries.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \"awaiting validation,\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7."}, {"path": ["2", "text"], "text": "M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \"vendor response pending,\" and the record link https://records.example/M4, but contained no update timestamp."}], "policy_evidence": [], "rules": [{"justification": "All unresolved items have appropriate owners, and exactly one required field instance is missing.", "target": "level_3_one_gap", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "All unresolved items have appropriate owners. The missing-field count is at most one but is not exactly one, so the nonnegative count is zero.", "target": "level_4_complete", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \"awaiting validation,\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7.", "negative_left": "The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \"awaiting validation,\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7.", "negative_right": "M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \"vendor response pending,\" the update timestamp 2026-09-17T17:46:00Z, and the record link https://records.example/M4.", "right": "M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \"vendor response pending,\" and the record link https://records.example/M4, but contained no update timestamp."}, "verifier_independent_model": false}, "family": "scale-diverse-031-001", "id": "scale-diverse-031-001-base", "input": {"questions": {"decision": {"criteria": {"level_1_unready": "Level 1 — An unresolved item has no appropriate owner; stop the handoff and assign ownership.", "level_2_major_gaps": "Level 2 — All unresolved items have owners, but two or more required evidence fields are missing; return for substantial revision.", "level_3_one_gap": "Level 3 — All unresolved items have appropriate owners, but exactly one required evidence field is missing; keep the routing and return for the missing evidence.", "level_4_complete": "Level 4 — Every unresolved item has an appropriate owner, current status, timestamp, and record link; accept the handoff as complete."}, "instructions": "Evaluate completeness and routing. Every unresolved item requires an appropriate owner, current status, update timestamp, and record link. Use the ordered readiness scale: Level 1 if any unresolved item lacks an owner; Level 2 if owners are assigned but at least two required evidence fields are missing; Level 3 if owners are assigned and exactly one required evidence field is missing; Level 4 if no required evidence is missing. Select the matching readiness level and action.", "type": "choice"}}, "state": [{"speaker": "Duty roster clerk", "text": "At 17:35, the duty roster assigned Inez Cole to validation work for Q7 and Dario Venn to vendor follow-up for M4; both assignments remained active through the handoff review."}, {"speaker": "Handoff recorder", "text": "The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \"awaiting validation,\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7."}, {"speaker": "Handoff recorder", "text": "M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \"vendor response pending,\" and the record link https://records.example/M4, but contained no update timestamp."}, {"speaker": "Review observer", "text": "At 18:06, the reviewer used the saved handoff snapshot for the completeness check and recorded that the roster assignments matched the work represented by the two entries."}]}, "method": "c2d", "provenance": {"source_id": "diverse-031", "source_is_synthetic": true, "source_sha256": "40a629b3bc5f233c285c76efbab5e221ba53802d4b3e1398846cfed366ebeec2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_3_one_gap"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same completeness-and-routing scope and do not alter or supplement the governing readiness policy. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only M4’s timestamp from absent to present, without conflicting with unchanged facts. Neither context includes a readiness level, prescribed output, rule table, proposition identifier, or label rationale; terms such as “appropriate owner” and “completeness check” are permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the aggregate missing-field-count propositions. A3 is factual rather than a policy classification. The base and counter assignments are realizable: the base can have exactly one missing non-owner field, while the counter can have none; A1 and A2 remain supported in both. Empty policy_evidence is correct because all governing readiness rules are already retained in the questions object, and no substantive rule from the original state is needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 establishes appropriate ownership for every unresolved item, and A3 establishes exactly one missing required field instance. These conditions are sufficient for Level 3. A2 is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "A2 entails that the nonnegative integer count of missing required field instances is zero or one, while refuted A3 entails that it is not exactly one. Therefore the count is zero. Together with A1, this is sufficient for Level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every item marked unresolved in the 18:00 handoff has an appropriate owner."}, {"id": "A2", "statement": "Across all items marked unresolved in the 18:00 handoff, the total number of missing required field instances—owner, current status, update timestamp, or record link—is at most one."}, {"id": "A3", "statement": "Across all items marked unresolved in the 18:00 handoff, the total number of missing required field instances—owner, current status, update timestamp, or record link—is exactly one."}], "base_state_json": "[{\"speaker\":\"Duty roster clerk\",\"text\":\"At 17:35, the duty roster assigned Inez Cole to validation work for Q7 and Dario Venn to vendor follow-up for M4; both assignments remained active through the handoff review.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \\\"awaiting validation,\\\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \\\"vendor response pending,\\\" and the record link https://records.example/M4, but contained no update timestamp.\"},{\"speaker\":\"Review observer\",\"text\":\"At 18:06, the reviewer used the saved handoff snapshot for the completeness check and recorded that the roster assignments matched the work represented by the two entries.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \"awaiting validation,\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7."}, {"path": ["2", "text"], "text": "M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \"vendor response pending,\" and the record link https://records.example/M4, but contained no update timestamp."}], "policy_evidence": [], "rules": [{"justification": "All unresolved items have appropriate owners, and exactly one required field instance is missing.", "target": "level_3_one_gap", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "All unresolved items have appropriate owners. The missing-field count is at most one but is not exactly one, so the nonnegative count is zero.", "target": "level_4_complete", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \"awaiting validation,\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7.", "negative_left": "The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \"awaiting validation,\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7.", "negative_right": "M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \"vendor response pending,\" the update timestamp 2026-09-17T17:46:00Z, and the record link https://records.example/M4.", "right": "M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \"vendor response pending,\" and the record link https://records.example/M4, but contained no update timestamp."}, "verifier_independent_model": false}, "family": "scale-diverse-031-001", "id": "scale-diverse-031-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_unready": "Level 1 — An unresolved item has no appropriate owner; stop the handoff and assign ownership.", "level_2_major_gaps": "Level 2 — All unresolved items have owners, but two or more required evidence fields are missing; return for substantial revision.", "level_3_one_gap": "Level 3 — All unresolved items have appropriate owners, but exactly one required evidence field is missing; keep the routing and return for the missing evidence.", "level_4_complete": "Level 4 — Every unresolved item has an appropriate owner, current status, timestamp, and record link; accept the handoff as complete."}, "instructions": "Evaluate completeness and routing. Every unresolved item requires an appropriate owner, current status, update timestamp, and record link. Use the ordered readiness scale: Level 1 if any unresolved item lacks an owner; Level 2 if owners are assigned but at least two required evidence fields are missing; Level 3 if owners are assigned and exactly one required evidence field is missing; Level 4 if no required evidence is missing. Select the matching readiness level and action.", "type": "choice"}}, "state": [{"speaker": "Duty roster clerk", "text": "At 17:35, the duty roster assigned Inez Cole to validation work for Q7 and Dario Venn to vendor follow-up for M4; both assignments remained active through the handoff review."}, {"speaker": "Handoff recorder", "text": "The 18:00 handoff on 2026-09-17 marked only items Q7 and M4 unresolved, and Q7's entry contained the appropriate owner Inez Cole, the current status \"awaiting validation,\" the update timestamp 2026-09-17T17:41:00Z, and the record link https://records.example/Q7."}, {"speaker": "Handoff recorder", "text": "M4's entry in the 18:00 handoff on 2026-09-17 contained the appropriate owner Dario Venn, the current status \"vendor response pending,\" the update timestamp 2026-09-17T17:46:00Z, and the record link https://records.example/M4."}, {"speaker": "Review observer", "text": "At 18:06, the reviewer used the saved handoff snapshot for the completeness check and recorded that the roster assignments matched the work represented by the two entries."}]}, "method": "c2d", "provenance": {"source_id": "diverse-031", "source_is_synthetic": true, "source_sha256": "40a629b3bc5f233c285c76efbab5e221ba53802d4b3e1398846cfed366ebeec2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_4_complete"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the full governing readiness policy, criteria, and request scope in both inputs. Both contexts remain bound to the 18:00 handoff review and its unresolved items, while permissible case observations and identifiers change. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes QN-83’s record-link field from blank to populated without conflicting with any unchanged assertion. Neither context contains a readiness level, answer code, rule table, proposition identifier, output instruction, or explicit label rationale; terms such as “appropriate owners” are permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the aggregate missing-field-count propositions. A3 is factual rather than a policy classification. The base and counter assignments are realizable: the base can have exactly one missing non-owner field, while the counter can have none; A1 and A2 remain supported in both. Empty policy_evidence is correct because all governing readiness rules are already retained in the questions object, and no substantive rule from the original state is needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 establishes appropriate ownership for every unresolved item, and A3 establishes exactly one missing required field instance. These conditions are sufficient for Level 3. A2 is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "A2 entails that the nonnegative integer count of missing required field instances is zero or one, while refuted A3 entails that it is not exactly one. Therefore the count is zero. Together with A1, this is sufficient for Level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every item marked unresolved in the 18:00 handoff has an appropriate owner."}, {"id": "A2", "statement": "Across all items marked unresolved in the 18:00 handoff, the total number of missing required field instances—owner, current status, update timestamp, or record link—is at most one."}, {"id": "A3", "statement": "Across all items marked unresolved in the 18:00 handoff, the total number of missing required field instances—owner, current status, update timestamp, or record link—is exactly one."}], "base_state_json": "[{\"speaker\":\"Evidence reviewer\",\"text\":\"The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin.\"},{\"speaker\":\"Evidence reviewer\",\"text\":\"In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending” and update timestamp 2026-09-16T17:51:00Z but has a blank record-link field.\"},{\"speaker\":\"Archive custodian\",\"text\":\"The review used the locked handoff export issued at 18:00 UTC. Later ticket edits, amendments, and replacement records were excluded, and the field entries were read directly from that archived snapshot.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin."}, {"path": ["1", "text"], "text": "In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending” and update timestamp 2026-09-16T17:51:00Z but has a blank record-link field."}], "policy_evidence": [], "rules": [{"justification": "All unresolved items have appropriate owners, and exactly one required field instance is missing.", "target": "level_3_one_gap", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "All unresolved items have appropriate owners. The missing-field count is at most one but is not exactly one, so the nonnegative count is zero.", "target": "level_4_complete", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin.", "negative_left": "The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin.", "negative_right": "In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending,” update timestamp 2026-09-16T17:51:00Z, and record link https://records.example/QN-83.", "right": "In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending” and update timestamp 2026-09-16T17:51:00Z but has a blank record-link field."}, "verifier_independent_model": false}, "family": "scale-diverse-031-004", "id": "scale-diverse-031-004-base", "input": {"questions": {"decision": {"criteria": {"level_1_unready": "Level 1 — An unresolved item has no appropriate owner; stop the handoff and assign ownership.", "level_2_major_gaps": "Level 2 — All unresolved items have owners, but two or more required evidence fields are missing; return for substantial revision.", "level_3_one_gap": "Level 3 — All unresolved items have appropriate owners, but exactly one required evidence field is missing; keep the routing and return for the missing evidence.", "level_4_complete": "Level 4 — Every unresolved item has an appropriate owner, current status, timestamp, and record link; accept the handoff as complete."}, "instructions": "Evaluate completeness and routing. Every unresolved item requires an appropriate owner, current status, update timestamp, and record link. Use the ordered readiness scale: Level 1 if any unresolved item lacks an owner; Level 2 if owners are assigned but at least two required evidence fields are missing; Level 3 if owners are assigned and exactly one required evidence field is missing; Level 4 if no required evidence is missing. Select the matching readiness level and action.", "type": "choice"}}, "state": [{"speaker": "Evidence reviewer", "text": "The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin."}, {"speaker": "Evidence reviewer", "text": "In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending” and update timestamp 2026-09-16T17:51:00Z but has a blank record-link field."}, {"speaker": "Archive custodian", "text": "The review used the locked handoff export issued at 18:00 UTC. Later ticket edits, amendments, and replacement records were excluded, and the field entries were read directly from that archived snapshot."}]}, "method": "c2d", "provenance": {"source_id": "diverse-031", "source_is_synthetic": true, "source_sha256": "40a629b3bc5f233c285c76efbab5e221ba53802d4b3e1398846cfed366ebeec2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_3_one_gap"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the full governing readiness policy, criteria, and request scope in both inputs. Both contexts remain bound to the 18:00 handoff review and its unresolved items, while permissible case observations and identifiers change. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes QN-83’s record-link field from blank to populated without conflicting with any unchanged assertion. Neither context contains a readiness level, answer code, rule table, proposition identifier, output instruction, or explicit label rationale; terms such as “appropriate owners” are permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the aggregate missing-field-count propositions. A3 is factual rather than a policy classification. The base and counter assignments are realizable: the base can have exactly one missing non-owner field, while the counter can have none; A1 and A2 remain supported in both. Empty policy_evidence is correct because all governing readiness rules are already retained in the questions object, and no substantive rule from the original state is needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 establishes appropriate ownership for every unresolved item, and A3 establishes exactly one missing required field instance. These conditions are sufficient for Level 3. A2 is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "A2 entails that the nonnegative integer count of missing required field instances is zero or one, while refuted A3 entails that it is not exactly one. Therefore the count is zero. Together with A1, this is sufficient for Level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every item marked unresolved in the 18:00 handoff has an appropriate owner."}, {"id": "A2", "statement": "Across all items marked unresolved in the 18:00 handoff, the total number of missing required field instances—owner, current status, update timestamp, or record link—is at most one."}, {"id": "A3", "statement": "Across all items marked unresolved in the 18:00 handoff, the total number of missing required field instances—owner, current status, update timestamp, or record link—is exactly one."}], "base_state_json": "[{\"speaker\":\"Evidence reviewer\",\"text\":\"The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin.\"},{\"speaker\":\"Evidence reviewer\",\"text\":\"In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending” and update timestamp 2026-09-16T17:51:00Z but has a blank record-link field.\"},{\"speaker\":\"Archive custodian\",\"text\":\"The review used the locked handoff export issued at 18:00 UTC. Later ticket edits, amendments, and replacement records were excluded, and the field entries were read directly from that archived snapshot.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin."}, {"path": ["1", "text"], "text": "In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending” and update timestamp 2026-09-16T17:51:00Z but has a blank record-link field."}], "policy_evidence": [], "rules": [{"justification": "All unresolved items have appropriate owners, and exactly one required field instance is missing.", "target": "level_3_one_gap", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "All unresolved items have appropriate owners. The missing-field count is at most one but is not exactly one, so the nonnegative count is zero.", "target": "level_4_complete", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin.", "negative_left": "The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin.", "negative_right": "In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending,” update timestamp 2026-09-16T17:51:00Z, and record link https://records.example/QN-83.", "right": "In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending” and update timestamp 2026-09-16T17:51:00Z but has a blank record-link field."}, "verifier_independent_model": false}, "family": "scale-diverse-031-004", "id": "scale-diverse-031-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_unready": "Level 1 — An unresolved item has no appropriate owner; stop the handoff and assign ownership.", "level_2_major_gaps": "Level 2 — All unresolved items have owners, but two or more required evidence fields are missing; return for substantial revision.", "level_3_one_gap": "Level 3 — All unresolved items have appropriate owners, but exactly one required evidence field is missing; keep the routing and return for the missing evidence.", "level_4_complete": "Level 4 — Every unresolved item has an appropriate owner, current status, timestamp, and record link; accept the handoff as complete."}, "instructions": "Evaluate completeness and routing. Every unresolved item requires an appropriate owner, current status, update timestamp, and record link. Use the ordered readiness scale: Level 1 if any unresolved item lacks an owner; Level 2 if owners are assigned but at least two required evidence fields are missing; Level 3 if owners are assigned and exactly one required evidence field is missing; Level 4 if no required evidence is missing. Select the matching readiness level and action.", "type": "choice"}}, "state": [{"speaker": "Evidence reviewer", "text": "The complete set of items marked unresolved in the 18:00 UTC handoff on 2026-09-16 is QN-47 and QN-83, whose owner fields respectively record appropriate owners Mira Sol and Pavel Ilyin."}, {"speaker": "Evidence reviewer", "text": "In that handoff, QN-47 records current status “awaiting validation,” update timestamp 2026-09-16T17:42:00Z, and record link https://records.example/QN-47, while QN-83 records current status “vendor response pending,” update timestamp 2026-09-16T17:51:00Z, and record link https://records.example/QN-83."}, {"speaker": "Archive custodian", "text": "The review used the locked handoff export issued at 18:00 UTC. Later ticket edits, amendments, and replacement records were excluded, and the field entries were read directly from that archived snapshot."}]}, "method": "c2d", "provenance": {"source_id": "diverse-031", "source_is_synthetic": true, "source_sha256": "40a629b3bc5f233c285c76efbab5e221ba53802d4b3e1398846cfed366ebeec2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_4_complete"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing reconciliation policy, including Inez’s required role and the absence of record precedence, while preserving the S-2, 19:00 handoff, and readiness-decision bindings. The two evidence spans are complete factual sentences, and the counterfactual changes only the exception register’s scope assertion, coherently creating a conflict without duplicating contradictory assertions from the same source. Neither context states a result, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff. At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff. No record other than the master scope record and the exception register governed whether S-2 applied to this handoff. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. Inez was unavailable and did not perform any reconciliation before the readiness rating. At 19:00, S-2 had an unresolved alarm routed to the controls engineer, but the appropriate owner had not accepted it. No temporary evidence exception covered that missing acceptance. By the handoff, the applicable scope of every item other than S-2 had been established.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as out of scope for the 19:00 handoff.", "right": "At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-032-001", "id": "scale-diverse-032-001-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff. At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff. No record other than the master scope record and the exception register governed whether S-2 applied to this handoff. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. Inez was unavailable and did not perform any reconciliation before the readiness rating. At 19:00, S-2 had an unresolved alarm routed to the controls engineer, but the appropriate owner had not accepted it. No temporary evidence exception covered that missing acceptance. By the handoff, the applicable scope of every item other than S-2 had been established."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing reconciliation policy, including Inez’s required role and the absence of record precedence, while preserving the S-2, 19:00 handoff, and readiness-decision bindings. The two evidence spans are complete factual sentences, and the counterfactual changes only the exception register’s scope assertion, coherently creating a conflict without duplicating contradictory assertions from the same source. Neither context states a result, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff. At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff. No record other than the master scope record and the exception register governed whether S-2 applied to this handoff. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. Inez was unavailable and did not perform any reconciliation before the readiness rating. At 19:00, S-2 had an unresolved alarm routed to the controls engineer, but the appropriate owner had not accepted it. No temporary evidence exception covered that missing acceptance. By the handoff, the applicable scope of every item other than S-2 had been established.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as out of scope for the 19:00 handoff.", "right": "At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-032-001", "id": "scale-diverse-032-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff. At 18:52, the operations-coordinator-approved exception register listed scrubber S-2 as out of scope for the 19:00 handoff. No record other than the master scope record and the exception register governed whether S-2 applied to this handoff. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. Inez was unavailable and did not perform any reconciliation before the readiness rating. At 19:00, S-2 had an unresolved alarm routed to the controls engineer, but the appropriate owner had not accepted it. No temporary evidence exception covered that missing acceptance. By the handoff, the applicable scope of every item other than S-2 had been established."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing reconciliation and non-precedence policy, as well as the S-2, 19:00 handoff, and Mara-to-Dev bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the exception register’s S-2 scope entry, creating an unreconciled conflict without duplicating contradictory measurements or assertions elsewhere. Neither context contains an answer label, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At the 19:00 handoff from outgoing shift lead Mara to incoming shift lead Dev, the master scope record had received its initial update at 18:45, and the S-2 entry did not change in the subsequent revision. The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as in scope for the 19:00 handoff. These documents are the complete set of records governing S-2’s scope. Applicable scope is established for every other handoff item. S-2 has an unresolved alarm, with current status and telemetry available, but the appropriate owner did not accept it before handoff. No temporary evidence exception covers the missing acceptance. Operations coordinator Inez performed no reconciliation before the readiness rating and was unavailable. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-032-002", "id": "scale-diverse-032-002-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At the 19:00 handoff from outgoing shift lead Mara to incoming shift lead Dev, the master scope record had received its initial update at 18:45, and the S-2 entry did not change in the subsequent revision. The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as in scope for the 19:00 handoff. These documents are the complete set of records governing S-2’s scope. Applicable scope is established for every other handoff item. S-2 has an unresolved alarm, with current status and telemetry available, but the appropriate owner did not accept it before handoff. No temporary evidence exception covers the missing acceptance. Operations coordinator Inez performed no reconciliation before the readiness rating and was unavailable. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing reconciliation and non-precedence policy, as well as the S-2, 19:00 handoff, and Mara-to-Dev bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the exception register’s S-2 scope entry, creating an unreconciled conflict without duplicating contradictory measurements or assertions elsewhere. Neither context contains an answer label, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At the 19:00 handoff from outgoing shift lead Mara to incoming shift lead Dev, the master scope record had received its initial update at 18:45, and the S-2 entry did not change in the subsequent revision. The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as in scope for the 19:00 handoff. These documents are the complete set of records governing S-2’s scope. Applicable scope is established for every other handoff item. S-2 has an unresolved alarm, with current status and telemetry available, but the appropriate owner did not accept it before handoff. No temporary evidence exception covers the missing acceptance. Operations coordinator Inez performed no reconciliation before the readiness rating and was unavailable. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-032-002", "id": "scale-diverse-032-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At the 19:00 handoff from outgoing shift lead Mara to incoming shift lead Dev, the master scope record had received its initial update at 18:45, and the S-2 entry did not change in the subsequent revision. The master scope record revised at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register signed at 18:53 lists scrubber S-2 as out of scope for the 19:00 handoff. These documents are the complete set of records governing S-2’s scope. Applicable scope is established for every other handoff item. S-2 has an unresolved alarm, with current status and telemetry available, but the appropriate owner did not accept it before handoff. No temporary evidence exception covers the missing acceptance. Operations coordinator Inez performed no reconciliation before the readiness rating and was unavailable. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing reconciliation policy, the 19:00 handoff decision scope, and the relevant entities. The two evidence spans are complete factual sentences; the counterfactual coherently changes only the exception register’s S-2 scope assertion, creating a conflict that remains consistent with the unchanged statement that Inez performed no reconciliation. Neither context includes a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, outgoing shift lead Mara submitted the handoff to incoming shift lead Dev. At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff. At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff. Operations coordinator Inez performed no reconciliation before the readiness rating. Scrubber S-2 had an unresolved alarm at handoff; its current status and timestamped telemetry artifact were available, but the appropriate owner had not accepted the alarm. No temporary evidence exception covered the missing owner acceptance. The applicable scope of every handoff item other than scrubber S-2 was established at 19:00. The master scope record and exception register were the complete set of records governing whether S-2 was in scope. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as out of scope for the 19:00 handoff.", "right": "At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-032-003", "id": "scale-diverse-032-003-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, outgoing shift lead Mara submitted the handoff to incoming shift lead Dev. At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff. At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff. Operations coordinator Inez performed no reconciliation before the readiness rating. Scrubber S-2 had an unresolved alarm at handoff; its current status and timestamped telemetry artifact were available, but the appropriate owner had not accepted the alarm. No temporary evidence exception covered the missing owner acceptance. The applicable scope of every handoff item other than scrubber S-2 was established at 19:00. The master scope record and exception register were the complete set of records governing whether S-2 was in scope. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing reconciliation policy, the 19:00 handoff decision scope, and the relevant entities. The two evidence spans are complete factual sentences; the counterfactual coherently changes only the exception register’s S-2 scope assertion, creating a conflict that remains consistent with the unchanged statement that Inez performed no reconciliation. Neither context includes a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, outgoing shift lead Mara submitted the handoff to incoming shift lead Dev. At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff. At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff. Operations coordinator Inez performed no reconciliation before the readiness rating. Scrubber S-2 had an unresolved alarm at handoff; its current status and timestamped telemetry artifact were available, but the appropriate owner had not accepted the alarm. No temporary evidence exception covered the missing owner acceptance. The applicable scope of every handoff item other than scrubber S-2 was established at 19:00. The master scope record and exception register were the complete set of records governing whether S-2 was in scope. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as out of scope for the 19:00 handoff.", "right": "At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-032-003", "id": "scale-diverse-032-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, outgoing shift lead Mara submitted the handoff to incoming shift lead Dev. At 18:45, the master scope record listed scrubber S-2 as in scope for the 19:00 handoff. At 18:47, the operations-coordinator-approved exception register listed scrubber S-2 as out of scope for the 19:00 handoff. Operations coordinator Inez performed no reconciliation before the readiness rating. Scrubber S-2 had an unresolved alarm at handoff; its current status and timestamped telemetry artifact were available, but the appropriate owner had not accepted the alarm. No temporary evidence exception covered the missing owner acceptance. The applicable scope of every handoff item other than scrubber S-2 was established at 19:00. The master scope record and exception register were the complete set of records governing whether S-2 was in scope. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions object and preserve the original-state policy requiring Inez to reconcile conflicting scope records, with neither record taking precedence. The decision remains bound to the 19:00 handoff, scrubber S-2, and the handoff-readiness question; changed record contents and timestamps are case observations rather than altered policy or bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently creates a conflict by changing the exception register from in scope to out of scope while leaving the master record in scope and stating that no reconciliation occurred. Neither context contains a gold answer, output instruction, answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"Concise field note: At 19:00, outgoing shift lead Mara submitted the handoff to incoming shift lead Dev. At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff. At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as in scope for the 19:00 handoff. Those documents are the complete set of records governing whether S-2 applies to this handoff. Inez was unavailable and performed no reconciliation before the readiness rating. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. At handoff, S-2 had an unresolved alarm, current status, and a timestamped telemetry artifact. The alarm had been routed to the controls engineer, but the appropriate owner had not accepted it. No temporary evidence exception covered the missing acceptance. At 19:00, the applicable scope of every other handoff item was established, and each such item had complete evidence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-032-004", "id": "scale-diverse-032-004-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "Concise field note: At 19:00, outgoing shift lead Mara submitted the handoff to incoming shift lead Dev. At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff. At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as in scope for the 19:00 handoff. Those documents are the complete set of records governing whether S-2 applies to this handoff. Inez was unavailable and performed no reconciliation before the readiness rating. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. At handoff, S-2 had an unresolved alarm, current status, and a timestamped telemetry artifact. The alarm had been routed to the controls engineer, but the appropriate owner had not accepted it. No temporary evidence exception covered the missing acceptance. At 19:00, the applicable scope of every other handoff item was established, and each such item had complete evidence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions object and preserve the original-state policy requiring Inez to reconcile conflicting scope records, with neither record taking precedence. The decision remains bound to the 19:00 handoff, scrubber S-2, and the handoff-readiness question; changed record contents and timestamps are case observations rather than altered policy or bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently creates a conflict by changing the exception register from in scope to out of scope while leaving the master record in scope and stating that no reconciliation occurred. Neither context contains a gold answer, output instruction, answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"Concise field note: At 19:00, outgoing shift lead Mara submitted the handoff to incoming shift lead Dev. At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff. At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as in scope for the 19:00 handoff. Those documents are the complete set of records governing whether S-2 applies to this handoff. Inez was unavailable and performed no reconciliation before the readiness rating. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. At handoff, S-2 had an unresolved alarm, current status, and a timestamped telemetry artifact. The alarm had been routed to the controls engineer, but the appropriate owner had not accepted it. No temporary evidence exception covered the missing acceptance. At 19:00, the applicable scope of every other handoff item was established, and each such item had complete evidence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-032-004", "id": "scale-diverse-032-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "Concise field note: At 19:00, outgoing shift lead Mara submitted the handoff to incoming shift lead Dev. At 18:45, the master scope record lists scrubber S-2 as in scope for the 19:00 handoff. At 18:52, the operations-coordinator-approved exception register lists scrubber S-2 as out of scope for the 19:00 handoff. Those documents are the complete set of records governing whether S-2 applies to this handoff. Inez was unavailable and performed no reconciliation before the readiness rating. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. At handoff, S-2 had an unresolved alarm, current status, and a timestamped telemetry artifact. The alarm had been routed to the controls engineer, but the appropriate owner had not accepted it. No temporary evidence exception covered the missing acceptance. At 19:00, the applicable scope of every other handoff item was established, and each such item had complete evidence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing routing and readiness policy verbatim and retain the same evaluated handoff, 06:00 shift change, C-7 alert, and requested determinations. The two focus spans are complete factual sentences. The counterfactual changes only Marco’s acknowledgment entry from an explicit acknowledgment to an explicit non-acknowledgment, which is coherent with his recorded ownership and all other unchanged evidence. Neither context includes an answer code, gold label, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"evidence\":[\"At 05:44, status-board export E was generated.\",\"At 05:47, export E was attached to the evaluated handoff scheduled for the 06:00 shift change.\",\"At 05:49, the board listed exactly one item unresolved for that shift change: C-7, an equipment alert.\",\"At 05:50, the duty roster identified Marco as the operations coordinator for the 06:00 shift change.\",\"At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7.\",\"At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly acknowledge C-7.”\",\"At 05:56, incoming lead Dana entered, “I explicitly accept the 06:00 shift.”\",\"The evaluated handoff occurred at the 06:00 shift change.\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "4"], "text": "At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7."}, {"path": ["evidence", "5"], "text": "At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly acknowledge C-7.”"}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7.", "negative_left": "At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7.", "negative_right": "At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly do not acknowledge C-7.”", "right": "At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly acknowledge C-7.”"}, "verifier_independent_model": false}, "family": "scale-diverse-033-001", "id": "scale-diverse-033-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["At 05:44, status-board export E was generated.", "At 05:47, export E was attached to the evaluated handoff scheduled for the 06:00 shift change.", "At 05:49, the board listed exactly one item unresolved for that shift change: C-7, an equipment alert.", "At 05:50, the duty roster identified Marco as the operations coordinator for the 06:00 shift change.", "At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7.", "At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly acknowledge C-7.”", "At 05:56, incoming lead Dana entered, “I explicitly accept the 06:00 shift.”", "The evaluated handoff occurred at the 06:00 shift change."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing routing and readiness policy verbatim and retain the same evaluated handoff, 06:00 shift change, C-7 alert, and requested determinations. The two focus spans are complete factual sentences. The counterfactual changes only Marco’s acknowledgment entry from an explicit acknowledgment to an explicit non-acknowledgment, which is coherent with his recorded ownership and all other unchanged evidence. Neither context includes an answer code, gold label, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"evidence\":[\"At 05:44, status-board export E was generated.\",\"At 05:47, export E was attached to the evaluated handoff scheduled for the 06:00 shift change.\",\"At 05:49, the board listed exactly one item unresolved for that shift change: C-7, an equipment alert.\",\"At 05:50, the duty roster identified Marco as the operations coordinator for the 06:00 shift change.\",\"At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7.\",\"At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly acknowledge C-7.”\",\"At 05:56, incoming lead Dana entered, “I explicitly accept the 06:00 shift.”\",\"The evaluated handoff occurred at the 06:00 shift change.\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "4"], "text": "At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7."}, {"path": ["evidence", "5"], "text": "At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly acknowledge C-7.”"}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7.", "negative_left": "At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7.", "negative_right": "At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly do not acknowledge C-7.”", "right": "At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly acknowledge C-7.”"}, "verifier_independent_model": false}, "family": "scale-diverse-033-001", "id": "scale-diverse-033-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["At 05:44, status-board export E was generated.", "At 05:47, export E was attached to the evaluated handoff scheduled for the 06:00 shift change.", "At 05:49, the board listed exactly one item unresolved for that shift change: C-7, an equipment alert.", "At 05:50, the duty roster identified Marco as the operations coordinator for the 06:00 shift change.", "At 05:51, the evaluated handoff record identifies Marco as the recorded owner of C-7.", "At 05:53, the evaluated handoff record attributes to Marco the entry, “I explicitly do not acknowledge C-7.”", "At 05:56, incoming lead Dana entered, “I explicitly accept the 06:00 shift.”", "The evaluated handoff occurred at the 06:00 shift change."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy verbatim and retain the same evaluated 06:00 handoff, C-7 entity, routing and acknowledgment paths, and request scope. The two focus spans are complete factual sentences. The counterfactual changes only Marco’s exact acknowledgment set from {C-7} to {C-9}; this is coherent with the unchanged facts because no other sentence asserts that Marco acknowledged C-7 or forbids acknowledgment of C-9. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"evidence\":[\"The audit concerns the evaluated handoff conducted at the 06:00 shift change.\",\"Status-board export E is attached to that handoff, and its generation timestamp is 05:44.\",\"The reconciliation log shows that the complete set of items unresolved at the shift change is exactly {C-7}; C-7 is an equipment alert.\",\"The staffing roster designates Marco as operations coordinator for the shift change.\",\"The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7.\",\"In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-7}.\",\"Incoming shift lead Eli explicitly accepts the 06:00 shift in the handoff record.\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "4"], "text": "The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7."}, {"path": ["evidence", "5"], "text": "In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-7}."}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7.", "negative_left": "The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7.", "negative_right": "In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-9}.", "right": "In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-7}."}, "verifier_independent_model": false}, "family": "scale-diverse-033-002", "id": "scale-diverse-033-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["The audit concerns the evaluated handoff conducted at the 06:00 shift change.", "Status-board export E is attached to that handoff, and its generation timestamp is 05:44.", "The reconciliation log shows that the complete set of items unresolved at the shift change is exactly {C-7}; C-7 is an equipment alert.", "The staffing roster designates Marco as operations coordinator for the shift change.", "The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7.", "In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-7}.", "Incoming shift lead Eli explicitly accepts the 06:00 shift in the handoff record."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy verbatim and retain the same evaluated 06:00 handoff, C-7 entity, routing and acknowledgment paths, and request scope. The two focus spans are complete factual sentences. The counterfactual changes only Marco’s exact acknowledgment set from {C-7} to {C-9}; this is coherent with the unchanged facts because no other sentence asserts that Marco acknowledged C-7 or forbids acknowledgment of C-9. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"evidence\":[\"The audit concerns the evaluated handoff conducted at the 06:00 shift change.\",\"Status-board export E is attached to that handoff, and its generation timestamp is 05:44.\",\"The reconciliation log shows that the complete set of items unresolved at the shift change is exactly {C-7}; C-7 is an equipment alert.\",\"The staffing roster designates Marco as operations coordinator for the shift change.\",\"The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7.\",\"In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-7}.\",\"Incoming shift lead Eli explicitly accepts the 06:00 shift in the handoff record.\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "4"], "text": "The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7."}, {"path": ["evidence", "5"], "text": "In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-7}."}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7.", "negative_left": "The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7.", "negative_right": "In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-9}.", "right": "In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-7}."}, "verifier_independent_model": false}, "family": "scale-diverse-033-002", "id": "scale-diverse-033-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["The audit concerns the evaluated handoff conducted at the 06:00 shift change.", "Status-board export E is attached to that handoff, and its generation timestamp is 05:44.", "The reconciliation log shows that the complete set of items unresolved at the shift change is exactly {C-7}; C-7 is an equipment alert.", "The staffing roster designates Marco as operations coordinator for the shift change.", "The evaluated handoff record for the 06:00 shift change identifies Marco as the recorded owner of C-7.", "In the evaluated handoff record for the 06:00 shift change, the complete set of items that Marco explicitly acknowledges is exactly {C-9}.", "Incoming shift lead Eli explicitly accepts the 06:00 shift in the handoff record."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing routing and readiness policy, while the unchanged questions object preserves the decision instructions and criteria. The request remains bound to the evaluated 06:00 handoff, C-7, its routing, and its owner acknowledgment. The two focus-evidence spans are complete factual sentences. The counterfactual coherently replaces Marco’s explicit acknowledgment with an exhaustive factual statement that no such acknowledgment appears, without conflicting with the unchanged ownership, export, unresolved-item, or acceptance facts. Neither context embeds an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"evidence\":[\"Review scope note: The evaluated handoff is the shift change scheduled for 06:00. The board materials identify status-board export E as attached to that handoff and show that E was generated at 05:44.\",\"Reconciliation note: At the shift change, the complete set of unresolved items is exactly {C-7}; no other item remains open. C-7 is classified as an equipment alert. The duty roster identifies Marco as the operations coordinator for that shift change.\",\"The evaluated 06:00 handoff record lists Marco as the owner of C-7.\",\"Within the evaluated 06:00 handoff record, an entry attributed to Marco states, “I explicitly acknowledge C-7.”\",\"Acceptance entry: Incoming shift lead Eli states, “I explicitly accept the 06:00 shift.”\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "2"], "text": "The evaluated 06:00 handoff record lists Marco as the owner of C-7."}, {"path": ["evidence", "3"], "text": "Within the evaluated 06:00 handoff record, an entry attributed to Marco states, “I explicitly acknowledge C-7.”"}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The evaluated 06:00 handoff record lists Marco as the owner of C-7.", "negative_left": "The evaluated 06:00 handoff record lists Marco as the owner of C-7.", "negative_right": "The complete set of entries attributed to Marco within the evaluated 06:00 handoff record contains no statement explicitly acknowledging C-7.", "right": "Within the evaluated 06:00 handoff record, an entry attributed to Marco states, “I explicitly acknowledge C-7.”"}, "verifier_independent_model": false}, "family": "scale-diverse-033-003", "id": "scale-diverse-033-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["Review scope note: The evaluated handoff is the shift change scheduled for 06:00. The board materials identify status-board export E as attached to that handoff and show that E was generated at 05:44.", "Reconciliation note: At the shift change, the complete set of unresolved items is exactly {C-7}; no other item remains open. C-7 is classified as an equipment alert. The duty roster identifies Marco as the operations coordinator for that shift change.", "The evaluated 06:00 handoff record lists Marco as the owner of C-7.", "Within the evaluated 06:00 handoff record, an entry attributed to Marco states, “I explicitly acknowledge C-7.”", "Acceptance entry: Incoming shift lead Eli states, “I explicitly accept the 06:00 shift.”"], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing routing and readiness policy, while the unchanged questions object preserves the decision instructions and criteria. The request remains bound to the evaluated 06:00 handoff, C-7, its routing, and its owner acknowledgment. The two focus-evidence spans are complete factual sentences. The counterfactual coherently replaces Marco’s explicit acknowledgment with an exhaustive factual statement that no such acknowledgment appears, without conflicting with the unchanged ownership, export, unresolved-item, or acceptance facts. Neither context embeds an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"evidence\":[\"Review scope note: The evaluated handoff is the shift change scheduled for 06:00. The board materials identify status-board export E as attached to that handoff and show that E was generated at 05:44.\",\"Reconciliation note: At the shift change, the complete set of unresolved items is exactly {C-7}; no other item remains open. C-7 is classified as an equipment alert. The duty roster identifies Marco as the operations coordinator for that shift change.\",\"The evaluated 06:00 handoff record lists Marco as the owner of C-7.\",\"Within the evaluated 06:00 handoff record, an entry attributed to Marco states, “I explicitly acknowledge C-7.”\",\"Acceptance entry: Incoming shift lead Eli states, “I explicitly accept the 06:00 shift.”\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "2"], "text": "The evaluated 06:00 handoff record lists Marco as the owner of C-7."}, {"path": ["evidence", "3"], "text": "Within the evaluated 06:00 handoff record, an entry attributed to Marco states, “I explicitly acknowledge C-7.”"}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The evaluated 06:00 handoff record lists Marco as the owner of C-7.", "negative_left": "The evaluated 06:00 handoff record lists Marco as the owner of C-7.", "negative_right": "The complete set of entries attributed to Marco within the evaluated 06:00 handoff record contains no statement explicitly acknowledging C-7.", "right": "Within the evaluated 06:00 handoff record, an entry attributed to Marco states, “I explicitly acknowledge C-7.”"}, "verifier_independent_model": false}, "family": "scale-diverse-033-003", "id": "scale-diverse-033-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["Review scope note: The evaluated handoff is the shift change scheduled for 06:00. The board materials identify status-board export E as attached to that handoff and show that E was generated at 05:44.", "Reconciliation note: At the shift change, the complete set of unresolved items is exactly {C-7}; no other item remains open. C-7 is classified as an equipment alert. The duty roster identifies Marco as the operations coordinator for that shift change.", "The evaluated 06:00 handoff record lists Marco as the owner of C-7.", "The complete set of entries attributed to Marco within the evaluated 06:00 handoff record contains no statement explicitly acknowledging C-7.", "Acceptance entry: Incoming shift lead Eli states, “I explicitly accept the 06:00 shift.”"], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original production-only scope, readiness requirements, handoff entity, ticket, item, and relevant timestamps without adding policy exceptions or defaults. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the acknowledgment actor from owner Mara Chen to Elias Navarro; this is coherent with Mara remaining the named owner and does not create a contradictory duplicate assertion. Neither context embeds a rating, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Shift recorder\",\"text\":\"The shift inventory contains exactly two production items: the completed filter swap and the Line 4 pump vibration. For each item, the handoff records a current status and a timestamped artifact. The filter swap is complete, with work order W-219 and its 17:30 completion photo; the pump is stable at 6.1 mm/s as of 17:42, with trend snapshot TS-144 attached.\"},{\"speaker\":\"Outgoing shift lead\",\"text\":\"The Line 4 pump vibration is the only unresolved production item in this handoff. Ticket M-882 records a 20:00 inspection as its next action.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Ticket audit\",\"text\":\"At 17:48, maintenance duty engineer Mara Chen recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action.\"},{\"speaker\":\"Incoming shift lead\",\"text\":\"The cafeteria freezer alarm and parking repaint notices are facilities announcements outside the production inventory.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["2", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item."}, {"path": ["3", "text"], "text": "At 17:48, maintenance duty engineer Mara Chen recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.", "negative_right": "At 17:48, maintenance duty engineer Elias Navarro recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action.", "right": "At 17:48, maintenance duty engineer Mara Chen recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action."}, "verifier_independent_model": false}, "family": "scale-diverse-034-001", "id": "scale-diverse-034-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Shift recorder", "text": "The shift inventory contains exactly two production items: the completed filter swap and the Line 4 pump vibration. For each item, the handoff records a current status and a timestamped artifact. The filter swap is complete, with work order W-219 and its 17:30 completion photo; the pump is stable at 6.1 mm/s as of 17:42, with trend snapshot TS-144 attached."}, {"speaker": "Outgoing shift lead", "text": "The Line 4 pump vibration is the only unresolved production item in this handoff. Ticket M-882 records a 20:00 inspection as its next action."}, {"speaker": "Operations coordinator", "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item."}, {"speaker": "Ticket audit", "text": "At 17:48, maintenance duty engineer Mara Chen recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action."}, {"speaker": "Incoming shift lead", "text": "The cafeteria freezer alarm and parking repaint notices are facilities announcements outside the production inventory."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original production-only scope, readiness requirements, handoff entity, ticket, item, and relevant timestamps without adding policy exceptions or defaults. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the acknowledgment actor from owner Mara Chen to Elias Navarro; this is coherent with Mara remaining the named owner and does not create a contradictory duplicate assertion. Neither context embeds a rating, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Shift recorder\",\"text\":\"The shift inventory contains exactly two production items: the completed filter swap and the Line 4 pump vibration. For each item, the handoff records a current status and a timestamped artifact. The filter swap is complete, with work order W-219 and its 17:30 completion photo; the pump is stable at 6.1 mm/s as of 17:42, with trend snapshot TS-144 attached.\"},{\"speaker\":\"Outgoing shift lead\",\"text\":\"The Line 4 pump vibration is the only unresolved production item in this handoff. Ticket M-882 records a 20:00 inspection as its next action.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Ticket audit\",\"text\":\"At 17:48, maintenance duty engineer Mara Chen recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action.\"},{\"speaker\":\"Incoming shift lead\",\"text\":\"The cafeteria freezer alarm and parking repaint notices are facilities announcements outside the production inventory.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["2", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item."}, {"path": ["3", "text"], "text": "At 17:48, maintenance duty engineer Mara Chen recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.", "negative_right": "At 17:48, maintenance duty engineer Elias Navarro recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action.", "right": "At 17:48, maintenance duty engineer Mara Chen recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action."}, "verifier_independent_model": false}, "family": "scale-diverse-034-001", "id": "scale-diverse-034-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Shift recorder", "text": "The shift inventory contains exactly two production items: the completed filter swap and the Line 4 pump vibration. For each item, the handoff records a current status and a timestamped artifact. The filter swap is complete, with work order W-219 and its 17:30 completion photo; the pump is stable at 6.1 mm/s as of 17:42, with trend snapshot TS-144 attached."}, {"speaker": "Outgoing shift lead", "text": "The Line 4 pump vibration is the only unresolved production item in this handoff. Ticket M-882 records a 20:00 inspection as its next action."}, {"speaker": "Operations coordinator", "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item."}, {"speaker": "Ticket audit", "text": "At 17:48, maintenance duty engineer Elias Navarro recorded on ticket M-882 the sole acknowledgment concerning the unresolved Line 4 pump vibration item's routing and next action."}, {"speaker": "Incoming shift lead", "text": "The cafeteria freezer alarm and parking repaint notices are facilities announcements outside the production inventory."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original handoff scope, production-item set, entities, timestamps, and readiness policy without adding exceptions or defaults. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the identity of the person who made the single 17:48 acknowledgment; it remains consistent with the unchanged statement that exactly one acknowledgment entry exists, while the named owner remains Mara Chen. Neither context states a readiness answer, code, rule table, proposition ID, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Evidence reviewer\",\"text\":\"The reconciliation lists exactly two in-scope production items for this handoff: the Line 4 pump vibration and the filter swap in W-219. No other production items are in scope.\"},{\"speaker\":\"Records clerk\",\"text\":\"The pump remains the only unresolved item. Its current status at 17:42 reports stable vibration at 6.1 mm/s, and timestamped trend snapshot TS-144 was captured at 17:43. The filter swap is resolved, has a current completed status, and has a completion photo timestamped 17:30 in W-219.\"},{\"speaker\":\"Ticket reviewer\",\"text\":\"Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Duty coordinator\",\"text\":\"The duty roster confirms that maintenance duty engineer is the appropriate owner role for this issue. M-882 records the 20:00 inspection as its next action. The acknowledgment register contains exactly one entry covering M-882’s routing and that inspection action, recorded on M-882 at 17:48.\"},{\"speaker\":\"Audit clerk\",\"text\":\"The 17:48 acknowledgment recorded on ticket M-882 was made by maintenance duty engineer Mara Chen.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["2", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item."}, {"path": ["4", "text"], "text": "The 17:48 acknowledgment recorded on ticket M-882 was made by maintenance duty engineer Mara Chen."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.", "negative_right": "The 17:48 acknowledgment recorded on ticket M-882 was made by operations supervisor Elias Grant.", "right": "The 17:48 acknowledgment recorded on ticket M-882 was made by maintenance duty engineer Mara Chen."}, "verifier_independent_model": false}, "family": "scale-diverse-034-002", "id": "scale-diverse-034-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Evidence reviewer", "text": "The reconciliation lists exactly two in-scope production items for this handoff: the Line 4 pump vibration and the filter swap in W-219. No other production items are in scope."}, {"speaker": "Records clerk", "text": "The pump remains the only unresolved item. Its current status at 17:42 reports stable vibration at 6.1 mm/s, and timestamped trend snapshot TS-144 was captured at 17:43. The filter swap is resolved, has a current completed status, and has a completion photo timestamped 17:30 in W-219."}, {"speaker": "Ticket reviewer", "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item."}, {"speaker": "Duty coordinator", "text": "The duty roster confirms that maintenance duty engineer is the appropriate owner role for this issue. M-882 records the 20:00 inspection as its next action. The acknowledgment register contains exactly one entry covering M-882’s routing and that inspection action, recorded on M-882 at 17:48."}, {"speaker": "Audit clerk", "text": "The 17:48 acknowledgment recorded on ticket M-882 was made by maintenance duty engineer Mara Chen."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original handoff scope, production-item set, entities, timestamps, and readiness policy without adding exceptions or defaults. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the identity of the person who made the single 17:48 acknowledgment; it remains consistent with the unchanged statement that exactly one acknowledgment entry exists, while the named owner remains Mara Chen. Neither context states a readiness answer, code, rule table, proposition ID, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Evidence reviewer\",\"text\":\"The reconciliation lists exactly two in-scope production items for this handoff: the Line 4 pump vibration and the filter swap in W-219. No other production items are in scope.\"},{\"speaker\":\"Records clerk\",\"text\":\"The pump remains the only unresolved item. Its current status at 17:42 reports stable vibration at 6.1 mm/s, and timestamped trend snapshot TS-144 was captured at 17:43. The filter swap is resolved, has a current completed status, and has a completion photo timestamped 17:30 in W-219.\"},{\"speaker\":\"Ticket reviewer\",\"text\":\"Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Duty coordinator\",\"text\":\"The duty roster confirms that maintenance duty engineer is the appropriate owner role for this issue. M-882 records the 20:00 inspection as its next action. The acknowledgment register contains exactly one entry covering M-882’s routing and that inspection action, recorded on M-882 at 17:48.\"},{\"speaker\":\"Audit clerk\",\"text\":\"The 17:48 acknowledgment recorded on ticket M-882 was made by maintenance duty engineer Mara Chen.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["2", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item."}, {"path": ["4", "text"], "text": "The 17:48 acknowledgment recorded on ticket M-882 was made by maintenance duty engineer Mara Chen."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item.", "negative_right": "The 17:48 acknowledgment recorded on ticket M-882 was made by operations supervisor Elias Grant.", "right": "The 17:48 acknowledgment recorded on ticket M-882 was made by maintenance duty engineer Mara Chen."}, "verifier_independent_model": false}, "family": "scale-diverse-034-002", "id": "scale-diverse-034-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Evidence reviewer", "text": "The reconciliation lists exactly two in-scope production items for this handoff: the Line 4 pump vibration and the filter swap in W-219. No other production items are in scope."}, {"speaker": "Records clerk", "text": "The pump remains the only unresolved item. Its current status at 17:42 reports stable vibration at 6.1 mm/s, and timestamped trend snapshot TS-144 was captured at 17:43. The filter swap is resolved, has a current completed status, and has a completion photo timestamped 17:30 in W-219."}, {"speaker": "Ticket reviewer", "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the owner of the unresolved Line 4 pump vibration item."}, {"speaker": "Duty coordinator", "text": "The duty roster confirms that maintenance duty engineer is the appropriate owner role for this issue. M-882 records the 20:00 inspection as its next action. The acknowledgment register contains exactly one entry covering M-882’s routing and that inspection action, recorded on M-882 at 17:48."}, {"speaker": "Audit clerk", "text": "The 17:48 acknowledgment recorded on ticket M-882 was made by operations supervisor Elias Grant."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing handoff policy, readiness-scoring scope, relevant entities, routing path, and scoring-time framing. The two evidence spans are complete factual sentences. The counterfactual changes only the registry mapping of T-47 from Facilities to Security; this coherently makes the recorded destination inconsistent with the unchanged responsibility assignment for refrigeration alarms without creating contradictory duplicate measurements or assertions. Neither context contains a gold score, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship rather than a policy conclusion or bundled final classification. The focus atom concerns the factual routing of item U. The base and counter assignments are jointly realizable under the state policy and differ only in whether U is routed to Facilities. The policy evidence correctly preserves the state-originating requirements needed to interpret the unchanged question, including the artifact, routing, paraphrase, and refrigeration-owner rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "With exactly one unresolved item, the conjunction establishes a timestamped status artifact, correct routing to Facilities for the refrigeration alarm, and an accurate own-word paraphrase of both status and next action. The sole owner's acknowledgment is explicitly refuted, so score 1 is entailed by the ordered rubric.", "rule_index": 0, "sound": true}, {"reason": "Facilities is the responsible owner for the sole refrigeration alarm, while routing to Facilities is explicitly refuted. Thus the unresolved item is routed to the wrong owner or no owner, which is sufficient for score 0 regardless of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At scoring time, the number of unresolved operational items in the handoff is exactly one."}, {"id": "a2", "statement": "The sole unresolved operational item U is a refrigeration alarm."}, {"id": "a3", "statement": "A status artifact for unresolved item U is attached to the handoff."}, {"id": "a4", "statement": "The attached status artifact for unresolved item U has a timestamp."}, {"id": "a5", "statement": "At scoring time, unresolved item U is routed to the Facilities team."}, {"id": "a6", "statement": "Incoming shift lead Dev's paraphrase accurately represents the status of unresolved item U."}, {"id": "a7", "statement": "Incoming shift lead Dev expresses the status of unresolved item U in his own words."}, {"id": "a8", "statement": "Incoming shift lead Dev's paraphrase accurately represents the next action for unresolved item U."}, {"id": "a9", "statement": "Incoming shift lead Dev expresses the next action for unresolved item U in his own words."}, {"id": "a10", "statement": "The owner of unresolved item U has acknowledged receipt of the item."}], "base_state_json": "\"Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count. Facilities handles refrigeration alarms. At 13:42 UTC, Mara recorded unresolved item U as a refrigeration alarm and attached a status artifact timestamped 2026-09-17 13:41 UTC. At 13:50 UTC, incoming shift lead Dev described in his own words that the alarm remained active and that the refrigeration equipment should be inspected before restart. A subsequent comparison with the artifact confirmed that Dev accurately represented both the current status and the next action. No owner acknowledgment of receipt had been received by scoring time. At 2026-09-17 14:00 UTC, U was the handoff’s only unresolved operational item. At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination. At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Facilities team.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination."}, {"path": [], "text": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Facilities team."}], "policy_evidence": [{"path": [], "text": "Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count."}, {"path": [], "text": "Facilities handles refrigeration alarms."}], "rules": [{"justification": "There is exactly one unresolved item; it has the required timestamped status artifact, is correctly routed to Facilities under the refrigeration-alarm policy, and has an accurate own-word paraphrase of both status and next action. The owner's explicitly absent acknowledgment therefore makes the handoff minimally ready.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}, {"justification": "The sole unresolved item is a refrigeration alarm, for which Facilities is the responsible owner, but the item is explicitly not routed to Facilities. This wrong routing is sufficient for a not-ready score despite the other documented requirements.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination.", "negative_left": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination.", "negative_right": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Security team.", "right": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Facilities team."}, "verifier_independent_model": false}, "family": "scale-diverse-035-001", "id": "scale-diverse-035-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one unresolved operational item lacks a required status artifact, is routed to the wrong or no owner, or lacks an accurate own-word paraphrase from the incoming lead.", "1 — Minimally ready: Every unresolved item has the required artifact, correct owner, and valid paraphrase, but at least one owner has not acknowledged receipt.", "2 — Provisionally ready: All required fields and owner acknowledgments are present, but at least one unresolved item lacks a follow-up time or completion checkpoint.", "3 — Ready: Required evidence, routing, paraphrases, acknowledgments, and follow-up checkpoints are complete; only minor nonoperational clarification remains.", "4 — Fully ready: Every item is fully documented, correctly routed, paraphrased, acknowledged, and scheduled, with no clarification or correction needed."], "instructions": "Rate the handoff’s readiness using the ordered rubric. Check the routing, required status evidence, and whether the incoming lead supplied a valid paraphrase.", "type": "score"}}, "state": "Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count. Facilities handles refrigeration alarms. At 13:42 UTC, Mara recorded unresolved item U as a refrigeration alarm and attached a status artifact timestamped 2026-09-17 13:41 UTC. At 13:50 UTC, incoming shift lead Dev described in his own words that the alarm remained active and that the refrigeration equipment should be inspected before restart. A subsequent comparison with the artifact confirmed that Dev accurately represented both the current status and the next action. No owner acknowledgment of receipt had been received by scoring time. At 2026-09-17 14:00 UTC, U was the handoff’s only unresolved operational item. At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination. At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Facilities team."}, "method": "c2d", "provenance": {"source_id": "diverse-035", "source_is_synthetic": true, "source_sha256": "ef2504c9a18e9a4c12f01c98be90595f7a44fb099e598d65aa259ad0418ac996", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing handoff policy, readiness-scoring scope, relevant entities, routing path, and scoring-time framing. The two evidence spans are complete factual sentences. The counterfactual changes only the registry mapping of T-47 from Facilities to Security; this coherently makes the recorded destination inconsistent with the unchanged responsibility assignment for refrigeration alarms without creating contradictory duplicate measurements or assertions. Neither context contains a gold score, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship rather than a policy conclusion or bundled final classification. The focus atom concerns the factual routing of item U. The base and counter assignments are jointly realizable under the state policy and differ only in whether U is routed to Facilities. The policy evidence correctly preserves the state-originating requirements needed to interpret the unchanged question, including the artifact, routing, paraphrase, and refrigeration-owner rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "With exactly one unresolved item, the conjunction establishes a timestamped status artifact, correct routing to Facilities for the refrigeration alarm, and an accurate own-word paraphrase of both status and next action. The sole owner's acknowledgment is explicitly refuted, so score 1 is entailed by the ordered rubric.", "rule_index": 0, "sound": true}, {"reason": "Facilities is the responsible owner for the sole refrigeration alarm, while routing to Facilities is explicitly refuted. Thus the unresolved item is routed to the wrong owner or no owner, which is sufficient for score 0 regardless of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At scoring time, the number of unresolved operational items in the handoff is exactly one."}, {"id": "a2", "statement": "The sole unresolved operational item U is a refrigeration alarm."}, {"id": "a3", "statement": "A status artifact for unresolved item U is attached to the handoff."}, {"id": "a4", "statement": "The attached status artifact for unresolved item U has a timestamp."}, {"id": "a5", "statement": "At scoring time, unresolved item U is routed to the Facilities team."}, {"id": "a6", "statement": "Incoming shift lead Dev's paraphrase accurately represents the status of unresolved item U."}, {"id": "a7", "statement": "Incoming shift lead Dev expresses the status of unresolved item U in his own words."}, {"id": "a8", "statement": "Incoming shift lead Dev's paraphrase accurately represents the next action for unresolved item U."}, {"id": "a9", "statement": "Incoming shift lead Dev expresses the next action for unresolved item U in his own words."}, {"id": "a10", "statement": "The owner of unresolved item U has acknowledged receipt of the item."}], "base_state_json": "\"Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count. Facilities handles refrigeration alarms. At 13:42 UTC, Mara recorded unresolved item U as a refrigeration alarm and attached a status artifact timestamped 2026-09-17 13:41 UTC. At 13:50 UTC, incoming shift lead Dev described in his own words that the alarm remained active and that the refrigeration equipment should be inspected before restart. A subsequent comparison with the artifact confirmed that Dev accurately represented both the current status and the next action. No owner acknowledgment of receipt had been received by scoring time. At 2026-09-17 14:00 UTC, U was the handoff’s only unresolved operational item. At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination. At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Facilities team.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination."}, {"path": [], "text": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Facilities team."}], "policy_evidence": [{"path": [], "text": "Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count."}, {"path": [], "text": "Facilities handles refrigeration alarms."}], "rules": [{"justification": "There is exactly one unresolved item; it has the required timestamped status artifact, is correctly routed to Facilities under the refrigeration-alarm policy, and has an accurate own-word paraphrase of both status and next action. The owner's explicitly absent acknowledgment therefore makes the handoff minimally ready.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}, {"justification": "The sole unresolved item is a refrigeration alarm, for which Facilities is the responsible owner, but the item is explicitly not routed to Facilities. This wrong routing is sufficient for a not-ready score despite the other documented requirements.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination.", "negative_left": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination.", "negative_right": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Security team.", "right": "At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Facilities team."}, "verifier_independent_model": false}, "family": "scale-diverse-035-001", "id": "scale-diverse-035-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one unresolved operational item lacks a required status artifact, is routed to the wrong or no owner, or lacks an accurate own-word paraphrase from the incoming lead.", "1 — Minimally ready: Every unresolved item has the required artifact, correct owner, and valid paraphrase, but at least one owner has not acknowledged receipt.", "2 — Provisionally ready: All required fields and owner acknowledgments are present, but at least one unresolved item lacks a follow-up time or completion checkpoint.", "3 — Ready: Required evidence, routing, paraphrases, acknowledgments, and follow-up checkpoints are complete; only minor nonoperational clarification remains.", "4 — Fully ready: Every item is fully documented, correctly routed, paraphrased, acknowledged, and scheduled, with no clarification or correction needed."], "instructions": "Rate the handoff’s readiness using the ordered rubric. Check the routing, required status evidence, and whether the incoming lead supplied a valid paraphrase.", "type": "score"}}, "state": "Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count. Facilities handles refrigeration alarms. At 13:42 UTC, Mara recorded unresolved item U as a refrigeration alarm and attached a status artifact timestamped 2026-09-17 13:41 UTC. At 13:50 UTC, incoming shift lead Dev described in his own words that the alarm remained active and that the refrigeration equipment should be inspected before restart. A subsequent comparison with the artifact confirmed that Dev accurately represented both the current status and the next action. No owner acknowledgment of receipt had been received by scoring time. At 2026-09-17 14:00 UTC, U was the handoff’s only unresolved operational item. At the scoring time of 2026-09-17 14:00 UTC, the handoff’s authoritative routing record lists team code T-47 as unresolved item U’s sole destination. At the scoring time of 2026-09-17 14:00 UTC, the handoff’s team registry uniquely maps code T-47 to the Security team."}, "method": "c2d", "provenance": {"source_id": "diverse-035", "source_is_synthetic": true, "source_sha256": "ef2504c9a18e9a4c12f01c98be90595f7a44fb099e598d65aa259ad0418ac996", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts retain the same handoff policy, request, entities, 18:00 decision time, and readiness-rubric scope. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Mara Ilyan’s staffing role from Operations coordinator to Inventory controller; this remains coherent with the unchanged ownership record and introduces no duplicate or contradictory measurement, count, or assertion. Neither context embeds a readiness level, answer code, rule table, proposition identifier, label rationale, or output directive.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"At 17:42, a timestamped clearance photo became the current status artifact for the routine Conveyor 4 restart check; its stated next action was a supervised restart.\",\"At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00.\",\"Mara Ilyan's staffing record shows that she served as the Operations coordinator continuously from 17:30 through 18:30.\",\"At 17:48, the Filler 2 pressure alarm, classified as the only critical equipment alarm, received a current trend report, an attached valve diagnostic result, and the next action ‘inspect after 18:15.’\",\"At 17:50, Jonas Reed was recorded as owner of the Conveyor 4 restart check; the staffing roster shows him serving as Incoming shift lead at 18:00.\",\"At 18:00, the unresolved-item register contained exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"At 18:00, the Incoming shift lead had not acknowledged the handoff; the acknowledgement field remained pending.\"],\"request\":\"Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00."}, {"path": ["evidence", "2"], "text": "Mara Ilyan's staffing record shows that she served as the Operations coordinator continuously from 17:30 through 18:30."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00.", "negative_left": "At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00.", "negative_right": "Mara Ilyan's staffing record shows that she served exclusively as the Inventory controller continuously from 17:30 through 18:30.", "right": "Mara Ilyan's staffing record shows that she served as the Operations coordinator continuously from 17:30 through 18:30."}, "verifier_independent_model": false}, "family": "scale-diverse-036-001", "id": "scale-diverse-036-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["At 17:42, a timestamped clearance photo became the current status artifact for the routine Conveyor 4 restart check; its stated next action was a supervised restart.", "At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00.", "Mara Ilyan's staffing record shows that she served as the Operations coordinator continuously from 17:30 through 18:30.", "At 17:48, the Filler 2 pressure alarm, classified as the only critical equipment alarm, received a current trend report, an attached valve diagnostic result, and the next action ‘inspect after 18:15.’", "At 17:50, Jonas Reed was recorded as owner of the Conveyor 4 restart check; the staffing roster shows him serving as Incoming shift lead at 18:00.", "At 18:00, the unresolved-item register contained exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.", "At 18:00, the Incoming shift lead had not acknowledged the handoff; the acknowledgement field remained pending."], "request": "Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts retain the same handoff policy, request, entities, 18:00 decision time, and readiness-rubric scope. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Mara Ilyan’s staffing role from Operations coordinator to Inventory controller; this remains coherent with the unchanged ownership record and introduces no duplicate or contradictory measurement, count, or assertion. Neither context embeds a readiness level, answer code, rule table, proposition identifier, label rationale, or output directive.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"At 17:42, a timestamped clearance photo became the current status artifact for the routine Conveyor 4 restart check; its stated next action was a supervised restart.\",\"At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00.\",\"Mara Ilyan's staffing record shows that she served as the Operations coordinator continuously from 17:30 through 18:30.\",\"At 17:48, the Filler 2 pressure alarm, classified as the only critical equipment alarm, received a current trend report, an attached valve diagnostic result, and the next action ‘inspect after 18:15.’\",\"At 17:50, Jonas Reed was recorded as owner of the Conveyor 4 restart check; the staffing roster shows him serving as Incoming shift lead at 18:00.\",\"At 18:00, the unresolved-item register contained exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"At 18:00, the Incoming shift lead had not acknowledged the handoff; the acknowledgement field remained pending.\"],\"request\":\"Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00."}, {"path": ["evidence", "2"], "text": "Mara Ilyan's staffing record shows that she served as the Operations coordinator continuously from 17:30 through 18:30."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00.", "negative_left": "At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00.", "negative_right": "Mara Ilyan's staffing record shows that she served exclusively as the Inventory controller continuously from 17:30 through 18:30.", "right": "Mara Ilyan's staffing record shows that she served as the Operations coordinator continuously from 17:30 through 18:30."}, "verifier_independent_model": false}, "family": "scale-diverse-036-001", "id": "scale-diverse-036-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["At 17:42, a timestamped clearance photo became the current status artifact for the routine Conveyor 4 restart check; its stated next action was a supervised restart.", "At 17:43, employee Mara Ilyan was entered in the Filler 2 pressure alarm's owner field, and the field remained unchanged through 18:00.", "Mara Ilyan's staffing record shows that she served exclusively as the Inventory controller continuously from 17:30 through 18:30.", "At 17:48, the Filler 2 pressure alarm, classified as the only critical equipment alarm, received a current trend report, an attached valve diagnostic result, and the next action ‘inspect after 18:15.’", "At 17:50, Jonas Reed was recorded as owner of the Conveyor 4 restart check; the staffing roster shows him serving as Incoming shift lead at 18:00.", "At 18:00, the unresolved-item register contained exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.", "At 18:00, the Incoming shift lead had not acknowledged the handoff; the acknowledgement field remained pending."], "request": "Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original readiness rubric, instructions, request, 18:00 handoff scope, and routing requirements verbatim. The two focus spans are complete factual sentences. The counterfactual changes Nila Voss’s role from Operations coordinator to inventory controller without conflicting with the unchanged identifier, item count, measurements, or acknowledgement evidence. Neither context supplies a selected readiness score, answer code, label rationale, proposition identifier, or classifier-output instruction beyond the preserved governing question and policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.\",\"1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.\",\"2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.\",\"3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.\",\"4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion.\"],\"instructions\":\"Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.\",\"type\":\"score\"}},\"state\":{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"The 18:00 register lists exactly two unresolved handoff items and no others: the Conveyor 4 restart check and the Filler 2 pressure alarm. The latter is marked as the only critical equipment alarm; the former is classified as a routine process follow-up.\",\"For Conveyor 4, a 17:48 status photo records that the obstruction is cleared, and the stated next action is to perform the restart check. Its owner field identifies Tomas Ilyan, who is serving as the Incoming shift lead at 18:00.\",\"For Filler 2, a current alarm trace records continuing pressure fluctuation, and the stated next action is an inspection after 18:15. An attached diagnostic result reports intermittent regulator drift.\",\"At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss.\",\"At 18:00, Nila Voss is serving as the Operations coordinator.\",\"The acknowledgement log states that, by 18:00, the Incoming shift lead has not acknowledged the handoff.\"],\"request\":\"Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["state", "evidence", "3"], "text": "At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss."}, {"path": ["state", "evidence", "4"], "text": "At 18:00, Nila Voss is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss.", "negative_right": "At 18:00, Nila Voss is serving as the inventory controller and is not serving as the Operations coordinator.", "right": "At 18:00, Nila Voss is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "scale-diverse-036-002", "id": "scale-diverse-036-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["The 18:00 register lists exactly two unresolved handoff items and no others: the Conveyor 4 restart check and the Filler 2 pressure alarm. The latter is marked as the only critical equipment alarm; the former is classified as a routine process follow-up.", "For Conveyor 4, a 17:48 status photo records that the obstruction is cleared, and the stated next action is to perform the restart check. Its owner field identifies Tomas Ilyan, who is serving as the Incoming shift lead at 18:00.", "For Filler 2, a current alarm trace records continuing pressure fluctuation, and the stated next action is an inspection after 18:15. An attached diagnostic result reports intermittent regulator drift.", "At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss.", "At 18:00, Nila Voss is serving as the Operations coordinator.", "The acknowledgement log states that, by 18:00, the Incoming shift lead has not acknowledged the handoff."], "request": "Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed."}}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original readiness rubric, instructions, request, 18:00 handoff scope, and routing requirements verbatim. The two focus spans are complete factual sentences. The counterfactual changes Nila Voss’s role from Operations coordinator to inventory controller without conflicting with the unchanged identifier, item count, measurements, or acknowledgement evidence. Neither context supplies a selected readiness score, answer code, label rationale, proposition identifier, or classifier-output instruction beyond the preserved governing question and policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.\",\"1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.\",\"2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.\",\"3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.\",\"4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion.\"],\"instructions\":\"Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.\",\"type\":\"score\"}},\"state\":{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"The 18:00 register lists exactly two unresolved handoff items and no others: the Conveyor 4 restart check and the Filler 2 pressure alarm. The latter is marked as the only critical equipment alarm; the former is classified as a routine process follow-up.\",\"For Conveyor 4, a 17:48 status photo records that the obstruction is cleared, and the stated next action is to perform the restart check. Its owner field identifies Tomas Ilyan, who is serving as the Incoming shift lead at 18:00.\",\"For Filler 2, a current alarm trace records continuing pressure fluctuation, and the stated next action is an inspection after 18:15. An attached diagnostic result reports intermittent regulator drift.\",\"At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss.\",\"At 18:00, Nila Voss is serving as the Operations coordinator.\",\"The acknowledgement log states that, by 18:00, the Incoming shift lead has not acknowledged the handoff.\"],\"request\":\"Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["state", "evidence", "3"], "text": "At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss."}, {"path": ["state", "evidence", "4"], "text": "At 18:00, Nila Voss is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss.", "negative_right": "At 18:00, Nila Voss is serving as the inventory controller and is not serving as the Operations coordinator.", "right": "At 18:00, Nila Voss is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "scale-diverse-036-002", "id": "scale-diverse-036-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["The 18:00 register lists exactly two unresolved handoff items and no others: the Conveyor 4 restart check and the Filler 2 pressure alarm. The latter is marked as the only critical equipment alarm; the former is classified as a routine process follow-up.", "For Conveyor 4, a 17:48 status photo records that the obstruction is cleared, and the stated next action is to perform the restart check. Its owner field identifies Tomas Ilyan, who is serving as the Incoming shift lead at 18:00.", "For Filler 2, a current alarm trace records continuing pressure fluctuation, and the stated next action is an inspection after 18:15. An attached diagnostic result reports intermittent regulator drift.", "At 18:00, the Filler 2 pressure alarm's owner field contains employee identifier Q-817, which uniquely identifies Nila Voss.", "At 18:00, Nila Voss is serving as the inventory controller and is not serving as the Operations coordinator.", "The acknowledgement log states that, by 18:00, the Incoming shift lead has not acknowledged the handoff."], "request": "Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed."}}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing handoff policy, request, 18:00 time binding, entities, and evaluation scope; the unchanged questions preserve the full rubric. The two focus spans are complete factual sentences. The counterfactual coherently changes PK-731’s sole role from Operations coordinator to Maintenance scheduler, making the critical alarm’s routing factually deficient without contradicting any unchanged observation. Neither context contains a readiness score, gold answer, output instruction, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731.\",\"At 18:00, the complete shift-role roster lists Operations coordinator as the sole role held by employee PK-731.\",\"The complete 18:00 unresolved-items register contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm. It classifies the restart check as a routine process follow-up and the pressure alarm as the only critical equipment alarm.\",\"The Conveyor 4 entry includes a current restart-test sheet and the next action to perform the restart check. Its owner field names Mara Ives, who is serving as Incoming shift lead.\",\"The Filler 2 entry includes a current pressure-trend chart, an attached valve diagnostic result, and the next action to inspect the regulator after 18:15.\",\"The acknowledgement log shows that the Incoming shift lead has not acknowledged the handoff by 18:00.\"],\"request\":\"Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731."}, {"path": ["evidence", "1"], "text": "At 18:00, the complete shift-role roster lists Operations coordinator as the sole role held by employee PK-731."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731.", "negative_right": "At 18:00, the complete shift-role roster lists Maintenance scheduler as the sole role held by employee PK-731.", "right": "At 18:00, the complete shift-role roster lists Operations coordinator as the sole role held by employee PK-731."}, "verifier_independent_model": false}, "family": "scale-diverse-036-003", "id": "scale-diverse-036-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731.", "At 18:00, the complete shift-role roster lists Operations coordinator as the sole role held by employee PK-731.", "The complete 18:00 unresolved-items register contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm. It classifies the restart check as a routine process follow-up and the pressure alarm as the only critical equipment alarm.", "The Conveyor 4 entry includes a current restart-test sheet and the next action to perform the restart check. Its owner field names Mara Ives, who is serving as Incoming shift lead.", "The Filler 2 entry includes a current pressure-trend chart, an attached valve diagnostic result, and the next action to inspect the regulator after 18:15.", "The acknowledgement log shows that the Incoming shift lead has not acknowledged the handoff by 18:00."], "request": "Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing handoff policy, request, 18:00 time binding, entities, and evaluation scope; the unchanged questions preserve the full rubric. The two focus spans are complete factual sentences. The counterfactual coherently changes PK-731’s sole role from Operations coordinator to Maintenance scheduler, making the critical alarm’s routing factually deficient without contradicting any unchanged observation. Neither context contains a readiness score, gold answer, output instruction, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731.\",\"At 18:00, the complete shift-role roster lists Operations coordinator as the sole role held by employee PK-731.\",\"The complete 18:00 unresolved-items register contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm. It classifies the restart check as a routine process follow-up and the pressure alarm as the only critical equipment alarm.\",\"The Conveyor 4 entry includes a current restart-test sheet and the next action to perform the restart check. Its owner field names Mara Ives, who is serving as Incoming shift lead.\",\"The Filler 2 entry includes a current pressure-trend chart, an attached valve diagnostic result, and the next action to inspect the regulator after 18:15.\",\"The acknowledgement log shows that the Incoming shift lead has not acknowledged the handoff by 18:00.\"],\"request\":\"Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731."}, {"path": ["evidence", "1"], "text": "At 18:00, the complete shift-role roster lists Operations coordinator as the sole role held by employee PK-731."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731.", "negative_right": "At 18:00, the complete shift-role roster lists Maintenance scheduler as the sole role held by employee PK-731.", "right": "At 18:00, the complete shift-role roster lists Operations coordinator as the sole role held by employee PK-731."}, "verifier_independent_model": false}, "family": "scale-diverse-036-003", "id": "scale-diverse-036-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["At 18:00, the Filler 2 pressure alarm's owner field displays employee identifier PK-731.", "At 18:00, the complete shift-role roster lists Maintenance scheduler as the sole role held by employee PK-731.", "The complete 18:00 unresolved-items register contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm. It classifies the restart check as a routine process follow-up and the pressure alarm as the only critical equipment alarm.", "The Conveyor 4 entry includes a current restart-test sheet and the next action to perform the restart check. Its owner field names Mara Ives, who is serving as Incoming shift lead.", "The Filler 2 entry includes a current pressure-trend chart, an attached valve diagnostic result, and the next action to inspect the regulator after 18:15.", "The acknowledgement log shows that the Incoming shift lead has not acknowledged the handoff by 18:00."], "request": "Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same readiness, confirmation, logistics, operating-hours, and written hiring-manager approval requirements, while preserving Mira Chen, the October 8 loop, and the relevant timing and decision scope. The two evidence spans are complete factual sentences. The counterfactual coherently changes only Elena Ruiz’s recorded role; the written approval can still exist without satisfying the unchanged requirement that its grantor be a hiring manager. Neither context states a classification, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"On October 7, the calendar record for Mira Chen’s proposed three-session video loop on October 8 documented the final session at 6:30–7:00 p.m. in the interviewer’s local time and flagged that timing as a possible disruption to the loop. By 8:30 a.m. on October 8, logistics for every session were complete, with conflict-free calendar holds and working video links. At 9:00 a.m., the written approval was issued before any confirmations occurred. At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor. At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a hiring-manager role. Mira confirmed the proposed loop at 2:30 p.m.; by 2:45 p.m., every scheduled interviewer had confirmed. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor."}, {"path": [], "text": "At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a hiring-manager role."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor.", "negative_left": "At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor.", "negative_right": "At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a recruiting-coordinator role and not a hiring-manager role.", "right": "At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a hiring-manager role."}, "verifier_independent_model": false}, "family": "scale-diverse-037-001", "id": "scale-diverse-037-001-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "On October 7, the calendar record for Mira Chen’s proposed three-session video loop on October 8 documented the final session at 6:30–7:00 p.m. in the interviewer’s local time and flagged that timing as a possible disruption to the loop. By 8:30 a.m. on October 8, logistics for every session were complete, with conflict-free calendar holds and working video links. At 9:00 a.m., the written approval was issued before any confirmations occurred. At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor. At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a hiring-manager role. Mira confirmed the proposed loop at 2:30 p.m.; by 2:45 p.m., every scheduled interviewer had confirmed. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same readiness, confirmation, logistics, operating-hours, and written hiring-manager approval requirements, while preserving Mira Chen, the October 8 loop, and the relevant timing and decision scope. The two evidence spans are complete factual sentences. The counterfactual coherently changes only Elena Ruiz’s recorded role; the written approval can still exist without satisfying the unchanged requirement that its grantor be a hiring manager. Neither context states a classification, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"On October 7, the calendar record for Mira Chen’s proposed three-session video loop on October 8 documented the final session at 6:30–7:00 p.m. in the interviewer’s local time and flagged that timing as a possible disruption to the loop. By 8:30 a.m. on October 8, logistics for every session were complete, with conflict-free calendar holds and working video links. At 9:00 a.m., the written approval was issued before any confirmations occurred. At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor. At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a hiring-manager role. Mira confirmed the proposed loop at 2:30 p.m.; by 2:45 p.m., every scheduled interviewer had confirmed. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor."}, {"path": [], "text": "At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a hiring-manager role."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor.", "negative_left": "At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor.", "negative_right": "At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a recruiting-coordinator role and not a hiring-manager role.", "right": "At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a hiring-manager role."}, "verifier_independent_model": false}, "family": "scale-diverse-037-001", "id": "scale-diverse-037-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "On October 7, the calendar record for Mira Chen’s proposed three-session video loop on October 8 documented the final session at 6:30–7:00 p.m. in the interviewer’s local time and flagged that timing as a possible disruption to the loop. By 8:30 a.m. on October 8, logistics for every session were complete, with conflict-free calendar holds and working video links. At 9:00 a.m., the written approval was issued before any confirmations occurred. At 2:15 p.m. on October 8, Mira Chen's loop record contained exactly one written approval for the final session's after-hours schedule, and that approval identified Elena Ruiz as its grantor. At 2:15 p.m. on October 8, the official personnel register listed Elena Ruiz as holding a recruiting-coordinator role and not a hiring-manager role. Mira confirmed the proposed loop at 2:30 p.m.; by 2:45 p.m., every scheduled interviewer had confirmed. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same readiness, scheduling-window, and written hiring-manager approval policy without adding exceptions or defaults. The candidate, October 8 loop, final-session timing, local-time basis, approval time, and relevant employee remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual changes only HC-731’s hiring-manager status at the time the approval was issued; an unauthorized person can coherently issue a document described as an approval, so this does not contradict the unchanged record. Neither context includes a classification label, answer code, rule table, proposition ID, output instruction, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Operational handoff: At 4:00 p.m. on October 7, Mira Chen confirmed the proposed three-session video loop on October 8, and every interviewer scheduled for that loop had confirmed. Logistics for every session are complete, with conflict-free calendar holds, delivered invitations, and tested video links. The calendar record documents the final session at 6:30–7:00 p.m. in the interviewer’s local time and identifies that timing as an issue that could disrupt the loop. Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731. At 2:15 p.m. on October 7, employee HC-731 held a hiring-manager role. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731."}, {"path": [], "text": "At 2:15 p.m. on October 7, employee HC-731 held a hiring-manager role."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731.", "negative_left": "Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731.", "negative_right": "At 2:15 p.m. on October 7, employee HC-731 held no hiring-manager role.", "right": "At 2:15 p.m. on October 7, employee HC-731 held a hiring-manager role."}, "verifier_independent_model": false}, "family": "scale-diverse-037-003", "id": "scale-diverse-037-003-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Operational handoff: At 4:00 p.m. on October 7, Mira Chen confirmed the proposed three-session video loop on October 8, and every interviewer scheduled for that loop had confirmed. Logistics for every session are complete, with conflict-free calendar holds, delivered invitations, and tested video links. The calendar record documents the final session at 6:30–7:00 p.m. in the interviewer’s local time and identifies that timing as an issue that could disrupt the loop. Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731. At 2:15 p.m. on October 7, employee HC-731 held a hiring-manager role. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same readiness, scheduling-window, and written hiring-manager approval policy without adding exceptions or defaults. The candidate, October 8 loop, final-session timing, local-time basis, approval time, and relevant employee remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual changes only HC-731’s hiring-manager status at the time the approval was issued; an unauthorized person can coherently issue a document described as an approval, so this does not contradict the unchanged record. Neither context includes a classification label, answer code, rule table, proposition ID, output instruction, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Operational handoff: At 4:00 p.m. on October 7, Mira Chen confirmed the proposed three-session video loop on October 8, and every interviewer scheduled for that loop had confirmed. Logistics for every session are complete, with conflict-free calendar holds, delivered invitations, and tested video links. The calendar record documents the final session at 6:30–7:00 p.m. in the interviewer’s local time and identifies that timing as an issue that could disrupt the loop. Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731. At 2:15 p.m. on October 7, employee HC-731 held a hiring-manager role. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731."}, {"path": [], "text": "At 2:15 p.m. on October 7, employee HC-731 held a hiring-manager role."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731.", "negative_left": "Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731.", "negative_right": "At 2:15 p.m. on October 7, employee HC-731 held no hiring-manager role.", "right": "At 2:15 p.m. on October 7, employee HC-731 held a hiring-manager role."}, "verifier_independent_model": false}, "family": "scale-diverse-037-003", "id": "scale-diverse-037-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Operational handoff: At 4:00 p.m. on October 7, Mira Chen confirmed the proposed three-session video loop on October 8, and every interviewer scheduled for that loop had confirmed. Logistics for every session are complete, with conflict-free calendar holds, delivered invitations, and tested video links. The calendar record documents the final session at 6:30–7:00 p.m. in the interviewer’s local time and identifies that timing as an issue that could disrupt the loop. Mira Chen's October 8 loop handoff record contains exactly one written approval for the final session's after-hours schedule, issued at 2:15 p.m. on October 7 by the person identified as employee HC-731. At 2:15 p.m. on October 7, employee HC-731 held no hiring-manager role. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing readiness and after-hours approval policy, while the unchanged questions preserve all classification and routing criteria. The candidate, October 8 loop, final-session timing, approval identity, and issuance-time role path remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes only Lila Navarro’s active role from hiring manager to recruiting coordinator, making the described writing insufficient as hiring-manager approval without creating a contradictory duplicate role assertion. Neither context states a classification, answer code, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Field note: Mira Chen confirmed the proposed three-session video loop on October 8, and every scheduled interviewer confirmed. Calendar holds, invitations, and working video links show that logistics are complete for every session. The final session runs from 6:30 to 7:00 p.m. in the interviewer’s local time. In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor. At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: hiring manager. The approval was issued October 7, before loop confirmation was recorded. The calendar record documents the final session’s after-hours timing, which the coordinator notes could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor."}, {"path": [], "text": "At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: hiring manager."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor.", "negative_left": "In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor.", "negative_right": "At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: recruiting coordinator.", "right": "At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: hiring manager."}, "verifier_independent_model": false}, "family": "scale-diverse-037-004", "id": "scale-diverse-037-004-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Field note: Mira Chen confirmed the proposed three-session video loop on October 8, and every scheduled interviewer confirmed. Calendar holds, invitations, and working video links show that logistics are complete for every session. The final session runs from 6:30 to 7:00 p.m. in the interviewer’s local time. In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor. At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: hiring manager. The approval was issued October 7, before loop confirmation was recorded. The calendar record documents the final session’s after-hours timing, which the coordinator notes could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing readiness and after-hours approval policy, while the unchanged questions preserve all classification and routing criteria. The candidate, October 8 loop, final-session timing, approval identity, and issuance-time role path remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes only Lila Navarro’s active role from hiring manager to recruiting coordinator, making the described writing insufficient as hiring-manager approval without creating a contradictory duplicate role assertion. Neither context states a classification, answer code, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Field note: Mira Chen confirmed the proposed three-session video loop on October 8, and every scheduled interviewer confirmed. Calendar holds, invitations, and working video links show that logistics are complete for every session. The final session runs from 6:30 to 7:00 p.m. in the interviewer’s local time. In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor. At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: hiring manager. The approval was issued October 7, before loop confirmation was recorded. The calendar record documents the final session’s after-hours timing, which the coordinator notes could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor."}, {"path": [], "text": "At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: hiring manager."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor.", "negative_left": "In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor.", "negative_right": "At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: recruiting coordinator.", "right": "At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: hiring manager."}, "verifier_independent_model": false}, "family": "scale-diverse-037-004", "id": "scale-diverse-037-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Field note: Mira Chen confirmed the proposed three-session video loop on October 8, and every scheduled interviewer confirmed. Calendar holds, invitations, and working video links show that logistics are complete for every session. The final session runs from 6:30 to 7:00 p.m. in the interviewer’s local time. In Mira Chen's October 8 loop record, the sole written approval for the final session's after-hours schedule identifies employee 5842, Lila Navarro, as its grantor. At the approval's issuance time, the personnel roster lists employee 5842, Lila Navarro, with exactly one active role: recruiting coordinator. The approval was issued October 7, before loop confirmation was recorded. The calendar record documents the final session’s after-hours timing, which the coordinator notes could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the required-seat and substitution-approval policy without adding exceptions, priorities, or missing-evidence defaults, and they retain the same candidate, October 8, 2026 loop, request scope, and relevant participant/seat bindings. The two focus spans are complete factual sentences; changing T-731 from “accepted” to “declined” is coherent with the unchanged evidence because record completeness, substitution approval, and the candidate’s schedule confirmation do not assert that Lena accepted, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified record and approval atoms remain single relations rather than bundled classifications. The focus atom concerns Lena’s acceptance of a particular seat, not policy. The base and counter assignments can differ only in Lena’s acceptance: a substitution may be approved and designated even though the replacement interviewer does not accept, and record presence does not imply acceptance. Policy evidence preserves the state-originating required-seat definition and hiring-manager-approval constraint; the unchanged questions already preserve all decision criteria and priorities. The extra request citation is unnecessary but does not omit or distort any needed state policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all required records are present, the candidate and all three required seats are confirmed, the substitution was approved and occurred, and no loop-structure change occurred. With higher-priority hold conditions excluded, the approved substitution is sufficient for moderate-risk confirmation.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes missing records, missing substitution approval, unresolved candidate availability, and unapproved loop-structure changes, while explicitly refuting acceptance by Lena, the substituted interviewer occupying the required technical seat. The priority policy therefore routes the hold to that interviewer.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present."}, {"id": "a2", "statement": "Collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00."}, {"id": "a3", "statement": "Hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00."}, {"id": "a4", "statement": "Replacement interviewer Lena accepted the required October 8 technical seat at 3:00."}, {"id": "a5", "statement": "Candidate Avery Cole confirmed the complete October 8 interview schedule."}, {"id": "a6", "statement": "Hiring manager Victor approved Lena’s substitution for Omar in the October 8 technical seat."}, {"id": "a7", "statement": "Lena substituted for Omar in the October 8 technical seat."}, {"id": "a8", "statement": "Every interviewer substitution in Avery Cole’s October 8 interview loop has hiring-manager approval."}, {"id": "a9", "statement": "No loop-structure change occurred in Avery Cole’s October 8 interview loop."}], "base_state_json": "{\"context\":\"The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.\",\"evidence\":[\"At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “accepted” for its 3:00 appointment on October 8, 2026.\",\"In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop.\",\"A reconciliation audit found every required calendar and confirmation record for Avery’s loop present.\",\"The ledger records collaboration interviewer Rosa’s acceptance of the October 8 collaboration seat at 1:00 and hiring manager Victor’s acceptance of the hiring-manager-review seat at 4:00.\",\"Avery confirmed the complete October 8 schedule.\",\"The substitution log states that Lena replaced Omar in the October 8 technical seat and that Victor approved this replacement.\",\"The completed approval audit shows that every interviewer substitution in the loop has hiring-manager approval.\",\"The structure comparison found no change to the loop’s required-seat structure.\"],\"request\":\"Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “accepted” for its 3:00 appointment on October 8, 2026."}, {"path": ["evidence", "1"], "text": "In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop."}], "policy_evidence": [{"path": ["context"], "text": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval."}, {"path": ["request"], "text": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}], "rules": [{"justification": "All required records are present, the candidate and all three required seats are confirmed, and the technical-seat substitution occurred with hiring-manager approval. The approved substitution makes the applicable confirmation outcome moderate risk.", "target": "confirm_moderate_risk", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "All required records and approvals are present and candidate availability is resolved, but the required technical seat is explicitly unaccepted. The priority policy therefore routes follow-up to that interviewer.", "target": "hold_route_interviewer", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “accepted” for its 3:00 appointment on October 8, 2026.", "negative_left": "At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “declined” for its 3:00 appointment on October 8, 2026.", "negative_right": "In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop.", "right": "In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop."}, "verifier_independent_model": false}, "family": "scale-diverse-038-002", "id": "scale-diverse-038-002-base", "input": {"questions": {"decision": {"criteria": {"confirm_low_risk": "Confirm now at low risk: every required seat and the candidate are confirmed, with no substitution or date change.", "confirm_moderate_risk": "Confirm now at moderate risk: every required seat and the candidate are confirmed, but an approved interviewer substitution or date change occurred.", "hold_route_candidate": "Do not confirm; scheduling risk is high because candidate availability remains unresolved after interviewer records are complete. Route to the candidate.", "hold_route_coordinator": "Do not confirm; scheduling risk is high because a required calendar or confirmation record is missing. Route to the recruiting coordinator.", "hold_route_hiring_manager": "Do not confirm; scheduling risk is high because a required substitution or loop-structure change lacks hiring-manager approval. Route to the hiring manager.", "hold_route_interviewer": "Do not confirm; scheduling risk is high because at least one required interviewer seat remains declined, unanswered, or otherwise unaccepted. Route to that interviewer."}, "instructions": "Apply this priority order: missing calendar or confirmation records routes to the recruiting coordinator; missing required substitution approval routes to the hiring manager; unresolved candidate availability routes to the candidate; an unaccepted required seat routes to that interviewer. Otherwise, confirm at moderate risk if an approved substitution or date change occurred, and at low risk if neither occurred. Select exactly one option.", "type": "choice"}}, "state": {"context": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.", "evidence": ["At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “accepted” for its 3:00 appointment on October 8, 2026.", "In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop.", "A reconciliation audit found every required calendar and confirmation record for Avery’s loop present.", "The ledger records collaboration interviewer Rosa’s acceptance of the October 8 collaboration seat at 1:00 and hiring manager Victor’s acceptance of the hiring-manager-review seat at 4:00.", "Avery confirmed the complete October 8 schedule.", "The substitution log states that Lena replaced Omar in the October 8 technical seat and that Victor approved this replacement.", "The completed approval audit shows that every interviewer substitution in the loop has hiring-manager approval.", "The structure comparison found no change to the loop’s required-seat structure."], "request": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}}, "method": "c2d", "provenance": {"source_id": "diverse-038", "source_is_synthetic": true, "source_sha256": "a38bfb9885128ee8923051462ea69e5f86cf88c8bf143e3228a1b6a65164072e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "confirm_moderate_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the required-seat and substitution-approval policy without adding exceptions, priorities, or missing-evidence defaults, and they retain the same candidate, October 8, 2026 loop, request scope, and relevant participant/seat bindings. The two focus spans are complete factual sentences; changing T-731 from “accepted” to “declined” is coherent with the unchanged evidence because record completeness, substitution approval, and the candidate’s schedule confirmation do not assert that Lena accepted, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified record and approval atoms remain single relations rather than bundled classifications. The focus atom concerns Lena’s acceptance of a particular seat, not policy. The base and counter assignments can differ only in Lena’s acceptance: a substitution may be approved and designated even though the replacement interviewer does not accept, and record presence does not imply acceptance. Policy evidence preserves the state-originating required-seat definition and hiring-manager-approval constraint; the unchanged questions already preserve all decision criteria and priorities. The extra request citation is unnecessary but does not omit or distort any needed state policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all required records are present, the candidate and all three required seats are confirmed, the substitution was approved and occurred, and no loop-structure change occurred. With higher-priority hold conditions excluded, the approved substitution is sufficient for moderate-risk confirmation.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes missing records, missing substitution approval, unresolved candidate availability, and unapproved loop-structure changes, while explicitly refuting acceptance by Lena, the substituted interviewer occupying the required technical seat. The priority policy therefore routes the hold to that interviewer.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present."}, {"id": "a2", "statement": "Collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00."}, {"id": "a3", "statement": "Hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00."}, {"id": "a4", "statement": "Replacement interviewer Lena accepted the required October 8 technical seat at 3:00."}, {"id": "a5", "statement": "Candidate Avery Cole confirmed the complete October 8 interview schedule."}, {"id": "a6", "statement": "Hiring manager Victor approved Lena’s substitution for Omar in the October 8 technical seat."}, {"id": "a7", "statement": "Lena substituted for Omar in the October 8 technical seat."}, {"id": "a8", "statement": "Every interviewer substitution in Avery Cole’s October 8 interview loop has hiring-manager approval."}, {"id": "a9", "statement": "No loop-structure change occurred in Avery Cole’s October 8 interview loop."}], "base_state_json": "{\"context\":\"The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.\",\"evidence\":[\"At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “accepted” for its 3:00 appointment on October 8, 2026.\",\"In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop.\",\"A reconciliation audit found every required calendar and confirmation record for Avery’s loop present.\",\"The ledger records collaboration interviewer Rosa’s acceptance of the October 8 collaboration seat at 1:00 and hiring manager Victor’s acceptance of the hiring-manager-review seat at 4:00.\",\"Avery confirmed the complete October 8 schedule.\",\"The substitution log states that Lena replaced Omar in the October 8 technical seat and that Victor approved this replacement.\",\"The completed approval audit shows that every interviewer substitution in the loop has hiring-manager approval.\",\"The structure comparison found no change to the loop’s required-seat structure.\"],\"request\":\"Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “accepted” for its 3:00 appointment on October 8, 2026."}, {"path": ["evidence", "1"], "text": "In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop."}], "policy_evidence": [{"path": ["context"], "text": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval."}, {"path": ["request"], "text": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}], "rules": [{"justification": "All required records are present, the candidate and all three required seats are confirmed, and the technical-seat substitution occurred with hiring-manager approval. The approved substitution makes the applicable confirmation outcome moderate risk.", "target": "confirm_moderate_risk", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "All required records and approvals are present and candidate availability is resolved, but the required technical seat is explicitly unaccepted. The priority policy therefore routes follow-up to that interviewer.", "target": "hold_route_interviewer", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “accepted” for its 3:00 appointment on October 8, 2026.", "negative_left": "At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “declined” for its 3:00 appointment on October 8, 2026.", "negative_right": "In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop.", "right": "In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop."}, "verifier_independent_model": false}, "family": "scale-diverse-038-002", "id": "scale-diverse-038-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"confirm_low_risk": "Confirm now at low risk: every required seat and the candidate are confirmed, with no substitution or date change.", "confirm_moderate_risk": "Confirm now at moderate risk: every required seat and the candidate are confirmed, but an approved interviewer substitution or date change occurred.", "hold_route_candidate": "Do not confirm; scheduling risk is high because candidate availability remains unresolved after interviewer records are complete. Route to the candidate.", "hold_route_coordinator": "Do not confirm; scheduling risk is high because a required calendar or confirmation record is missing. Route to the recruiting coordinator.", "hold_route_hiring_manager": "Do not confirm; scheduling risk is high because a required substitution or loop-structure change lacks hiring-manager approval. Route to the hiring manager.", "hold_route_interviewer": "Do not confirm; scheduling risk is high because at least one required interviewer seat remains declined, unanswered, or otherwise unaccepted. Route to that interviewer."}, "instructions": "Apply this priority order: missing calendar or confirmation records routes to the recruiting coordinator; missing required substitution approval routes to the hiring manager; unresolved candidate availability routes to the candidate; an unaccepted required seat routes to that interviewer. Otherwise, confirm at moderate risk if an approved substitution or date change occurred, and at low risk if neither occurred. Select exactly one option.", "type": "choice"}}, "state": {"context": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.", "evidence": ["At 09:12 on October 7, 2026, the scheduling ledger showed confirmation record T-731 with the status “declined” for its 3:00 appointment on October 8, 2026.", "In the scheduling ledger, confirmation record T-731 uniquely belongs to replacement interviewer Lena and covers the required technical seat in candidate Avery Cole’s October 8, 2026 interview loop.", "A reconciliation audit found every required calendar and confirmation record for Avery’s loop present.", "The ledger records collaboration interviewer Rosa’s acceptance of the October 8 collaboration seat at 1:00 and hiring manager Victor’s acceptance of the hiring-manager-review seat at 4:00.", "Avery confirmed the complete October 8 schedule.", "The substitution log states that Lena replaced Omar in the October 8 technical seat and that Victor approved this replacement.", "The completed approval audit shows that every interviewer substitution in the loop has hiring-manager approval.", "The structure comparison found no change to the loop’s required-seat structure."], "request": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}}, "method": "c2d", "provenance": {"source_id": "diverse-038", "source_is_synthetic": true, "source_sha256": "a38bfb9885128ee8923051462ea69e5f86cf88c8bf143e3228a1b6a65164072e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "hold_route_interviewer"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the required-seat and substitution-approval policy without adding exceptions or defaults, and they preserve Avery Cole, the October 8 loop, and the original request scope. The two focus spans are complete factual sentences. Changing Lena’s sole response from “Accept” to “Decline” is coherent with the unchanged facts: the invitation, approved substitution, complete records, and candidate confirmation do not assert that Lena accepted. Neither context embeds a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified record and approval atoms remain single relations rather than bundled classifications. The focus atom concerns Lena’s acceptance of a particular seat, not policy. The base and counter assignments can differ only in Lena’s acceptance: a substitution may be approved and designated even though the replacement interviewer does not accept, and record presence does not imply acceptance. Policy evidence preserves the state-originating required-seat definition and hiring-manager-approval constraint; the unchanged questions already preserve all decision criteria and priorities. The extra request citation is unnecessary but does not omit or distort any needed state policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all required records are present, the candidate and all three required seats are confirmed, the substitution was approved and occurred, and no loop-structure change occurred. With higher-priority hold conditions excluded, the approved substitution is sufficient for moderate-risk confirmation.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes missing records, missing substitution approval, unresolved candidate availability, and unapproved loop-structure changes, while explicitly refuting acceptance by Lena, the substituted interviewer occupying the required technical seat. The priority policy therefore routes the hold to that interviewer.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present."}, {"id": "a2", "statement": "Collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00."}, {"id": "a3", "statement": "Hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00."}, {"id": "a4", "statement": "Replacement interviewer Lena accepted the required October 8 technical seat at 3:00."}, {"id": "a5", "statement": "Candidate Avery Cole confirmed the complete October 8 interview schedule."}, {"id": "a6", "statement": "Hiring manager Victor approved Lena’s substitution for Omar in the October 8 technical seat."}, {"id": "a7", "statement": "Lena substituted for Omar in the October 8 technical seat."}, {"id": "a8", "statement": "Every interviewer substitution in Avery Cole’s October 8 interview loop has hiring-manager approval."}, {"id": "a9", "statement": "No loop-structure change occurred in Avery Cole’s October 8 interview loop."}], "base_state_json": "{\"context\":\"The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.\",\"evidence\":[\"Operational handoff audit: every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present.\",\"Calendar responses show collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00.\",\"Calendar responses show hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00.\",\"Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00.\",\"Lena’s only response to invitation H-3008 was the selection “Accept” at 9:12 on October 6.\",\"The substitution log records that Lena replaced Omar in the October 8 technical seat, with Victor’s approval.\",\"Compliance review found this was the loop’s only interviewer substitution and that every interviewer substitution had hiring-manager approval.\",\"No loop-structure change occurred in Avery Cole’s October 8 interview loop.\",\"Candidate confirmation: Avery Cole confirmed the complete October 8 interview schedule.\"],\"request\":\"Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "3"], "text": "Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00."}, {"path": ["evidence", "4"], "text": "Lena’s only response to invitation H-3008 was the selection “Accept” at 9:12 on October 6."}], "policy_evidence": [{"path": ["context"], "text": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval."}, {"path": ["request"], "text": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}], "rules": [{"justification": "All required records are present, the candidate and all three required seats are confirmed, and the technical-seat substitution occurred with hiring-manager approval. The approved substitution makes the applicable confirmation outcome moderate risk.", "target": "confirm_moderate_risk", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "All required records and approvals are present and candidate availability is resolved, but the required technical seat is explicitly unaccepted. The priority policy therefore routes follow-up to that interviewer.", "target": "hold_route_interviewer", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00.", "negative_left": "Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00.", "negative_right": "Lena’s only response to invitation H-3008 was the selection “Decline” at 9:12 on October 6.", "right": "Lena’s only response to invitation H-3008 was the selection “Accept” at 9:12 on October 6."}, "verifier_independent_model": false}, "family": "scale-diverse-038-003", "id": "scale-diverse-038-003-base", "input": {"questions": {"decision": {"criteria": {"confirm_low_risk": "Confirm now at low risk: every required seat and the candidate are confirmed, with no substitution or date change.", "confirm_moderate_risk": "Confirm now at moderate risk: every required seat and the candidate are confirmed, but an approved interviewer substitution or date change occurred.", "hold_route_candidate": "Do not confirm; scheduling risk is high because candidate availability remains unresolved after interviewer records are complete. Route to the candidate.", "hold_route_coordinator": "Do not confirm; scheduling risk is high because a required calendar or confirmation record is missing. Route to the recruiting coordinator.", "hold_route_hiring_manager": "Do not confirm; scheduling risk is high because a required substitution or loop-structure change lacks hiring-manager approval. Route to the hiring manager.", "hold_route_interviewer": "Do not confirm; scheduling risk is high because at least one required interviewer seat remains declined, unanswered, or otherwise unaccepted. Route to that interviewer."}, "instructions": "Apply this priority order: missing calendar or confirmation records routes to the recruiting coordinator; missing required substitution approval routes to the hiring manager; unresolved candidate availability routes to the candidate; an unaccepted required seat routes to that interviewer. Otherwise, confirm at moderate risk if an approved substitution or date change occurred, and at low risk if neither occurred. Select exactly one option.", "type": "choice"}}, "state": {"context": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.", "evidence": ["Operational handoff audit: every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present.", "Calendar responses show collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00.", "Calendar responses show hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00.", "Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00.", "Lena’s only response to invitation H-3008 was the selection “Accept” at 9:12 on October 6.", "The substitution log records that Lena replaced Omar in the October 8 technical seat, with Victor’s approval.", "Compliance review found this was the loop’s only interviewer substitution and that every interviewer substitution had hiring-manager approval.", "No loop-structure change occurred in Avery Cole’s October 8 interview loop.", "Candidate confirmation: Avery Cole confirmed the complete October 8 interview schedule."], "request": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}}, "method": "c2d", "provenance": {"source_id": "diverse-038", "source_is_synthetic": true, "source_sha256": "a38bfb9885128ee8923051462ea69e5f86cf88c8bf143e3228a1b6a65164072e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "confirm_moderate_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the required-seat and substitution-approval policy without adding exceptions or defaults, and they preserve Avery Cole, the October 8 loop, and the original request scope. The two focus spans are complete factual sentences. Changing Lena’s sole response from “Accept” to “Decline” is coherent with the unchanged facts: the invitation, approved substitution, complete records, and candidate confirmation do not assert that Lena accepted. Neither context embeds a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified record and approval atoms remain single relations rather than bundled classifications. The focus atom concerns Lena’s acceptance of a particular seat, not policy. The base and counter assignments can differ only in Lena’s acceptance: a substitution may be approved and designated even though the replacement interviewer does not accept, and record presence does not imply acceptance. Policy evidence preserves the state-originating required-seat definition and hiring-manager-approval constraint; the unchanged questions already preserve all decision criteria and priorities. The extra request citation is unnecessary but does not omit or distort any needed state policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all required records are present, the candidate and all three required seats are confirmed, the substitution was approved and occurred, and no loop-structure change occurred. With higher-priority hold conditions excluded, the approved substitution is sufficient for moderate-risk confirmation.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes missing records, missing substitution approval, unresolved candidate availability, and unapproved loop-structure changes, while explicitly refuting acceptance by Lena, the substituted interviewer occupying the required technical seat. The priority policy therefore routes the hold to that interviewer.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present."}, {"id": "a2", "statement": "Collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00."}, {"id": "a3", "statement": "Hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00."}, {"id": "a4", "statement": "Replacement interviewer Lena accepted the required October 8 technical seat at 3:00."}, {"id": "a5", "statement": "Candidate Avery Cole confirmed the complete October 8 interview schedule."}, {"id": "a6", "statement": "Hiring manager Victor approved Lena’s substitution for Omar in the October 8 technical seat."}, {"id": "a7", "statement": "Lena substituted for Omar in the October 8 technical seat."}, {"id": "a8", "statement": "Every interviewer substitution in Avery Cole’s October 8 interview loop has hiring-manager approval."}, {"id": "a9", "statement": "No loop-structure change occurred in Avery Cole’s October 8 interview loop."}], "base_state_json": "{\"context\":\"The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.\",\"evidence\":[\"Operational handoff audit: every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present.\",\"Calendar responses show collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00.\",\"Calendar responses show hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00.\",\"Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00.\",\"Lena’s only response to invitation H-3008 was the selection “Accept” at 9:12 on October 6.\",\"The substitution log records that Lena replaced Omar in the October 8 technical seat, with Victor’s approval.\",\"Compliance review found this was the loop’s only interviewer substitution and that every interviewer substitution had hiring-manager approval.\",\"No loop-structure change occurred in Avery Cole’s October 8 interview loop.\",\"Candidate confirmation: Avery Cole confirmed the complete October 8 interview schedule.\"],\"request\":\"Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "3"], "text": "Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00."}, {"path": ["evidence", "4"], "text": "Lena’s only response to invitation H-3008 was the selection “Accept” at 9:12 on October 6."}], "policy_evidence": [{"path": ["context"], "text": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval."}, {"path": ["request"], "text": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}], "rules": [{"justification": "All required records are present, the candidate and all three required seats are confirmed, and the technical-seat substitution occurred with hiring-manager approval. The approved substitution makes the applicable confirmation outcome moderate risk.", "target": "confirm_moderate_risk", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "All required records and approvals are present and candidate availability is resolved, but the required technical seat is explicitly unaccepted. The priority policy therefore routes follow-up to that interviewer.", "target": "hold_route_interviewer", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00.", "negative_left": "Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00.", "negative_right": "Lena’s only response to invitation H-3008 was the selection “Decline” at 9:12 on October 6.", "right": "Lena’s only response to invitation H-3008 was the selection “Accept” at 9:12 on October 6."}, "verifier_independent_model": false}, "family": "scale-diverse-038-003", "id": "scale-diverse-038-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"confirm_low_risk": "Confirm now at low risk: every required seat and the candidate are confirmed, with no substitution or date change.", "confirm_moderate_risk": "Confirm now at moderate risk: every required seat and the candidate are confirmed, but an approved interviewer substitution or date change occurred.", "hold_route_candidate": "Do not confirm; scheduling risk is high because candidate availability remains unresolved after interviewer records are complete. Route to the candidate.", "hold_route_coordinator": "Do not confirm; scheduling risk is high because a required calendar or confirmation record is missing. Route to the recruiting coordinator.", "hold_route_hiring_manager": "Do not confirm; scheduling risk is high because a required substitution or loop-structure change lacks hiring-manager approval. Route to the hiring manager.", "hold_route_interviewer": "Do not confirm; scheduling risk is high because at least one required interviewer seat remains declined, unanswered, or otherwise unaccepted. Route to that interviewer."}, "instructions": "Apply this priority order: missing calendar or confirmation records routes to the recruiting coordinator; missing required substitution approval routes to the hiring manager; unresolved candidate availability routes to the candidate; an unaccepted required seat routes to that interviewer. Otherwise, confirm at moderate risk if an approved substitution or date change occurred, and at low risk if neither occurred. Select exactly one option.", "type": "choice"}}, "state": {"context": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.", "evidence": ["Operational handoff audit: every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present.", "Calendar responses show collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00.", "Calendar responses show hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00.", "Invitation H-3008 in candidate Avery Cole’s interview loop was for replacement interviewer Lena’s required October 8 technical seat at 3:00.", "Lena’s only response to invitation H-3008 was the selection “Decline” at 9:12 on October 6.", "The substitution log records that Lena replaced Omar in the October 8 technical seat, with Victor’s approval.", "Compliance review found this was the loop’s only interviewer substitution and that every interviewer substitution had hiring-manager approval.", "No loop-structure change occurred in Avery Cole’s October 8 interview loop.", "Candidate confirmation: Avery Cole confirmed the complete October 8 interview schedule."], "request": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}}, "method": "c2d", "provenance": {"source_id": "diverse-038", "source_is_synthetic": true, "source_sha256": "a38bfb9885128ee8923051462ea69e5f86cf88c8bf143e3228a1b6a65164072e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "hold_route_interviewer"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing seat and substitution-approval policy, while the unchanged questions preserve the full decision criteria and priority order. The request, candidate, October 8 loop, required technical-seat event, time, and replacement interviewer remain bound consistently. The two focus spans are complete factual sentences. Changing Lena’s logged response from accepted to declined is a coherent single observational change: the audit states that records are present, not that every response is accepted, and Avery’s schedule confirmation does not contradict Lena’s later-described declined status. Neither context contains an answer code, prescribed classifier output, rule table, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified record and approval atoms remain single relations rather than bundled classifications. The focus atom concerns Lena’s acceptance of a particular seat, not policy. The base and counter assignments can differ only in Lena’s acceptance: a substitution may be approved and designated even though the replacement interviewer does not accept, and record presence does not imply acceptance. Policy evidence preserves the state-originating required-seat definition and hiring-manager-approval constraint; the unchanged questions already preserve all decision criteria and priorities. The extra request citation is unnecessary but does not omit or distort any needed state policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all required records are present, the candidate and all three required seats are confirmed, the substitution was approved and occurred, and no loop-structure change occurred. With higher-priority hold conditions excluded, the approved substitution is sufficient for moderate-risk confirmation.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes missing records, missing substitution approval, unresolved candidate availability, and unapproved loop-structure changes, while explicitly refuting acceptance by Lena, the substituted interviewer occupying the required technical seat. The priority policy therefore routes the hold to that interviewer.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present."}, {"id": "a2", "statement": "Collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00."}, {"id": "a3", "statement": "Hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00."}, {"id": "a4", "statement": "Replacement interviewer Lena accepted the required October 8 technical seat at 3:00."}, {"id": "a5", "statement": "Candidate Avery Cole confirmed the complete October 8 interview schedule."}, {"id": "a6", "statement": "Hiring manager Victor approved Lena’s substitution for Omar in the October 8 technical seat."}, {"id": "a7", "statement": "Lena substituted for Omar in the October 8 technical seat."}, {"id": "a8", "statement": "Every interviewer substitution in Avery Cole’s October 8 interview loop has hiring-manager approval."}, {"id": "a9", "statement": "No loop-structure change occurred in Avery Cole’s October 8 interview loop."}], "base_state_json": "{\"context\":\"The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.\",\"evidence\":[\"Coordinator Nia’s final audit found every required calendar and confirmation record present for Avery Cole’s October 8 interview loop.\",\"Rosa’s calendar response shows acceptance of the October 8 collaboration seat at 1:00.\",\"Victor’s calendar response shows acceptance of the October 8 hiring-manager-review seat at 4:00.\",\"In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer.\",\"At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as accepted.\",\"The Lena-for-Omar assignment is the loop’s only interviewer substitution, and hiring manager Victor approved it.\",\"Avery confirmed the complete October 8 interview schedule after the substitution was recorded.\",\"The loop remained on October 8, and no change was made to its three-seat structure.\"],\"request\":\"Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "3"], "text": "In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer."}, {"path": ["evidence", "4"], "text": "At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as accepted."}], "policy_evidence": [{"path": ["context"], "text": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval."}, {"path": ["request"], "text": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}], "rules": [{"justification": "All required records are present, the candidate and all three required seats are confirmed, and the technical-seat substitution occurred with hiring-manager approval. The approved substitution makes the applicable confirmation outcome moderate risk.", "target": "confirm_moderate_risk", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "All required records and approvals are present and candidate availability is resolved, but the required technical seat is explicitly unaccepted. The priority policy therefore routes follow-up to that interviewer.", "target": "hold_route_interviewer", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer.", "negative_left": "In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer.", "negative_right": "At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as declined.", "right": "At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as accepted."}, "verifier_independent_model": false}, "family": "scale-diverse-038-004", "id": "scale-diverse-038-004-base", "input": {"questions": {"decision": {"criteria": {"confirm_low_risk": "Confirm now at low risk: every required seat and the candidate are confirmed, with no substitution or date change.", "confirm_moderate_risk": "Confirm now at moderate risk: every required seat and the candidate are confirmed, but an approved interviewer substitution or date change occurred.", "hold_route_candidate": "Do not confirm; scheduling risk is high because candidate availability remains unresolved after interviewer records are complete. Route to the candidate.", "hold_route_coordinator": "Do not confirm; scheduling risk is high because a required calendar or confirmation record is missing. Route to the recruiting coordinator.", "hold_route_hiring_manager": "Do not confirm; scheduling risk is high because a required substitution or loop-structure change lacks hiring-manager approval. Route to the hiring manager.", "hold_route_interviewer": "Do not confirm; scheduling risk is high because at least one required interviewer seat remains declined, unanswered, or otherwise unaccepted. Route to that interviewer."}, "instructions": "Apply this priority order: missing calendar or confirmation records routes to the recruiting coordinator; missing required substitution approval routes to the hiring manager; unresolved candidate availability routes to the candidate; an unaccepted required seat routes to that interviewer. Otherwise, confirm at moderate risk if an approved substitution or date change occurred, and at low risk if neither occurred. Select exactly one option.", "type": "choice"}}, "state": {"context": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.", "evidence": ["Coordinator Nia’s final audit found every required calendar and confirmation record present for Avery Cole’s October 8 interview loop.", "Rosa’s calendar response shows acceptance of the October 8 collaboration seat at 1:00.", "Victor’s calendar response shows acceptance of the October 8 hiring-manager-review seat at 4:00.", "In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer.", "At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as accepted.", "The Lena-for-Omar assignment is the loop’s only interviewer substitution, and hiring manager Victor approved it.", "Avery confirmed the complete October 8 interview schedule after the substitution was recorded.", "The loop remained on October 8, and no change was made to its three-seat structure."], "request": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}}, "method": "c2d", "provenance": {"source_id": "diverse-038", "source_is_synthetic": true, "source_sha256": "a38bfb9885128ee8923051462ea69e5f86cf88c8bf143e3228a1b6a65164072e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "confirm_moderate_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing seat and substitution-approval policy, while the unchanged questions preserve the full decision criteria and priority order. The request, candidate, October 8 loop, required technical-seat event, time, and replacement interviewer remain bound consistently. The two focus spans are complete factual sentences. Changing Lena’s logged response from accepted to declined is a coherent single observational change: the audit states that records are present, not that every response is accepted, and Avery’s schedule confirmation does not contradict Lena’s later-described declined status. Neither context contains an answer code, prescribed classifier output, rule table, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified record and approval atoms remain single relations rather than bundled classifications. The focus atom concerns Lena’s acceptance of a particular seat, not policy. The base and counter assignments can differ only in Lena’s acceptance: a substitution may be approved and designated even though the replacement interviewer does not accept, and record presence does not imply acceptance. Policy evidence preserves the state-originating required-seat definition and hiring-manager-approval constraint; the unchanged questions already preserve all decision criteria and priorities. The extra request citation is unnecessary but does not omit or distort any needed state policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all required records are present, the candidate and all three required seats are confirmed, the substitution was approved and occurred, and no loop-structure change occurred. With higher-priority hold conditions excluded, the approved substitution is sufficient for moderate-risk confirmation.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes missing records, missing substitution approval, unresolved candidate availability, and unapproved loop-structure changes, while explicitly refuting acceptance by Lena, the substituted interviewer occupying the required technical seat. The priority policy therefore routes the hold to that interviewer.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every required calendar or confirmation record for candidate Avery Cole’s October 8 interview loop is present."}, {"id": "a2", "statement": "Collaboration interviewer Rosa accepted the October 8 collaboration seat at 1:00."}, {"id": "a3", "statement": "Hiring manager Victor accepted the October 8 hiring-manager-review seat at 4:00."}, {"id": "a4", "statement": "Replacement interviewer Lena accepted the required October 8 technical seat at 3:00."}, {"id": "a5", "statement": "Candidate Avery Cole confirmed the complete October 8 interview schedule."}, {"id": "a6", "statement": "Hiring manager Victor approved Lena’s substitution for Omar in the October 8 technical seat."}, {"id": "a7", "statement": "Lena substituted for Omar in the October 8 technical seat."}, {"id": "a8", "statement": "Every interviewer substitution in Avery Cole’s October 8 interview loop has hiring-manager approval."}, {"id": "a9", "statement": "No loop-structure change occurred in Avery Cole’s October 8 interview loop."}], "base_state_json": "{\"context\":\"The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.\",\"evidence\":[\"Coordinator Nia’s final audit found every required calendar and confirmation record present for Avery Cole’s October 8 interview loop.\",\"Rosa’s calendar response shows acceptance of the October 8 collaboration seat at 1:00.\",\"Victor’s calendar response shows acceptance of the October 8 hiring-manager-review seat at 4:00.\",\"In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer.\",\"At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as accepted.\",\"The Lena-for-Omar assignment is the loop’s only interviewer substitution, and hiring manager Victor approved it.\",\"Avery confirmed the complete October 8 interview schedule after the substitution was recorded.\",\"The loop remained on October 8, and no change was made to its three-seat structure.\"],\"request\":\"Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "3"], "text": "In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer."}, {"path": ["evidence", "4"], "text": "At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as accepted."}], "policy_evidence": [{"path": ["context"], "text": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval."}, {"path": ["request"], "text": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}], "rules": [{"justification": "All required records are present, the candidate and all three required seats are confirmed, and the technical-seat substitution occurred with hiring-manager approval. The approved substitution makes the applicable confirmation outcome moderate risk.", "target": "confirm_moderate_risk", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "All required records and approvals are present and candidate availability is resolved, but the required technical seat is explicitly unaccepted. The priority policy therefore routes follow-up to that interviewer.", "target": "hold_route_interviewer", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer.", "negative_left": "In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer.", "negative_right": "At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as declined.", "right": "At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as accepted."}, "verifier_independent_model": false}, "family": "scale-diverse-038-004", "id": "scale-diverse-038-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"confirm_low_risk": "Confirm now at low risk: every required seat and the candidate are confirmed, with no substitution or date change.", "confirm_moderate_risk": "Confirm now at moderate risk: every required seat and the candidate are confirmed, but an approved interviewer substitution or date change occurred.", "hold_route_candidate": "Do not confirm; scheduling risk is high because candidate availability remains unresolved after interviewer records are complete. Route to the candidate.", "hold_route_coordinator": "Do not confirm; scheduling risk is high because a required calendar or confirmation record is missing. Route to the recruiting coordinator.", "hold_route_hiring_manager": "Do not confirm; scheduling risk is high because a required substitution or loop-structure change lacks hiring-manager approval. Route to the hiring manager.", "hold_route_interviewer": "Do not confirm; scheduling risk is high because at least one required interviewer seat remains declined, unanswered, or otherwise unaccepted. Route to that interviewer."}, "instructions": "Apply this priority order: missing calendar or confirmation records routes to the recruiting coordinator; missing required substitution approval routes to the hiring manager; unresolved candidate availability routes to the candidate; an unaccepted required seat routes to that interviewer. Otherwise, confirm at moderate risk if an approved substitution or date change occurred, and at low risk if neither occurred. Select exactly one option.", "type": "choice"}}, "state": {"context": "The required seats are technical, collaboration, and hiring-manager review. Substitutes are allowed only with hiring-manager approval.", "evidence": ["Coordinator Nia’s final audit found every required calendar and confirmation record present for Avery Cole’s October 8 interview loop.", "Rosa’s calendar response shows acceptance of the October 8 collaboration seat at 1:00.", "Victor’s calendar response shows acceptance of the October 8 hiring-manager-review seat at 4:00.", "In Avery Cole’s interview ledger, event Q-47 is the required October 8 technical seat at 3:00, with Lena assigned as the replacement interviewer.", "At 9:14 on October 3, event Q-47’s response log recorded Lena’s status as declined.", "The Lena-for-Omar assignment is the loop’s only interviewer substitution, and hiring manager Victor approved it.", "Avery confirmed the complete October 8 interview schedule after the substitution was recorded.", "The loop remained on October 8, and no change was made to its three-seat structure."], "request": "Determine whether the loop is ready to confirm, rate its scheduling risk, and identify any participant requiring follow-up."}}, "method": "c2d", "provenance": {"source_id": "diverse-038", "source_is_synthetic": true, "source_sha256": "a38bfb9885128ee8923051462ea69e5f86cf88c8bf143e3228a1b6a65164072e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "hold_route_interviewer"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing confirmation rule without adding exceptions or defaults and preserve the decision’s Nadia Cole/October 13, 2026 loop and Priya Shah bindings. The two evidence spans are complete factual sentences. The counterfactual changes only Priya’s current signed response from confirmation to nonconfirmation; this remains consistent with the roster, calendar, candidate, manager, and invitation facts. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Scheduling log\",\"text\":\"On October 11, 2026, the finalized roster identified Priya Shah as a required interviewer for Nadia Cole’s October 13 interview loop.\"},{\"speaker\":\"Coordinator\",\"text\":\"By 16:30 UTC on October 12, Nadia Cole had confirmed all blocks in her October 13, 2026 interview loop.\"},{\"speaker\":\"Approval log\",\"text\":\"The hiring manager confirmed Nadia Cole’s October 13, 2026 interview loop at 16:45 UTC.\"},{\"speaker\":\"Roster status\",\"text\":\"Every required interviewer other than Priya Shah had confirmed Nadia Cole’s October 13, 2026 interview loop by 17:00 UTC.\"},{\"speaker\":\"Calendar review\",\"text\":\"A calendar check completed at 17:30 UTC showed that Priya Shah had no conflict during her scheduled interview block on October 13, 2026.\"},{\"speaker\":\"Signed-response record\",\"text\":\"At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I confirm this interview loop.”\"},{\"speaker\":\"Invitation registry\",\"text\":\"At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["5", "text"], "text": "At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I confirm this interview loop.”"}, {"path": ["6", "text"], "text": "At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I confirm this interview loop.”", "negative_left": "At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I have not confirmed this interview loop.”", "negative_right": "At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop.", "right": "At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop."}, "verifier_independent_model": false}, "family": "scale-diverse-039-001", "id": "scale-diverse-039-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Scheduling log", "text": "On October 11, 2026, the finalized roster identified Priya Shah as a required interviewer for Nadia Cole’s October 13 interview loop."}, {"speaker": "Coordinator", "text": "By 16:30 UTC on October 12, Nadia Cole had confirmed all blocks in her October 13, 2026 interview loop."}, {"speaker": "Approval log", "text": "The hiring manager confirmed Nadia Cole’s October 13, 2026 interview loop at 16:45 UTC."}, {"speaker": "Roster status", "text": "Every required interviewer other than Priya Shah had confirmed Nadia Cole’s October 13, 2026 interview loop by 17:00 UTC."}, {"speaker": "Calendar review", "text": "A calendar check completed at 17:30 UTC showed that Priya Shah had no conflict during her scheduled interview block on October 13, 2026."}, {"speaker": "Signed-response record", "text": "At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I confirm this interview loop.”"}, {"speaker": "Invitation registry", "text": "At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing confirmation rule without adding exceptions or defaults and preserve the decision’s Nadia Cole/October 13, 2026 loop and Priya Shah bindings. The two evidence spans are complete factual sentences. The counterfactual changes only Priya’s current signed response from confirmation to nonconfirmation; this remains consistent with the roster, calendar, candidate, manager, and invitation facts. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Scheduling log\",\"text\":\"On October 11, 2026, the finalized roster identified Priya Shah as a required interviewer for Nadia Cole’s October 13 interview loop.\"},{\"speaker\":\"Coordinator\",\"text\":\"By 16:30 UTC on October 12, Nadia Cole had confirmed all blocks in her October 13, 2026 interview loop.\"},{\"speaker\":\"Approval log\",\"text\":\"The hiring manager confirmed Nadia Cole’s October 13, 2026 interview loop at 16:45 UTC.\"},{\"speaker\":\"Roster status\",\"text\":\"Every required interviewer other than Priya Shah had confirmed Nadia Cole’s October 13, 2026 interview loop by 17:00 UTC.\"},{\"speaker\":\"Calendar review\",\"text\":\"A calendar check completed at 17:30 UTC showed that Priya Shah had no conflict during her scheduled interview block on October 13, 2026.\"},{\"speaker\":\"Signed-response record\",\"text\":\"At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I confirm this interview loop.”\"},{\"speaker\":\"Invitation registry\",\"text\":\"At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["5", "text"], "text": "At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I confirm this interview loop.”"}, {"path": ["6", "text"], "text": "At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I confirm this interview loop.”", "negative_left": "At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I have not confirmed this interview loop.”", "negative_right": "At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop.", "right": "At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop."}, "verifier_independent_model": false}, "family": "scale-diverse-039-001", "id": "scale-diverse-039-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Scheduling log", "text": "On October 11, 2026, the finalized roster identified Priya Shah as a required interviewer for Nadia Cole’s October 13 interview loop."}, {"speaker": "Coordinator", "text": "By 16:30 UTC on October 12, Nadia Cole had confirmed all blocks in her October 13, 2026 interview loop."}, {"speaker": "Approval log", "text": "The hiring manager confirmed Nadia Cole’s October 13, 2026 interview loop at 16:45 UTC."}, {"speaker": "Roster status", "text": "Every required interviewer other than Priya Shah had confirmed Nadia Cole’s October 13, 2026 interview loop by 17:00 UTC."}, {"speaker": "Calendar review", "text": "A calendar check completed at 17:30 UTC showed that Priya Shah had no conflict during her scheduled interview block on October 13, 2026."}, {"speaker": "Signed-response record", "text": "At 18:00 UTC on October 12, 2026, Priya Shah's current signed response for invitation Q7-418 read, “I have not confirmed this interview loop.”"}, {"speaker": "Invitation registry", "text": "At 18:00 UTC on October 12, 2026, invitation Q7-418 uniquely identified Nadia Cole's October 13, 2026 interview loop."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing readiness criteria without adding exceptions or defaults, and they preserve the decision’s bindings to Nadia Cole’s October 13, 2026 loop and Priya Shah’s current response. The two evidence spans are complete factual sentences. The counterfactual changes only HN-584’s selected status from “confirmed” to “not confirmed,” which remains consistent with the statements that all required interviewers other than Priya confirmed and that Priya has no calendar conflict. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Scheduling system\",\"text\":\"At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop.\"},{\"speaker\":\"Response record\",\"text\":\"Response record HN-584 contains the selected status “confirmed” and was successfully saved at 14:25 UTC on September 29, 2026.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"Nadia Cole has confirmed her October 13, 2026 interview loop. The hiring manager for that loop has confirmed it as well. Priya Shah is listed as a required interviewer, and every required interviewer other than Priya Shah has confirmed the current loop.\"},{\"speaker\":\"Calendar audit\",\"text\":\"Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop."}, {"path": ["1", "text"], "text": "Response record HN-584 contains the selected status “confirmed” and was successfully saved at 14:25 UTC on September 29, 2026."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop.", "negative_left": "At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop.", "negative_right": "Response record HN-584 contains the selected status “not confirmed” and was successfully saved at 14:25 UTC on September 29, 2026.", "right": "Response record HN-584 contains the selected status “confirmed” and was successfully saved at 14:25 UTC on September 29, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-039-003", "id": "scale-diverse-039-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Scheduling system", "text": "At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop."}, {"speaker": "Response record", "text": "Response record HN-584 contains the selected status “confirmed” and was successfully saved at 14:25 UTC on September 29, 2026."}, {"speaker": "Recruiting coordinator", "text": "Nadia Cole has confirmed her October 13, 2026 interview loop. The hiring manager for that loop has confirmed it as well. Priya Shah is listed as a required interviewer, and every required interviewer other than Priya Shah has confirmed the current loop."}, {"speaker": "Calendar audit", "text": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing readiness criteria without adding exceptions or defaults, and they preserve the decision’s bindings to Nadia Cole’s October 13, 2026 loop and Priya Shah’s current response. The two evidence spans are complete factual sentences. The counterfactual changes only HN-584’s selected status from “confirmed” to “not confirmed,” which remains consistent with the statements that all required interviewers other than Priya confirmed and that Priya has no calendar conflict. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Scheduling system\",\"text\":\"At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop.\"},{\"speaker\":\"Response record\",\"text\":\"Response record HN-584 contains the selected status “confirmed” and was successfully saved at 14:25 UTC on September 29, 2026.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"Nadia Cole has confirmed her October 13, 2026 interview loop. The hiring manager for that loop has confirmed it as well. Priya Shah is listed as a required interviewer, and every required interviewer other than Priya Shah has confirmed the current loop.\"},{\"speaker\":\"Calendar audit\",\"text\":\"Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop."}, {"path": ["1", "text"], "text": "Response record HN-584 contains the selected status “confirmed” and was successfully saved at 14:25 UTC on September 29, 2026."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop.", "negative_left": "At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop.", "negative_right": "Response record HN-584 contains the selected status “not confirmed” and was successfully saved at 14:25 UTC on September 29, 2026.", "right": "Response record HN-584 contains the selected status “confirmed” and was successfully saved at 14:25 UTC on September 29, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-039-003", "id": "scale-diverse-039-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Scheduling system", "text": "At 14:26 UTC on September 29, 2026, the scheduling handoff log identifies response record HN-584 as Priya Shah's current response for Nadia Cole's October 13, 2026 interview loop."}, {"speaker": "Response record", "text": "Response record HN-584 contains the selected status “not confirmed” and was successfully saved at 14:25 UTC on September 29, 2026."}, {"speaker": "Recruiting coordinator", "text": "Nadia Cole has confirmed her October 13, 2026 interview loop. The hiring manager for that loop has confirmed it as well. Priya Shah is listed as a required interviewer, and every required interviewer other than Priya Shah has confirmed the current loop."}, {"speaker": "Calendar audit", "text": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original rubric, instructions, request, participants, date, time, and assessment scope. The two focus spans are complete factual sentences about Tomas’s calendar notation and its record-specific legend. The counterfactual changes only the meaning of Q7 from documented availability to undocumented availability; this is coherent with the sole-record assertion and with the absence of an explicit conflict. Neither context embeds a selected risk level, answer code, proposition ID, classifier instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship; the universally scoped negative atoms concern single relationship types over explicit sets. A3 is a factual calendar-documentation relation rather than a policy conclusion. The base and counter assignments are jointly realizable with only A3 changing: C-T can be the sole record documenting Tomas as clear in the base and can explicitly fail to do so in the counter, while no other record documents clearance and no conflict is documented. The policy evidence preserves the state-originated participant bindings and confirmation prerequisites; the unchanged questions automatically preserve the full risk rubric and routing instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction documents the candidate’s exact-time acceptance, clear calendars for all required participants, and the hiring manager’s panel approval. It also excludes explicit declines, calendar conflicts, and candidate-time incompatibility. These conditions are sufficient for Low risk. A9 is unnecessary but does not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "A3 refuted together with A9 supported establishes that no current record documents Tomas as clear. The other required confirmation, calendar records, and approval are documented, while explicit high-risk conditions are excluded. Thus exactly one required record is missing, which is sufficient for Moderate risk and routing the missing item to Tomas while the coordinator holds confirmation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Leo’s acceptance of the October 8, 1:00–2:30 p.m. ET interview time is explicitly documented."}, {"id": "A2", "statement": "Interviewer Priya’s calendar is explicitly documented as clear for the full October 8, 1:00–2:30 p.m. ET interview interval."}, {"id": "A3", "statement": "Current calendar record C-T explicitly documents required interviewer Tomas’s calendar as clear for the full October 8, 1:00–2:30 p.m. ET interview interval."}, {"id": "A4", "statement": "Hiring manager Erin’s calendar is explicitly documented as clear for the full October 8, 1:00–2:30 p.m. ET interview interval."}, {"id": "A5", "statement": "Hiring manager Erin’s approval of Priya and Tomas as the final panel for Leo’s October 8, 1:00–2:30 p.m. ET interview loop is explicitly documented."}, {"id": "A6", "statement": "No required person among Leo, Priya, Tomas, and Erin has explicitly declined the October 8, 1:00–2:30 p.m. ET interview loop."}, {"id": "A7", "statement": "No explicit calendar conflict is documented for any required participant among Priya, Tomas, and Erin during the October 8, 1:00–2:30 p.m. ET interview interval."}, {"id": "A8", "statement": "Leo has no explicitly stated time incompatibility with the October 8, 1:00–2:30 p.m. ET interview loop."}, {"id": "A9", "statement": "Apart from current calendar record C-T, no current calendar record explicitly documents Tomas’s calendar as clear for the full October 8, 1:00–2:30 p.m. ET interview interval."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"Low risk: All required confirmations, clear calendars, and approvals are explicitly documented. The loop is ready to confirm; route the final invitation to the recruiting coordinator.\",\"Moderate risk: No explicit conflict exists, but one required confirmation, calendar record, or approval is missing or tentative. Route the missing item to the candidate, relevant interviewer, or hiring manager, while the recruiting coordinator holds confirmation.\",\"High risk: An explicit decline, calendar conflict, incompatible candidate time, or at least two missing required records prevent confirmation. Route the case to the recruiting coordinator to reopen scheduling with the affected candidate, interviewer, or hiring manager.\"],\"instructions\":\"Select exactly one scheduling-risk level. Use only the explicit case evidence. Assess candidate confirmation, calendar availability, and panel approval; route the next action as specified by the selected level.\",\"type\":\"score\"}},\"state\":{\"context\":\"Recruiting coordinator Maya is preparing a three-person interview loop for candidate Leo on October 8, 1:00–2:30 p.m. ET. The required participants are interviewer Priya, interviewer Tomas, and hiring manager Erin. The loop may be confirmed only when the candidate has accepted the exact time, every participant’s calendar is clear, and the hiring manager has approved the panel.\",\"evidence\":[\"Leo’s signed reply accepts the October 8, 1:00–2:30 p.m. ET interview time and states no time incompatibility.\",\"Current records explicitly show Priya’s and Erin’s calendars clear for the entire October 8, 1:00–2:30 p.m. ET interval.\",\"The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7.\",\"The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar clear for the full listed interval.”\",\"The current-record index contains no calendar record for Tomas other than C-T.\",\"Erin’s signed approval names Priya and Tomas as the final panel for Leo’s October 8, 1:00–2:30 p.m. ET loop.\",\"The handoff audit records no decline by Leo, Priya, Tomas, or Erin and no explicit calendar conflict for Priya, Tomas, or Erin during the scheduled interval.\"],\"request\":\"Rate scheduling risk and identify the appropriate next owner under the rubric.\"}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["state", "evidence", "2"], "text": "The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7."}, {"path": ["state", "evidence", "3"], "text": "The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar clear for the full listed interval.”"}], "policy_evidence": [{"path": ["context"], "text": "Recruiting coordinator Maya is preparing a three-person interview loop for candidate Leo on October 8, 1:00–2:30 p.m. ET. The required participants are interviewer Priya, interviewer Tomas, and hiring manager Erin."}, {"path": ["context"], "text": "The loop may be confirmed only when the candidate has accepted the exact time, every participant’s calendar is clear, and the hiring manager has approved the panel."}, {"path": ["request"], "text": "Rate scheduling risk and identify the appropriate next owner under the rubric."}], "rules": [{"justification": "The candidate’s exact-time acceptance, every required participant’s clear calendar, and the hiring manager’s panel approval are explicitly documented. Explicit declines, calendar conflicts, and candidate-time incompatibility are excluded. The loop is therefore Low risk, and the final invitation is routed to the recruiting coordinator.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Current calendar record C-T does not document Tomas’s calendar as clear, and no other current calendar record documents it. All other required confirmations, calendar records, and approvals are documented, while explicit declines, conflicts, and candidate-time incompatibility are excluded. Exactly one required calendar record is missing, so the case is Moderate risk; the missing item is routed to Tomas while the recruiting coordinator holds confirmation.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7.", "negative_left": "The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7.", "negative_right": "The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar availability not documented for the full listed interval.”", "right": "The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar clear for the full listed interval.”"}, "verifier_independent_model": false}, "family": "scale-diverse-041-003", "id": "scale-diverse-041-003-base", "input": {"questions": {"decision": {"criteria": ["Low risk: All required confirmations, clear calendars, and approvals are explicitly documented. The loop is ready to confirm; route the final invitation to the recruiting coordinator.", "Moderate risk: No explicit conflict exists, but one required confirmation, calendar record, or approval is missing or tentative. Route the missing item to the candidate, relevant interviewer, or hiring manager, while the recruiting coordinator holds confirmation.", "High risk: An explicit decline, calendar conflict, incompatible candidate time, or at least two missing required records prevent confirmation. Route the case to the recruiting coordinator to reopen scheduling with the affected candidate, interviewer, or hiring manager."], "instructions": "Select exactly one scheduling-risk level. Use only the explicit case evidence. Assess candidate confirmation, calendar availability, and panel approval; route the next action as specified by the selected level.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["Low risk: All required confirmations, clear calendars, and approvals are explicitly documented. The loop is ready to confirm; route the final invitation to the recruiting coordinator.", "Moderate risk: No explicit conflict exists, but one required confirmation, calendar record, or approval is missing or tentative. Route the missing item to the candidate, relevant interviewer, or hiring manager, while the recruiting coordinator holds confirmation.", "High risk: An explicit decline, calendar conflict, incompatible candidate time, or at least two missing required records prevent confirmation. Route the case to the recruiting coordinator to reopen scheduling with the affected candidate, interviewer, or hiring manager."], "instructions": "Select exactly one scheduling-risk level. Use only the explicit case evidence. Assess candidate confirmation, calendar availability, and panel approval; route the next action as specified by the selected level.", "type": "score"}}, "state": {"context": "Recruiting coordinator Maya is preparing a three-person interview loop for candidate Leo on October 8, 1:00–2:30 p.m. ET. The required participants are interviewer Priya, interviewer Tomas, and hiring manager Erin. The loop may be confirmed only when the candidate has accepted the exact time, every participant’s calendar is clear, and the hiring manager has approved the panel.", "evidence": ["Leo’s signed reply accepts the October 8, 1:00–2:30 p.m. ET interview time and states no time incompatibility.", "Current records explicitly show Priya’s and Erin’s calendars clear for the entire October 8, 1:00–2:30 p.m. ET interval.", "The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7.", "The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar clear for the full listed interval.”", "The current-record index contains no calendar record for Tomas other than C-T.", "Erin’s signed approval names Priya and Tomas as the final panel for Leo’s October 8, 1:00–2:30 p.m. ET loop.", "The handoff audit records no decline by Leo, Priya, Tomas, or Erin and no explicit calendar conflict for Priya, Tomas, or Erin during the scheduled interval."], "request": "Rate scheduling risk and identify the appropriate next owner under the rubric."}}}, "method": "c2d", "provenance": {"source_id": "diverse-041", "source_is_synthetic": true, "source_sha256": "9f426a8a1021d8f9971666e05b8f43ed20034ecc08c3ee4342a7119e99152b9c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original rubric, instructions, request, participants, date, time, and assessment scope. The two focus spans are complete factual sentences about Tomas’s calendar notation and its record-specific legend. The counterfactual changes only the meaning of Q7 from documented availability to undocumented availability; this is coherent with the sole-record assertion and with the absence of an explicit conflict. Neither context embeds a selected risk level, answer code, proposition ID, classifier instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship; the universally scoped negative atoms concern single relationship types over explicit sets. A3 is a factual calendar-documentation relation rather than a policy conclusion. The base and counter assignments are jointly realizable with only A3 changing: C-T can be the sole record documenting Tomas as clear in the base and can explicitly fail to do so in the counter, while no other record documents clearance and no conflict is documented. The policy evidence preserves the state-originated participant bindings and confirmation prerequisites; the unchanged questions automatically preserve the full risk rubric and routing instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction documents the candidate’s exact-time acceptance, clear calendars for all required participants, and the hiring manager’s panel approval. It also excludes explicit declines, calendar conflicts, and candidate-time incompatibility. These conditions are sufficient for Low risk. A9 is unnecessary but does not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "A3 refuted together with A9 supported establishes that no current record documents Tomas as clear. The other required confirmation, calendar records, and approval are documented, while explicit high-risk conditions are excluded. Thus exactly one required record is missing, which is sufficient for Moderate risk and routing the missing item to Tomas while the coordinator holds confirmation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Leo’s acceptance of the October 8, 1:00–2:30 p.m. ET interview time is explicitly documented."}, {"id": "A2", "statement": "Interviewer Priya’s calendar is explicitly documented as clear for the full October 8, 1:00–2:30 p.m. ET interview interval."}, {"id": "A3", "statement": "Current calendar record C-T explicitly documents required interviewer Tomas’s calendar as clear for the full October 8, 1:00–2:30 p.m. ET interview interval."}, {"id": "A4", "statement": "Hiring manager Erin’s calendar is explicitly documented as clear for the full October 8, 1:00–2:30 p.m. ET interview interval."}, {"id": "A5", "statement": "Hiring manager Erin’s approval of Priya and Tomas as the final panel for Leo’s October 8, 1:00–2:30 p.m. ET interview loop is explicitly documented."}, {"id": "A6", "statement": "No required person among Leo, Priya, Tomas, and Erin has explicitly declined the October 8, 1:00–2:30 p.m. ET interview loop."}, {"id": "A7", "statement": "No explicit calendar conflict is documented for any required participant among Priya, Tomas, and Erin during the October 8, 1:00–2:30 p.m. ET interview interval."}, {"id": "A8", "statement": "Leo has no explicitly stated time incompatibility with the October 8, 1:00–2:30 p.m. ET interview loop."}, {"id": "A9", "statement": "Apart from current calendar record C-T, no current calendar record explicitly documents Tomas’s calendar as clear for the full October 8, 1:00–2:30 p.m. ET interview interval."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"Low risk: All required confirmations, clear calendars, and approvals are explicitly documented. The loop is ready to confirm; route the final invitation to the recruiting coordinator.\",\"Moderate risk: No explicit conflict exists, but one required confirmation, calendar record, or approval is missing or tentative. Route the missing item to the candidate, relevant interviewer, or hiring manager, while the recruiting coordinator holds confirmation.\",\"High risk: An explicit decline, calendar conflict, incompatible candidate time, or at least two missing required records prevent confirmation. Route the case to the recruiting coordinator to reopen scheduling with the affected candidate, interviewer, or hiring manager.\"],\"instructions\":\"Select exactly one scheduling-risk level. Use only the explicit case evidence. Assess candidate confirmation, calendar availability, and panel approval; route the next action as specified by the selected level.\",\"type\":\"score\"}},\"state\":{\"context\":\"Recruiting coordinator Maya is preparing a three-person interview loop for candidate Leo on October 8, 1:00–2:30 p.m. ET. The required participants are interviewer Priya, interviewer Tomas, and hiring manager Erin. The loop may be confirmed only when the candidate has accepted the exact time, every participant’s calendar is clear, and the hiring manager has approved the panel.\",\"evidence\":[\"Leo’s signed reply accepts the October 8, 1:00–2:30 p.m. ET interview time and states no time incompatibility.\",\"Current records explicitly show Priya’s and Erin’s calendars clear for the entire October 8, 1:00–2:30 p.m. ET interval.\",\"The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7.\",\"The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar clear for the full listed interval.”\",\"The current-record index contains no calendar record for Tomas other than C-T.\",\"Erin’s signed approval names Priya and Tomas as the final panel for Leo’s October 8, 1:00–2:30 p.m. ET loop.\",\"The handoff audit records no decline by Leo, Priya, Tomas, or Erin and no explicit calendar conflict for Priya, Tomas, or Erin during the scheduled interval.\"],\"request\":\"Rate scheduling risk and identify the appropriate next owner under the rubric.\"}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["state", "evidence", "2"], "text": "The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7."}, {"path": ["state", "evidence", "3"], "text": "The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar clear for the full listed interval.”"}], "policy_evidence": [{"path": ["context"], "text": "Recruiting coordinator Maya is preparing a three-person interview loop for candidate Leo on October 8, 1:00–2:30 p.m. ET. The required participants are interviewer Priya, interviewer Tomas, and hiring manager Erin."}, {"path": ["context"], "text": "The loop may be confirmed only when the candidate has accepted the exact time, every participant’s calendar is clear, and the hiring manager has approved the panel."}, {"path": ["request"], "text": "Rate scheduling risk and identify the appropriate next owner under the rubric."}], "rules": [{"justification": "The candidate’s exact-time acceptance, every required participant’s clear calendar, and the hiring manager’s panel approval are explicitly documented. Explicit declines, calendar conflicts, and candidate-time incompatibility are excluded. The loop is therefore Low risk, and the final invitation is routed to the recruiting coordinator.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Current calendar record C-T does not document Tomas’s calendar as clear, and no other current calendar record documents it. All other required confirmations, calendar records, and approvals are documented, while explicit declines, conflicts, and candidate-time incompatibility are excluded. Exactly one required calendar record is missing, so the case is Moderate risk; the missing item is routed to Tomas while the recruiting coordinator holds confirmation.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7.", "negative_left": "The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7.", "negative_right": "The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar availability not documented for the full listed interval.”", "right": "The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar clear for the full listed interval.”"}, "verifier_independent_model": false}, "family": "scale-diverse-041-003", "id": "scale-diverse-041-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["Low risk: All required confirmations, clear calendars, and approvals are explicitly documented. The loop is ready to confirm; route the final invitation to the recruiting coordinator.", "Moderate risk: No explicit conflict exists, but one required confirmation, calendar record, or approval is missing or tentative. Route the missing item to the candidate, relevant interviewer, or hiring manager, while the recruiting coordinator holds confirmation.", "High risk: An explicit decline, calendar conflict, incompatible candidate time, or at least two missing required records prevent confirmation. Route the case to the recruiting coordinator to reopen scheduling with the affected candidate, interviewer, or hiring manager."], "instructions": "Select exactly one scheduling-risk level. Use only the explicit case evidence. Assess candidate confirmation, calendar availability, and panel approval; route the next action as specified by the selected level.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["Low risk: All required confirmations, clear calendars, and approvals are explicitly documented. The loop is ready to confirm; route the final invitation to the recruiting coordinator.", "Moderate risk: No explicit conflict exists, but one required confirmation, calendar record, or approval is missing or tentative. Route the missing item to the candidate, relevant interviewer, or hiring manager, while the recruiting coordinator holds confirmation.", "High risk: An explicit decline, calendar conflict, incompatible candidate time, or at least two missing required records prevent confirmation. Route the case to the recruiting coordinator to reopen scheduling with the affected candidate, interviewer, or hiring manager."], "instructions": "Select exactly one scheduling-risk level. Use only the explicit case evidence. Assess candidate confirmation, calendar availability, and panel approval; route the next action as specified by the selected level.", "type": "score"}}, "state": {"context": "Recruiting coordinator Maya is preparing a three-person interview loop for candidate Leo on October 8, 1:00–2:30 p.m. ET. The required participants are interviewer Priya, interviewer Tomas, and hiring manager Erin. The loop may be confirmed only when the candidate has accepted the exact time, every participant’s calendar is clear, and the hiring manager has approved the panel.", "evidence": ["Leo’s signed reply accepts the October 8, 1:00–2:30 p.m. ET interview time and states no time incompatibility.", "Current records explicitly show Priya’s and Erin’s calendars clear for the entire October 8, 1:00–2:30 p.m. ET interval.", "The sole availability notation for required interviewer Tomas in current calendar record C-T covers October 8, 1:00–2:30 p.m. ET and bears status symbol Q7.", "The embedded legend in current calendar record C-T states that status symbol Q7 means “calendar availability not documented for the full listed interval.”", "The current-record index contains no calendar record for Tomas other than C-T.", "Erin’s signed approval names Priya and Tomas as the final panel for Leo’s October 8, 1:00–2:30 p.m. ET loop.", "The handoff audit records no decline by Leo, Priya, Tomas, or Erin and no explicit calendar conflict for Priya, Tomas, or Erin during the scheduled interval."], "request": "Rate scheduling risk and identify the appropriate next owner under the rubric."}}}, "method": "c2d", "provenance": {"source_id": "diverse-041", "source_is_synthetic": true, "source_sha256": "9f426a8a1021d8f9971666e05b8f43ed20034ecc08c3ee4342a7119e99152b9c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scheduling policy, scope, required participants, and October 8, 10:00–11:30 a.m. loop binding. The two evidence spans are complete factual sentences. The counterfactual changes only Priya Shah’s sole archived message from explicit confirmation to non-confirmation; this is consistent with the unchanged statement that no conflict or decline is known because declining to confirm is not necessarily declining attendance. Neither context includes a risk label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set and universal no-conflict/no-decline propositions remain atomic. The focus atom concerns Priya's written confirmation, not a policy conclusion. Base and counter assignments are realizable with only that confirmation fact changing. Empty policy_evidence is correct because the governing scheduling rules are entirely in the questions object, which is automatically retained; the state contributes case observations rather than additional policy needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact required-participant set is established; every member explicitly confirmed; and known conflicts and declines are excluded. This is sufficient for Low risk under the preserved question policy.", "rule_index": 0, "sound": true}, {"reason": "Given the exact required-participant set, the candidate and hiring manager confirmed while Priya's confirmation is refuted, so exactly one required confirmation is missing. Known conflicts and declines are excluded, making Moderate risk sufficient. The table may validly abstain on cases where missing confirmation is merely unknown.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"id": "a2", "statement": "The candidate provided explicit written confirmation for the full proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a3", "statement": "Hiring manager Mateo Ruiz provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a4", "statement": "Interviewer Priya Shah provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a5", "statement": "No required participant has a known conflict with the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a6", "statement": "No required participant has a known decline for the proposed October 8, 10:00–11:30 a.m. loop."}], "base_state_json": "[{\"speaker\":\"Scheduling record\",\"text\":\"On October 1, the finalized plan listed exactly three required participants for the proposed October 8, 10:00–11:30 a.m. loop: the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.\"},{\"speaker\":\"Candidate correspondence\",\"text\":\"Later that day, the candidate sent an explicit written confirmation covering the full proposed October 8, 10:00–11:30 a.m. loop.\"},{\"speaker\":\"Hiring manager correspondence\",\"text\":\"On October 2, Mateo Ruiz sent an explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop.\"},{\"speaker\":\"Records review\",\"text\":\"The subsequent calendar and response review found no known conflict and no known decline for any required participant.\"},{\"speaker\":\"Archive record\",\"text\":\"The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2.\"},{\"speaker\":\"Message transcription\",\"text\":\"The exact full text of message PS-417 is “I explicitly confirm my attendance for the full proposed October 8, 10:00–11:30 a.m. loop.”\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["4", "text"], "text": "The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2."}, {"path": ["5", "text"], "text": "The exact full text of message PS-417 is “I explicitly confirm my attendance for the full proposed October 8, 10:00–11:30 a.m. loop.”"}], "policy_evidence": [], "rules": [{"justification": "Exactly all required participants explicitly confirmed the proposed loop. No required participant has a known conflict or decline, which also excludes any recorded calendar conflict and the competing High-risk triggers.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Priya Shah is exactly the one required participant lacking explicit confirmation. No required participant has a known conflict or decline, so neither High-risk trigger applies and the Moderate-risk criterion governs.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2.", "negative_left": "The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2.", "negative_right": "The exact full text of message PS-417 is “I do not confirm my attendance for the proposed October 8, 10:00–11:30 a.m. loop.”", "right": "The exact full text of message PS-417 is “I explicitly confirm my attendance for the full proposed October 8, 10:00–11:30 a.m. loop.”"}, "verifier_independent_model": false}, "family": "scale-diverse-042-001", "id": "scale-diverse-042-001-base", "input": {"questions": {"decision": {"criteria": ["Low risk: Ready to confirm because the candidate and every required interviewer or hiring manager explicitly confirmed, with no recorded calendar conflict.", "Moderate risk: Not ready to confirm because exactly one required participant lacks explicit confirmation, but no participant has a known conflict; route follow-up to the unconfirmed participant.", "High risk: Not ready to confirm because a required participant has a known conflict or decline, or because two or more required confirmations are missing; escalate to the recruiting coordinator and hiring manager for replanning."], "instructions": "Rate scheduling risk and readiness under this policy: every required participant must provide an explicit written confirmation; an open calendar is not confirmation. If exactly one required confirmation is missing and no conflict is known, route follow-up to that participant. Select the single best level.", "type": "score"}}, "state": [{"speaker": "Scheduling record", "text": "On October 1, the finalized plan listed exactly three required participants for the proposed October 8, 10:00–11:30 a.m. loop: the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"speaker": "Candidate correspondence", "text": "Later that day, the candidate sent an explicit written confirmation covering the full proposed October 8, 10:00–11:30 a.m. loop."}, {"speaker": "Hiring manager correspondence", "text": "On October 2, Mateo Ruiz sent an explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"speaker": "Records review", "text": "The subsequent calendar and response review found no known conflict and no known decline for any required participant."}, {"speaker": "Archive record", "text": "The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2."}, {"speaker": "Message transcription", "text": "The exact full text of message PS-417 is “I explicitly confirm my attendance for the full proposed October 8, 10:00–11:30 a.m. loop.”"}]}, "method": "c2d", "provenance": {"source_id": "diverse-042", "source_is_synthetic": true, "source_sha256": "ecee670d01b33980a9f3e996c8338710a94c55ed2c82ca837c4c5b6102412664", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scheduling policy, scope, required participants, and October 8, 10:00–11:30 a.m. loop binding. The two evidence spans are complete factual sentences. The counterfactual changes only Priya Shah’s sole archived message from explicit confirmation to non-confirmation; this is consistent with the unchanged statement that no conflict or decline is known because declining to confirm is not necessarily declining attendance. Neither context includes a risk label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set and universal no-conflict/no-decline propositions remain atomic. The focus atom concerns Priya's written confirmation, not a policy conclusion. Base and counter assignments are realizable with only that confirmation fact changing. Empty policy_evidence is correct because the governing scheduling rules are entirely in the questions object, which is automatically retained; the state contributes case observations rather than additional policy needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact required-participant set is established; every member explicitly confirmed; and known conflicts and declines are excluded. This is sufficient for Low risk under the preserved question policy.", "rule_index": 0, "sound": true}, {"reason": "Given the exact required-participant set, the candidate and hiring manager confirmed while Priya's confirmation is refuted, so exactly one required confirmation is missing. Known conflicts and declines are excluded, making Moderate risk sufficient. The table may validly abstain on cases where missing confirmation is merely unknown.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"id": "a2", "statement": "The candidate provided explicit written confirmation for the full proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a3", "statement": "Hiring manager Mateo Ruiz provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a4", "statement": "Interviewer Priya Shah provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a5", "statement": "No required participant has a known conflict with the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a6", "statement": "No required participant has a known decline for the proposed October 8, 10:00–11:30 a.m. loop."}], "base_state_json": "[{\"speaker\":\"Scheduling record\",\"text\":\"On October 1, the finalized plan listed exactly three required participants for the proposed October 8, 10:00–11:30 a.m. loop: the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.\"},{\"speaker\":\"Candidate correspondence\",\"text\":\"Later that day, the candidate sent an explicit written confirmation covering the full proposed October 8, 10:00–11:30 a.m. loop.\"},{\"speaker\":\"Hiring manager correspondence\",\"text\":\"On October 2, Mateo Ruiz sent an explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop.\"},{\"speaker\":\"Records review\",\"text\":\"The subsequent calendar and response review found no known conflict and no known decline for any required participant.\"},{\"speaker\":\"Archive record\",\"text\":\"The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2.\"},{\"speaker\":\"Message transcription\",\"text\":\"The exact full text of message PS-417 is “I explicitly confirm my attendance for the full proposed October 8, 10:00–11:30 a.m. loop.”\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["4", "text"], "text": "The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2."}, {"path": ["5", "text"], "text": "The exact full text of message PS-417 is “I explicitly confirm my attendance for the full proposed October 8, 10:00–11:30 a.m. loop.”"}], "policy_evidence": [], "rules": [{"justification": "Exactly all required participants explicitly confirmed the proposed loop. No required participant has a known conflict or decline, which also excludes any recorded calendar conflict and the competing High-risk triggers.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Priya Shah is exactly the one required participant lacking explicit confirmation. No required participant has a known conflict or decline, so neither High-risk trigger applies and the Moderate-risk criterion governs.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2.", "negative_left": "The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2.", "negative_right": "The exact full text of message PS-417 is “I do not confirm my attendance for the proposed October 8, 10:00–11:30 a.m. loop.”", "right": "The exact full text of message PS-417 is “I explicitly confirm my attendance for the full proposed October 8, 10:00–11:30 a.m. loop.”"}, "verifier_independent_model": false}, "family": "scale-diverse-042-001", "id": "scale-diverse-042-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Low risk: Ready to confirm because the candidate and every required interviewer or hiring manager explicitly confirmed, with no recorded calendar conflict.", "Moderate risk: Not ready to confirm because exactly one required participant lacks explicit confirmation, but no participant has a known conflict; route follow-up to the unconfirmed participant.", "High risk: Not ready to confirm because a required participant has a known conflict or decline, or because two or more required confirmations are missing; escalate to the recruiting coordinator and hiring manager for replanning."], "instructions": "Rate scheduling risk and readiness under this policy: every required participant must provide an explicit written confirmation; an open calendar is not confirmation. If exactly one required confirmation is missing and no conflict is known, route follow-up to that participant. Select the single best level.", "type": "score"}}, "state": [{"speaker": "Scheduling record", "text": "On October 1, the finalized plan listed exactly three required participants for the proposed October 8, 10:00–11:30 a.m. loop: the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"speaker": "Candidate correspondence", "text": "Later that day, the candidate sent an explicit written confirmation covering the full proposed October 8, 10:00–11:30 a.m. loop."}, {"speaker": "Hiring manager correspondence", "text": "On October 2, Mateo Ruiz sent an explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"speaker": "Records review", "text": "The subsequent calendar and response review found no known conflict and no known decline for any required participant."}, {"speaker": "Archive record", "text": "The complete archive of all written communications concerning the proposed October 8, 10:00–11:30 a.m. loop contains exactly one item authored by interviewer Priya Shah: message PS-417, sent at 2:14 p.m. on October 2."}, {"speaker": "Message transcription", "text": "The exact full text of message PS-417 is “I do not confirm my attendance for the proposed October 8, 10:00–11:30 a.m. loop.”"}]}, "method": "c2d", "provenance": {"source_id": "diverse-042", "source_is_synthetic": true, "source_sha256": "ecee670d01b33980a9f3e996c8338710a94c55ed2c82ca837c4c5b6102412664", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scheduling policy, criteria, and follow-up requirements; neither context alters or adds policy. Both contexts retain the same October 8, 10:00–11:30 a.m. loop and the same required candidate, Priya Shah/P-47, and Mateo Ruiz bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes only Priya/P-47’s confirmation status, leaving exactly the candidate and Mateo confirmed with no known conflict or decline, and it creates no duplicate contradiction within that context. Neither context includes an answer label, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set and universal no-conflict/no-decline propositions remain atomic. The focus atom concerns Priya's written confirmation, not a policy conclusion. Base and counter assignments are realizable with only that confirmation fact changing. Empty policy_evidence is correct because the governing scheduling rules are entirely in the questions object, which is automatically retained; the state contributes case observations rather than additional policy needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact required-participant set is established; every member explicitly confirmed; and known conflicts and declines are excluded. This is sufficient for Low risk under the preserved question policy.", "rule_index": 0, "sound": true}, {"reason": "Given the exact required-participant set, the candidate and hiring manager confirmed while Priya's confirmation is refuted, so exactly one required confirmation is missing. Known conflicts and declines are excluded, making Moderate risk sufficient. The table may validly abstain on cases where missing confirmation is merely unknown.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"id": "a2", "statement": "The candidate provided explicit written confirmation for the full proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a3", "statement": "Hiring manager Mateo Ruiz provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a4", "statement": "Interviewer Priya Shah provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a5", "statement": "No required participant has a known conflict with the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a6", "statement": "No required participant has a known decline for the proposed October 8, 10:00–11:30 a.m. loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"This concise field note consolidates the attendance roster, written-response log, calendar-conflict report, and decline register for the single proposed loop. The review covers only required participants and records associated with the stated time window.\"},{\"speaker\":\"Records review\",\"text\":\"For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline.\"},{\"speaker\":\"Register entry\",\"text\":\"The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records an explicit written confirmation from participant P-47 at 2:16 p.m. on October 3.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline."}, {"path": ["2", "text"], "text": "The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records an explicit written confirmation from participant P-47 at 2:16 p.m. on October 3."}], "policy_evidence": [], "rules": [{"justification": "Exactly all required participants explicitly confirmed the proposed loop. No required participant has a known conflict or decline, which also excludes any recorded calendar conflict and the competing High-risk triggers.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Priya Shah is exactly the one required participant lacking explicit confirmation. No required participant has a known conflict or decline, so neither High-risk trigger applies and the Moderate-risk criterion governs.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline.", "negative_left": "For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline.", "negative_right": "The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records no explicit written confirmation from participant P-47.", "right": "The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records an explicit written confirmation from participant P-47 at 2:16 p.m. on October 3."}, "verifier_independent_model": false}, "family": "scale-diverse-042-004", "id": "scale-diverse-042-004-base", "input": {"questions": {"decision": {"criteria": ["Low risk: Ready to confirm because the candidate and every required interviewer or hiring manager explicitly confirmed, with no recorded calendar conflict.", "Moderate risk: Not ready to confirm because exactly one required participant lacks explicit confirmation, but no participant has a known conflict; route follow-up to the unconfirmed participant.", "High risk: Not ready to confirm because a required participant has a known conflict or decline, or because two or more required confirmations are missing; escalate to the recruiting coordinator and hiring manager for replanning."], "instructions": "Rate scheduling risk and readiness under this policy: every required participant must provide an explicit written confirmation; an open calendar is not confirmation. If exactly one required confirmation is missing and no conflict is known, route follow-up to that participant. Select the single best level.", "type": "score"}}, "state": [{"speaker": "Recruiting coordinator", "text": "This concise field note consolidates the attendance roster, written-response log, calendar-conflict report, and decline register for the single proposed loop. The review covers only required participants and records associated with the stated time window."}, {"speaker": "Records review", "text": "For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline."}, {"speaker": "Register entry", "text": "The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records an explicit written confirmation from participant P-47 at 2:16 p.m. on October 3."}]}, "method": "c2d", "provenance": {"source_id": "diverse-042", "source_is_synthetic": true, "source_sha256": "ecee670d01b33980a9f3e996c8338710a94c55ed2c82ca837c4c5b6102412664", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scheduling policy, criteria, and follow-up requirements; neither context alters or adds policy. Both contexts retain the same October 8, 10:00–11:30 a.m. loop and the same required candidate, Priya Shah/P-47, and Mateo Ruiz bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes only Priya/P-47’s confirmation status, leaving exactly the candidate and Mateo confirmed with no known conflict or decline, and it creates no duplicate contradiction within that context. Neither context includes an answer label, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set and universal no-conflict/no-decline propositions remain atomic. The focus atom concerns Priya's written confirmation, not a policy conclusion. Base and counter assignments are realizable with only that confirmation fact changing. Empty policy_evidence is correct because the governing scheduling rules are entirely in the questions object, which is automatically retained; the state contributes case observations rather than additional policy needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact required-participant set is established; every member explicitly confirmed; and known conflicts and declines are excluded. This is sufficient for Low risk under the preserved question policy.", "rule_index": 0, "sound": true}, {"reason": "Given the exact required-participant set, the candidate and hiring manager confirmed while Priya's confirmation is refuted, so exactly one required confirmation is missing. Known conflicts and declines are excluded, making Moderate risk sufficient. The table may validly abstain on cases where missing confirmation is merely unknown.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"id": "a2", "statement": "The candidate provided explicit written confirmation for the full proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a3", "statement": "Hiring manager Mateo Ruiz provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a4", "statement": "Interviewer Priya Shah provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a5", "statement": "No required participant has a known conflict with the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a6", "statement": "No required participant has a known decline for the proposed October 8, 10:00–11:30 a.m. loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"This concise field note consolidates the attendance roster, written-response log, calendar-conflict report, and decline register for the single proposed loop. The review covers only required participants and records associated with the stated time window.\"},{\"speaker\":\"Records review\",\"text\":\"For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline.\"},{\"speaker\":\"Register entry\",\"text\":\"The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records an explicit written confirmation from participant P-47 at 2:16 p.m. on October 3.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline."}, {"path": ["2", "text"], "text": "The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records an explicit written confirmation from participant P-47 at 2:16 p.m. on October 3."}], "policy_evidence": [], "rules": [{"justification": "Exactly all required participants explicitly confirmed the proposed loop. No required participant has a known conflict or decline, which also excludes any recorded calendar conflict and the competing High-risk triggers.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Priya Shah is exactly the one required participant lacking explicit confirmation. No required participant has a known conflict or decline, so neither High-risk trigger applies and the Moderate-risk criterion governs.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline.", "negative_left": "For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline.", "negative_right": "The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records no explicit written confirmation from participant P-47.", "right": "The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records an explicit written confirmation from participant P-47 at 2:16 p.m. on October 3."}, "verifier_independent_model": false}, "family": "scale-diverse-042-004", "id": "scale-diverse-042-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["Low risk: Ready to confirm because the candidate and every required interviewer or hiring manager explicitly confirmed, with no recorded calendar conflict.", "Moderate risk: Not ready to confirm because exactly one required participant lacks explicit confirmation, but no participant has a known conflict; route follow-up to the unconfirmed participant.", "High risk: Not ready to confirm because a required participant has a known conflict or decline, or because two or more required confirmations are missing; escalate to the recruiting coordinator and hiring manager for replanning."], "instructions": "Rate scheduling risk and readiness under this policy: every required participant must provide an explicit written confirmation; an open calendar is not confirmation. If exactly one required confirmation is missing and no conflict is known, route follow-up to that participant. Select the single best level.", "type": "score"}}, "state": [{"speaker": "Recruiting coordinator", "text": "This concise field note consolidates the attendance roster, written-response log, calendar-conflict report, and decline register for the single proposed loop. The review covers only required participants and records associated with the stated time window."}, {"speaker": "Records review", "text": "For the proposed October 8, 10:00–11:30 a.m. loop, the required participants are exactly the candidate, interviewer Priya Shah, registered as participant P-47, and hiring manager Mateo Ruiz; the candidate and Mateo Ruiz submitted explicit written confirmations, and no required participant has a known conflict or decline."}, {"speaker": "Register entry", "text": "The complete confirmation register for the proposed October 8, 10:00–11:30 a.m. loop records no explicit written confirmation from participant P-47."}]}, "method": "c2d", "provenance": {"source_id": "diverse-042", "source_is_synthetic": true, "source_sha256": "ecee670d01b33980a9f3e996c8338710a94c55ed2c82ca837c4c5b6102412664", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing impact, routing, and readiness policies, while the unchanged questions preserve all choice criteria and instructions. The request remains bound to Ivo’s audit-export change request, and both focus spans are complete factual sentences. The counterfactual changes only S-47’s attachment status; saying it existed and described the schedule effect does not contradict saying it was not attached. Neither context embeds an answer choice, code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns whether quantified schedule evidence is attached rather than a policy conclusion. The base and counter assignments are realizable with only that attachment fact changing. The policy evidence accurately preserves the impact thresholds, owner bindings, and readiness rule originating in the original state; rules already contained in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Eight added workdays and $12,000 establish moderate impact and therefore delivery-lead review. The conjunction also establishes a rationale plus attached quantified schedule and resource evidence, which is sufficient for readiness.", "rule_index": 0, "sound": true}, {"reason": "Eight added workdays and $12,000 establish moderate impact and therefore delivery-lead review. Refutation of attached quantified schedule evidence establishes that a required readiness component is missing, so the request is not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Ivo’s audit-export change request includes a rationale explaining its need."}, {"id": "a2", "statement": "Ivo’s audit-export change request adds eight workdays."}, {"id": "a3", "statement": "Ivo’s audit-export change request requires $12,000 in contractor support."}, {"id": "a4", "statement": "Ivo’s audit-export change request has attached quantified schedule evidence."}, {"id": "a5", "statement": "Ivo’s audit-export change request has attached quantified resource evidence."}], "base_state_json": "{\"case_note\":\"Ivo submitted an audit-export change request for review.\",\"evidence\":[\"On 12 May 2026, Ivo included a rationale explaining that the audit export was needed to satisfy a new contractual reporting clause.\",\"At 15:45 UTC on 14 May 2026, cost sheet C-19 was attached to the request and quantified the required contractor support at $12,000.\",\"At 16:00 UTC on 14 May 2026, worksheet S-47 was attached to Ivo’s audit-export change request.\",\"At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays.\"],\"policy\":[\"Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000.\",\"Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes.\",\"A request is ready only with a rationale and attached, quantified schedule and resource evidence.\"],\"request\":\"Route and classify Ivo’s audit-export change request.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 16:00 UTC on 14 May 2026, worksheet S-47 was attached to Ivo’s audit-export change request."}, {"path": ["evidence", "3"], "text": "At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays."}], "policy_evidence": [{"path": ["context"], "text": "Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000."}, {"path": ["context"], "text": "Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes."}, {"path": ["context"], "text": "A request is ready only with a rationale and attached, quantified schedule and resource evidence."}], "rules": [{"justification": "Eight added workdays and $12,000 are moderate impact, which is reviewed by the delivery lead. The rationale and both forms of attached quantified evidence are present, so the request is ready.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Eight added workdays and $12,000 are moderate impact, which is reviewed by the delivery lead. Although the rationale and quantified resource evidence are present, required attached quantified schedule evidence is missing, so the request is not ready.", "target": "delivery_lead_not_ready_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:00 UTC on 14 May 2026, worksheet S-47 was attached to Ivo’s audit-export change request.", "negative_left": "At 16:00 UTC on 14 May 2026, worksheet S-47 was not attached to Ivo’s audit-export change request.", "negative_right": "At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays.", "right": "At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays."}, "verifier_independent_model": false}, "family": "scale-diverse-043-005", "id": "scale-diverse-043-005-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_moderate": "Delivery lead review; not ready; moderate impact. Use only when impact is 3–10 days or $5,001–$25,000 but required rationale or attached quantified evidence is missing.", "delivery_lead_ready_moderate": "Delivery lead review; ready; moderate impact. Use only when impact is 3–10 days or $5,001–$25,000 and both rationale and attached quantified evidence are present.", "none_of_above": "Use only when the correct owner, readiness, and severity combination is not represented by any other option.", "project_manager_not_ready_low": "Project manager review; not ready; low impact. Use only when impact is at most two days and $5,000 but required rationale or attached quantified evidence is missing.", "project_manager_ready_low": "Project manager review; ready; low impact. Use only when the request has the required rationale and evidence and adds at most two days and $5,000.", "project_sponsor_ready_severe": "Project sponsor review; ready; severe impact. Use only when the request has the required rationale and evidence and exceeds ten days or $25,000."}, "instructions": "Select the single option that correctly identifies the review owner, readiness, and delivery-impact severity. Resolve pronouns and references using their surrounding context. The categories in the options are exact combinations; choose none_of_above only if no listed combination applies.", "type": "choice"}}, "state": {"case_note": "Ivo submitted an audit-export change request for review.", "evidence": ["On 12 May 2026, Ivo included a rationale explaining that the audit export was needed to satisfy a new contractual reporting clause.", "At 15:45 UTC on 14 May 2026, cost sheet C-19 was attached to the request and quantified the required contractor support at $12,000.", "At 16:00 UTC on 14 May 2026, worksheet S-47 was attached to Ivo’s audit-export change request.", "At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays."], "policy": ["Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000.", "Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes.", "A request is ready only with a rationale and attached, quantified schedule and resource evidence."], "request": "Route and classify Ivo’s audit-export change request."}}, "method": "c2d", "provenance": {"source_id": "diverse-043", "source_is_synthetic": true, "source_sha256": "1e08b3d5ffbc4028fbb9e3181f5514c4bafda4745f42969ea60a2dd3f1360142", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing impact, routing, and readiness policies, while the unchanged questions preserve all choice criteria and instructions. The request remains bound to Ivo’s audit-export change request, and both focus spans are complete factual sentences. The counterfactual changes only S-47’s attachment status; saying it existed and described the schedule effect does not contradict saying it was not attached. Neither context embeds an answer choice, code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns whether quantified schedule evidence is attached rather than a policy conclusion. The base and counter assignments are realizable with only that attachment fact changing. The policy evidence accurately preserves the impact thresholds, owner bindings, and readiness rule originating in the original state; rules already contained in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Eight added workdays and $12,000 establish moderate impact and therefore delivery-lead review. The conjunction also establishes a rationale plus attached quantified schedule and resource evidence, which is sufficient for readiness.", "rule_index": 0, "sound": true}, {"reason": "Eight added workdays and $12,000 establish moderate impact and therefore delivery-lead review. Refutation of attached quantified schedule evidence establishes that a required readiness component is missing, so the request is not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Ivo’s audit-export change request includes a rationale explaining its need."}, {"id": "a2", "statement": "Ivo’s audit-export change request adds eight workdays."}, {"id": "a3", "statement": "Ivo’s audit-export change request requires $12,000 in contractor support."}, {"id": "a4", "statement": "Ivo’s audit-export change request has attached quantified schedule evidence."}, {"id": "a5", "statement": "Ivo’s audit-export change request has attached quantified resource evidence."}], "base_state_json": "{\"case_note\":\"Ivo submitted an audit-export change request for review.\",\"evidence\":[\"On 12 May 2026, Ivo included a rationale explaining that the audit export was needed to satisfy a new contractual reporting clause.\",\"At 15:45 UTC on 14 May 2026, cost sheet C-19 was attached to the request and quantified the required contractor support at $12,000.\",\"At 16:00 UTC on 14 May 2026, worksheet S-47 was attached to Ivo’s audit-export change request.\",\"At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays.\"],\"policy\":[\"Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000.\",\"Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes.\",\"A request is ready only with a rationale and attached, quantified schedule and resource evidence.\"],\"request\":\"Route and classify Ivo’s audit-export change request.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 16:00 UTC on 14 May 2026, worksheet S-47 was attached to Ivo’s audit-export change request."}, {"path": ["evidence", "3"], "text": "At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays."}], "policy_evidence": [{"path": ["context"], "text": "Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000."}, {"path": ["context"], "text": "Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes."}, {"path": ["context"], "text": "A request is ready only with a rationale and attached, quantified schedule and resource evidence."}], "rules": [{"justification": "Eight added workdays and $12,000 are moderate impact, which is reviewed by the delivery lead. The rationale and both forms of attached quantified evidence are present, so the request is ready.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Eight added workdays and $12,000 are moderate impact, which is reviewed by the delivery lead. Although the rationale and quantified resource evidence are present, required attached quantified schedule evidence is missing, so the request is not ready.", "target": "delivery_lead_not_ready_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:00 UTC on 14 May 2026, worksheet S-47 was attached to Ivo’s audit-export change request.", "negative_left": "At 16:00 UTC on 14 May 2026, worksheet S-47 was not attached to Ivo’s audit-export change request.", "negative_right": "At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays.", "right": "At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays."}, "verifier_independent_model": false}, "family": "scale-diverse-043-005", "id": "scale-diverse-043-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_moderate": "Delivery lead review; not ready; moderate impact. Use only when impact is 3–10 days or $5,001–$25,000 but required rationale or attached quantified evidence is missing.", "delivery_lead_ready_moderate": "Delivery lead review; ready; moderate impact. Use only when impact is 3–10 days or $5,001–$25,000 and both rationale and attached quantified evidence are present.", "none_of_above": "Use only when the correct owner, readiness, and severity combination is not represented by any other option.", "project_manager_not_ready_low": "Project manager review; not ready; low impact. Use only when impact is at most two days and $5,000 but required rationale or attached quantified evidence is missing.", "project_manager_ready_low": "Project manager review; ready; low impact. Use only when the request has the required rationale and evidence and adds at most two days and $5,000.", "project_sponsor_ready_severe": "Project sponsor review; ready; severe impact. Use only when the request has the required rationale and evidence and exceeds ten days or $25,000."}, "instructions": "Select the single option that correctly identifies the review owner, readiness, and delivery-impact severity. Resolve pronouns and references using their surrounding context. The categories in the options are exact combinations; choose none_of_above only if no listed combination applies.", "type": "choice"}}, "state": {"case_note": "Ivo submitted an audit-export change request for review.", "evidence": ["On 12 May 2026, Ivo included a rationale explaining that the audit export was needed to satisfy a new contractual reporting clause.", "At 15:45 UTC on 14 May 2026, cost sheet C-19 was attached to the request and quantified the required contractor support at $12,000.", "At 16:00 UTC on 14 May 2026, worksheet S-47 was not attached to Ivo’s audit-export change request.", "At 16:00 UTC on 14 May 2026, worksheet S-47 was the only existing document that described the request’s schedule effect, stating that implementation would add eight workdays."], "policy": ["Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000.", "Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes.", "A request is ready only with a rationale and attached, quantified schedule and resource evidence."], "request": "Route and classify Ivo’s audit-export change request."}}, "method": "c2d", "provenance": {"source_id": "diverse-043", "source_is_synthetic": true, "source_sha256": "1e08b3d5ffbc4028fbb9e3181f5514c4bafda4745f42969ea60a2dd3f1360142", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_not_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original classification policy and request scope without adding exceptions, priorities, or missing-evidence defaults. They remain bound to an internal-beta schedule-change request and vary only case observations: the proposed date and corresponding business-day count. The focus evidence consists of exactly two complete factual sentences. The counterfactual’s 12 October proposal and five-business-day calendar count are coherent with the unchanged 5 October scheduled date and do not duplicate a conflicting measurement. Neither context contains an answer choice, output instruction, rule table, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change log\",\"text\":\"When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss.\"},{\"speaker\":\"Schedule update\",\"text\":\"At 14:10 UTC on 17 September 2026, request IB-47 proposed 13 October 2026 as the internal-beta date, and the project calendar recorded six business days from 5 October 2026 to 13 October 2026.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss."}, {"path": ["1", "text"], "text": "At 14:10 UTC on 17 September 2026, request IB-47 proposed 13 October 2026 as the internal-beta date, and the project calendar recorded six business days from 5 October 2026 to 13 October 2026."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss.", "negative_left": "When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss.", "negative_right": "At 14:10 UTC on 17 September 2026, request IB-47 proposed 12 October 2026 as the internal-beta date, and the project calendar recorded five business days from 5 October 2026 to 12 October 2026.", "right": "At 14:10 UTC on 17 September 2026, request IB-47 proposed 13 October 2026 as the internal-beta date, and the project calendar recorded six business days from 5 October 2026 to 13 October 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-044-001", "id": "scale-diverse-044-001-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change log", "text": "When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss."}, {"speaker": "Schedule update", "text": "At 14:10 UTC on 17 September 2026, request IB-47 proposed 13 October 2026 as the internal-beta date, and the project calendar recorded six business days from 5 October 2026 to 13 October 2026."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original classification policy and request scope without adding exceptions, priorities, or missing-evidence defaults. They remain bound to an internal-beta schedule-change request and vary only case observations: the proposed date and corresponding business-day count. The focus evidence consists of exactly two complete factual sentences. The counterfactual’s 12 October proposal and five-business-day calendar count are coherent with the unchanged 5 October scheduled date and do not duplicate a conflicting measurement. Neither context contains an answer choice, output instruction, rule table, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change log\",\"text\":\"When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss.\"},{\"speaker\":\"Schedule update\",\"text\":\"At 14:10 UTC on 17 September 2026, request IB-47 proposed 13 October 2026 as the internal-beta date, and the project calendar recorded six business days from 5 October 2026 to 13 October 2026.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss."}, {"path": ["1", "text"], "text": "At 14:10 UTC on 17 September 2026, request IB-47 proposed 13 October 2026 as the internal-beta date, and the project calendar recorded six business days from 5 October 2026 to 13 October 2026."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss.", "negative_left": "When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss.", "negative_right": "At 14:10 UTC on 17 September 2026, request IB-47 proposed 12 October 2026 as the internal-beta date, and the project calendar recorded five business days from 5 October 2026 to 12 October 2026.", "right": "At 14:10 UTC on 17 September 2026, request IB-47 proposed 13 October 2026 as the internal-beta date, and the project calendar recorded six business days from 5 October 2026 to 13 October 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-044-001", "id": "scale-diverse-044-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change log", "text": "When internal-beta change request IB-47 was logged at 09:20 UTC on 17 September 2026, its currently scheduled internal-beta date was 5 October 2026, its file contained a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner, and its review ledger recorded no budget change, no scope-baseline change, and no contractual milestone miss."}, {"speaker": "Schedule update", "text": "At 14:10 UTC on 17 September 2026, request IB-47 proposed 12 October 2026 as the internal-beta date, and the project calendar recorded five business days from 5 October 2026 to 12 October 2026."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the request, decision scope, responsible entity, and relevant timing/path bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes Mara’s paraphrase from an 11-minute result to a 7-minute result while leaving the attached report’s 11-minute result unchanged; these are distinct source and paraphrase assertions rather than contradictory duplicate measurements. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users. At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users. At 09:20, Mara identified the proposed addition as a scope change and included a rationale explaining that analysts needed a faster way to locate relevant reports. At 09:35, the delivery lead estimated that the change would require five additional weeks and cost an additional $60,000. At 09:42, Mara added those quantified schedule and cost impacts to the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users."}, {"path": [], "text": "At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users.", "negative_left": "At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users.", "negative_right": "At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 7 minutes among 240 pilot users.", "right": "At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users."}, "verifier_independent_model": false}, "family": "scale-diverse-045-001", "id": "scale-diverse-045-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users. At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users. At 09:20, Mara identified the proposed addition as a scope change and included a rationale explaining that analysts needed a faster way to locate relevant reports. At 09:35, the delivery lead estimated that the change would require five additional weeks and cost an additional $60,000. At 09:42, Mara added those quantified schedule and cost impacts to the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the request, decision scope, responsible entity, and relevant timing/path bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes Mara’s paraphrase from an 11-minute result to a 7-minute result while leaving the attached report’s 11-minute result unchanged; these are distinct source and paraphrase assertions rather than contradictory duplicate measurements. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users. At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users. At 09:20, Mara identified the proposed addition as a scope change and included a rationale explaining that analysts needed a faster way to locate relevant reports. At 09:35, the delivery lead estimated that the change would require five additional weeks and cost an additional $60,000. At 09:42, Mara added those quantified schedule and cost impacts to the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users."}, {"path": [], "text": "At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users.", "negative_left": "At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users.", "negative_right": "At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 7 minutes among 240 pilot users.", "right": "At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users."}, "verifier_independent_model": false}, "family": "scale-diverse-045-001", "id": "scale-diverse-045-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "At 09:00 UTC on 14 August 2026, the pilot report attached to Mara's request to add advanced filtering to the reporting module recorded one finding: advanced filtering reduced the median report-search time from 18 minutes to 11 minutes among 240 pilot users. At 09:12 UTC on 14 August 2026, Mara's evidence paraphrase in the advanced-filtering request stated that advanced filtering reduced the median report-search time from 18 minutes to 7 minutes among 240 pilot users. At 09:20, Mara identified the proposed addition as a scope change and included a rationale explaining that analysts needed a faster way to locate relevant reports. At 09:35, the delivery lead estimated that the change would require five additional weeks and cost an additional $60,000. At 09:42, Mara added those quantified schedule and cost impacts to the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the full governing policy and the request, requester, project-manager decision path, sponsor-routing scope, and delivery-impact facts. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the pilot finding from 31 to 39 minutes while leaving Mara’s paraphrase at 31 minutes; these are separately attributed assertions rather than contradictory duplicate measurements. Neither context contains an explicit answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Operational handoff note: Mara asks to add advanced filtering to the reporting module, and the request is logged as a scope change. Her submission includes a rationale explaining that easier filtering would reduce report-preparation effort. It also includes quantified delivery impact: the delivery lead estimates five additional weeks and an additional cost of $60,000. The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes. Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes. The project manager is reviewing the submission at this handoff. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes."}, {"path": [], "text": "Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes.", "negative_left": "The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 39 minutes.", "negative_right": "Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes.", "right": "Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes."}, "verifier_independent_model": false}, "family": "scale-diverse-045-003", "id": "scale-diverse-045-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Operational handoff note: Mara asks to add advanced filtering to the reporting module, and the request is logged as a scope change. Her submission includes a rationale explaining that easier filtering would reduce report-preparation effort. It also includes quantified delivery impact: the delivery lead estimates five additional weeks and an additional cost of $60,000. The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes. Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes. The project manager is reviewing the submission at this handoff. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the full governing policy and the request, requester, project-manager decision path, sponsor-routing scope, and delivery-impact facts. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the pilot finding from 31 to 39 minutes while leaving Mara’s paraphrase at 31 minutes; these are separately attributed assertions rather than contradictory duplicate measurements. Neither context contains an explicit answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Operational handoff note: Mara asks to add advanced filtering to the reporting module, and the request is logged as a scope change. Her submission includes a rationale explaining that easier filtering would reduce report-preparation effort. It also includes quantified delivery impact: the delivery lead estimates five additional weeks and an additional cost of $60,000. The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes. Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes. The project manager is reviewing the submission at this handoff. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes."}, {"path": [], "text": "Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes.", "negative_left": "The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 39 minutes.", "negative_right": "Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes.", "right": "Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes."}, "verifier_independent_model": false}, "family": "scale-diverse-045-003", "id": "scale-diverse-045-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Operational handoff note: Mara asks to add advanced filtering to the reporting module, and the request is logged as a scope change. Her submission includes a rationale explaining that easier filtering would reduce report-preparation effort. It also includes quantified delivery impact: the delivery lead estimates five additional weeks and an additional cost of $60,000. The pilot report attached to Mara's advanced-filtering request at the 2026-09-14 16:20 UTC handoff contains one outcome finding: cohort AF-27's median report-preparation time fell from 46 minutes to 39 minutes. Mara's evidence paraphrase in the advanced-filtering request at the 2026-09-14 16:20 UTC handoff consists solely of the statement that cohort AF-27's median report-preparation time fell from 46 minutes to 31 minutes. The project manager is reviewing the submission at this handoff. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing routing, readiness, and delivery-impact policy verbatim and retain the same dashboard request, entity, review task, and relevant timing path. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the post-change launch estimate from business day 255 to 248; it does not conflict with the unchanged pre-change estimate or validation evidence. Neither context contains a gold answer, output instruction, answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual delay-range proposition rather than a policy statement. The base and counter assignments are realizable with only that range proposition changing: a validated, quantified delay can be inside or outside 11–20 business days while all other facts remain fixed. The policy evidence accurately preserves the governing ownership, readiness, and impact rules from the original state; instructions and decision criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes sponsor ownership because the request adds a committed deliverable, readiness because rationale, quantified timeline and resource impacts, and Delivery lead validation are all present, and High impact because the delay is within 11–20 business days. The exact $48,000 cost does not conflict with that delay-based High classification.", "rule_index": 0, "sound": true}, {"reason": "The conjunction still establishes sponsor ownership and readiness. Refutation of the 11–20-business-day delay range excludes the delay basis for High, while the exact $48,000 added cost excludes the $50,000–$100,000 cost basis. Therefore the request does not meet the specified High-impact classification.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The dashboard change request adds the supplier-audit dashboard."}, {"id": "a2", "statement": "The supplier-audit dashboard is a committed deliverable under the signed procurement commitment."}, {"id": "a3", "statement": "The dashboard change request states a rationale."}, {"id": "a4", "statement": "The dashboard change request has a quantified timeline impact."}, {"id": "a5", "statement": "The dashboard change request has a quantified resource impact."}, {"id": "a6", "statement": "The Delivery lead validated the dashboard change request's timeline-impact estimate."}, {"id": "a7", "statement": "The Delivery lead validated the dashboard change request's resource-impact estimate."}, {"id": "a8", "statement": "The dashboard change request's estimated launch delay is within 11–20 business days."}, {"id": "a9", "statement": "The dashboard change request's estimated added cost is $48,000."}], "base_state_json": "{\"context\":\"At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.\",\"evidence\":[\"On 7 April 2026, the procurement officer confirmed that the signed procurement commitment lists the supplier-audit dashboard as a committed deliverable.\",\"At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240.\",\"At 10:00 UTC, the requester submitted the dashboard change request to add that dashboard, stating that the addition would fulfill the signed procurement commitment.\",\"At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 255 in the same ledger.\",\"The Project manager documented exactly $48,000 in added contractor cost as the request's quantified resource impact.\",\"At noon, the Delivery lead checked the dependency plan and validated both the ledger-based timeline-impact estimate and the resource-impact estimate.\"],\"request\":\"Triage the dashboard change request for review.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240."}, {"path": ["evidence", "3"], "text": "At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 255 in the same ledger."}], "policy_evidence": [{"path": ["context"], "text": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost."}], "rules": [{"justification": "Adding a committed deliverable makes the Project sponsor the owner. The stated rationale, quantified impacts, and Delivery lead validation make the request ready. An 11–20-business-day delay gives it High delivery impact; the observed $48,000 cost does not introduce a competing cost band.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The request remains sponsor-owned and ready, but it has neither an 11–20-business-day delay nor a $50,000–$100,000 added cost. It therefore does not have High delivery impact.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240.", "negative_left": "At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240.", "negative_right": "At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 248 in the same ledger.", "right": "At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 255 in the same ledger."}, "verifier_independent_model": false}, "family": "scale-diverse-046-001", "id": "scale-diverse-046-001-base", "input": {"questions": {"decision": {"criteria": {"false": "The request has a different owner, is not ready for review, or does not have High delivery impact.", "true": "The request is sponsor-owned, ready for review, and has High delivery impact."}, "instructions": "Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.", "type": "noul"}}, "state": {"context": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.", "evidence": ["On 7 April 2026, the procurement officer confirmed that the signed procurement commitment lists the supplier-audit dashboard as a committed deliverable.", "At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240.", "At 10:00 UTC, the requester submitted the dashboard change request to add that dashboard, stating that the addition would fulfill the signed procurement commitment.", "At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 255 in the same ledger.", "The Project manager documented exactly $48,000 in added contractor cost as the request's quantified resource impact.", "At noon, the Delivery lead checked the dependency plan and validated both the ledger-based timeline-impact estimate and the resource-impact estimate."], "request": "Triage the dashboard change request for review."}}, "method": "c2d", "provenance": {"source_id": "diverse-046", "source_is_synthetic": true, "source_sha256": "43f34b10f9d9ec2b8f7bd8ba11230d8a6bfd24bac2a882d42ea803c135c9fe6a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing routing, readiness, and delivery-impact policy verbatim and retain the same dashboard request, entity, review task, and relevant timing path. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the post-change launch estimate from business day 255 to 248; it does not conflict with the unchanged pre-change estimate or validation evidence. Neither context contains a gold answer, output instruction, answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual delay-range proposition rather than a policy statement. The base and counter assignments are realizable with only that range proposition changing: a validated, quantified delay can be inside or outside 11–20 business days while all other facts remain fixed. The policy evidence accurately preserves the governing ownership, readiness, and impact rules from the original state; instructions and decision criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes sponsor ownership because the request adds a committed deliverable, readiness because rationale, quantified timeline and resource impacts, and Delivery lead validation are all present, and High impact because the delay is within 11–20 business days. The exact $48,000 cost does not conflict with that delay-based High classification.", "rule_index": 0, "sound": true}, {"reason": "The conjunction still establishes sponsor ownership and readiness. Refutation of the 11–20-business-day delay range excludes the delay basis for High, while the exact $48,000 added cost excludes the $50,000–$100,000 cost basis. Therefore the request does not meet the specified High-impact classification.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The dashboard change request adds the supplier-audit dashboard."}, {"id": "a2", "statement": "The supplier-audit dashboard is a committed deliverable under the signed procurement commitment."}, {"id": "a3", "statement": "The dashboard change request states a rationale."}, {"id": "a4", "statement": "The dashboard change request has a quantified timeline impact."}, {"id": "a5", "statement": "The dashboard change request has a quantified resource impact."}, {"id": "a6", "statement": "The Delivery lead validated the dashboard change request's timeline-impact estimate."}, {"id": "a7", "statement": "The Delivery lead validated the dashboard change request's resource-impact estimate."}, {"id": "a8", "statement": "The dashboard change request's estimated launch delay is within 11–20 business days."}, {"id": "a9", "statement": "The dashboard change request's estimated added cost is $48,000."}], "base_state_json": "{\"context\":\"At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.\",\"evidence\":[\"On 7 April 2026, the procurement officer confirmed that the signed procurement commitment lists the supplier-audit dashboard as a committed deliverable.\",\"At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240.\",\"At 10:00 UTC, the requester submitted the dashboard change request to add that dashboard, stating that the addition would fulfill the signed procurement commitment.\",\"At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 255 in the same ledger.\",\"The Project manager documented exactly $48,000 in added contractor cost as the request's quantified resource impact.\",\"At noon, the Delivery lead checked the dependency plan and validated both the ledger-based timeline-impact estimate and the resource-impact estimate.\"],\"request\":\"Triage the dashboard change request for review.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240."}, {"path": ["evidence", "3"], "text": "At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 255 in the same ledger."}], "policy_evidence": [{"path": ["context"], "text": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost."}], "rules": [{"justification": "Adding a committed deliverable makes the Project sponsor the owner. The stated rationale, quantified impacts, and Delivery lead validation make the request ready. An 11–20-business-day delay gives it High delivery impact; the observed $48,000 cost does not introduce a competing cost band.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The request remains sponsor-owned and ready, but it has neither an 11–20-business-day delay nor a $50,000–$100,000 added cost. It therefore does not have High delivery impact.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240.", "negative_left": "At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240.", "negative_right": "At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 248 in the same ledger.", "right": "At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 255 in the same ledger."}, "verifier_independent_model": false}, "family": "scale-diverse-046-001", "id": "scale-diverse-046-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request has a different owner, is not ready for review, or does not have High delivery impact.", "true": "The request is sponsor-owned, ready for review, and has High delivery impact."}, "instructions": "Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.", "type": "noul"}}, "state": {"context": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.", "evidence": ["On 7 April 2026, the procurement officer confirmed that the signed procurement commitment lists the supplier-audit dashboard as a committed deliverable.", "At 09:00 UTC on 8 April 2026, Northstar Works' consecutively numbered business-day ledger recorded the supplier-audit dashboard's launch estimate before the dashboard change request as business day 240.", "At 10:00 UTC, the requester submitted the dashboard change request to add that dashboard, stating that the addition would fulfill the signed procurement commitment.", "At 11:00 UTC on 8 April 2026, the dashboard change request recorded the supplier-audit dashboard's estimated launch after the change as business day 248 in the same ledger.", "The Project manager documented exactly $48,000 in added contractor cost as the request's quantified resource impact.", "At noon, the Delivery lead checked the dependency plan and validated both the ledger-based timeline-impact estimate and the resource-impact estimate."], "request": "Triage the dashboard change request for review."}}, "method": "c2d", "provenance": {"source_id": "diverse-046", "source_is_synthetic": true, "source_sha256": "43f34b10f9d9ec2b8f7bd8ba11230d8a6bfd24bac2a882d42ea803c135c9fe6a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing routing, readiness, and delivery-impact policy without alteration or added defaults. The request remains bound to the same dashboard change, decision criteria, and review path; the two identified focus spans are complete factual sentences. The counterfactual coherently changes the audit-view verification estimate from 8 to 16 business days while retaining the exclusive, sequential, nonoverlapping two-increment structure, so it creates no duplicate or contradictory measurement. Neither context contains an answer, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual delay-range proposition rather than a policy statement. The base and counter assignments are realizable with only that range proposition changing: a validated, quantified delay can be inside or outside 11–20 business days while all other facts remain fixed. The policy evidence accurately preserves the governing ownership, readiness, and impact rules from the original state; instructions and decision criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes sponsor ownership because the request adds a committed deliverable, readiness because rationale, quantified timeline and resource impacts, and Delivery lead validation are all present, and High impact because the delay is within 11–20 business days. The exact $48,000 cost does not conflict with that delay-based High classification.", "rule_index": 0, "sound": true}, {"reason": "The conjunction still establishes sponsor ownership and readiness. Refutation of the 11–20-business-day delay range excludes the delay basis for High, while the exact $48,000 added cost excludes the $50,000–$100,000 cost basis. Therefore the request does not meet the specified High-impact classification.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The dashboard change request adds the supplier-audit dashboard."}, {"id": "a2", "statement": "The supplier-audit dashboard is a committed deliverable under the signed procurement commitment."}, {"id": "a3", "statement": "The dashboard change request states a rationale."}, {"id": "a4", "statement": "The dashboard change request has a quantified timeline impact."}, {"id": "a5", "statement": "The dashboard change request has a quantified resource impact."}, {"id": "a6", "statement": "The Delivery lead validated the dashboard change request's timeline-impact estimate."}, {"id": "a7", "statement": "The Delivery lead validated the dashboard change request's resource-impact estimate."}, {"id": "a8", "statement": "The dashboard change request's estimated launch delay is within 11–20 business days."}, {"id": "a9", "statement": "The dashboard change request's estimated added cost is $48,000."}], "base_state_json": "{\"context\":\"At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.\",\"request\":\"Determine whether the dashboard change request should be routed to the Project sponsor as ready for review with High delivery impact.\",\"evidence\":[\"The dashboard change request dated 17 September 2026 adds the supplier-audit dashboard.\",\"The signed procurement commitment lists the supplier-audit dashboard as a committed deliverable.\",\"The request states its rationale: the addition is needed to satisfy the signed procurement commitment.\",\"The request records quantified timeline and resource impacts, including an estimated $48,000 in added contractor cost.\",\"The Delivery lead separately validated the request's timeline-impact estimate and resource-impact estimate after reviewing the dependency and staffing records.\",\"In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days.\",\"The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 8 business days.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "5"], "text": "In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days."}, {"path": ["evidence", "6"], "text": "The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 8 business days."}], "policy_evidence": [{"path": ["context"], "text": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost."}], "rules": [{"justification": "Adding a committed deliverable makes the Project sponsor the owner. The stated rationale, quantified impacts, and Delivery lead validation make the request ready. An 11–20-business-day delay gives it High delivery impact; the observed $48,000 cost does not introduce a competing cost band.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The request remains sponsor-owned and ready, but it has neither an 11–20-business-day delay nor a $50,000–$100,000 added cost. It therefore does not have High delivery impact.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days.", "negative_left": "In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days.", "negative_right": "The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 16 business days.", "right": "The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 8 business days."}, "verifier_independent_model": false}, "family": "scale-diverse-046-002", "id": "scale-diverse-046-002-base", "input": {"questions": {"decision": {"criteria": {"false": "The request has a different owner, is not ready for review, or does not have High delivery impact.", "true": "The request is sponsor-owned, ready for review, and has High delivery impact."}, "instructions": "Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.", "type": "noul"}}, "state": {"context": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.", "evidence": ["The dashboard change request dated 17 September 2026 adds the supplier-audit dashboard.", "The signed procurement commitment lists the supplier-audit dashboard as a committed deliverable.", "The request states its rationale: the addition is needed to satisfy the signed procurement commitment.", "The request records quantified timeline and resource impacts, including an estimated $48,000 in added contractor cost.", "The Delivery lead separately validated the request's timeline-impact estimate and resource-impact estimate after reviewing the dependency and staffing records.", "In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days.", "The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 8 business days."], "request": "Determine whether the dashboard change request should be routed to the Project sponsor as ready for review with High delivery impact."}}, "method": "c2d", "provenance": {"source_id": "diverse-046", "source_is_synthetic": true, "source_sha256": "43f34b10f9d9ec2b8f7bd8ba11230d8a6bfd24bac2a882d42ea803c135c9fe6a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing routing, readiness, and delivery-impact policy without alteration or added defaults. The request remains bound to the same dashboard change, decision criteria, and review path; the two identified focus spans are complete factual sentences. The counterfactual coherently changes the audit-view verification estimate from 8 to 16 business days while retaining the exclusive, sequential, nonoverlapping two-increment structure, so it creates no duplicate or contradictory measurement. Neither context contains an answer, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual delay-range proposition rather than a policy statement. The base and counter assignments are realizable with only that range proposition changing: a validated, quantified delay can be inside or outside 11–20 business days while all other facts remain fixed. The policy evidence accurately preserves the governing ownership, readiness, and impact rules from the original state; instructions and decision criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes sponsor ownership because the request adds a committed deliverable, readiness because rationale, quantified timeline and resource impacts, and Delivery lead validation are all present, and High impact because the delay is within 11–20 business days. The exact $48,000 cost does not conflict with that delay-based High classification.", "rule_index": 0, "sound": true}, {"reason": "The conjunction still establishes sponsor ownership and readiness. Refutation of the 11–20-business-day delay range excludes the delay basis for High, while the exact $48,000 added cost excludes the $50,000–$100,000 cost basis. Therefore the request does not meet the specified High-impact classification.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The dashboard change request adds the supplier-audit dashboard."}, {"id": "a2", "statement": "The supplier-audit dashboard is a committed deliverable under the signed procurement commitment."}, {"id": "a3", "statement": "The dashboard change request states a rationale."}, {"id": "a4", "statement": "The dashboard change request has a quantified timeline impact."}, {"id": "a5", "statement": "The dashboard change request has a quantified resource impact."}, {"id": "a6", "statement": "The Delivery lead validated the dashboard change request's timeline-impact estimate."}, {"id": "a7", "statement": "The Delivery lead validated the dashboard change request's resource-impact estimate."}, {"id": "a8", "statement": "The dashboard change request's estimated launch delay is within 11–20 business days."}, {"id": "a9", "statement": "The dashboard change request's estimated added cost is $48,000."}], "base_state_json": "{\"context\":\"At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.\",\"request\":\"Determine whether the dashboard change request should be routed to the Project sponsor as ready for review with High delivery impact.\",\"evidence\":[\"The dashboard change request dated 17 September 2026 adds the supplier-audit dashboard.\",\"The signed procurement commitment lists the supplier-audit dashboard as a committed deliverable.\",\"The request states its rationale: the addition is needed to satisfy the signed procurement commitment.\",\"The request records quantified timeline and resource impacts, including an estimated $48,000 in added contractor cost.\",\"The Delivery lead separately validated the request's timeline-impact estimate and resource-impact estimate after reviewing the dependency and staffing records.\",\"In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days.\",\"The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 8 business days.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "5"], "text": "In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days."}, {"path": ["evidence", "6"], "text": "The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 8 business days."}], "policy_evidence": [{"path": ["context"], "text": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost."}], "rules": [{"justification": "Adding a committed deliverable makes the Project sponsor the owner. The stated rationale, quantified impacts, and Delivery lead validation make the request ready. An 11–20-business-day delay gives it High delivery impact; the observed $48,000 cost does not introduce a competing cost band.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The request remains sponsor-owned and ready, but it has neither an 11–20-business-day delay nor a $50,000–$100,000 added cost. It therefore does not have High delivery impact.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days.", "negative_left": "In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days.", "negative_right": "The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 16 business days.", "right": "The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 8 business days."}, "verifier_independent_model": false}, "family": "scale-diverse-046-002", "id": "scale-diverse-046-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request has a different owner, is not ready for review, or does not have High delivery impact.", "true": "The request is sponsor-owned, ready for review, and has High delivery impact."}, "instructions": "Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.", "type": "noul"}}, "state": {"context": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.", "evidence": ["The dashboard change request dated 17 September 2026 adds the supplier-audit dashboard.", "The signed procurement commitment lists the supplier-audit dashboard as a committed deliverable.", "The request states its rationale: the addition is needed to satisfy the signed procurement commitment.", "The request records quantified timeline and resource impacts, including an estimated $48,000 in added contractor cost.", "The Delivery lead separately validated the request's timeline-impact estimate and resource-impact estimate after reviewing the dependency and staffing records.", "In Northstar Works' dashboard change request dated 17 September 2026, the estimated launch delay consists exclusively of two sequential, nonoverlapping increments—supplier-data remapping and audit-view verification—and supplier-data remapping is estimated to take 7 business days.", "The audit-view verification increment in Northstar Works' dashboard change request dated 17 September 2026 is estimated to take 16 business days."], "request": "Determine whether the dashboard change request should be routed to the Project sponsor as ready for review with High delivery impact."}}, "method": "c2d", "provenance": {"source_id": "diverse-046", "source_is_synthetic": true, "source_sha256": "43f34b10f9d9ec2b8f7bd8ba11230d8a6bfd24bac2a882d42ea803c135c9fe6a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy, decision criteria, request, dashboard entity, calendar, and timestamp bindings. Each evidence span is a complete factual sentence; the counterfactual changes only the certified business-day count from 14 to 9, without creating an internal duplicate or contradiction, and neither context embeds an answer or classifier-output instruction beyond the original question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual delay-range proposition rather than a policy statement. The base and counter assignments are realizable with only that range proposition changing: a validated, quantified delay can be inside or outside 11–20 business days while all other facts remain fixed. The policy evidence accurately preserves the governing ownership, readiness, and impact rules from the original state; instructions and decision criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes sponsor ownership because the request adds a committed deliverable, readiness because rationale, quantified timeline and resource impacts, and Delivery lead validation are all present, and High impact because the delay is within 11–20 business days. The exact $48,000 cost does not conflict with that delay-based High classification.", "rule_index": 0, "sound": true}, {"reason": "The conjunction still establishes sponsor ownership and readiness. Refutation of the 11–20-business-day delay range excludes the delay basis for High, while the exact $48,000 added cost excludes the $50,000–$100,000 cost basis. Therefore the request does not meet the specified High-impact classification.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The dashboard change request adds the supplier-audit dashboard."}, {"id": "a2", "statement": "The supplier-audit dashboard is a committed deliverable under the signed procurement commitment."}, {"id": "a3", "statement": "The dashboard change request states a rationale."}, {"id": "a4", "statement": "The dashboard change request has a quantified timeline impact."}, {"id": "a5", "statement": "The dashboard change request has a quantified resource impact."}, {"id": "a6", "statement": "The Delivery lead validated the dashboard change request's timeline-impact estimate."}, {"id": "a7", "statement": "The Delivery lead validated the dashboard change request's resource-impact estimate."}, {"id": "a8", "statement": "The dashboard change request's estimated launch delay is within 11–20 business days."}, {"id": "a9", "statement": "The dashboard change request's estimated added cost is $48,000."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"The request has a different owner, is not ready for review, or does not have High delivery impact.\",\"true\":\"The request is sponsor-owned, ready for review, and has High delivery impact.\"},\"instructions\":\"Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.\",\"type\":\"noul\"}},\"state\":{\"context\":\"At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.\",\"handoff\":\"The signed procurement commitment identifies the supplier-audit dashboard as a committed deliverable. The dashboard change request adds that dashboard and states that satisfying the signed procurement commitment is the rationale for the addition. It quantifies the resource impact as $48,000 in added contractor cost. The Delivery lead reviewed the dependency and staffing plans and validated both the timeline-impact estimate and the resource-impact estimate.\",\"evidence\":[\"The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031.\",\"The certified ledger for Northstar Works calendar NW-27 records 14 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031.\"],\"request\":\"Triage the dashboard change request for review.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031."}, {"path": ["state", "evidence", "1"], "text": "The certified ledger for Northstar Works calendar NW-27 records 14 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031."}], "policy_evidence": [{"path": ["context"], "text": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost."}], "rules": [{"justification": "Adding a committed deliverable makes the Project sponsor the owner. The stated rationale, quantified impacts, and Delivery lead validation make the request ready. An 11–20-business-day delay gives it High delivery impact; the observed $48,000 cost does not introduce a competing cost band.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The request remains sponsor-owned and ready, but it has neither an 11–20-business-day delay nor a $50,000–$100,000 added cost. It therefore does not have High delivery impact.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031.", "negative_left": "The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031.", "negative_right": "The certified ledger for Northstar Works calendar NW-27 records 9 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031.", "right": "The certified ledger for Northstar Works calendar NW-27 records 14 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031."}, "verifier_independent_model": false}, "family": "scale-diverse-046-003", "id": "scale-diverse-046-003-base", "input": {"questions": {"decision": {"criteria": {"false": "The request has a different owner, is not ready for review, or does not have High delivery impact.", "true": "The request is sponsor-owned, ready for review, and has High delivery impact."}, "instructions": "Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "The request has a different owner, is not ready for review, or does not have High delivery impact.", "true": "The request is sponsor-owned, ready for review, and has High delivery impact."}, "instructions": "Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.", "type": "noul"}}, "state": {"context": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.", "evidence": ["The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031.", "The certified ledger for Northstar Works calendar NW-27 records 14 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031."], "handoff": "The signed procurement commitment identifies the supplier-audit dashboard as a committed deliverable. The dashboard change request adds that dashboard and states that satisfying the signed procurement commitment is the rationale for the addition. It quantifies the resource impact as $48,000 in added contractor cost. The Delivery lead reviewed the dependency and staffing plans and validated both the timeline-impact estimate and the resource-impact estimate.", "request": "Triage the dashboard change request for review."}}}, "method": "c2d", "provenance": {"source_id": "diverse-046", "source_is_synthetic": true, "source_sha256": "43f34b10f9d9ec2b8f7bd8ba11230d8a6bfd24bac2a882d42ea803c135c9fe6a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy, decision criteria, request, dashboard entity, calendar, and timestamp bindings. Each evidence span is a complete factual sentence; the counterfactual changes only the certified business-day count from 14 to 9, without creating an internal duplicate or contradiction, and neither context embeds an answer or classifier-output instruction beyond the original question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual delay-range proposition rather than a policy statement. The base and counter assignments are realizable with only that range proposition changing: a validated, quantified delay can be inside or outside 11–20 business days while all other facts remain fixed. The policy evidence accurately preserves the governing ownership, readiness, and impact rules from the original state; instructions and decision criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes sponsor ownership because the request adds a committed deliverable, readiness because rationale, quantified timeline and resource impacts, and Delivery lead validation are all present, and High impact because the delay is within 11–20 business days. The exact $48,000 cost does not conflict with that delay-based High classification.", "rule_index": 0, "sound": true}, {"reason": "The conjunction still establishes sponsor ownership and readiness. Refutation of the 11–20-business-day delay range excludes the delay basis for High, while the exact $48,000 added cost excludes the $50,000–$100,000 cost basis. Therefore the request does not meet the specified High-impact classification.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The dashboard change request adds the supplier-audit dashboard."}, {"id": "a2", "statement": "The supplier-audit dashboard is a committed deliverable under the signed procurement commitment."}, {"id": "a3", "statement": "The dashboard change request states a rationale."}, {"id": "a4", "statement": "The dashboard change request has a quantified timeline impact."}, {"id": "a5", "statement": "The dashboard change request has a quantified resource impact."}, {"id": "a6", "statement": "The Delivery lead validated the dashboard change request's timeline-impact estimate."}, {"id": "a7", "statement": "The Delivery lead validated the dashboard change request's resource-impact estimate."}, {"id": "a8", "statement": "The dashboard change request's estimated launch delay is within 11–20 business days."}, {"id": "a9", "statement": "The dashboard change request's estimated added cost is $48,000."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"The request has a different owner, is not ready for review, or does not have High delivery impact.\",\"true\":\"The request is sponsor-owned, ready for review, and has High delivery impact.\"},\"instructions\":\"Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.\",\"type\":\"noul\"}},\"state\":{\"context\":\"At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.\",\"handoff\":\"The signed procurement commitment identifies the supplier-audit dashboard as a committed deliverable. The dashboard change request adds that dashboard and states that satisfying the signed procurement commitment is the rationale for the addition. It quantifies the resource impact as $48,000 in added contractor cost. The Delivery lead reviewed the dependency and staffing plans and validated both the timeline-impact estimate and the resource-impact estimate.\",\"evidence\":[\"The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031.\",\"The certified ledger for Northstar Works calendar NW-27 records 14 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031.\"],\"request\":\"Triage the dashboard change request for review.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031."}, {"path": ["state", "evidence", "1"], "text": "The certified ledger for Northstar Works calendar NW-27 records 14 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031."}], "policy_evidence": [{"path": ["context"], "text": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost."}], "rules": [{"justification": "Adding a committed deliverable makes the Project sponsor the owner. The stated rationale, quantified impacts, and Delivery lead validation make the request ready. An 11–20-business-day delay gives it High delivery impact; the observed $48,000 cost does not introduce a competing cost band.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The request remains sponsor-owned and ready, but it has neither an 11–20-business-day delay nor a $50,000–$100,000 added cost. It therefore does not have High delivery impact.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031.", "negative_left": "The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031.", "negative_right": "The certified ledger for Northstar Works calendar NW-27 records 9 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031.", "right": "The certified ledger for Northstar Works calendar NW-27 records 14 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031."}, "verifier_independent_model": false}, "family": "scale-diverse-046-003", "id": "scale-diverse-046-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request has a different owner, is not ready for review, or does not have High delivery impact.", "true": "The request is sponsor-owned, ready for review, and has High delivery impact."}, "instructions": "Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "The request has a different owner, is not ready for review, or does not have High delivery impact.", "true": "The request is sponsor-owned, ready for review, and has High delivery impact."}, "instructions": "Decide whether the request should be routed to the Project sponsor as ready for review with High delivery impact. Answer Yes or No using the stated routing, readiness, and severity rules.", "type": "noul"}}, "state": {"context": "At Northstar Works, scope changes go to the Project sponsor if they add or remove a committed deliverable or delay launch by more than 10 business days; otherwise, the Project manager owns review. A request is ready only with a stated rationale, quantified timeline and resource impacts, and Delivery lead validation. Delivery impact is ordered Low, Moderate, High, Critical; High means an 11–20-day delay or $50,000–$100,000 added cost.", "evidence": ["The dashboard change request calculates its estimated launch delay on Northstar Works calendar NW-27 from the baseline launch timestamp of 09:00 on March 3, 2031, to the revised launch timestamp of 09:00 on March 21, 2031.", "The certified ledger for Northstar Works calendar NW-27 records 9 business days as elapsing from 09:00 on March 3, 2031, to 09:00 on March 21, 2031."], "handoff": "The signed procurement commitment identifies the supplier-audit dashboard as a committed deliverable. The dashboard change request adds that dashboard and states that satisfying the signed procurement commitment is the rationale for the addition. It quantifies the resource impact as $48,000 in added contractor cost. The Delivery lead reviewed the dependency and staffing plans and validated both the timeline-impact estimate and the resource-impact estimate.", "request": "Triage the dashboard change request for review."}}}, "method": "c2d", "provenance": {"source_id": "diverse-046", "source_is_synthetic": true, "source_sha256": "43f34b10f9d9ec2b8f7bd8ba11230d8a6bfd24bac2a882d42ea803c135c9fe6a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original delivery-impact criteria, readiness requirements, and review-ownership rule without adding exceptions or defaults. They remain tied to the same encryption-and-key-rotation request and quantified schedule, cost, staffing, and contractual-milestone path; the changed completion date is a permissible case observation. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the projected completion date, remains consistent with the stated ten-business-day delay because no contradictory baseline date is supplied, and creates no duplicate conflicting measurement. Neither context contains a score, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Contract administrator\",\"text\":\"The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026.\"},{\"speaker\":\"Impact analyst\",\"text\":\"The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 30 October 2026.\"},{\"speaker\":\"Configuration-control officer\",\"text\":\"On 15 September 2026, configuration-control staff verified that the signed contract and dated impact analysis concern the same pending request. The analysis is the final applicable impact assessment, and the file contains no superseding assessment or contract amendment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026."}, {"path": ["1", "text"], "text": "The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 30 October 2026."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026.", "negative_left": "The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026.", "negative_right": "The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 23 October 2026.", "right": "The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 30 October 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-047-001", "id": "scale-diverse-047-001-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Contract administrator", "text": "The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026."}, {"speaker": "Impact analyst", "text": "The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 30 October 2026."}, {"speaker": "Configuration-control officer", "text": "On 15 September 2026, configuration-control staff verified that the signed contract and dated impact analysis concern the same pending request. The analysis is the final applicable impact assessment, and the file contains no superseding assessment or contract amendment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original delivery-impact criteria, readiness requirements, and review-ownership rule without adding exceptions or defaults. They remain tied to the same encryption-and-key-rotation request and quantified schedule, cost, staffing, and contractual-milestone path; the changed completion date is a permissible case observation. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the projected completion date, remains consistent with the stated ten-business-day delay because no contradictory baseline date is supplied, and creates no duplicate conflicting measurement. Neither context contains a score, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Contract administrator\",\"text\":\"The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026.\"},{\"speaker\":\"Impact analyst\",\"text\":\"The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 30 October 2026.\"},{\"speaker\":\"Configuration-control officer\",\"text\":\"On 15 September 2026, configuration-control staff verified that the signed contract and dated impact analysis concern the same pending request. The analysis is the final applicable impact assessment, and the file contains no superseding assessment or contract amendment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026."}, {"path": ["1", "text"], "text": "The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 30 October 2026."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026.", "negative_left": "The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026.", "negative_right": "The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 23 October 2026.", "right": "The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 30 October 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-047-001", "id": "scale-diverse-047-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Contract administrator", "text": "The contract signed on 2 September 2026 contains exactly one contractual milestone for the encryption-and-key-rotation change request and stipulates that the milestone is breached if and only if production deployment is not completed by 5:00 p.m. UTC on 26 October 2026."}, {"speaker": "Impact analyst", "text": "The impact analysis dated 14 September 2026 documents the encryption-and-key-rotation rationale and quantifies implementation as an exact 10-business-day delay, an exact $15,000 increase against the $300,000 baseline, and the assignment of two engineers, placing production-deployment completion at 5:00 p.m. UTC on 23 October 2026."}, {"speaker": "Configuration-control officer", "text": "On 15 September 2026, configuration-control staff verified that the signed contract and dated impact analysis concern the same pending request. The analysis is the final applicable impact assessment, and the file contains no superseding assessment or contract amendment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original readiness, ownership, and delivery-impact policies remain unchanged and apply equally to both contexts. The request scope and decision path are preserved, while only the case-specific handoff date changes. The two focus-evidence spans are complete factual sentences. Moving the handoff from 22 October to 20 October is coherent with the unchanged 21 October contractual deadline and creates no duplicate or contradictory measurement. Neither context contains an answer code, explicit classification, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change coordinator\",\"text\":\"The review packet documents the rationale for the encryption-and-key-rotation request: a compliance audit identified missing field-level encryption in the payment export.\"},{\"speaker\":\"Planning analyst\",\"text\":\"The applicable schedule impact is quantified as a delay of exactly 10 business days.\"},{\"speaker\":\"Finance analyst\",\"text\":\"The applicable cost impact is quantified as exactly $15,000, equal to exactly 5% of the $300,000 baseline.\"},{\"speaker\":\"Engineering manager\",\"text\":\"The applicable resource impact is quantified as the temporary assignment of two engineers for implementation and validation.\"},{\"speaker\":\"Handoff lead\",\"text\":\"For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 22 October 2026, and that handoff is the project's sole contractual milestone.\"},{\"speaker\":\"Contract administrator\",\"text\":\"The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["4", "text"], "text": "For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 22 October 2026, and that handoff is the project's sole contractual milestone."}, {"path": ["5", "text"], "text": "The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 22 October 2026, and that handoff is the project's sole contractual milestone.", "negative_left": "For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 20 October 2026, and that handoff is the project's sole contractual milestone.", "negative_right": "The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026.", "right": "The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-047-003", "id": "scale-diverse-047-003-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change coordinator", "text": "The review packet documents the rationale for the encryption-and-key-rotation request: a compliance audit identified missing field-level encryption in the payment export."}, {"speaker": "Planning analyst", "text": "The applicable schedule impact is quantified as a delay of exactly 10 business days."}, {"speaker": "Finance analyst", "text": "The applicable cost impact is quantified as exactly $15,000, equal to exactly 5% of the $300,000 baseline."}, {"speaker": "Engineering manager", "text": "The applicable resource impact is quantified as the temporary assignment of two engineers for implementation and validation."}, {"speaker": "Handoff lead", "text": "For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 22 October 2026, and that handoff is the project's sole contractual milestone."}, {"speaker": "Contract administrator", "text": "The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original readiness, ownership, and delivery-impact policies remain unchanged and apply equally to both contexts. The request scope and decision path are preserved, while only the case-specific handoff date changes. The two focus-evidence spans are complete factual sentences. Moving the handoff from 22 October to 20 October is coherent with the unchanged 21 October contractual deadline and creates no duplicate or contradictory measurement. Neither context contains an answer code, explicit classification, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change coordinator\",\"text\":\"The review packet documents the rationale for the encryption-and-key-rotation request: a compliance audit identified missing field-level encryption in the payment export.\"},{\"speaker\":\"Planning analyst\",\"text\":\"The applicable schedule impact is quantified as a delay of exactly 10 business days.\"},{\"speaker\":\"Finance analyst\",\"text\":\"The applicable cost impact is quantified as exactly $15,000, equal to exactly 5% of the $300,000 baseline.\"},{\"speaker\":\"Engineering manager\",\"text\":\"The applicable resource impact is quantified as the temporary assignment of two engineers for implementation and validation.\"},{\"speaker\":\"Handoff lead\",\"text\":\"For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 22 October 2026, and that handoff is the project's sole contractual milestone.\"},{\"speaker\":\"Contract administrator\",\"text\":\"The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["4", "text"], "text": "For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 22 October 2026, and that handoff is the project's sole contractual milestone."}, {"path": ["5", "text"], "text": "The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 22 October 2026, and that handoff is the project's sole contractual milestone.", "negative_left": "For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 20 October 2026, and that handoff is the project's sole contractual milestone.", "negative_right": "The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026.", "right": "The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-047-003", "id": "scale-diverse-047-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change coordinator", "text": "The review packet documents the rationale for the encryption-and-key-rotation request: a compliance audit identified missing field-level encryption in the payment export."}, {"speaker": "Planning analyst", "text": "The applicable schedule impact is quantified as a delay of exactly 10 business days."}, {"speaker": "Finance analyst", "text": "The applicable cost impact is quantified as exactly $15,000, equal to exactly 5% of the $300,000 baseline."}, {"speaker": "Engineering manager", "text": "The applicable resource impact is quantified as the temporary assignment of two engineers for implementation and validation."}, {"speaker": "Handoff lead", "text": "For Project Nimbus, implementing the encryption-and-key-rotation change request would move the operational handoff to 17:00 UTC on 20 October 2026, and that handoff is the project's sole contractual milestone."}, {"speaker": "Contract administrator", "text": "The signed Project Nimbus contract requires the operational handoff to occur no later than 17:00 UTC on 21 October 2026."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing criteria and request condition without adding exceptions, priorities, or missing-evidence defaults. Lina, Birch’s API, the 12 September deadline, the certification timing path, and the request’s delivery effects remain bound consistently; only the case observation for the certification day changes from 18 to 8. The two evidence spans are complete factual sentences. The counterfactual date of 8 September 2026 coherently precedes the unchanged 12 September deadline and does not conflict with another certification measurement. Neither context embeds an impact-level answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026. The vendor's confirmation record gives 18 as the day of the month for Birch's API certification. The request states that it remains active when certification occurs no later than the recorded deadline and is withdrawn when certification does not occur by that deadline. Lina identifies accelerating Birch’s mobile release as the rationale. The proposed change’s greatest stated milestone delay is three business days, and its longest stated resource reassignment is two weeks. It would not alter the core project outcome or require additional funding. The requested QA staff transfer has not been initiated. No schedule update, spending, or scope change for Lina’s request has been initiated.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026."}, {"path": [], "text": "The vendor's confirmation record gives 18 as the day of the month for Birch's API certification."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026.", "negative_left": "Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026.", "negative_right": "The vendor's confirmation record gives 8 as the day of the month for Birch's API certification.", "right": "The vendor's confirmation record gives 18 as the day of the month for Birch's API certification."}, "verifier_independent_model": false}, "family": "scale-diverse-048-002", "id": "scale-diverse-048-002-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026. The vendor's confirmation record gives 18 as the day of the month for Birch's API certification. The request states that it remains active when certification occurs no later than the recorded deadline and is withdrawn when certification does not occur by that deadline. Lina identifies accelerating Birch’s mobile release as the rationale. The proposed change’s greatest stated milestone delay is three business days, and its longest stated resource reassignment is two weeks. It would not alter the core project outcome or require additional funding. The requested QA staff transfer has not been initiated. No schedule update, spending, or scope change for Lina’s request has been initiated."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing criteria and request condition without adding exceptions, priorities, or missing-evidence defaults. Lina, Birch’s API, the 12 September deadline, the certification timing path, and the request’s delivery effects remain bound consistently; only the case observation for the certification day changes from 18 to 8. The two evidence spans are complete factual sentences. The counterfactual date of 8 September 2026 coherently precedes the unchanged 12 September deadline and does not conflict with another certification measurement. Neither context embeds an impact-level answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026. The vendor's confirmation record gives 18 as the day of the month for Birch's API certification. The request states that it remains active when certification occurs no later than the recorded deadline and is withdrawn when certification does not occur by that deadline. Lina identifies accelerating Birch’s mobile release as the rationale. The proposed change’s greatest stated milestone delay is three business days, and its longest stated resource reassignment is two weeks. It would not alter the core project outcome or require additional funding. The requested QA staff transfer has not been initiated. No schedule update, spending, or scope change for Lina’s request has been initiated.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026."}, {"path": [], "text": "The vendor's confirmation record gives 18 as the day of the month for Birch's API certification."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026.", "negative_left": "Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026.", "negative_right": "The vendor's confirmation record gives 8 as the day of the month for Birch's API certification.", "right": "The vendor's confirmation record gives 18 as the day of the month for Birch's API certification."}, "verifier_independent_model": false}, "family": "scale-diverse-048-002", "id": "scale-diverse-048-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Lina's request records its 12 September deadline and the vendor's confirmed certification date for Birch's API as dates in September 2026. The vendor's confirmation record gives 8 as the day of the month for Birch's API certification. The request states that it remains active when certification occurs no later than the recorded deadline and is withdrawn when certification does not occur by that deadline. Lina identifies accelerating Birch’s mobile release as the rationale. The proposed change’s greatest stated milestone delay is three business days, and its longest stated resource reassignment is two weeks. It would not alter the core project outcome or require additional funding. The requested QA staff transfer has not been initiated. No schedule update, spending, or scope change for Lina’s request has been initiated."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve Lina’s Birch request, the API-certification trigger and 12 September deadline, the proposed Cedar QA reassignment, and the governing active/withdrawn condition without adding policy exceptions or output instructions. The two evidence spans are complete factual sentences; changing the confirmed certification date from 16 September to 10 September is coherent with the otherwise unchanged facts, and neither context contains a gold score, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Operational handoff note for Lina’s proposed Project Birch change. The vendor's signed operational handoff for Birch's API records its confirmed certification date as 16 September 2026. Lina's request records the certification deadline for Birch's API as 12 September 2026. The request says it remains active when certification occurs no later than that recorded deadline and is withdrawn when the vendor does not certify the API by it. Lina gives accelerating Birch’s mobile release as her rationale. The greatest milestone delay stated for the proposal is three business days, and the longest resource reassignment stated is two weeks. The proposal would not alter the core project outcome and requires no additional funding. The requested QA staff transfer from Project Cedar has not been initiated. No schedule update, spending, or scope change for the proposal has been initiated.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The vendor's signed operational handoff for Birch's API records its confirmed certification date as 16 September 2026."}, {"path": [], "text": "Lina's request records the certification deadline for Birch's API as 12 September 2026."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The vendor's signed operational handoff for Birch's API records its confirmed certification date as 16 September 2026.", "negative_left": "The vendor's signed operational handoff for Birch's API records its confirmed certification date as 10 September 2026.", "negative_right": "Lina's request records the certification deadline for Birch's API as 12 September 2026.", "right": "Lina's request records the certification deadline for Birch's API as 12 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-048-003", "id": "scale-diverse-048-003-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Operational handoff note for Lina’s proposed Project Birch change. The vendor's signed operational handoff for Birch's API records its confirmed certification date as 16 September 2026. Lina's request records the certification deadline for Birch's API as 12 September 2026. The request says it remains active when certification occurs no later than that recorded deadline and is withdrawn when the vendor does not certify the API by it. Lina gives accelerating Birch’s mobile release as her rationale. The greatest milestone delay stated for the proposal is three business days, and the longest resource reassignment stated is two weeks. The proposal would not alter the core project outcome and requires no additional funding. The requested QA staff transfer from Project Cedar has not been initiated. No schedule update, spending, or scope change for the proposal has been initiated."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve Lina’s Birch request, the API-certification trigger and 12 September deadline, the proposed Cedar QA reassignment, and the governing active/withdrawn condition without adding policy exceptions or output instructions. The two evidence spans are complete factual sentences; changing the confirmed certification date from 16 September to 10 September is coherent with the otherwise unchanged facts, and neither context contains a gold score, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Operational handoff note for Lina’s proposed Project Birch change. The vendor's signed operational handoff for Birch's API records its confirmed certification date as 16 September 2026. Lina's request records the certification deadline for Birch's API as 12 September 2026. The request says it remains active when certification occurs no later than that recorded deadline and is withdrawn when the vendor does not certify the API by it. Lina gives accelerating Birch’s mobile release as her rationale. The greatest milestone delay stated for the proposal is three business days, and the longest resource reassignment stated is two weeks. The proposal would not alter the core project outcome and requires no additional funding. The requested QA staff transfer from Project Cedar has not been initiated. No schedule update, spending, or scope change for the proposal has been initiated.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The vendor's signed operational handoff for Birch's API records its confirmed certification date as 16 September 2026."}, {"path": [], "text": "Lina's request records the certification deadline for Birch's API as 12 September 2026."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The vendor's signed operational handoff for Birch's API records its confirmed certification date as 16 September 2026.", "negative_left": "The vendor's signed operational handoff for Birch's API records its confirmed certification date as 10 September 2026.", "negative_right": "Lina's request records the certification deadline for Birch's API as 12 September 2026.", "right": "Lina's request records the certification deadline for Birch's API as 12 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-048-003", "id": "scale-diverse-048-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Operational handoff note for Lina’s proposed Project Birch change. The vendor's signed operational handoff for Birch's API records its confirmed certification date as 10 September 2026. Lina's request records the certification deadline for Birch's API as 12 September 2026. The request says it remains active when certification occurs no later than that recorded deadline and is withdrawn when the vendor does not certify the API by it. Lina gives accelerating Birch’s mobile release as her rationale. The greatest milestone delay stated for the proposal is three business days, and the longest resource reassignment stated is two weeks. The proposal would not alter the core project outcome and requires no additional funding. The requested QA staff transfer from Project Cedar has not been initiated. No schedule update, spending, or scope change for the proposal has been initiated."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the acceptance, routing, and quality policies and keep the question focused on the submitted completed “Quarterly Export” work package at its acceptance review. The two evidence spans are complete factual sentences. The counterfactual coherently changes D-9’s custody code so it no longer matches A-17’s; this does not conflict with the unchanged one-code-per-document and no-code-sharing ledger assertions. Neither context embeds an answer option, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "full_context_fact_states": {"base": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "counterfactual": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "remove_left": {"attached_registry_document_identity": "unknown"}, "remove_right": {"attached_registry_document_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"attached_registry_document_identity": "unknown"}, "negative_pair": {"attached_registry_document_identity": "refuted"}, "negative_sentence": {"attached_registry_document_identity": "unknown"}, "positive_pair": {"attached_registry_document_identity": "supported"}, "right": {"attached_registry_document_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual status, identity, attachment, or explicitly scoped universal relationship. The focus is the factual identity of two documents, not a policy conclusion. The base and counter assignments are jointly realizable with only that identity changing: in the counter, D-9 can be valid but unattached while A-17 is attached but not valid. The policy evidence correctly cites substantive rules from the original state; governing material in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A test defect is present. A-17 is attached and identical to D-9, which is a valid signed business approval, so valid approval is attached. The criteria therefore require rejection, routing only to the delivery owner, and a medium rating.", "rule_index": 0, "sound": true}, {"reason": "A test defect is present. If any valid signed approval were attached, the two universal restrictions would make it both D-9 and A-17, contradicting the supported non-identity of those documents. Thus no valid signed approval is attached, so both deficiencies are present and the none_of_above outcome is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "test_defect", "statement": "At the acceptance review of the submitted completed “Quarterly Export” work package, at least one of the 12 required tests failed."}, {"id": "a17_attached", "statement": "At the acceptance review, document A-17 is attached to the submitted completed “Quarterly Export” work package."}, {"id": "a17_only_attached_approval_candidate", "statement": "At the acceptance review, every business-approval document attached to the submitted completed “Quarterly Export” work package is document A-17."}, {"id": "d9_valid_approval", "statement": "At the acceptance review, registry document D-9 is a valid signed business approval for the submitted completed “Quarterly Export” work package."}, {"id": "d9_only_valid_approval", "statement": "At the acceptance review, every valid signed business approval for the submitted completed “Quarterly Export” work package is registry document D-9."}, {"id": "attached_registry_document_identity", "statement": "Document A-17 attached to the submitted completed “Quarterly Export” work package is the same document as registry document D-9."}], "base_state_json": "\"At the acceptance review, the signed test report records that 11 required tests passed and the CSV export test failed. The controlled attachment manifest lists A-17 as attached to the submitted completed “Quarterly Export” work package and identifies it as the only attached business-approval document. The approval registry records D-9 as a valid signed business approval for that work package and confirms that no other document is a valid signed business approval for it.\\n\\nAt the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314. In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-7314.\\n\\nAcceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both.\"", "base_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}], "counter_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}], "focus_atom": "attached_registry_document_identity", "focus_evidence": [{"path": [], "text": "At the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314."}, {"path": [], "text": "In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-7314."}], "policy_evidence": [{"path": [], "text": "Acceptance requires all 12 tests to pass and a signed business approval to be attached."}, {"path": [], "text": "Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver."}, {"path": [], "text": "Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}], "rules": [{"justification": "The failed test establishes a test defect. Because attached document A-17 is identical to D-9, which is a valid signed business approval for this work package, valid approval is attached. The rubric therefore requires rejection, routing only to the delivery owner, and a medium rating.", "target": "reject_route_delivery_owner_medium", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}]}, {"justification": "The failed test establishes a test defect. A-17 is the only attached business-approval document, D-9 is the only valid signed approval for this work package, and the two documents are explicitly different; therefore no valid signed business approval is attached. Both deficiencies are present, requiring rejection, routing to both responsible roles, and a low rating.", "target": "none_of_above", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314.", "negative_left": "At the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314.", "negative_right": "In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-8826.", "right": "In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-7314."}, "verifier_independent_model": false}, "family": "scale-diverse-055-002", "id": "scale-diverse-055-002-base", "input": {"questions": {"decision": {"criteria": {"accept_high_quality": "Accept and rate high only when all 12 tests pass and valid signed business approval is attached; no corrective routing is required.", "none_of_above": "Choose when both a test defect and missing valid business approval are present; reject, route to both responsible roles, and rate completion quality low.", "reject_route_business_approver_medium": "Reject, route only to the business approver, and rate medium when all 12 tests pass but valid signed business approval is missing.", "reject_route_delivery_owner_medium": "Reject, route only to the delivery owner, and rate medium when at least one test fails but valid signed business approval is attached."}, "instructions": "Select the single option whose rubric matches the evidence, including the acceptance decision, responsible routing, and completion-quality rating.", "type": "choice"}}, "state": "At the acceptance review, the signed test report records that 11 required tests passed and the CSV export test failed. The controlled attachment manifest lists A-17 as attached to the submitted completed “Quarterly Export” work package and identifies it as the only attached business-approval document. The approval registry records D-9 as a valid signed business approval for that work package and confirms that no other document is a valid signed business approval for it.\n\nAt the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314. In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-7314.\n\nAcceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}, "method": "c2d", "provenance": {"source_id": "diverse-055", "source_is_synthetic": true, "source_sha256": "16b728b9f77e4bb47a594dca1c07cd310327dbb290c6117b5f7d03d56d1e3e36", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "reject_route_delivery_owner_medium"}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the acceptance, routing, and quality policies and keep the question focused on the submitted completed “Quarterly Export” work package at its acceptance review. The two evidence spans are complete factual sentences. The counterfactual coherently changes D-9’s custody code so it no longer matches A-17’s; this does not conflict with the unchanged one-code-per-document and no-code-sharing ledger assertions. Neither context embeds an answer option, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "full_context_fact_states": {"base": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "counterfactual": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "remove_left": {"attached_registry_document_identity": "unknown"}, "remove_right": {"attached_registry_document_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"attached_registry_document_identity": "unknown"}, "negative_pair": {"attached_registry_document_identity": "refuted"}, "negative_sentence": {"attached_registry_document_identity": "unknown"}, "positive_pair": {"attached_registry_document_identity": "supported"}, "right": {"attached_registry_document_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual status, identity, attachment, or explicitly scoped universal relationship. The focus is the factual identity of two documents, not a policy conclusion. The base and counter assignments are jointly realizable with only that identity changing: in the counter, D-9 can be valid but unattached while A-17 is attached but not valid. The policy evidence correctly cites substantive rules from the original state; governing material in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A test defect is present. A-17 is attached and identical to D-9, which is a valid signed business approval, so valid approval is attached. The criteria therefore require rejection, routing only to the delivery owner, and a medium rating.", "rule_index": 0, "sound": true}, {"reason": "A test defect is present. If any valid signed approval were attached, the two universal restrictions would make it both D-9 and A-17, contradicting the supported non-identity of those documents. Thus no valid signed approval is attached, so both deficiencies are present and the none_of_above outcome is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "test_defect", "statement": "At the acceptance review of the submitted completed “Quarterly Export” work package, at least one of the 12 required tests failed."}, {"id": "a17_attached", "statement": "At the acceptance review, document A-17 is attached to the submitted completed “Quarterly Export” work package."}, {"id": "a17_only_attached_approval_candidate", "statement": "At the acceptance review, every business-approval document attached to the submitted completed “Quarterly Export” work package is document A-17."}, {"id": "d9_valid_approval", "statement": "At the acceptance review, registry document D-9 is a valid signed business approval for the submitted completed “Quarterly Export” work package."}, {"id": "d9_only_valid_approval", "statement": "At the acceptance review, every valid signed business approval for the submitted completed “Quarterly Export” work package is registry document D-9."}, {"id": "attached_registry_document_identity", "statement": "Document A-17 attached to the submitted completed “Quarterly Export” work package is the same document as registry document D-9."}], "base_state_json": "\"At the acceptance review, the signed test report records that 11 required tests passed and the CSV export test failed. The controlled attachment manifest lists A-17 as attached to the submitted completed “Quarterly Export” work package and identifies it as the only attached business-approval document. The approval registry records D-9 as a valid signed business approval for that work package and confirms that no other document is a valid signed business approval for it.\\n\\nAt the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314. In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-7314.\\n\\nAcceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both.\"", "base_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}], "counter_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}], "focus_atom": "attached_registry_document_identity", "focus_evidence": [{"path": [], "text": "At the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314."}, {"path": [], "text": "In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-7314."}], "policy_evidence": [{"path": [], "text": "Acceptance requires all 12 tests to pass and a signed business approval to be attached."}, {"path": [], "text": "Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver."}, {"path": [], "text": "Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}], "rules": [{"justification": "The failed test establishes a test defect. Because attached document A-17 is identical to D-9, which is a valid signed business approval for this work package, valid approval is attached. The rubric therefore requires rejection, routing only to the delivery owner, and a medium rating.", "target": "reject_route_delivery_owner_medium", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}]}, {"justification": "The failed test establishes a test defect. A-17 is the only attached business-approval document, D-9 is the only valid signed approval for this work package, and the two documents are explicitly different; therefore no valid signed business approval is attached. Both deficiencies are present, requiring rejection, routing to both responsible roles, and a low rating.", "target": "none_of_above", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314.", "negative_left": "At the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314.", "negative_right": "In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-8826.", "right": "In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-7314."}, "verifier_independent_model": false}, "family": "scale-diverse-055-002", "id": "scale-diverse-055-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_high_quality": "Accept and rate high only when all 12 tests pass and valid signed business approval is attached; no corrective routing is required.", "none_of_above": "Choose when both a test defect and missing valid business approval are present; reject, route to both responsible roles, and rate completion quality low.", "reject_route_business_approver_medium": "Reject, route only to the business approver, and rate medium when all 12 tests pass but valid signed business approval is missing.", "reject_route_delivery_owner_medium": "Reject, route only to the delivery owner, and rate medium when at least one test fails but valid signed business approval is attached."}, "instructions": "Select the single option whose rubric matches the evidence, including the acceptance decision, responsible routing, and completion-quality rating.", "type": "choice"}}, "state": "At the acceptance review, the signed test report records that 11 required tests passed and the CSV export test failed. The controlled attachment manifest lists A-17 as attached to the submitted completed “Quarterly Export” work package and identifies it as the only attached business-approval document. The approval registry records D-9 as a valid signed business approval for that work package and confirms that no other document is a valid signed business approval for it.\n\nAt the 2026-05-14 16:22 UTC acceptance review, the finalized reconciliation ledger for the submitted completed “Quarterly Export” work package assigned exactly one custody code to every listed document and no custody code to more than one listed document, and its entry for attached document A-17 gave code QX-7314. In that same finalized reconciliation ledger at the 2026-05-14 16:22 UTC acceptance review, registry document D-9 was listed with custody code QX-8826.\n\nAcceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}, "method": "c2d", "provenance": {"source_id": "diverse-055", "source_is_synthetic": true, "source_sha256": "16b728b9f77e4bb47a594dca1c07cd310327dbb290c6117b5f7d03d56d1e3e36", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing acceptance, routing, and quality policies from the original state, while the unchanged questions preserve the complete decision rubric. They maintain the same work package, review stage, criteria, roles, and handoff time. The two evidence spans are complete factual sentences. The counterfactual changes only D-9’s immutable identifier, coherently distinguishing it from attached A-17 under the stated one-to-one identifier assignment without creating duplicate or contradictory measurements. Neither context supplies an answer choice, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "full_context_fact_states": {"base": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "counterfactual": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "remove_left": {"attached_registry_document_identity": "unknown"}, "remove_right": {"attached_registry_document_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"attached_registry_document_identity": "unknown"}, "negative_pair": {"attached_registry_document_identity": "refuted"}, "negative_sentence": {"attached_registry_document_identity": "unknown"}, "positive_pair": {"attached_registry_document_identity": "supported"}, "right": {"attached_registry_document_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual status, identity, attachment, or explicitly scoped universal relationship. The focus is the factual identity of two documents, not a policy conclusion. The base and counter assignments are jointly realizable with only that identity changing: in the counter, D-9 can be valid but unattached while A-17 is attached but not valid. The policy evidence correctly cites substantive rules from the original state; governing material in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A test defect is present. A-17 is attached and identical to D-9, which is a valid signed business approval, so valid approval is attached. The criteria therefore require rejection, routing only to the delivery owner, and a medium rating.", "rule_index": 0, "sound": true}, {"reason": "A test defect is present. If any valid signed approval were attached, the two universal restrictions would make it both D-9 and A-17, contradicting the supported non-identity of those documents. Thus no valid signed approval is attached, so both deficiencies are present and the none_of_above outcome is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "test_defect", "statement": "At the acceptance review of the submitted completed “Quarterly Export” work package, at least one of the 12 required tests failed."}, {"id": "a17_attached", "statement": "At the acceptance review, document A-17 is attached to the submitted completed “Quarterly Export” work package."}, {"id": "a17_only_attached_approval_candidate", "statement": "At the acceptance review, every business-approval document attached to the submitted completed “Quarterly Export” work package is document A-17."}, {"id": "d9_valid_approval", "statement": "At the acceptance review, registry document D-9 is a valid signed business approval for the submitted completed “Quarterly Export” work package."}, {"id": "d9_only_valid_approval", "statement": "At the acceptance review, every valid signed business approval for the submitted completed “Quarterly Export” work package is registry document D-9."}, {"id": "attached_registry_document_identity", "statement": "Document A-17 attached to the submitted completed “Quarterly Export” work package is the same document as registry document D-9."}], "base_state_json": "\"At acceptance review, the project coordinator submitted the completed “Quarterly Export” work package. The signed test report covers all 12 required tests and records 11 passes; the CSV export test failed. The attachment manifest lists document A-17 as the sole business-approval document attached to the submitted completed work package. Registry document D-9 is a valid signed business approval for that work package, and the registry records no other valid signed business approval for it. At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff. At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff. Acceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both.\"", "base_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}], "counter_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}], "focus_atom": "attached_registry_document_identity", "focus_evidence": [{"path": [], "text": "At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff."}, {"path": [], "text": "At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff."}], "policy_evidence": [{"path": [], "text": "Acceptance requires all 12 tests to pass and a signed business approval to be attached."}, {"path": [], "text": "Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver."}, {"path": [], "text": "Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}], "rules": [{"justification": "The failed test establishes a test defect. Because attached document A-17 is identical to D-9, which is a valid signed business approval for this work package, valid approval is attached. The rubric therefore requires rejection, routing only to the delivery owner, and a medium rating.", "target": "reject_route_delivery_owner_medium", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}]}, {"justification": "The failed test establishes a test defect. A-17 is the only attached business-approval document, D-9 is the only valid signed approval for this work package, and the two documents are explicitly different; therefore no valid signed business approval is attached. Both deficiencies are present, requiring rejection, routing to both responsible roles, and a low rating.", "target": "none_of_above", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff.", "negative_left": "At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff.", "negative_right": "At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-6194, assigned one-to-one among all documents in that handoff.", "right": "At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-055-003", "id": "scale-diverse-055-003-base", "input": {"questions": {"decision": {"criteria": {"accept_high_quality": "Accept and rate high only when all 12 tests pass and valid signed business approval is attached; no corrective routing is required.", "none_of_above": "Choose when both a test defect and missing valid business approval are present; reject, route to both responsible roles, and rate completion quality low.", "reject_route_business_approver_medium": "Reject, route only to the business approver, and rate medium when all 12 tests pass but valid signed business approval is missing.", "reject_route_delivery_owner_medium": "Reject, route only to the delivery owner, and rate medium when at least one test fails but valid signed business approval is attached."}, "instructions": "Select the single option whose rubric matches the evidence, including the acceptance decision, responsible routing, and completion-quality rating.", "type": "choice"}}, "state": "At acceptance review, the project coordinator submitted the completed “Quarterly Export” work package. The signed test report covers all 12 required tests and records 11 passes; the CSV export test failed. The attachment manifest lists document A-17 as the sole business-approval document attached to the submitted completed work package. Registry document D-9 is a valid signed business approval for that work package, and the registry records no other valid signed business approval for it. At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff. At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff. Acceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}, "method": "c2d", "provenance": {"source_id": "diverse-055", "source_is_synthetic": true, "source_sha256": "16b728b9f77e4bb47a594dca1c07cd310327dbb290c6117b5f7d03d56d1e3e36", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "reject_route_delivery_owner_medium"}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing acceptance, routing, and quality policies from the original state, while the unchanged questions preserve the complete decision rubric. They maintain the same work package, review stage, criteria, roles, and handoff time. The two evidence spans are complete factual sentences. The counterfactual changes only D-9’s immutable identifier, coherently distinguishing it from attached A-17 under the stated one-to-one identifier assignment without creating duplicate or contradictory measurements. Neither context supplies an answer choice, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "full_context_fact_states": {"base": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "counterfactual": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "remove_left": {"attached_registry_document_identity": "unknown"}, "remove_right": {"attached_registry_document_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"attached_registry_document_identity": "unknown"}, "negative_pair": {"attached_registry_document_identity": "refuted"}, "negative_sentence": {"attached_registry_document_identity": "unknown"}, "positive_pair": {"attached_registry_document_identity": "supported"}, "right": {"attached_registry_document_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual status, identity, attachment, or explicitly scoped universal relationship. The focus is the factual identity of two documents, not a policy conclusion. The base and counter assignments are jointly realizable with only that identity changing: in the counter, D-9 can be valid but unattached while A-17 is attached but not valid. The policy evidence correctly cites substantive rules from the original state; governing material in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A test defect is present. A-17 is attached and identical to D-9, which is a valid signed business approval, so valid approval is attached. The criteria therefore require rejection, routing only to the delivery owner, and a medium rating.", "rule_index": 0, "sound": true}, {"reason": "A test defect is present. If any valid signed approval were attached, the two universal restrictions would make it both D-9 and A-17, contradicting the supported non-identity of those documents. Thus no valid signed approval is attached, so both deficiencies are present and the none_of_above outcome is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "test_defect", "statement": "At the acceptance review of the submitted completed “Quarterly Export” work package, at least one of the 12 required tests failed."}, {"id": "a17_attached", "statement": "At the acceptance review, document A-17 is attached to the submitted completed “Quarterly Export” work package."}, {"id": "a17_only_attached_approval_candidate", "statement": "At the acceptance review, every business-approval document attached to the submitted completed “Quarterly Export” work package is document A-17."}, {"id": "d9_valid_approval", "statement": "At the acceptance review, registry document D-9 is a valid signed business approval for the submitted completed “Quarterly Export” work package."}, {"id": "d9_only_valid_approval", "statement": "At the acceptance review, every valid signed business approval for the submitted completed “Quarterly Export” work package is registry document D-9."}, {"id": "attached_registry_document_identity", "statement": "Document A-17 attached to the submitted completed “Quarterly Export” work package is the same document as registry document D-9."}], "base_state_json": "\"At acceptance review, the project coordinator submitted the completed “Quarterly Export” work package. The signed test report covers all 12 required tests and records 11 passes; the CSV export test failed. The attachment manifest lists document A-17 as the sole business-approval document attached to the submitted completed work package. Registry document D-9 is a valid signed business approval for that work package, and the registry records no other valid signed business approval for it. At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff. At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff. Acceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both.\"", "base_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}], "counter_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}], "focus_atom": "attached_registry_document_identity", "focus_evidence": [{"path": [], "text": "At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff."}, {"path": [], "text": "At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff."}], "policy_evidence": [{"path": [], "text": "Acceptance requires all 12 tests to pass and a signed business approval to be attached."}, {"path": [], "text": "Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver."}, {"path": [], "text": "Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}], "rules": [{"justification": "The failed test establishes a test defect. Because attached document A-17 is identical to D-9, which is a valid signed business approval for this work package, valid approval is attached. The rubric therefore requires rejection, routing only to the delivery owner, and a medium rating.", "target": "reject_route_delivery_owner_medium", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}]}, {"justification": "The failed test establishes a test defect. A-17 is the only attached business-approval document, D-9 is the only valid signed approval for this work package, and the two documents are explicitly different; therefore no valid signed business approval is attached. Both deficiencies are present, requiring rejection, routing to both responsible roles, and a low rating.", "target": "none_of_above", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff.", "negative_left": "At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff.", "negative_right": "At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-6194, assigned one-to-one among all documents in that handoff.", "right": "At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff."}, "verifier_independent_model": false}, "family": "scale-diverse-055-003", "id": "scale-diverse-055-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_high_quality": "Accept and rate high only when all 12 tests pass and valid signed business approval is attached; no corrective routing is required.", "none_of_above": "Choose when both a test defect and missing valid business approval are present; reject, route to both responsible roles, and rate completion quality low.", "reject_route_business_approver_medium": "Reject, route only to the business approver, and rate medium when all 12 tests pass but valid signed business approval is missing.", "reject_route_delivery_owner_medium": "Reject, route only to the delivery owner, and rate medium when at least one test fails but valid signed business approval is attached."}, "instructions": "Select the single option whose rubric matches the evidence, including the acceptance decision, responsible routing, and completion-quality rating.", "type": "choice"}}, "state": "At acceptance review, the project coordinator submitted the completed “Quarterly Export” work package. The signed test report covers all 12 required tests and records 11 passes; the CSV export test failed. The attachment manifest lists document A-17 as the sole business-approval document attached to the submitted completed work package. Registry document D-9 is a valid signed business approval for that work package, and the registry records no other valid signed business approval for it. At the 2026-06-30 14:32 UTC acceptance-review handoff, document A-17 attached to the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-4827, assigned one-to-one among all documents in that handoff. At the 2026-06-30 14:32 UTC acceptance-review handoff, registry document D-9 for the submitted completed “Quarterly Export” work package bore the single immutable document identifier HN-6194, assigned one-to-one among all documents in that handoff. Acceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}, "method": "c2d", "provenance": {"source_id": "diverse-055", "source_is_synthetic": true, "source_sha256": "16b728b9f77e4bb47a594dca1c07cd310327dbb290c6117b5f7d03d56d1e3e36", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing acceptance, routing, and quality policies, supplemented by the unchanged questions object. They preserve the work-package and acceptance-review bindings. The two evidence spans are complete factual sentences. The counterfactual changes only D-9’s fingerprint; this is coherent with A-17’s sole fingerprint and the fingerprint-uniqueness assertion, with no contradictory duplicate measurement. Neither context contains a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "full_context_fact_states": {"base": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "counterfactual": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "remove_left": {"attached_registry_document_identity": "unknown"}, "remove_right": {"attached_registry_document_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"attached_registry_document_identity": "unknown"}, "negative_pair": {"attached_registry_document_identity": "refuted"}, "negative_sentence": {"attached_registry_document_identity": "unknown"}, "positive_pair": {"attached_registry_document_identity": "supported"}, "right": {"attached_registry_document_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual status, identity, attachment, or explicitly scoped universal relationship. The focus is the factual identity of two documents, not a policy conclusion. The base and counter assignments are jointly realizable with only that identity changing: in the counter, D-9 can be valid but unattached while A-17 is attached but not valid. The policy evidence correctly cites substantive rules from the original state; governing material in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A test defect is present. A-17 is attached and identical to D-9, which is a valid signed business approval, so valid approval is attached. The criteria therefore require rejection, routing only to the delivery owner, and a medium rating.", "rule_index": 0, "sound": true}, {"reason": "A test defect is present. If any valid signed approval were attached, the two universal restrictions would make it both D-9 and A-17, contradicting the supported non-identity of those documents. Thus no valid signed approval is attached, so both deficiencies are present and the none_of_above outcome is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "test_defect", "statement": "At the acceptance review of the submitted completed “Quarterly Export” work package, at least one of the 12 required tests failed."}, {"id": "a17_attached", "statement": "At the acceptance review, document A-17 is attached to the submitted completed “Quarterly Export” work package."}, {"id": "a17_only_attached_approval_candidate", "statement": "At the acceptance review, every business-approval document attached to the submitted completed “Quarterly Export” work package is document A-17."}, {"id": "d9_valid_approval", "statement": "At the acceptance review, registry document D-9 is a valid signed business approval for the submitted completed “Quarterly Export” work package."}, {"id": "d9_only_valid_approval", "statement": "At the acceptance review, every valid signed business approval for the submitted completed “Quarterly Export” work package is registry document D-9."}, {"id": "attached_registry_document_identity", "statement": "Document A-17 attached to the submitted completed “Quarterly Export” work package is the same document as registry document D-9."}], "base_state_json": "\"Field note—The signed test report for the submitted completed “Quarterly Export” work package records eleven passes and a CSV export failure among the 12 required tests. The attachment inventory lists A-17 and states that every attached business-approval document is A-17. The approval registry identifies D-9 as a valid signed business approval for the work package and states that every valid signed business approval for it is D-9. At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore. At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-8041. Acceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both.\"", "base_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}], "counter_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}], "focus_atom": "attached_registry_document_identity", "focus_evidence": [{"path": [], "text": "At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore."}, {"path": [], "text": "At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-8041."}], "policy_evidence": [{"path": [], "text": "Acceptance requires all 12 tests to pass and a signed business approval to be attached."}, {"path": [], "text": "Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver."}, {"path": [], "text": "Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}], "rules": [{"justification": "The failed test establishes a test defect. Because attached document A-17 is identical to D-9, which is a valid signed business approval for this work package, valid approval is attached. The rubric therefore requires rejection, routing only to the delivery owner, and a medium rating.", "target": "reject_route_delivery_owner_medium", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}]}, {"justification": "The failed test establishes a test defect. A-17 is the only attached business-approval document, D-9 is the only valid signed approval for this work package, and the two documents are explicitly different; therefore no valid signed business approval is attached. Both deficiencies are present, requiring rejection, routing to both responsible roles, and a low rating.", "target": "none_of_above", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore.", "negative_left": "At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore.", "negative_right": "At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-6173.", "right": "At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-8041."}, "verifier_independent_model": false}, "family": "scale-diverse-055-004", "id": "scale-diverse-055-004-base", "input": {"questions": {"decision": {"criteria": {"accept_high_quality": "Accept and rate high only when all 12 tests pass and valid signed business approval is attached; no corrective routing is required.", "none_of_above": "Choose when both a test defect and missing valid business approval are present; reject, route to both responsible roles, and rate completion quality low.", "reject_route_business_approver_medium": "Reject, route only to the business approver, and rate medium when all 12 tests pass but valid signed business approval is missing.", "reject_route_delivery_owner_medium": "Reject, route only to the delivery owner, and rate medium when at least one test fails but valid signed business approval is attached."}, "instructions": "Select the single option whose rubric matches the evidence, including the acceptance decision, responsible routing, and completion-quality rating.", "type": "choice"}}, "state": "Field note—The signed test report for the submitted completed “Quarterly Export” work package records eleven passes and a CSV export failure among the 12 required tests. The attachment inventory lists A-17 and states that every attached business-approval document is A-17. The approval registry identifies D-9 as a valid signed business approval for the work package and states that every valid signed business approval for it is D-9. At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore. At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-8041. Acceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}, "method": "c2d", "provenance": {"source_id": "diverse-055", "source_is_synthetic": true, "source_sha256": "16b728b9f77e4bb47a594dca1c07cd310327dbb290c6117b5f7d03d56d1e3e36", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "reject_route_delivery_owner_medium"}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing acceptance, routing, and quality policies, supplemented by the unchanged questions object. They preserve the work-package and acceptance-review bindings. The two evidence spans are complete factual sentences. The counterfactual changes only D-9’s fingerprint; this is coherent with A-17’s sole fingerprint and the fingerprint-uniqueness assertion, with no contradictory duplicate measurement. Neither context contains a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "full_context_fact_states": {"base": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "supported", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "counterfactual": {"a17_attached": "supported", "a17_only_attached_approval_candidate": "supported", "attached_registry_document_identity": "refuted", "d9_only_valid_approval": "supported", "d9_valid_approval": "supported", "test_defect": "supported"}, "remove_left": {"attached_registry_document_identity": "unknown"}, "remove_right": {"attached_registry_document_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"attached_registry_document_identity": "unknown"}, "negative_pair": {"attached_registry_document_identity": "refuted"}, "negative_sentence": {"attached_registry_document_identity": "unknown"}, "positive_pair": {"attached_registry_document_identity": "supported"}, "right": {"attached_registry_document_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual status, identity, attachment, or explicitly scoped universal relationship. The focus is the factual identity of two documents, not a policy conclusion. The base and counter assignments are jointly realizable with only that identity changing: in the counter, D-9 can be valid but unattached while A-17 is attached but not valid. The policy evidence correctly cites substantive rules from the original state; governing material in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A test defect is present. A-17 is attached and identical to D-9, which is a valid signed business approval, so valid approval is attached. The criteria therefore require rejection, routing only to the delivery owner, and a medium rating.", "rule_index": 0, "sound": true}, {"reason": "A test defect is present. If any valid signed approval were attached, the two universal restrictions would make it both D-9 and A-17, contradicting the supported non-identity of those documents. Thus no valid signed approval is attached, so both deficiencies are present and the none_of_above outcome is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "test_defect", "statement": "At the acceptance review of the submitted completed “Quarterly Export” work package, at least one of the 12 required tests failed."}, {"id": "a17_attached", "statement": "At the acceptance review, document A-17 is attached to the submitted completed “Quarterly Export” work package."}, {"id": "a17_only_attached_approval_candidate", "statement": "At the acceptance review, every business-approval document attached to the submitted completed “Quarterly Export” work package is document A-17."}, {"id": "d9_valid_approval", "statement": "At the acceptance review, registry document D-9 is a valid signed business approval for the submitted completed “Quarterly Export” work package."}, {"id": "d9_only_valid_approval", "statement": "At the acceptance review, every valid signed business approval for the submitted completed “Quarterly Export” work package is registry document D-9."}, {"id": "attached_registry_document_identity", "statement": "Document A-17 attached to the submitted completed “Quarterly Export” work package is the same document as registry document D-9."}], "base_state_json": "\"Field note—The signed test report for the submitted completed “Quarterly Export” work package records eleven passes and a CSV export failure among the 12 required tests. The attachment inventory lists A-17 and states that every attached business-approval document is A-17. The approval registry identifies D-9 as a valid signed business approval for the work package and states that every valid signed business approval for it is D-9. At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore. At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-8041. Acceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both.\"", "base_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}], "counter_states": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}], "focus_atom": "attached_registry_document_identity", "focus_evidence": [{"path": [], "text": "At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore."}, {"path": [], "text": "At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-8041."}], "policy_evidence": [{"path": [], "text": "Acceptance requires all 12 tests to pass and a signed business approval to be attached."}, {"path": [], "text": "Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver."}, {"path": [], "text": "Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}], "rules": [{"justification": "The failed test establishes a test defect. Because attached document A-17 is identical to D-9, which is a valid signed business approval for this work package, valid approval is attached. The rubric therefore requires rejection, routing only to the delivery owner, and a medium rating.", "target": "reject_route_delivery_owner_medium", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "supported"}]}, {"justification": "The failed test establishes a test defect. A-17 is the only attached business-approval document, D-9 is the only valid signed approval for this work package, and the two documents are explicitly different; therefore no valid signed business approval is attached. Both deficiencies are present, requiring rejection, routing to both responsible roles, and a low rating.", "target": "none_of_above", "when": [{"atom_id": "test_defect", "state": "supported"}, {"atom_id": "a17_attached", "state": "supported"}, {"atom_id": "a17_only_attached_approval_candidate", "state": "supported"}, {"atom_id": "d9_valid_approval", "state": "supported"}, {"atom_id": "d9_only_valid_approval", "state": "supported"}, {"atom_id": "attached_registry_document_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore.", "negative_left": "At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore.", "negative_right": "At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-6173.", "right": "At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-8041."}, "verifier_independent_model": false}, "family": "scale-diverse-055-004", "id": "scale-diverse-055-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_high_quality": "Accept and rate high only when all 12 tests pass and valid signed business approval is attached; no corrective routing is required.", "none_of_above": "Choose when both a test defect and missing valid business approval are present; reject, route to both responsible roles, and rate completion quality low.", "reject_route_business_approver_medium": "Reject, route only to the business approver, and rate medium when all 12 tests pass but valid signed business approval is missing.", "reject_route_delivery_owner_medium": "Reject, route only to the delivery owner, and rate medium when at least one test fails but valid signed business approval is attached."}, "instructions": "Select the single option whose rubric matches the evidence, including the acceptance decision, responsible routing, and completion-quality rating.", "type": "choice"}}, "state": "Field note—The signed test report for the submitted completed “Quarterly Export” work package records eleven passes and a CSV export failure among the 12 required tests. The attachment inventory lists A-17 and states that every attached business-approval document is A-17. The approval registry identifies D-9 as a valid signed business approval for the work package and states that every valid signed business approval for it is D-9. At the 2026-09-17 acceptance review, document A-17 attached to the submitted completed “Quarterly Export” work package bore fingerprint QX-8041 as its sole fingerprint, and each document examined at that review bore exactly one fingerprint that no other examined document bore. At the 2026-09-17 acceptance review, registry document D-9 for the submitted completed “Quarterly Export” work package bore fingerprint QX-6173. Acceptance requires all 12 tests to pass and a signed business approval to be attached. Failed tests must be routed to the delivery owner, while missing sign-off must be routed to the business approver. Completion quality is high only when both criteria pass, medium for exactly one deficiency, and low for both."}, "method": "c2d", "provenance": {"source_id": "diverse-055", "source_is_synthetic": true, "source_sha256": "16b728b9f77e4bb47a594dca1c07cd310327dbb290c6117b5f7d03d56d1e3e36", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the complete acceptance and rating policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the WP-17 decision, cutoff, and yes/no request. The two focus-evidence spans are complete factual sentences. The counterfactual coherently replaces the ledger’s recorded approval with a complete-ledger assertion that Elena Voss did not approve, without conflicting with any unchanged observation. Neither context supplies a gold answer, answer code, rule table, proposition ID, label rationale, or directive selecting an answer; “Answer yes or no” only preserves the original response format.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Records clerk\",\"text\":\"At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver.\"},{\"speaker\":\"Test lead\",\"text\":\"By 15:00 UTC, the final WP-17 report showed that all 12 acceptance tests had passed. The consolidated technical evidence package was complete, and the defect closure log showed zero remaining defects.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"At 15:20 UTC, I completed the Quality review and signed off WP-17 in the review register.\"},{\"speaker\":\"Ledger custodian\",\"text\":\"The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 records Elena Voss signing the entry “APPROVE” at 15:42 UTC that day.\"},{\"speaker\":\"Project coordinator\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"},{\"speaker\":\"Decision chair\",\"text\":\"Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["1", "text"], "text": "At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver."}, {"path": ["4", "text"], "text": "The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 records Elena Voss signing the entry “APPROVE” at 15:42 UTC that day."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver.", "negative_left": "At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver.", "negative_right": "The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 contains no approval signed by Elena Voss.", "right": "The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 records Elena Voss signing the entry “APPROVE” at 15:42 UTC that day."}, "verifier_independent_model": false}, "family": "scale-diverse-057-001", "id": "scale-diverse-057-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Project coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Records clerk", "text": "At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver."}, {"speaker": "Test lead", "text": "By 15:00 UTC, the final WP-17 report showed that all 12 acceptance tests had passed. The consolidated technical evidence package was complete, and the defect closure log showed zero remaining defects."}, {"speaker": "Quality reviewer", "text": "At 15:20 UTC, I completed the Quality review and signed off WP-17 in the review register."}, {"speaker": "Ledger custodian", "text": "The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 records Elena Voss signing the entry “APPROVE” at 15:42 UTC that day."}, {"speaker": "Project coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}, {"speaker": "Decision chair", "text": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the complete acceptance and rating policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the WP-17 decision, cutoff, and yes/no request. The two focus-evidence spans are complete factual sentences. The counterfactual coherently replaces the ledger’s recorded approval with a complete-ledger assertion that Elena Voss did not approve, without conflicting with any unchanged observation. Neither context supplies a gold answer, answer code, rule table, proposition ID, label rationale, or directive selecting an answer; “Answer yes or no” only preserves the original response format.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Records clerk\",\"text\":\"At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver.\"},{\"speaker\":\"Test lead\",\"text\":\"By 15:00 UTC, the final WP-17 report showed that all 12 acceptance tests had passed. The consolidated technical evidence package was complete, and the defect closure log showed zero remaining defects.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"At 15:20 UTC, I completed the Quality review and signed off WP-17 in the review register.\"},{\"speaker\":\"Ledger custodian\",\"text\":\"The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 records Elena Voss signing the entry “APPROVE” at 15:42 UTC that day.\"},{\"speaker\":\"Project coordinator\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"},{\"speaker\":\"Decision chair\",\"text\":\"Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["1", "text"], "text": "At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver."}, {"path": ["4", "text"], "text": "The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 records Elena Voss signing the entry “APPROVE” at 15:42 UTC that day."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver.", "negative_left": "At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver.", "negative_right": "The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 contains no approval signed by Elena Voss.", "right": "The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 records Elena Voss signing the entry “APPROVE” at 15:42 UTC that day."}, "verifier_independent_model": false}, "family": "scale-diverse-057-001", "id": "scale-diverse-057-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Project coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Records clerk", "text": "At the WP-17 readiness cutoff of 16:00 UTC on 14 September 2026, the assignment register listed Elena Voss as WP-17's sole Business approver."}, {"speaker": "Test lead", "text": "By 15:00 UTC, the final WP-17 report showed that all 12 acceptance tests had passed. The consolidated technical evidence package was complete, and the defect closure log showed zero remaining defects."}, {"speaker": "Quality reviewer", "text": "At 15:20 UTC, I completed the Quality review and signed off WP-17 in the review register."}, {"speaker": "Ledger custodian", "text": "The complete WP-17 approval ledger through 16:00 UTC on 14 September 2026 contains no approval signed by Elena Voss."}, {"speaker": "Project coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}, {"speaker": "Decision chair", "text": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full acceptance and rating policy without adding exceptions, priorities, or missing-evidence defaults. They preserve the WP-17 readiness decision, criteria, scope, and relevant approver/time path; the changed observation is confined to Mara Voss’s recorded decision. The two focus-evidence spans are complete factual sentences. The counterfactual’s single changed sentence coherently replaces the approval record with a decline and absence of approval, without conflicting with the unchanged assignment or other evidence. Neither context contains an answer code, proposition ID, rule table, explicit gold answer, or classifier-output instruction; the outcome terminology appears only as part of the governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Technical verifier\",\"text\":\"All 12 acceptance tests for WP-17 passed, and the technical evidence for WP-17 is complete.\"},{\"speaker\":\"Quality records officer\",\"text\":\"The Quality reviewer signed off on WP-17, and the final defect ledger records zero remaining defects.\"},{\"speaker\":\"Assignment registrar\",\"text\":\"The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver.\"},{\"speaker\":\"Decision-audit custodian\",\"text\":\"The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Approve” at 13:42 UTC and applying her authenticated signature to that selection.\"},{\"speaker\":\"Acceptance policy\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Project coordinator\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["2", "text"], "text": "The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver."}, {"path": ["3", "text"], "text": "The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Approve” at 13:42 UTC and applying her authenticated signature to that selection."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver.", "negative_left": "The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver.", "negative_right": "The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Decline” at 13:42 UTC and contains no “Approve” selection or approval signature by her.", "right": "The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Approve” at 13:42 UTC and applying her authenticated signature to that selection."}, "verifier_independent_model": false}, "family": "scale-diverse-057-002", "id": "scale-diverse-057-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Technical verifier", "text": "All 12 acceptance tests for WP-17 passed, and the technical evidence for WP-17 is complete."}, {"speaker": "Quality records officer", "text": "The Quality reviewer signed off on WP-17, and the final defect ledger records zero remaining defects."}, {"speaker": "Assignment registrar", "text": "The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver."}, {"speaker": "Decision-audit custodian", "text": "The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Approve” at 13:42 UTC and applying her authenticated signature to that selection."}, {"speaker": "Acceptance policy", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Project coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full acceptance and rating policy without adding exceptions, priorities, or missing-evidence defaults. They preserve the WP-17 readiness decision, criteria, scope, and relevant approver/time path; the changed observation is confined to Mara Voss’s recorded decision. The two focus-evidence spans are complete factual sentences. The counterfactual’s single changed sentence coherently replaces the approval record with a decline and absence of approval, without conflicting with the unchanged assignment or other evidence. Neither context contains an answer code, proposition ID, rule table, explicit gold answer, or classifier-output instruction; the outcome terminology appears only as part of the governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Technical verifier\",\"text\":\"All 12 acceptance tests for WP-17 passed, and the technical evidence for WP-17 is complete.\"},{\"speaker\":\"Quality records officer\",\"text\":\"The Quality reviewer signed off on WP-17, and the final defect ledger records zero remaining defects.\"},{\"speaker\":\"Assignment registrar\",\"text\":\"The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver.\"},{\"speaker\":\"Decision-audit custodian\",\"text\":\"The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Approve” at 13:42 UTC and applying her authenticated signature to that selection.\"},{\"speaker\":\"Acceptance policy\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Project coordinator\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["2", "text"], "text": "The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver."}, {"path": ["3", "text"], "text": "The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Approve” at 13:42 UTC and applying her authenticated signature to that selection."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver.", "negative_left": "The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver.", "negative_right": "The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Decline” at 13:42 UTC and contains no “Approve” selection or approval signature by her.", "right": "The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Approve” at 13:42 UTC and applying her authenticated signature to that selection."}, "verifier_independent_model": false}, "family": "scale-diverse-057-002", "id": "scale-diverse-057-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Technical verifier", "text": "All 12 acceptance tests for WP-17 passed, and the technical evidence for WP-17 is complete."}, {"speaker": "Quality records officer", "text": "The Quality reviewer signed off on WP-17, and the final defect ledger records zero remaining defects."}, {"speaker": "Assignment registrar", "text": "The WP-17 assignment register frozen at 09:00 UTC on 14 May 2026 names Mara Voss as WP-17's sole assigned Business approver."}, {"speaker": "Decision-audit custodian", "text": "The complete WP-17 decision audit log through 17:00 UTC on 14 May 2026 records Mara Voss selecting “Decline” at 13:42 UTC and contains no “Approve” selection or approval signature by her."}, {"speaker": "Acceptance policy", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Project coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original acceptance and quality-rating policies without adding exceptions, priorities, or missing-evidence defaults, and they retain the same Atlas work package, decision, criteria, and timing. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the signer from AT-417 to AT-862: AT-862 can sign the totals record while not being the Business approver, and AT-417 can be the Business approver without having provided sign-off. Neither context contains a gold answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"At 14:50 UTC on 18 September 2026, the Quality reviewer finalized the log for all 12 specified Atlas tests; each test had passing status, and no defect remained at the 16:00 acceptance decision time. At 15:10 UTC, the Delivery owner wrote, “If Finance approves the totals, submit the Atlas reporting work package for acceptance.” Finance approved the Atlas totals at 15:30 UTC, meeting that condition before the decision. At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-417 as its signer. At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver. By the decision time, no person other than the signer identified on that dated record had provided Business approver sign-off for the package. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-417 as its signer."}, {"path": [], "text": "At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-417 as its signer.", "negative_left": "At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-862 as its signer.", "negative_right": "At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver.", "right": "At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver."}, "verifier_independent_model": false}, "family": "scale-diverse-058-001", "id": "scale-diverse-058-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "At 14:50 UTC on 18 September 2026, the Quality reviewer finalized the log for all 12 specified Atlas tests; each test had passing status, and no defect remained at the 16:00 acceptance decision time. At 15:10 UTC, the Delivery owner wrote, “If Finance approves the totals, submit the Atlas reporting work package for acceptance.” Finance approved the Atlas totals at 15:30 UTC, meeting that condition before the decision. At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-417 as its signer. At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver. By the decision time, no person other than the signer identified on that dated record had provided Business approver sign-off for the package. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original acceptance and quality-rating policies without adding exceptions, priorities, or missing-evidence defaults, and they retain the same Atlas work package, decision, criteria, and timing. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the signer from AT-417 to AT-862: AT-862 can sign the totals record while not being the Business approver, and AT-417 can be the Business approver without having provided sign-off. Neither context contains a gold answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"At 14:50 UTC on 18 September 2026, the Quality reviewer finalized the log for all 12 specified Atlas tests; each test had passing status, and no defect remained at the 16:00 acceptance decision time. At 15:10 UTC, the Delivery owner wrote, “If Finance approves the totals, submit the Atlas reporting work package for acceptance.” Finance approved the Atlas totals at 15:30 UTC, meeting that condition before the decision. At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-417 as its signer. At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver. By the decision time, no person other than the signer identified on that dated record had provided Business approver sign-off for the package. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-417 as its signer."}, {"path": [], "text": "At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-417 as its signer.", "negative_left": "At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-862 as its signer.", "negative_right": "At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver.", "right": "At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver."}, "verifier_independent_model": false}, "family": "scale-diverse-058-001", "id": "scale-diverse-058-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "At 14:50 UTC on 18 September 2026, the Quality reviewer finalized the log for all 12 specified Atlas tests; each test had passing status, and no defect remained at the 16:00 acceptance decision time. At 15:10 UTC, the Delivery owner wrote, “If Finance approves the totals, submit the Atlas reporting work package for acceptance.” Finance approved the Atlas totals at 15:30 UTC, meeting that condition before the decision. At 15:40 UTC on 18 September 2026, the dated approval record for the Atlas totals identified employee AT-862 as its signer. At 16:00 UTC on 18 September 2026, employee AT-417 was the Business approver for the Atlas reporting work package, whereas employee AT-862 was a different person and was not its Business approver. By the decision time, no person other than the signer identified on that dated record had provided Business approver sign-off for the package. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same acceptance and quality-rating policies and preserve the Atlas work package, decision scope, criteria, and 2026-08-21 14:00 UTC decision time. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the Business approver from Rowan Ilyas to Priya Sen while retaining Rowan as the approval-record signer; the audit statement excludes sign-off by anyone other than Rowan but does not itself assert that Rowan supplied valid Business approver sign-off. Neither context embeds a gold answer, code, rule table, proposition identifier, label rationale, or extra output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"Evidence reconciliation at the acceptance decision time of 2026-08-21 14:00 UTC found that the final Quality log gave passing status to each of the 12 specified Atlas tests and recorded no remaining defect. The Delivery owner had expressed intent to submit the Atlas reporting work package for acceptance if Finance approved the totals, and Finance had approved them by that time. A dated, signed approval record for the totals was present in the acceptance file. The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas. From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Rowan Ilyas was the Business approver for the Atlas reporting work package. The sign-off audit showed that no person other than the signer recorded on that approval had provided Business approver sign-off by the decision time.\\n\\nAcceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met.\\n\\nRate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas."}, {"path": [], "text": "From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Rowan Ilyas was the Business approver for the Atlas reporting work package."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas.", "negative_left": "The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas.", "negative_right": "From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Priya Sen was the sole Business approver for the Atlas reporting work package, and Rowan Ilyas was not a Business approver for it.", "right": "From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Rowan Ilyas was the Business approver for the Atlas reporting work package."}, "verifier_independent_model": false}, "family": "scale-diverse-058-002", "id": "scale-diverse-058-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "Evidence reconciliation at the acceptance decision time of 2026-08-21 14:00 UTC found that the final Quality log gave passing status to each of the 12 specified Atlas tests and recorded no remaining defect. The Delivery owner had expressed intent to submit the Atlas reporting work package for acceptance if Finance approved the totals, and Finance had approved them by that time. A dated, signed approval record for the totals was present in the acceptance file. The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas. From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Rowan Ilyas was the Business approver for the Atlas reporting work package. The sign-off audit showed that no person other than the signer recorded on that approval had provided Business approver sign-off by the decision time.\n\nAcceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met.\n\nRate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same acceptance and quality-rating policies and preserve the Atlas work package, decision scope, criteria, and 2026-08-21 14:00 UTC decision time. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the Business approver from Rowan Ilyas to Priya Sen while retaining Rowan as the approval-record signer; the audit statement excludes sign-off by anyone other than Rowan but does not itself assert that Rowan supplied valid Business approver sign-off. Neither context embeds a gold answer, code, rule table, proposition identifier, label rationale, or extra output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"Evidence reconciliation at the acceptance decision time of 2026-08-21 14:00 UTC found that the final Quality log gave passing status to each of the 12 specified Atlas tests and recorded no remaining defect. The Delivery owner had expressed intent to submit the Atlas reporting work package for acceptance if Finance approved the totals, and Finance had approved them by that time. A dated, signed approval record for the totals was present in the acceptance file. The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas. From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Rowan Ilyas was the Business approver for the Atlas reporting work package. The sign-off audit showed that no person other than the signer recorded on that approval had provided Business approver sign-off by the decision time.\\n\\nAcceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met.\\n\\nRate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas."}, {"path": [], "text": "From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Rowan Ilyas was the Business approver for the Atlas reporting work package."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas.", "negative_left": "The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas.", "negative_right": "From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Priya Sen was the sole Business approver for the Atlas reporting work package, and Rowan Ilyas was not a Business approver for it.", "right": "From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Rowan Ilyas was the Business approver for the Atlas reporting work package."}, "verifier_independent_model": false}, "family": "scale-diverse-058-002", "id": "scale-diverse-058-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "Evidence reconciliation at the acceptance decision time of 2026-08-21 14:00 UTC found that the final Quality log gave passing status to each of the 12 specified Atlas tests and recorded no remaining defect. The Delivery owner had expressed intent to submit the Atlas reporting work package for acceptance if Finance approved the totals, and Finance had approved them by that time. A dated, signed approval record for the totals was present in the acceptance file. The signer field on the Atlas totals approval dated 2026-08-19 identifies Rowan Ilyas. From 2026-08-19 through the acceptance decision at 14:00 UTC on 2026-08-21, Priya Sen was the sole Business approver for the Atlas reporting work package, and Rowan Ilyas was not a Business approver for it. The sign-off audit showed that no person other than the signer recorded on that approval had provided Business approver sign-off by the decision time.\n\nAcceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met.\n\nRate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance, conditional-intent, and quality-rating policies, while the unchanged questions preserve the request and scoring criteria. The Atlas work package, decision time, relevant records, and signer-to-role comparison remain bound to the same entities and paths. The two evidence spans are complete factual sentences. The counterfactual coherently changes the sole Business approver from Lina Voss to the distinct Omar Pell without creating duplicate role holders or conflicting measurements. Neither context contains an explicit answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"Operational handoff at 16:00 UTC on 14 September 2026: The final Atlas test log listed all 12 specified tests as passing and confirmed that no defect remained. Before the decision time, the Delivery owner wrote, “If Finance approves the totals, submit the Atlas reporting work package for acceptance.” Finance approved the Atlas totals before the handoff and linked the dated, signed approval record. The sign-off audit found no Business approver sign-off from any person other than the record’s sole signer. At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417. At 16:00 UTC on 14 September 2026, the Atlas role register identified Lina Voss, employee AV-417, as the sole person holding the Business approver role. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417."}, {"path": [], "text": "At 16:00 UTC on 14 September 2026, the Atlas role register identified Lina Voss, employee AV-417, as the sole person holding the Business approver role."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417.", "negative_left": "At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417.", "negative_right": "At 16:00 UTC on 14 September 2026, the Atlas role register identified Omar Pell, a person distinct from Lina Voss, as the sole person holding the Business approver role.", "right": "At 16:00 UTC on 14 September 2026, the Atlas role register identified Lina Voss, employee AV-417, as the sole person holding the Business approver role."}, "verifier_independent_model": false}, "family": "scale-diverse-058-003", "id": "scale-diverse-058-003-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "Operational handoff at 16:00 UTC on 14 September 2026: The final Atlas test log listed all 12 specified tests as passing and confirmed that no defect remained. Before the decision time, the Delivery owner wrote, “If Finance approves the totals, submit the Atlas reporting work package for acceptance.” Finance approved the Atlas totals before the handoff and linked the dated, signed approval record. The sign-off audit found no Business approver sign-off from any person other than the record’s sole signer. At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417. At 16:00 UTC on 14 September 2026, the Atlas role register identified Lina Voss, employee AV-417, as the sole person holding the Business approver role. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance, conditional-intent, and quality-rating policies, while the unchanged questions preserve the request and scoring criteria. The Atlas work package, decision time, relevant records, and signer-to-role comparison remain bound to the same entities and paths. The two evidence spans are complete factual sentences. The counterfactual coherently changes the sole Business approver from Lina Voss to the distinct Omar Pell without creating duplicate role holders or conflicting measurements. Neither context contains an explicit answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"Operational handoff at 16:00 UTC on 14 September 2026: The final Atlas test log listed all 12 specified tests as passing and confirmed that no defect remained. Before the decision time, the Delivery owner wrote, “If Finance approves the totals, submit the Atlas reporting work package for acceptance.” Finance approved the Atlas totals before the handoff and linked the dated, signed approval record. The sign-off audit found no Business approver sign-off from any person other than the record’s sole signer. At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417. At 16:00 UTC on 14 September 2026, the Atlas role register identified Lina Voss, employee AV-417, as the sole person holding the Business approver role. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417."}, {"path": [], "text": "At 16:00 UTC on 14 September 2026, the Atlas role register identified Lina Voss, employee AV-417, as the sole person holding the Business approver role."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417.", "negative_left": "At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417.", "negative_right": "At 16:00 UTC on 14 September 2026, the Atlas role register identified Omar Pell, a person distinct from Lina Voss, as the sole person holding the Business approver role.", "right": "At 16:00 UTC on 14 September 2026, the Atlas role register identified Lina Voss, employee AV-417, as the sole person holding the Business approver role."}, "verifier_independent_model": false}, "family": "scale-diverse-058-003", "id": "scale-diverse-058-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "Operational handoff at 16:00 UTC on 14 September 2026: The final Atlas test log listed all 12 specified tests as passing and confirmed that no defect remained. Before the decision time, the Delivery owner wrote, “If Finance approves the totals, submit the Atlas reporting work package for acceptance.” Finance approved the Atlas totals before the handoff and linked the dated, signed approval record. The sign-off audit found no Business approver sign-off from any person other than the record’s sole signer. At 16:00 UTC on 14 September 2026, the dated approval record for the Atlas totals identified its sole signer as Lina Voss, employee AV-417. At 16:00 UTC on 14 September 2026, the Atlas role register identified Omar Pell, a person distinct from Lina Voss, as the sole person holding the Business approver role. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the final-build testing, matching-approval, role-routing, and newer-build supersession policies, while preserving the Orion package, final-build decision scope, and request. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the test-ledger observation from one failed test to all tests passing and remains consistent with the approval and clean submission register. Neither context states a completion level, acceptance answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A2's quantification over the explicit set of 20 tests does not make it a bundle. A1 is a factual test-result proposition rather than a policy conclusion. The base and counter assignments are realizable with only A1 changing: the base can have one or more recorded failures, while the counter can have all recorded results be passes, with current-build approval and no evidence or documentation issues in both. Policy evidence preserves the state-originated final-build acceptance requirement, role routing, and newer-build supersession rule. Rules appearing only in the retained questions object need not be duplicated. The rule table may validly abstain on moderate-completion scenarios.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A1 entails that at least one mandatory test failed on the final build. The explicit criterion assigns level 0 whenever any mandatory final-build test fails, regardless of the other atoms.", "rule_index": 0, "sound": true}, {"reason": "Refuted A1 excludes every failed mandatory test, while supported A2 establishes that every mandatory test has a pass-or-fail result; together these entail that all 20 passed. A3 establishes approval naming the final build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions are sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At least one of the 20 mandatory tests failed on the build submitted as final for the Orion payroll-export package."}, {"id": "A2", "statement": "Each of the 20 mandatory tests has a recorded pass-or-fail result on the build submitted as final for the Orion payroll-export package."}, {"id": "A3", "statement": "The business approval names the build submitted as final for the Orion payroll-export package."}, {"id": "A4", "statement": "An unresolved evidence issue remains for the build submitted as final for the Orion payroll-export package."}, {"id": "A5", "statement": "An unresolved documentation issue remains for the build submitted as final for the Orion payroll-export package."}], "base_state_json": "{\"context\":\"Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.\",\"case_note\":[\"At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731.\",\"At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded FAIL for test 14 and PASS for each of tests 1–13 and 15–20.\",\"The coordinator treated the signed approval, submission register, and result ledger as authenticated closeout records for the package.\"],\"request\":\"Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["case_note", "0"], "text": "At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731."}, {"path": ["case_note", "1"], "text": "At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded FAIL for test 14 and PASS for each of tests 1–13 and 15–20."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build."}, {"path": ["context"], "text": "The delivery owner fixes product defects; the business approver renews outdated approvals."}, {"path": ["context"], "text": "Evidence for a newer build supersedes evidence for earlier builds."}, {"path": ["request"], "text": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}], "rules": [{"justification": "A supported A1 establishes a failed mandatory test on the final build. The low-completion criterion therefore requires non-acceptance, and the failed-test blocker is routed to the delivery owner because that role fixes product defects.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}]}, {"justification": "A2 establishes a pass-or-fail result for every mandatory test, while refuted A1 excludes any failed result, so all 20 mandatory tests pass on the final build. A3 supplies approval naming that build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions satisfy high completion.", "target": "2", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731.", "negative_left": "At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731.", "negative_right": "At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded PASS for each of tests 1–20.", "right": "At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded FAIL for test 14 and PASS for each of tests 1–13 and 15–20."}, "verifier_independent_model": false}, "family": "scale-diverse-059-001", "id": "scale-diverse-059-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Low completion: Do not accept. At least one mandatory final-build test fails or required final-build approval is missing; route each blocker to its responsible role.", "1 — Moderate completion: Conditionally accept only when all mandatory final-build tests pass and approval exists, but minor non-blocking documentation corrections remain.", "2 — High completion: Accept fully. Every mandatory test passes on the final build, the business approval names that build, and no unresolved evidence or documentation issues remain."], "instructions": "Apply the latest-build rule and the acceptance criteria. Choose exactly one ordered level. Identify acceptance status and route any failed test or missing current-build approval.", "type": "score"}}, "state": {"case_note": ["At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731.", "At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded FAIL for test 14 and PASS for each of tests 1–13 and 15–20.", "The coordinator treated the signed approval, submission register, and result ledger as authenticated closeout records for the package."], "context": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.", "request": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}}, "method": "c2d", "provenance": {"source_id": "diverse-059", "source_is_synthetic": true, "source_sha256": "f8f6806cfdd18388512db0984c3eb1386d711a91bbc8e11806a11e30573c54e8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the final-build testing, matching-approval, role-routing, and newer-build supersession policies, while preserving the Orion package, final-build decision scope, and request. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the test-ledger observation from one failed test to all tests passing and remains consistent with the approval and clean submission register. Neither context states a completion level, acceptance answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A2's quantification over the explicit set of 20 tests does not make it a bundle. A1 is a factual test-result proposition rather than a policy conclusion. The base and counter assignments are realizable with only A1 changing: the base can have one or more recorded failures, while the counter can have all recorded results be passes, with current-build approval and no evidence or documentation issues in both. Policy evidence preserves the state-originated final-build acceptance requirement, role routing, and newer-build supersession rule. Rules appearing only in the retained questions object need not be duplicated. The rule table may validly abstain on moderate-completion scenarios.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A1 entails that at least one mandatory test failed on the final build. The explicit criterion assigns level 0 whenever any mandatory final-build test fails, regardless of the other atoms.", "rule_index": 0, "sound": true}, {"reason": "Refuted A1 excludes every failed mandatory test, while supported A2 establishes that every mandatory test has a pass-or-fail result; together these entail that all 20 passed. A3 establishes approval naming the final build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions are sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At least one of the 20 mandatory tests failed on the build submitted as final for the Orion payroll-export package."}, {"id": "A2", "statement": "Each of the 20 mandatory tests has a recorded pass-or-fail result on the build submitted as final for the Orion payroll-export package."}, {"id": "A3", "statement": "The business approval names the build submitted as final for the Orion payroll-export package."}, {"id": "A4", "statement": "An unresolved evidence issue remains for the build submitted as final for the Orion payroll-export package."}, {"id": "A5", "statement": "An unresolved documentation issue remains for the build submitted as final for the Orion payroll-export package."}], "base_state_json": "{\"context\":\"Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.\",\"case_note\":[\"At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731.\",\"At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded FAIL for test 14 and PASS for each of tests 1–13 and 15–20.\",\"The coordinator treated the signed approval, submission register, and result ledger as authenticated closeout records for the package.\"],\"request\":\"Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["case_note", "0"], "text": "At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731."}, {"path": ["case_note", "1"], "text": "At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded FAIL for test 14 and PASS for each of tests 1–13 and 15–20."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build."}, {"path": ["context"], "text": "The delivery owner fixes product defects; the business approver renews outdated approvals."}, {"path": ["context"], "text": "Evidence for a newer build supersedes evidence for earlier builds."}, {"path": ["request"], "text": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}], "rules": [{"justification": "A supported A1 establishes a failed mandatory test on the final build. The low-completion criterion therefore requires non-acceptance, and the failed-test blocker is routed to the delivery owner because that role fixes product defects.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}]}, {"justification": "A2 establishes a pass-or-fail result for every mandatory test, while refuted A1 excludes any failed result, so all 20 mandatory tests pass on the final build. A3 supplies approval naming that build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions satisfy high completion.", "target": "2", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731.", "negative_left": "At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731.", "negative_right": "At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded PASS for each of tests 1–20.", "right": "At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded FAIL for test 14 and PASS for each of tests 1–13 and 15–20."}, "verifier_independent_model": false}, "family": "scale-diverse-059-001", "id": "scale-diverse-059-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low completion: Do not accept. At least one mandatory final-build test fails or required final-build approval is missing; route each blocker to its responsible role.", "1 — Moderate completion: Conditionally accept only when all mandatory final-build tests pass and approval exists, but minor non-blocking documentation corrections remain.", "2 — High completion: Accept fully. Every mandatory test passes on the final build, the business approval names that build, and no unresolved evidence or documentation issues remain."], "instructions": "Apply the latest-build rule and the acceptance criteria. Choose exactly one ordered level. Identify acceptance status and route any failed test or missing current-build approval.", "type": "score"}}, "state": {"case_note": ["At 16:00 UTC on 14 August 2026, Orion payroll-export package build OX-731 was submitted as final, the business approval named OX-731, and the submission register showed no unresolved evidence or documentation issues for OX-731.", "At 15:42 UTC on 14 August 2026, the result ledger for the complete set of 20 mandatory tests on Orion payroll-export package build OX-731 recorded PASS for each of tests 1–20.", "The coordinator treated the signed approval, submission register, and result ledger as authenticated closeout records for the package."], "context": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.", "request": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}}, "method": "c2d", "provenance": {"source_id": "diverse-059", "source_is_synthetic": true, "source_sha256": "f8f6806cfdd18388512db0984c3eb1386d711a91bbc8e11806a11e30573c54e8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same acceptance requirements, role responsibilities, latest-build rule, request scope, Orion package, final-build path, and relevant build/time bindings. Each context contains exactly two complete factual evidence sentences. The counterfactual changes only Q20 from code 4 (fail) to code 7 (pass), creating no duplicate or contradictory measurement. Neither context states a completion level, acceptance decision, routing result, proposition ID, classifier instruction, or gold answer; the ledger-code mapping is a factual record needed to interpret the measurements rather than an output answer code.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A2's quantification over the explicit set of 20 tests does not make it a bundle. A1 is a factual test-result proposition rather than a policy conclusion. The base and counter assignments are realizable with only A1 changing: the base can have one or more recorded failures, while the counter can have all recorded results be passes, with current-build approval and no evidence or documentation issues in both. Policy evidence preserves the state-originated final-build acceptance requirement, role routing, and newer-build supersession rule. Rules appearing only in the retained questions object need not be duplicated. The rule table may validly abstain on moderate-completion scenarios.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A1 entails that at least one mandatory test failed on the final build. The explicit criterion assigns level 0 whenever any mandatory final-build test fails, regardless of the other atoms.", "rule_index": 0, "sound": true}, {"reason": "Refuted A1 excludes every failed mandatory test, while supported A2 establishes that every mandatory test has a pass-or-fail result; together these entail that all 20 passed. A3 establishes approval naming the final build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions are sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At least one of the 20 mandatory tests failed on the build submitted as final for the Orion payroll-export package."}, {"id": "A2", "statement": "Each of the 20 mandatory tests has a recorded pass-or-fail result on the build submitted as final for the Orion payroll-export package."}, {"id": "A3", "statement": "The business approval names the build submitted as final for the Orion payroll-export package."}, {"id": "A4", "statement": "An unresolved evidence issue remains for the build submitted as final for the Orion payroll-export package."}, {"id": "A5", "statement": "An unresolved documentation issue remains for the build submitted as final for the Orion payroll-export package."}], "base_state_json": "{\"context\":\"Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.\",\"evidence\":[\"The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731.\",\"The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=4.\"],\"request\":\"Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731."}, {"path": ["evidence", "1"], "text": "The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=4."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build."}, {"path": ["context"], "text": "The delivery owner fixes product defects; the business approver renews outdated approvals."}, {"path": ["context"], "text": "Evidence for a newer build supersedes evidence for earlier builds."}, {"path": ["request"], "text": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}], "rules": [{"justification": "A supported A1 establishes a failed mandatory test on the final build. The low-completion criterion therefore requires non-acceptance, and the failed-test blocker is routed to the delivery owner because that role fixes product defects.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}]}, {"justification": "A2 establishes a pass-or-fail result for every mandatory test, while refuted A1 excludes any failed result, so all 20 mandatory tests pass on the final build. A3 supplies approval naming that build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions satisfy high completion.", "target": "2", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731.", "negative_left": "The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731.", "negative_right": "The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=7.", "right": "The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=4."}, "verifier_independent_model": false}, "family": "scale-diverse-059-002", "id": "scale-diverse-059-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Low completion: Do not accept. At least one mandatory final-build test fails or required final-build approval is missing; route each blocker to its responsible role.", "1 — Moderate completion: Conditionally accept only when all mandatory final-build tests pass and approval exists, but minor non-blocking documentation corrections remain.", "2 — High completion: Accept fully. Every mandatory test passes on the final build, the business approval names that build, and no unresolved evidence or documentation issues remain."], "instructions": "Apply the latest-build rule and the acceptance criteria. Choose exactly one ordered level. Identify acceptance status and route any failed test or missing current-build approval.", "type": "score"}}, "state": {"context": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.", "evidence": ["The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731.", "The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=4."], "request": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}}, "method": "c2d", "provenance": {"source_id": "diverse-059", "source_is_synthetic": true, "source_sha256": "f8f6806cfdd18388512db0984c3eb1386d711a91bbc8e11806a11e30573c54e8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same acceptance requirements, role responsibilities, latest-build rule, request scope, Orion package, final-build path, and relevant build/time bindings. Each context contains exactly two complete factual evidence sentences. The counterfactual changes only Q20 from code 4 (fail) to code 7 (pass), creating no duplicate or contradictory measurement. Neither context states a completion level, acceptance decision, routing result, proposition ID, classifier instruction, or gold answer; the ledger-code mapping is a factual record needed to interpret the measurements rather than an output answer code.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A2's quantification over the explicit set of 20 tests does not make it a bundle. A1 is a factual test-result proposition rather than a policy conclusion. The base and counter assignments are realizable with only A1 changing: the base can have one or more recorded failures, while the counter can have all recorded results be passes, with current-build approval and no evidence or documentation issues in both. Policy evidence preserves the state-originated final-build acceptance requirement, role routing, and newer-build supersession rule. Rules appearing only in the retained questions object need not be duplicated. The rule table may validly abstain on moderate-completion scenarios.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A1 entails that at least one mandatory test failed on the final build. The explicit criterion assigns level 0 whenever any mandatory final-build test fails, regardless of the other atoms.", "rule_index": 0, "sound": true}, {"reason": "Refuted A1 excludes every failed mandatory test, while supported A2 establishes that every mandatory test has a pass-or-fail result; together these entail that all 20 passed. A3 establishes approval naming the final build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions are sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At least one of the 20 mandatory tests failed on the build submitted as final for the Orion payroll-export package."}, {"id": "A2", "statement": "Each of the 20 mandatory tests has a recorded pass-or-fail result on the build submitted as final for the Orion payroll-export package."}, {"id": "A3", "statement": "The business approval names the build submitted as final for the Orion payroll-export package."}, {"id": "A4", "statement": "An unresolved evidence issue remains for the build submitted as final for the Orion payroll-export package."}, {"id": "A5", "statement": "An unresolved documentation issue remains for the build submitted as final for the Orion payroll-export package."}], "base_state_json": "{\"context\":\"Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.\",\"evidence\":[\"The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731.\",\"The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=4.\"],\"request\":\"Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731."}, {"path": ["evidence", "1"], "text": "The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=4."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build."}, {"path": ["context"], "text": "The delivery owner fixes product defects; the business approver renews outdated approvals."}, {"path": ["context"], "text": "Evidence for a newer build supersedes evidence for earlier builds."}, {"path": ["request"], "text": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}], "rules": [{"justification": "A supported A1 establishes a failed mandatory test on the final build. The low-completion criterion therefore requires non-acceptance, and the failed-test blocker is routed to the delivery owner because that role fixes product defects.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}]}, {"justification": "A2 establishes a pass-or-fail result for every mandatory test, while refuted A1 excludes any failed result, so all 20 mandatory tests pass on the final build. A3 supplies approval naming that build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions satisfy high completion.", "target": "2", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731.", "negative_left": "The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731.", "negative_right": "The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=7.", "right": "The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=4."}, "verifier_independent_model": false}, "family": "scale-diverse-059-002", "id": "scale-diverse-059-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low completion: Do not accept. At least one mandatory final-build test fails or required final-build approval is missing; route each blocker to its responsible role.", "1 — Moderate completion: Conditionally accept only when all mandatory final-build tests pass and approval exists, but minor non-blocking documentation corrections remain.", "2 — High completion: Accept fully. Every mandatory test passes on the final build, the business approval names that build, and no unresolved evidence or documentation issues remain."], "instructions": "Apply the latest-build rule and the acceptance criteria. Choose exactly one ordered level. Identify acceptance status and route any failed test or missing current-build approval.", "type": "score"}}, "state": {"context": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.", "evidence": ["The Orion release record signed at 2026-08-14T16:20:00Z identifies build OPE-731 as the build submitted as final for the Orion payroll-export package, defines its 20 mandatory tests as Q01 through Q20, maps execution-ledger code 7 to pass and code 4 to fail, contains business approval naming OPE-731, and records zero unresolved evidence issues and zero unresolved documentation issues for OPE-731.", "The execution ledger timestamped 2026-08-14T15:55:00Z for build OPE-731 contains exactly these entries: Q01=7, Q02=7, Q03=7, Q04=7, Q05=7, Q06=7, Q07=7, Q08=7, Q09=7, Q10=7, Q11=7, Q12=7, Q13=7, Q14=7, Q15=7, Q16=7, Q17=7, Q18=7, Q19=7, and Q20=7."], "request": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}}, "method": "c2d", "provenance": {"source_id": "diverse-059", "source_is_synthetic": true, "source_sha256": "f8f6806cfdd18388512db0984c3eb1386d711a91bbc8e11806a11e30573c54e8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the governing acceptance, responsibility, and latest-build policies, with the unchanged questions object retaining all scoring criteria and instructions. The request and Orion payroll-export/final-build bindings remain unchanged. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the finalized OX-731 test result from T20 failing to all 20 tests passing, without conflicting with the approval, audit, or documentation facts. Neither context embeds a score, gold answer, answer code, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A2's quantification over the explicit set of 20 tests does not make it a bundle. A1 is a factual test-result proposition rather than a policy conclusion. The base and counter assignments are realizable with only A1 changing: the base can have one or more recorded failures, while the counter can have all recorded results be passes, with current-build approval and no evidence or documentation issues in both. Policy evidence preserves the state-originated final-build acceptance requirement, role routing, and newer-build supersession rule. Rules appearing only in the retained questions object need not be duplicated. The rule table may validly abstain on moderate-completion scenarios.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A1 entails that at least one mandatory test failed on the final build. The explicit criterion assigns level 0 whenever any mandatory final-build test fails, regardless of the other atoms.", "rule_index": 0, "sound": true}, {"reason": "Refuted A1 excludes every failed mandatory test, while supported A2 establishes that every mandatory test has a pass-or-fail result; together these entail that all 20 passed. A3 establishes approval naming the final build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions are sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At least one of the 20 mandatory tests failed on the build submitted as final for the Orion payroll-export package."}, {"id": "A2", "statement": "Each of the 20 mandatory tests has a recorded pass-or-fail result on the build submitted as final for the Orion payroll-export package."}, {"id": "A3", "statement": "The business approval names the build submitted as final for the Orion payroll-export package."}, {"id": "A4", "statement": "An unresolved evidence issue remains for the build submitted as final for the Orion payroll-export package."}, {"id": "A5", "statement": "An unresolved documentation issue remains for the build submitted as final for the Orion payroll-export package."}], "base_state_json": "{\"context\":\"Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.\",\"evidence\":[\"The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20.\",\"The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T19 and a failure for T20.\",\"The signed business approval dated 14 September 2026 specifically names Orion payroll-export build OX-731.\",\"The release-record audit closed after confirming that all expected evidence was present and that no evidence questions remained unresolved for OX-731.\",\"The documentation review also closed with every correction completed and no documentation issue left unresolved for OX-731.\"],\"request\":\"Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20."}, {"path": ["evidence", "1"], "text": "The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T19 and a failure for T20."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build."}, {"path": ["context"], "text": "The delivery owner fixes product defects; the business approver renews outdated approvals."}, {"path": ["context"], "text": "Evidence for a newer build supersedes evidence for earlier builds."}, {"path": ["request"], "text": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}], "rules": [{"justification": "A supported A1 establishes a failed mandatory test on the final build. The low-completion criterion therefore requires non-acceptance, and the failed-test blocker is routed to the delivery owner because that role fixes product defects.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}]}, {"justification": "A2 establishes a pass-or-fail result for every mandatory test, while refuted A1 excludes any failed result, so all 20 mandatory tests pass on the final build. A3 supplies approval naming that build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions satisfy high completion.", "target": "2", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20.", "negative_left": "The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20.", "negative_right": "The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T20.", "right": "The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T19 and a failure for T20."}, "verifier_independent_model": false}, "family": "scale-diverse-059-003", "id": "scale-diverse-059-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Low completion: Do not accept. At least one mandatory final-build test fails or required final-build approval is missing; route each blocker to its responsible role.", "1 — Moderate completion: Conditionally accept only when all mandatory final-build tests pass and approval exists, but minor non-blocking documentation corrections remain.", "2 — High completion: Accept fully. Every mandatory test passes on the final build, the business approval names that build, and no unresolved evidence or documentation issues remain."], "instructions": "Apply the latest-build rule and the acceptance criteria. Choose exactly one ordered level. Identify acceptance status and route any failed test or missing current-build approval.", "type": "score"}}, "state": {"context": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.", "evidence": ["The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20.", "The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T19 and a failure for T20.", "The signed business approval dated 14 September 2026 specifically names Orion payroll-export build OX-731.", "The release-record audit closed after confirming that all expected evidence was present and that no evidence questions remained unresolved for OX-731.", "The documentation review also closed with every correction completed and no documentation issue left unresolved for OX-731."], "request": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}}, "method": "c2d", "provenance": {"source_id": "diverse-059", "source_is_synthetic": true, "source_sha256": "f8f6806cfdd18388512db0984c3eb1386d711a91bbc8e11806a11e30573c54e8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the governing acceptance, responsibility, and latest-build policies, with the unchanged questions object retaining all scoring criteria and instructions. The request and Orion payroll-export/final-build bindings remain unchanged. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the finalized OX-731 test result from T20 failing to all 20 tests passing, without conflicting with the approval, audit, or documentation facts. Neither context embeds a score, gold answer, answer code, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A2's quantification over the explicit set of 20 tests does not make it a bundle. A1 is a factual test-result proposition rather than a policy conclusion. The base and counter assignments are realizable with only A1 changing: the base can have one or more recorded failures, while the counter can have all recorded results be passes, with current-build approval and no evidence or documentation issues in both. Policy evidence preserves the state-originated final-build acceptance requirement, role routing, and newer-build supersession rule. Rules appearing only in the retained questions object need not be duplicated. The rule table may validly abstain on moderate-completion scenarios.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A1 entails that at least one mandatory test failed on the final build. The explicit criterion assigns level 0 whenever any mandatory final-build test fails, regardless of the other atoms.", "rule_index": 0, "sound": true}, {"reason": "Refuted A1 excludes every failed mandatory test, while supported A2 establishes that every mandatory test has a pass-or-fail result; together these entail that all 20 passed. A3 establishes approval naming the final build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions are sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At least one of the 20 mandatory tests failed on the build submitted as final for the Orion payroll-export package."}, {"id": "A2", "statement": "Each of the 20 mandatory tests has a recorded pass-or-fail result on the build submitted as final for the Orion payroll-export package."}, {"id": "A3", "statement": "The business approval names the build submitted as final for the Orion payroll-export package."}, {"id": "A4", "statement": "An unresolved evidence issue remains for the build submitted as final for the Orion payroll-export package."}, {"id": "A5", "statement": "An unresolved documentation issue remains for the build submitted as final for the Orion payroll-export package."}], "base_state_json": "{\"context\":\"Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.\",\"evidence\":[\"The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20.\",\"The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T19 and a failure for T20.\",\"The signed business approval dated 14 September 2026 specifically names Orion payroll-export build OX-731.\",\"The release-record audit closed after confirming that all expected evidence was present and that no evidence questions remained unresolved for OX-731.\",\"The documentation review also closed with every correction completed and no documentation issue left unresolved for OX-731.\"],\"request\":\"Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20."}, {"path": ["evidence", "1"], "text": "The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T19 and a failure for T20."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build."}, {"path": ["context"], "text": "The delivery owner fixes product defects; the business approver renews outdated approvals."}, {"path": ["context"], "text": "Evidence for a newer build supersedes evidence for earlier builds."}, {"path": ["request"], "text": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}], "rules": [{"justification": "A supported A1 establishes a failed mandatory test on the final build. The low-completion criterion therefore requires non-acceptance, and the failed-test blocker is routed to the delivery owner because that role fixes product defects.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}]}, {"justification": "A2 establishes a pass-or-fail result for every mandatory test, while refuted A1 excludes any failed result, so all 20 mandatory tests pass on the final build. A3 supplies approval naming that build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions satisfy high completion.", "target": "2", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20.", "negative_left": "The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20.", "negative_right": "The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T20.", "right": "The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T19 and a failure for T20."}, "verifier_independent_model": false}, "family": "scale-diverse-059-003", "id": "scale-diverse-059-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low completion: Do not accept. At least one mandatory final-build test fails or required final-build approval is missing; route each blocker to its responsible role.", "1 — Moderate completion: Conditionally accept only when all mandatory final-build tests pass and approval exists, but minor non-blocking documentation corrections remain.", "2 — High completion: Accept fully. Every mandatory test passes on the final build, the business approval names that build, and no unresolved evidence or documentation issues remain."], "instructions": "Apply the latest-build rule and the acceptance criteria. Choose exactly one ordered level. Identify acceptance status and route any failed test or missing current-build approval.", "type": "score"}}, "state": {"context": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds.", "evidence": ["The Orion payroll-export operational handoff at 16:00 UTC on 14 September 2026 designated build OX-731 as the build submitted as final and identified its complete set of 20 mandatory tests as T01 through T20.", "The finalized test ledger for Orion payroll-export build OX-731 at 15:42 UTC on 14 September 2026 records passes for T01 through T20.", "The signed business approval dated 14 September 2026 specifically names Orion payroll-export build OX-731.", "The release-record audit closed after confirming that all expected evidence was present and that no evidence questions remained unresolved for OX-731.", "The documentation review also closed with every correction completed and no documentation issue left unresolved for OX-731."], "request": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}}, "method": "c2d", "provenance": {"source_id": "diverse-059", "source_is_synthetic": true, "source_sha256": "f8f6806cfdd18388512db0984c3eb1386d711a91bbc8e11806a11e30573c54e8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same acceptance policy and preserve the build 4.2, test-count, approval-scope, and date bindings. The two evidence spans are complete factual sentences. Changing BA-1’s recorded scope code from S-731 to S-846 is coherent with the unchanged fact that S-731 is the sole valid code; it creates an invalidly scoped artifact rather than a contradictory duplicate measurement. Neither context contains an answer code, gold label, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"},{\"speaker\":\"Records clerk\",\"text\":\"On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable.\"},{\"speaker\":\"Approval clerk\",\"text\":\"On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-731.\"},{\"speaker\":\"Submission custodian\",\"text\":\"The submission log closed on 9 May with BA-1 as the only business-approval artifact received.\"},{\"speaker\":\"Test auditor\",\"text\":\"On 10 May, I checked the controlled test-run metadata for build 4.2. It verifies a failing result for the specified CSV-export test. The same metadata verifies passing results for each of the other 19 specified tests. No replacement run or QA test approval was submitted.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable."}, {"path": ["2", "text"], "text": "On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-731."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable.", "negative_left": "On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable.", "negative_right": "On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-846.", "right": "On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-731."}, "verifier_independent_model": false}, "family": "scale-diverse-060-001", "id": "scale-diverse-060-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Project coordinator", "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}, {"speaker": "Records clerk", "text": "On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable."}, {"speaker": "Approval clerk", "text": "On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-731."}, {"speaker": "Submission custodian", "text": "The submission log closed on 9 May with BA-1 as the only business-approval artifact received."}, {"speaker": "Test auditor", "text": "On 10 May, I checked the controlled test-run metadata for build 4.2. It verifies a failing result for the specified CSV-export test. The same metadata verifies passing results for each of the other 19 specified tests. No replacement run or QA test approval was submitted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same acceptance policy and preserve the build 4.2, test-count, approval-scope, and date bindings. The two evidence spans are complete factual sentences. Changing BA-1’s recorded scope code from S-731 to S-846 is coherent with the unchanged fact that S-731 is the sole valid code; it creates an invalidly scoped artifact rather than a contradictory duplicate measurement. Neither context contains an answer code, gold label, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"},{\"speaker\":\"Records clerk\",\"text\":\"On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable.\"},{\"speaker\":\"Approval clerk\",\"text\":\"On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-731.\"},{\"speaker\":\"Submission custodian\",\"text\":\"The submission log closed on 9 May with BA-1 as the only business-approval artifact received.\"},{\"speaker\":\"Test auditor\",\"text\":\"On 10 May, I checked the controlled test-run metadata for build 4.2. It verifies a failing result for the specified CSV-export test. The same metadata verifies passing results for each of the other 19 specified tests. No replacement run or QA test approval was submitted.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable."}, {"path": ["2", "text"], "text": "On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-731."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable.", "negative_left": "On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable.", "negative_right": "On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-846.", "right": "On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-731."}, "verifier_independent_model": false}, "family": "scale-diverse-060-001", "id": "scale-diverse-060-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Project coordinator", "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}, {"speaker": "Records clerk", "text": "On 8 May 2026, the controlled release register assigned the build 4.2 deliverable the sole valid approval-scope code S-731 and recorded that every other scope code denotes a different deliverable."}, {"speaker": "Approval clerk", "text": "On 9 May 2026, the approval-scope field of signed business-approval artifact BA-1 contained the code S-846."}, {"speaker": "Submission custodian", "text": "The submission log closed on 9 May with BA-1 as the only business-approval artifact received."}, {"speaker": "Test auditor", "text": "On 10 May, I checked the controlled test-run metadata for build 4.2. It verifies a failing result for the specified CSV-export test. The same metadata verifies passing results for each of the other 19 specified tests. No replacement run or QA test approval was submitted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring rubric, while both contexts retain the governing acceptance policy and required remediation routing from the original state. The build, deliverable, test set, and 2026-09-17 handoff bindings remain fixed; only BA-1's observed approval-scope identifier changes. The two focus spans are complete factual sentences. In the counterfactual, BA-1 scopes HP-904 while the deliverable is HP-731, and the inventory statement does not claim that BA-1 itself corresponds to the deliverable, so there is no contradiction. Neither context contains a gold score, answer code, proposition identifier, output instruction, or impermissible label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The locked run register for build 4.2 records a verified failing result for the specified CSV-export test. Each of the other 19 specified tests has an individual verified passing result. QA approval has not yet been signed.\"},{\"speaker\":\"Records custodian\",\"text\":\"At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-731.\"},{\"speaker\":\"Configuration controller\",\"text\":\"At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731.\"},{\"speaker\":\"Approval clerk\",\"text\":\"BA-1 bears the business approver's signature. The submission inventory confirms that no submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable.\"},{\"speaker\":\"Handoff lead\",\"text\":\"The Delivery owner has the CSV-export correction or rerun task. Once new evidence is posted, the Quality reviewer is assigned to verify it and provide QA approval.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-731."}, {"path": ["3", "text"], "text": "At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-731.", "negative_left": "At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-904.", "negative_right": "At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731.", "right": "At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731."}, "verifier_independent_model": false}, "family": "scale-diverse-060-003", "id": "scale-diverse-060-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Project coordinator", "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}, {"speaker": "Quality reviewer", "text": "The locked run register for build 4.2 records a verified failing result for the specified CSV-export test. Each of the other 19 specified tests has an individual verified passing result. QA approval has not yet been signed."}, {"speaker": "Records custodian", "text": "At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-731."}, {"speaker": "Configuration controller", "text": "At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731."}, {"speaker": "Approval clerk", "text": "BA-1 bears the business approver's signature. The submission inventory confirms that no submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}, {"speaker": "Handoff lead", "text": "The Delivery owner has the CSV-export correction or rerun task. Once new evidence is posted, the Quality reviewer is assigned to verify it and provide QA approval."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring rubric, while both contexts retain the governing acceptance policy and required remediation routing from the original state. The build, deliverable, test set, and 2026-09-17 handoff bindings remain fixed; only BA-1's observed approval-scope identifier changes. The two focus spans are complete factual sentences. In the counterfactual, BA-1 scopes HP-904 while the deliverable is HP-731, and the inventory statement does not claim that BA-1 itself corresponds to the deliverable, so there is no contradiction. Neither context contains a gold score, answer code, proposition identifier, output instruction, or impermissible label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The locked run register for build 4.2 records a verified failing result for the specified CSV-export test. Each of the other 19 specified tests has an individual verified passing result. QA approval has not yet been signed.\"},{\"speaker\":\"Records custodian\",\"text\":\"At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-731.\"},{\"speaker\":\"Configuration controller\",\"text\":\"At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731.\"},{\"speaker\":\"Approval clerk\",\"text\":\"BA-1 bears the business approver's signature. The submission inventory confirms that no submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable.\"},{\"speaker\":\"Handoff lead\",\"text\":\"The Delivery owner has the CSV-export correction or rerun task. Once new evidence is posted, the Quality reviewer is assigned to verify it and provide QA approval.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-731."}, {"path": ["3", "text"], "text": "At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-731.", "negative_left": "At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-904.", "negative_right": "At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731.", "right": "At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731."}, "verifier_independent_model": false}, "family": "scale-diverse-060-003", "id": "scale-diverse-060-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Project coordinator", "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}, {"speaker": "Quality reviewer", "text": "The locked run register for build 4.2 records a verified failing result for the specified CSV-export test. Each of the other 19 specified tests has an individual verified passing result. QA approval has not yet been signed."}, {"speaker": "Records custodian", "text": "At the 2026-09-17 operational handoff, BA-1's approval-scope field contained only the unique deliverable identifier HP-904."}, {"speaker": "Configuration controller", "text": "At the 2026-09-17 operational handoff, the build 4.2 deliverable's unique identifier was HP-731."}, {"speaker": "Approval clerk", "text": "BA-1 bears the business approver's signature. The submission inventory confirms that no submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}, {"speaker": "Handoff lead", "text": "The Delivery owner has the CSV-export correction or rerun task. Once new evidence is posted, the Quality reviewer is assigned to verify it and provide QA approval."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same rubric statement, request, learner, assessment, diagnostic-item designations, and relevant item binding; only the item 3 keyed value changes. The two focus-evidence spans are complete factual sentences. The counterfactual remains consistent: item 3 changes from a match to a mismatch, while the stated outcomes for items 1, 2, 4, and 5 remain unchanged. Neither context includes an answer option, urgency result, classifier instruction, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.\",\"chronology\":[\"At 09:10 on 12 March 2026, the course instructor opened Miko’s five-item check; the file contained a recorded response and an answer-key entry for every item.\",\"At 09:12, the file identified items 1 and 5 as diagnostic and the other items as nondiagnostic, with Miko’s shown work attached for both diagnostic items.\",\"At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check.\",\"At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “38 square centimeters.”\",\"Comparison of the remaining entries found that Miko’s responses to items 1 and 5 did not match their keyed answers, while the responses to items 2 and 4 matched theirs.\",\"The diagnostic work showed different wrong methods: item 1 added side lengths, whereas item 5 multiplied one listed side length by itself.\",\"At 09:22, the instructor confirmed that Miko’s lesson record was available and that the learning support tutor remained available.\"],\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["chronology", "2"], "text": "At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check."}, {"path": ["chronology", "3"], "text": "At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “38 square centimeters.”"}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check.", "negative_left": "At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check.", "negative_right": "At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “43 square centimeters.”", "right": "At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “38 square centimeters.”"}, "verifier_independent_model": false}, "family": "scale-diverse-061-001", "id": "scale-diverse-061-001-base", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"chronology": ["At 09:10 on 12 March 2026, the course instructor opened Miko’s five-item check; the file contained a recorded response and an answer-key entry for every item.", "At 09:12, the file identified items 1 and 5 as diagnostic and the other items as nondiagnostic, with Miko’s shown work attached for both diagnostic items.", "At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check.", "At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “38 square centimeters.”", "Comparison of the remaining entries found that Miko’s responses to items 1 and 5 did not match their keyed answers, while the responses to items 2 and 4 matched theirs.", "The diagnostic work showed different wrong methods: item 1 added side lengths, whereas item 5 multiplied one listed side length by itself.", "At 09:22, the instructor confirmed that Miko’s lesson record was available and that the learning support tutor remained available."], "context": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.", "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "B — Not yet mastered; practice set; U1"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same rubric statement, request, learner, assessment, diagnostic-item designations, and relevant item binding; only the item 3 keyed value changes. The two focus-evidence spans are complete factual sentences. The counterfactual remains consistent: item 3 changes from a match to a mismatch, while the stated outcomes for items 1, 2, 4, and 5 remain unchanged. Neither context includes an answer option, urgency result, classifier instruction, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.\",\"chronology\":[\"At 09:10 on 12 March 2026, the course instructor opened Miko’s five-item check; the file contained a recorded response and an answer-key entry for every item.\",\"At 09:12, the file identified items 1 and 5 as diagnostic and the other items as nondiagnostic, with Miko’s shown work attached for both diagnostic items.\",\"At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check.\",\"At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “38 square centimeters.”\",\"Comparison of the remaining entries found that Miko’s responses to items 1 and 5 did not match their keyed answers, while the responses to items 2 and 4 matched theirs.\",\"The diagnostic work showed different wrong methods: item 1 added side lengths, whereas item 5 multiplied one listed side length by itself.\",\"At 09:22, the instructor confirmed that Miko’s lesson record was available and that the learning support tutor remained available.\"],\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["chronology", "2"], "text": "At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check."}, {"path": ["chronology", "3"], "text": "At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “38 square centimeters.”"}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check.", "negative_left": "At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check.", "negative_right": "At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “43 square centimeters.”", "right": "At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “38 square centimeters.”"}, "verifier_independent_model": false}, "family": "scale-diverse-061-001", "id": "scale-diverse-061-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"chronology": ["At 09:10 on 12 March 2026, the course instructor opened Miko’s five-item check; the file contained a recorded response and an answer-key entry for every item.", "At 09:12, the file identified items 1 and 5 as diagnostic and the other items as nondiagnostic, with Miko’s shown work attached for both diagnostic items.", "At 09:14 on 12 March 2026, Miko entered “38 square centimeters” as the response to item 3 of the five-item check.", "At 09:20 on 12 March 2026, the answer key for item 3 of Miko's five-item check listed “43 square centimeters.”", "Comparison of the remaining entries found that Miko’s responses to items 1 and 5 did not match their keyed answers, while the responses to items 2 and 4 matched theirs.", "The diagnostic work showed different wrong methods: item 1 added side lengths, whereas item 5 multiplied one listed side length by itself.", "At 09:22, the instructor confirmed that Miko’s lesson record was available and that the learning support tutor remained available."], "context": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.", "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Not mastered; tutor remediation; U3"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same decision criteria, rule ordering, request, learner, five-item assessment, diagnostic-item designations, and relevant lesson-record scope. The two focus spans are complete factual sentences. The counterfactual changes only item 3’s finalized key from option D to option B; this makes item 3 incorrect while remaining consistent with the unchanged reconciliation statements about items 1, 2, 4, and 5. References to options D and B are ordinary item-response and answer-key facts, not embedded classifier outputs or decision-label leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"A — Mastered; independent extension; U0\":\"Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.\",\"B — Not yet mastered; practice set; U1\":\"Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.\",\"C — Not mastered; tutor remediation; U3\":\"Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.\",\"D — Not mastered; instructor misconception mini-lesson; U2\":\"Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.\",\"E — Decision deferred; evidence completion; no urgency assigned\":\"Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing.\"},\"instructions\":\"Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.\",\"type\":\"choice\"}},\"state\":{\"context\":\"The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.\",\"evidence\":[\"The QF-27 audit contains a recorded response and a finalized answer-key entry for each of its five items. Items 1 and 5 are the only diagnostic items, and shown work is recorded for both.\",\"At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D.\",\"The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option D.\",\"Reconciliation found that Miko's responses for items 1 and 5 did not match their respective key entries, while the responses for items 2 and 4 matched theirs.\",\"Item 1 work adds all side lengths; item 5 work multiplies adjacent dimensions, so the wrong methods shown on the two diagnostic items are different.\",\"Miko's lesson record is available and confirms attendance at the area-versus-perimeter lesson.\"],\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D."}, {"path": ["state", "evidence", "2"], "text": "The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option D."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D.", "negative_left": "At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D.", "negative_right": "The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option B.", "right": "The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option D."}, "verifier_independent_model": false}, "family": "scale-diverse-061-002", "id": "scale-diverse-061-002-base", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.", "evidence": ["The QF-27 audit contains a recorded response and a finalized answer-key entry for each of its five items. Items 1 and 5 are the only diagnostic items, and shown work is recorded for both.", "At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D.", "The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option D.", "Reconciliation found that Miko's responses for items 1 and 5 did not match their respective key entries, while the responses for items 2 and 4 matched theirs.", "Item 1 work adds all side lengths; item 5 work multiplies adjacent dimensions, so the wrong methods shown on the two diagnostic items are different.", "Miko's lesson record is available and confirms attendance at the area-versus-perimeter lesson."], "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "B — Not yet mastered; practice set; U1"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same decision criteria, rule ordering, request, learner, five-item assessment, diagnostic-item designations, and relevant lesson-record scope. The two focus spans are complete factual sentences. The counterfactual changes only item 3’s finalized key from option D to option B; this makes item 3 incorrect while remaining consistent with the unchanged reconciliation statements about items 1, 2, 4, and 5. References to options D and B are ordinary item-response and answer-key facts, not embedded classifier outputs or decision-label leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"A — Mastered; independent extension; U0\":\"Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.\",\"B — Not yet mastered; practice set; U1\":\"Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.\",\"C — Not mastered; tutor remediation; U3\":\"Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.\",\"D — Not mastered; instructor misconception mini-lesson; U2\":\"Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.\",\"E — Decision deferred; evidence completion; no urgency assigned\":\"Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing.\"},\"instructions\":\"Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.\",\"type\":\"choice\"}},\"state\":{\"context\":\"The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.\",\"evidence\":[\"The QF-27 audit contains a recorded response and a finalized answer-key entry for each of its five items. Items 1 and 5 are the only diagnostic items, and shown work is recorded for both.\",\"At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D.\",\"The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option D.\",\"Reconciliation found that Miko's responses for items 1 and 5 did not match their respective key entries, while the responses for items 2 and 4 matched theirs.\",\"Item 1 work adds all side lengths; item 5 work multiplies adjacent dimensions, so the wrong methods shown on the two diagnostic items are different.\",\"Miko's lesson record is available and confirms attendance at the area-versus-perimeter lesson.\"],\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D."}, {"path": ["state", "evidence", "2"], "text": "The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option D."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D.", "negative_left": "At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D.", "negative_right": "The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option B.", "right": "The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option D."}, "verifier_independent_model": false}, "family": "scale-diverse-061-002", "id": "scale-diverse-061-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.", "evidence": ["The QF-27 audit contains a recorded response and a finalized answer-key entry for each of its five items. Items 1 and 5 are the only diagnostic items, and shown work is recorded for both.", "At 10:14 UTC on 12 September 2026, Miko's submitted response for item 3 of five-item check QF-27 was recorded as option D.", "The finalized answer key for item 3 of five-item check QF-27, timestamped 09:40 UTC on 12 September 2026, specifies option B.", "Reconciliation found that Miko's responses for items 1 and 5 did not match their respective key entries, while the responses for items 2 and 4 matched theirs.", "Item 1 work adds all side lengths; item 5 work multiplies adjacent dimensions, so the wrong methods shown on the two diagnostic items are different.", "Miko's lesson record is available and confirms attendance at the area-versus-perimeter lesson."], "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Not mastered; tutor remediation; U3"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full decision rubric, while both contexts retain the original-state policy concerning repeated diagnostic methods. Miko, the five-item assessment, and the requested mastery/route/urgency decision remain bound consistently. The two focus spans are complete factual sentences. The counterfactual changes only item 3’s key entry from 86 to 68 square centimeters and creates no duplicate or conflicting measurement within its own context. Neither context states a decision option, urgency result, answer code, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.\",\"evidence\":[\"At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”\",\"The answer key transferred at the 2026-09-17 assessment handoff specifies “86 square centimeters” for item 3.\",\"The transferred key also specifies item 1 as 24 square centimeters, item 2 as 20 centimeters, item 4 as 22 centimeters, and item 5 as 16 square centimeters.\",\"Miko’s other recorded responses are 20 centimeters for item 1, 20 centimeters for item 2, 22 centimeters for item 4, and 12 square centimeters for item 5.\",\"Items 1 and 5 are the only diagnostic items; items 2, 3, and 4 are marked non-diagnostic. Miko’s shown work adds all side lengths on item 1 but subtracts one side length from another on item 5, so the two wrong methods differ.\",\"The handoff contains a response and key entry for every item, diagnostic designations and shown work as noted, and Miko’s available lesson record confirming attendance at the area-versus-perimeter lesson.\"],\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "0"], "text": "At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”"}, {"path": ["evidence", "1"], "text": "The answer key transferred at the 2026-09-17 assessment handoff specifies “86 square centimeters” for item 3."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”", "negative_left": "At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”", "negative_right": "The answer key transferred at the 2026-09-17 assessment handoff specifies “68 square centimeters” for item 3.", "right": "The answer key transferred at the 2026-09-17 assessment handoff specifies “86 square centimeters” for item 3."}, "verifier_independent_model": false}, "family": "scale-diverse-061-003", "id": "scale-diverse-061-003-base", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.", "evidence": ["At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”", "The answer key transferred at the 2026-09-17 assessment handoff specifies “86 square centimeters” for item 3.", "The transferred key also specifies item 1 as 24 square centimeters, item 2 as 20 centimeters, item 4 as 22 centimeters, and item 5 as 16 square centimeters.", "Miko’s other recorded responses are 20 centimeters for item 1, 20 centimeters for item 2, 22 centimeters for item 4, and 12 square centimeters for item 5.", "Items 1 and 5 are the only diagnostic items; items 2, 3, and 4 are marked non-diagnostic. Miko’s shown work adds all side lengths on item 1 but subtracts one side length from another on item 5, so the two wrong methods differ.", "The handoff contains a response and key entry for every item, diagnostic designations and shown work as noted, and Miko’s available lesson record confirming attendance at the area-versus-perimeter lesson."], "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "B — Not yet mastered; practice set; U1"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full decision rubric, while both contexts retain the original-state policy concerning repeated diagnostic methods. Miko, the five-item assessment, and the requested mastery/route/urgency decision remain bound consistently. The two focus spans are complete factual sentences. The counterfactual changes only item 3’s key entry from 86 to 68 square centimeters and creates no duplicate or conflicting measurement within its own context. Neither context states a decision option, urgency result, answer code, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.\",\"evidence\":[\"At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”\",\"The answer key transferred at the 2026-09-17 assessment handoff specifies “86 square centimeters” for item 3.\",\"The transferred key also specifies item 1 as 24 square centimeters, item 2 as 20 centimeters, item 4 as 22 centimeters, and item 5 as 16 square centimeters.\",\"Miko’s other recorded responses are 20 centimeters for item 1, 20 centimeters for item 2, 22 centimeters for item 4, and 12 square centimeters for item 5.\",\"Items 1 and 5 are the only diagnostic items; items 2, 3, and 4 are marked non-diagnostic. Miko’s shown work adds all side lengths on item 1 but subtracts one side length from another on item 5, so the two wrong methods differ.\",\"The handoff contains a response and key entry for every item, diagnostic designations and shown work as noted, and Miko’s available lesson record confirming attendance at the area-versus-perimeter lesson.\"],\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "0"], "text": "At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”"}, {"path": ["evidence", "1"], "text": "The answer key transferred at the 2026-09-17 assessment handoff specifies “86 square centimeters” for item 3."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”", "negative_left": "At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”", "negative_right": "The answer key transferred at the 2026-09-17 assessment handoff specifies “68 square centimeters” for item 3.", "right": "The answer key transferred at the 2026-09-17 assessment handoff specifies “86 square centimeters” for item 3."}, "verifier_independent_model": false}, "family": "scale-diverse-061-003", "id": "scale-diverse-061-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.", "evidence": ["At the 2026-09-17 assessment handoff, Miko's recorded response to item 3 was “86 square centimeters.”", "The answer key transferred at the 2026-09-17 assessment handoff specifies “68 square centimeters” for item 3.", "The transferred key also specifies item 1 as 24 square centimeters, item 2 as 20 centimeters, item 4 as 22 centimeters, and item 5 as 16 square centimeters.", "Miko’s other recorded responses are 20 centimeters for item 1, 20 centimeters for item 2, 22 centimeters for item 4, and 12 square centimeters for item 5.", "Items 1 and 5 are the only diagnostic items; items 2, 3, and 4 are marked non-diagnostic. Miko’s shown work adds all side lengths on item 1 but subtracts one side length from another on item 5, so the two wrong methods differ.", "The handoff contains a response and key entry for every item, diagnostic designations and shown work as noted, and Miko’s available lesson record confirming attendance at the area-versus-perimeter lesson."], "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Not mastered; tutor remediation; U3"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual changes only the affirmative-version answer for QZ-417 item 4 from “C” to “B”; it leaves the rubric, five-item quiz scope, item path, entity, and timestamps unchanged. The two focus-evidence spans are complete factual sentences, the changed answer remains consistent with the 4/5 score and sole item-4 error, and neither context contains a rubric-option code, explicit decision, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a8 is one consistency relation over an explicit set of records rather than a bundle of policy requirements. The focus a4 is factual. The base and counter assignments can both be realized while changing only whether the item-4 response matches the affirmative-version answer; the response can remain incorrect in either case, preserving the 4/5 score and all other atom states. The policy evidence is an accurate citation from the original state. No additional state-derived rule is needed because the complete governing rubric remains available verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. The rule also establishes that item 4 is the negation probe, the selected response matches its affirmative-version answer, and all required records are present and consistent. These conditions sufficiently entail U2_negation_review and exclude EVIDENCE_HOLD and the competing scored outcomes.", "rule_index": 0, "sound": true}, {"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. With item 4 established as the negation probe and the affirmative-version match refuted, the U2 exception does not apply and U1_boundary_mastery is entailed. Record presence and consistency exclude EVIDENCE_HOLD.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The learner answered exactly four of the five quiz items correctly."}, {"id": "a2", "statement": "The learner answered quiz item 4 incorrectly."}, {"id": "a3", "statement": "Quiz item 4 is the negation probe."}, {"id": "a4", "statement": "The learner’s selected response on quiz item 4 matches the answer to the affirmative version of quiz item 4."}, {"id": "a5", "statement": "The learner’s quiz responses are present."}, {"id": "a6", "statement": "The answer key is present."}, {"id": "a7", "statement": "The lesson record is present."}, {"id": "a8", "statement": "The learner’s quiz responses, answer key, and lesson record contain no mutually inconsistent information relevant to scoring or routing this quiz."}], "base_state_json": "[{\"speaker\":\"Quiz system\",\"text\":\"At 10:12 UTC on 8 May 2026, QZ-417 opened with five items. Its submission audit later confirmed that all five learner-response fields were populated and retained.\"},{\"speaker\":\"Submission log\",\"text\":\"At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417.\"},{\"speaker\":\"Scoring log\",\"text\":\"At 10:16 UTC, the complete answer key scored the submission at exactly four correct answers out of five. Item 4 was the sole incorrect response. The item wording asked which claim was NOT supported, making it the quiz’s negation probe.\"},{\"speaker\":\"Answer-key audit\",\"text\":\"At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “C” as the answer to the affirmative version of quiz item 4.\"},{\"speaker\":\"Records officer\",\"text\":\"The lesson record was present and showed completion of the affirmative-claims activity. A cross-check of the learner responses, answer key, scoring log, and lesson record found no mutually inconsistent information relevant to scoring or routing QZ-417.\"},{\"speaker\":\"Course instructor\",\"text\":\"Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417."}, {"path": ["3", "text"], "text": "At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “C” as the answer to the affirmative version of quiz item 4."}], "policy_evidence": [{"path": ["3", "text"], "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}], "rules": [{"justification": "The records are present and consistent, the score is 4/5, item 4 is necessarily the sole miss, item 4 is the negation probe, and the selected response matches its affirmative-version answer.", "target": "U2_negation_review", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The records are present and consistent, the score is 4/5, and item 4 is necessarily the sole miss; although it is the negation probe, the selected response does not match its affirmative-version answer.", "target": "U1_boundary_mastery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417.", "negative_left": "At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417.", "negative_right": "At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “B” as the answer to the affirmative version of quiz item 4.", "right": "At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “C” as the answer to the affirmative version of quiz item 4."}, "verifier_independent_model": false}, "family": "scale-diverse-062-001", "id": "scale-diverse-062-001-base", "input": {"questions": {"decision": {"criteria": {"EVIDENCE_HOLD": "No urgency assigned: Responses, answer key, or lesson record are missing or inconsistent, so scoring or routing cannot yet be completed.", "U0_full_mastery": "Urgency 0: Score is 5/5. Verify mastery; assign no review or support role.", "U1_boundary_mastery": "Urgency 1: Score is 4/5, but the sole miss either is not the negation probe or does not match the affirmative-version answer. Verify mastery; assign an optional independent recap.", "U2_negation_review": "Urgency 2: Score is 4/5, the only miss is the negation probe, and the chosen response matches the affirmative-version answer. Do not verify mastery; route to the course instructor’s NOT/EXCEPT contrast sort before the next quiz.", "U3_tutor_reteach": "Urgency 3: Score is exactly 3/5, regardless of error type. Do not verify mastery; route to a learning support tutor for a full guided practice set before the next quiz.", "U4_intensive_support": "Urgency 4: Score is 0–2/5. Do not verify mastery; route first to the course instructor for reteaching and then to the learning support tutor for supervised practice."}, "instructions": "Select the single rubric option that correctly verifies mastery status, assigns the prescribed route, and gives the ordered intervention urgency. Lower urgency numbers mean less urgent intervention.", "type": "choice"}}, "state": [{"speaker": "Quiz system", "text": "At 10:12 UTC on 8 May 2026, QZ-417 opened with five items. Its submission audit later confirmed that all five learner-response fields were populated and retained."}, {"speaker": "Submission log", "text": "At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417."}, {"speaker": "Scoring log", "text": "At 10:16 UTC, the complete answer key scored the submission at exactly four correct answers out of five. Item 4 was the sole incorrect response. The item wording asked which claim was NOT supported, making it the quiz’s negation probe."}, {"speaker": "Answer-key audit", "text": "At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “C” as the answer to the affirmative version of quiz item 4."}, {"speaker": "Records officer", "text": "The lesson record was present and showed completion of the affirmative-claims activity. A cross-check of the learner responses, answer key, scoring log, and lesson record found no mutually inconsistent information relevant to scoring or routing QZ-417."}, {"speaker": "Course instructor", "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}]}, "method": "c2d", "provenance": {"source_id": "diverse-062", "source_is_synthetic": true, "source_sha256": "571895f9ad2bd8e1cdf4993a3c848224bd6a8603b7160d6c25bc7527551b1fa8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "U2_negation_review"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual changes only the affirmative-version answer for QZ-417 item 4 from “C” to “B”; it leaves the rubric, five-item quiz scope, item path, entity, and timestamps unchanged. The two focus-evidence spans are complete factual sentences, the changed answer remains consistent with the 4/5 score and sole item-4 error, and neither context contains a rubric-option code, explicit decision, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a8 is one consistency relation over an explicit set of records rather than a bundle of policy requirements. The focus a4 is factual. The base and counter assignments can both be realized while changing only whether the item-4 response matches the affirmative-version answer; the response can remain incorrect in either case, preserving the 4/5 score and all other atom states. The policy evidence is an accurate citation from the original state. No additional state-derived rule is needed because the complete governing rubric remains available verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. The rule also establishes that item 4 is the negation probe, the selected response matches its affirmative-version answer, and all required records are present and consistent. These conditions sufficiently entail U2_negation_review and exclude EVIDENCE_HOLD and the competing scored outcomes.", "rule_index": 0, "sound": true}, {"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. With item 4 established as the negation probe and the affirmative-version match refuted, the U2 exception does not apply and U1_boundary_mastery is entailed. Record presence and consistency exclude EVIDENCE_HOLD.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The learner answered exactly four of the five quiz items correctly."}, {"id": "a2", "statement": "The learner answered quiz item 4 incorrectly."}, {"id": "a3", "statement": "Quiz item 4 is the negation probe."}, {"id": "a4", "statement": "The learner’s selected response on quiz item 4 matches the answer to the affirmative version of quiz item 4."}, {"id": "a5", "statement": "The learner’s quiz responses are present."}, {"id": "a6", "statement": "The answer key is present."}, {"id": "a7", "statement": "The lesson record is present."}, {"id": "a8", "statement": "The learner’s quiz responses, answer key, and lesson record contain no mutually inconsistent information relevant to scoring or routing this quiz."}], "base_state_json": "[{\"speaker\":\"Quiz system\",\"text\":\"At 10:12 UTC on 8 May 2026, QZ-417 opened with five items. Its submission audit later confirmed that all five learner-response fields were populated and retained.\"},{\"speaker\":\"Submission log\",\"text\":\"At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417.\"},{\"speaker\":\"Scoring log\",\"text\":\"At 10:16 UTC, the complete answer key scored the submission at exactly four correct answers out of five. Item 4 was the sole incorrect response. The item wording asked which claim was NOT supported, making it the quiz’s negation probe.\"},{\"speaker\":\"Answer-key audit\",\"text\":\"At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “C” as the answer to the affirmative version of quiz item 4.\"},{\"speaker\":\"Records officer\",\"text\":\"The lesson record was present and showed completion of the affirmative-claims activity. A cross-check of the learner responses, answer key, scoring log, and lesson record found no mutually inconsistent information relevant to scoring or routing QZ-417.\"},{\"speaker\":\"Course instructor\",\"text\":\"Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417."}, {"path": ["3", "text"], "text": "At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “C” as the answer to the affirmative version of quiz item 4."}], "policy_evidence": [{"path": ["3", "text"], "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}], "rules": [{"justification": "The records are present and consistent, the score is 4/5, item 4 is necessarily the sole miss, item 4 is the negation probe, and the selected response matches its affirmative-version answer.", "target": "U2_negation_review", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The records are present and consistent, the score is 4/5, and item 4 is necessarily the sole miss; although it is the negation probe, the selected response does not match its affirmative-version answer.", "target": "U1_boundary_mastery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417.", "negative_left": "At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417.", "negative_right": "At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “B” as the answer to the affirmative version of quiz item 4.", "right": "At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “C” as the answer to the affirmative version of quiz item 4."}, "verifier_independent_model": false}, "family": "scale-diverse-062-001", "id": "scale-diverse-062-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"EVIDENCE_HOLD": "No urgency assigned: Responses, answer key, or lesson record are missing or inconsistent, so scoring or routing cannot yet be completed.", "U0_full_mastery": "Urgency 0: Score is 5/5. Verify mastery; assign no review or support role.", "U1_boundary_mastery": "Urgency 1: Score is 4/5, but the sole miss either is not the negation probe or does not match the affirmative-version answer. Verify mastery; assign an optional independent recap.", "U2_negation_review": "Urgency 2: Score is 4/5, the only miss is the negation probe, and the chosen response matches the affirmative-version answer. Do not verify mastery; route to the course instructor’s NOT/EXCEPT contrast sort before the next quiz.", "U3_tutor_reteach": "Urgency 3: Score is exactly 3/5, regardless of error type. Do not verify mastery; route to a learning support tutor for a full guided practice set before the next quiz.", "U4_intensive_support": "Urgency 4: Score is 0–2/5. Do not verify mastery; route first to the course instructor for reteaching and then to the learning support tutor for supervised practice."}, "instructions": "Select the single rubric option that correctly verifies mastery status, assigns the prescribed route, and gives the ordered intervention urgency. Lower urgency numbers mean less urgent intervention.", "type": "choice"}}, "state": [{"speaker": "Quiz system", "text": "At 10:12 UTC on 8 May 2026, QZ-417 opened with five items. Its submission audit later confirmed that all five learner-response fields were populated and retained."}, {"speaker": "Submission log", "text": "At 10:14:22 UTC on 8 May 2026, the learner submitted response “C” for quiz item 4 of the five-item quiz identified as QZ-417."}, {"speaker": "Scoring log", "text": "At 10:16 UTC, the complete answer key scored the submission at exactly four correct answers out of five. Item 4 was the sole incorrect response. The item wording asked which claim was NOT supported, making it the quiz’s negation probe."}, {"speaker": "Answer-key audit", "text": "At 10:16:05 UTC on 8 May 2026, the answer key for QZ-417 recorded response “B” as the answer to the affirmative version of quiz item 4."}, {"speaker": "Records officer", "text": "The lesson record was present and showed completion of the affirmative-claims activity. A cross-check of the learner responses, answer key, scoring log, and lesson record found no mutually inconsistent information relevant to scoring or routing QZ-417."}, {"speaker": "Course instructor", "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}]}, "method": "c2d", "provenance": {"source_id": "diverse-062", "source_is_synthetic": true, "source_sha256": "571895f9ad2bd8e1cdf4993a3c848224bd6a8603b7160d6c25bc7527551b1fa8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "U1_boundary_mastery"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same rubric statement and preserve the learner, quiz item, score, timing, and routing scope. The two evidence spans are complete factual sentences. Changing only the affirmative-version answer from C to B is coherent with the unchanged 4/5 score and item 4 being the sole incorrect response, and it introduces no duplicate contradiction. Neither context contains an answer label, code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a8 is one consistency relation over an explicit set of records rather than a bundle of policy requirements. The focus a4 is factual. The base and counter assignments can both be realized while changing only whether the item-4 response matches the affirmative-version answer; the response can remain incorrect in either case, preserving the 4/5 score and all other atom states. The policy evidence is an accurate citation from the original state. No additional state-derived rule is needed because the complete governing rubric remains available verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. The rule also establishes that item 4 is the negation probe, the selected response matches its affirmative-version answer, and all required records are present and consistent. These conditions sufficiently entail U2_negation_review and exclude EVIDENCE_HOLD and the competing scored outcomes.", "rule_index": 0, "sound": true}, {"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. With item 4 established as the negation probe and the affirmative-version match refuted, the U2 exception does not apply and U1_boundary_mastery is entailed. Record presence and consistency exclude EVIDENCE_HOLD.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The learner answered exactly four of the five quiz items correctly."}, {"id": "a2", "statement": "The learner answered quiz item 4 incorrectly."}, {"id": "a3", "statement": "Quiz item 4 is the negation probe."}, {"id": "a4", "statement": "The learner’s selected response on quiz item 4 matches the answer to the affirmative version of quiz item 4."}, {"id": "a5", "statement": "The learner’s quiz responses are present."}, {"id": "a6", "statement": "The answer key is present."}, {"id": "a7", "statement": "The lesson record is present."}, {"id": "a8", "statement": "The learner’s quiz responses, answer key, and lesson record contain no mutually inconsistent information relevant to scoring or routing this quiz."}], "base_state_json": "[{\"speaker\":\"Records coordinator\",\"text\":\"The archived assessment packet contains all five learner responses, the complete answer key, and lesson record LR-2011. Reconciliation found no mutually inconsistent information relevant to scoring or routing.\"},{\"speaker\":\"Scoring auditor\",\"text\":\"The learner answered exactly four of the five items correctly; item 4 was the sole incorrect response.\"},{\"speaker\":\"Assessment editor\",\"text\":\"Quiz item 4 is the negation probe: it asks which claim is NOT supported.\"},{\"speaker\":\"Submission log\",\"text\":\"In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C.\"},{\"speaker\":\"Answer-key log\",\"text\":\"In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option C.\"},{\"speaker\":\"Course instructor\",\"text\":\"Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["3", "text"], "text": "In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C."}, {"path": ["4", "text"], "text": "In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option C."}], "policy_evidence": [{"path": ["3", "text"], "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}], "rules": [{"justification": "The records are present and consistent, the score is 4/5, item 4 is necessarily the sole miss, item 4 is the negation probe, and the selected response matches its affirmative-version answer.", "target": "U2_negation_review", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The records are present and consistent, the score is 4/5, and item 4 is necessarily the sole miss; although it is the negation probe, the selected response does not match its affirmative-version answer.", "target": "U1_boundary_mastery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C.", "negative_left": "In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C.", "negative_right": "In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option B.", "right": "In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option C."}, "verifier_independent_model": false}, "family": "scale-diverse-062-002", "id": "scale-diverse-062-002-base", "input": {"questions": {"decision": {"criteria": {"EVIDENCE_HOLD": "No urgency assigned: Responses, answer key, or lesson record are missing or inconsistent, so scoring or routing cannot yet be completed.", "U0_full_mastery": "Urgency 0: Score is 5/5. Verify mastery; assign no review or support role.", "U1_boundary_mastery": "Urgency 1: Score is 4/5, but the sole miss either is not the negation probe or does not match the affirmative-version answer. Verify mastery; assign an optional independent recap.", "U2_negation_review": "Urgency 2: Score is 4/5, the only miss is the negation probe, and the chosen response matches the affirmative-version answer. Do not verify mastery; route to the course instructor’s NOT/EXCEPT contrast sort before the next quiz.", "U3_tutor_reteach": "Urgency 3: Score is exactly 3/5, regardless of error type. Do not verify mastery; route to a learning support tutor for a full guided practice set before the next quiz.", "U4_intensive_support": "Urgency 4: Score is 0–2/5. Do not verify mastery; route first to the course instructor for reteaching and then to the learning support tutor for supervised practice."}, "instructions": "Select the single rubric option that correctly verifies mastery status, assigns the prescribed route, and gives the ordered intervention urgency. Lower urgency numbers mean less urgent intervention.", "type": "choice"}}, "state": [{"speaker": "Records coordinator", "text": "The archived assessment packet contains all five learner responses, the complete answer key, and lesson record LR-2011. Reconciliation found no mutually inconsistent information relevant to scoring or routing."}, {"speaker": "Scoring auditor", "text": "The learner answered exactly four of the five items correctly; item 4 was the sole incorrect response."}, {"speaker": "Assessment editor", "text": "Quiz item 4 is the negation probe: it asks which claim is NOT supported."}, {"speaker": "Submission log", "text": "In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C."}, {"speaker": "Answer-key log", "text": "In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option C."}, {"speaker": "Course instructor", "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}]}, "method": "c2d", "provenance": {"source_id": "diverse-062", "source_is_synthetic": true, "source_sha256": "571895f9ad2bd8e1cdf4993a3c848224bd6a8603b7160d6c25bc7527551b1fa8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "U2_negation_review"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same rubric statement and preserve the learner, quiz item, score, timing, and routing scope. The two evidence spans are complete factual sentences. Changing only the affirmative-version answer from C to B is coherent with the unchanged 4/5 score and item 4 being the sole incorrect response, and it introduces no duplicate contradiction. Neither context contains an answer label, code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a8 is one consistency relation over an explicit set of records rather than a bundle of policy requirements. The focus a4 is factual. The base and counter assignments can both be realized while changing only whether the item-4 response matches the affirmative-version answer; the response can remain incorrect in either case, preserving the 4/5 score and all other atom states. The policy evidence is an accurate citation from the original state. No additional state-derived rule is needed because the complete governing rubric remains available verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. The rule also establishes that item 4 is the negation probe, the selected response matches its affirmative-version answer, and all required records are present and consistent. These conditions sufficiently entail U2_negation_review and exclude EVIDENCE_HOLD and the competing scored outcomes.", "rule_index": 0, "sound": true}, {"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. With item 4 established as the negation probe and the affirmative-version match refuted, the U2 exception does not apply and U1_boundary_mastery is entailed. Record presence and consistency exclude EVIDENCE_HOLD.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The learner answered exactly four of the five quiz items correctly."}, {"id": "a2", "statement": "The learner answered quiz item 4 incorrectly."}, {"id": "a3", "statement": "Quiz item 4 is the negation probe."}, {"id": "a4", "statement": "The learner’s selected response on quiz item 4 matches the answer to the affirmative version of quiz item 4."}, {"id": "a5", "statement": "The learner’s quiz responses are present."}, {"id": "a6", "statement": "The answer key is present."}, {"id": "a7", "statement": "The lesson record is present."}, {"id": "a8", "statement": "The learner’s quiz responses, answer key, and lesson record contain no mutually inconsistent information relevant to scoring or routing this quiz."}], "base_state_json": "[{\"speaker\":\"Records coordinator\",\"text\":\"The archived assessment packet contains all five learner responses, the complete answer key, and lesson record LR-2011. Reconciliation found no mutually inconsistent information relevant to scoring or routing.\"},{\"speaker\":\"Scoring auditor\",\"text\":\"The learner answered exactly four of the five items correctly; item 4 was the sole incorrect response.\"},{\"speaker\":\"Assessment editor\",\"text\":\"Quiz item 4 is the negation probe: it asks which claim is NOT supported.\"},{\"speaker\":\"Submission log\",\"text\":\"In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C.\"},{\"speaker\":\"Answer-key log\",\"text\":\"In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option C.\"},{\"speaker\":\"Course instructor\",\"text\":\"Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["3", "text"], "text": "In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C."}, {"path": ["4", "text"], "text": "In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option C."}], "policy_evidence": [{"path": ["3", "text"], "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}], "rules": [{"justification": "The records are present and consistent, the score is 4/5, item 4 is necessarily the sole miss, item 4 is the negation probe, and the selected response matches its affirmative-version answer.", "target": "U2_negation_review", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The records are present and consistent, the score is 4/5, and item 4 is necessarily the sole miss; although it is the negation probe, the selected response does not match its affirmative-version answer.", "target": "U1_boundary_mastery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C.", "negative_left": "In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C.", "negative_right": "In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option B.", "right": "In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option C."}, "verifier_independent_model": false}, "family": "scale-diverse-062-002", "id": "scale-diverse-062-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"EVIDENCE_HOLD": "No urgency assigned: Responses, answer key, or lesson record are missing or inconsistent, so scoring or routing cannot yet be completed.", "U0_full_mastery": "Urgency 0: Score is 5/5. Verify mastery; assign no review or support role.", "U1_boundary_mastery": "Urgency 1: Score is 4/5, but the sole miss either is not the negation probe or does not match the affirmative-version answer. Verify mastery; assign an optional independent recap.", "U2_negation_review": "Urgency 2: Score is 4/5, the only miss is the negation probe, and the chosen response matches the affirmative-version answer. Do not verify mastery; route to the course instructor’s NOT/EXCEPT contrast sort before the next quiz.", "U3_tutor_reteach": "Urgency 3: Score is exactly 3/5, regardless of error type. Do not verify mastery; route to a learning support tutor for a full guided practice set before the next quiz.", "U4_intensive_support": "Urgency 4: Score is 0–2/5. Do not verify mastery; route first to the course instructor for reteaching and then to the learning support tutor for supervised practice."}, "instructions": "Select the single rubric option that correctly verifies mastery status, assigns the prescribed route, and gives the ordered intervention urgency. Lower urgency numbers mean less urgent intervention.", "type": "choice"}}, "state": [{"speaker": "Records coordinator", "text": "The archived assessment packet contains all five learner responses, the complete answer key, and lesson record LR-2011. Reconciliation found no mutually inconsistent information relevant to scoring or routing."}, {"speaker": "Scoring auditor", "text": "The learner answered exactly four of the five items correctly; item 4 was the sole incorrect response."}, {"speaker": "Assessment editor", "text": "Quiz item 4 is the negation probe: it asks which claim is NOT supported."}, {"speaker": "Submission log", "text": "In the learner’s quiz submission Q-2011, recorded at 14:06 UTC on 17 September 2026, the selected response for quiz item 4 is option C."}, {"speaker": "Answer-key log", "text": "In the answer key for quiz Q-2011, issued at 09:00 UTC on 17 September 2026, the answer to the affirmative version of quiz item 4 is option B."}, {"speaker": "Course instructor", "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}]}, "method": "c2d", "provenance": {"source_id": "diverse-062", "source_is_synthetic": true, "source_sha256": "571895f9ad2bd8e1cdf4993a3c848224bd6a8603b7160d6c25bc7527551b1fa8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "U1_boundary_mastery"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete decision rubric, and both contexts retain the original-state rubric statement without altering or adding policy. The same quiz, learner response, item, score, and review scope are maintained; only the affirmative-version answer changes from H to J in one complete factual sentence. Both evidence spans are complete factual sentences. In the counterfactual, selecting H remains coherent with item 4 being the sole incorrect response while J is the affirmative-version answer, and no duplicate assertion contradicts that change. Neither context embeds a decision label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a8 is one consistency relation over an explicit set of records rather than a bundle of policy requirements. The focus a4 is factual. The base and counter assignments can both be realized while changing only whether the item-4 response matches the affirmative-version answer; the response can remain incorrect in either case, preserving the 4/5 score and all other atom states. The policy evidence is an accurate citation from the original state. No additional state-derived rule is needed because the complete governing rubric remains available verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. The rule also establishes that item 4 is the negation probe, the selected response matches its affirmative-version answer, and all required records are present and consistent. These conditions sufficiently entail U2_negation_review and exclude EVIDENCE_HOLD and the competing scored outcomes.", "rule_index": 0, "sound": true}, {"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. With item 4 established as the negation probe and the affirmative-version match refuted, the U2 exception does not apply and U1_boundary_mastery is entailed. Record presence and consistency exclude EVIDENCE_HOLD.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The learner answered exactly four of the five quiz items correctly."}, {"id": "a2", "statement": "The learner answered quiz item 4 incorrectly."}, {"id": "a3", "statement": "Quiz item 4 is the negation probe."}, {"id": "a4", "statement": "The learner’s selected response on quiz item 4 matches the answer to the affirmative version of quiz item 4."}, {"id": "a5", "statement": "The learner’s quiz responses are present."}, {"id": "a6", "statement": "The answer key is present."}, {"id": "a7", "statement": "The lesson record is present."}, {"id": "a8", "statement": "The learner’s quiz responses, answer key, and lesson record contain no mutually inconsistent information relevant to scoring or routing this quiz."}], "base_state_json": "[{\"speaker\":\"Operations coordinator\",\"text\":\"Operational handoff for quiz QZ-681: the scoring packet has been locked for routing review.\"},{\"speaker\":\"Records custodian\",\"text\":\"The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H.\"},{\"speaker\":\"Assessment lead\",\"text\":\"The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option H as the answer to the affirmative version of quiz item 4.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The packet contains the learner’s complete quiz responses, the signed answer key, and the lesson record. Reconciliation confirms that exactly four of the five items were answered correctly and that item 4 was answered incorrectly. Item 4 is designated as the negation probe. The three records contain no mutually inconsistent information relevant to scoring or routing this quiz.\"},{\"speaker\":\"Course instructor\",\"text\":\"Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H."}, {"path": ["2", "text"], "text": "The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option H as the answer to the affirmative version of quiz item 4."}], "policy_evidence": [{"path": ["3", "text"], "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}], "rules": [{"justification": "The records are present and consistent, the score is 4/5, item 4 is necessarily the sole miss, item 4 is the negation probe, and the selected response matches its affirmative-version answer.", "target": "U2_negation_review", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The records are present and consistent, the score is 4/5, and item 4 is necessarily the sole miss; although it is the negation probe, the selected response does not match its affirmative-version answer.", "target": "U1_boundary_mastery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H.", "negative_left": "The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H.", "negative_right": "The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option J as the answer to the affirmative version of quiz item 4.", "right": "The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option H as the answer to the affirmative version of quiz item 4."}, "verifier_independent_model": false}, "family": "scale-diverse-062-003", "id": "scale-diverse-062-003-base", "input": {"questions": {"decision": {"criteria": {"EVIDENCE_HOLD": "No urgency assigned: Responses, answer key, or lesson record are missing or inconsistent, so scoring or routing cannot yet be completed.", "U0_full_mastery": "Urgency 0: Score is 5/5. Verify mastery; assign no review or support role.", "U1_boundary_mastery": "Urgency 1: Score is 4/5, but the sole miss either is not the negation probe or does not match the affirmative-version answer. Verify mastery; assign an optional independent recap.", "U2_negation_review": "Urgency 2: Score is 4/5, the only miss is the negation probe, and the chosen response matches the affirmative-version answer. Do not verify mastery; route to the course instructor’s NOT/EXCEPT contrast sort before the next quiz.", "U3_tutor_reteach": "Urgency 3: Score is exactly 3/5, regardless of error type. Do not verify mastery; route to a learning support tutor for a full guided practice set before the next quiz.", "U4_intensive_support": "Urgency 4: Score is 0–2/5. Do not verify mastery; route first to the course instructor for reteaching and then to the learning support tutor for supervised practice."}, "instructions": "Select the single rubric option that correctly verifies mastery status, assigns the prescribed route, and gives the ordered intervention urgency. Lower urgency numbers mean less urgent intervention.", "type": "choice"}}, "state": [{"speaker": "Operations coordinator", "text": "Operational handoff for quiz QZ-681: the scoring packet has been locked for routing review."}, {"speaker": "Records custodian", "text": "The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H."}, {"speaker": "Assessment lead", "text": "The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option H as the answer to the affirmative version of quiz item 4."}, {"speaker": "Quality reviewer", "text": "The packet contains the learner’s complete quiz responses, the signed answer key, and the lesson record. Reconciliation confirms that exactly four of the five items were answered correctly and that item 4 was answered incorrectly. Item 4 is designated as the negation probe. The three records contain no mutually inconsistent information relevant to scoring or routing this quiz."}, {"speaker": "Course instructor", "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}]}, "method": "c2d", "provenance": {"source_id": "diverse-062", "source_is_synthetic": true, "source_sha256": "571895f9ad2bd8e1cdf4993a3c848224bd6a8603b7160d6c25bc7527551b1fa8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "U2_negation_review"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete decision rubric, and both contexts retain the original-state rubric statement without altering or adding policy. The same quiz, learner response, item, score, and review scope are maintained; only the affirmative-version answer changes from H to J in one complete factual sentence. Both evidence spans are complete factual sentences. In the counterfactual, selecting H remains coherent with item 4 being the sole incorrect response while J is the affirmative-version answer, and no duplicate assertion contradicts that change. Neither context embeds a decision label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a8 is one consistency relation over an explicit set of records rather than a bundle of policy requirements. The focus a4 is factual. The base and counter assignments can both be realized while changing only whether the item-4 response matches the affirmative-version answer; the response can remain incorrect in either case, preserving the 4/5 score and all other atom states. The policy evidence is an accurate citation from the original state. No additional state-derived rule is needed because the complete governing rubric remains available verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. The rule also establishes that item 4 is the negation probe, the selected response matches its affirmative-version answer, and all required records are present and consistent. These conditions sufficiently entail U2_negation_review and exclude EVIDENCE_HOLD and the competing scored outcomes.", "rule_index": 0, "sound": true}, {"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. With item 4 established as the negation probe and the affirmative-version match refuted, the U2 exception does not apply and U1_boundary_mastery is entailed. Record presence and consistency exclude EVIDENCE_HOLD.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The learner answered exactly four of the five quiz items correctly."}, {"id": "a2", "statement": "The learner answered quiz item 4 incorrectly."}, {"id": "a3", "statement": "Quiz item 4 is the negation probe."}, {"id": "a4", "statement": "The learner’s selected response on quiz item 4 matches the answer to the affirmative version of quiz item 4."}, {"id": "a5", "statement": "The learner’s quiz responses are present."}, {"id": "a6", "statement": "The answer key is present."}, {"id": "a7", "statement": "The lesson record is present."}, {"id": "a8", "statement": "The learner’s quiz responses, answer key, and lesson record contain no mutually inconsistent information relevant to scoring or routing this quiz."}], "base_state_json": "[{\"speaker\":\"Operations coordinator\",\"text\":\"Operational handoff for quiz QZ-681: the scoring packet has been locked for routing review.\"},{\"speaker\":\"Records custodian\",\"text\":\"The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H.\"},{\"speaker\":\"Assessment lead\",\"text\":\"The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option H as the answer to the affirmative version of quiz item 4.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The packet contains the learner’s complete quiz responses, the signed answer key, and the lesson record. Reconciliation confirms that exactly four of the five items were answered correctly and that item 4 was answered incorrectly. Item 4 is designated as the negation probe. The three records contain no mutually inconsistent information relevant to scoring or routing this quiz.\"},{\"speaker\":\"Course instructor\",\"text\":\"Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H."}, {"path": ["2", "text"], "text": "The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option H as the answer to the affirmative version of quiz item 4."}], "policy_evidence": [{"path": ["3", "text"], "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}], "rules": [{"justification": "The records are present and consistent, the score is 4/5, item 4 is necessarily the sole miss, item 4 is the negation probe, and the selected response matches its affirmative-version answer.", "target": "U2_negation_review", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The records are present and consistent, the score is 4/5, and item 4 is necessarily the sole miss; although it is the negation probe, the selected response does not match its affirmative-version answer.", "target": "U1_boundary_mastery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H.", "negative_left": "The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H.", "negative_right": "The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option J as the answer to the affirmative version of quiz item 4.", "right": "The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option H as the answer to the affirmative version of quiz item 4."}, "verifier_independent_model": false}, "family": "scale-diverse-062-003", "id": "scale-diverse-062-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"EVIDENCE_HOLD": "No urgency assigned: Responses, answer key, or lesson record are missing or inconsistent, so scoring or routing cannot yet be completed.", "U0_full_mastery": "Urgency 0: Score is 5/5. Verify mastery; assign no review or support role.", "U1_boundary_mastery": "Urgency 1: Score is 4/5, but the sole miss either is not the negation probe or does not match the affirmative-version answer. Verify mastery; assign an optional independent recap.", "U2_negation_review": "Urgency 2: Score is 4/5, the only miss is the negation probe, and the chosen response matches the affirmative-version answer. Do not verify mastery; route to the course instructor’s NOT/EXCEPT contrast sort before the next quiz.", "U3_tutor_reteach": "Urgency 3: Score is exactly 3/5, regardless of error type. Do not verify mastery; route to a learning support tutor for a full guided practice set before the next quiz.", "U4_intensive_support": "Urgency 4: Score is 0–2/5. Do not verify mastery; route first to the course instructor for reteaching and then to the learning support tutor for supervised practice."}, "instructions": "Select the single rubric option that correctly verifies mastery status, assigns the prescribed route, and gives the ordered intervention urgency. Lower urgency numbers mean less urgent intervention.", "type": "choice"}}, "state": [{"speaker": "Operations coordinator", "text": "Operational handoff for quiz QZ-681: the scoring packet has been locked for routing review."}, {"speaker": "Records custodian", "text": "The locked response export for quiz QZ-681, timestamped 2026-04-18 at 14:32 UTC, records the learner’s selected response on quiz item 4 as option H."}, {"speaker": "Assessment lead", "text": "The signed answer key for quiz QZ-681, timestamped 2026-04-18 at 14:35 UTC, records option J as the answer to the affirmative version of quiz item 4."}, {"speaker": "Quality reviewer", "text": "The packet contains the learner’s complete quiz responses, the signed answer key, and the lesson record. Reconciliation confirms that exactly four of the five items were answered correctly and that item 4 was answered incorrectly. Item 4 is designated as the negation probe. The three records contain no mutually inconsistent information relevant to scoring or routing this quiz."}, {"speaker": "Course instructor", "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}]}, "method": "c2d", "provenance": {"source_id": "diverse-062", "source_is_synthetic": true, "source_sha256": "571895f9ad2bd8e1cdf4993a3c848224bd6a8603b7160d6c25bc7527551b1fa8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "U1_boundary_mastery"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same mastery, routing, and urgency policies and preserve Mina, the five-item fraction quiz, the proposed triage, and the relevant date/time path. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes item 1’s recorded choice from left to right; this yields two correct diagnostic responses while remaining consistent with exactly two correct answers overall, since no unchanged sentence specifies additional correct responses. Neither context states the decision’s gold answer, an answer code, proposition ID, rule table, label rationale, or an output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "full_context_fact_states": {"base": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "counterfactual": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "remove_left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "remove_right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "negative_pair": {"pattern_on_at_least_2_diagnostic_items": "refuted"}, "negative_sentence": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "positive_pair": {"pattern_on_at_least_2_diagnostic_items": "supported"}, "right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express single factual relationships, and the focus is the factual presence or absence of a response pattern rather than a policy conclusion. Both focus assignments are realizable while holding the 2/5 score fixed. The policy evidence preserves all substantive state-originating rules needed to apply the unchanged question; omitted details such as guided-practice performance are not needed for these rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly 2/5 correct necessarily fails the stated 4/5 mastery threshold. With the misconception pattern present on at least two diagnostic items, Mina’s condition is met, requiring the tutor’s Fraction-Strip Comparison Lab at urgency 3; therefore the proposed instructor conference and urgency 2 triage is incorrect.", "rule_index": 0, "sound": true}, {"reason": "Exactly 2/5 correct establishes nonmastery. Refutation of the at-least-two pattern proposition entails that Mina’s condition is not met, so the otherwise branch requires the instructor correction conference at urgency 2, making every component of the proposal correct.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "quiz_score_2_of_5", "statement": "Mina answered exactly 2 of the 5 items on this five-item fraction quiz correctly."}, {"id": "pattern_on_at_least_2_diagnostic_items", "statement": "Mina’s responses on this five-item fraction quiz exhibit the “larger denominator means larger fraction” pattern on at least 2 of diagnostic items 1, 2, and 4."}], "base_state_json": "\"The pending decision is whether to mark Mina not mastered, route her to the course instructor’s correction conference, and assign urgency 2. At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the left fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4. At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest).\"", "base_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}], "counter_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}], "focus_atom": "pattern_on_at_least_2_diagnostic_items", "focus_evidence": [{"path": [], "text": "At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the left fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4."}, {"path": [], "text": "At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4."}], "policy_evidence": [{"path": [], "text": "Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”"}, {"path": [], "text": "Policy defines mastery as 4/5 plus 2/3 diagnostic items correct."}, {"path": [], "text": "Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2."}, {"path": [], "text": "Urgency runs from 1 (lowest) to 3 (highest)."}], "rules": [{"justification": "A score of 2/5 establishes nonmastery. Exhibiting the stated pattern on at least two diagnostic items meets Mina’s condition, so policy requires the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3 rather than the proposed instructor conference at urgency 2.", "target": "false", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}]}, {"justification": "A score of 2/5 establishes nonmastery. Explicit refutation of exhibiting the pattern on at least two diagnostic items establishes that Mina’s condition is not met, so policy requires the instructor correction conference at urgency 2, matching the proposal.", "target": "true", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the left fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4.", "negative_left": "At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the right fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4.", "negative_right": "At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4.", "right": "At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4."}, "verifier_independent_model": false}, "family": "scale-diverse-063-001", "id": "scale-diverse-063-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — Mina is not mastered, but she must be routed to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3.", "true": "Yes — the proposed not-mastered classification, instructor conference, and urgency 2 assignment are all correct."}, "instructions": "Decide whether this proposed triage is correct: mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Answer Yes or No according to the stated policy.", "type": "noul"}}, "state": "The pending decision is whether to mark Mina not mastered, route her to the course instructor’s correction conference, and assign urgency 2. At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the left fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4. At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest)."}, "method": "c2d", "provenance": {"source_id": "diverse-063", "source_is_synthetic": true, "source_sha256": "1ab2940f7e11265b1fc7b64fecf6943e6d0f6fc9c6402c436749bbc8483bd245", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same mastery, routing, and urgency policies and preserve Mina, the five-item fraction quiz, the proposed triage, and the relevant date/time path. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes item 1’s recorded choice from left to right; this yields two correct diagnostic responses while remaining consistent with exactly two correct answers overall, since no unchanged sentence specifies additional correct responses. Neither context states the decision’s gold answer, an answer code, proposition ID, rule table, label rationale, or an output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "full_context_fact_states": {"base": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "counterfactual": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "remove_left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "remove_right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "negative_pair": {"pattern_on_at_least_2_diagnostic_items": "refuted"}, "negative_sentence": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "positive_pair": {"pattern_on_at_least_2_diagnostic_items": "supported"}, "right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express single factual relationships, and the focus is the factual presence or absence of a response pattern rather than a policy conclusion. Both focus assignments are realizable while holding the 2/5 score fixed. The policy evidence preserves all substantive state-originating rules needed to apply the unchanged question; omitted details such as guided-practice performance are not needed for these rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly 2/5 correct necessarily fails the stated 4/5 mastery threshold. With the misconception pattern present on at least two diagnostic items, Mina’s condition is met, requiring the tutor’s Fraction-Strip Comparison Lab at urgency 3; therefore the proposed instructor conference and urgency 2 triage is incorrect.", "rule_index": 0, "sound": true}, {"reason": "Exactly 2/5 correct establishes nonmastery. Refutation of the at-least-two pattern proposition entails that Mina’s condition is not met, so the otherwise branch requires the instructor correction conference at urgency 2, making every component of the proposal correct.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "quiz_score_2_of_5", "statement": "Mina answered exactly 2 of the 5 items on this five-item fraction quiz correctly."}, {"id": "pattern_on_at_least_2_diagnostic_items", "statement": "Mina’s responses on this five-item fraction quiz exhibit the “larger denominator means larger fraction” pattern on at least 2 of diagnostic items 1, 2, and 4."}], "base_state_json": "\"The pending decision is whether to mark Mina not mastered, route her to the course instructor’s correction conference, and assign urgency 2. At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the left fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4. At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest).\"", "base_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}], "counter_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}], "focus_atom": "pattern_on_at_least_2_diagnostic_items", "focus_evidence": [{"path": [], "text": "At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the left fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4."}, {"path": [], "text": "At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4."}], "policy_evidence": [{"path": [], "text": "Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”"}, {"path": [], "text": "Policy defines mastery as 4/5 plus 2/3 diagnostic items correct."}, {"path": [], "text": "Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2."}, {"path": [], "text": "Urgency runs from 1 (lowest) to 3 (highest)."}], "rules": [{"justification": "A score of 2/5 establishes nonmastery. Exhibiting the stated pattern on at least two diagnostic items meets Mina’s condition, so policy requires the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3 rather than the proposed instructor conference at urgency 2.", "target": "false", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}]}, {"justification": "A score of 2/5 establishes nonmastery. Explicit refutation of exhibiting the pattern on at least two diagnostic items establishes that Mina’s condition is not met, so policy requires the instructor correction conference at urgency 2, matching the proposal.", "target": "true", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the left fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4.", "negative_left": "At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the right fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4.", "negative_right": "At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4.", "right": "At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4."}, "verifier_independent_model": false}, "family": "scale-diverse-063-001", "id": "scale-diverse-063-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — Mina is not mastered, but she must be routed to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3.", "true": "Yes — the proposed not-mastered classification, instructor conference, and urgency 2 assignment are all correct."}, "instructions": "Decide whether this proposed triage is correct: mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Answer Yes or No according to the stated policy.", "type": "noul"}}, "state": "The pending decision is whether to mark Mina not mastered, route her to the course instructor’s correction conference, and assign urgency 2. At 09:20 on 14 March 2026, the audit record for Mina’s five-item fraction quiz showed exactly two correct answers and recorded that she had circled the right fraction on diagnostic item 1, the right fraction on diagnostic item 2, and the left fraction on diagnostic item 4. At 08:55 on 14 March 2026, Mina’s five-item fraction quiz instructed her to circle the larger fraction and displayed 3/8 on the left and 3/5 on the right in item 1, 4/9 on the left and 4/7 on the right in item 2, and 2/11 on the left and 2/6 on the right in item 4. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest)."}, "method": "c2d", "provenance": {"source_id": "diverse-063", "source_is_synthetic": true, "source_sha256": "1ab2940f7e11265b1fc7b64fecf6943e6d0f6fc9c6402c436749bbc8483bd245", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same mastery, routing, and urgency policies and preserve Mina, the five-item fraction quiz, the diagnostic-item path, and the proposed triage under review. The two evidence spans are complete factual sentences. The counterfactual changes only Mina’s recorded diagnostic selections from B/A/F to C/D/E; this is coherent with the unchanged 2-of-5 total because no item-answer mapping creates a contradictory duplicate measurement. Neither context states a gold Yes/No answer, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "full_context_fact_states": {"base": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "counterfactual": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "remove_left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "remove_right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "negative_pair": {"pattern_on_at_least_2_diagnostic_items": "refuted"}, "negative_sentence": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "positive_pair": {"pattern_on_at_least_2_diagnostic_items": "supported"}, "right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express single factual relationships, and the focus is the factual presence or absence of a response pattern rather than a policy conclusion. Both focus assignments are realizable while holding the 2/5 score fixed. The policy evidence preserves all substantive state-originating rules needed to apply the unchanged question; omitted details such as guided-practice performance are not needed for these rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly 2/5 correct necessarily fails the stated 4/5 mastery threshold. With the misconception pattern present on at least two diagnostic items, Mina’s condition is met, requiring the tutor’s Fraction-Strip Comparison Lab at urgency 3; therefore the proposed instructor conference and urgency 2 triage is incorrect.", "rule_index": 0, "sound": true}, {"reason": "Exactly 2/5 correct establishes nonmastery. Refutation of the at-least-two pattern proposition entails that Mina’s condition is not met, so the otherwise branch requires the instructor correction conference at urgency 2, making every component of the proposal correct.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "quiz_score_2_of_5", "statement": "Mina answered exactly 2 of the 5 items on this five-item fraction quiz correctly."}, {"id": "pattern_on_at_least_2_diagnostic_items", "statement": "Mina’s responses on this five-item fraction quiz exhibit the “larger denominator means larger fraction” pattern on at least 2 of diagnostic items 1, 2, and 4."}], "base_state_json": "\"An operational reviewer is checking the recorded disposition for Mina’s five-item fraction quiz Q-3008. The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as B, A, and F, respectively. On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively. The reviewer must assess a proposal to mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest).\"", "base_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}], "counter_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}], "focus_atom": "pattern_on_at_least_2_diagnostic_items", "focus_evidence": [{"path": [], "text": "The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as B, A, and F, respectively."}, {"path": [], "text": "On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively."}], "policy_evidence": [{"path": [], "text": "Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”"}, {"path": [], "text": "Policy defines mastery as 4/5 plus 2/3 diagnostic items correct."}, {"path": [], "text": "Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2."}, {"path": [], "text": "Urgency runs from 1 (lowest) to 3 (highest)."}], "rules": [{"justification": "A score of 2/5 establishes nonmastery. Exhibiting the stated pattern on at least two diagnostic items meets Mina’s condition, so policy requires the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3 rather than the proposed instructor conference at urgency 2.", "target": "false", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}]}, {"justification": "A score of 2/5 establishes nonmastery. Explicit refutation of exhibiting the pattern on at least two diagnostic items establishes that Mina’s condition is not met, so policy requires the instructor correction conference at urgency 2, matching the proposal.", "target": "true", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}]}]}, "verified_pair": {"left": "The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as B, A, and F, respectively.", "negative_left": "The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as C, D, and E, respectively.", "negative_right": "On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively.", "right": "On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively."}, "verifier_independent_model": false}, "family": "scale-diverse-063-003", "id": "scale-diverse-063-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — Mina is not mastered, but she must be routed to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3.", "true": "Yes — the proposed not-mastered classification, instructor conference, and urgency 2 assignment are all correct."}, "instructions": "Decide whether this proposed triage is correct: mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Answer Yes or No according to the stated policy.", "type": "noul"}}, "state": "An operational reviewer is checking the recorded disposition for Mina’s five-item fraction quiz Q-3008. The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as B, A, and F, respectively. On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively. The reviewer must assess a proposal to mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest)."}, "method": "c2d", "provenance": {"source_id": "diverse-063", "source_is_synthetic": true, "source_sha256": "1ab2940f7e11265b1fc7b64fecf6943e6d0f6fc9c6402c436749bbc8483bd245", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same mastery, routing, and urgency policies and preserve Mina, the five-item fraction quiz, the diagnostic-item path, and the proposed triage under review. The two evidence spans are complete factual sentences. The counterfactual changes only Mina’s recorded diagnostic selections from B/A/F to C/D/E; this is coherent with the unchanged 2-of-5 total because no item-answer mapping creates a contradictory duplicate measurement. Neither context states a gold Yes/No answer, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "full_context_fact_states": {"base": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "counterfactual": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "remove_left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "remove_right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "negative_pair": {"pattern_on_at_least_2_diagnostic_items": "refuted"}, "negative_sentence": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "positive_pair": {"pattern_on_at_least_2_diagnostic_items": "supported"}, "right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express single factual relationships, and the focus is the factual presence or absence of a response pattern rather than a policy conclusion. Both focus assignments are realizable while holding the 2/5 score fixed. The policy evidence preserves all substantive state-originating rules needed to apply the unchanged question; omitted details such as guided-practice performance are not needed for these rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly 2/5 correct necessarily fails the stated 4/5 mastery threshold. With the misconception pattern present on at least two diagnostic items, Mina’s condition is met, requiring the tutor’s Fraction-Strip Comparison Lab at urgency 3; therefore the proposed instructor conference and urgency 2 triage is incorrect.", "rule_index": 0, "sound": true}, {"reason": "Exactly 2/5 correct establishes nonmastery. Refutation of the at-least-two pattern proposition entails that Mina’s condition is not met, so the otherwise branch requires the instructor correction conference at urgency 2, making every component of the proposal correct.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "quiz_score_2_of_5", "statement": "Mina answered exactly 2 of the 5 items on this five-item fraction quiz correctly."}, {"id": "pattern_on_at_least_2_diagnostic_items", "statement": "Mina’s responses on this five-item fraction quiz exhibit the “larger denominator means larger fraction” pattern on at least 2 of diagnostic items 1, 2, and 4."}], "base_state_json": "\"An operational reviewer is checking the recorded disposition for Mina’s five-item fraction quiz Q-3008. The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as B, A, and F, respectively. On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively. The reviewer must assess a proposal to mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest).\"", "base_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}], "counter_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}], "focus_atom": "pattern_on_at_least_2_diagnostic_items", "focus_evidence": [{"path": [], "text": "The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as B, A, and F, respectively."}, {"path": [], "text": "On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively."}], "policy_evidence": [{"path": [], "text": "Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”"}, {"path": [], "text": "Policy defines mastery as 4/5 plus 2/3 diagnostic items correct."}, {"path": [], "text": "Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2."}, {"path": [], "text": "Urgency runs from 1 (lowest) to 3 (highest)."}], "rules": [{"justification": "A score of 2/5 establishes nonmastery. Exhibiting the stated pattern on at least two diagnostic items meets Mina’s condition, so policy requires the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3 rather than the proposed instructor conference at urgency 2.", "target": "false", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}]}, {"justification": "A score of 2/5 establishes nonmastery. Explicit refutation of exhibiting the pattern on at least two diagnostic items establishes that Mina’s condition is not met, so policy requires the instructor correction conference at urgency 2, matching the proposal.", "target": "true", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}]}]}, "verified_pair": {"left": "The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as B, A, and F, respectively.", "negative_left": "The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as C, D, and E, respectively.", "negative_right": "On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively.", "right": "On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively."}, "verifier_independent_model": false}, "family": "scale-diverse-063-003", "id": "scale-diverse-063-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — Mina is not mastered, but she must be routed to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3.", "true": "Yes — the proposed not-mastered classification, instructor conference, and urgency 2 assignment are all correct."}, "instructions": "Decide whether this proposed triage is correct: mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Answer Yes or No according to the stated policy.", "type": "noul"}}, "state": "An operational reviewer is checking the recorded disposition for Mina’s five-item fraction quiz Q-3008. The 14 September 2026 handoff record for Mina’s five-item fraction quiz Q-3008 records exactly 2 of 5 answers as correct and lists her selections on diagnostic items 1, 2, and 4 as C, D, and E, respectively. On five-item fraction quiz Q-3008, the sole option expressing “larger denominator means larger fraction” on diagnostic items 1, 2, and 4 is B, D, and F, respectively. The reviewer must assess a proposal to mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest)."}, "method": "c2d", "provenance": {"source_id": "diverse-063", "source_is_synthetic": true, "source_sha256": "1ab2940f7e11265b1fc7b64fecf6943e6d0f6fc9c6402c436749bbc8483bd245", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing mastery, routing, and urgency policies from the original state, while the unchanged questions preserve the request and scoring criteria. Mina, the five-item fraction quiz, the proposed triage path, and relevant item bindings remain unchanged. The two evidence spans are complete factual sentences. The counterfactual coherently changes only item 4 from exhibiting the pattern to not exhibiting it, without conflicting with the item 1 or item 2 findings or the two-correct total. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "full_context_fact_states": {"base": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "counterfactual": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "remove_left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "remove_right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "negative_pair": {"pattern_on_at_least_2_diagnostic_items": "refuted"}, "negative_sentence": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "positive_pair": {"pattern_on_at_least_2_diagnostic_items": "supported"}, "right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express single factual relationships, and the focus is the factual presence or absence of a response pattern rather than a policy conclusion. Both focus assignments are realizable while holding the 2/5 score fixed. The policy evidence preserves all substantive state-originating rules needed to apply the unchanged question; omitted details such as guided-practice performance are not needed for these rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly 2/5 correct necessarily fails the stated 4/5 mastery threshold. With the misconception pattern present on at least two diagnostic items, Mina’s condition is met, requiring the tutor’s Fraction-Strip Comparison Lab at urgency 3; therefore the proposed instructor conference and urgency 2 triage is incorrect.", "rule_index": 0, "sound": true}, {"reason": "Exactly 2/5 correct establishes nonmastery. Refutation of the at-least-two pattern proposition entails that Mina’s condition is not met, so the otherwise branch requires the instructor correction conference at urgency 2, making every component of the proposal correct.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "quiz_score_2_of_5", "statement": "Mina answered exactly 2 of the 5 items on this five-item fraction quiz correctly."}, {"id": "pattern_on_at_least_2_diagnostic_items", "statement": "Mina’s responses on this five-item fraction quiz exhibit the “larger denominator means larger fraction” pattern on at least 2 of diagnostic items 1, 2, and 4."}], "base_state_json": "\"Field note: A separate answer-key reconciliation established that Mina answered exactly 2 of the 5 items on her five-item fraction quiz correctly. At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2. At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 4. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest).\"", "base_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}], "counter_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}], "focus_atom": "pattern_on_at_least_2_diagnostic_items", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2."}, {"path": [], "text": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 4."}], "policy_evidence": [{"path": [], "text": "Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”"}, {"path": [], "text": "Policy defines mastery as 4/5 plus 2/3 diagnostic items correct."}, {"path": [], "text": "Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2."}, {"path": [], "text": "Urgency runs from 1 (lowest) to 3 (highest)."}], "rules": [{"justification": "A score of 2/5 establishes nonmastery. Exhibiting the stated pattern on at least two diagnostic items meets Mina’s condition, so policy requires the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3 rather than the proposed instructor conference at urgency 2.", "target": "false", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}]}, {"justification": "A score of 2/5 establishes nonmastery. Explicit refutation of exhibiting the pattern on at least two diagnostic items establishes that Mina’s condition is not met, so policy requires the instructor correction conference at urgency 2, matching the proposal.", "target": "true", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2.", "negative_left": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2.", "negative_right": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found no “larger denominator means larger fraction” pattern on diagnostic item 4.", "right": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 4."}, "verifier_independent_model": false}, "family": "scale-diverse-063-004", "id": "scale-diverse-063-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — Mina is not mastered, but she must be routed to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3.", "true": "Yes — the proposed not-mastered classification, instructor conference, and urgency 2 assignment are all correct."}, "instructions": "Decide whether this proposed triage is correct: mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Answer Yes or No according to the stated policy.", "type": "noul"}}, "state": "Field note: A separate answer-key reconciliation established that Mina answered exactly 2 of the 5 items on her five-item fraction quiz correctly. At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2. At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 4. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest)."}, "method": "c2d", "provenance": {"source_id": "diverse-063", "source_is_synthetic": true, "source_sha256": "1ab2940f7e11265b1fc7b64fecf6943e6d0f6fc9c6402c436749bbc8483bd245", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing mastery, routing, and urgency policies from the original state, while the unchanged questions preserve the request and scoring criteria. Mina, the five-item fraction quiz, the proposed triage path, and relevant item bindings remain unchanged. The two evidence spans are complete factual sentences. The counterfactual coherently changes only item 4 from exhibiting the pattern to not exhibiting it, without conflicting with the item 1 or item 2 findings or the two-correct total. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "full_context_fact_states": {"base": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "counterfactual": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "remove_left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "remove_right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "negative_pair": {"pattern_on_at_least_2_diagnostic_items": "refuted"}, "negative_sentence": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "positive_pair": {"pattern_on_at_least_2_diagnostic_items": "supported"}, "right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express single factual relationships, and the focus is the factual presence or absence of a response pattern rather than a policy conclusion. Both focus assignments are realizable while holding the 2/5 score fixed. The policy evidence preserves all substantive state-originating rules needed to apply the unchanged question; omitted details such as guided-practice performance are not needed for these rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly 2/5 correct necessarily fails the stated 4/5 mastery threshold. With the misconception pattern present on at least two diagnostic items, Mina’s condition is met, requiring the tutor’s Fraction-Strip Comparison Lab at urgency 3; therefore the proposed instructor conference and urgency 2 triage is incorrect.", "rule_index": 0, "sound": true}, {"reason": "Exactly 2/5 correct establishes nonmastery. Refutation of the at-least-two pattern proposition entails that Mina’s condition is not met, so the otherwise branch requires the instructor correction conference at urgency 2, making every component of the proposal correct.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "quiz_score_2_of_5", "statement": "Mina answered exactly 2 of the 5 items on this five-item fraction quiz correctly."}, {"id": "pattern_on_at_least_2_diagnostic_items", "statement": "Mina’s responses on this five-item fraction quiz exhibit the “larger denominator means larger fraction” pattern on at least 2 of diagnostic items 1, 2, and 4."}], "base_state_json": "\"Field note: A separate answer-key reconciliation established that Mina answered exactly 2 of the 5 items on her five-item fraction quiz correctly. At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2. At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 4. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest).\"", "base_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}], "counter_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}], "focus_atom": "pattern_on_at_least_2_diagnostic_items", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2."}, {"path": [], "text": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 4."}], "policy_evidence": [{"path": [], "text": "Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”"}, {"path": [], "text": "Policy defines mastery as 4/5 plus 2/3 diagnostic items correct."}, {"path": [], "text": "Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2."}, {"path": [], "text": "Urgency runs from 1 (lowest) to 3 (highest)."}], "rules": [{"justification": "A score of 2/5 establishes nonmastery. Exhibiting the stated pattern on at least two diagnostic items meets Mina’s condition, so policy requires the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3 rather than the proposed instructor conference at urgency 2.", "target": "false", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}]}, {"justification": "A score of 2/5 establishes nonmastery. Explicit refutation of exhibiting the pattern on at least two diagnostic items establishes that Mina’s condition is not met, so policy requires the instructor correction conference at urgency 2, matching the proposal.", "target": "true", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2.", "negative_left": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2.", "negative_right": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found no “larger denominator means larger fraction” pattern on diagnostic item 4.", "right": "At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 4."}, "verifier_independent_model": false}, "family": "scale-diverse-063-004", "id": "scale-diverse-063-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — Mina is not mastered, but she must be routed to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3.", "true": "Yes — the proposed not-mastered classification, instructor conference, and urgency 2 assignment are all correct."}, "instructions": "Decide whether this proposed triage is correct: mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Answer Yes or No according to the stated policy.", "type": "noul"}}, "state": "Field note: A separate answer-key reconciliation established that Mina answered exactly 2 of the 5 items on her five-item fraction quiz correctly. At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found the “larger denominator means larger fraction” pattern on diagnostic item 1 but not on diagnostic item 2. At 14:20 UTC on 2026-09-17, the scoring audit of Mina’s five-item fraction quiz found no “larger denominator means larger fraction” pattern on diagnostic item 4. Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.” Policy defines mastery as 4/5 plus 2/3 diagnostic items correct. Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2. Urgency runs from 1 (lowest) to 3 (highest)."}, "method": "c2d", "provenance": {"source_id": "diverse-063", "source_is_synthetic": true, "source_sha256": "1ab2940f7e11265b1fc7b64fecf6943e6d0f6fc9c6402c436749bbc8483bd245", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same mastery, temporal-update, and urgency policies and preserve Mira, the requested combined classification/route, and the newest-quiz binding. Each context contains exactly two complete factual evidence sentences. The counterfactual coherently changes the amber-marked-answer count from two to one without conflicting with the unchanged 2-of-4 score or prior-review statement. Neither constructed context embeds a gold answer, output code, proposition identifier, classifier instruction, or impermissible policy exception.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is factual rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the newest quiz has exactly two denominator-adding errors. The policy evidence correctly preserves the governing rules originating in the original state; criteria and instructions from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 2/4 score on the newest independent quiz is below the 3/4 mastery threshold, so Mira is not yet mastered. Exactly two denominator-adding errors prescribe urgency 2 and course-instructor fraction-strip review next class, while refutation of prior urgency-2 review excludes the recurrence-based urgency-3 condition.", "rule_index": 0, "sound": true}, {"reason": "A 2/4 score still entails not-yet-mastered, but refutation of exactly two denominator-adding errors entails that the exact condition for the requested urgency-2 route is absent. Other possible error counts would yield a different route or urgency, or no specified urgency, so the requested combined outcome is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira's newest independent quiz score is 2 correct answers out of 4 answers."}, {"id": "A2", "statement": "Mira's newest independent quiz contains exactly two denominator-adding errors."}, {"id": "A3", "statement": "Mira received urgency-2 review before her newest independent quiz."}], "base_state_json": "{\"context\":\"In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours).\",\"evidence\":[\"For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz.\",\"The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly two amber-marked answers.\"],\"request\":\"Verify whether Mira is not yet mastered and should be routed to course-instructor fraction-strip review at urgency 2.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A2", "focus_evidence": [{"path": ["evidence", "0"], "text": "For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz."}, {"path": ["evidence", "1"], "text": "The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly two amber-marked answers."}], "policy_evidence": [{"path": ["context"], "text": "In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours)."}], "rules": [{"justification": "A score of 2/4 on the overriding newest independent quiz is below the 3/4 mastery threshold. Exactly two denominator-adding errors require course-instructor fraction-strip review next class at urgency 2, and the absence of prior urgency-2 review excludes the recurrence-based urgency-3 outcome.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}, {"justification": "Although 2/4 remains below the mastery threshold, an explicitly non-two count of denominator-adding errors cannot satisfy the rubric's exact condition for the requested urgency-2 route; therefore the requested combined classification, route, and urgency level is not correct.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz.", "negative_left": "For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz.", "negative_right": "The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly one amber-marked answer.", "right": "The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly two amber-marked answers."}, "verifier_independent_model": false}, "family": "scale-diverse-064-001", "id": "scale-diverse-064-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not apply that combined mastery classification, route, and urgency level.", "true": "Yes — classify the skill as not yet mastered and route Mira to course-instructor fraction-strip review next class at urgency 2."}, "instructions": "Decide whether the requested routing and urgency are correct under the stated temporal-update rule and triage rubric. Answer yes or no.", "type": "noul"}}, "state": {"context": "In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours).", "evidence": ["For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz.", "The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly two amber-marked answers."], "request": "Verify whether Mira is not yet mastered and should be routed to course-instructor fraction-strip review at urgency 2."}}, "method": "c2d", "provenance": {"source_id": "diverse-064", "source_is_synthetic": true, "source_sha256": "8ada39e27611333fb04e1025d0b473127b9b817485d2f2c06367d2b62788d1a3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same mastery, temporal-update, and urgency policies and preserve Mira, the requested combined classification/route, and the newest-quiz binding. Each context contains exactly two complete factual evidence sentences. The counterfactual coherently changes the amber-marked-answer count from two to one without conflicting with the unchanged 2-of-4 score or prior-review statement. Neither constructed context embeds a gold answer, output code, proposition identifier, classifier instruction, or impermissible policy exception.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is factual rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the newest quiz has exactly two denominator-adding errors. The policy evidence correctly preserves the governing rules originating in the original state; criteria and instructions from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 2/4 score on the newest independent quiz is below the 3/4 mastery threshold, so Mira is not yet mastered. Exactly two denominator-adding errors prescribe urgency 2 and course-instructor fraction-strip review next class, while refutation of prior urgency-2 review excludes the recurrence-based urgency-3 condition.", "rule_index": 0, "sound": true}, {"reason": "A 2/4 score still entails not-yet-mastered, but refutation of exactly two denominator-adding errors entails that the exact condition for the requested urgency-2 route is absent. Other possible error counts would yield a different route or urgency, or no specified urgency, so the requested combined outcome is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira's newest independent quiz score is 2 correct answers out of 4 answers."}, {"id": "A2", "statement": "Mira's newest independent quiz contains exactly two denominator-adding errors."}, {"id": "A3", "statement": "Mira received urgency-2 review before her newest independent quiz."}], "base_state_json": "{\"context\":\"In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours).\",\"evidence\":[\"For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz.\",\"The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly two amber-marked answers.\"],\"request\":\"Verify whether Mira is not yet mastered and should be routed to course-instructor fraction-strip review at urgency 2.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A2", "focus_evidence": [{"path": ["evidence", "0"], "text": "For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz."}, {"path": ["evidence", "1"], "text": "The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly two amber-marked answers."}], "policy_evidence": [{"path": ["context"], "text": "In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours)."}], "rules": [{"justification": "A score of 2/4 on the overriding newest independent quiz is below the 3/4 mastery threshold. Exactly two denominator-adding errors require course-instructor fraction-strip review next class at urgency 2, and the absence of prior urgency-2 review excludes the recurrence-based urgency-3 outcome.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}, {"justification": "Although 2/4 remains below the mastery threshold, an explicitly non-two count of denominator-adding errors cannot satisfy the rubric's exact condition for the requested urgency-2 route; therefore the requested combined classification, route, and urgency level is not correct.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz.", "negative_left": "For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz.", "negative_right": "The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly one amber-marked answer.", "right": "The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly two amber-marked answers."}, "verifier_independent_model": false}, "family": "scale-diverse-064-001", "id": "scale-diverse-064-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not apply that combined mastery classification, route, and urgency level.", "true": "Yes — classify the skill as not yet mastered and route Mira to course-instructor fraction-strip review next class at urgency 2."}, "instructions": "Decide whether the requested routing and urgency are correct under the stated temporal-update rule and triage rubric. Answer yes or no.", "type": "noul"}}, "state": {"context": "In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours).", "evidence": ["For Mira's newest independent quiz completed at 10:00 on 14 September 2026, the audit counted one denominator-adding error for each amber-marked answer and none for any other answer, and Mira had received no urgency-2 review before that quiz.", "The final ledger for Mira's newest independent quiz completed at 10:00 on 14 September 2026 records 2 correct answers out of 4 answers and exactly one amber-marked answer."], "request": "Verify whether Mira is not yet mastered and should be routed to course-instructor fraction-strip review at urgency 2."}}, "method": "c2d", "provenance": {"source_id": "diverse-064", "source_is_synthetic": true, "source_sha256": "8ada39e27611333fb04e1025d0b473127b9b817485d2f2c06367d2b62788d1a3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing mastery, temporal-update, routing, and urgency policies without adding exceptions or defaults, while preserving Mira, the targeted skill, the newest-quiz path, and the requested urgency-2 route. Each evidence item is a complete factual sentence. In the base context, R41 and R42 are both on the quiz and are denominator-adding errors; in the counterfactual, R42 is no longer on the quiz, leaving only R41 as that error type. The counterfactual's 2/4 score remains coherent because the other incorrect response need not be a denominator-adding error. Neither context states a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is factual rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the newest quiz has exactly two denominator-adding errors. The policy evidence correctly preserves the governing rules originating in the original state; criteria and instructions from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 2/4 score on the newest independent quiz is below the 3/4 mastery threshold, so Mira is not yet mastered. Exactly two denominator-adding errors prescribe urgency 2 and course-instructor fraction-strip review next class, while refutation of prior urgency-2 review excludes the recurrence-based urgency-3 condition.", "rule_index": 0, "sound": true}, {"reason": "A 2/4 score still entails not-yet-mastered, but refutation of exactly two denominator-adding errors entails that the exact condition for the requested urgency-2 route is absent. Other possible error counts would yield a different route or urgency, or no specified urgency, so the requested combined outcome is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira's newest independent quiz score is 2 correct answers out of 4 answers."}, {"id": "A2", "statement": "Mira's newest independent quiz contains exactly two denominator-adding errors."}, {"id": "A3", "statement": "Mira received urgency-2 review before her newest independent quiz."}], "base_state_json": "{\"context\":\"In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours).\",\"evidence\":[\"Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R42, R43, and R44, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review.\",\"Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A2", "focus_evidence": [{"path": ["evidence", "0"], "text": "Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R42, R43, and R44, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review."}, {"path": ["evidence", "1"], "text": "Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors."}], "policy_evidence": [{"path": ["context"], "text": "In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours)."}], "rules": [{"justification": "A score of 2/4 on the overriding newest independent quiz is below the 3/4 mastery threshold. Exactly two denominator-adding errors require course-instructor fraction-strip review next class at urgency 2, and the absence of prior urgency-2 review excludes the recurrence-based urgency-3 outcome.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}, {"justification": "Although 2/4 remains below the mastery threshold, an explicitly non-two count of denominator-adding errors cannot satisfy the rubric's exact condition for the requested urgency-2 route; therefore the requested combined classification, route, and urgency level is not correct.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R42, R43, and R44, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review.", "negative_left": "Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R43, R44, and R45, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review.", "negative_right": "Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors.", "right": "Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors."}, "verifier_independent_model": false}, "family": "scale-diverse-064-002", "id": "scale-diverse-064-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not apply that combined mastery classification, route, and urgency level.", "true": "Yes — classify the skill as not yet mastered and route Mira to course-instructor fraction-strip review next class at urgency 2."}, "instructions": "Decide whether the requested routing and urgency are correct under the stated temporal-update rule and triage rubric. Answer yes or no.", "type": "noul"}}, "state": {"context": "In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours).", "evidence": ["Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R42, R43, and R44, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review.", "Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors."]}}, "method": "c2d", "provenance": {"source_id": "diverse-064", "source_is_synthetic": true, "source_sha256": "8ada39e27611333fb04e1025d0b473127b9b817485d2f2c06367d2b62788d1a3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing mastery, temporal-update, routing, and urgency policies without adding exceptions or defaults, while preserving Mira, the targeted skill, the newest-quiz path, and the requested urgency-2 route. Each evidence item is a complete factual sentence. In the base context, R41 and R42 are both on the quiz and are denominator-adding errors; in the counterfactual, R42 is no longer on the quiz, leaving only R41 as that error type. The counterfactual's 2/4 score remains coherent because the other incorrect response need not be a denominator-adding error. Neither context states a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is factual rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the newest quiz has exactly two denominator-adding errors. The policy evidence correctly preserves the governing rules originating in the original state; criteria and instructions from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 2/4 score on the newest independent quiz is below the 3/4 mastery threshold, so Mira is not yet mastered. Exactly two denominator-adding errors prescribe urgency 2 and course-instructor fraction-strip review next class, while refutation of prior urgency-2 review excludes the recurrence-based urgency-3 condition.", "rule_index": 0, "sound": true}, {"reason": "A 2/4 score still entails not-yet-mastered, but refutation of exactly two denominator-adding errors entails that the exact condition for the requested urgency-2 route is absent. Other possible error counts would yield a different route or urgency, or no specified urgency, so the requested combined outcome is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira's newest independent quiz score is 2 correct answers out of 4 answers."}, {"id": "A2", "statement": "Mira's newest independent quiz contains exactly two denominator-adding errors."}, {"id": "A3", "statement": "Mira received urgency-2 review before her newest independent quiz."}], "base_state_json": "{\"context\":\"In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours).\",\"evidence\":[\"Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R42, R43, and R44, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review.\",\"Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A2", "focus_evidence": [{"path": ["evidence", "0"], "text": "Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R42, R43, and R44, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review."}, {"path": ["evidence", "1"], "text": "Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors."}], "policy_evidence": [{"path": ["context"], "text": "In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours)."}], "rules": [{"justification": "A score of 2/4 on the overriding newest independent quiz is below the 3/4 mastery threshold. Exactly two denominator-adding errors require course-instructor fraction-strip review next class at urgency 2, and the absence of prior urgency-2 review excludes the recurrence-based urgency-3 outcome.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}, {"justification": "Although 2/4 remains below the mastery threshold, an explicitly non-two count of denominator-adding errors cannot satisfy the rubric's exact condition for the requested urgency-2 route; therefore the requested combined classification, route, and urgency level is not correct.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R42, R43, and R44, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review.", "negative_left": "Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R43, R44, and R45, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review.", "negative_right": "Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors.", "right": "Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors."}, "verifier_independent_model": false}, "family": "scale-diverse-064-002", "id": "scale-diverse-064-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not apply that combined mastery classification, route, and urgency level.", "true": "Yes — classify the skill as not yet mastered and route Mira to course-instructor fraction-strip review next class at urgency 2."}, "instructions": "Decide whether the requested routing and urgency are correct under the stated temporal-update rule and triage rubric. Answer yes or no.", "type": "noul"}}, "state": {"context": "In this course, the newest independent quiz overrides older checks. Mastery of like-denominator addition requires at least 3/4 correct and no denominator-adding error. One such error is urgency 1 (self-review cards); exactly two is urgency 2 (course-instructor fraction-strip review next class); three or more, or recurrence after urgency 2, is urgency 3 (learning-support tutor within 24 hours).", "evidence": ["Mira's newest independent quiz, submitted at 10:15 on 14 September 2026, comprised exactly response records R41, R43, R44, and R45, received a score of 2 correct answers out of 4 answers, and was completed before Mira had received any urgency-2 review.", "Mira's response-record audit codes R41 and R42 as denominator-adding errors and codes R43, R44, and R45 as not denominator-adding errors."]}}, "method": "c2d", "provenance": {"source_id": "diverse-064", "source_is_synthetic": true, "source_sha256": "8ada39e27611333fb04e1025d0b473127b9b817485d2f2c06367d2b62788d1a3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same learner, prior-lesson/current-quiz path, decision request, and authoritative-key/routing-rubric framework without adding policy exceptions or output instructions. The two evidence spans are complete factual sentences. The counterfactual changes only the prior lesson’s documented misconception; it does not contradict the separately documented current-quiz misconception, because the records concern different events. Neither context states an intervention level, support route, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Course records coordinator\",\"text\":\"The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater.\"},{\"speaker\":\"Assessment reviewer\",\"text\":\"The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater.\"},{\"speaker\":\"Course instructor\",\"text\":\"These finalized records contain no pending corrections or score changes. Using the routing rubric, decide the learner’s intervention level and support route.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}, {"path": ["1", "text"], "text": "The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater.", "negative_left": "The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with fewer such digits is always greater.", "negative_right": "The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater.", "right": "The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}, "verifier_independent_model": false}, "family": "scale-diverse-065-001", "id": "scale-diverse-065-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Course records coordinator", "text": "The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}, {"speaker": "Assessment reviewer", "text": "The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}, {"speaker": "Course instructor", "text": "These finalized records contain no pending corrections or score changes. Using the routing rubric, decide the learner’s intervention level and support route."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same learner, prior-lesson/current-quiz path, decision request, and authoritative-key/routing-rubric framework without adding policy exceptions or output instructions. The two evidence spans are complete factual sentences. The counterfactual changes only the prior lesson’s documented misconception; it does not contradict the separately documented current-quiz misconception, because the records concern different events. Neither context states an intervention level, support route, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Course records coordinator\",\"text\":\"The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater.\"},{\"speaker\":\"Assessment reviewer\",\"text\":\"The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater.\"},{\"speaker\":\"Course instructor\",\"text\":\"These finalized records contain no pending corrections or score changes. Using the routing rubric, decide the learner’s intervention level and support route.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}, {"path": ["1", "text"], "text": "The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater.", "negative_left": "The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with fewer such digits is always greater.", "negative_right": "The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater.", "right": "The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}, "verifier_independent_model": false}, "family": "scale-diverse-065-001", "id": "scale-diverse-065-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Course records coordinator", "text": "The learner’s Tuesday, September 8, 2026 prior lesson record identifies its sole documented misconception as the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with fewer such digits is always greater."}, {"speaker": "Assessment reviewer", "text": "The authoritative answer key marks exactly 3 of the learner’s 5 answers on the Friday, September 11, 2026 decimal-comparison quiz as correct, and the quiz review attributes both incorrect answers to its sole documented misconception: the rule that, when two decimals have different numbers of digits after the decimal point, the numeral with more such digits is always greater."}, {"speaker": "Course instructor", "text": "These finalized records contain no pending corrections or score changes. Using the routing rubric, decide the learner’s intervention level and support route."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete intervention rubric, scoring criteria, authority rule, and highest-applicable-level instruction; neither context alters that policy. Both contexts retain the same learner, decimal-comparison quiz, Friday September 11, 2026 assessment, Tuesday September 8, 2026 prior record, and intervention-routing request. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the prior lesson’s documented misconception without contradicting the finalized current-quiz findings or any unchanged assertion. Neither context states an intervention level, support-route answer, answer code, proposition ID, rule table, or label rationale; the generic request to apply the supplied rubric does not reveal an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Assessment records coordinator\",\"text\":\"The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value.\"},{\"speaker\":\"Lesson records coordinator\",\"text\":\"The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as choosing the decimal with more written digits as the larger value regardless of place value.\"},{\"speaker\":\"Program administrator\",\"text\":\"Both dated records are finalized for this learner. No scoring appeal, correction request, or record amendment is pending, and the intervention-routing handoff remains open.\"},{\"speaker\":\"Course instructor\",\"text\":\"Please use the supplied intervention rubric to select the single highest applicable level and its prescribed support route so scheduling can proceed.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value."}, {"path": ["1", "text"], "text": "The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as choosing the decimal with more written digits as the larger value regardless of place value."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value.", "negative_left": "The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value.", "negative_right": "The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as treating a decimal’s tenths digit as less significant than its hundredths digit.", "right": "The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as choosing the decimal with more written digits as the larger value regardless of place value."}, "verifier_independent_model": false}, "family": "scale-diverse-065-003", "id": "scale-diverse-065-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Assessment records coordinator", "text": "The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value."}, {"speaker": "Lesson records coordinator", "text": "The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as choosing the decimal with more written digits as the larger value regardless of place value."}, {"speaker": "Program administrator", "text": "Both dated records are finalized for this learner. No scoring appeal, correction request, or record amendment is pending, and the intervention-routing handoff remains open."}, {"speaker": "Course instructor", "text": "Please use the supplied intervention rubric to select the single highest applicable level and its prescribed support route so scheduling can proceed."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete intervention rubric, scoring criteria, authority rule, and highest-applicable-level instruction; neither context alters that policy. Both contexts retain the same learner, decimal-comparison quiz, Friday September 11, 2026 assessment, Tuesday September 8, 2026 prior record, and intervention-routing request. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the prior lesson’s documented misconception without contradicting the finalized current-quiz findings or any unchanged assertion. Neither context states an intervention level, support-route answer, answer code, proposition ID, rule table, or label rationale; the generic request to apply the supplied rubric does not reveal an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Assessment records coordinator\",\"text\":\"The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value.\"},{\"speaker\":\"Lesson records coordinator\",\"text\":\"The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as choosing the decimal with more written digits as the larger value regardless of place value.\"},{\"speaker\":\"Program administrator\",\"text\":\"Both dated records are finalized for this learner. No scoring appeal, correction request, or record amendment is pending, and the intervention-routing handoff remains open.\"},{\"speaker\":\"Course instructor\",\"text\":\"Please use the supplied intervention rubric to select the single highest applicable level and its prescribed support route so scheduling can proceed.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value."}, {"path": ["1", "text"], "text": "The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as choosing the decimal with more written digits as the larger value regardless of place value."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value.", "negative_left": "The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value.", "negative_right": "The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as treating a decimal’s tenths digit as less significant than its hundredths digit.", "right": "The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as choosing the decimal with more written digits as the larger value regardless of place value."}, "verifier_independent_model": false}, "family": "scale-diverse-065-003", "id": "scale-diverse-065-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Assessment records coordinator", "text": "The authoritative answer-key review of the learner’s Friday, September 11, 2026 decimal-comparison quiz recorded exactly 3 correct answers out of 5 and identified the sole misconception behind both errors as choosing the decimal with more written digits as the larger value regardless of place value."}, {"speaker": "Lesson records coordinator", "text": "The learner’s Tuesday, September 8, 2026 prior lesson record identified its sole documented misconception as treating a decimal’s tenths digit as less significant than its hundredths digit."}, {"speaker": "Program administrator", "text": "Both dated records are finalized for this learner. No scoring appeal, correction request, or record amendment is pending, and the intervention-routing handoff remains open."}, {"speaker": "Course instructor", "text": "Please use the supplied intervention rubric to select the single highest applicable level and its prescribed support route so scheduling can proceed."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same mastery threshold and do not contradict or modify the ordered routing rubric or its missing-evidence rule; those policies remain supplied by the original question. The learner, quiz, decision scope, and dated evidence path remain fixed, while the counterfactual changes only the observational Q4 answer-key value. The two focus spans are complete factual sentences. The changed key makes Q4 incorrect without conflicting with the missing Q5 or the diagnostic finding, and neither context includes a gold routing result, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns quiz correctness rather than policy. Both assignments are realizable: with exactly four responses present, the base can have four correct while the counter has exactly three correct, with no repeated misconception in either. The policy evidence is a valid quotation from the original state; all other governing rubric details remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly four responses are present and at least four recorded responses match the key, so four correct responses are established and the mastery threshold is verifiable. The no-repeated-misconception condition is also established. The missing fifth response therefore does not trigger level 3.", "rule_index": 0, "sound": true}, {"reason": "Exactly four responses are present; at least three are correct, while at least four correct is refuted. Thus exactly three are correct and one response is missing. Because the missing response could raise the total to four, mastery cannot be verified, so the explicit missing-evidence instruction requires level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly four of Nia's five Q1–Q5 quiz responses are present in the assembled evidence available for the routing decision."}, {"id": "a2", "statement": "At least three of Nia's recorded Q1–Q5 quiz responses match the corresponding entries in the official answer key."}, {"id": "a3", "statement": "At least four of Nia's recorded Q1–Q5 quiz responses match the corresponding entries in the official answer key."}, {"id": "a4", "statement": "No single misconception is exhibited by more than one of Nia's completed Q1–Q5 quiz items."}], "base_state_json": "\"On 14 March 2026, the course instructor assembled Nia’s available evidence from the Equivalent Fractions quiz for a routing decision and notified the learning support tutor that the review was underway. The evidence packet was then locked against later synchronization changes. At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5. The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = C, and Q5 = B. At 09:18, an item-level diagnostic review found that no misconception category appeared on more than one completed quiz item. Mastery requires at least four correct answers out of five. The instructor recorded that this locked packet was the evidence available for the decision.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5."}, {"path": [], "text": "The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = C, and Q5 = B."}], "policy_evidence": [{"path": [], "text": "Mastery requires at least four correct answers out of five."}], "rules": [{"justification": "Four recorded responses match the key, so the four-of-five threshold is verified despite the fifth response being missing; no misconception is repeated.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "Exactly four responses are present, while at least three but fewer than four match the key. Thus exactly three recorded responses are correct, and the missing fifth response could determine whether the four-of-five threshold is met, requiring evidence recovery.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5.", "negative_left": "At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5.", "negative_right": "The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = D, and Q5 = B.", "right": "The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = C, and Q5 = B."}, "verifier_independent_model": false}, "family": "scale-diverse-066-001", "id": "scale-diverse-066-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Verified mastery: At least four of five quiz responses are present and correct, with no repeated misconception. Continue to the next lesson without review.", "1 — Minor review: Exactly three responses are correct and all five are present, with only one isolated error type. Assign the instructor’s 10-minute equivalent-fractions recap before the next lesson.", "2 — Targeted support: Two or fewer responses are correct, all five are present, or two completed items show the same misconception. Route the learner to the learning support tutor’s guided fraction-model activity within two school days.", "3 — Evidence-recovery intervention: Response or lesson evidence is missing such that the four-of-five mastery threshold cannot be verified. The course instructor must recover or re-administer the missing items today and reserve a tutor check-in until mastery is established."], "instructions": "Route Nia using the ordered intervention rubric. Missing quiz fields are not counted as correct or incorrect. If missing evidence prevents verification against the mastery threshold, apply level 3 even when the available answers are correct.", "type": "score"}}, "state": "On 14 March 2026, the course instructor assembled Nia’s available evidence from the Equivalent Fractions quiz for a routing decision and notified the learning support tutor that the review was underway. The evidence packet was then locked against later synchronization changes. At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5. The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = C, and Q5 = B. At 09:18, an item-level diagnostic review found that no misconception category appeared on more than one completed quiz item. Mastery requires at least four correct answers out of five. The instructor recorded that this locked packet was the evidence available for the decision."}, "method": "c2d", "provenance": {"source_id": "diverse-066", "source_is_synthetic": true, "source_sha256": "85c8b11cc2d06cf41ebf7ed1afed75a5cc32f4a7c9a066382b853344e98b9542", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same mastery threshold and do not contradict or modify the ordered routing rubric or its missing-evidence rule; those policies remain supplied by the original question. The learner, quiz, decision scope, and dated evidence path remain fixed, while the counterfactual changes only the observational Q4 answer-key value. The two focus spans are complete factual sentences. The changed key makes Q4 incorrect without conflicting with the missing Q5 or the diagnostic finding, and neither context includes a gold routing result, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns quiz correctness rather than policy. Both assignments are realizable: with exactly four responses present, the base can have four correct while the counter has exactly three correct, with no repeated misconception in either. The policy evidence is a valid quotation from the original state; all other governing rubric details remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly four responses are present and at least four recorded responses match the key, so four correct responses are established and the mastery threshold is verifiable. The no-repeated-misconception condition is also established. The missing fifth response therefore does not trigger level 3.", "rule_index": 0, "sound": true}, {"reason": "Exactly four responses are present; at least three are correct, while at least four correct is refuted. Thus exactly three are correct and one response is missing. Because the missing response could raise the total to four, mastery cannot be verified, so the explicit missing-evidence instruction requires level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly four of Nia's five Q1–Q5 quiz responses are present in the assembled evidence available for the routing decision."}, {"id": "a2", "statement": "At least three of Nia's recorded Q1–Q5 quiz responses match the corresponding entries in the official answer key."}, {"id": "a3", "statement": "At least four of Nia's recorded Q1–Q5 quiz responses match the corresponding entries in the official answer key."}, {"id": "a4", "statement": "No single misconception is exhibited by more than one of Nia's completed Q1–Q5 quiz items."}], "base_state_json": "\"On 14 March 2026, the course instructor assembled Nia’s available evidence from the Equivalent Fractions quiz for a routing decision and notified the learning support tutor that the review was underway. The evidence packet was then locked against later synchronization changes. At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5. The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = C, and Q5 = B. At 09:18, an item-level diagnostic review found that no misconception category appeared on more than one completed quiz item. Mastery requires at least four correct answers out of five. The instructor recorded that this locked packet was the evidence available for the decision.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5."}, {"path": [], "text": "The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = C, and Q5 = B."}], "policy_evidence": [{"path": [], "text": "Mastery requires at least four correct answers out of five."}], "rules": [{"justification": "Four recorded responses match the key, so the four-of-five threshold is verified despite the fifth response being missing; no misconception is repeated.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "Exactly four responses are present, while at least three but fewer than four match the key. Thus exactly three recorded responses are correct, and the missing fifth response could determine whether the four-of-five threshold is met, requiring evidence recovery.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5.", "negative_left": "At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5.", "negative_right": "The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = D, and Q5 = B.", "right": "The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = C, and Q5 = B."}, "verifier_independent_model": false}, "family": "scale-diverse-066-001", "id": "scale-diverse-066-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Verified mastery: At least four of five quiz responses are present and correct, with no repeated misconception. Continue to the next lesson without review.", "1 — Minor review: Exactly three responses are correct and all five are present, with only one isolated error type. Assign the instructor’s 10-minute equivalent-fractions recap before the next lesson.", "2 — Targeted support: Two or fewer responses are correct, all five are present, or two completed items show the same misconception. Route the learner to the learning support tutor’s guided fraction-model activity within two school days.", "3 — Evidence-recovery intervention: Response or lesson evidence is missing such that the four-of-five mastery threshold cannot be verified. The course instructor must recover or re-administer the missing items today and reserve a tutor check-in until mastery is established."], "instructions": "Route Nia using the ordered intervention rubric. Missing quiz fields are not counted as correct or incorrect. If missing evidence prevents verification against the mastery threshold, apply level 3 even when the available answers are correct.", "type": "score"}}, "state": "On 14 March 2026, the course instructor assembled Nia’s available evidence from the Equivalent Fractions quiz for a routing decision and notified the learning support tutor that the review was underway. The evidence packet was then locked against later synchronization changes. At 09:10 on 14 March 2026, Nia's assembled quiz record showed Q1 = B, Q2 = D, Q3 = A, and Q4 = C, with no response recorded for Q5. The official answer key issued at 08:00 on 14 March 2026 listed Q1 = B, Q2 = D, Q3 = A, Q4 = D, and Q5 = B. At 09:18, an item-level diagnostic review found that no misconception category appeared on more than one completed quiz item. Mastery requires at least four correct answers out of five. The instructor recorded that this locked packet was the evidence available for the decision."}, "method": "c2d", "provenance": {"source_id": "diverse-066", "source_is_synthetic": true, "source_sha256": "85c8b11cc2d06cf41ebf7ed1afed75a5cc32f4a7c9a066382b853344e98b9542", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the governing four-of-five mastery threshold, while the unchanged questions preserve the complete ordered routing rubric and missing-evidence instructions. They remain bound to Nia, the same Q1–Q5 Equivalent Fractions quiz, and the same routing decision. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the official Q4 key from 3/10 to 2/9, without creating a duplicate or internally contradictory measurement. Neither context includes a score code, gold route, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns quiz correctness rather than policy. Both assignments are realizable: with exactly four responses present, the base can have four correct while the counter has exactly three correct, with no repeated misconception in either. The policy evidence is a valid quotation from the original state; all other governing rubric details remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly four responses are present and at least four recorded responses match the key, so four correct responses are established and the mastery threshold is verifiable. The no-repeated-misconception condition is also established. The missing fifth response therefore does not trigger level 3.", "rule_index": 0, "sound": true}, {"reason": "Exactly four responses are present; at least three are correct, while at least four correct is refuted. Thus exactly three are correct and one response is missing. Because the missing response could raise the total to four, mastery cannot be verified, so the explicit missing-evidence instruction requires level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly four of Nia's five Q1–Q5 quiz responses are present in the assembled evidence available for the routing decision."}, {"id": "a2", "statement": "At least three of Nia's recorded Q1–Q5 quiz responses match the corresponding entries in the official answer key."}, {"id": "a3", "statement": "At least four of Nia's recorded Q1–Q5 quiz responses match the corresponding entries in the official answer key."}, {"id": "a4", "statement": "No single misconception is exhibited by more than one of Nia's completed Q1–Q5 quiz items."}], "base_state_json": "\"The course instructor is reviewing the frozen operational handoff for Nia’s fictional “Equivalent Fractions” quiz; the learning support tutor remains available if the resulting route requires a check-in. No later response record, alternate key, or additional item-level analysis is available for this decision. In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item. The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 3/10, and Q5 = 9/14. Mastery requires at least four correct answers out of five.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item."}, {"path": [], "text": "The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 3/10, and Q5 = 9/14."}], "policy_evidence": [{"path": [], "text": "Mastery requires at least four correct answers out of five."}], "rules": [{"justification": "Four recorded responses match the key, so the four-of-five threshold is verified despite the fifth response being missing; no misconception is repeated.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "Exactly four responses are present, while at least three but fewer than four match the key. Thus exactly three recorded responses are correct, and the missing fifth response could determine whether the four-of-five threshold is met, requiring evidence recovery.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item.", "negative_left": "In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item.", "negative_right": "The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 2/9, and Q5 = 9/14.", "right": "The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 3/10, and Q5 = 9/14."}, "verifier_independent_model": false}, "family": "scale-diverse-066-003", "id": "scale-diverse-066-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Verified mastery: At least four of five quiz responses are present and correct, with no repeated misconception. Continue to the next lesson without review.", "1 — Minor review: Exactly three responses are correct and all five are present, with only one isolated error type. Assign the instructor’s 10-minute equivalent-fractions recap before the next lesson.", "2 — Targeted support: Two or fewer responses are correct, all five are present, or two completed items show the same misconception. Route the learner to the learning support tutor’s guided fraction-model activity within two school days.", "3 — Evidence-recovery intervention: Response or lesson evidence is missing such that the four-of-five mastery threshold cannot be verified. The course instructor must recover or re-administer the missing items today and reserve a tutor check-in until mastery is established."], "instructions": "Route Nia using the ordered intervention rubric. Missing quiz fields are not counted as correct or incorrect. If missing evidence prevents verification against the mastery threshold, apply level 3 even when the available answers are correct.", "type": "score"}}, "state": "The course instructor is reviewing the frozen operational handoff for Nia’s fictional “Equivalent Fractions” quiz; the learning support tutor remains available if the resulting route requires a check-in. No later response record, alternate key, or additional item-level analysis is available for this decision. In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item. The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 3/10, and Q5 = 9/14. Mastery requires at least four correct answers out of five."}, "method": "c2d", "provenance": {"source_id": "diverse-066", "source_is_synthetic": true, "source_sha256": "85c8b11cc2d06cf41ebf7ed1afed75a5cc32f4a7c9a066382b853344e98b9542", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the governing four-of-five mastery threshold, while the unchanged questions preserve the complete ordered routing rubric and missing-evidence instructions. They remain bound to Nia, the same Q1–Q5 Equivalent Fractions quiz, and the same routing decision. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the official Q4 key from 3/10 to 2/9, without creating a duplicate or internally contradictory measurement. Neither context includes a score code, gold route, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns quiz correctness rather than policy. Both assignments are realizable: with exactly four responses present, the base can have four correct while the counter has exactly three correct, with no repeated misconception in either. The policy evidence is a valid quotation from the original state; all other governing rubric details remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly four responses are present and at least four recorded responses match the key, so four correct responses are established and the mastery threshold is verifiable. The no-repeated-misconception condition is also established. The missing fifth response therefore does not trigger level 3.", "rule_index": 0, "sound": true}, {"reason": "Exactly four responses are present; at least three are correct, while at least four correct is refuted. Thus exactly three are correct and one response is missing. Because the missing response could raise the total to four, mastery cannot be verified, so the explicit missing-evidence instruction requires level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly four of Nia's five Q1–Q5 quiz responses are present in the assembled evidence available for the routing decision."}, {"id": "a2", "statement": "At least three of Nia's recorded Q1–Q5 quiz responses match the corresponding entries in the official answer key."}, {"id": "a3", "statement": "At least four of Nia's recorded Q1–Q5 quiz responses match the corresponding entries in the official answer key."}, {"id": "a4", "statement": "No single misconception is exhibited by more than one of Nia's completed Q1–Q5 quiz items."}], "base_state_json": "\"The course instructor is reviewing the frozen operational handoff for Nia’s fictional “Equivalent Fractions” quiz; the learning support tutor remains available if the resulting route requires a check-in. No later response record, alternate key, or additional item-level analysis is available for this decision. In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item. The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 3/10, and Q5 = 9/14. Mastery requires at least four correct answers out of five.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item."}, {"path": [], "text": "The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 3/10, and Q5 = 9/14."}], "policy_evidence": [{"path": [], "text": "Mastery requires at least four correct answers out of five."}], "rules": [{"justification": "Four recorded responses match the key, so the four-of-five threshold is verified despite the fifth response being missing; no misconception is repeated.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "Exactly four responses are present, while at least three but fewer than four match the key. Thus exactly three recorded responses are correct, and the missing fifth response could determine whether the four-of-five threshold is met, requiring evidence recovery.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item.", "negative_left": "In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item.", "negative_right": "The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 2/9, and Q5 = 9/14.", "right": "The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 3/10, and Q5 = 9/14."}, "verifier_independent_model": false}, "family": "scale-diverse-066-003", "id": "scale-diverse-066-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Verified mastery: At least four of five quiz responses are present and correct, with no repeated misconception. Continue to the next lesson without review.", "1 — Minor review: Exactly three responses are correct and all five are present, with only one isolated error type. Assign the instructor’s 10-minute equivalent-fractions recap before the next lesson.", "2 — Targeted support: Two or fewer responses are correct, all five are present, or two completed items show the same misconception. Route the learner to the learning support tutor’s guided fraction-model activity within two school days.", "3 — Evidence-recovery intervention: Response or lesson evidence is missing such that the four-of-five mastery threshold cannot be verified. The course instructor must recover or re-administer the missing items today and reserve a tutor check-in until mastery is established."], "instructions": "Route Nia using the ordered intervention rubric. Missing quiz fields are not counted as correct or incorrect. If missing evidence prevents verification against the mastery threshold, apply level 3 even when the available answers are correct.", "type": "score"}}, "state": "The course instructor is reviewing the frozen operational handoff for Nia’s fictional “Equivalent Fractions” quiz; the learning support tutor remains available if the resulting route requires a check-in. No later response record, alternate key, or additional item-level analysis is available for this decision. In the routing-evidence handoff frozen at 14:20 UTC on 31 March 2026, Nia's recorded responses were Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, and Q4 = 3/10, with Q5 absent, and the item-level review found no misconception exhibited on more than one completed item. The official answer key governing Nia's Q1–Q5 quiz in that handoff lists Q1 = 7/12, Q2 = 5/8, Q3 = 11/15, Q4 = 2/9, and Q5 = 9/14. Mastery requires at least four correct answers out of five."}, "method": "c2d", "provenance": {"source_id": "diverse-066", "source_is_synthetic": true, "source_sha256": "85c8b11cc2d06cf41ebf7ed1afed75a5cc32f4a7c9a066382b853344e98b9542", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the normal prerequisite, C+-with-instructor-support exception, and future-dependent-request policy without adding exceptions or defaults. They preserve Leo, Applied Statistics, the registrar review path, and the stated review date; the only counterfactual change is the recorded Statistics I symbol from Q to V. The two evidence spans are complete factual sentences, and the scale statement makes the grade change coherent without creating duplicate or contradictory measurements. Neither context includes an answer code, selected outcome, proposition identifier, rule table, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns Leo’s recorded grade rather than a policy conclusion. The base and counter assignments can both occur under the policy while differing only in whether the grade is at least C+; file completeness, support, unconditional intent, and a below-75 placement score can remain fixed. The policy evidence preserves the substantive state-originating prerequisite/exception and future-dependent-request rules; observations about Leo’s particular case need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes an unconditional request, a complete file, at least a C+ grade, and written instructor support. Those facts are sufficient for approve_ready_registrar under the stated criterion. A placement score below 75 does not negate the C+-plus-support exception.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an unconditional request and refutes a grade of at least C+, so the request lacks the minimum C+ required by the denial criterion. The below-75 placement score also excludes qualification through the normal placement route. Instructor support does not cure the missing minimum grade.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Leo’s request to enter Applied Statistics is unconditional at Mina’s review."}, {"id": "a2", "statement": "Leo’s submitted standard file for the Applied Statistics request is complete at Mina’s review."}, {"id": "a3", "statement": "The Statistics I grade recorded for Leo ranks at least C+ on the registrar’s letter-grade scale."}, {"id": "a4", "statement": "Leo’s submitted file contains written instructor support for his admission to Applied Statistics."}, {"id": "a5", "statement": "Leo’s placement score for Applied Statistics is below 75 at Mina’s review."}], "base_state_json": "\"Registrar Mina completed an evidence reconciliation for Leo’s Applied Statistics request on 14 March 2026. Leo’s signed request instructs the registrar to process his admission without any condition, unresolved contingency, or future trigger. Mina confirms that the submitted standard file is complete. It includes a written statement from Leo’s instructor supporting his admission to Applied Statistics. The official placement result recorded for this review is 68, below 75. At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol Q. The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+. The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support. Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol Q."}, {"path": [], "text": "The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+."}], "policy_evidence": [{"path": [], "text": "The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support."}, {"path": [], "text": "Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}], "rules": [{"justification": "The request is unconditional and complete, and Leo has both the minimum Statistics I grade and written instructor support required for exception approval. His sub-75 placement score does not defeat the expressly permitted C+-plus-support exception.", "target": "approve_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The request is unconditional, but Leo lacks the minimum C+ and also has no qualifying placement score; written instructor support alone cannot satisfy either the normal prerequisite or the exception requirements.", "target": "deny_not_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol Q.", "negative_left": "At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol V.", "negative_right": "The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+.", "right": "The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+."}, "verifier_independent_model": false}, "family": "scale-diverse-073-002", "id": "scale-diverse-073-002-base", "input": {"questions": {"decision": {"criteria": {"approve_ready_registrar": "Recommend yes, rate Ready, and let registrar staff finalize. Use only for an unconditional request with a complete file, at least a C+ in Statistics I, and written instructor support.", "deny_not_ready_registrar": "Recommend no, rate Not Ready, and let registrar staff close the request. Use only for an unconditional request that lacks either the minimum C+ or instructor support.", "none_of_above": "Make no yes/no recommendation, rate Provisional, and route to the course coordinator. Use for a request whose intent depends on a future event or condition that has not yet resolved."}, "instructions": "Select the single processing outcome that matches the evidence and policy. Apply the conditional-intent rule even if the ordinary exception criteria would otherwise support a decision.", "type": "choice"}}, "state": "Registrar Mina completed an evidence reconciliation for Leo’s Applied Statistics request on 14 March 2026. Leo’s signed request instructs the registrar to process his admission without any condition, unresolved contingency, or future trigger. Mina confirms that the submitted standard file is complete. It includes a written statement from Leo’s instructor supporting his admission to Applied Statistics. The official placement result recorded for this review is 68, below 75. At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol Q. The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+. The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support. Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}, "method": "c2d", "provenance": {"source_id": "diverse-073", "source_is_synthetic": true, "source_sha256": "3d8b5af4be06b90ac301446c5958d7d599067d0a272006d010e26fd0847ef133", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_ready_registrar"}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the normal prerequisite, C+-with-instructor-support exception, and future-dependent-request policy without adding exceptions or defaults. They preserve Leo, Applied Statistics, the registrar review path, and the stated review date; the only counterfactual change is the recorded Statistics I symbol from Q to V. The two evidence spans are complete factual sentences, and the scale statement makes the grade change coherent without creating duplicate or contradictory measurements. Neither context includes an answer code, selected outcome, proposition identifier, rule table, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns Leo’s recorded grade rather than a policy conclusion. The base and counter assignments can both occur under the policy while differing only in whether the grade is at least C+; file completeness, support, unconditional intent, and a below-75 placement score can remain fixed. The policy evidence preserves the substantive state-originating prerequisite/exception and future-dependent-request rules; observations about Leo’s particular case need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes an unconditional request, a complete file, at least a C+ grade, and written instructor support. Those facts are sufficient for approve_ready_registrar under the stated criterion. A placement score below 75 does not negate the C+-plus-support exception.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an unconditional request and refutes a grade of at least C+, so the request lacks the minimum C+ required by the denial criterion. The below-75 placement score also excludes qualification through the normal placement route. Instructor support does not cure the missing minimum grade.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Leo’s request to enter Applied Statistics is unconditional at Mina’s review."}, {"id": "a2", "statement": "Leo’s submitted standard file for the Applied Statistics request is complete at Mina’s review."}, {"id": "a3", "statement": "The Statistics I grade recorded for Leo ranks at least C+ on the registrar’s letter-grade scale."}, {"id": "a4", "statement": "Leo’s submitted file contains written instructor support for his admission to Applied Statistics."}, {"id": "a5", "statement": "Leo’s placement score for Applied Statistics is below 75 at Mina’s review."}], "base_state_json": "\"Registrar Mina completed an evidence reconciliation for Leo’s Applied Statistics request on 14 March 2026. Leo’s signed request instructs the registrar to process his admission without any condition, unresolved contingency, or future trigger. Mina confirms that the submitted standard file is complete. It includes a written statement from Leo’s instructor supporting his admission to Applied Statistics. The official placement result recorded for this review is 68, below 75. At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol Q. The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+. The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support. Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol Q."}, {"path": [], "text": "The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+."}], "policy_evidence": [{"path": [], "text": "The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support."}, {"path": [], "text": "Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}], "rules": [{"justification": "The request is unconditional and complete, and Leo has both the minimum Statistics I grade and written instructor support required for exception approval. His sub-75 placement score does not defeat the expressly permitted C+-plus-support exception.", "target": "approve_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The request is unconditional, but Leo lacks the minimum C+ and also has no qualifying placement score; written instructor support alone cannot satisfy either the normal prerequisite or the exception requirements.", "target": "deny_not_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol Q.", "negative_left": "At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol V.", "negative_right": "The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+.", "right": "The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+."}, "verifier_independent_model": false}, "family": "scale-diverse-073-002", "id": "scale-diverse-073-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_ready_registrar": "Recommend yes, rate Ready, and let registrar staff finalize. Use only for an unconditional request with a complete file, at least a C+ in Statistics I, and written instructor support.", "deny_not_ready_registrar": "Recommend no, rate Not Ready, and let registrar staff close the request. Use only for an unconditional request that lacks either the minimum C+ or instructor support.", "none_of_above": "Make no yes/no recommendation, rate Provisional, and route to the course coordinator. Use for a request whose intent depends on a future event or condition that has not yet resolved."}, "instructions": "Select the single processing outcome that matches the evidence and policy. Apply the conditional-intent rule even if the ordinary exception criteria would otherwise support a decision.", "type": "choice"}}, "state": "Registrar Mina completed an evidence reconciliation for Leo’s Applied Statistics request on 14 March 2026. Leo’s signed request instructs the registrar to process his admission without any condition, unresolved contingency, or future trigger. Mina confirms that the submitted standard file is complete. It includes a written statement from Leo’s instructor supporting his admission to Applied Statistics. The official placement result recorded for this review is 68, below 75. At Mina’s review on 14 March 2026, the Statistics I grade field in Leo’s registrar record displayed the letter-grade symbol V. The registrar’s letter-grade scale applicable to Leo’s record on 14 March 2026 places Q one tier above C+ and V one tier below C+. The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support. Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}, "method": "c2d", "provenance": {"source_id": "diverse-073", "source_is_synthetic": true, "source_sha256": "3d8b5af4be06b90ac301446c5958d7d599067d0a272006d010e26fd0847ef133", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "deny_not_ready_registrar"}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing prerequisite, exception, and future-dependent-request policies, while the unchanged questions preserve all outcome criteria and instructions. Leo, Applied Statistics, Mina’s review, and the relevant registrar-grade path remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes the locked grade symbol from Q7 to N4 while retaining a scale that places Q7 above C+ and N4 below C+, with no contradictory duplicate measurement. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction; the outcome terminology appears only as natural governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns Leo’s recorded grade rather than a policy conclusion. The base and counter assignments can both occur under the policy while differing only in whether the grade is at least C+; file completeness, support, unconditional intent, and a below-75 placement score can remain fixed. The policy evidence preserves the substantive state-originating prerequisite/exception and future-dependent-request rules; observations about Leo’s particular case need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes an unconditional request, a complete file, at least a C+ grade, and written instructor support. Those facts are sufficient for approve_ready_registrar under the stated criterion. A placement score below 75 does not negate the C+-plus-support exception.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an unconditional request and refutes a grade of at least C+, so the request lacks the minimum C+ required by the denial criterion. The below-75 placement score also excludes qualification through the normal placement route. Instructor support does not cure the missing minimum grade.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Leo’s request to enter Applied Statistics is unconditional at Mina’s review."}, {"id": "a2", "statement": "Leo’s submitted standard file for the Applied Statistics request is complete at Mina’s review."}, {"id": "a3", "statement": "The Statistics I grade recorded for Leo ranks at least C+ on the registrar’s letter-grade scale."}, {"id": "a4", "statement": "Leo’s submitted file contains written instructor support for his admission to Applied Statistics."}, {"id": "a5", "statement": "Leo’s placement score for Applied Statistics is below 75 at Mina’s review."}], "base_state_json": "\"Operational handoff — Mina is reviewing Leo’s Applied Statistics request. His signed request authorizes immediate processing and contains no future event, condition, or cancellation clause. The submitted standard file is complete and includes the request, registrar record, placement result, and a written note from the instructor supporting Leo’s admission. The recorded placement result is 68.\\n\\nAt Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol Q7. The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+.\\n\\nThe normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support.\\n\\nPolicy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol Q7."}, {"path": [], "text": "The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+."}], "policy_evidence": [{"path": [], "text": "The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support."}, {"path": [], "text": "Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}], "rules": [{"justification": "The request is unconditional and complete, and Leo has both the minimum Statistics I grade and written instructor support required for exception approval. His sub-75 placement score does not defeat the expressly permitted C+-plus-support exception.", "target": "approve_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The request is unconditional, but Leo lacks the minimum C+ and also has no qualifying placement score; written instructor support alone cannot satisfy either the normal prerequisite or the exception requirements.", "target": "deny_not_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol Q7.", "negative_left": "At Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol N4.", "negative_right": "The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+.", "right": "The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+."}, "verifier_independent_model": false}, "family": "scale-diverse-073-003", "id": "scale-diverse-073-003-base", "input": {"questions": {"decision": {"criteria": {"approve_ready_registrar": "Recommend yes, rate Ready, and let registrar staff finalize. Use only for an unconditional request with a complete file, at least a C+ in Statistics I, and written instructor support.", "deny_not_ready_registrar": "Recommend no, rate Not Ready, and let registrar staff close the request. Use only for an unconditional request that lacks either the minimum C+ or instructor support.", "none_of_above": "Make no yes/no recommendation, rate Provisional, and route to the course coordinator. Use for a request whose intent depends on a future event or condition that has not yet resolved."}, "instructions": "Select the single processing outcome that matches the evidence and policy. Apply the conditional-intent rule even if the ordinary exception criteria would otherwise support a decision.", "type": "choice"}}, "state": "Operational handoff — Mina is reviewing Leo’s Applied Statistics request. His signed request authorizes immediate processing and contains no future event, condition, or cancellation clause. The submitted standard file is complete and includes the request, registrar record, placement result, and a written note from the instructor supporting Leo’s admission. The recorded placement result is 68.\n\nAt Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol Q7. The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+.\n\nThe normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support.\n\nPolicy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}, "method": "c2d", "provenance": {"source_id": "diverse-073", "source_is_synthetic": true, "source_sha256": "3d8b5af4be06b90ac301446c5958d7d599067d0a272006d010e26fd0847ef133", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_ready_registrar"}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing prerequisite, exception, and future-dependent-request policies, while the unchanged questions preserve all outcome criteria and instructions. Leo, Applied Statistics, Mina’s review, and the relevant registrar-grade path remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes the locked grade symbol from Q7 to N4 while retaining a scale that places Q7 above C+ and N4 below C+, with no contradictory duplicate measurement. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction; the outcome terminology appears only as natural governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns Leo’s recorded grade rather than a policy conclusion. The base and counter assignments can both occur under the policy while differing only in whether the grade is at least C+; file completeness, support, unconditional intent, and a below-75 placement score can remain fixed. The policy evidence preserves the substantive state-originating prerequisite/exception and future-dependent-request rules; observations about Leo’s particular case need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes an unconditional request, a complete file, at least a C+ grade, and written instructor support. Those facts are sufficient for approve_ready_registrar under the stated criterion. A placement score below 75 does not negate the C+-plus-support exception.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an unconditional request and refutes a grade of at least C+, so the request lacks the minimum C+ required by the denial criterion. The below-75 placement score also excludes qualification through the normal placement route. Instructor support does not cure the missing minimum grade.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Leo’s request to enter Applied Statistics is unconditional at Mina’s review."}, {"id": "a2", "statement": "Leo’s submitted standard file for the Applied Statistics request is complete at Mina’s review."}, {"id": "a3", "statement": "The Statistics I grade recorded for Leo ranks at least C+ on the registrar’s letter-grade scale."}, {"id": "a4", "statement": "Leo’s submitted file contains written instructor support for his admission to Applied Statistics."}, {"id": "a5", "statement": "Leo’s placement score for Applied Statistics is below 75 at Mina’s review."}], "base_state_json": "\"Operational handoff — Mina is reviewing Leo’s Applied Statistics request. His signed request authorizes immediate processing and contains no future event, condition, or cancellation clause. The submitted standard file is complete and includes the request, registrar record, placement result, and a written note from the instructor supporting Leo’s admission. The recorded placement result is 68.\\n\\nAt Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol Q7. The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+.\\n\\nThe normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support.\\n\\nPolicy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol Q7."}, {"path": [], "text": "The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+."}], "policy_evidence": [{"path": [], "text": "The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support."}, {"path": [], "text": "Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}], "rules": [{"justification": "The request is unconditional and complete, and Leo has both the minimum Statistics I grade and written instructor support required for exception approval. His sub-75 placement score does not defeat the expressly permitted C+-plus-support exception.", "target": "approve_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The request is unconditional, but Leo lacks the minimum C+ and also has no qualifying placement score; written instructor support alone cannot satisfy either the normal prerequisite or the exception requirements.", "target": "deny_not_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol Q7.", "negative_left": "At Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol N4.", "negative_right": "The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+.", "right": "The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+."}, "verifier_independent_model": false}, "family": "scale-diverse-073-003", "id": "scale-diverse-073-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_ready_registrar": "Recommend yes, rate Ready, and let registrar staff finalize. Use only for an unconditional request with a complete file, at least a C+ in Statistics I, and written instructor support.", "deny_not_ready_registrar": "Recommend no, rate Not Ready, and let registrar staff close the request. Use only for an unconditional request that lacks either the minimum C+ or instructor support.", "none_of_above": "Make no yes/no recommendation, rate Provisional, and route to the course coordinator. Use for a request whose intent depends on a future event or condition that has not yet resolved."}, "instructions": "Select the single processing outcome that matches the evidence and policy. Apply the conditional-intent rule even if the ordinary exception criteria would otherwise support a decision.", "type": "choice"}}, "state": "Operational handoff — Mina is reviewing Leo’s Applied Statistics request. His signed request authorizes immediate processing and contains no future event, condition, or cancellation clause. The submitted standard file is complete and includes the request, registrar record, placement result, and a written note from the instructor supporting Leo’s admission. The recorded placement result is 68.\n\nAt Mina’s review handoff at 14:20 UTC on 12 March 2026, the registrar’s locked Statistics I record for Leo displays grade symbol N4. The registrar’s letter-grade scale in force at that handoff places grade symbol Q7 above C+ and grade symbol N4 below C+.\n\nThe normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support.\n\nPolicy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}, "method": "c2d", "provenance": {"source_id": "diverse-073", "source_is_synthetic": true, "source_sha256": "3d8b5af4be06b90ac301446c5958d7d599067d0a272006d010e26fd0847ef133", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "deny_not_ready_registrar"}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original governing policy and the unchanged questions preserve all choice criteria and instructions. Mira, the August current-term placement result, entry Q-317, the latest instructor position, and the requested outputs remain bound consistently. The two focus spans are complete factual sentences. Changing Q-317’s score from 86 to 74 creates a coherent counterfactual without a duplicate conflicting score, and neither context embeds a gold answer, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual placement-threshold result rather than a policy conclusion. The base and counter assignments differ only on that focus and can be realized by changing the placement score from at least 80 to below 80. The policy evidence preserves the substantive state-originating policy, including latest-update precedence; the unchanged questions automatically preserve the outcome criteria. Both proposed rules include the required outcome exclusions and routing conditions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes completeness, passing grade and placement thresholds, absence of current instructor support, and non-equivalent coursework. These conditions are sufficient for no recommendation, coordinator routing, and Moderate readiness.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completeness, failure of the placement threshold, and non-equivalent coursework as an unusual coordinator-routing condition. These conditions are sufficient for no recommendation, coordinator routing, and Low readiness.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira's final transcript is present in the file."}, {"id": "a2", "statement": "Mira's current-term placement result is present in the file."}, {"id": "a3", "statement": "Mira's latest instructor position is present in the file."}, {"id": "a4", "statement": "Mira's final-transcript grade for Applied Programming is B or better."}, {"id": "a5", "statement": "Applied Programming is related to the study for which Mira seeks entry."}, {"id": "a6", "statement": "Mira's August current-term placement result is at least 80."}, {"id": "a7", "statement": "Mira's latest instructor position supports entry."}, {"id": "a8", "statement": "Mira's Applied Programming coursework is non-equivalent."}], "base_state_json": "{\"context\":\"Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete.\",\"evidence\":[\"The file audit confirms Mira’s final transcript is present. It lists Applied Programming with a final grade of B+ and identifies the course as related to her intended study; the records office flags the coursework as non-equivalent.\",\"The testing handoff assigns Mira's August current-term placement result to entry Q-317.\",\"Entry Q-317 records a placement score of 86 points.\",\"updates supersede earlier notes.\",\"The instructor-position section is present: a September 2 note supported entry, while the September 14 update withdrew that support after a lab diagnostic.\"],\"request\":\"Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["evidence", "1"], "text": "The testing handoff assigns Mira's August current-term placement result to entry Q-317."}, {"path": ["evidence", "2"], "text": "Entry Q-317 records a placement score of 86 points."}], "policy_evidence": [{"path": ["context"], "text": "Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete."}, {"path": ["evidence", "3"], "text": "updates supersede earlier notes."}, {"path": ["request"], "text": "Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence."}], "rules": [{"justification": "All required documents are present, both academic thresholds pass, current instructor support is absent, and non-equivalent coursework requires coordinator routing.", "target": "complete_no_coordinator_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All required documents are present, the placement threshold fails, and non-equivalent coursework supplies the unusual condition requiring coordinator routing.", "target": "complete_no_coordinator_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The testing handoff assigns Mira's August current-term placement result to entry Q-317.", "negative_left": "The testing handoff assigns Mira's August current-term placement result to entry Q-317.", "negative_right": "Entry Q-317 records a placement score of 74 points.", "right": "Entry Q-317 records a placement score of 86 points."}, "verifier_independent_model": false}, "family": "scale-diverse-074-003", "id": "scale-diverse-074-003-base", "input": {"questions": {"decision": {"criteria": {"complete_no_coordinator_low": "Use when the file is complete but the grade or placement threshold fails, and an unusual routing condition exists: recommend no, route to the course coordinator, and rate Low.", "complete_no_coordinator_moderate": "Use when the file is complete and academic thresholds pass, but current instructor support is absent, with non-equivalent coursework or a changed position: recommend no, route to the course coordinator, and rate Moderate.", "complete_yes_registrar_high": "Use when the file is complete, every approval test passes, and no unusual routing condition exists: recommend yes, route to registrar staff, and rate High.", "incomplete_adviser_indeterminate": "Use when any required document is absent: mark incomplete, make no approval recommendation, route to the academic adviser, and rate Indeterminate."}, "instructions": "Choose the single option whose complete outcome rubric matches the policy and evidence as currently updated.", "type": "choice"}}, "state": {"context": "Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete.", "evidence": ["The file audit confirms Mira’s final transcript is present. It lists Applied Programming with a final grade of B+ and identifies the course as related to her intended study; the records office flags the coursework as non-equivalent.", "The testing handoff assigns Mira's August current-term placement result to entry Q-317.", "Entry Q-317 records a placement score of 86 points.", "updates supersede earlier notes.", "The instructor-position section is present: a September 2 note supported entry, while the September 14 update withdrew that support after a lab diagnostic."], "request": "Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-074", "source_is_synthetic": true, "source_sha256": "f6b73c0c2df392cc17b3d3a55974caeaa07084ae400a30235bea57a187f39bcf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_no_coordinator_moderate"}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original governing policy and the unchanged questions preserve all choice criteria and instructions. Mira, the August current-term placement result, entry Q-317, the latest instructor position, and the requested outputs remain bound consistently. The two focus spans are complete factual sentences. Changing Q-317’s score from 86 to 74 creates a coherent counterfactual without a duplicate conflicting score, and neither context embeds a gold answer, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual placement-threshold result rather than a policy conclusion. The base and counter assignments differ only on that focus and can be realized by changing the placement score from at least 80 to below 80. The policy evidence preserves the substantive state-originating policy, including latest-update precedence; the unchanged questions automatically preserve the outcome criteria. Both proposed rules include the required outcome exclusions and routing conditions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes completeness, passing grade and placement thresholds, absence of current instructor support, and non-equivalent coursework. These conditions are sufficient for no recommendation, coordinator routing, and Moderate readiness.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completeness, failure of the placement threshold, and non-equivalent coursework as an unusual coordinator-routing condition. These conditions are sufficient for no recommendation, coordinator routing, and Low readiness.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira's final transcript is present in the file."}, {"id": "a2", "statement": "Mira's current-term placement result is present in the file."}, {"id": "a3", "statement": "Mira's latest instructor position is present in the file."}, {"id": "a4", "statement": "Mira's final-transcript grade for Applied Programming is B or better."}, {"id": "a5", "statement": "Applied Programming is related to the study for which Mira seeks entry."}, {"id": "a6", "statement": "Mira's August current-term placement result is at least 80."}, {"id": "a7", "statement": "Mira's latest instructor position supports entry."}, {"id": "a8", "statement": "Mira's Applied Programming coursework is non-equivalent."}], "base_state_json": "{\"context\":\"Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete.\",\"evidence\":[\"The file audit confirms Mira’s final transcript is present. It lists Applied Programming with a final grade of B+ and identifies the course as related to her intended study; the records office flags the coursework as non-equivalent.\",\"The testing handoff assigns Mira's August current-term placement result to entry Q-317.\",\"Entry Q-317 records a placement score of 86 points.\",\"updates supersede earlier notes.\",\"The instructor-position section is present: a September 2 note supported entry, while the September 14 update withdrew that support after a lab diagnostic.\"],\"request\":\"Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["evidence", "1"], "text": "The testing handoff assigns Mira's August current-term placement result to entry Q-317."}, {"path": ["evidence", "2"], "text": "Entry Q-317 records a placement score of 86 points."}], "policy_evidence": [{"path": ["context"], "text": "Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete."}, {"path": ["evidence", "3"], "text": "updates supersede earlier notes."}, {"path": ["request"], "text": "Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence."}], "rules": [{"justification": "All required documents are present, both academic thresholds pass, current instructor support is absent, and non-equivalent coursework requires coordinator routing.", "target": "complete_no_coordinator_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All required documents are present, the placement threshold fails, and non-equivalent coursework supplies the unusual condition requiring coordinator routing.", "target": "complete_no_coordinator_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The testing handoff assigns Mira's August current-term placement result to entry Q-317.", "negative_left": "The testing handoff assigns Mira's August current-term placement result to entry Q-317.", "negative_right": "Entry Q-317 records a placement score of 74 points.", "right": "Entry Q-317 records a placement score of 86 points."}, "verifier_independent_model": false}, "family": "scale-diverse-074-003", "id": "scale-diverse-074-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_no_coordinator_low": "Use when the file is complete but the grade or placement threshold fails, and an unusual routing condition exists: recommend no, route to the course coordinator, and rate Low.", "complete_no_coordinator_moderate": "Use when the file is complete and academic thresholds pass, but current instructor support is absent, with non-equivalent coursework or a changed position: recommend no, route to the course coordinator, and rate Moderate.", "complete_yes_registrar_high": "Use when the file is complete, every approval test passes, and no unusual routing condition exists: recommend yes, route to registrar staff, and rate High.", "incomplete_adviser_indeterminate": "Use when any required document is absent: mark incomplete, make no approval recommendation, route to the academic adviser, and rate Indeterminate."}, "instructions": "Choose the single option whose complete outcome rubric matches the policy and evidence as currently updated.", "type": "choice"}}, "state": {"context": "Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete.", "evidence": ["The file audit confirms Mira’s final transcript is present. It lists Applied Programming with a final grade of B+ and identifies the course as related to her intended study; the records office flags the coursework as non-equivalent.", "The testing handoff assigns Mira's August current-term placement result to entry Q-317.", "Entry Q-317 records a placement score of 74 points.", "updates supersede earlier notes.", "The instructor-position section is present: a September 2 note supported entry, while the September 14 update withdrew that support after a lab diagnostic."], "request": "Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-074", "source_is_synthetic": true, "source_sha256": "f6b73c0c2df392cc17b3d3a55974caeaa07084ae400a30235bea57a187f39bcf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_no_coordinator_low"}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same EDU 340 exception requirements, conflict-routing rule, readiness rubric, registrar authority, applicant, and decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only the testing-center export from 80 to 73; this coherently conflicts with the distinct verified portal score of 80 and introduces no duplicate assertion. Neither context embeds an answer code, gold label, proposition ID, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Records auditor\",\"text\":\"At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80.\"},{\"speaker\":\"Testing-center liaison\",\"text\":\"At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 80 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80."}, {"path": ["1", "text"], "text": "At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 80 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80.", "negative_left": "At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80.", "negative_right": "At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 73 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission.", "right": "At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 80 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission."}, "verifier_independent_model": false}, "family": "scale-diverse-075-001", "id": "scale-diverse-075-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Records auditor", "text": "At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80."}, {"speaker": "Testing-center liaison", "text": "At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 80 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same EDU 340 exception requirements, conflict-routing rule, readiness rubric, registrar authority, applicant, and decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only the testing-center export from 80 to 73; this coherently conflicts with the distinct verified portal score of 80 and introduces no duplicate assertion. Neither context embeds an answer code, gold label, proposition ID, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Records auditor\",\"text\":\"At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80.\"},{\"speaker\":\"Testing-center liaison\",\"text\":\"At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 80 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80."}, {"path": ["1", "text"], "text": "At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 80 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80.", "negative_left": "At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80.", "negative_right": "At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 73 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission.", "right": "At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 80 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission."}, "verifier_independent_model": false}, "family": "scale-diverse-075-001", "id": "scale-diverse-075-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Records auditor", "text": "At 14:06 UTC on 18 February 2026, an audit of the complete exception-request file for the student, EDU 340 applicant 5724, found that its only quantitative placement-test score reports were the testing portal report and the testing-center export dated 14 February 2026, and that the portal report had verified status and recorded a score of 80."}, {"speaker": "Testing-center liaison", "text": "At 14:11 UTC on 18 February 2026, the testing-center export dated 14 February 2026 recorded a score of 73 for EDU 340 applicant 5724's quantitative placement test, and the signed instructor note in that student's exception-request file recommended admission."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original EDU 340 eligibility and routing policies without adding exceptions or priorities, and they preserve the question’s student, course, registrar-routing, and prerequisite-exception bindings. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the testing-center score from 80 to 76; this coherently conflicts with the unchanged portal score of 80 and does not create a duplicate assertion about the export’s value. Neither context contains a gold answer, output instruction, answer code, proposition identifier, or non-policy label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Records analyst\",\"text\":\"As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission.\"},{\"speaker\":\"Testing-center record\",\"text\":\"The testing-center export dated 2026-09-11 reports a score of 80 for the student's EDU 340 quantitative placement test.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission."}, {"path": ["1", "text"], "text": "The testing-center export dated 2026-09-11 reports a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission.", "negative_left": "As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission.", "negative_right": "The testing-center export dated 2026-09-11 reports a score of 76 for the student's EDU 340 quantitative placement test.", "right": "The testing-center export dated 2026-09-11 reports a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "scale-diverse-075-002", "id": "scale-diverse-075-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Records analyst", "text": "As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission."}, {"speaker": "Testing-center record", "text": "The testing-center export dated 2026-09-11 reports a score of 80 for the student's EDU 340 quantitative placement test."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original EDU 340 eligibility and routing policies without adding exceptions or priorities, and they preserve the question’s student, course, registrar-routing, and prerequisite-exception bindings. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the testing-center score from 80 to 76; this coherently conflicts with the unchanged portal score of 80 and does not create a duplicate assertion about the export’s value. Neither context contains a gold answer, output instruction, answer code, proposition identifier, or non-policy label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Records analyst\",\"text\":\"As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission.\"},{\"speaker\":\"Testing-center record\",\"text\":\"The testing-center export dated 2026-09-11 reports a score of 80 for the student's EDU 340 quantitative placement test.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission."}, {"path": ["1", "text"], "text": "The testing-center export dated 2026-09-11 reports a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission.", "negative_left": "As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission.", "negative_right": "The testing-center export dated 2026-09-11 reports a score of 76 for the student's EDU 340 quantitative placement test.", "right": "The testing-center export dated 2026-09-11 reports a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "scale-diverse-075-002", "id": "scale-diverse-075-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Records analyst", "text": "As of 2026-09-12 at 14:32 UTC, the testing portal recorded a verified score of 80 for the student's EDU 340 quantitative placement test, and the complete exception-request file contained no other quantitative placement-test score reports besides that portal report and a testing-center export dated 2026-09-11, along with a signed instructor note recommending admission."}, {"speaker": "Testing-center record", "text": "The testing-center export dated 2026-09-11 reports a score of 76 for the student's EDU 340 quantitative placement test."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original EDU 340 threshold, recommendation requirement, conflict-routing rule, staff authority limit, and coordinator readiness rubric without adding exceptions or defaults. The same student request, course, registrar decision scope, and handoff timing remain bound. The two focus-evidence spans are complete factual observation sentences. In the counterfactual, the portal remains a verified 80 while the separately identified testing-center export changes to 74; this creates a coherent two-report conflict rather than a contradictory duplicate measurement. Neither constructed context states the requested yes/no answer, an answer code, proposition identifier, label rationale, or an output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Registrar handoff\",\"text\":\"At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission.\"},{\"speaker\":\"Records clerk\",\"text\":\"The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 80.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission."}, {"path": ["1", "text"], "text": "The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission.", "negative_left": "At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission.", "negative_right": "The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 74.", "right": "The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 80."}, "verifier_independent_model": false}, "family": "scale-diverse-075-003", "id": "scale-diverse-075-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Registrar handoff", "text": "At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission."}, {"speaker": "Records clerk", "text": "The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 80."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original EDU 340 threshold, recommendation requirement, conflict-routing rule, staff authority limit, and coordinator readiness rubric without adding exceptions or defaults. The same student request, course, registrar decision scope, and handoff timing remain bound. The two focus-evidence spans are complete factual observation sentences. In the counterfactual, the portal remains a verified 80 while the separately identified testing-center export changes to 74; this creates a coherent two-report conflict rather than a contradictory duplicate measurement. Neither constructed context states the requested yes/no answer, an answer code, proposition identifier, label rationale, or an output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Registrar handoff\",\"text\":\"At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission.\"},{\"speaker\":\"Records clerk\",\"text\":\"The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 80.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission."}, {"path": ["1", "text"], "text": "The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission.", "negative_left": "At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission.", "negative_right": "The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 74.", "right": "The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 80."}, "verifier_independent_model": false}, "family": "scale-diverse-075-003", "id": "scale-diverse-075-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Registrar handoff", "text": "At the registrar’s 14:00 UTC handoff on 16 September 2026, the complete documents for the student’s EDU 340 exception request contained exactly two quantitative placement-test score reports—the testing portal report, marked verified and showing 80, and the testing-center export dated 15 September 2026—plus a signed instructor note recommending admission."}, {"speaker": "Records clerk", "text": "The testing-center export dated 15 September 2026 records the student’s EDU 340 quantitative placement-test score as 74."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing registrar and coordinator policies, while the unchanged questions preserve the request and decision criteria. The EDU 340 request, registrar decision-maker, direct-recommendation issue, document scope, and score-report identities remain bound consistently. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the portal score from 80 to 67 while retaining the export score of 80, producing a non-duplicative conflict consistent with the unchanged policy. Neither context contains a gold answer, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Records auditor\",\"text\":\"The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 80, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission.\"},{\"speaker\":\"Testing-center records officer\",\"text\":\"The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 80, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission."}, {"path": ["1", "text"], "text": "The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 80, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission.", "negative_left": "The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 67, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission.", "negative_right": "The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request.", "right": "The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request."}, "verifier_independent_model": false}, "family": "scale-diverse-075-004", "id": "scale-diverse-075-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Records auditor", "text": "The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 80, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission."}, {"speaker": "Testing-center records officer", "text": "The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing registrar and coordinator policies, while the unchanged questions preserve the request and decision criteria. The EDU 340 request, registrar decision-maker, direct-recommendation issue, document scope, and score-report identities remain bound consistently. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the portal score from 80 to 67 while retaining the export score of 80, producing a non-duplicative conflict consistent with the unchanged policy. Neither context contains a gold answer, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Records auditor\",\"text\":\"The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 80, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission.\"},{\"speaker\":\"Testing-center records officer\",\"text\":\"The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 80, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission."}, {"path": ["1", "text"], "text": "The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 80, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission.", "negative_left": "The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 67, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission.", "negative_right": "The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request.", "right": "The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request."}, "verifier_independent_model": false}, "family": "scale-diverse-075-004", "id": "scale-diverse-075-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Records auditor", "text": "The student's complete EDU 340 exception documents contain exactly two quantitative placement-test score reports—the verified testing portal report, which records 67, and the testing-center export dated 2026-08-14—and include a signed instructor note recommending admission."}, {"speaker": "Testing-center records officer", "text": "The testing-center export dated 2026-08-14 records a quantitative placement-test score of 80 for the student's EDU 340 exception request."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same prerequisite requirements, routing rules, and readiness definitions, while preserving Maya Chen, the MAT 120-to-MAT 210 exception request, and the dated placement-report record. The counterfactual changes only field Beta from a pending result to a numerical score of 82; this is coherent with the unchanged statement that there is no evidence of a score below 75. The two focus-evidence spans are complete factual sentences. Neither context states a required classifier output, answer code, proposition ID, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"On 14 September 2026, Maya Chen requested an exception to the MAT 120 prerequisite for MAT 210. Intake staff confirmed that her submitted prerequisite-exception packet contained her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in the cited algebra course, which the equivalency register identifies as equivalent to MAT 120. The coordinator's note states that Maya appears prepared. At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta. At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “result pending,” with no other marks in either field. Final packet review found no documented evidence that Maya's placement result was below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta."}, {"path": [], "text": "At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “result pending,” with no other marks in either field."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta.", "negative_left": "At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta.", "negative_right": "At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “placement result: 82,” with no other marks in either field.", "right": "At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “result pending,” with no other marks in either field."}, "verifier_independent_model": false}, "family": "scale-diverse-076-001", "id": "scale-diverse-076-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "On 14 September 2026, Maya Chen requested an exception to the MAT 120 prerequisite for MAT 210. Intake staff confirmed that her submitted prerequisite-exception packet contained her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in the cited algebra course, which the equivalency register identifies as equivalent to MAT 120. The coordinator's note states that Maya appears prepared. At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta. At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “result pending,” with no other marks in either field. Final packet review found no documented evidence that Maya's placement result was below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same prerequisite requirements, routing rules, and readiness definitions, while preserving Maya Chen, the MAT 120-to-MAT 210 exception request, and the dated placement-report record. The counterfactual changes only field Beta from a pending result to a numerical score of 82; this is coherent with the unchanged statement that there is no evidence of a score below 75. The two focus-evidence spans are complete factual sentences. Neither context states a required classifier output, answer code, proposition ID, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"On 14 September 2026, Maya Chen requested an exception to the MAT 120 prerequisite for MAT 210. Intake staff confirmed that her submitted prerequisite-exception packet contained her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in the cited algebra course, which the equivalency register identifies as equivalent to MAT 120. The coordinator's note states that Maya appears prepared. At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta. At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “result pending,” with no other marks in either field. Final packet review found no documented evidence that Maya's placement result was below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta."}, {"path": [], "text": "At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “result pending,” with no other marks in either field."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta.", "negative_left": "At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta.", "negative_right": "At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “placement result: 82,” with no other marks in either field.", "right": "At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “result pending,” with no other marks in either field."}, "verifier_independent_model": false}, "family": "scale-diverse-076-001", "id": "scale-diverse-076-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "On 14 September 2026, Maya Chen requested an exception to the MAT 120 prerequisite for MAT 210. Intake staff confirmed that her submitted prerequisite-exception packet contained her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in the cited algebra course, which the equivalency register identifies as equivalent to MAT 120. The coordinator's note states that Maya appears prepared. At 09:10 UTC on 14 September 2026, the registrar certified that Maya Chen's official placement report was record QZ-481, whose entire placement-score section comprised fields Alpha and Beta. At 09:12 UTC on 14 September 2026, an archive inspection found that field Alpha of record QZ-481 read “assessment completed” and field Beta read “placement result: 82,” with no other marks in either field. Final packet review found no documented evidence that Maya's placement result was below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same prerequisite requirements, routing rules, and readiness ordering, while preserving Maya Chen’s MAT 210 prerequisite-exception request and the relevant handoff time. The two evidence spans are complete factual sentences. The counterfactual coherently replaces the nonnumeric P-17 content with an official score of 82, without conflicting with the statement that there is no documented result below 75. Neither context explicitly supplies a classification output, answer code, rule table, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"Operational handoff for Maya Chen’s request seeking an exception to the MAT 120 prerequisite for MAT 210. The submitted MAT 210 prerequisite-exception packet inventory lists her transcript. The transcript records a B+ in the cited algebra course, which the registrar’s equivalency table identifies as equivalent to MAT 120. The inventory also lists the academic adviser’s request form and a course-coordinator note. That note states that Maya Chen appears prepared. At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value. At that handoff, field P-17 contained only the words “manual review pending,” with no digits or number words. Packet review found no documented evidence that her placement result is below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value."}, {"path": [], "text": "At that handoff, field P-17 contained only the words “manual review pending,” with no digits or number words."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value.", "negative_left": "At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value.", "negative_right": "At that handoff, field P-17 contained only the decimal integer “82.”", "right": "At that handoff, field P-17 contained only the words “manual review pending,” with no digits or number words."}, "verifier_independent_model": false}, "family": "scale-diverse-076-003", "id": "scale-diverse-076-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "Operational handoff for Maya Chen’s request seeking an exception to the MAT 120 prerequisite for MAT 210. The submitted MAT 210 prerequisite-exception packet inventory lists her transcript. The transcript records a B+ in the cited algebra course, which the registrar’s equivalency table identifies as equivalent to MAT 120. The inventory also lists the academic adviser’s request form and a course-coordinator note. That note states that Maya Chen appears prepared. At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value. At that handoff, field P-17 contained only the words “manual review pending,” with no digits or number words. Packet review found no documented evidence that her placement result is below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same prerequisite requirements, routing rules, and readiness ordering, while preserving Maya Chen’s MAT 210 prerequisite-exception request and the relevant handoff time. The two evidence spans are complete factual sentences. The counterfactual coherently replaces the nonnumeric P-17 content with an official score of 82, without conflicting with the statement that there is no documented result below 75. Neither context explicitly supplies a classification output, answer code, rule table, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"Operational handoff for Maya Chen’s request seeking an exception to the MAT 120 prerequisite for MAT 210. The submitted MAT 210 prerequisite-exception packet inventory lists her transcript. The transcript records a B+ in the cited algebra course, which the registrar’s equivalency table identifies as equivalent to MAT 120. The inventory also lists the academic adviser’s request form and a course-coordinator note. That note states that Maya Chen appears prepared. At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value. At that handoff, field P-17 contained only the words “manual review pending,” with no digits or number words. Packet review found no documented evidence that her placement result is below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value."}, {"path": [], "text": "At that handoff, field P-17 contained only the words “manual review pending,” with no digits or number words."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value.", "negative_left": "At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value.", "negative_right": "At that handoff, field P-17 contained only the decimal integer “82.”", "right": "At that handoff, field P-17 contained only the words “manual review pending,” with no digits or number words."}, "verifier_independent_model": false}, "family": "scale-diverse-076-003", "id": "scale-diverse-076-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "Operational handoff for Maya Chen’s request seeking an exception to the MAT 120 prerequisite for MAT 210. The submitted MAT 210 prerequisite-exception packet inventory lists her transcript. The transcript records a B+ in the cited algebra course, which the registrar’s equivalency table identifies as equivalent to MAT 120. The inventory also lists the academic adviser’s request form and a course-coordinator note. That note states that Maya Chen appears prepared. At the 14:20 operational handoff on 17 September 2026, the registrar certified that field P-17 contains Maya Chen's official placement report's numerical placement-score value, if any, and that no other part of the report can contain that value. At that handoff, field P-17 contained only the decimal integer “82.” Packet review found no documented evidence that her placement result is below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original prerequisite, placement-score threshold and 12-month limit, coordinator authority, panel routing, documentation requirements, decision criteria, instructions, request scope, student, course, and relevant time relationship. The two focus spans are complete factual sentences. The counterfactual changes only the CHEM 220 start date from March 10 to March 20, 2026; this is coherent with the unchanged March 14, 2025 report date and introduces no contradictory duplicate assertion. Neither context embeds a selected readiness level, recommendation, routing answer, answer code, proposition ID, or classifier-output instruction beyond the preserved governing policy and task instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.\",\"Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.\",\"Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed.\"],\"instructions\":\"Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.\",\"type\":\"score\"}},\"state\":{\"context\":\"Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The exception file includes an official transcript showing that Mina received a D+ in CHEM 110.\",\"The file also includes an official placement report stating a score of 82.\",\"Mina Cho's official placement report is dated March 14, 2025.\",\"Mina Cho's requested CHEM 220 term starts on March 10, 2026.\",\"A CHEM 220 instructor note supporting Mina’s enrollment is included in the exception file.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["state", "evidence", "2"], "text": "Mina Cho's official placement report is dated March 14, 2025."}, {"path": ["state", "evidence", "3"], "text": "Mina Cho's requested CHEM 220 term starts on March 10, 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 14, 2025.", "negative_left": "Mina Cho's official placement report is dated March 14, 2025.", "negative_right": "Mina Cho's requested CHEM 220 term starts on March 20, 2026.", "right": "Mina Cho's requested CHEM 220 term starts on March 10, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-077-001", "id": "scale-diverse-077-001-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The exception file includes an official transcript showing that Mina received a D+ in CHEM 110.", "The file also includes an official placement report stating a score of 82.", "Mina Cho's official placement report is dated March 14, 2025.", "Mina Cho's requested CHEM 220 term starts on March 10, 2026.", "A CHEM 220 instructor note supporting Mina’s enrollment is included in the exception file."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original prerequisite, placement-score threshold and 12-month limit, coordinator authority, panel routing, documentation requirements, decision criteria, instructions, request scope, student, course, and relevant time relationship. The two focus spans are complete factual sentences. The counterfactual changes only the CHEM 220 start date from March 10 to March 20, 2026; this is coherent with the unchanged March 14, 2025 report date and introduces no contradictory duplicate assertion. Neither context embeds a selected readiness level, recommendation, routing answer, answer code, proposition ID, or classifier-output instruction beyond the preserved governing policy and task instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.\",\"Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.\",\"Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed.\"],\"instructions\":\"Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.\",\"type\":\"score\"}},\"state\":{\"context\":\"Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The exception file includes an official transcript showing that Mina received a D+ in CHEM 110.\",\"The file also includes an official placement report stating a score of 82.\",\"Mina Cho's official placement report is dated March 14, 2025.\",\"Mina Cho's requested CHEM 220 term starts on March 10, 2026.\",\"A CHEM 220 instructor note supporting Mina’s enrollment is included in the exception file.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["state", "evidence", "2"], "text": "Mina Cho's official placement report is dated March 14, 2025."}, {"path": ["state", "evidence", "3"], "text": "Mina Cho's requested CHEM 220 term starts on March 10, 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 14, 2025.", "negative_left": "Mina Cho's official placement report is dated March 14, 2025.", "negative_right": "Mina Cho's requested CHEM 220 term starts on March 20, 2026.", "right": "Mina Cho's requested CHEM 220 term starts on March 10, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-077-001", "id": "scale-diverse-077-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The exception file includes an official transcript showing that Mina received a D+ in CHEM 110.", "The file also includes an official placement report stating a score of 82.", "Mina Cho's official placement report is dated March 14, 2025.", "Mina Cho's requested CHEM 220 term starts on March 20, 2026.", "A CHEM 220 instructor note supporting Mina’s enrollment is included in the exception file."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same prerequisite, placement-score timing and scope limits, completeness requirements, request, and Mina/CHEM 220/date bindings. The two focus spans are complete factual sentences, and changing the report date from May 19, 2025 to April 27, 2025 coherently moves it from within 12 months of the May 11, 2026 term start to outside that period without conflicting with other evidence. Neither constructed context embeds a gold answer, output code, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":{\"case_note\":\"Mina Cho submitted a prerequisite-exception request for enrollment in CHEM 220. Registrar staff reconciled the student identifiers across the submitted materials and confirmed that every document belongs to Mina’s exception file. The coordinator is reviewing the assembled record under the following requirements.\",\"policy\":[\"The normal prerequisite is CHEM 110 with at least C.\",\"A placement score of 75 or higher substitutes only when dated within 12 months.\",\"The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel.\",\"A complete file requires an official transcript, dated placement report, and instructor note.\"]},\"evidence\":[\"The exception file contains an official transcript showing that Mina earned D+ in CHEM 110.\",\"The exception file contains an official placement report showing a score of 82.\",\"The exception file contains a CHEM 220 instructor note supporting Mina’s enrollment based on her laboratory work.\",\"Mina Cho's official placement report states May 19, 2025, as its report date.\",\"Mina Cho's requested CHEM 220 term starts on May 11, 2026.\",\"Registrar staff verified the transcript and placement report as official and confirmed that the instructor note is the version submitted for CHEM 220 review.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mina Cho's official placement report states May 19, 2025, as its report date."}, {"path": ["evidence", "4"], "text": "Mina Cho's requested CHEM 220 term starts on May 11, 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report states May 19, 2025, as its report date.", "negative_left": "Mina Cho's official placement report states April 27, 2025, as its report date.", "negative_right": "Mina Cho's requested CHEM 220 term starts on May 11, 2026.", "right": "Mina Cho's requested CHEM 220 term starts on May 11, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-077-002", "id": "scale-diverse-077-002-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": {"case_note": "Mina Cho submitted a prerequisite-exception request for enrollment in CHEM 220. Registrar staff reconciled the student identifiers across the submitted materials and confirmed that every document belongs to Mina’s exception file. The coordinator is reviewing the assembled record under the following requirements.", "policy": ["The normal prerequisite is CHEM 110 with at least C.", "A placement score of 75 or higher substitutes only when dated within 12 months.", "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel.", "A complete file requires an official transcript, dated placement report, and instructor note."]}, "evidence": ["The exception file contains an official transcript showing that Mina earned D+ in CHEM 110.", "The exception file contains an official placement report showing a score of 82.", "The exception file contains a CHEM 220 instructor note supporting Mina’s enrollment based on her laboratory work.", "Mina Cho's official placement report states May 19, 2025, as its report date.", "Mina Cho's requested CHEM 220 term starts on May 11, 2026.", "Registrar staff verified the transcript and placement report as official and confirmed that the instructor note is the version submitted for CHEM 220 review."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same prerequisite, placement-score timing and scope limits, completeness requirements, request, and Mina/CHEM 220/date bindings. The two focus spans are complete factual sentences, and changing the report date from May 19, 2025 to April 27, 2025 coherently moves it from within 12 months of the May 11, 2026 term start to outside that period without conflicting with other evidence. Neither constructed context embeds a gold answer, output code, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":{\"case_note\":\"Mina Cho submitted a prerequisite-exception request for enrollment in CHEM 220. Registrar staff reconciled the student identifiers across the submitted materials and confirmed that every document belongs to Mina’s exception file. The coordinator is reviewing the assembled record under the following requirements.\",\"policy\":[\"The normal prerequisite is CHEM 110 with at least C.\",\"A placement score of 75 or higher substitutes only when dated within 12 months.\",\"The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel.\",\"A complete file requires an official transcript, dated placement report, and instructor note.\"]},\"evidence\":[\"The exception file contains an official transcript showing that Mina earned D+ in CHEM 110.\",\"The exception file contains an official placement report showing a score of 82.\",\"The exception file contains a CHEM 220 instructor note supporting Mina’s enrollment based on her laboratory work.\",\"Mina Cho's official placement report states May 19, 2025, as its report date.\",\"Mina Cho's requested CHEM 220 term starts on May 11, 2026.\",\"Registrar staff verified the transcript and placement report as official and confirmed that the instructor note is the version submitted for CHEM 220 review.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mina Cho's official placement report states May 19, 2025, as its report date."}, {"path": ["evidence", "4"], "text": "Mina Cho's requested CHEM 220 term starts on May 11, 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report states May 19, 2025, as its report date.", "negative_left": "Mina Cho's official placement report states April 27, 2025, as its report date.", "negative_right": "Mina Cho's requested CHEM 220 term starts on May 11, 2026.", "right": "Mina Cho's requested CHEM 220 term starts on May 11, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-077-002", "id": "scale-diverse-077-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": {"case_note": "Mina Cho submitted a prerequisite-exception request for enrollment in CHEM 220. Registrar staff reconciled the student identifiers across the submitted materials and confirmed that every document belongs to Mina’s exception file. The coordinator is reviewing the assembled record under the following requirements.", "policy": ["The normal prerequisite is CHEM 110 with at least C.", "A placement score of 75 or higher substitutes only when dated within 12 months.", "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel.", "A complete file requires an official transcript, dated placement report, and instructor note."]}, "evidence": ["The exception file contains an official transcript showing that Mina earned D+ in CHEM 110.", "The exception file contains an official placement report showing a score of 82.", "The exception file contains a CHEM 220 instructor note supporting Mina’s enrollment based on her laboratory work.", "Mina Cho's official placement report states April 27, 2025, as its report date.", "Mina Cho's requested CHEM 220 term starts on May 11, 2026.", "Registrar staff verified the transcript and placement report as official and confirmed that the instructor note is the version submitted for CHEM 220 review."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions and preserve the original governing prerequisite, placement-score threshold, 12-month validity limit, coordinator scope restriction, panel routing, and file requirements. The request remains bound to Mina Cho’s CHEM 220 approval readiness and routing. The two focus spans are complete factual sentences. The counterfactual changes only the requested term start date from December 2, 2025, to February 3, 2026; this is coherent with the unchanged January 18, 2025 placement-report date and creates no duplicate or contradictory measurement. Neither context embeds a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.\",\"Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.\",\"Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed.\"],\"instructions\":\"Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.\",\"type\":\"score\"}},\"state\":{\"context\":\"CHEM 220 exception handoff for Mina Cho. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The coordinator’s intake inventory lists Mina’s official transcript, official placement report, and CHEM 220 instructor note as received.\",\"Mina Cho's official placement report states January 18, 2025, as its date.\",\"Mina Cho's requested CHEM 220 term starts on December 2, 2025.\",\"The official transcript records a D+ for Mina in CHEM 110.\",\"The official placement report records a score of 82.\",\"The CHEM 220 instructor note reports strong laboratory skills and supports Mina’s enrollment.\",\"Registrar staff verified the transcript and placement report as official and confirmed that the instructor note belongs to this request.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "Mina Cho's official placement report states January 18, 2025, as its date."}, {"path": ["state", "evidence", "2"], "text": "Mina Cho's requested CHEM 220 term starts on December 2, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report states January 18, 2025, as its date.", "negative_left": "Mina Cho's official placement report states January 18, 2025, as its date.", "negative_right": "Mina Cho's requested CHEM 220 term starts on February 3, 2026.", "right": "Mina Cho's requested CHEM 220 term starts on December 2, 2025."}, "verifier_independent_model": false}, "family": "scale-diverse-077-003", "id": "scale-diverse-077-003-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "CHEM 220 exception handoff for Mina Cho. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The coordinator’s intake inventory lists Mina’s official transcript, official placement report, and CHEM 220 instructor note as received.", "Mina Cho's official placement report states January 18, 2025, as its date.", "Mina Cho's requested CHEM 220 term starts on December 2, 2025.", "The official transcript records a D+ for Mina in CHEM 110.", "The official placement report records a score of 82.", "The CHEM 220 instructor note reports strong laboratory skills and supports Mina’s enrollment.", "Registrar staff verified the transcript and placement report as official and confirmed that the instructor note belongs to this request."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions and preserve the original governing prerequisite, placement-score threshold, 12-month validity limit, coordinator scope restriction, panel routing, and file requirements. The request remains bound to Mina Cho’s CHEM 220 approval readiness and routing. The two focus spans are complete factual sentences. The counterfactual changes only the requested term start date from December 2, 2025, to February 3, 2026; this is coherent with the unchanged January 18, 2025 placement-report date and creates no duplicate or contradictory measurement. Neither context embeds a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.\",\"Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.\",\"Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed.\"],\"instructions\":\"Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.\",\"type\":\"score\"}},\"state\":{\"context\":\"CHEM 220 exception handoff for Mina Cho. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The coordinator’s intake inventory lists Mina’s official transcript, official placement report, and CHEM 220 instructor note as received.\",\"Mina Cho's official placement report states January 18, 2025, as its date.\",\"Mina Cho's requested CHEM 220 term starts on December 2, 2025.\",\"The official transcript records a D+ for Mina in CHEM 110.\",\"The official placement report records a score of 82.\",\"The CHEM 220 instructor note reports strong laboratory skills and supports Mina’s enrollment.\",\"Registrar staff verified the transcript and placement report as official and confirmed that the instructor note belongs to this request.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "Mina Cho's official placement report states January 18, 2025, as its date."}, {"path": ["state", "evidence", "2"], "text": "Mina Cho's requested CHEM 220 term starts on December 2, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report states January 18, 2025, as its date.", "negative_left": "Mina Cho's official placement report states January 18, 2025, as its date.", "negative_right": "Mina Cho's requested CHEM 220 term starts on February 3, 2026.", "right": "Mina Cho's requested CHEM 220 term starts on December 2, 2025."}, "verifier_independent_model": false}, "family": "scale-diverse-077-003", "id": "scale-diverse-077-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "CHEM 220 exception handoff for Mina Cho. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The coordinator’s intake inventory lists Mina’s official transcript, official placement report, and CHEM 220 instructor note as received.", "Mina Cho's official placement report states January 18, 2025, as its date.", "Mina Cho's requested CHEM 220 term starts on February 3, 2026.", "The official transcript records a D+ for Mina in CHEM 110.", "The official placement report records a score of 82.", "The CHEM 220 instructor note reports strong laboratory skills and supports Mina’s enrollment.", "Registrar staff verified the transcript and placement report as official and confirmed that the instructor note belongs to this request."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing rule and preserve the Maya Chen–Calculus I readiness decision, including packet completeness, exception recommendation, and routing scope. The two evidence spans are complete factual sentences; the counterfactual changes only Professor Ibarra’s recorded answer from “Yes” to “No,” which remains coherent because neither context otherwise states the note’s substantive position. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar reviewer\",\"text\":\"On September 4, 2026, I reconciled Maya Chen’s Calculus I readiness packet with the document-control ledger. The packet contains her transcript, her official placement result, and Professor Ibarra’s signed instructor note. Each document bears Maya’s student identifier, and the signature was authenticated during intake.\"},{\"speaker\":\"Records analyst\",\"text\":\"The official placement result in Maya’s packet records a Calculus I placement score of 79. No superseding result, missing page, or mismatched student record was found during the reconciliation.\"},{\"speaker\":\"Recording log\",\"text\":\"At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “Yes.”\"},{\"speaker\":\"Interview index\",\"text\":\"The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “Yes.”"}, {"path": ["3", "text"], "text": "The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”"}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “Yes.”", "negative_left": "At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “No.”", "negative_right": "The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”", "right": "The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”"}, "verifier_independent_model": false}, "family": "scale-diverse-078-002", "id": "scale-diverse-078-002-base", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar reviewer", "text": "On September 4, 2026, I reconciled Maya Chen’s Calculus I readiness packet with the document-control ledger. The packet contains her transcript, her official placement result, and Professor Ibarra’s signed instructor note. Each document bears Maya’s student identifier, and the signature was authenticated during intake."}, {"speaker": "Records analyst", "text": "The official placement result in Maya’s packet records a Calculus I placement score of 79. No superseding result, missing page, or mismatched student record was found during the reconciliation."}, {"speaker": "Recording log", "text": "At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “Yes.”"}, {"speaker": "Interview index", "text": "The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”"}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing rule and preserve the Maya Chen–Calculus I readiness decision, including packet completeness, exception recommendation, and routing scope. The two evidence spans are complete factual sentences; the counterfactual changes only Professor Ibarra’s recorded answer from “Yes” to “No,” which remains coherent because neither context otherwise states the note’s substantive position. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar reviewer\",\"text\":\"On September 4, 2026, I reconciled Maya Chen’s Calculus I readiness packet with the document-control ledger. The packet contains her transcript, her official placement result, and Professor Ibarra’s signed instructor note. Each document bears Maya’s student identifier, and the signature was authenticated during intake.\"},{\"speaker\":\"Records analyst\",\"text\":\"The official placement result in Maya’s packet records a Calculus I placement score of 79. No superseding result, missing page, or mismatched student record was found during the reconciliation.\"},{\"speaker\":\"Recording log\",\"text\":\"At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “Yes.”\"},{\"speaker\":\"Interview index\",\"text\":\"The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “Yes.”"}, {"path": ["3", "text"], "text": "The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”"}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “Yes.”", "negative_left": "At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “No.”", "negative_right": "The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”", "right": "The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”"}, "verifier_independent_model": false}, "family": "scale-diverse-078-002", "id": "scale-diverse-078-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar reviewer", "text": "On September 4, 2026, I reconciled Maya Chen’s Calculus I readiness packet with the document-control ledger. The packet contains her transcript, her official placement result, and Professor Ibarra’s signed instructor note. Each document bears Maya’s student identifier, and the signature was authenticated during intake."}, {"speaker": "Records analyst", "text": "The official placement result in Maya’s packet records a Calculus I placement score of 79. No superseding result, missing page, or mismatched student record was found during the reconciliation."}, {"speaker": "Recording log", "text": "At 10:14:32 on September 3, 2026, Professor Ibarra answered the immediately preceding yes-or-no question in a recorded interview with the single word “No.”"}, {"speaker": "Interview index", "text": "The question immediately preceding Professor Ibarra’s 10:14:32 answer in the September 3, 2026 recorded interview was, “Do you support Maya Chen’s enrollment in Calculus I?”"}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full decision criteria, while both contexts retain the original governing course rule without adding exceptions, priorities, or missing-evidence defaults. Both contexts remain bound to Maya Chen’s Calculus I packet and provide facts relevant to completeness, score, instructor support, exception recommendation, and routing. The two focus-evidence spans are complete factual sentences; the form legend is an evidentiary decoding fact rather than a classifier instruction or governing-policy definition. The counterfactual coherently changes only the recorded form response from E17 to N42, with the unchanged legend supporting that change and no duplicate contradictory assertion. Neither context embeds a readiness label, gold answer, classifier output instruction, proposition identifier, or decision rationale; E17 and N42 are ordinary source-document response codes.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Readiness intake specialist\",\"text\":\"Operational handoff for Maya Chen’s Calculus I readiness review: the packet inventory separately lists a transcript, a placement-result report, and Professor Ibarra’s signed instructor note. File checks found each listed document present and readable.\"},{\"speaker\":\"Assessment clerk\",\"text\":\"The placement-result report identifies Maya Chen, names Calculus I, and records a score of 79.\"},{\"speaker\":\"Form audit log\",\"text\":\"At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code E17 from form edition CI-2026-4.\"},{\"speaker\":\"Forms reference\",\"text\":\"The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code E17 from form edition CI-2026-4."}, {"path": ["3", "text"], "text": "The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code E17 from form edition CI-2026-4.", "negative_left": "At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code N42 from form edition CI-2026-4.", "negative_right": "The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”", "right": "The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”"}, "verifier_independent_model": false}, "family": "scale-diverse-078-003", "id": "scale-diverse-078-003-base", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Readiness intake specialist", "text": "Operational handoff for Maya Chen’s Calculus I readiness review: the packet inventory separately lists a transcript, a placement-result report, and Professor Ibarra’s signed instructor note. File checks found each listed document present and readable."}, {"speaker": "Assessment clerk", "text": "The placement-result report identifies Maya Chen, names Calculus I, and records a score of 79."}, {"speaker": "Form audit log", "text": "At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code E17 from form edition CI-2026-4."}, {"speaker": "Forms reference", "text": "The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”"}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full decision criteria, while both contexts retain the original governing course rule without adding exceptions, priorities, or missing-evidence defaults. Both contexts remain bound to Maya Chen’s Calculus I packet and provide facts relevant to completeness, score, instructor support, exception recommendation, and routing. The two focus-evidence spans are complete factual sentences; the form legend is an evidentiary decoding fact rather than a classifier instruction or governing-policy definition. The counterfactual coherently changes only the recorded form response from E17 to N42, with the unchanged legend supporting that change and no duplicate contradictory assertion. Neither context embeds a readiness label, gold answer, classifier output instruction, proposition identifier, or decision rationale; E17 and N42 are ordinary source-document response codes.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Readiness intake specialist\",\"text\":\"Operational handoff for Maya Chen’s Calculus I readiness review: the packet inventory separately lists a transcript, a placement-result report, and Professor Ibarra’s signed instructor note. File checks found each listed document present and readable.\"},{\"speaker\":\"Assessment clerk\",\"text\":\"The placement-result report identifies Maya Chen, names Calculus I, and records a score of 79.\"},{\"speaker\":\"Form audit log\",\"text\":\"At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code E17 from form edition CI-2026-4.\"},{\"speaker\":\"Forms reference\",\"text\":\"The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code E17 from form edition CI-2026-4."}, {"path": ["3", "text"], "text": "The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code E17 from form edition CI-2026-4.", "negative_left": "At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code N42 from form edition CI-2026-4.", "negative_right": "The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”", "right": "The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”"}, "verifier_independent_model": false}, "family": "scale-diverse-078-003", "id": "scale-diverse-078-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Readiness intake specialist", "text": "Operational handoff for Maya Chen’s Calculus I readiness review: the packet inventory separately lists a transcript, a placement-result report, and Professor Ibarra’s signed instructor note. File checks found each listed document present and readable."}, {"speaker": "Assessment clerk", "text": "The placement-result report identifies Maya Chen, names Calculus I, and records a score of 79."}, {"speaker": "Form audit log", "text": "At 14:20 UTC on 8 September 2026, Professor Ibarra's signed Calculus I readiness note for Maya Chen recorded response code N42 from form edition CI-2026-4."}, {"speaker": "Forms reference", "text": "The printed response legend on form edition CI-2026-4 states that E17 means “I endorse the named student's enrollment in the listed course” and N42 means “I do not endorse the named student's enrollment in the listed course.”"}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all decision criteria, while both contexts retain the governing packet, score, instructor-support, exception, and routing rules from the original state without alteration or added defaults. Both contexts remain bound to Maya Chen’s Calculus I readiness packet and the same note, score, documents, and timestamp. The two focus-evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual changes only the note’s support assertion from support to non-support, which is coherent with the unchanged authorship, signature, packet inventory, score, and policy statements and creates no contradictory duplicate assertion within that context. Neither context includes a gold label, answer code, proposition identifier, rule table, label rationale, or output instruction; the included policy terminology is permissible.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Document audit\",\"text\":\"Maya Chen’s Calculus I readiness packet inventory lists her transcript, the official placement-result report, and note CI-407. The placement report records a score of 79.\"},{\"speaker\":\"Timestamp record\",\"text\":\"At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen.\"},{\"speaker\":\"Note transcription\",\"text\":\"Calculus I note CI-407 contains the sentence, “I support this student's enrollment in Calculus I.”\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen."}, {"path": ["2", "text"], "text": "Calculus I note CI-407 contains the sentence, “I support this student's enrollment in Calculus I.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen.", "negative_left": "At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen.", "negative_right": "Calculus I note CI-407 contains the sentence, “I do not support this student's enrollment in Calculus I.”", "right": "Calculus I note CI-407 contains the sentence, “I support this student's enrollment in Calculus I.”"}, "verifier_independent_model": false}, "family": "scale-diverse-078-004", "id": "scale-diverse-078-004-base", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Document audit", "text": "Maya Chen’s Calculus I readiness packet inventory lists her transcript, the official placement-result report, and note CI-407. The placement report records a score of 79."}, {"speaker": "Timestamp record", "text": "At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen."}, {"speaker": "Note transcription", "text": "Calculus I note CI-407 contains the sentence, “I support this student's enrollment in Calculus I.”"}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all decision criteria, while both contexts retain the governing packet, score, instructor-support, exception, and routing rules from the original state without alteration or added defaults. Both contexts remain bound to Maya Chen’s Calculus I readiness packet and the same note, score, documents, and timestamp. The two focus-evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual changes only the note’s support assertion from support to non-support, which is coherent with the unchanged authorship, signature, packet inventory, score, and policy statements and creates no contradictory duplicate assertion within that context. Neither context includes a gold label, answer code, proposition identifier, rule table, label rationale, or output instruction; the included policy terminology is permissible.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Document audit\",\"text\":\"Maya Chen’s Calculus I readiness packet inventory lists her transcript, the official placement-result report, and note CI-407. The placement report records a score of 79.\"},{\"speaker\":\"Timestamp record\",\"text\":\"At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen.\"},{\"speaker\":\"Note transcription\",\"text\":\"Calculus I note CI-407 contains the sentence, “I support this student's enrollment in Calculus I.”\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen."}, {"path": ["2", "text"], "text": "Calculus I note CI-407 contains the sentence, “I support this student's enrollment in Calculus I.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen.", "negative_left": "At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen.", "negative_right": "Calculus I note CI-407 contains the sentence, “I do not support this student's enrollment in Calculus I.”", "right": "Calculus I note CI-407 contains the sentence, “I support this student's enrollment in Calculus I.”"}, "verifier_independent_model": false}, "family": "scale-diverse-078-004", "id": "scale-diverse-078-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Document audit", "text": "Maya Chen’s Calculus I readiness packet inventory lists her transcript, the official placement-result report, and note CI-407. The placement report records a score of 79."}, {"speaker": "Timestamp record", "text": "At 14:20 UTC on 6 August 2026, Professor Ibarra authored and signed Calculus I note CI-407 concerning Maya Chen."}, {"speaker": "Note transcription", "text": "Calculus I note CI-407 contains the sentence, “I do not support this student's enrollment in Calculus I.”"}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing booking, conflict-routing, organizer, capacity, accessibility, ranking, and walk-in policies, while the unchanged questions object preserves the request and criteria. The same R-214/R-219 entities and relevant ledger-versus-attachment paths remain in scope; the two evidence spans are complete factual sentences. Changing R-219’s attachment status from “confirmed” to “pending sync” makes it agree with the unchanged live-ledger measurement rather than creating a contradiction. Neither context includes an answer choice, output instruction, proposition identifier, rule table, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"Today's live-ledger check lists booking R-214 as confirmed for Cedar, for six people, from 2:00–4:00.\"},{\"speaker\":\"Accessibility coordinator\",\"text\":\"Accessibility intake notes that one attendee in R-214's group uses a wheelchair.\"},{\"speaker\":\"System check\",\"text\":\"At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”\"},{\"speaker\":\"Attachment check\",\"text\":\"At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “confirmed.”\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Library booking assistant\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation.\"},{\"speaker\":\"Room monitor\",\"text\":\"During R-214's 2:00–4:00 interval, walk-ins are occupying Cedar. Maple and Birch are each free throughout 2:00–4:00.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["2", "text"], "text": "At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”"}, {"path": ["3", "text"], "text": "At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “confirmed.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”", "negative_left": "At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”", "negative_right": "At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “pending sync.”", "right": "At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “confirmed.”"}, "verifier_independent_model": false}, "family": "scale-diverse-085-004", "id": "scale-diverse-085-004-base", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Booking clerk", "text": "Today's live-ledger check lists booking R-214 as confirmed for Cedar, for six people, from 2:00–4:00."}, {"speaker": "Accessibility coordinator", "text": "Accessibility intake notes that one attendee in R-214's group uses a wheelchair."}, {"speaker": "System check", "text": "At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”"}, {"speaker": "Attachment check", "text": "At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “confirmed.”"}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Library booking assistant", "text": "Cedar has walk-ins, who must yield to a valid reservation."}, {"speaker": "Room monitor", "text": "During R-214's 2:00–4:00 interval, walk-ins are occupying Cedar. Maple and Birch are each free throughout 2:00–4:00."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_valid_cedar_route_conflict"}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing booking, conflict-routing, organizer, capacity, accessibility, ranking, and walk-in policies, while the unchanged questions object preserves the request and criteria. The same R-214/R-219 entities and relevant ledger-versus-attachment paths remain in scope; the two evidence spans are complete factual sentences. Changing R-219’s attachment status from “confirmed” to “pending sync” makes it agree with the unchanged live-ledger measurement rather than creating a contradiction. Neither context includes an answer choice, output instruction, proposition identifier, rule table, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"Today's live-ledger check lists booking R-214 as confirmed for Cedar, for six people, from 2:00–4:00.\"},{\"speaker\":\"Accessibility coordinator\",\"text\":\"Accessibility intake notes that one attendee in R-214's group uses a wheelchair.\"},{\"speaker\":\"System check\",\"text\":\"At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”\"},{\"speaker\":\"Attachment check\",\"text\":\"At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “confirmed.”\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Library booking assistant\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation.\"},{\"speaker\":\"Room monitor\",\"text\":\"During R-214's 2:00–4:00 interval, walk-ins are occupying Cedar. Maple and Birch are each free throughout 2:00–4:00.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["2", "text"], "text": "At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”"}, {"path": ["3", "text"], "text": "At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “confirmed.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”", "negative_left": "At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”", "negative_right": "At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “pending sync.”", "right": "At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “confirmed.”"}, "verifier_independent_model": false}, "family": "scale-diverse-085-004", "id": "scale-diverse-085-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Booking clerk", "text": "Today's live-ledger check lists booking R-214 as confirmed for Cedar, for six people, from 2:00–4:00."}, {"speaker": "Accessibility coordinator", "text": "Accessibility intake notes that one attendee in R-214's group uses a wheelchair."}, {"speaker": "System check", "text": "At 10:15 on 17 September 2026, the live ledger recorded R-219's booking status as “pending sync.”"}, {"speaker": "Attachment check", "text": "At 10:15 on 17 September 2026, the audit attachment for R-219 recorded its booking status as “pending sync.”"}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Library booking assistant", "text": "Cedar has walk-ins, who must yield to a valid reservation."}, {"speaker": "Room monitor", "text": "During R-214's 2:00–4:00 interval, walk-ins are occupying Cedar. Maple and Birch are each free throughout 2:00–4:00."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_reject_pending_booking"}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts retain the original validity, overlap, referral, suitability, and capacity-ranking policies without adding exceptions or defaults. They preserve the relevant entities and booking intervals; the changed submission timestamps and field values are case observations rather than changes to the decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only the uniquely matched record’s populated accessibility value, with no conflicting duplicate assertion. Neither context contains a gold option, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled outcome classification or policy proposition. The focus atom is the factual value of Owen's accessibility field. The base and counter assignments differ only on that focus and are jointly realizable: the field can respectively record wheelchair access required or no wheelchair access required. Policy evidence correctly preserves the substantive rules originating in the original state—validity, overlap priority, referral for missing data, and alternative suitability/ranking—while the unchanged questions automatically preserve their own instructions and option definitions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes active cards and sufficient Cedar capacity for both bookings, Nia's earlier overlapping valid booking, and populated card/accessibility information, so Nia keeps Cedar and no referral is required. Owen requires wheelchair access; Oak is accessible and sufficiently large, while Birch is not accessible, so Oak ranks ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both bookings as valid, Nia as the earlier overlapping valid Cedar claimant, and no missing relevant fields. Because Owen's populated accessibility field has exactly the two stated possible values and 'wheelchair access required' is refuted, it records no wheelchair-access requirement. Both alternatives meet capacity, and Birch is smaller than Oak, so Birch ranks first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The card-status field on Nia's Cedar booking for 2:00–4:00 records that Nia's library card is active."}, {"id": "a2", "statement": "Cedar's capacity is at least Nia's party size of six for Nia's booking from 2:00–4:00."}, {"id": "a3", "statement": "The accessibility field on Nia's Cedar booking for 2:00–4:00 records that wheelchair-accessible space is required."}, {"id": "a4", "statement": "Cedar is wheelchair accessible."}, {"id": "a5", "statement": "The card-status field on Owen's Cedar booking for 3:00–4:00 records that Owen's library card is active."}, {"id": "a6", "statement": "Cedar's capacity is at least Owen's party size of four for Owen's booking from 3:00–4:00."}, {"id": "a7", "statement": "Nia's Cedar booking for 2:00–4:00 overlaps Owen's Cedar booking for 3:00–4:00."}, {"id": "a8", "statement": "Nia submitted her Cedar booking for 2:00–4:00 earlier than Owen submitted his Cedar booking for 3:00–4:00."}, {"id": "a9", "statement": "The accessibility field in the booking record whose confirmation code matches Owen's Cedar confirmation code for 3:00–4:00 contains a recorded value."}, {"id": "a10", "statement": "The populated accessibility field for Owen's Cedar booking for 3:00–4:00 permits exactly the values “wheelchair access required” and “no wheelchair access required”."}, {"id": "a11", "statement": "The accessibility value in the booking record whose confirmation code matches Owen's Cedar confirmation code for 3:00–4:00 is “wheelchair access required”."}, {"id": "a12", "statement": "Birch's capacity is at least Owen's party size of four for 3:00–4:00."}, {"id": "a13", "statement": "Oak's capacity is at least Owen's party size of four for 3:00–4:00."}, {"id": "a14", "statement": "Birch is not wheelchair accessible."}, {"id": "a15", "statement": "Oak is wheelchair accessible."}, {"id": "a16", "statement": "Birch has a smaller capacity than Oak."}, {"id": "a17", "statement": "Nia's and Owen's bookings are the only Cedar bookings involved in the overlap from 3:00–4:00."}], "base_state_json": "\"On 14 May 2026, Nia submitted Cedar for 2:00–4:00 and six people at 1:05 p.m.; Owen submitted Cedar for 3:00–4:00 and four people at 1:17 p.m. Their bookings overlap from 3:00–4:00 and are the only Cedar bookings involved then. Both card-status fields record active library cards. Cedar holds six and is wheelchair accessible; Nia’s accessibility field requires wheelchair-accessible space. At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00. At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “wheelchair access required”. The matched record’s accessibility field is populated and permits exactly “wheelchair access required” and “no wheelchair access required”. From 3:00–4:00, Birch holds four and is not wheelchair accessible; Oak holds eight and is wheelchair accessible. Birch has smaller capacity than Oak. Validity requires an active card and sufficient capacity; confirmation is insufficient. The earlier valid booking wins an overlap. Missing card or accessibility data must go to the branch supervisor. Suitable alternatives meet capacity and required access, then rank by smallest capacity.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00."}, {"path": [], "text": "At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “wheelchair access required”."}], "policy_evidence": [{"path": [], "text": "Validity requires an active card and sufficient capacity; confirmation is insufficient."}, {"path": [], "text": "The earlier valid booking wins an overlap."}, {"path": [], "text": "Missing card or accessibility data must go to the branch supervisor."}, {"path": [], "text": "Suitable alternatives meet capacity and required access, then rank by smallest capacity."}], "rules": [{"justification": "Both bookings have active cards and sufficient Cedar capacity, so both are valid. Nia submitted the earlier overlapping booking and keeps Cedar. All relevant card-status and accessibility fields are populated, so no referral is required. Owen's joined booking record explicitly requires wheelchair access; Oak meets that requirement and Birch does not, so Oak ranks ahead of Birch.", "target": "both_valid_oak_first", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}]}, {"justification": "Both bookings have active cards and sufficient Cedar capacity, so both are valid. Nia submitted the earlier overlapping booking and keeps Cedar. Owen's accessibility field is populated and has exactly two permitted values; because it is explicitly not “wheelchair access required,” it records “no wheelchair access required.” Thus no relevant field is missing and no referral is required. Birch and Oak both meet Owen's capacity and access requirements, and smaller Birch ranks ahead of Oak.", "target": "both_valid_birch_first", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}]}]}, "verified_pair": {"left": "At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00.", "negative_left": "At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00.", "negative_right": "At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “no wheelchair access required”.", "right": "At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “wheelchair access required”."}, "verifier_independent_model": false}, "family": "scale-diverse-086-001", "id": "scale-diverse-086-001-base", "input": {"questions": {"decision": {"criteria": {"booking_assistant_followup": "Treat Nia as valid and Owen as unresolved, let Nia keep Cedar, defer the alternative ranking, and have the booking assistant resolve the missing fields.", "both_valid_birch_first": "Treat both bookings as valid, let Nia keep Cedar, rank Birch ahead of Oak for Owen, and make no referral.", "both_valid_oak_first": "Treat both bookings as valid, let Nia keep Cedar, rank Oak ahead of Birch for Owen, and make no referral.", "none_of_above": "Use when Nia is valid and keeps Cedar, Owen remains unresolved and must be referred to the branch supervisor, and Birch versus Oak cannot yet be ranked because Owen’s accessibility requirement is missing.", "owen_invalid_no_referral": "Treat Nia as valid and Owen as invalid solely because his fields are blank; let Nia keep Cedar and make no referral.", "supervisor_birch_first": "Treat Nia as valid and Owen as unresolved, let Nia keep Cedar, refer Owen to the branch supervisor, but rank Birch ahead of Oak despite the missing accessibility information."}, "instructions": "Choose the one option whose complete resolution follows the stated rules. Verify each booking, resolve the Cedar conflict, route missing information to the correct role, and rank alternatives only when the evidence permits it.", "type": "choice"}}, "state": "On 14 May 2026, Nia submitted Cedar for 2:00–4:00 and six people at 1:05 p.m.; Owen submitted Cedar for 3:00–4:00 and four people at 1:17 p.m. Their bookings overlap from 3:00–4:00 and are the only Cedar bookings involved then. Both card-status fields record active library cards. Cedar holds six and is wheelchair accessible; Nia’s accessibility field requires wheelchair-accessible space. At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00. At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “wheelchair access required”. The matched record’s accessibility field is populated and permits exactly “wheelchair access required” and “no wheelchair access required”. From 3:00–4:00, Birch holds four and is not wheelchair accessible; Oak holds eight and is wheelchair accessible. Birch has smaller capacity than Oak. Validity requires an active card and sufficient capacity; confirmation is insufficient. The earlier valid booking wins an overlap. Missing card or accessibility data must go to the branch supervisor. Suitable alternatives meet capacity and required access, then rank by smallest capacity."}, "method": "c2d", "provenance": {"source_id": "diverse-086", "source_is_synthetic": true, "source_sha256": "8aecd0306b639bb744c12aa14e629e60aebe5d14711a9ed9466e496aee983946", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "both_valid_oak_first"}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts retain the original validity, overlap, referral, suitability, and capacity-ranking policies without adding exceptions or defaults. They preserve the relevant entities and booking intervals; the changed submission timestamps and field values are case observations rather than changes to the decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only the uniquely matched record’s populated accessibility value, with no conflicting duplicate assertion. Neither context contains a gold option, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled outcome classification or policy proposition. The focus atom is the factual value of Owen's accessibility field. The base and counter assignments differ only on that focus and are jointly realizable: the field can respectively record wheelchair access required or no wheelchair access required. Policy evidence correctly preserves the substantive rules originating in the original state—validity, overlap priority, referral for missing data, and alternative suitability/ranking—while the unchanged questions automatically preserve their own instructions and option definitions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes active cards and sufficient Cedar capacity for both bookings, Nia's earlier overlapping valid booking, and populated card/accessibility information, so Nia keeps Cedar and no referral is required. Owen requires wheelchair access; Oak is accessible and sufficiently large, while Birch is not accessible, so Oak ranks ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both bookings as valid, Nia as the earlier overlapping valid Cedar claimant, and no missing relevant fields. Because Owen's populated accessibility field has exactly the two stated possible values and 'wheelchair access required' is refuted, it records no wheelchair-access requirement. Both alternatives meet capacity, and Birch is smaller than Oak, so Birch ranks first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The card-status field on Nia's Cedar booking for 2:00–4:00 records that Nia's library card is active."}, {"id": "a2", "statement": "Cedar's capacity is at least Nia's party size of six for Nia's booking from 2:00–4:00."}, {"id": "a3", "statement": "The accessibility field on Nia's Cedar booking for 2:00–4:00 records that wheelchair-accessible space is required."}, {"id": "a4", "statement": "Cedar is wheelchair accessible."}, {"id": "a5", "statement": "The card-status field on Owen's Cedar booking for 3:00–4:00 records that Owen's library card is active."}, {"id": "a6", "statement": "Cedar's capacity is at least Owen's party size of four for Owen's booking from 3:00–4:00."}, {"id": "a7", "statement": "Nia's Cedar booking for 2:00–4:00 overlaps Owen's Cedar booking for 3:00–4:00."}, {"id": "a8", "statement": "Nia submitted her Cedar booking for 2:00–4:00 earlier than Owen submitted his Cedar booking for 3:00–4:00."}, {"id": "a9", "statement": "The accessibility field in the booking record whose confirmation code matches Owen's Cedar confirmation code for 3:00–4:00 contains a recorded value."}, {"id": "a10", "statement": "The populated accessibility field for Owen's Cedar booking for 3:00–4:00 permits exactly the values “wheelchair access required” and “no wheelchair access required”."}, {"id": "a11", "statement": "The accessibility value in the booking record whose confirmation code matches Owen's Cedar confirmation code for 3:00–4:00 is “wheelchair access required”."}, {"id": "a12", "statement": "Birch's capacity is at least Owen's party size of four for 3:00–4:00."}, {"id": "a13", "statement": "Oak's capacity is at least Owen's party size of four for 3:00–4:00."}, {"id": "a14", "statement": "Birch is not wheelchair accessible."}, {"id": "a15", "statement": "Oak is wheelchair accessible."}, {"id": "a16", "statement": "Birch has a smaller capacity than Oak."}, {"id": "a17", "statement": "Nia's and Owen's bookings are the only Cedar bookings involved in the overlap from 3:00–4:00."}], "base_state_json": "\"On 14 May 2026, Nia submitted Cedar for 2:00–4:00 and six people at 1:05 p.m.; Owen submitted Cedar for 3:00–4:00 and four people at 1:17 p.m. Their bookings overlap from 3:00–4:00 and are the only Cedar bookings involved then. Both card-status fields record active library cards. Cedar holds six and is wheelchair accessible; Nia’s accessibility field requires wheelchair-accessible space. At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00. At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “wheelchair access required”. The matched record’s accessibility field is populated and permits exactly “wheelchair access required” and “no wheelchair access required”. From 3:00–4:00, Birch holds four and is not wheelchair accessible; Oak holds eight and is wheelchair accessible. Birch has smaller capacity than Oak. Validity requires an active card and sufficient capacity; confirmation is insufficient. The earlier valid booking wins an overlap. Missing card or accessibility data must go to the branch supervisor. Suitable alternatives meet capacity and required access, then rank by smallest capacity.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00."}, {"path": [], "text": "At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “wheelchair access required”."}], "policy_evidence": [{"path": [], "text": "Validity requires an active card and sufficient capacity; confirmation is insufficient."}, {"path": [], "text": "The earlier valid booking wins an overlap."}, {"path": [], "text": "Missing card or accessibility data must go to the branch supervisor."}, {"path": [], "text": "Suitable alternatives meet capacity and required access, then rank by smallest capacity."}], "rules": [{"justification": "Both bookings have active cards and sufficient Cedar capacity, so both are valid. Nia submitted the earlier overlapping booking and keeps Cedar. All relevant card-status and accessibility fields are populated, so no referral is required. Owen's joined booking record explicitly requires wheelchair access; Oak meets that requirement and Birch does not, so Oak ranks ahead of Birch.", "target": "both_valid_oak_first", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}]}, {"justification": "Both bookings have active cards and sufficient Cedar capacity, so both are valid. Nia submitted the earlier overlapping booking and keeps Cedar. Owen's accessibility field is populated and has exactly two permitted values; because it is explicitly not “wheelchair access required,” it records “no wheelchair access required.” Thus no relevant field is missing and no referral is required. Birch and Oak both meet Owen's capacity and access requirements, and smaller Birch ranks ahead of Oak.", "target": "both_valid_birch_first", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}]}]}, "verified_pair": {"left": "At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00.", "negative_left": "At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00.", "negative_right": "At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “no wheelchair access required”.", "right": "At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “wheelchair access required”."}, "verifier_independent_model": false}, "family": "scale-diverse-086-001", "id": "scale-diverse-086-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"booking_assistant_followup": "Treat Nia as valid and Owen as unresolved, let Nia keep Cedar, defer the alternative ranking, and have the booking assistant resolve the missing fields.", "both_valid_birch_first": "Treat both bookings as valid, let Nia keep Cedar, rank Birch ahead of Oak for Owen, and make no referral.", "both_valid_oak_first": "Treat both bookings as valid, let Nia keep Cedar, rank Oak ahead of Birch for Owen, and make no referral.", "none_of_above": "Use when Nia is valid and keeps Cedar, Owen remains unresolved and must be referred to the branch supervisor, and Birch versus Oak cannot yet be ranked because Owen’s accessibility requirement is missing.", "owen_invalid_no_referral": "Treat Nia as valid and Owen as invalid solely because his fields are blank; let Nia keep Cedar and make no referral.", "supervisor_birch_first": "Treat Nia as valid and Owen as unresolved, let Nia keep Cedar, refer Owen to the branch supervisor, but rank Birch ahead of Oak despite the missing accessibility information."}, "instructions": "Choose the one option whose complete resolution follows the stated rules. Verify each booking, resolve the Cedar conflict, route missing information to the correct role, and rank alternatives only when the evidence permits it.", "type": "choice"}}, "state": "On 14 May 2026, Nia submitted Cedar for 2:00–4:00 and six people at 1:05 p.m.; Owen submitted Cedar for 3:00–4:00 and four people at 1:17 p.m. Their bookings overlap from 3:00–4:00 and are the only Cedar bookings involved then. Both card-status fields record active library cards. Cedar holds six and is wheelchair accessible; Nia’s accessibility field requires wheelchair-accessible space. At 1:18 p.m. on 14 May 2026, the Cedar reservation system assigned confirmation code CDR-6842 to Owen's booking for 3:00–4:00. At 1:24 p.m. on 14 May 2026, the booking database contained exactly one record with confirmation code CDR-6842, and that record's accessibility value was “no wheelchair access required”. The matched record’s accessibility field is populated and permits exactly “wheelchair access required” and “no wheelchair access required”. From 3:00–4:00, Birch holds four and is not wheelchair accessible; Oak holds eight and is wheelchair accessible. Birch has smaller capacity than Oak. Validity requires an active card and sufficient capacity; confirmation is insufficient. The earlier valid booking wins an overlap. Missing card or accessibility data must go to the branch supervisor. Suitable alternatives meet capacity and required access, then rank by smallest capacity."}, "method": "c2d", "provenance": {"source_id": "diverse-086", "source_is_synthetic": true, "source_sha256": "8aecd0306b639bb744c12aa14e629e60aebe5d14711a9ed9466e496aee983946", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "both_valid_birch_first"}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original validity, conflict-resolution, referral, and alternative-ranking policies without adding exceptions or defaults. The entities, Cedar conflict, 3:00–4:00 overlap, and decision scope remain fixed; only Owen’s reconciled accessibility value changes. The two evidence spans are complete factual sentences. The counterfactual value is compatible with the unchanged active-card, capacity, and room facts; saying both alternatives can accommodate four is consistent with Birch having a smaller total capacity than Oak. Neither context includes an answer option, code, proposition ID, output instruction, or explicit gold resolution.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled outcome classification or policy proposition. The focus atom is the factual value of Owen's accessibility field. The base and counter assignments differ only on that focus and are jointly realizable: the field can respectively record wheelchair access required or no wheelchair access required. Policy evidence correctly preserves the substantive rules originating in the original state—validity, overlap priority, referral for missing data, and alternative suitability/ranking—while the unchanged questions automatically preserve their own instructions and option definitions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes active cards and sufficient Cedar capacity for both bookings, Nia's earlier overlapping valid booking, and populated card/accessibility information, so Nia keeps Cedar and no referral is required. Owen requires wheelchair access; Oak is accessible and sufficiently large, while Birch is not accessible, so Oak ranks ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both bookings as valid, Nia as the earlier overlapping valid Cedar claimant, and no missing relevant fields. Because Owen's populated accessibility field has exactly the two stated possible values and 'wheelchair access required' is refuted, it records no wheelchair-access requirement. Both alternatives meet capacity, and Birch is smaller than Oak, so Birch ranks first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The card-status field on Nia's Cedar booking for 2:00–4:00 records that Nia's library card is active."}, {"id": "a2", "statement": "Cedar's capacity is at least Nia's party size of six for Nia's booking from 2:00–4:00."}, {"id": "a3", "statement": "The accessibility field on Nia's Cedar booking for 2:00–4:00 records that wheelchair-accessible space is required."}, {"id": "a4", "statement": "Cedar is wheelchair accessible."}, {"id": "a5", "statement": "The card-status field on Owen's Cedar booking for 3:00–4:00 records that Owen's library card is active."}, {"id": "a6", "statement": "Cedar's capacity is at least Owen's party size of four for Owen's booking from 3:00–4:00."}, {"id": "a7", "statement": "Nia's Cedar booking for 2:00–4:00 overlaps Owen's Cedar booking for 3:00–4:00."}, {"id": "a8", "statement": "Nia submitted her Cedar booking for 2:00–4:00 earlier than Owen submitted his Cedar booking for 3:00–4:00."}, {"id": "a9", "statement": "The accessibility field in the booking record whose confirmation code matches Owen's Cedar confirmation code for 3:00–4:00 contains a recorded value."}, {"id": "a10", "statement": "The populated accessibility field for Owen's Cedar booking for 3:00–4:00 permits exactly the values “wheelchair access required” and “no wheelchair access required”."}, {"id": "a11", "statement": "The accessibility value in the booking record whose confirmation code matches Owen's Cedar confirmation code for 3:00–4:00 is “wheelchair access required”."}, {"id": "a12", "statement": "Birch's capacity is at least Owen's party size of four for 3:00–4:00."}, {"id": "a13", "statement": "Oak's capacity is at least Owen's party size of four for 3:00–4:00."}, {"id": "a14", "statement": "Birch is not wheelchair accessible."}, {"id": "a15", "statement": "Oak is wheelchair accessible."}, {"id": "a16", "statement": "Birch has a smaller capacity than Oak."}, {"id": "a17", "statement": "Nia's and Owen's bookings are the only Cedar bookings involved in the overlap from 3:00–4:00."}], "base_state_json": "\"Reconciled records show Nia requested Cedar from 2:00–4:00 for six, submitted before Owen, and has an active-card entry plus a requirement for wheelchair-accessible space. Owen requested Cedar from 3:00–4:00 for four and has an active-card entry. Cedar holds six and is wheelchair accessible, making its capacity sufficient for both parties. Their requests overlap from 3:00–4:00 and are the only Cedar bookings involved in that overlap. Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817. The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “wheelchair access required”. That populated field permits exactly two values: “wheelchair access required” and “no wheelchair access required”. During 3:00–4:00, Birch and Oak can each accommodate four; Birch is not wheelchair accessible, Oak is wheelchair accessible, and Birch has smaller capacity than Oak. Validity requires an active card and sufficient capacity; confirmation is insufficient. The earlier valid booking wins an overlap. Missing card or accessibility data must go to the branch supervisor. Suitable alternatives meet capacity and required access, then rank by smallest capacity.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817."}, {"path": [], "text": "The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “wheelchair access required”."}], "policy_evidence": [{"path": [], "text": "Validity requires an active card and sufficient capacity; confirmation is insufficient."}, {"path": [], "text": "The earlier valid booking wins an overlap."}, {"path": [], "text": "Missing card or accessibility data must go to the branch supervisor."}, {"path": [], "text": "Suitable alternatives meet capacity and required access, then rank by smallest capacity."}], "rules": [{"justification": "Both bookings have active cards and sufficient Cedar capacity, so both are valid. Nia submitted the earlier overlapping booking and keeps Cedar. All relevant card-status and accessibility fields are populated, so no referral is required. Owen's joined booking record explicitly requires wheelchair access; Oak meets that requirement and Birch does not, so Oak ranks ahead of Birch.", "target": "both_valid_oak_first", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}]}, {"justification": "Both bookings have active cards and sufficient Cedar capacity, so both are valid. Nia submitted the earlier overlapping booking and keeps Cedar. Owen's accessibility field is populated and has exactly two permitted values; because it is explicitly not “wheelchair access required,” it records “no wheelchair access required.” Thus no relevant field is missing and no referral is required. Birch and Oak both meet Owen's capacity and access requirements, and smaller Birch ranks ahead of Oak.", "target": "both_valid_birch_first", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}]}]}, "verified_pair": {"left": "Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817.", "negative_left": "Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817.", "negative_right": "The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “no wheelchair access required”.", "right": "The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “wheelchair access required”."}, "verifier_independent_model": false}, "family": "scale-diverse-086-002", "id": "scale-diverse-086-002-base", "input": {"questions": {"decision": {"criteria": {"booking_assistant_followup": "Treat Nia as valid and Owen as unresolved, let Nia keep Cedar, defer the alternative ranking, and have the booking assistant resolve the missing fields.", "both_valid_birch_first": "Treat both bookings as valid, let Nia keep Cedar, rank Birch ahead of Oak for Owen, and make no referral.", "both_valid_oak_first": "Treat both bookings as valid, let Nia keep Cedar, rank Oak ahead of Birch for Owen, and make no referral.", "none_of_above": "Use when Nia is valid and keeps Cedar, Owen remains unresolved and must be referred to the branch supervisor, and Birch versus Oak cannot yet be ranked because Owen’s accessibility requirement is missing.", "owen_invalid_no_referral": "Treat Nia as valid and Owen as invalid solely because his fields are blank; let Nia keep Cedar and make no referral.", "supervisor_birch_first": "Treat Nia as valid and Owen as unresolved, let Nia keep Cedar, refer Owen to the branch supervisor, but rank Birch ahead of Oak despite the missing accessibility information."}, "instructions": "Choose the one option whose complete resolution follows the stated rules. Verify each booking, resolve the Cedar conflict, route missing information to the correct role, and rank alternatives only when the evidence permits it.", "type": "choice"}}, "state": "Reconciled records show Nia requested Cedar from 2:00–4:00 for six, submitted before Owen, and has an active-card entry plus a requirement for wheelchair-accessible space. Owen requested Cedar from 3:00–4:00 for four and has an active-card entry. Cedar holds six and is wheelchair accessible, making its capacity sufficient for both parties. Their requests overlap from 3:00–4:00 and are the only Cedar bookings involved in that overlap. Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817. The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “wheelchair access required”. That populated field permits exactly two values: “wheelchair access required” and “no wheelchair access required”. During 3:00–4:00, Birch and Oak can each accommodate four; Birch is not wheelchair accessible, Oak is wheelchair accessible, and Birch has smaller capacity than Oak. Validity requires an active card and sufficient capacity; confirmation is insufficient. The earlier valid booking wins an overlap. Missing card or accessibility data must go to the branch supervisor. Suitable alternatives meet capacity and required access, then rank by smallest capacity."}, "method": "c2d", "provenance": {"source_id": "diverse-086", "source_is_synthetic": true, "source_sha256": "8aecd0306b639bb744c12aa14e629e60aebe5d14711a9ed9466e496aee983946", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "both_valid_oak_first"}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original validity, conflict-resolution, referral, and alternative-ranking policies without adding exceptions or defaults. The entities, Cedar conflict, 3:00–4:00 overlap, and decision scope remain fixed; only Owen’s reconciled accessibility value changes. The two evidence spans are complete factual sentences. The counterfactual value is compatible with the unchanged active-card, capacity, and room facts; saying both alternatives can accommodate four is consistent with Birch having a smaller total capacity than Oak. Neither context includes an answer option, code, proposition ID, output instruction, or explicit gold resolution.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled outcome classification or policy proposition. The focus atom is the factual value of Owen's accessibility field. The base and counter assignments differ only on that focus and are jointly realizable: the field can respectively record wheelchair access required or no wheelchair access required. Policy evidence correctly preserves the substantive rules originating in the original state—validity, overlap priority, referral for missing data, and alternative suitability/ranking—while the unchanged questions automatically preserve their own instructions and option definitions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes active cards and sufficient Cedar capacity for both bookings, Nia's earlier overlapping valid booking, and populated card/accessibility information, so Nia keeps Cedar and no referral is required. Owen requires wheelchair access; Oak is accessible and sufficiently large, while Birch is not accessible, so Oak ranks ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both bookings as valid, Nia as the earlier overlapping valid Cedar claimant, and no missing relevant fields. Because Owen's populated accessibility field has exactly the two stated possible values and 'wheelchair access required' is refuted, it records no wheelchair-access requirement. Both alternatives meet capacity, and Birch is smaller than Oak, so Birch ranks first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The card-status field on Nia's Cedar booking for 2:00–4:00 records that Nia's library card is active."}, {"id": "a2", "statement": "Cedar's capacity is at least Nia's party size of six for Nia's booking from 2:00–4:00."}, {"id": "a3", "statement": "The accessibility field on Nia's Cedar booking for 2:00–4:00 records that wheelchair-accessible space is required."}, {"id": "a4", "statement": "Cedar is wheelchair accessible."}, {"id": "a5", "statement": "The card-status field on Owen's Cedar booking for 3:00–4:00 records that Owen's library card is active."}, {"id": "a6", "statement": "Cedar's capacity is at least Owen's party size of four for Owen's booking from 3:00–4:00."}, {"id": "a7", "statement": "Nia's Cedar booking for 2:00–4:00 overlaps Owen's Cedar booking for 3:00–4:00."}, {"id": "a8", "statement": "Nia submitted her Cedar booking for 2:00–4:00 earlier than Owen submitted his Cedar booking for 3:00–4:00."}, {"id": "a9", "statement": "The accessibility field in the booking record whose confirmation code matches Owen's Cedar confirmation code for 3:00–4:00 contains a recorded value."}, {"id": "a10", "statement": "The populated accessibility field for Owen's Cedar booking for 3:00–4:00 permits exactly the values “wheelchair access required” and “no wheelchair access required”."}, {"id": "a11", "statement": "The accessibility value in the booking record whose confirmation code matches Owen's Cedar confirmation code for 3:00–4:00 is “wheelchair access required”."}, {"id": "a12", "statement": "Birch's capacity is at least Owen's party size of four for 3:00–4:00."}, {"id": "a13", "statement": "Oak's capacity is at least Owen's party size of four for 3:00–4:00."}, {"id": "a14", "statement": "Birch is not wheelchair accessible."}, {"id": "a15", "statement": "Oak is wheelchair accessible."}, {"id": "a16", "statement": "Birch has a smaller capacity than Oak."}, {"id": "a17", "statement": "Nia's and Owen's bookings are the only Cedar bookings involved in the overlap from 3:00–4:00."}], "base_state_json": "\"Reconciled records show Nia requested Cedar from 2:00–4:00 for six, submitted before Owen, and has an active-card entry plus a requirement for wheelchair-accessible space. Owen requested Cedar from 3:00–4:00 for four and has an active-card entry. Cedar holds six and is wheelchair accessible, making its capacity sufficient for both parties. Their requests overlap from 3:00–4:00 and are the only Cedar bookings involved in that overlap. Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817. The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “wheelchair access required”. That populated field permits exactly two values: “wheelchair access required” and “no wheelchair access required”. During 3:00–4:00, Birch and Oak can each accommodate four; Birch is not wheelchair accessible, Oak is wheelchair accessible, and Birch has smaller capacity than Oak. Validity requires an active card and sufficient capacity; confirmation is insufficient. The earlier valid booking wins an overlap. Missing card or accessibility data must go to the branch supervisor. Suitable alternatives meet capacity and required access, then rank by smallest capacity.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817."}, {"path": [], "text": "The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “wheelchair access required”."}], "policy_evidence": [{"path": [], "text": "Validity requires an active card and sufficient capacity; confirmation is insufficient."}, {"path": [], "text": "The earlier valid booking wins an overlap."}, {"path": [], "text": "Missing card or accessibility data must go to the branch supervisor."}, {"path": [], "text": "Suitable alternatives meet capacity and required access, then rank by smallest capacity."}], "rules": [{"justification": "Both bookings have active cards and sufficient Cedar capacity, so both are valid. Nia submitted the earlier overlapping booking and keeps Cedar. All relevant card-status and accessibility fields are populated, so no referral is required. Owen's joined booking record explicitly requires wheelchair access; Oak meets that requirement and Birch does not, so Oak ranks ahead of Birch.", "target": "both_valid_oak_first", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}]}, {"justification": "Both bookings have active cards and sufficient Cedar capacity, so both are valid. Nia submitted the earlier overlapping booking and keeps Cedar. Owen's accessibility field is populated and has exactly two permitted values; because it is explicitly not “wheelchair access required,” it records “no wheelchair access required.” Thus no relevant field is missing and no referral is required. Birch and Oak both meet Owen's capacity and access requirements, and smaller Birch ranks ahead of Oak.", "target": "both_valid_birch_first", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}]}]}, "verified_pair": {"left": "Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817.", "negative_left": "Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817.", "negative_right": "The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “no wheelchair access required”.", "right": "The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “wheelchair access required”."}, "verifier_independent_model": false}, "family": "scale-diverse-086-002", "id": "scale-diverse-086-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"booking_assistant_followup": "Treat Nia as valid and Owen as unresolved, let Nia keep Cedar, defer the alternative ranking, and have the booking assistant resolve the missing fields.", "both_valid_birch_first": "Treat both bookings as valid, let Nia keep Cedar, rank Birch ahead of Oak for Owen, and make no referral.", "both_valid_oak_first": "Treat both bookings as valid, let Nia keep Cedar, rank Oak ahead of Birch for Owen, and make no referral.", "none_of_above": "Use when Nia is valid and keeps Cedar, Owen remains unresolved and must be referred to the branch supervisor, and Birch versus Oak cannot yet be ranked because Owen’s accessibility requirement is missing.", "owen_invalid_no_referral": "Treat Nia as valid and Owen as invalid solely because his fields are blank; let Nia keep Cedar and make no referral.", "supervisor_birch_first": "Treat Nia as valid and Owen as unresolved, let Nia keep Cedar, refer Owen to the branch supervisor, but rank Birch ahead of Oak despite the missing accessibility information."}, "instructions": "Choose the one option whose complete resolution follows the stated rules. Verify each booking, resolve the Cedar conflict, route missing information to the correct role, and rank alternatives only when the evidence permits it.", "type": "choice"}}, "state": "Reconciled records show Nia requested Cedar from 2:00–4:00 for six, submitted before Owen, and has an active-card entry plus a requirement for wheelchair-accessible space. Owen requested Cedar from 3:00–4:00 for four and has an active-card entry. Cedar holds six and is wheelchair accessible, making its capacity sufficient for both parties. Their requests overlap from 3:00–4:00 and are the only Cedar bookings involved in that overlap. Owen's Cedar booking for 3:00–4:00 has the confirmation code CD-5817. The accessibility field in the booking record identified by confirmation code CD-5817 contains the value “no wheelchair access required”. That populated field permits exactly two values: “wheelchair access required” and “no wheelchair access required”. During 3:00–4:00, Birch and Oak can each accommodate four; Birch is not wheelchair accessible, Oak is wheelchair accessible, and Birch has smaller capacity than Oak. Validity requires an active card and sufficient capacity; confirmation is insufficient. The earlier valid booking wins an overlap. Missing card or accessibility data must go to the branch supervisor. Suitable alternatives meet capacity and required access, then rank by smallest capacity."}, "method": "c2d", "provenance": {"source_id": "diverse-086", "source_is_synthetic": true, "source_sha256": "8aecd0306b639bb744c12aa14e629e60aebe5d14711a9ed9466e496aee983946", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "both_valid_birch_first"}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking-validity and alternative-ranking policies and preserve the Priya/Leo, Cedar, 2–3 interval, decision-time, and reassignment bindings. The two evidence spans are complete factual record sentences. The counterfactual changes only Birch’s unused-seat count from 12 to 4, which is coherent with the unchanged accommodation and equipment facts and creates no duplicate conflicting measurement within that context. Neither context includes an answer code, rule table, proposition ID, label rationale, output instruction, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At 09:01 on 17 September 2026, Priya submitted a Cedar request for the 2–3 interval.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At 09:04, Leo submitted a Cedar request for the same interval, requested a display, and reported that his group had no accessibility need.\"},{\"speaker\":\"Policy record\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Facilities record\",\"text\":\"At 09:10, staff confirmed that Maple and Birch could each accommodate Leo’s group and that a display was available in each room.\"},{\"speaker\":\"Reassignment record\",\"text\":\"At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7.\"},{\"speaker\":\"Reassignment record\",\"text\":\"At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 12.\"},{\"speaker\":\"Review clerk\",\"text\":\"At the current decision time, Priya’s request was valid, and every other valid Cedar request overlapping her 2–3 interval had been submitted later. Leo had exactly two no-shows during the preceding 30 days, and no branch supervisor had approved his Cedar request before this decision.\"},{\"speaker\":\"Policy record\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["4", "text"], "text": "At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7."}, {"path": ["5", "text"], "text": "At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 12."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7.", "negative_left": "At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7.", "negative_right": "At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 4.", "right": "At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 12."}, "verifier_independent_model": false}, "family": "scale-diverse-088-001", "id": "scale-diverse-088-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At 09:01 on 17 September 2026, Priya submitted a Cedar request for the 2–3 interval."}, {"speaker": "Booking clerk", "text": "At 09:04, Leo submitted a Cedar request for the same interval, requested a display, and reported that his group had no accessibility need."}, {"speaker": "Policy record", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Facilities record", "text": "At 09:10, staff confirmed that Maple and Birch could each accommodate Leo’s group and that a display was available in each room."}, {"speaker": "Reassignment record", "text": "At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7."}, {"speaker": "Reassignment record", "text": "At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 12."}, {"speaker": "Review clerk", "text": "At the current decision time, Priya’s request was valid, and every other valid Cedar request overlapping her 2–3 interval had been submitted later. Leo had exactly two no-shows during the preceding 30 days, and no branch supervisor had approved his Cedar request before this decision."}, {"speaker": "Policy record", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking-validity and alternative-ranking policies and preserve the Priya/Leo, Cedar, 2–3 interval, decision-time, and reassignment bindings. The two evidence spans are complete factual record sentences. The counterfactual changes only Birch’s unused-seat count from 12 to 4, which is coherent with the unchanged accommodation and equipment facts and creates no duplicate conflicting measurement within that context. Neither context includes an answer code, rule table, proposition ID, label rationale, output instruction, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At 09:01 on 17 September 2026, Priya submitted a Cedar request for the 2–3 interval.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At 09:04, Leo submitted a Cedar request for the same interval, requested a display, and reported that his group had no accessibility need.\"},{\"speaker\":\"Policy record\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Facilities record\",\"text\":\"At 09:10, staff confirmed that Maple and Birch could each accommodate Leo’s group and that a display was available in each room.\"},{\"speaker\":\"Reassignment record\",\"text\":\"At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7.\"},{\"speaker\":\"Reassignment record\",\"text\":\"At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 12.\"},{\"speaker\":\"Review clerk\",\"text\":\"At the current decision time, Priya’s request was valid, and every other valid Cedar request overlapping her 2–3 interval had been submitted later. Leo had exactly two no-shows during the preceding 30 days, and no branch supervisor had approved his Cedar request before this decision.\"},{\"speaker\":\"Policy record\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["4", "text"], "text": "At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7."}, {"path": ["5", "text"], "text": "At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 12."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7.", "negative_left": "At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7.", "negative_right": "At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 4.", "right": "At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 12."}, "verifier_independent_model": false}, "family": "scale-diverse-088-001", "id": "scale-diverse-088-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At 09:01 on 17 September 2026, Priya submitted a Cedar request for the 2–3 interval."}, {"speaker": "Booking clerk", "text": "At 09:04, Leo submitted a Cedar request for the same interval, requested a display, and reported that his group had no accessibility need."}, {"speaker": "Policy record", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Facilities record", "text": "At 09:10, staff confirmed that Maple and Birch could each accommodate Leo’s group and that a display was available in each room."}, {"speaker": "Reassignment record", "text": "At 09:12 on 17 September 2026, the finalized reassignment record lists Maple’s unused-seat count for Leo’s group during the described 2–3 interval as 7."}, {"speaker": "Reassignment record", "text": "At 09:14 on 17 September 2026, the finalized reassignment record lists Birch’s unused-seat count for Leo’s group during the described 2–3 interval as 4."}, {"speaker": "Review clerk", "text": "At the current decision time, Priya’s request was valid, and every other valid Cedar request overlapping her 2–3 interval had been submitted later. Leo had exactly two no-shows during the preceding 30 days, and no branch supervisor had approved his Cedar request before this decision."}, {"speaker": "Policy record", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking-validity and reassignment policies and preserve the relevant entities, rooms, 2–3 interval, decision time, and requested resolution path. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Maple’s unused-seat measurement from 6 to 14; this is coherent with the unchanged statement that Maple can hold Leo’s group and does not duplicate or contradict another Maple measurement. Neither context contains an explicit gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking auditor\",\"text\":\"At the current decision time, Priya’s Cedar request for 2–3 is valid. The submission ledger shows that every other valid Cedar request overlapping that interval was filed after hers. Leo also requested Cedar for the same 2–3 interval.\"},{\"speaker\":\"Compliance reviewer\",\"text\":\"Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his Cedar request before this decision.\"},{\"speaker\":\"Reassignment record\",\"text\":\"Leo’s group has five people, requested a display, and has no accessibility need. Capacity checks confirm that Maple and Birch can each hold the group, and each room has a display available for 2–3.\"},{\"speaker\":\"Booking policy\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Seating-record verifier\",\"text\":\"At the current booking-validity decision time, the verified seating record shows 6 unused seats in Maple after seating Leo’s group for the described 2–3 interval.\"},{\"speaker\":\"Seating-record verifier\",\"text\":\"At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval.\"},{\"speaker\":\"Reassignment policy\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["4", "text"], "text": "At the current booking-validity decision time, the verified seating record shows 6 unused seats in Maple after seating Leo’s group for the described 2–3 interval."}, {"path": ["5", "text"], "text": "At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the current booking-validity decision time, the verified seating record shows 6 unused seats in Maple after seating Leo’s group for the described 2–3 interval.", "negative_left": "At the current booking-validity decision time, the verified seating record shows 14 unused seats in Maple after seating Leo’s group for the described 2–3 interval.", "negative_right": "At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval.", "right": "At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval."}, "verifier_independent_model": false}, "family": "scale-diverse-088-002", "id": "scale-diverse-088-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking auditor", "text": "At the current decision time, Priya’s Cedar request for 2–3 is valid. The submission ledger shows that every other valid Cedar request overlapping that interval was filed after hers. Leo also requested Cedar for the same 2–3 interval."}, {"speaker": "Compliance reviewer", "text": "Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his Cedar request before this decision."}, {"speaker": "Reassignment record", "text": "Leo’s group has five people, requested a display, and has no accessibility need. Capacity checks confirm that Maple and Birch can each hold the group, and each room has a display available for 2–3."}, {"speaker": "Booking policy", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Seating-record verifier", "text": "At the current booking-validity decision time, the verified seating record shows 6 unused seats in Maple after seating Leo’s group for the described 2–3 interval."}, {"speaker": "Seating-record verifier", "text": "At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval."}, {"speaker": "Reassignment policy", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking-validity and reassignment policies and preserve the relevant entities, rooms, 2–3 interval, decision time, and requested resolution path. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Maple’s unused-seat measurement from 6 to 14; this is coherent with the unchanged statement that Maple can hold Leo’s group and does not duplicate or contradict another Maple measurement. Neither context contains an explicit gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking auditor\",\"text\":\"At the current decision time, Priya’s Cedar request for 2–3 is valid. The submission ledger shows that every other valid Cedar request overlapping that interval was filed after hers. Leo also requested Cedar for the same 2–3 interval.\"},{\"speaker\":\"Compliance reviewer\",\"text\":\"Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his Cedar request before this decision.\"},{\"speaker\":\"Reassignment record\",\"text\":\"Leo’s group has five people, requested a display, and has no accessibility need. Capacity checks confirm that Maple and Birch can each hold the group, and each room has a display available for 2–3.\"},{\"speaker\":\"Booking policy\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Seating-record verifier\",\"text\":\"At the current booking-validity decision time, the verified seating record shows 6 unused seats in Maple after seating Leo’s group for the described 2–3 interval.\"},{\"speaker\":\"Seating-record verifier\",\"text\":\"At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval.\"},{\"speaker\":\"Reassignment policy\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["4", "text"], "text": "At the current booking-validity decision time, the verified seating record shows 6 unused seats in Maple after seating Leo’s group for the described 2–3 interval."}, {"path": ["5", "text"], "text": "At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the current booking-validity decision time, the verified seating record shows 6 unused seats in Maple after seating Leo’s group for the described 2–3 interval.", "negative_left": "At the current booking-validity decision time, the verified seating record shows 14 unused seats in Maple after seating Leo’s group for the described 2–3 interval.", "negative_right": "At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval.", "right": "At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval."}, "verifier_independent_model": false}, "family": "scale-diverse-088-002", "id": "scale-diverse-088-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking auditor", "text": "At the current decision time, Priya’s Cedar request for 2–3 is valid. The submission ledger shows that every other valid Cedar request overlapping that interval was filed after hers. Leo also requested Cedar for the same 2–3 interval."}, {"speaker": "Compliance reviewer", "text": "Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his Cedar request before this decision."}, {"speaker": "Reassignment record", "text": "Leo’s group has five people, requested a display, and has no accessibility need. Capacity checks confirm that Maple and Birch can each hold the group, and each room has a display available for 2–3."}, {"speaker": "Booking policy", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Seating-record verifier", "text": "At the current booking-validity decision time, the verified seating record shows 14 unused seats in Maple after seating Leo’s group for the described 2–3 interval."}, {"speaker": "Seating-record verifier", "text": "At the current booking-validity decision time, the verified seating record shows 11 unused seats in Birch after seating Leo’s group for the described 2–3 interval."}, {"speaker": "Reassignment policy", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking-validity and reassignment policies without adding exceptions, priorities, or missing-evidence defaults. The question remains bound to Priya and Leo’s overlapping Cedar requests for 2–3, supervisor routing, and the Maple-versus-Birch ranking. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Birch’s usable-seat count from 24 to 17; this is consistent with the unchanged 14-person group and creates no duplicate contradictory measurement within that context. Neither context states a yes/no gold label, answer code, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking coordinator\",\"text\":\"At the 2026-10-12 1:45 p.m. decision handoff, Priya’s Cedar request for 2–3 was valid. Leo’s Cedar request overlaps that interval. The ledger confirms every other valid Cedar request overlapping it was submitted after Priya’s request.\"},{\"speaker\":\"Compliance reviewer\",\"text\":\"Leo had exactly two no-shows during the 30 days preceding this decision. No branch supervisor had approved his Cedar request before the decision.\"},{\"speaker\":\"Capacity handoff—Maple\",\"text\":\"At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval.\"},{\"speaker\":\"Capacity handoff—Birch\",\"text\":\"At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 24 usable seats for that interval.\"},{\"speaker\":\"Facilities coordinator\",\"text\":\"Leo requested a display for the described 2–3 booking and his group has no accessibility need. Maple and Birch each have a display available for that reassignment.\"},{\"speaker\":\"Booking policy\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Reassignment policy\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval."}, {"path": ["3", "text"], "text": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 24 usable seats for that interval."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval.", "negative_left": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval.", "negative_right": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 17 usable seats for that interval.", "right": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 24 usable seats for that interval."}, "verifier_independent_model": false}, "family": "scale-diverse-088-003", "id": "scale-diverse-088-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking coordinator", "text": "At the 2026-10-12 1:45 p.m. decision handoff, Priya’s Cedar request for 2–3 was valid. Leo’s Cedar request overlaps that interval. The ledger confirms every other valid Cedar request overlapping it was submitted after Priya’s request."}, {"speaker": "Compliance reviewer", "text": "Leo had exactly two no-shows during the 30 days preceding this decision. No branch supervisor had approved his Cedar request before the decision."}, {"speaker": "Capacity handoff—Maple", "text": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval."}, {"speaker": "Capacity handoff—Birch", "text": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 24 usable seats for that interval."}, {"speaker": "Facilities coordinator", "text": "Leo requested a display for the described 2–3 booking and his group has no accessibility need. Maple and Birch each have a display available for that reassignment."}, {"speaker": "Booking policy", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Reassignment policy", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking-validity and reassignment policies without adding exceptions, priorities, or missing-evidence defaults. The question remains bound to Priya and Leo’s overlapping Cedar requests for 2–3, supervisor routing, and the Maple-versus-Birch ranking. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Birch’s usable-seat count from 24 to 17; this is consistent with the unchanged 14-person group and creates no duplicate contradictory measurement within that context. Neither context states a yes/no gold label, answer code, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking coordinator\",\"text\":\"At the 2026-10-12 1:45 p.m. decision handoff, Priya’s Cedar request for 2–3 was valid. Leo’s Cedar request overlaps that interval. The ledger confirms every other valid Cedar request overlapping it was submitted after Priya’s request.\"},{\"speaker\":\"Compliance reviewer\",\"text\":\"Leo had exactly two no-shows during the 30 days preceding this decision. No branch supervisor had approved his Cedar request before the decision.\"},{\"speaker\":\"Capacity handoff—Maple\",\"text\":\"At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval.\"},{\"speaker\":\"Capacity handoff—Birch\",\"text\":\"At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 24 usable seats for that interval.\"},{\"speaker\":\"Facilities coordinator\",\"text\":\"Leo requested a display for the described 2–3 booking and his group has no accessibility need. Maple and Birch each have a display available for that reassignment.\"},{\"speaker\":\"Booking policy\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Reassignment policy\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval."}, {"path": ["3", "text"], "text": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 24 usable seats for that interval."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval.", "negative_left": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval.", "negative_right": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 17 usable seats for that interval.", "right": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 24 usable seats for that interval."}, "verifier_independent_model": false}, "family": "scale-diverse-088-003", "id": "scale-diverse-088-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking coordinator", "text": "At the 2026-10-12 1:45 p.m. decision handoff, Priya’s Cedar request for 2–3 was valid. Leo’s Cedar request overlaps that interval. The ledger confirms every other valid Cedar request overlapping it was submitted after Priya’s request."}, {"speaker": "Compliance reviewer", "text": "Leo had exactly two no-shows during the 30 days preceding this decision. No branch supervisor had approved his Cedar request before the decision."}, {"speaker": "Capacity handoff—Maple", "text": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Maple had 19 usable seats for that interval."}, {"speaker": "Capacity handoff—Birch", "text": "At the 2026-10-12 1:45 p.m. operational handoff, Leo’s group for the described 2–3 reassignment had 14 members, and Birch had 17 usable seats for that interval."}, {"speaker": "Facilities coordinator", "text": "Leo requested a display for the described 2–3 booking and his group has no accessibility need. Maple and Birch each have a display available for that reassignment."}, {"speaker": "Booking policy", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Reassignment policy", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking and reassignment policies when assessed with the unchanged questions object. They preserve the relevant people, Cedar request, Maple/Birch alternatives, described 2–3 interval, and current decision-time bindings. The two evidence spans are complete factual sentences. The counterfactual changes only Birch’s unused-seat count from nine to two; this remains consistent with Birch being large enough for Leo’s group and with the assertion that the two rooms’ unused-seat counts are unequal. Neither context contains an answer code, explicit classifier instruction, rule table, proposition identifier, or stated gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Field note\",\"text\":\"At the current decision time, Priya’s Cedar request for the described 2–3 interval is valid. Every other valid Cedar request overlapping that interval was submitted later than Priya’s request. Leo’s Cedar request overlaps Priya’s during that interval.\"},{\"speaker\":\"Booking policy\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Booking record\",\"text\":\"Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision. No branch supervisor approved his Cedar request before that decision. His group has no accessibility need, and he requested a display for the described 2–3 booking.\"},{\"speaker\":\"Capacity record\",\"text\":\"At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment.\"},{\"speaker\":\"Capacity record\",\"text\":\"At the current booking-validity decision time, Birch has nine unused seats for Leo’s group during the described 2–3 reassignment.\"},{\"speaker\":\"Facilities check\",\"text\":\"For the described 2–3 reassignment, Maple’s and Birch’s capacities are each at least Leo’s group size. A display is available in each room. Their recorded unused-seat counts are unequal.\"},{\"speaker\":\"Reassignment policy\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["3", "text"], "text": "At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment."}, {"path": ["4", "text"], "text": "At the current booking-validity decision time, Birch has nine unused seats for Leo’s group during the described 2–3 reassignment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment.", "negative_left": "At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment.", "negative_right": "At the current booking-validity decision time, Birch has two unused seats for Leo’s group during the described 2–3 reassignment.", "right": "At the current booking-validity decision time, Birch has nine unused seats for Leo’s group during the described 2–3 reassignment."}, "verifier_independent_model": false}, "family": "scale-diverse-088-004", "id": "scale-diverse-088-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Field note", "text": "At the current decision time, Priya’s Cedar request for the described 2–3 interval is valid. Every other valid Cedar request overlapping that interval was submitted later than Priya’s request. Leo’s Cedar request overlaps Priya’s during that interval."}, {"speaker": "Booking policy", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Booking record", "text": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision. No branch supervisor approved his Cedar request before that decision. His group has no accessibility need, and he requested a display for the described 2–3 booking."}, {"speaker": "Capacity record", "text": "At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment."}, {"speaker": "Capacity record", "text": "At the current booking-validity decision time, Birch has nine unused seats for Leo’s group during the described 2–3 reassignment."}, {"speaker": "Facilities check", "text": "For the described 2–3 reassignment, Maple’s and Birch’s capacities are each at least Leo’s group size. A display is available in each room. Their recorded unused-seat counts are unequal."}, {"speaker": "Reassignment policy", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking and reassignment policies when assessed with the unchanged questions object. They preserve the relevant people, Cedar request, Maple/Birch alternatives, described 2–3 interval, and current decision-time bindings. The two evidence spans are complete factual sentences. The counterfactual changes only Birch’s unused-seat count from nine to two; this remains consistent with Birch being large enough for Leo’s group and with the assertion that the two rooms’ unused-seat counts are unequal. Neither context contains an answer code, explicit classifier instruction, rule table, proposition identifier, or stated gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Field note\",\"text\":\"At the current decision time, Priya’s Cedar request for the described 2–3 interval is valid. Every other valid Cedar request overlapping that interval was submitted later than Priya’s request. Leo’s Cedar request overlaps Priya’s during that interval.\"},{\"speaker\":\"Booking policy\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Booking record\",\"text\":\"Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision. No branch supervisor approved his Cedar request before that decision. His group has no accessibility need, and he requested a display for the described 2–3 booking.\"},{\"speaker\":\"Capacity record\",\"text\":\"At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment.\"},{\"speaker\":\"Capacity record\",\"text\":\"At the current booking-validity decision time, Birch has nine unused seats for Leo’s group during the described 2–3 reassignment.\"},{\"speaker\":\"Facilities check\",\"text\":\"For the described 2–3 reassignment, Maple’s and Birch’s capacities are each at least Leo’s group size. A display is available in each room. Their recorded unused-seat counts are unequal.\"},{\"speaker\":\"Reassignment policy\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["3", "text"], "text": "At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment."}, {"path": ["4", "text"], "text": "At the current booking-validity decision time, Birch has nine unused seats for Leo’s group during the described 2–3 reassignment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment.", "negative_left": "At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment.", "negative_right": "At the current booking-validity decision time, Birch has two unused seats for Leo’s group during the described 2–3 reassignment.", "right": "At the current booking-validity decision time, Birch has nine unused seats for Leo’s group during the described 2–3 reassignment."}, "verifier_independent_model": false}, "family": "scale-diverse-088-004", "id": "scale-diverse-088-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Field note", "text": "At the current decision time, Priya’s Cedar request for the described 2–3 interval is valid. Every other valid Cedar request overlapping that interval was submitted later than Priya’s request. Leo’s Cedar request overlaps Priya’s during that interval."}, {"speaker": "Booking policy", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Booking record", "text": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision. No branch supervisor approved his Cedar request before that decision. His group has no accessibility need, and he requested a display for the described 2–3 booking."}, {"speaker": "Capacity record", "text": "At the current booking-validity decision time, Maple has four unused seats for Leo’s group during the described 2–3 reassignment."}, {"speaker": "Capacity record", "text": "At the current booking-validity decision time, Birch has two unused seats for Leo’s group during the described 2–3 reassignment."}, {"speaker": "Facilities check", "text": "For the described 2–3 reassignment, Maple’s and Birch’s capacities are each at least Leo’s group size. A display is available in each room. Their recorded unused-seat counts are unequal."}, {"speaker": "Reassignment policy", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the cancellation, capacity, alternative-ranking, verification, and supervisor-routing policies without adding exceptions or defaults. They preserve the relevant entities and decision scope, and the counterfactual changes only Birch’s recorded rank from 8 to 2; this may represent a policy-incorrect decision but is not internally contradictory. The two evidence spans are complete factual sentences. Neither context includes a score, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation; the universally quantified log atoms are not bundled classifications. A4 is a factual ranking relation rather than a policy statement. The base and counter assignments differ only on A4 and are realizable: an assistant can make all other listed decisions identically while ranking Elm above Birch in one case and failing to do so in the other. Policy evidence correctly cites the substantive booking, capacity, ranking, verification, and escalation rules originating in the original state. Rules and scoring criteria already present in the retained questions object do not need to be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails every level-4 requirement: Mina is validated and assigned Cedar, Omar’s booking is voided, Elm is ranked above Birch, every disputed log is routed to the branch supervisor, and the charger note is not relied upon. The additional log-handling atoms do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "The conjunction has both booking outcomes correct, preserves the required disputed-log routing, and excludes reliance on the charger note, while refuting the required Elm-over-Birch ranking. Thus exactly the ranking secondary requirement is wrong or omitted, which is sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assistant validates Mina’s booking."}, {"id": "A2", "statement": "The assistant assigns Mina’s group to Cedar."}, {"id": "A3", "statement": "The assistant voids Omar’s booking."}, {"id": "A4", "statement": "In the assistant’s alternative-room ranking for Omar, Elm has a numerically lower rank position than Birch."}, {"id": "A5", "statement": "Every disputed log is routed by the assistant to the branch supervisor."}, {"id": "A6", "statement": "Every log routed by the assistant to the branch supervisor is disputed."}, {"id": "A7", "statement": "Every undisputed booking log is verified by the booking assistant."}, {"id": "A8", "statement": "The assistant relies on the charger note when making the booking decision."}], "base_state_json": "\"On 18 September 2026, the assistant validated Mina’s booking and assigned her group to Cedar. The assistant then voided Omar’s booking after recording his late arrival. At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3. At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 8. During the subsequent review, the assistant routed every disputed log to the branch supervisor, while no undisputed log was routed there. Every undisputed booking log was verified by the booking assistant. The decision record confirms that the assistant did not consult or rely on the charger note. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": [], "text": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3."}, {"path": [], "text": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 8."}], "policy_evidence": [{"path": [], "text": "Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed."}, {"path": [], "text": "Booking assistants verify logs; only disputed records go to the branch supervisor."}], "rules": [{"justification": "All required booking outcomes, the Elm-over-Birch ranking, log handling, and the charger-note exclusion are correct, satisfying the fully correct and complete criterion.", "target": "4", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Both bookings and log routing are correct, and the charger note is not used, but the Elm-over-Birch ranking is wrong; this is exactly one incorrect secondary requirement.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3.", "negative_left": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3.", "negative_right": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 2.", "right": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 8."}, "verifier_independent_model": false}, "family": "scale-diverse-089-001", "id": "scale-diverse-089-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Fully incorrect: reverses or ignores the booking rules, such as removing Mina, honoring Omar’s canceled reservation, or assigning a room that cannot hold the group.", "1 — Mostly incorrect: recognizes only one relevant fact but reaches incorrect booking-validity conclusions or provides unsupported room assignment and routing decisions.", "2 — Partly correct: correctly resolves only one booking, or resolves both bookings but gives neither the required alternative ranking nor the correct route for a disputed record.", "3 — Mostly correct: correctly validates Mina’s booking and voids Omar’s booking, but omits or gets wrong exactly one secondary requirement—either ranking Elm above Birch or routing a log dispute to the branch supervisor.", "4 — Fully correct and complete: validates Mina in Cedar, voids Omar’s late booking, ranks Elm above Birch for Omar under the capacity rule, routes any disputed log to the branch supervisor, and does not rely on the charger-note distractor."], "instructions": "Score the library booking assistant’s decision for completeness and correctness using the five ordered levels. Verify both bookings, the alternative-room ranking, and staff routing. Treat unrelated facts as distractors.", "type": "score"}}, "state": "On 18 September 2026, the assistant validated Mina’s booking and assigned her group to Cedar. The assistant then voided Omar’s booking after recording his late arrival. At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3. At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 8. During the subsequent review, the assistant routed every disputed log to the branch supervisor, while no undisputed log was routed there. Every undisputed booking log was verified by the booking assistant. The decision record confirms that the assistant did not consult or rely on the charger note. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor."}, "method": "c2d", "provenance": {"source_id": "diverse-089", "source_is_synthetic": true, "source_sha256": "09465f5cddfe8d4e40880c4b0b843bcad80d4656a8a7c38d24afa69f4f78a0e2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the cancellation, capacity, alternative-ranking, verification, and supervisor-routing policies without adding exceptions or defaults. They preserve the relevant entities and decision scope, and the counterfactual changes only Birch’s recorded rank from 8 to 2; this may represent a policy-incorrect decision but is not internally contradictory. The two evidence spans are complete factual sentences. Neither context includes a score, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation; the universally quantified log atoms are not bundled classifications. A4 is a factual ranking relation rather than a policy statement. The base and counter assignments differ only on A4 and are realizable: an assistant can make all other listed decisions identically while ranking Elm above Birch in one case and failing to do so in the other. Policy evidence correctly cites the substantive booking, capacity, ranking, verification, and escalation rules originating in the original state. Rules and scoring criteria already present in the retained questions object do not need to be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails every level-4 requirement: Mina is validated and assigned Cedar, Omar’s booking is voided, Elm is ranked above Birch, every disputed log is routed to the branch supervisor, and the charger note is not relied upon. The additional log-handling atoms do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "The conjunction has both booking outcomes correct, preserves the required disputed-log routing, and excludes reliance on the charger note, while refuting the required Elm-over-Birch ranking. Thus exactly the ranking secondary requirement is wrong or omitted, which is sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assistant validates Mina’s booking."}, {"id": "A2", "statement": "The assistant assigns Mina’s group to Cedar."}, {"id": "A3", "statement": "The assistant voids Omar’s booking."}, {"id": "A4", "statement": "In the assistant’s alternative-room ranking for Omar, Elm has a numerically lower rank position than Birch."}, {"id": "A5", "statement": "Every disputed log is routed by the assistant to the branch supervisor."}, {"id": "A6", "statement": "Every log routed by the assistant to the branch supervisor is disputed."}, {"id": "A7", "statement": "Every undisputed booking log is verified by the booking assistant."}, {"id": "A8", "statement": "The assistant relies on the charger note when making the booking decision."}], "base_state_json": "\"On 18 September 2026, the assistant validated Mina’s booking and assigned her group to Cedar. The assistant then voided Omar’s booking after recording his late arrival. At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3. At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 8. During the subsequent review, the assistant routed every disputed log to the branch supervisor, while no undisputed log was routed there. Every undisputed booking log was verified by the booking assistant. The decision record confirms that the assistant did not consult or rely on the charger note. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": [], "text": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3."}, {"path": [], "text": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 8."}], "policy_evidence": [{"path": [], "text": "Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed."}, {"path": [], "text": "Booking assistants verify logs; only disputed records go to the branch supervisor."}], "rules": [{"justification": "All required booking outcomes, the Elm-over-Birch ranking, log handling, and the charger-note exclusion are correct, satisfying the fully correct and complete criterion.", "target": "4", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Both bookings and log routing are correct, and the charger note is not used, but the Elm-over-Birch ranking is wrong; this is exactly one incorrect secondary requirement.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3.", "negative_left": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3.", "negative_right": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 2.", "right": "At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 8."}, "verifier_independent_model": false}, "family": "scale-diverse-089-001", "id": "scale-diverse-089-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Fully incorrect: reverses or ignores the booking rules, such as removing Mina, honoring Omar’s canceled reservation, or assigning a room that cannot hold the group.", "1 — Mostly incorrect: recognizes only one relevant fact but reaches incorrect booking-validity conclusions or provides unsupported room assignment and routing decisions.", "2 — Partly correct: correctly resolves only one booking, or resolves both bookings but gives neither the required alternative ranking nor the correct route for a disputed record.", "3 — Mostly correct: correctly validates Mina’s booking and voids Omar’s booking, but omits or gets wrong exactly one secondary requirement—either ranking Elm above Birch or routing a log dispute to the branch supervisor.", "4 — Fully correct and complete: validates Mina in Cedar, voids Omar’s late booking, ranks Elm above Birch for Omar under the capacity rule, routes any disputed log to the branch supervisor, and does not rely on the charger-note distractor."], "instructions": "Score the library booking assistant’s decision for completeness and correctness using the five ordered levels. Verify both bookings, the alternative-room ranking, and staff routing. Treat unrelated facts as distractors.", "type": "score"}}, "state": "On 18 September 2026, the assistant validated Mina’s booking and assigned her group to Cedar. The assistant then voided Omar’s booking after recording his late arrival. At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Elm at rank position 3. At 14:07 on 18 September 2026, the assistant’s finalized alternative-room ranking for Omar recorded Birch at rank position 2. During the subsequent review, the assistant routed every disputed log to the branch supervisor, while no undisputed log was routed there. Every undisputed booking log was verified by the booking assistant. The decision record confirms that the assistant did not consult or rely on the charger note. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor."}, "method": "c2d", "provenance": {"source_id": "diverse-089", "source_is_synthetic": true, "source_sha256": "09465f5cddfe8d4e40880c4b0b843bcad80d4656a8a7c38d24afa69f4f78a0e2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing booking, capacity, ranking, verification, and dispute-routing policies, while the unchanged questions preserve the full scoring criteria. Omar, the alternative-room ranking, the September 17, 2026 timestamp, and the relevant decision scope remain bound consistently. The focus evidence contains exactly two complete factual sentences. The counterfactual changes only Birch’s rank from 7 to 2; Elm at rank 3 and Birch at rank 2 are distinct, coherent ranking assertions, even though the changed ranking may affect the eventual score. Neither context includes a score, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation; the universally quantified log atoms are not bundled classifications. A4 is a factual ranking relation rather than a policy statement. The base and counter assignments differ only on A4 and are realizable: an assistant can make all other listed decisions identically while ranking Elm above Birch in one case and failing to do so in the other. Policy evidence correctly cites the substantive booking, capacity, ranking, verification, and escalation rules originating in the original state. Rules and scoring criteria already present in the retained questions object do not need to be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails every level-4 requirement: Mina is validated and assigned Cedar, Omar’s booking is voided, Elm is ranked above Birch, every disputed log is routed to the branch supervisor, and the charger note is not relied upon. The additional log-handling atoms do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "The conjunction has both booking outcomes correct, preserves the required disputed-log routing, and excludes reliance on the charger note, while refuting the required Elm-over-Birch ranking. Thus exactly the ranking secondary requirement is wrong or omitted, which is sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assistant validates Mina’s booking."}, {"id": "A2", "statement": "The assistant assigns Mina’s group to Cedar."}, {"id": "A3", "statement": "The assistant voids Omar’s booking."}, {"id": "A4", "statement": "In the assistant’s alternative-room ranking for Omar, Elm has a numerically lower rank position than Birch."}, {"id": "A5", "statement": "Every disputed log is routed by the assistant to the branch supervisor."}, {"id": "A6", "statement": "Every log routed by the assistant to the branch supervisor is disputed."}, {"id": "A7", "statement": "Every undisputed booking log is verified by the booking assistant."}, {"id": "A8", "statement": "The assistant relies on the charger note when making the booking decision."}], "base_state_json": "\"Field note, September 17, 2026: Mina reserved Cedar from 2:00 to 4:00 for her five-person group. Cedar holds six people, and Mina checked in at 2:07. Omar reserved a room beginning at 2:00 but arrived at 2:18. The assistant validated Mina’s booking and assigned her group to Cedar. The assistant voided Omar’s booking. The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3. The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 7. Every disputed log is routed by the assistant to the branch supervisor, and every log routed there is disputed. The booking assistant verifies every undisputed booking log. Omar’s charger note was excluded from the decision materials and was not relied upon. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": [], "text": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3."}, {"path": [], "text": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 7."}], "policy_evidence": [{"path": [], "text": "Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed."}, {"path": [], "text": "Booking assistants verify logs; only disputed records go to the branch supervisor."}], "rules": [{"justification": "All required booking outcomes, the Elm-over-Birch ranking, log handling, and the charger-note exclusion are correct, satisfying the fully correct and complete criterion.", "target": "4", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Both bookings and log routing are correct, and the charger note is not used, but the Elm-over-Birch ranking is wrong; this is exactly one incorrect secondary requirement.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3.", "negative_left": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3.", "negative_right": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 2.", "right": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 7."}, "verifier_independent_model": false}, "family": "scale-diverse-089-004", "id": "scale-diverse-089-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Fully incorrect: reverses or ignores the booking rules, such as removing Mina, honoring Omar’s canceled reservation, or assigning a room that cannot hold the group.", "1 — Mostly incorrect: recognizes only one relevant fact but reaches incorrect booking-validity conclusions or provides unsupported room assignment and routing decisions.", "2 — Partly correct: correctly resolves only one booking, or resolves both bookings but gives neither the required alternative ranking nor the correct route for a disputed record.", "3 — Mostly correct: correctly validates Mina’s booking and voids Omar’s booking, but omits or gets wrong exactly one secondary requirement—either ranking Elm above Birch or routing a log dispute to the branch supervisor.", "4 — Fully correct and complete: validates Mina in Cedar, voids Omar’s late booking, ranks Elm above Birch for Omar under the capacity rule, routes any disputed log to the branch supervisor, and does not rely on the charger-note distractor."], "instructions": "Score the library booking assistant’s decision for completeness and correctness using the five ordered levels. Verify both bookings, the alternative-room ranking, and staff routing. Treat unrelated facts as distractors.", "type": "score"}}, "state": "Field note, September 17, 2026: Mina reserved Cedar from 2:00 to 4:00 for her five-person group. Cedar holds six people, and Mina checked in at 2:07. Omar reserved a room beginning at 2:00 but arrived at 2:18. The assistant validated Mina’s booking and assigned her group to Cedar. The assistant voided Omar’s booking. The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3. The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 7. Every disputed log is routed by the assistant to the branch supervisor, and every log routed there is disputed. The booking assistant verifies every undisputed booking log. Omar’s charger note was excluded from the decision materials and was not relied upon. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor."}, "method": "c2d", "provenance": {"source_id": "diverse-089", "source_is_synthetic": true, "source_sha256": "09465f5cddfe8d4e40880c4b0b843bcad80d4656a8a7c38d24afa69f4f78a0e2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing booking, capacity, ranking, verification, and dispute-routing policies, while the unchanged questions preserve the full scoring criteria. Omar, the alternative-room ranking, the September 17, 2026 timestamp, and the relevant decision scope remain bound consistently. The focus evidence contains exactly two complete factual sentences. The counterfactual changes only Birch’s rank from 7 to 2; Elm at rank 3 and Birch at rank 2 are distinct, coherent ranking assertions, even though the changed ranking may affect the eventual score. Neither context includes a score, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation; the universally quantified log atoms are not bundled classifications. A4 is a factual ranking relation rather than a policy statement. The base and counter assignments differ only on A4 and are realizable: an assistant can make all other listed decisions identically while ranking Elm above Birch in one case and failing to do so in the other. Policy evidence correctly cites the substantive booking, capacity, ranking, verification, and escalation rules originating in the original state. Rules and scoring criteria already present in the retained questions object do not need to be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails every level-4 requirement: Mina is validated and assigned Cedar, Omar’s booking is voided, Elm is ranked above Birch, every disputed log is routed to the branch supervisor, and the charger note is not relied upon. The additional log-handling atoms do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "The conjunction has both booking outcomes correct, preserves the required disputed-log routing, and excludes reliance on the charger note, while refuting the required Elm-over-Birch ranking. Thus exactly the ranking secondary requirement is wrong or omitted, which is sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assistant validates Mina’s booking."}, {"id": "A2", "statement": "The assistant assigns Mina’s group to Cedar."}, {"id": "A3", "statement": "The assistant voids Omar’s booking."}, {"id": "A4", "statement": "In the assistant’s alternative-room ranking for Omar, Elm has a numerically lower rank position than Birch."}, {"id": "A5", "statement": "Every disputed log is routed by the assistant to the branch supervisor."}, {"id": "A6", "statement": "Every log routed by the assistant to the branch supervisor is disputed."}, {"id": "A7", "statement": "Every undisputed booking log is verified by the booking assistant."}, {"id": "A8", "statement": "The assistant relies on the charger note when making the booking decision."}], "base_state_json": "\"Field note, September 17, 2026: Mina reserved Cedar from 2:00 to 4:00 for her five-person group. Cedar holds six people, and Mina checked in at 2:07. Omar reserved a room beginning at 2:00 but arrived at 2:18. The assistant validated Mina’s booking and assigned her group to Cedar. The assistant voided Omar’s booking. The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3. The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 7. Every disputed log is routed by the assistant to the branch supervisor, and every log routed there is disputed. The booking assistant verifies every undisputed booking log. Omar’s charger note was excluded from the decision materials and was not relied upon. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": [], "text": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3."}, {"path": [], "text": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 7."}], "policy_evidence": [{"path": [], "text": "Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed."}, {"path": [], "text": "Booking assistants verify logs; only disputed records go to the branch supervisor."}], "rules": [{"justification": "All required booking outcomes, the Elm-over-Birch ranking, log handling, and the charger-note exclusion are correct, satisfying the fully correct and complete criterion.", "target": "4", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Both bookings and log routing are correct, and the charger note is not used, but the Elm-over-Birch ranking is wrong; this is exactly one incorrect secondary requirement.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3.", "negative_left": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3.", "negative_right": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 2.", "right": "The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 7."}, "verifier_independent_model": false}, "family": "scale-diverse-089-004", "id": "scale-diverse-089-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Fully incorrect: reverses or ignores the booking rules, such as removing Mina, honoring Omar’s canceled reservation, or assigning a room that cannot hold the group.", "1 — Mostly incorrect: recognizes only one relevant fact but reaches incorrect booking-validity conclusions or provides unsupported room assignment and routing decisions.", "2 — Partly correct: correctly resolves only one booking, or resolves both bookings but gives neither the required alternative ranking nor the correct route for a disputed record.", "3 — Mostly correct: correctly validates Mina’s booking and voids Omar’s booking, but omits or gets wrong exactly one secondary requirement—either ranking Elm above Birch or routing a log dispute to the branch supervisor.", "4 — Fully correct and complete: validates Mina in Cedar, voids Omar’s late booking, ranks Elm above Birch for Omar under the capacity rule, routes any disputed log to the branch supervisor, and does not rely on the charger-note distractor."], "instructions": "Score the library booking assistant’s decision for completeness and correctness using the five ordered levels. Verify both bookings, the alternative-room ranking, and staff routing. Treat unrelated facts as distractors.", "type": "score"}}, "state": "Field note, September 17, 2026: Mina reserved Cedar from 2:00 to 4:00 for her five-person group. Cedar holds six people, and Mina checked in at 2:07. Omar reserved a room beginning at 2:00 but arrived at 2:18. The assistant validated Mina’s booking and assigned her group to Cedar. The assistant voided Omar’s booking. The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Elm rank position 3. The assistant's alternative-room ranking for Omar, finalized at 14:20 UTC on September 17, 2026, assigns Birch rank position 2. Every disputed log is routed by the assistant to the branch supervisor, and every log routed there is disputed. The booking assistant verifies every undisputed booking log. Omar’s charger note was excluded from the decision materials and was not relied upon. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor."}, "method": "c2d", "provenance": {"source_id": "diverse-089", "source_is_synthetic": true, "source_sha256": "09465f5cddfe8d4e40880c4b0b843bcad80d4656a8a7c38d24afa69f4f78a0e2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the booking rules, request scope, entities, Oak booking, and Tuesday 14:00–16:00 binding. The two focus spans are complete factual sentences describing separate record displays. The counterfactual coherently introduces a discrepancy between two record sources; the statement that reconciliation found no other timestamp conflict expressly allows this focal conflict, and Ada can still precede Ben under either displayed time. Neither context embeds an answer code, routing result, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"active_accessibility_breach": "refuted", "active_safety_breach": "refuted", "ada_attendance_within_oak_capacity": "supported", "ada_precedes_ben": "supported", "ada_timestamp_match": "supported", "ada_update_before_use": "supported", "ambiguous_applicable_rule": "refuted", "deliberate_rule_refusal": "refuted", "escalating_confrontation": "refuted", "juniper_capacity_sufficient": "supported", "juniper_first_ranked": "supported", "juniper_full_coverage": "supported", "missing_accessibility_fact": "refuted", "missing_attendance_fact": "refuted", "missing_timing_fact": "refuted", "oak_bookings_overlap": "supported", "other_reservation_record_contradiction": "refuted", "other_timestamp_contradiction": "refuted", "requested_policy_exception": "refuted", "room_data_contradiction": "refuted"}, "full_context_fact_states": {"base": {"active_accessibility_breach": "refuted", "active_safety_breach": "refuted", "ada_attendance_within_oak_capacity": "supported", "ada_precedes_ben": "supported", "ada_timestamp_match": "supported", "ada_update_before_use": "supported", "ambiguous_applicable_rule": "refuted", "deliberate_rule_refusal": "refuted", "escalating_confrontation": "refuted", "juniper_capacity_sufficient": "supported", "juniper_first_ranked": "supported", "juniper_full_coverage": "supported", "missing_accessibility_fact": "refuted", "missing_attendance_fact": "refuted", "missing_timing_fact": "refuted", "oak_bookings_overlap": "supported", "other_reservation_record_contradiction": "refuted", "other_timestamp_contradiction": "refuted", "requested_policy_exception": "refuted", "room_data_contradiction": "refuted"}, "counterfactual": {"active_accessibility_breach": "refuted", "active_safety_breach": "refuted", "ada_attendance_within_oak_capacity": "supported", "ada_precedes_ben": "supported", "ada_timestamp_match": "refuted", "ada_update_before_use": "supported", "ambiguous_applicable_rule": "refuted", "deliberate_rule_refusal": "refuted", "escalating_confrontation": "refuted", "juniper_capacity_sufficient": "supported", "juniper_first_ranked": "supported", "juniper_full_coverage": "supported", "missing_accessibility_fact": "refuted", "missing_attendance_fact": "refuted", "missing_timing_fact": "refuted", "oak_bookings_overlap": "supported", "other_reservation_record_contradiction": "refuted", "other_timestamp_contradiction": "refuted", "requested_policy_exception": "refuted", "room_data_contradiction": "refuted"}, "remove_left": {"ada_timestamp_match": "unknown"}, "remove_right": {"ada_timestamp_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"ada_timestamp_match": "unknown"}, "negative_pair": {"ada_timestamp_match": "refuted"}, "negative_sentence": {"ada_timestamp_match": "unknown"}, "positive_pair": {"ada_timestamp_match": "supported"}, "right": {"ada_timestamp_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual condition or relationship; the ranking atom is a factual ranking rather than a routing classification. The focus is the factual equality of two timestamps. The base and counter assignments can differ only in that equality: the counter can contain unequal record and transaction-log timestamps while Ada still precedes Ben according to the operative booking record. Policy evidence correctly cites the substantive rules originating in the original state, while the routing criteria and lowest-applicable-level instruction remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a routine resolution: Ada has the earlier overlapping reservation, her timely attendance update fits Oak, and Juniper is a sufficient, fully available, first-ranked alternative for Ben. It also explicitly refutes every stated trigger for clarification, record audit, supervisory judgment, or immediate intervention.", "rule_index": 0, "sound": true}, {"reason": "Refuting equality between two submission timestamps for the same Ada reservation entails contradictory timestamp records and therefore the level-2 audit trigger. The conjunction excludes lower level 1 and the supervisory/intervention outcomes, while retaining enough facts for the underlying booking and alternative analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ada_timestamp_match", "statement": "The submission timestamp for Ada’s Tuesday 14:00–16:00 Oak reservation in the booking record equals the submission timestamp for that same reservation in the system transaction log."}, {"id": "other_timestamp_contradiction", "statement": "A contradictory timestamp exists among the booking-system records relevant to Ada’s or Ben’s Tuesday 14:00–16:00 booking, excluding the two timestamp entries compared in ada_timestamp_match."}, {"id": "room_data_contradiction", "statement": "A contradiction exists among the recorded room data relevant to Ada’s or Ben’s Tuesday 14:00–16:00 booking or Ben’s listed alternatives."}, {"id": "other_reservation_record_contradiction", "statement": "A contradiction other than a timestamp or room-data contradiction exists among the reservation records relevant to Ada’s or Ben’s Tuesday 14:00–16:00 booking."}, {"id": "missing_attendance_fact", "statement": "A student group organizer must supply an additional attendance fact before the booking assistant can resolve the Tuesday 14:00–16:00 case."}, {"id": "missing_timing_fact", "statement": "A student group organizer must supply an additional timing fact before the booking assistant can resolve the Tuesday 14:00–16:00 case."}, {"id": "missing_accessibility_fact", "statement": "A student group organizer must supply an additional accessibility fact before the booking assistant can resolve the Tuesday 14:00–16:00 case."}, {"id": "ambiguous_applicable_rule", "statement": "An applicable rule governing the Tuesday 14:00–16:00 case is ambiguous and requires branch-supervisor judgment."}, {"id": "requested_policy_exception", "statement": "Ada or Ben requests a policy exception for the Tuesday 14:00–16:00 case."}, {"id": "active_safety_breach", "statement": "An active safety breach exists in the Tuesday 14:00–16:00 case."}, {"id": "active_accessibility_breach", "statement": "An active accessibility breach exists in the Tuesday 14:00–16:00 case."}, {"id": "deliberate_rule_refusal", "statement": "Ada or Ben deliberately refuses an applicable rule in the Tuesday 14:00–16:00 case."}, {"id": "escalating_confrontation", "statement": "The confrontation between Ada and Ben is escalating in the Tuesday 14:00–16:00 case."}, {"id": "oak_bookings_overlap", "statement": "Ada’s and Ben’s Oak reservations cover the same Tuesday 14:00–16:00 period."}, {"id": "ada_precedes_ben", "statement": "Ada’s Oak reservation was submitted before Ben’s Oak reservation."}, {"id": "ada_attendance_within_oak_capacity", "statement": "Ada’s updated attendance for Tuesday 14:00–16:00 does not exceed Oak’s capacity."}, {"id": "ada_update_before_use", "statement": "Ada requested her attendance update before the Tuesday 14:00 start of room use."}, {"id": "juniper_full_coverage", "statement": "Juniper is free throughout Ben’s Tuesday 14:00–16:00 requested period."}, {"id": "juniper_capacity_sufficient", "statement": "Juniper’s capacity is at least Ben’s attendance for Tuesday 14:00–16:00."}, {"id": "juniper_first_ranked", "statement": "Juniper ranks first among Ben’s alternatives under full coverage followed by least spare capacity."}], "base_state_json": "{\"context\":\"At Westmere Library, student organizers Ada and Ben dispute step-free Oak for Tuesday 14:00–16:00; the library booking assistant applies rules, and exceptions go to the branch supervisor.\",\"evidence\":[\"The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.\",\"The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.\",\"Rules: earliest valid reservation wins; attendance may be updated before use if capacity allows; overlaps are invalid. The assistant may make routine updates.\",\"For Ben, Juniper holds 6 and is free throughout; Maple holds 10 and is free throughout; Cedar holds 6 but is free only 15:00–16:00. Rank full coverage first, then least spare capacity.\",\"Both Oak reservations cover Tuesday 14:00–16:00, and Ada’s was submitted before Ben’s. Ada updated attendance to 8 at 13:30; Oak holds 8. Ben lists 5 attendees.\",\"Reconciliation found no other timestamp conflict, inconsistent room data, or other reservation-record discrepancy.\",\"The organizers supplied all needed attendance, timing, and accessibility facts. Oak and Juniper meet the stated accessibility needs, with no active safety or accessibility breach.\",\"The applicable rules are unambiguous. Neither organizer requests an exception or deliberately refuses a rule, and their calm discussion is not escalating.\"],\"request\":\"Decide each booking’s validity, rank Ben’s alternatives, and select the required routing level.\"}", "base_states": [{"atom_id": "ada_timestamp_match", "state": "supported"}, {"atom_id": "other_timestamp_contradiction", "state": "refuted"}, {"atom_id": "room_data_contradiction", "state": "refuted"}, {"atom_id": "other_reservation_record_contradiction", "state": "refuted"}, {"atom_id": "missing_attendance_fact", "state": "refuted"}, {"atom_id": "missing_timing_fact", "state": "refuted"}, {"atom_id": "missing_accessibility_fact", "state": "refuted"}, {"atom_id": "ambiguous_applicable_rule", "state": "refuted"}, {"atom_id": "requested_policy_exception", "state": "refuted"}, {"atom_id": "active_safety_breach", "state": "refuted"}, {"atom_id": "active_accessibility_breach", "state": "refuted"}, {"atom_id": "deliberate_rule_refusal", "state": "refuted"}, {"atom_id": "escalating_confrontation", "state": "refuted"}, {"atom_id": "oak_bookings_overlap", "state": "supported"}, {"atom_id": "ada_precedes_ben", "state": "supported"}, {"atom_id": "ada_attendance_within_oak_capacity", "state": "supported"}, {"atom_id": "ada_update_before_use", "state": "supported"}, {"atom_id": "juniper_full_coverage", "state": "supported"}, {"atom_id": "juniper_capacity_sufficient", "state": "supported"}, {"atom_id": "juniper_first_ranked", "state": "supported"}], "counter_states": [{"atom_id": "ada_timestamp_match", "state": "refuted"}, {"atom_id": "other_timestamp_contradiction", "state": "refuted"}, {"atom_id": "room_data_contradiction", "state": "refuted"}, {"atom_id": "other_reservation_record_contradiction", "state": "refuted"}, {"atom_id": "missing_attendance_fact", "state": "refuted"}, {"atom_id": "missing_timing_fact", "state": "refuted"}, {"atom_id": "missing_accessibility_fact", "state": "refuted"}, {"atom_id": "ambiguous_applicable_rule", "state": "refuted"}, {"atom_id": "requested_policy_exception", "state": "refuted"}, {"atom_id": "active_safety_breach", "state": "refuted"}, {"atom_id": "active_accessibility_breach", "state": "refuted"}, {"atom_id": "deliberate_rule_refusal", "state": "refuted"}, {"atom_id": "escalating_confrontation", "state": "refuted"}, {"atom_id": "oak_bookings_overlap", "state": "supported"}, {"atom_id": "ada_precedes_ben", "state": "supported"}, {"atom_id": "ada_attendance_within_oak_capacity", "state": "supported"}, {"atom_id": "ada_update_before_use", "state": "supported"}, {"atom_id": "juniper_full_coverage", "state": "supported"}, {"atom_id": "juniper_capacity_sufficient", "state": "supported"}, {"atom_id": "juniper_first_ranked", "state": "supported"}], "focus_atom": "ada_timestamp_match", "focus_evidence": [{"path": ["evidence", "0"], "text": "The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC."}, {"path": ["evidence", "1"], "text": "The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC."}], "policy_evidence": [{"path": ["context"], "text": "At Westmere Library, student organizers Ada and Ben dispute step-free Oak for Tuesday 14:00–16:00; the library booking assistant applies rules, and exceptions go to the branch supervisor."}, {"path": ["evidence", "2"], "text": "Rules: earliest valid reservation wins; attendance may be updated before use if capacity allows; overlaps are invalid. The assistant may make routine updates."}, {"path": ["evidence", "3"], "text": "For Ben, Juniper holds 6 and is free throughout; Maple holds 10 and is free throughout; Cedar holds 6 but is free only 15:00–16:00. Rank full coverage first, then least spare capacity."}], "rules": [{"justification": "The records agree and no other audit trigger exists. Ada’s earlier overlapping reservation wins, her before-use attendance update fits Oak, and Ben can be assigned the sufficiently large, fully available, first-ranked alternative. No organizer clarification, supervisory judgment, or immediate-intervention trigger applies, so the assistant can complete a routine resolution.", "target": "0", "when": [{"atom_id": "ada_timestamp_match", "state": "supported"}, {"atom_id": "other_timestamp_contradiction", "state": "refuted"}, {"atom_id": "room_data_contradiction", "state": "refuted"}, {"atom_id": "other_reservation_record_contradiction", "state": "refuted"}, {"atom_id": "missing_attendance_fact", "state": "refuted"}, {"atom_id": "missing_timing_fact", "state": "refuted"}, {"atom_id": "missing_accessibility_fact", "state": "refuted"}, {"atom_id": "ambiguous_applicable_rule", "state": "refuted"}, {"atom_id": "requested_policy_exception", "state": "refuted"}, {"atom_id": "active_safety_breach", "state": "refuted"}, {"atom_id": "active_accessibility_breach", "state": "refuted"}, {"atom_id": "deliberate_rule_refusal", "state": "refuted"}, {"atom_id": "escalating_confrontation", "state": "refuted"}, {"atom_id": "oak_bookings_overlap", "state": "supported"}, {"atom_id": "ada_precedes_ben", "state": "supported"}, {"atom_id": "ada_attendance_within_oak_capacity", "state": "supported"}, {"atom_id": "ada_update_before_use", "state": "supported"}, {"atom_id": "juniper_full_coverage", "state": "supported"}, {"atom_id": "juniper_capacity_sufficient", "state": "supported"}, {"atom_id": "juniper_first_ranked", "state": "supported"}]}, {"justification": "The booking record and transaction log contain different submission timestamps for Ada’s same reservation, creating a contradictory-timestamp audit trigger. No organizer-clarification, supervisory-judgment, or immediate-intervention condition applies, so the required routing level is 2.", "target": "2", "when": [{"atom_id": "ada_timestamp_match", "state": "refuted"}, {"atom_id": "other_timestamp_contradiction", "state": "refuted"}, {"atom_id": "room_data_contradiction", "state": "refuted"}, {"atom_id": "other_reservation_record_contradiction", "state": "refuted"}, {"atom_id": "missing_attendance_fact", "state": "refuted"}, {"atom_id": "missing_timing_fact", "state": "refuted"}, {"atom_id": "missing_accessibility_fact", "state": "refuted"}, {"atom_id": "ambiguous_applicable_rule", "state": "refuted"}, {"atom_id": "requested_policy_exception", "state": "refuted"}, {"atom_id": "active_safety_breach", "state": "refuted"}, {"atom_id": "active_accessibility_breach", "state": "refuted"}, {"atom_id": "deliberate_rule_refusal", "state": "refuted"}, {"atom_id": "escalating_confrontation", "state": "refuted"}, {"atom_id": "oak_bookings_overlap", "state": "supported"}, {"atom_id": "ada_precedes_ben", "state": "supported"}, {"atom_id": "ada_attendance_within_oak_capacity", "state": "supported"}, {"atom_id": "ada_update_before_use", "state": "supported"}, {"atom_id": "juniper_full_coverage", "state": "supported"}, {"atom_id": "juniper_capacity_sufficient", "state": "supported"}, {"atom_id": "juniper_first_ranked", "state": "supported"}]}]}, "verified_pair": {"left": "The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.", "negative_left": "The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.", "negative_right": "The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:38:41 UTC.", "right": "The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-090-002", "id": "scale-diverse-090-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Fully rule-determined routine resolution: the booking assistant can verify validity, make any permitted attendance update, assign an alternative, and close the conflict without supervisor involvement.", "1 — Organizer clarification required: a missing attendance, timing, or accessibility fact must be supplied by a student group organizer before the booking assistant can resolve the case.", "2 — Booking-record audit required: contradictory timestamps, room data, or reservation records require the library booking assistant to pause assignment and investigate the system record.", "3 — Supervisory judgment required: an ambiguous rule or requested policy exception must be decided by the branch supervisor before a room can be assigned.", "4 — Immediate supervisory intervention required: an active safety or accessibility breach, deliberate rule refusal, or escalating confrontation requires the branch supervisor to halt room use and intervene."], "instructions": "Interpret the organizer’s wording by meaning, not exact phrase matching. Verify both bookings, apply the stated ranking rule to Ben’s alternatives, and choose the lowest applicable routing level.", "type": "score"}}, "state": {"context": "At Westmere Library, student organizers Ada and Ben dispute step-free Oak for Tuesday 14:00–16:00; the library booking assistant applies rules, and exceptions go to the branch supervisor.", "evidence": ["The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.", "The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.", "Rules: earliest valid reservation wins; attendance may be updated before use if capacity allows; overlaps are invalid. The assistant may make routine updates.", "For Ben, Juniper holds 6 and is free throughout; Maple holds 10 and is free throughout; Cedar holds 6 but is free only 15:00–16:00. Rank full coverage first, then least spare capacity.", "Both Oak reservations cover Tuesday 14:00–16:00, and Ada’s was submitted before Ben’s. Ada updated attendance to 8 at 13:30; Oak holds 8. Ben lists 5 attendees.", "Reconciliation found no other timestamp conflict, inconsistent room data, or other reservation-record discrepancy.", "The organizers supplied all needed attendance, timing, and accessibility facts. Oak and Juniper meet the stated accessibility needs, with no active safety or accessibility breach.", "The applicable rules are unambiguous. Neither organizer requests an exception or deliberately refuses a rule, and their calm discussion is not escalating."], "request": "Decide each booking’s validity, rank Ben’s alternatives, and select the required routing level."}}, "method": "c2d", "provenance": {"source_id": "diverse-090", "source_is_synthetic": true, "source_sha256": "5392349843cf9e40b4b571441be872a1eae00f7f2941188abdec5a749bff63a8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the booking rules, request scope, entities, Oak booking, and Tuesday 14:00–16:00 binding. The two focus spans are complete factual sentences describing separate record displays. The counterfactual coherently introduces a discrepancy between two record sources; the statement that reconciliation found no other timestamp conflict expressly allows this focal conflict, and Ada can still precede Ben under either displayed time. Neither context embeds an answer code, routing result, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"active_accessibility_breach": "refuted", "active_safety_breach": "refuted", "ada_attendance_within_oak_capacity": "supported", "ada_precedes_ben": "supported", "ada_timestamp_match": "refuted", "ada_update_before_use": "supported", "ambiguous_applicable_rule": "refuted", "deliberate_rule_refusal": "refuted", "escalating_confrontation": "refuted", "juniper_capacity_sufficient": "supported", "juniper_first_ranked": "supported", "juniper_full_coverage": "supported", "missing_accessibility_fact": "refuted", "missing_attendance_fact": "refuted", "missing_timing_fact": "refuted", "oak_bookings_overlap": "supported", "other_reservation_record_contradiction": "refuted", "other_timestamp_contradiction": "refuted", "requested_policy_exception": "refuted", "room_data_contradiction": "refuted"}, "full_context_fact_states": {"base": {"active_accessibility_breach": "refuted", "active_safety_breach": "refuted", "ada_attendance_within_oak_capacity": "supported", "ada_precedes_ben": "supported", "ada_timestamp_match": "supported", "ada_update_before_use": "supported", "ambiguous_applicable_rule": "refuted", "deliberate_rule_refusal": "refuted", "escalating_confrontation": "refuted", "juniper_capacity_sufficient": "supported", "juniper_first_ranked": "supported", "juniper_full_coverage": "supported", "missing_accessibility_fact": "refuted", "missing_attendance_fact": "refuted", "missing_timing_fact": "refuted", "oak_bookings_overlap": "supported", "other_reservation_record_contradiction": "refuted", "other_timestamp_contradiction": "refuted", "requested_policy_exception": "refuted", "room_data_contradiction": "refuted"}, "counterfactual": {"active_accessibility_breach": "refuted", "active_safety_breach": "refuted", "ada_attendance_within_oak_capacity": "supported", "ada_precedes_ben": "supported", "ada_timestamp_match": "refuted", "ada_update_before_use": "supported", "ambiguous_applicable_rule": "refuted", "deliberate_rule_refusal": "refuted", "escalating_confrontation": "refuted", "juniper_capacity_sufficient": "supported", "juniper_first_ranked": "supported", "juniper_full_coverage": "supported", "missing_accessibility_fact": "refuted", "missing_attendance_fact": "refuted", "missing_timing_fact": "refuted", "oak_bookings_overlap": "supported", "other_reservation_record_contradiction": "refuted", "other_timestamp_contradiction": "refuted", "requested_policy_exception": "refuted", "room_data_contradiction": "refuted"}, "remove_left": {"ada_timestamp_match": "unknown"}, "remove_right": {"ada_timestamp_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"ada_timestamp_match": "unknown"}, "negative_pair": {"ada_timestamp_match": "refuted"}, "negative_sentence": {"ada_timestamp_match": "unknown"}, "positive_pair": {"ada_timestamp_match": "supported"}, "right": {"ada_timestamp_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual condition or relationship; the ranking atom is a factual ranking rather than a routing classification. The focus is the factual equality of two timestamps. The base and counter assignments can differ only in that equality: the counter can contain unequal record and transaction-log timestamps while Ada still precedes Ben according to the operative booking record. Policy evidence correctly cites the substantive rules originating in the original state, while the routing criteria and lowest-applicable-level instruction remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a routine resolution: Ada has the earlier overlapping reservation, her timely attendance update fits Oak, and Juniper is a sufficient, fully available, first-ranked alternative for Ben. It also explicitly refutes every stated trigger for clarification, record audit, supervisory judgment, or immediate intervention.", "rule_index": 0, "sound": true}, {"reason": "Refuting equality between two submission timestamps for the same Ada reservation entails contradictory timestamp records and therefore the level-2 audit trigger. The conjunction excludes lower level 1 and the supervisory/intervention outcomes, while retaining enough facts for the underlying booking and alternative analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ada_timestamp_match", "statement": "The submission timestamp for Ada’s Tuesday 14:00–16:00 Oak reservation in the booking record equals the submission timestamp for that same reservation in the system transaction log."}, {"id": "other_timestamp_contradiction", "statement": "A contradictory timestamp exists among the booking-system records relevant to Ada’s or Ben’s Tuesday 14:00–16:00 booking, excluding the two timestamp entries compared in ada_timestamp_match."}, {"id": "room_data_contradiction", "statement": "A contradiction exists among the recorded room data relevant to Ada’s or Ben’s Tuesday 14:00–16:00 booking or Ben’s listed alternatives."}, {"id": "other_reservation_record_contradiction", "statement": "A contradiction other than a timestamp or room-data contradiction exists among the reservation records relevant to Ada’s or Ben’s Tuesday 14:00–16:00 booking."}, {"id": "missing_attendance_fact", "statement": "A student group organizer must supply an additional attendance fact before the booking assistant can resolve the Tuesday 14:00–16:00 case."}, {"id": "missing_timing_fact", "statement": "A student group organizer must supply an additional timing fact before the booking assistant can resolve the Tuesday 14:00–16:00 case."}, {"id": "missing_accessibility_fact", "statement": "A student group organizer must supply an additional accessibility fact before the booking assistant can resolve the Tuesday 14:00–16:00 case."}, {"id": "ambiguous_applicable_rule", "statement": "An applicable rule governing the Tuesday 14:00–16:00 case is ambiguous and requires branch-supervisor judgment."}, {"id": "requested_policy_exception", "statement": "Ada or Ben requests a policy exception for the Tuesday 14:00–16:00 case."}, {"id": "active_safety_breach", "statement": "An active safety breach exists in the Tuesday 14:00–16:00 case."}, {"id": "active_accessibility_breach", "statement": "An active accessibility breach exists in the Tuesday 14:00–16:00 case."}, {"id": "deliberate_rule_refusal", "statement": "Ada or Ben deliberately refuses an applicable rule in the Tuesday 14:00–16:00 case."}, {"id": "escalating_confrontation", "statement": "The confrontation between Ada and Ben is escalating in the Tuesday 14:00–16:00 case."}, {"id": "oak_bookings_overlap", "statement": "Ada’s and Ben’s Oak reservations cover the same Tuesday 14:00–16:00 period."}, {"id": "ada_precedes_ben", "statement": "Ada’s Oak reservation was submitted before Ben’s Oak reservation."}, {"id": "ada_attendance_within_oak_capacity", "statement": "Ada’s updated attendance for Tuesday 14:00–16:00 does not exceed Oak’s capacity."}, {"id": "ada_update_before_use", "statement": "Ada requested her attendance update before the Tuesday 14:00 start of room use."}, {"id": "juniper_full_coverage", "statement": "Juniper is free throughout Ben’s Tuesday 14:00–16:00 requested period."}, {"id": "juniper_capacity_sufficient", "statement": "Juniper’s capacity is at least Ben’s attendance for Tuesday 14:00–16:00."}, {"id": "juniper_first_ranked", "statement": "Juniper ranks first among Ben’s alternatives under full coverage followed by least spare capacity."}], "base_state_json": "{\"context\":\"At Westmere Library, student organizers Ada and Ben dispute step-free Oak for Tuesday 14:00–16:00; the library booking assistant applies rules, and exceptions go to the branch supervisor.\",\"evidence\":[\"The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.\",\"The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.\",\"Rules: earliest valid reservation wins; attendance may be updated before use if capacity allows; overlaps are invalid. The assistant may make routine updates.\",\"For Ben, Juniper holds 6 and is free throughout; Maple holds 10 and is free throughout; Cedar holds 6 but is free only 15:00–16:00. Rank full coverage first, then least spare capacity.\",\"Both Oak reservations cover Tuesday 14:00–16:00, and Ada’s was submitted before Ben’s. Ada updated attendance to 8 at 13:30; Oak holds 8. Ben lists 5 attendees.\",\"Reconciliation found no other timestamp conflict, inconsistent room data, or other reservation-record discrepancy.\",\"The organizers supplied all needed attendance, timing, and accessibility facts. Oak and Juniper meet the stated accessibility needs, with no active safety or accessibility breach.\",\"The applicable rules are unambiguous. Neither organizer requests an exception or deliberately refuses a rule, and their calm discussion is not escalating.\"],\"request\":\"Decide each booking’s validity, rank Ben’s alternatives, and select the required routing level.\"}", "base_states": [{"atom_id": "ada_timestamp_match", "state": "supported"}, {"atom_id": "other_timestamp_contradiction", "state": "refuted"}, {"atom_id": "room_data_contradiction", "state": "refuted"}, {"atom_id": "other_reservation_record_contradiction", "state": "refuted"}, {"atom_id": "missing_attendance_fact", "state": "refuted"}, {"atom_id": "missing_timing_fact", "state": "refuted"}, {"atom_id": "missing_accessibility_fact", "state": "refuted"}, {"atom_id": "ambiguous_applicable_rule", "state": "refuted"}, {"atom_id": "requested_policy_exception", "state": "refuted"}, {"atom_id": "active_safety_breach", "state": "refuted"}, {"atom_id": "active_accessibility_breach", "state": "refuted"}, {"atom_id": "deliberate_rule_refusal", "state": "refuted"}, {"atom_id": "escalating_confrontation", "state": "refuted"}, {"atom_id": "oak_bookings_overlap", "state": "supported"}, {"atom_id": "ada_precedes_ben", "state": "supported"}, {"atom_id": "ada_attendance_within_oak_capacity", "state": "supported"}, {"atom_id": "ada_update_before_use", "state": "supported"}, {"atom_id": "juniper_full_coverage", "state": "supported"}, {"atom_id": "juniper_capacity_sufficient", "state": "supported"}, {"atom_id": "juniper_first_ranked", "state": "supported"}], "counter_states": [{"atom_id": "ada_timestamp_match", "state": "refuted"}, {"atom_id": "other_timestamp_contradiction", "state": "refuted"}, {"atom_id": "room_data_contradiction", "state": "refuted"}, {"atom_id": "other_reservation_record_contradiction", "state": "refuted"}, {"atom_id": "missing_attendance_fact", "state": "refuted"}, {"atom_id": "missing_timing_fact", "state": "refuted"}, {"atom_id": "missing_accessibility_fact", "state": "refuted"}, {"atom_id": "ambiguous_applicable_rule", "state": "refuted"}, {"atom_id": "requested_policy_exception", "state": "refuted"}, {"atom_id": "active_safety_breach", "state": "refuted"}, {"atom_id": "active_accessibility_breach", "state": "refuted"}, {"atom_id": "deliberate_rule_refusal", "state": "refuted"}, {"atom_id": "escalating_confrontation", "state": "refuted"}, {"atom_id": "oak_bookings_overlap", "state": "supported"}, {"atom_id": "ada_precedes_ben", "state": "supported"}, {"atom_id": "ada_attendance_within_oak_capacity", "state": "supported"}, {"atom_id": "ada_update_before_use", "state": "supported"}, {"atom_id": "juniper_full_coverage", "state": "supported"}, {"atom_id": "juniper_capacity_sufficient", "state": "supported"}, {"atom_id": "juniper_first_ranked", "state": "supported"}], "focus_atom": "ada_timestamp_match", "focus_evidence": [{"path": ["evidence", "0"], "text": "The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC."}, {"path": ["evidence", "1"], "text": "The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC."}], "policy_evidence": [{"path": ["context"], "text": "At Westmere Library, student organizers Ada and Ben dispute step-free Oak for Tuesday 14:00–16:00; the library booking assistant applies rules, and exceptions go to the branch supervisor."}, {"path": ["evidence", "2"], "text": "Rules: earliest valid reservation wins; attendance may be updated before use if capacity allows; overlaps are invalid. The assistant may make routine updates."}, {"path": ["evidence", "3"], "text": "For Ben, Juniper holds 6 and is free throughout; Maple holds 10 and is free throughout; Cedar holds 6 but is free only 15:00–16:00. Rank full coverage first, then least spare capacity."}], "rules": [{"justification": "The records agree and no other audit trigger exists. Ada’s earlier overlapping reservation wins, her before-use attendance update fits Oak, and Ben can be assigned the sufficiently large, fully available, first-ranked alternative. No organizer clarification, supervisory judgment, or immediate-intervention trigger applies, so the assistant can complete a routine resolution.", "target": "0", "when": [{"atom_id": "ada_timestamp_match", "state": "supported"}, {"atom_id": "other_timestamp_contradiction", "state": "refuted"}, {"atom_id": "room_data_contradiction", "state": "refuted"}, {"atom_id": "other_reservation_record_contradiction", "state": "refuted"}, {"atom_id": "missing_attendance_fact", "state": "refuted"}, {"atom_id": "missing_timing_fact", "state": "refuted"}, {"atom_id": "missing_accessibility_fact", "state": "refuted"}, {"atom_id": "ambiguous_applicable_rule", "state": "refuted"}, {"atom_id": "requested_policy_exception", "state": "refuted"}, {"atom_id": "active_safety_breach", "state": "refuted"}, {"atom_id": "active_accessibility_breach", "state": "refuted"}, {"atom_id": "deliberate_rule_refusal", "state": "refuted"}, {"atom_id": "escalating_confrontation", "state": "refuted"}, {"atom_id": "oak_bookings_overlap", "state": "supported"}, {"atom_id": "ada_precedes_ben", "state": "supported"}, {"atom_id": "ada_attendance_within_oak_capacity", "state": "supported"}, {"atom_id": "ada_update_before_use", "state": "supported"}, {"atom_id": "juniper_full_coverage", "state": "supported"}, {"atom_id": "juniper_capacity_sufficient", "state": "supported"}, {"atom_id": "juniper_first_ranked", "state": "supported"}]}, {"justification": "The booking record and transaction log contain different submission timestamps for Ada’s same reservation, creating a contradictory-timestamp audit trigger. No organizer-clarification, supervisory-judgment, or immediate-intervention condition applies, so the required routing level is 2.", "target": "2", "when": [{"atom_id": "ada_timestamp_match", "state": "refuted"}, {"atom_id": "other_timestamp_contradiction", "state": "refuted"}, {"atom_id": "room_data_contradiction", "state": "refuted"}, {"atom_id": "other_reservation_record_contradiction", "state": "refuted"}, {"atom_id": "missing_attendance_fact", "state": "refuted"}, {"atom_id": "missing_timing_fact", "state": "refuted"}, {"atom_id": "missing_accessibility_fact", "state": "refuted"}, {"atom_id": "ambiguous_applicable_rule", "state": "refuted"}, {"atom_id": "requested_policy_exception", "state": "refuted"}, {"atom_id": "active_safety_breach", "state": "refuted"}, {"atom_id": "active_accessibility_breach", "state": "refuted"}, {"atom_id": "deliberate_rule_refusal", "state": "refuted"}, {"atom_id": "escalating_confrontation", "state": "refuted"}, {"atom_id": "oak_bookings_overlap", "state": "supported"}, {"atom_id": "ada_precedes_ben", "state": "supported"}, {"atom_id": "ada_attendance_within_oak_capacity", "state": "supported"}, {"atom_id": "ada_update_before_use", "state": "supported"}, {"atom_id": "juniper_full_coverage", "state": "supported"}, {"atom_id": "juniper_capacity_sufficient", "state": "supported"}, {"atom_id": "juniper_first_ranked", "state": "supported"}]}]}, "verified_pair": {"left": "The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.", "negative_left": "The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.", "negative_right": "The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:38:41 UTC.", "right": "The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-090-002", "id": "scale-diverse-090-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Fully rule-determined routine resolution: the booking assistant can verify validity, make any permitted attendance update, assign an alternative, and close the conflict without supervisor involvement.", "1 — Organizer clarification required: a missing attendance, timing, or accessibility fact must be supplied by a student group organizer before the booking assistant can resolve the case.", "2 — Booking-record audit required: contradictory timestamps, room data, or reservation records require the library booking assistant to pause assignment and investigate the system record.", "3 — Supervisory judgment required: an ambiguous rule or requested policy exception must be decided by the branch supervisor before a room can be assigned.", "4 — Immediate supervisory intervention required: an active safety or accessibility breach, deliberate rule refusal, or escalating confrontation requires the branch supervisor to halt room use and intervene."], "instructions": "Interpret the organizer’s wording by meaning, not exact phrase matching. Verify both bookings, apply the stated ranking rule to Ben’s alternatives, and choose the lowest applicable routing level.", "type": "score"}}, "state": {"context": "At Westmere Library, student organizers Ada and Ben dispute step-free Oak for Tuesday 14:00–16:00; the library booking assistant applies rules, and exceptions go to the branch supervisor.", "evidence": ["The booking-record entry for Ada’s Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:37:12 UTC.", "The system transaction-log entry for that same Ada Tuesday 14:00–16:00 Oak reservation displays a submission timestamp of 2026-04-06 08:38:41 UTC.", "Rules: earliest valid reservation wins; attendance may be updated before use if capacity allows; overlaps are invalid. The assistant may make routine updates.", "For Ben, Juniper holds 6 and is free throughout; Maple holds 10 and is free throughout; Cedar holds 6 but is free only 15:00–16:00. Rank full coverage first, then least spare capacity.", "Both Oak reservations cover Tuesday 14:00–16:00, and Ada’s was submitted before Ben’s. Ada updated attendance to 8 at 13:30; Oak holds 8. Ben lists 5 attendees.", "Reconciliation found no other timestamp conflict, inconsistent room data, or other reservation-record discrepancy.", "The organizers supplied all needed attendance, timing, and accessibility facts. Oak and Juniper meet the stated accessibility needs, with no active safety or accessibility breach.", "The applicable rules are unambiguous. Neither organizer requests an exception or deliberately refuses a rule, and their calm discussion is not escalating."], "request": "Decide each booking’s validity, rank Ben’s alternatives, and select the required routing level."}}, "method": "c2d", "provenance": {"source_id": "diverse-090", "source_is_synthetic": true, "source_sha256": "5392349843cf9e40b4b571441be872a1eae00f7f2941188abdec5a749bff63a8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all decision criteria, while both contexts retain the original 12-minute minimum and the proof requirements for confirmation and rejection. Mara, the proposed Orwick transfer itinerary, train legs, and relevant timetable path remain fixed. The two focus-evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual coherently changes only the recorded interval from 17 minutes to 9 minutes, without conflicting with the stated train departure times or any duplicate interval measurement. Neither context contains an explicit answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal claims over the explicitly identified itinerary remain atomic. A5 is a factual timetable relation rather than a policy proposition. The base and counter assignments can differ only in whether the transfer interval is at least 12 minutes while all other facts remain fixed, so both are realizable. Policy evidence correctly preserves the state-originating 12-minute threshold and the rule governing unproven connections; requirements already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a transfer itinerary, valid ticket acceptance, available capacity, operating services, and timetable proof that the interval meets the 12-minute minimum. It also excludes an identified direct replacement, cancellation/capacity problems, and the represented unresolved platform or service-control issues. These facts are sufficient for confirm_feasible_connection.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A5 entails that the defined transfer interval is not at least 12 minutes, hence is shorter than the required minimum. The remaining conditions establish the proposed transfer and exclude the represented competing ticket, capacity, cancellation, direct-replacement, platform, and service-control outcomes. This is sufficient for reject_infeasible_connection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed itinerary requires Mara to transfer at Orwick from the 14:05 Brindle Central–Orwick train to the 15:10 Orwick–Lanton train."}, {"id": "A2", "statement": "Mara's flexible ticket is valid on every train in the proposed same-day Brindle Central–Orwick–Lanton itinerary."}, {"id": "A3", "statement": "Every train in the proposed Brindle Central–Orwick–Lanton itinerary has an available seat for Mara."}, {"id": "A4", "statement": "Every train in the proposed Brindle Central–Orwick–Lanton itinerary is operating rather than canceled."}, {"id": "A5", "statement": "The timetable interval from the 14:05 Brindle Central–Orwick train's arrival at Orwick to the 15:10 Orwick–Lanton train's departure is at least 12 minutes."}, {"id": "A6", "statement": "An unresolved platform issue affects the proposed transfer at Orwick."}, {"id": "A7", "statement": "An unresolved service-control issue affects the proposed Brindle Central–Orwick–Lanton itinerary."}, {"id": "A8", "statement": "A direct same-day Brindle Central–Lanton replacement for Mara has been identified."}], "base_state_json": "\"Field note: The proposed same-day itinerary requires Mara to take the 14:05 Brindle Central–Orwick train and transfer at Orwick to the 15:10 Orwick–Lanton train. Mara’s flexible ticket is valid on every train in that itinerary, and each train is operating with an available seat for her. Checks found no unresolved platform issue affecting the Orwick transfer and no unresolved service-control issue affecting the itinerary. No direct same-day Brindle Central–Lanton replacement for Mara has been identified. In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick. In official timetable record OT-4005, marker E8 occurs 17 minutes after marker E7. Orwick requires at least 12 minutes for this platform change. Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick."}, {"path": [], "text": "In official timetable record OT-4005, marker E8 occurs 17 minutes after marker E7."}], "policy_evidence": [{"path": [], "text": "Orwick requires at least 12 minutes for this platform change"}, {"path": [], "text": "Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible."}], "rules": [{"justification": "The proposed transfer has established ticket validity and capacity, both services are operating, timetable observations prove that the connection meets the 12-minute minimum, and no competing direct-replacement or operations-support condition is present.", "target": "confirm_feasible_connection", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Timetable observations prove that the proposed Orwick transfer interval is shorter than the required 12 minutes, while the unchanged facts exclude ticket-support, operations-support, and direct-replacement outcomes.", "target": "reject_infeasible_connection", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick.", "negative_left": "In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick.", "negative_right": "In official timetable record OT-4005, marker E8 occurs 9 minutes after marker E7.", "right": "In official timetable record OT-4005, marker E8 occurs 17 minutes after marker E7."}, "verifier_independent_model": false}, "family": "scale-diverse-091-004", "id": "scale-diverse-091-004-base", "input": {"questions": {"decision": {"criteria": {"confirm_direct_replacement": "Confirm a replacement requiring no transfer, with ticket validity, seat availability, and arrival evidence established.", "confirm_feasible_connection": "Confirm a transfer itinerary because ticket validity and capacity are established and timetable evidence proves the transfer meets or exceeds the stated minimum.", "none_of_above": "Choose when no substantive classification is supported, including when required timetable evidence is missing and neither feasibility nor infeasibility can be proved.", "reject_infeasible_connection": "Reject the itinerary because timetable evidence proves the transfer interval is shorter than the stated minimum.", "route_operations_support": "Route the traveler to operations support because the evidence identifies unavailable capacity, a canceled replacement, or an unresolved platform or service-control issue.", "route_ticket_support": "Route the traveler to ticket support because the evidence identifies invalid, restricted, or uncertain ticket acceptance."}, "instructions": "Classify the proposed Orwick itinerary’s rebooking readiness using only the supplied evidence. Select exactly one option. The categories are mutually exclusive: confirmation and rejection require proof, while support routing requires an identified operational or ticket issue.", "type": "choice"}}, "state": "Field note: The proposed same-day itinerary requires Mara to take the 14:05 Brindle Central–Orwick train and transfer at Orwick to the 15:10 Orwick–Lanton train. Mara’s flexible ticket is valid on every train in that itinerary, and each train is operating with an available seat for her. Checks found no unresolved platform issue affecting the Orwick transfer and no unresolved service-control issue affecting the itinerary. No direct same-day Brindle Central–Lanton replacement for Mara has been identified. In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick. In official timetable record OT-4005, marker E8 occurs 17 minutes after marker E7. Orwick requires at least 12 minutes for this platform change. Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible."}, "method": "c2d", "provenance": {"source_id": "diverse-091", "source_is_synthetic": true, "source_sha256": "24b2a9b037b6ca4bf030e0c3a3fddbd327ca5f64017e628b81cca5e587443007", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "confirm_feasible_connection"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all decision criteria, while both contexts retain the original 12-minute minimum and the proof requirements for confirmation and rejection. Mara, the proposed Orwick transfer itinerary, train legs, and relevant timetable path remain fixed. The two focus-evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual coherently changes only the recorded interval from 17 minutes to 9 minutes, without conflicting with the stated train departure times or any duplicate interval measurement. Neither context contains an explicit answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal claims over the explicitly identified itinerary remain atomic. A5 is a factual timetable relation rather than a policy proposition. The base and counter assignments can differ only in whether the transfer interval is at least 12 minutes while all other facts remain fixed, so both are realizable. Policy evidence correctly preserves the state-originating 12-minute threshold and the rule governing unproven connections; requirements already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a transfer itinerary, valid ticket acceptance, available capacity, operating services, and timetable proof that the interval meets the 12-minute minimum. It also excludes an identified direct replacement, cancellation/capacity problems, and the represented unresolved platform or service-control issues. These facts are sufficient for confirm_feasible_connection.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A5 entails that the defined transfer interval is not at least 12 minutes, hence is shorter than the required minimum. The remaining conditions establish the proposed transfer and exclude the represented competing ticket, capacity, cancellation, direct-replacement, platform, and service-control outcomes. This is sufficient for reject_infeasible_connection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed itinerary requires Mara to transfer at Orwick from the 14:05 Brindle Central–Orwick train to the 15:10 Orwick–Lanton train."}, {"id": "A2", "statement": "Mara's flexible ticket is valid on every train in the proposed same-day Brindle Central–Orwick–Lanton itinerary."}, {"id": "A3", "statement": "Every train in the proposed Brindle Central–Orwick–Lanton itinerary has an available seat for Mara."}, {"id": "A4", "statement": "Every train in the proposed Brindle Central–Orwick–Lanton itinerary is operating rather than canceled."}, {"id": "A5", "statement": "The timetable interval from the 14:05 Brindle Central–Orwick train's arrival at Orwick to the 15:10 Orwick–Lanton train's departure is at least 12 minutes."}, {"id": "A6", "statement": "An unresolved platform issue affects the proposed transfer at Orwick."}, {"id": "A7", "statement": "An unresolved service-control issue affects the proposed Brindle Central–Orwick–Lanton itinerary."}, {"id": "A8", "statement": "A direct same-day Brindle Central–Lanton replacement for Mara has been identified."}], "base_state_json": "\"Field note: The proposed same-day itinerary requires Mara to take the 14:05 Brindle Central–Orwick train and transfer at Orwick to the 15:10 Orwick–Lanton train. Mara’s flexible ticket is valid on every train in that itinerary, and each train is operating with an available seat for her. Checks found no unresolved platform issue affecting the Orwick transfer and no unresolved service-control issue affecting the itinerary. No direct same-day Brindle Central–Lanton replacement for Mara has been identified. In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick. In official timetable record OT-4005, marker E8 occurs 17 minutes after marker E7. Orwick requires at least 12 minutes for this platform change. Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick."}, {"path": [], "text": "In official timetable record OT-4005, marker E8 occurs 17 minutes after marker E7."}], "policy_evidence": [{"path": [], "text": "Orwick requires at least 12 minutes for this platform change"}, {"path": [], "text": "Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible."}], "rules": [{"justification": "The proposed transfer has established ticket validity and capacity, both services are operating, timetable observations prove that the connection meets the 12-minute minimum, and no competing direct-replacement or operations-support condition is present.", "target": "confirm_feasible_connection", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Timetable observations prove that the proposed Orwick transfer interval is shorter than the required 12 minutes, while the unchanged facts exclude ticket-support, operations-support, and direct-replacement outcomes.", "target": "reject_infeasible_connection", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick.", "negative_left": "In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick.", "negative_right": "In official timetable record OT-4005, marker E8 occurs 9 minutes after marker E7.", "right": "In official timetable record OT-4005, marker E8 occurs 17 minutes after marker E7."}, "verifier_independent_model": false}, "family": "scale-diverse-091-004", "id": "scale-diverse-091-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"confirm_direct_replacement": "Confirm a replacement requiring no transfer, with ticket validity, seat availability, and arrival evidence established.", "confirm_feasible_connection": "Confirm a transfer itinerary because ticket validity and capacity are established and timetable evidence proves the transfer meets or exceeds the stated minimum.", "none_of_above": "Choose when no substantive classification is supported, including when required timetable evidence is missing and neither feasibility nor infeasibility can be proved.", "reject_infeasible_connection": "Reject the itinerary because timetable evidence proves the transfer interval is shorter than the stated minimum.", "route_operations_support": "Route the traveler to operations support because the evidence identifies unavailable capacity, a canceled replacement, or an unresolved platform or service-control issue.", "route_ticket_support": "Route the traveler to ticket support because the evidence identifies invalid, restricted, or uncertain ticket acceptance."}, "instructions": "Classify the proposed Orwick itinerary’s rebooking readiness using only the supplied evidence. Select exactly one option. The categories are mutually exclusive: confirmation and rejection require proof, while support routing requires an identified operational or ticket issue.", "type": "choice"}}, "state": "Field note: The proposed same-day itinerary requires Mara to take the 14:05 Brindle Central–Orwick train and transfer at Orwick to the 15:10 Orwick–Lanton train. Mara’s flexible ticket is valid on every train in that itinerary, and each train is operating with an available seat for her. Checks found no unresolved platform issue affecting the Orwick transfer and no unresolved service-control issue affecting the itinerary. No direct same-day Brindle Central–Lanton replacement for Mara has been identified. In the official timetable record OT-4005 for Mara's proposed itinerary on 22 September 2026, marker E7 denotes the 14:05 Brindle Central–Orwick train's arrival at Orwick and marker E8 denotes the 15:10 Orwick–Lanton train's departure from Orwick. In official timetable record OT-4005, marker E8 occurs 9 minutes after marker E7. Orwick requires at least 12 minutes for this platform change. Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible."}, "method": "c2d", "provenance": {"source_id": "diverse-091", "source_is_synthetic": true, "source_sha256": "24b2a9b037b6ca4bf030e0c3a3fddbd327ca5f64017e628b81cca5e587443007", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "reject_infeasible_connection"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing 12-minute minimum and proof requirements, while the unchanged questions preserve all classification criteria. Mara, the proposed Brindle Central–Orwick–Lanton itinerary, services, and date remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual changes one interval from 6 minutes to 3 minutes, yielding a coherent total change from 13 to 10 minutes without conflicting duplicate measurements. Neither context includes an answer label, code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal claims over the explicitly identified itinerary remain atomic. A5 is a factual timetable relation rather than a policy proposition. The base and counter assignments can differ only in whether the transfer interval is at least 12 minutes while all other facts remain fixed, so both are realizable. Policy evidence correctly preserves the state-originating 12-minute threshold and the rule governing unproven connections; requirements already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a transfer itinerary, valid ticket acceptance, available capacity, operating services, and timetable proof that the interval meets the 12-minute minimum. It also excludes an identified direct replacement, cancellation/capacity problems, and the represented unresolved platform or service-control issues. These facts are sufficient for confirm_feasible_connection.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A5 entails that the defined transfer interval is not at least 12 minutes, hence is shorter than the required minimum. The remaining conditions establish the proposed transfer and exclude the represented competing ticket, capacity, cancellation, direct-replacement, platform, and service-control outcomes. This is sufficient for reject_infeasible_connection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed itinerary requires Mara to transfer at Orwick from the 14:05 Brindle Central–Orwick train to the 15:10 Orwick–Lanton train."}, {"id": "A2", "statement": "Mara's flexible ticket is valid on every train in the proposed same-day Brindle Central–Orwick–Lanton itinerary."}, {"id": "A3", "statement": "Every train in the proposed Brindle Central–Orwick–Lanton itinerary has an available seat for Mara."}, {"id": "A4", "statement": "Every train in the proposed Brindle Central–Orwick–Lanton itinerary is operating rather than canceled."}, {"id": "A5", "statement": "The timetable interval from the 14:05 Brindle Central–Orwick train's arrival at Orwick to the 15:10 Orwick–Lanton train's departure is at least 12 minutes."}, {"id": "A6", "statement": "An unresolved platform issue affects the proposed transfer at Orwick."}, {"id": "A7", "statement": "An unresolved service-control issue affects the proposed Brindle Central–Orwick–Lanton itinerary."}, {"id": "A8", "statement": "A direct same-day Brindle Central–Lanton replacement for Mara has been identified."}], "base_state_json": "\"A rebooking review concerns Mara’s proposed same-day journey: the 14:05 Brindle Central–Orwick train, followed by a transfer to the 15:10 Orwick–Lanton train. Her flexible ticket is valid on every train in that itinerary, and a seat is available for her on each train. Both proposed services are operating and are not canceled. Staff found no unresolved platform issue at Orwick and no unresolved service-control issue affecting the itinerary. No direct same-day Brindle Central–Lanton replacement has been identified.\\n\\nIn the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 6 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick. In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41.\\n\\nOperations guidance states: Orwick requires at least 12 minutes for this platform change. Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "In the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 6 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick."}, {"path": [], "text": "In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41."}], "policy_evidence": [{"path": [], "text": "Orwick requires at least 12 minutes for this platform change"}, {"path": [], "text": "Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible."}], "rules": [{"justification": "The proposed transfer has established ticket validity and capacity, both services are operating, timetable observations prove that the connection meets the 12-minute minimum, and no competing direct-replacement or operations-support condition is present.", "target": "confirm_feasible_connection", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Timetable observations prove that the proposed Orwick transfer interval is shorter than the required 12 minutes, while the unchanged facts exclude ticket-support, operations-support, and direct-replacement outcomes.", "target": "reject_infeasible_connection", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "In the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 6 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick.", "negative_left": "In the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 3 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick.", "negative_right": "In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41.", "right": "In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41."}, "verifier_independent_model": false}, "family": "scale-diverse-091-006", "id": "scale-diverse-091-006-base", "input": {"questions": {"decision": {"criteria": {"confirm_direct_replacement": "Confirm a replacement requiring no transfer, with ticket validity, seat availability, and arrival evidence established.", "confirm_feasible_connection": "Confirm a transfer itinerary because ticket validity and capacity are established and timetable evidence proves the transfer meets or exceeds the stated minimum.", "none_of_above": "Choose when no substantive classification is supported, including when required timetable evidence is missing and neither feasibility nor infeasibility can be proved.", "reject_infeasible_connection": "Reject the itinerary because timetable evidence proves the transfer interval is shorter than the stated minimum.", "route_operations_support": "Route the traveler to operations support because the evidence identifies unavailable capacity, a canceled replacement, or an unresolved platform or service-control issue.", "route_ticket_support": "Route the traveler to ticket support because the evidence identifies invalid, restricted, or uncertain ticket acceptance."}, "instructions": "Classify the proposed Orwick itinerary’s rebooking readiness using only the supplied evidence. Select exactly one option. The categories are mutually exclusive: confirmation and rejection require proof, while support routing requires an identified operational or ticket issue.", "type": "choice"}}, "state": "A rebooking review concerns Mara’s proposed same-day journey: the 14:05 Brindle Central–Orwick train, followed by a transfer to the 15:10 Orwick–Lanton train. Her flexible ticket is valid on every train in that itinerary, and a seat is available for her on each train. Both proposed services are operating and are not canceled. Staff found no unresolved platform issue at Orwick and no unresolved service-control issue affecting the itinerary. No direct same-day Brindle Central–Lanton replacement has been identified.\n\nIn the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 6 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick. In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41.\n\nOperations guidance states: Orwick requires at least 12 minutes for this platform change. Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible."}, "method": "c2d", "provenance": {"source_id": "diverse-091", "source_is_synthetic": true, "source_sha256": "24b2a9b037b6ca4bf030e0c3a3fddbd327ca5f64017e628b81cca5e587443007", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "confirm_feasible_connection"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing 12-minute minimum and proof requirements, while the unchanged questions preserve all classification criteria. Mara, the proposed Brindle Central–Orwick–Lanton itinerary, services, and date remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual changes one interval from 6 minutes to 3 minutes, yielding a coherent total change from 13 to 10 minutes without conflicting duplicate measurements. Neither context includes an answer label, code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal claims over the explicitly identified itinerary remain atomic. A5 is a factual timetable relation rather than a policy proposition. The base and counter assignments can differ only in whether the transfer interval is at least 12 minutes while all other facts remain fixed, so both are realizable. Policy evidence correctly preserves the state-originating 12-minute threshold and the rule governing unproven connections; requirements already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a transfer itinerary, valid ticket acceptance, available capacity, operating services, and timetable proof that the interval meets the 12-minute minimum. It also excludes an identified direct replacement, cancellation/capacity problems, and the represented unresolved platform or service-control issues. These facts are sufficient for confirm_feasible_connection.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A5 entails that the defined transfer interval is not at least 12 minutes, hence is shorter than the required minimum. The remaining conditions establish the proposed transfer and exclude the represented competing ticket, capacity, cancellation, direct-replacement, platform, and service-control outcomes. This is sufficient for reject_infeasible_connection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed itinerary requires Mara to transfer at Orwick from the 14:05 Brindle Central–Orwick train to the 15:10 Orwick–Lanton train."}, {"id": "A2", "statement": "Mara's flexible ticket is valid on every train in the proposed same-day Brindle Central–Orwick–Lanton itinerary."}, {"id": "A3", "statement": "Every train in the proposed Brindle Central–Orwick–Lanton itinerary has an available seat for Mara."}, {"id": "A4", "statement": "Every train in the proposed Brindle Central–Orwick–Lanton itinerary is operating rather than canceled."}, {"id": "A5", "statement": "The timetable interval from the 14:05 Brindle Central–Orwick train's arrival at Orwick to the 15:10 Orwick–Lanton train's departure is at least 12 minutes."}, {"id": "A6", "statement": "An unresolved platform issue affects the proposed transfer at Orwick."}, {"id": "A7", "statement": "An unresolved service-control issue affects the proposed Brindle Central–Orwick–Lanton itinerary."}, {"id": "A8", "statement": "A direct same-day Brindle Central–Lanton replacement for Mara has been identified."}], "base_state_json": "\"A rebooking review concerns Mara’s proposed same-day journey: the 14:05 Brindle Central–Orwick train, followed by a transfer to the 15:10 Orwick–Lanton train. Her flexible ticket is valid on every train in that itinerary, and a seat is available for her on each train. Both proposed services are operating and are not canceled. Staff found no unresolved platform issue at Orwick and no unresolved service-control issue affecting the itinerary. No direct same-day Brindle Central–Lanton replacement has been identified.\\n\\nIn the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 6 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick. In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41.\\n\\nOperations guidance states: Orwick requires at least 12 minutes for this platform change. Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "In the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 6 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick."}, {"path": [], "text": "In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41."}], "policy_evidence": [{"path": [], "text": "Orwick requires at least 12 minutes for this platform change"}, {"path": [], "text": "Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible."}], "rules": [{"justification": "The proposed transfer has established ticket validity and capacity, both services are operating, timetable observations prove that the connection meets the 12-minute minimum, and no competing direct-replacement or operations-support condition is present.", "target": "confirm_feasible_connection", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Timetable observations prove that the proposed Orwick transfer interval is shorter than the required 12 minutes, while the unchanged facts exclude ticket-support, operations-support, and direct-replacement outcomes.", "target": "reject_infeasible_connection", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "In the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 6 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick.", "negative_left": "In the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 3 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick.", "negative_right": "In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41.", "right": "In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41."}, "verifier_independent_model": false}, "family": "scale-diverse-091-006", "id": "scale-diverse-091-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"confirm_direct_replacement": "Confirm a replacement requiring no transfer, with ticket validity, seat availability, and arrival evidence established.", "confirm_feasible_connection": "Confirm a transfer itinerary because ticket validity and capacity are established and timetable evidence proves the transfer meets or exceeds the stated minimum.", "none_of_above": "Choose when no substantive classification is supported, including when required timetable evidence is missing and neither feasibility nor infeasibility can be proved.", "reject_infeasible_connection": "Reject the itinerary because timetable evidence proves the transfer interval is shorter than the stated minimum.", "route_operations_support": "Route the traveler to operations support because the evidence identifies unavailable capacity, a canceled replacement, or an unresolved platform or service-control issue.", "route_ticket_support": "Route the traveler to ticket support because the evidence identifies invalid, restricted, or uncertain ticket acceptance."}, "instructions": "Classify the proposed Orwick itinerary’s rebooking readiness using only the supplied evidence. Select exactly one option. The categories are mutually exclusive: confirmation and rejection require proof, while support routing requires an identified operational or ticket issue.", "type": "choice"}}, "state": "A rebooking review concerns Mara’s proposed same-day journey: the 14:05 Brindle Central–Orwick train, followed by a transfer to the 15:10 Orwick–Lanton train. Her flexible ticket is valid on every train in that itinerary, and a seat is available for her on each train. Both proposed services are operating and are not canceled. Staff found no unresolved platform issue at Orwick and no unresolved service-control issue affecting the itinerary. No direct same-day Brindle Central–Lanton replacement has been identified.\n\nIn the timetable record for Mara's proposed itinerary on 17 September 2026, reconciliation marker R-41 occurs exactly 3 minutes after the 14:05 Brindle Central–Orwick train's arrival at Orwick. In the timetable record for Mara's proposed itinerary on 17 September 2026, the 15:10 Orwick–Lanton train's departure occurs exactly 7 minutes after reconciliation marker R-41.\n\nOperations guidance states: Orwick requires at least 12 minutes for this platform change. Policy permits confirmation only when timetable evidence proves the minimum connection; an unproven connection must not be rejected as infeasible."}, "method": "c2d", "provenance": {"source_id": "diverse-091", "source_is_synthetic": true, "source_sha256": "24b2a9b037b6ca4bf030e0c3a3fddbd327ca5f64017e628b81cca5e587443007", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "reject_infeasible_connection"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original ranking rule, connection threshold and held same-platform exception without adding priorities or defaults, and the request, options, traveler, route, and time bindings remain unchanged. The two focus spans are complete factual sentences; the counterfactual coherently changes the accessibility finding for mandatory passage P-44, with no conflicting duplicate assertion, and neither context reveals a correct choice or directs the classifier toward one.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"A_B_C\":\"A is least suitable, B is intermediate, and C is most suitable.\",\"A_C_B\":\"A is least suitable, C is intermediate, and B is most suitable.\",\"C_A_B\":\"C is least suitable, A is intermediate, and B is most suitable.\"},\"instructions\":\"Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.\",\"type\":\"choice\"}},\"context\":\"After the Northport–Lake City cancellation, agent Ivo reviewed three rebooking options for Mara. Mara said she cannot use stairs, and Ivo verified that options A, B, and C were all available to her.\",\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order.\",\"At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse each of passages P-31, P-44, and P-58 from end to end without using stairs.\",\"Option A requires Mara to use stairs and arrives at 16:10. Option B permits travel without using stairs and arrives at 16:30. It has exactly one connection: a seven-minute, same-platform transfer that train operations coordinator Sela confirmed would be held. Option C is direct and arrives at 17:20.\"],\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order."}, {"path": ["evidence", "3"], "text": "At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse each of passages P-31, P-44, and P-58 from end to end without using stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order.", "negative_left": "At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order.", "negative_right": "At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse passage P-44 from end to end only by using stairs.", "right": "At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse each of passages P-31, P-44, and P-58 from end to end without using stairs."}, "verifier_independent_model": false}, "family": "scale-diverse-092-001", "id": "scale-diverse-092-001-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"context": "After the Northport–Lake City cancellation, agent Ivo reviewed three rebooking options for Mara. Mara said she cannot use stairs, and Ivo verified that options A, B, and C were all available to her.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order.", "At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse each of passages P-31, P-44, and P-58 from end to end without using stairs.", "Option A requires Mara to use stairs and arrives at 16:10. Option B permits travel without using stairs and arrives at 16:30. It has exactly one connection: a seven-minute, same-platform transfer that train operations coordinator Sela confirmed would be held. Option C is direct and arrives at 17:20."], "questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original ranking rule, connection threshold and held same-platform exception without adding priorities or defaults, and the request, options, traveler, route, and time bindings remain unchanged. The two focus spans are complete factual sentences; the counterfactual coherently changes the accessibility finding for mandatory passage P-44, with no conflicting duplicate assertion, and neither context reveals a correct choice or directs the classifier toward one.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"A_B_C\":\"A is least suitable, B is intermediate, and C is most suitable.\",\"A_C_B\":\"A is least suitable, C is intermediate, and B is most suitable.\",\"C_A_B\":\"C is least suitable, A is intermediate, and B is most suitable.\"},\"instructions\":\"Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.\",\"type\":\"choice\"}},\"context\":\"After the Northport–Lake City cancellation, agent Ivo reviewed three rebooking options for Mara. Mara said she cannot use stairs, and Ivo verified that options A, B, and C were all available to her.\",\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order.\",\"At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse each of passages P-31, P-44, and P-58 from end to end without using stairs.\",\"Option A requires Mara to use stairs and arrives at 16:10. Option B permits travel without using stairs and arrives at 16:30. It has exactly one connection: a seven-minute, same-platform transfer that train operations coordinator Sela confirmed would be held. Option C is direct and arrives at 17:20.\"],\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order."}, {"path": ["evidence", "3"], "text": "At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse each of passages P-31, P-44, and P-58 from end to end without using stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order.", "negative_left": "At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order.", "negative_right": "At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse passage P-44 from end to end only by using stairs.", "right": "At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse each of passages P-31, P-44, and P-58 from end to end without using stairs."}, "verifier_independent_model": false}, "family": "scale-diverse-092-001", "id": "scale-diverse-092-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"context": "After the Northport–Lake City cancellation, agent Ivo reviewed three rebooking options for Mara. Mara said she cannot use stairs, and Ivo verified that options A, B, and C were all available to her.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "At 08:12 on 17 September 2026, the rebooking record assigned Mara to option C with itinerary IC-47, whose complete route consists solely of mandatory passages P-31, P-44, and P-58 in that order.", "At 07:40 on 17 September 2026, an on-site accessibility inspection recorded that Mara could traverse passage P-44 from end to end only by using stairs.", "Option A requires Mara to use stairs and arrives at 16:10. Option B permits travel without using stairs and arrives at 16:30. It has exactly one connection: a seven-minute, same-platform transfer that train operations coordinator Sela confirmed would be held. Option C is direct and arrives at 17:20."], "questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same ranking rule, connection scope and held-connection exception, as well as Mara, options A/B/C, and the same ranking request. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the route-accessibility observation for option C and remains consistent with C being direct, since a direct route may still contain stairs. Neither constructed context states an ordering, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\":\"Mara cannot use stairs. Options A, B, and C are all available to her for rebooking.\",\"handoff\":[\"The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27.\",\"Route L-27 contains a continuous path from O-17 to D-17 consisting solely of level corridors and operating elevators.\"],\"option_facts\":[\"Option A requires Mara to use stairs and arrives earlier than option C.\",\"Option B permits Mara to travel without using stairs and arrives earlier than option C.\",\"Option B has exactly one connection. It is on the same platform and lasts at least 7 minutes.\",\"Train operations coordinator Sela confirmed that option B's connection would be held.\",\"Option C is direct.\"],\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\"],\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["handoff", "0"], "text": "The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27."}, {"path": ["handoff", "1"], "text": "Route L-27 contains a continuous path from O-17 to D-17 consisting solely of level corridors and operating elevators."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27.", "negative_left": "The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27.", "negative_right": "Every continuous path within route L-27 from O-17 to D-17 includes at least one staircase.", "right": "Route L-27 contains a continuous path from O-17 to D-17 consisting solely of level corridors and operating elevators."}, "verifier_independent_model": false}, "family": "scale-diverse-092-003", "id": "scale-diverse-092-003-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara cannot use stairs. Options A, B, and C are all available to her for rebooking.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."], "handoff": ["The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27.", "Route L-27 contains a continuous path from O-17 to D-17 consisting solely of level corridors and operating elevators."], "option_facts": ["Option A requires Mara to use stairs and arrives earlier than option C.", "Option B permits Mara to travel without using stairs and arrives earlier than option C.", "Option B has exactly one connection. It is on the same platform and lasts at least 7 minutes.", "Train operations coordinator Sela confirmed that option B's connection would be held.", "Option C is direct."], "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same ranking rule, connection scope and held-connection exception, as well as Mara, options A/B/C, and the same ranking request. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the route-accessibility observation for option C and remains consistent with C being direct, since a direct route may still contain stairs. Neither constructed context states an ordering, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\":\"Mara cannot use stairs. Options A, B, and C are all available to her for rebooking.\",\"handoff\":[\"The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27.\",\"Route L-27 contains a continuous path from O-17 to D-17 consisting solely of level corridors and operating elevators.\"],\"option_facts\":[\"Option A requires Mara to use stairs and arrives earlier than option C.\",\"Option B permits Mara to travel without using stairs and arrives earlier than option C.\",\"Option B has exactly one connection. It is on the same platform and lasts at least 7 minutes.\",\"Train operations coordinator Sela confirmed that option B's connection would be held.\",\"Option C is direct.\"],\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\"],\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["handoff", "0"], "text": "The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27."}, {"path": ["handoff", "1"], "text": "Route L-27 contains a continuous path from O-17 to D-17 consisting solely of level corridors and operating elevators."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27.", "negative_left": "The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27.", "negative_right": "Every continuous path within route L-27 from O-17 to D-17 includes at least one staircase.", "right": "Route L-27 contains a continuous path from O-17 to D-17 consisting solely of level corridors and operating elevators."}, "verifier_independent_model": false}, "family": "scale-diverse-092-003", "id": "scale-diverse-092-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara cannot use stairs. Options A, B, and C are all available to her for rebooking.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."], "handoff": ["The 16:40 operational handoff assigned Mara's option C itinerary as travel from origin point O-17 to destination point D-17 by any available continuous path within route L-27.", "Every continuous path within route L-27 from O-17 to D-17 includes at least one staircase."], "option_facts": ["Option A requires Mara to use stairs and arrives earlier than option C.", "Option B permits Mara to travel without using stairs and arrives earlier than option C.", "Option B has exactly one connection. It is on the same platform and lasts at least 7 minutes.", "Train operations coordinator Sela confirmed that option B's connection would be held.", "Option C is direct."], "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing ranking and connection rules, the same request, and the Mara/A/B/C rebooking bindings. The two focus spans are complete factual sentences: one links route Q-47 to option C, and the other supplies the base accessibility observation. The counterfactual changes only that accessibility observation; it does not conflict with C being direct, since a direct route may still require stairs, and it introduces no duplicate measurement. Neither context states a choice code, final ordering, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"context\":\"Field note: Ivo reviewed the three rebooking options with Mara after the cancellation.\",\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\",\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"Mara stated that she cannot use stairs.\",\"Options A, B, and C are each available to Mara for rebooking.\",\"Option A requires stair use and arrives at 16:10.\",\"Option B permits travel without stairs and arrives at 16:30.\",\"Option B has exactly one connection, a same-platform transfer from 15:20 to 15:27. Sela, the train operations coordinator, confirmed that it would be held.\",\"Option C is direct and arrives at 17:20.\",\"At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara.\",\"At 14:31 on 6 May 2026, a complete end-to-end inspection recorded a continuous path through route Q-47 that Mara can travel without using any stairs.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "8"], "text": "At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara."}, {"path": ["evidence", "9"], "text": "At 14:31 on 6 May 2026, a complete end-to-end inspection recorded a continuous path through route Q-47 that Mara can travel without using any stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara.", "negative_left": "At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara.", "negative_right": "At 14:31 on 6 May 2026, a complete end-to-end inspection recorded that every continuous path through route Q-47 that Mara can travel requires her to use stairs.", "right": "At 14:31 on 6 May 2026, a complete end-to-end inspection recorded a continuous path through route Q-47 that Mara can travel without using any stairs."}, "verifier_independent_model": false}, "family": "scale-diverse-092-004", "id": "scale-diverse-092-004-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"context": "Field note: Ivo reviewed the three rebooking options with Mara after the cancellation.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "Mara stated that she cannot use stairs.", "Options A, B, and C are each available to Mara for rebooking.", "Option A requires stair use and arrives at 16:10.", "Option B permits travel without stairs and arrives at 16:30.", "Option B has exactly one connection, a same-platform transfer from 15:20 to 15:27. Sela, the train operations coordinator, confirmed that it would be held.", "Option C is direct and arrives at 17:20.", "At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara.", "At 14:31 on 6 May 2026, a complete end-to-end inspection recorded a continuous path through route Q-47 that Mara can travel without using any stairs."], "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing ranking and connection rules, the same request, and the Mara/A/B/C rebooking bindings. The two focus spans are complete factual sentences: one links route Q-47 to option C, and the other supplies the base accessibility observation. The counterfactual changes only that accessibility observation; it does not conflict with C being direct, since a direct route may still require stairs, and it introduces no duplicate measurement. Neither context states a choice code, final ordering, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"context\":\"Field note: Ivo reviewed the three rebooking options with Mara after the cancellation.\",\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\",\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"Mara stated that she cannot use stairs.\",\"Options A, B, and C are each available to Mara for rebooking.\",\"Option A requires stair use and arrives at 16:10.\",\"Option B permits travel without stairs and arrives at 16:30.\",\"Option B has exactly one connection, a same-platform transfer from 15:20 to 15:27. Sela, the train operations coordinator, confirmed that it would be held.\",\"Option C is direct and arrives at 17:20.\",\"At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara.\",\"At 14:31 on 6 May 2026, a complete end-to-end inspection recorded a continuous path through route Q-47 that Mara can travel without using any stairs.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "8"], "text": "At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara."}, {"path": ["evidence", "9"], "text": "At 14:31 on 6 May 2026, a complete end-to-end inspection recorded a continuous path through route Q-47 that Mara can travel without using any stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara.", "negative_left": "At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara.", "negative_right": "At 14:31 on 6 May 2026, a complete end-to-end inspection recorded that every continuous path through route Q-47 that Mara can travel requires her to use stairs.", "right": "At 14:31 on 6 May 2026, a complete end-to-end inspection recorded a continuous path through route Q-47 that Mara can travel without using any stairs."}, "verifier_independent_model": false}, "family": "scale-diverse-092-004", "id": "scale-diverse-092-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"context": "Field note: Ivo reviewed the three rebooking options with Mara after the cancellation.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "Mara stated that she cannot use stairs.", "Options A, B, and C are each available to Mara for rebooking.", "Option A requires stair use and arrives at 16:10.", "Option B permits travel without stairs and arrives at 16:30.", "Option B has exactly one connection, a same-platform transfer from 15:20 to 15:27. Sela, the train operations coordinator, confirmed that it would be held.", "Option C is direct and arrives at 17:20.", "At 14:22 on 6 May 2026, the finalized rebooking record assigned route Q-47, and no other route, to option C for Mara.", "At 14:31 on 6 May 2026, a complete end-to-end inspection recorded that every continuous path through route Q-47 that Mara can travel requires her to use stairs."], "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original ranking rule and connection exception, while the unchanged questions object preserves the instructions and choice criteria. Mara, options A/B/C, itinerary 742, the ranking request, and relevant dated observations remain consistently bound. The two focus spans are complete factual sentences. Replacing Mara’s stair-free completion with the finding that every complete path through itinerary 742 contains a staircase is coherent with the unchanged facts: option C being direct does not imply that its station path is stair-free, and assignment or availability does not establish feasibility for Mara. Neither context contains a gold ordering, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"context\":\"Following a cancellation, agent Ivo reviewed three rebooking options with Mara using Sela’s live operations information.\",\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"At 08:25, all three options remained available for Mara, who stated that she could not use stairs.\",\"At 08:30, Ivo recorded that option A required a stair-only transfer.\",\"At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara.\",\"At 09:05, the records showed that option B permitted stair-free travel and had exactly one connection: a seven-minute, same-platform connection that Sela confirmed would be held. Sela was the train operations coordinator.\",\"At 09:10, the timetable identified option C as direct and showed both A and B arriving earlier than C.\",\"At 10:18 on 17 September 2026, Mara completed itinerary 742 from origin to destination using only level passageways and elevator E-26, without entering a staircase.\"],\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara."}, {"path": ["evidence", "7"], "text": "At 10:18 on 17 September 2026, Mara completed itinerary 742 from origin to destination using only level passageways and elevator E-26, without entering a staircase."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara.", "negative_left": "At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara.", "negative_right": "At 10:18 on 17 September 2026, inspectors found that every complete origin-to-destination path through itinerary 742 included at least one staircase.", "right": "At 10:18 on 17 September 2026, Mara completed itinerary 742 from origin to destination using only level passageways and elevator E-26, without entering a staircase."}, "verifier_independent_model": false}, "family": "scale-diverse-092-005", "id": "scale-diverse-092-005-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"context": "Following a cancellation, agent Ivo reviewed three rebooking options with Mara using Sela’s live operations information.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "At 08:25, all three options remained available for Mara, who stated that she could not use stairs.", "At 08:30, Ivo recorded that option A required a stair-only transfer.", "At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara.", "At 09:05, the records showed that option B permitted stair-free travel and had exactly one connection: a seven-minute, same-platform connection that Sela confirmed would be held. Sela was the train operations coordinator.", "At 09:10, the timetable identified option C as direct and showed both A and B arriving earlier than C.", "At 10:18 on 17 September 2026, Mara completed itinerary 742 from origin to destination using only level passageways and elevator E-26, without entering a staircase."], "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original ranking rule and connection exception, while the unchanged questions object preserves the instructions and choice criteria. Mara, options A/B/C, itinerary 742, the ranking request, and relevant dated observations remain consistently bound. The two focus spans are complete factual sentences. Replacing Mara’s stair-free completion with the finding that every complete path through itinerary 742 contains a staircase is coherent with the unchanged facts: option C being direct does not imply that its station path is stair-free, and assignment or availability does not establish feasibility for Mara. Neither context contains a gold ordering, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"context\":\"Following a cancellation, agent Ivo reviewed three rebooking options with Mara using Sela’s live operations information.\",\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"At 08:25, all three options remained available for Mara, who stated that she could not use stairs.\",\"At 08:30, Ivo recorded that option A required a stair-only transfer.\",\"At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara.\",\"At 09:05, the records showed that option B permitted stair-free travel and had exactly one connection: a seven-minute, same-platform connection that Sela confirmed would be held. Sela was the train operations coordinator.\",\"At 09:10, the timetable identified option C as direct and showed both A and B arriving earlier than C.\",\"At 10:18 on 17 September 2026, Mara completed itinerary 742 from origin to destination using only level passageways and elevator E-26, without entering a staircase.\"],\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara."}, {"path": ["evidence", "7"], "text": "At 10:18 on 17 September 2026, Mara completed itinerary 742 from origin to destination using only level passageways and elevator E-26, without entering a staircase."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara.", "negative_left": "At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara.", "negative_right": "At 10:18 on 17 September 2026, inspectors found that every complete origin-to-destination path through itinerary 742 included at least one staircase.", "right": "At 10:18 on 17 September 2026, Mara completed itinerary 742 from origin to destination using only level passageways and elevator E-26, without entering a staircase."}, "verifier_independent_model": false}, "family": "scale-diverse-092-005", "id": "scale-diverse-092-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"context": "Following a cancellation, agent Ivo reviewed three rebooking options with Mara using Sela’s live operations information.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "At 08:25, all three options remained available for Mara, who stated that she could not use stairs.", "At 08:30, Ivo recorded that option A required a stair-only transfer.", "At 08:42 on 17 September 2026, the rebooking log assigned itinerary 742 to option C for Mara.", "At 09:05, the records showed that option B permitted stair-free travel and had exactly one connection: a seven-minute, same-platform connection that Sela confirmed would be held. Sela was the train operations coordinator.", "At 09:10, the timetable identified option C as direct and showed both A and B arriving earlier than C.", "At 10:18 on 17 September 2026, inspectors found that every complete origin-to-destination path through itinerary 742 included at least one staircase."], "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original ranking rule and connection exception verbatim, while the unchanged questions object preserves the choice criteria and instructions. The request continues to rank Mara’s options A, B, and C, and the changed observations do not alter those bindings. The two focus spans are complete factual sentences. The counterfactual changes only the accessibility finding for itinerary IX-482/option C; a direct route may still contain stairs, so it does not conflict with the unchanged evidence. Neither context states a complete ordering, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"context\":\"Following a cancellation, agent Ivo reconciled Mara’s rebooking record with operations and accessibility reports dated 17 September 2026.\",\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara.\",\"At 09:18 on 17 September 2026, an accessibility inspection found that itinerary IX-482 included a complete route available to Mara with no stairs anywhere along it.\",\"Options A, B, and C are each available to Mara for rebooking.\",\"Mara cannot use stairs, and option A requires her to use them.\",\"Option B permits Mara to travel without stairs and has exactly one connection. That connection is on the same platform and lasts at least seven minutes.\",\"Sela is a train operations coordinator, and she confirmed that option B’s connection would be held.\",\"Option C is direct.\",\"The arrival comparison shows that B arrives earlier than C and A arrives earlier than C.\"],\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara."}, {"path": ["evidence", "3"], "text": "At 09:18 on 17 September 2026, an accessibility inspection found that itinerary IX-482 included a complete route available to Mara with no stairs anywhere along it."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara.", "negative_left": "At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara.", "negative_right": "At 09:18 on 17 September 2026, an accessibility inspection found that every complete route available to Mara on itinerary IX-482 included stairs.", "right": "At 09:18 on 17 September 2026, an accessibility inspection found that itinerary IX-482 included a complete route available to Mara with no stairs anywhere along it."}, "verifier_independent_model": false}, "family": "scale-diverse-092-006", "id": "scale-diverse-092-006-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"context": "Following a cancellation, agent Ivo reconciled Mara’s rebooking record with operations and accessibility reports dated 17 September 2026.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara.", "At 09:18 on 17 September 2026, an accessibility inspection found that itinerary IX-482 included a complete route available to Mara with no stairs anywhere along it.", "Options A, B, and C are each available to Mara for rebooking.", "Mara cannot use stairs, and option A requires her to use them.", "Option B permits Mara to travel without stairs and has exactly one connection. That connection is on the same platform and lasts at least seven minutes.", "Sela is a train operations coordinator, and she confirmed that option B’s connection would be held.", "Option C is direct.", "The arrival comparison shows that B arrives earlier than C and A arrives earlier than C."], "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original ranking rule and connection exception verbatim, while the unchanged questions object preserves the choice criteria and instructions. The request continues to rank Mara’s options A, B, and C, and the changed observations do not alter those bindings. The two focus spans are complete factual sentences. The counterfactual changes only the accessibility finding for itinerary IX-482/option C; a direct route may still contain stairs, so it does not conflict with the unchanged evidence. Neither context states a complete ordering, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"context\":\"Following a cancellation, agent Ivo reconciled Mara’s rebooking record with operations and accessibility reports dated 17 September 2026.\",\"evidence\":[\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara.\",\"At 09:18 on 17 September 2026, an accessibility inspection found that itinerary IX-482 included a complete route available to Mara with no stairs anywhere along it.\",\"Options A, B, and C are each available to Mara for rebooking.\",\"Mara cannot use stairs, and option A requires her to use them.\",\"Option B permits Mara to travel without stairs and has exactly one connection. That connection is on the same platform and lasts at least seven minutes.\",\"Sela is a train operations coordinator, and she confirmed that option B’s connection would be held.\",\"Option C is direct.\",\"The arrival comparison shows that B arrives earlier than C and A arrives earlier than C.\"],\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara."}, {"path": ["evidence", "3"], "text": "At 09:18 on 17 September 2026, an accessibility inspection found that itinerary IX-482 included a complete route available to Mara with no stairs anywhere along it."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara.", "negative_left": "At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara.", "negative_right": "At 09:18 on 17 September 2026, an accessibility inspection found that every complete route available to Mara on itinerary IX-482 included stairs.", "right": "At 09:18 on 17 September 2026, an accessibility inspection found that itinerary IX-482 included a complete route available to Mara with no stairs anywhere along it."}, "verifier_independent_model": false}, "family": "scale-diverse-092-006", "id": "scale-diverse-092-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"context": "Following a cancellation, agent Ivo reconciled Mara’s rebooking record with operations and accessibility reports dated 17 September 2026.", "evidence": ["Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "At 09:12 on 17 September 2026, the rebooking record assigned itinerary IX-482 to option C for Mara.", "At 09:18 on 17 September 2026, an accessibility inspection found that every complete route available to Mara on itinerary IX-482 included stairs.", "Options A, B, and C are each available to Mara for rebooking.", "Mara cannot use stairs, and option A requires her to use them.", "Option B permits Mara to travel without stairs and has exactly one connection. That connection is on the same platform and lasts at least seven minutes.", "Sela is a train operations coordinator, and she confirmed that option B’s connection would be held.", "Option C is direct.", "The arrival comparison shows that B arrives earlier than C and A arrives earlier than C."], "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler’s one-transfer and 18:00-arrival constraints, Option A’s Alderwick–Bellford–Seaborne path, the identified Coastliner C88 service, and Bellford’s same-platform and cross-platform minimum connection rules. The counterfactual changes only C88’s Bellford departure platform from 9 to 4, without creating a duplicate or contradictory measurement. The two evidence spans are complete factual sentences, and neither context includes an answer, code, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"My canceled through-ticket ran from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking agent\",\"text\":\"Option A uses only two legs: R214 from Alderwick to Bellford, followed by Coastliner C88 from Bellford to Seaborne, with the sole transfer at Bellford. Coastliner C88 is scheduled to reach Seaborne at 17:20.\"},{\"speaker\":\"platform record\",\"text\":\"In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4.\"},{\"speaker\":\"platform record\",\"text\":\"In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 9.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please book Option A only if it meets my constraints and Bellford's applicable minimum connection time.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4."}, {"path": ["3", "text"], "text": "In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 9."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4.", "negative_left": "In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4.", "negative_right": "In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 4.", "right": "In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 9."}, "verifier_independent_model": false}, "family": "scale-diverse-093-001", "id": "scale-diverse-093-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My canceled through-ticket ran from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking agent", "text": "Option A uses only two legs: R214 from Alderwick to Bellford, followed by Coastliner C88 from Bellford to Seaborne, with the sole transfer at Bellford. Coastliner C88 is scheduled to reach Seaborne at 17:20."}, {"speaker": "platform record", "text": "In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4."}, {"speaker": "platform record", "text": "In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 9."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please book Option A only if it meets my constraints and Bellford's applicable minimum connection time."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler’s one-transfer and 18:00-arrival constraints, Option A’s Alderwick–Bellford–Seaborne path, the identified Coastliner C88 service, and Bellford’s same-platform and cross-platform minimum connection rules. The counterfactual changes only C88’s Bellford departure platform from 9 to 4, without creating a duplicate or contradictory measurement. The two evidence spans are complete factual sentences, and neither context includes an answer, code, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"My canceled through-ticket ran from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking agent\",\"text\":\"Option A uses only two legs: R214 from Alderwick to Bellford, followed by Coastliner C88 from Bellford to Seaborne, with the sole transfer at Bellford. Coastliner C88 is scheduled to reach Seaborne at 17:20.\"},{\"speaker\":\"platform record\",\"text\":\"In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4.\"},{\"speaker\":\"platform record\",\"text\":\"In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 9.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please book Option A only if it meets my constraints and Bellford's applicable minimum connection time.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4."}, {"path": ["3", "text"], "text": "In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 9."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4.", "negative_left": "In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4.", "negative_right": "In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 4.", "right": "In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 9."}, "verifier_independent_model": false}, "family": "scale-diverse-093-001", "id": "scale-diverse-093-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My canceled through-ticket ran from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking agent", "text": "Option A uses only two legs: R214 from Alderwick to Bellford, followed by Coastliner C88 from Bellford to Seaborne, with the sole transfer at Bellford. Coastliner C88 is scheduled to reach Seaborne at 17:20."}, {"speaker": "platform record", "text": "In Option A at Bellford, R214's scheduled 16:02 arrival is assigned to platform 4."}, {"speaker": "platform record", "text": "In Option A at Bellford, Coastliner C88's scheduled 16:09 departure is assigned to platform 4."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please book Option A only if it meets my constraints and Bellford's applicable minimum connection time."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler’s one-transfer and 18:00-arrival constraints, the Option A route and train identities, and the Bellford minimum-connection policy, including the exact-minimum rule. The two evidence spans are complete factual log sentences. The counterfactual changes only C88’s recorded departure platform from 7 to 4, which is coherent with the unchanged seven-minute interval and creates no duplicate conflicting record. Neither context contains a gold label, answer code, rationale, proposition identifier, or output directive.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking handoff\",\"text\":\"Option A for 14 October 2026 consists of R214 from Alderwick to Bellford followed by Coastliner C88 from Bellford to Seaborne. Bellford is the itinerary’s sole transfer point. C88 is scheduled to reach Seaborne at 17:20. The timetable allows exactly seven minutes between R214’s scheduled 16:02 arrival and C88’s scheduled 16:09 departure.\"},{\"speaker\":\"Bellford platform log\",\"text\":\"For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4.\"},{\"speaker\":\"Bellford platform log\",\"text\":\"For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 7.\"},{\"speaker\":\"operations policy\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"booking instruction\",\"text\":\"Assess the requested itinerary using the finalized log and the Bellford connection rule. A connection that exactly meets the applicable minimum is acceptable.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4."}, {"path": ["3", "text"], "text": "For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4.", "negative_left": "For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4.", "negative_right": "For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 4.", "right": "For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "scale-diverse-093-003", "id": "scale-diverse-093-003-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking handoff", "text": "Option A for 14 October 2026 consists of R214 from Alderwick to Bellford followed by Coastliner C88 from Bellford to Seaborne. Bellford is the itinerary’s sole transfer point. C88 is scheduled to reach Seaborne at 17:20. The timetable allows exactly seven minutes between R214’s scheduled 16:02 arrival and C88’s scheduled 16:09 departure."}, {"speaker": "Bellford platform log", "text": "For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4."}, {"speaker": "Bellford platform log", "text": "For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 7."}, {"speaker": "operations policy", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "booking instruction", "text": "Assess the requested itinerary using the finalized log and the Bellford connection rule. A connection that exactly meets the applicable minimum is acceptable."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler’s one-transfer and 18:00-arrival constraints, the Option A route and train identities, and the Bellford minimum-connection policy, including the exact-minimum rule. The two evidence spans are complete factual log sentences. The counterfactual changes only C88’s recorded departure platform from 7 to 4, which is coherent with the unchanged seven-minute interval and creates no duplicate conflicting record. Neither context contains a gold label, answer code, rationale, proposition identifier, or output directive.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking handoff\",\"text\":\"Option A for 14 October 2026 consists of R214 from Alderwick to Bellford followed by Coastliner C88 from Bellford to Seaborne. Bellford is the itinerary’s sole transfer point. C88 is scheduled to reach Seaborne at 17:20. The timetable allows exactly seven minutes between R214’s scheduled 16:02 arrival and C88’s scheduled 16:09 departure.\"},{\"speaker\":\"Bellford platform log\",\"text\":\"For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4.\"},{\"speaker\":\"Bellford platform log\",\"text\":\"For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 7.\"},{\"speaker\":\"operations policy\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"booking instruction\",\"text\":\"Assess the requested itinerary using the finalized log and the Bellford connection rule. A connection that exactly meets the applicable minimum is acceptable.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4."}, {"path": ["3", "text"], "text": "For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4.", "negative_left": "For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4.", "negative_right": "For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 4.", "right": "For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "scale-diverse-093-003", "id": "scale-diverse-093-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking handoff", "text": "Option A for 14 October 2026 consists of R214 from Alderwick to Bellford followed by Coastliner C88 from Bellford to Seaborne. Bellford is the itinerary’s sole transfer point. C88 is scheduled to reach Seaborne at 17:20. The timetable allows exactly seven minutes between R214’s scheduled 16:02 arrival and C88’s scheduled 16:09 departure."}, {"speaker": "Bellford platform log", "text": "For Option A on 14 October 2026, Bellford's finalized platform log records R214 arriving at 16:02 on platform 4."}, {"speaker": "Bellford platform log", "text": "For Option A on 14 October 2026, Bellford's finalized platform log records Coastliner C88 departing at 16:09 from platform 4."}, {"speaker": "operations policy", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "booking instruction", "text": "Assess the requested itinerary using the finalized log and the Bellford connection rule. A connection that exactly meets the applicable minimum is acceptable."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler’s arrival and transfer constraints and Bellford’s same-platform and cross-platform minimum connection rules, while the unchanged questions preserve the decision criteria and exact-minimum instruction. The same Option A, trains, route, connection, destination, and timing bindings are maintained. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes C88’s departure from platform 9 to platform 14, matching R214’s arrival platform, without conflicting with the seven-minute interval or any unchanged fact. Neither context embeds a gold label, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking clerk\",\"text\":\"After the cancellation, Option A uses R214 from Alderwick to Bellford and then Coastliner C88 from Bellford to Seaborne. There are no further changes, so the itinerary requires exactly one transfer. C88 is scheduled to reach Seaborne at 17:20. The scheduled connection interval at Bellford is exactly seven minutes.\"},{\"speaker\":\"platform controller\",\"text\":\"Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14.\"},{\"speaker\":\"platform controller\",\"text\":\"Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 9.\"},{\"speaker\":\"operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please check Option A against my limits and Bellford's applicable minimum connection time before booking it.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14."}, {"path": ["3", "text"], "text": "Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 9."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14.", "negative_left": "Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14.", "negative_right": "Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 14.", "right": "Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 9."}, "verifier_independent_model": false}, "family": "scale-diverse-093-004", "id": "scale-diverse-093-004-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking clerk", "text": "After the cancellation, Option A uses R214 from Alderwick to Bellford and then Coastliner C88 from Bellford to Seaborne. There are no further changes, so the itinerary requires exactly one transfer. C88 is scheduled to reach Seaborne at 17:20. The scheduled connection interval at Bellford is exactly seven minutes."}, {"speaker": "platform controller", "text": "Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14."}, {"speaker": "platform controller", "text": "Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 9."}, {"speaker": "operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please check Option A against my limits and Bellford's applicable minimum connection time before booking it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler’s arrival and transfer constraints and Bellford’s same-platform and cross-platform minimum connection rules, while the unchanged questions preserve the decision criteria and exact-minimum instruction. The same Option A, trains, route, connection, destination, and timing bindings are maintained. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes C88’s departure from platform 9 to platform 14, matching R214’s arrival platform, without conflicting with the seven-minute interval or any unchanged fact. Neither context embeds a gold label, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking clerk\",\"text\":\"After the cancellation, Option A uses R214 from Alderwick to Bellford and then Coastliner C88 from Bellford to Seaborne. There are no further changes, so the itinerary requires exactly one transfer. C88 is scheduled to reach Seaborne at 17:20. The scheduled connection interval at Bellford is exactly seven minutes.\"},{\"speaker\":\"platform controller\",\"text\":\"Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14.\"},{\"speaker\":\"platform controller\",\"text\":\"Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 9.\"},{\"speaker\":\"operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please check Option A against my limits and Bellford's applicable minimum connection time before booking it.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14."}, {"path": ["3", "text"], "text": "Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 9."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14.", "negative_left": "Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14.", "negative_right": "Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 14.", "right": "Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 9."}, "verifier_independent_model": false}, "family": "scale-diverse-093-004", "id": "scale-diverse-093-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking clerk", "text": "After the cancellation, Option A uses R214 from Alderwick to Bellford and then Coastliner C88 from Bellford to Seaborne. There are no further changes, so the itinerary requires exactly one transfer. C88 is scheduled to reach Seaborne at 17:20. The scheduled connection interval at Bellford is exactly seven minutes."}, {"speaker": "platform controller", "text": "Under Option A, Bellford's finalized platform sheet records R214's scheduled 16:02 arrival at platform 14."}, {"speaker": "platform controller", "text": "Under Option A, Bellford's finalized platform sheet records Coastliner C88's scheduled 16:09 departure from platform 14."}, {"speaker": "operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please check Option A against my limits and Bellford's applicable minimum connection time before booking it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same feasibility and ranking policy, request scope, traveler, proposed service, route, destination, and 18:30 deadline binding. The two evidence spans are complete factual sentences. The counterfactual changes only the proposed train’s arrival from eleven minutes before Q7 to six minutes after Q7, without conflicting with any unchanged measurement or assertion. Neither context embeds an answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A5 is one exhaustive-set fact rather than an unrelated bundle. A1 is a factual timing relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the proposed train's relation to the deadline changing. Policy evidence correctly preserves the state-originating constraints, minimum interchange requirement, ranking rule, and support fallback; observations such as the particular ticket and itinerary details need not be preserved as policy evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the proposed train meets the deadline, ticket-validity, transfer, and stair-avoidance requirements. A5 makes the listed alternatives exhaustive; A6 excludes the 17:55 option for exceeding Mara's transfer limit, and A7 excludes the 18:05 option for violating the applicable minimum interchange time. Thus no more suitable feasible alternative remains.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 entails that the proposed train arrives after Mara's 18:30 deadline. It therefore violates an explicit constraint and cannot receive the true outcome, irrespective of the alternatives.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled arrival time at Bellhaven is no later than Mara's 18:30 latest-arrival time."}, {"id": "A2", "statement": "Mara's valid flexible Alderwick-to-Bellhaven ticket is valid for travel on the proposed 15:35 direct Alderwick-to-Bellhaven train."}, {"id": "A3", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train requires zero transfers."}, {"id": "A4", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven itinerary permits Mara to travel without using stairs."}, {"id": "A5", "statement": "The complete set of available replacement itineraries for Mara's canceled 15:10 Alderwick-to-Bellhaven train consists of the proposed 15:35 direct train, the itinerary arriving at 17:55, and the itinerary arriving at 18:05."}, {"id": "A6", "statement": "The Alderwick-to-Bellhaven itinerary arriving at 17:55 requires two transfers."}, {"id": "A7", "statement": "The seven-minute interchange of the Alderwick-to-Bellhaven itinerary arriving at 18:05 is shorter than the applicable ten-minute minimum interchange time."}], "base_state_json": "\"Rebooking case note: Mara’s 15:10 Alderwick-to-Bellhaven train was canceled. The complete replacement inventory contains only the proposed 15:35 direct train, an itinerary arriving at 17:55, and an itinerary arriving at 18:05. Mara’s valid flexible Alderwick-to-Bellhaven ticket is accepted on the proposed train, which requires no transfers and provides a step-free route throughout. The 17:55 itinerary requires two transfers. The 18:05 itinerary has a seven-minute interchange; the applicable standard is the train operations coordinator’s published ten-minute minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs.\\n\\nThe finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival eleven minutes before reference instant Q7. Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven.\\n\\nPolicy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival eleven minutes before reference instant Q7."}, {"path": [], "text": "Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven."}], "policy_evidence": [{"path": [], "text": "She must arrive by 18:30, use no more than one transfer, and avoid stairs."}, {"path": [], "text": "the train operations coordinator’s published ten-minute minimum"}, {"path": [], "text": "Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}], "rules": [{"justification": "The proposed train is ticket-valid, arrives by Mara's deadline, has no transfers, and avoids stairs. The exhaustive alternatives are infeasible: the 17:55 itinerary exceeds the one-transfer limit, and the 18:05 itinerary has an interchange below the applicable minimum. Therefore the proposed train is the most suitable feasible option.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Refutation of A1 entails that the proposed train arrives after Mara's 18:30 latest-arrival time. It therefore violates her arrival constraint and must not be used, regardless of the unchanged deficiencies of the other available itineraries.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival eleven minutes before reference instant Q7.", "negative_left": "The finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival six minutes after reference instant Q7.", "negative_right": "Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven.", "right": "Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven."}, "verifier_independent_model": false}, "family": "scale-diverse-094-002", "id": "scale-diverse-094-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not use the proposed train; route Mara to support because it is infeasible, violates her constraints, or a more suitable compliant option exists.", "true": "Yes — rebook Mara onto the proposed 15:35 direct train because it is the most suitable feasible option under her constraints and the stated policy."}, "instructions": "Decide whether the agent should rebook Mara onto the proposed 15:35 direct train rather than route her to support. Answer yes or no using the stated feasibility and ranking policy.", "type": "noul"}}, "state": "Rebooking case note: Mara’s 15:10 Alderwick-to-Bellhaven train was canceled. The complete replacement inventory contains only the proposed 15:35 direct train, an itinerary arriving at 17:55, and an itinerary arriving at 18:05. Mara’s valid flexible Alderwick-to-Bellhaven ticket is accepted on the proposed train, which requires no transfers and provides a step-free route throughout. The 17:55 itinerary requires two transfers. The 18:05 itinerary has a seven-minute interchange; the applicable standard is the train operations coordinator’s published ten-minute minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs.\n\nThe finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival eleven minutes before reference instant Q7. Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven.\n\nPolicy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}, "method": "c2d", "provenance": {"source_id": "diverse-094", "source_is_synthetic": true, "source_sha256": "7c5946099125dac256dc280abea7b852aa82ad5f37dd617c38d101dd520d00f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same feasibility and ranking policy, request scope, traveler, proposed service, route, destination, and 18:30 deadline binding. The two evidence spans are complete factual sentences. The counterfactual changes only the proposed train’s arrival from eleven minutes before Q7 to six minutes after Q7, without conflicting with any unchanged measurement or assertion. Neither context embeds an answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A5 is one exhaustive-set fact rather than an unrelated bundle. A1 is a factual timing relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the proposed train's relation to the deadline changing. Policy evidence correctly preserves the state-originating constraints, minimum interchange requirement, ranking rule, and support fallback; observations such as the particular ticket and itinerary details need not be preserved as policy evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the proposed train meets the deadline, ticket-validity, transfer, and stair-avoidance requirements. A5 makes the listed alternatives exhaustive; A6 excludes the 17:55 option for exceeding Mara's transfer limit, and A7 excludes the 18:05 option for violating the applicable minimum interchange time. Thus no more suitable feasible alternative remains.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 entails that the proposed train arrives after Mara's 18:30 deadline. It therefore violates an explicit constraint and cannot receive the true outcome, irrespective of the alternatives.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled arrival time at Bellhaven is no later than Mara's 18:30 latest-arrival time."}, {"id": "A2", "statement": "Mara's valid flexible Alderwick-to-Bellhaven ticket is valid for travel on the proposed 15:35 direct Alderwick-to-Bellhaven train."}, {"id": "A3", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train requires zero transfers."}, {"id": "A4", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven itinerary permits Mara to travel without using stairs."}, {"id": "A5", "statement": "The complete set of available replacement itineraries for Mara's canceled 15:10 Alderwick-to-Bellhaven train consists of the proposed 15:35 direct train, the itinerary arriving at 17:55, and the itinerary arriving at 18:05."}, {"id": "A6", "statement": "The Alderwick-to-Bellhaven itinerary arriving at 17:55 requires two transfers."}, {"id": "A7", "statement": "The seven-minute interchange of the Alderwick-to-Bellhaven itinerary arriving at 18:05 is shorter than the applicable ten-minute minimum interchange time."}], "base_state_json": "\"Rebooking case note: Mara’s 15:10 Alderwick-to-Bellhaven train was canceled. The complete replacement inventory contains only the proposed 15:35 direct train, an itinerary arriving at 17:55, and an itinerary arriving at 18:05. Mara’s valid flexible Alderwick-to-Bellhaven ticket is accepted on the proposed train, which requires no transfers and provides a step-free route throughout. The 17:55 itinerary requires two transfers. The 18:05 itinerary has a seven-minute interchange; the applicable standard is the train operations coordinator’s published ten-minute minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs.\\n\\nThe finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival eleven minutes before reference instant Q7. Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven.\\n\\nPolicy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival eleven minutes before reference instant Q7."}, {"path": [], "text": "Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven."}], "policy_evidence": [{"path": [], "text": "She must arrive by 18:30, use no more than one transfer, and avoid stairs."}, {"path": [], "text": "the train operations coordinator’s published ten-minute minimum"}, {"path": [], "text": "Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}], "rules": [{"justification": "The proposed train is ticket-valid, arrives by Mara's deadline, has no transfers, and avoids stairs. The exhaustive alternatives are infeasible: the 17:55 itinerary exceeds the one-transfer limit, and the 18:05 itinerary has an interchange below the applicable minimum. Therefore the proposed train is the most suitable feasible option.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Refutation of A1 entails that the proposed train arrives after Mara's 18:30 latest-arrival time. It therefore violates her arrival constraint and must not be used, regardless of the unchanged deficiencies of the other available itineraries.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival eleven minutes before reference instant Q7.", "negative_left": "The finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival six minutes after reference instant Q7.", "negative_right": "Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven.", "right": "Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven."}, "verifier_independent_model": false}, "family": "scale-diverse-094-002", "id": "scale-diverse-094-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not use the proposed train; route Mara to support because it is infeasible, violates her constraints, or a more suitable compliant option exists.", "true": "Yes — rebook Mara onto the proposed 15:35 direct train because it is the most suitable feasible option under her constraints and the stated policy."}, "instructions": "Decide whether the agent should rebook Mara onto the proposed 15:35 direct train rather than route her to support. Answer yes or no using the stated feasibility and ranking policy.", "type": "noul"}}, "state": "Rebooking case note: Mara’s 15:10 Alderwick-to-Bellhaven train was canceled. The complete replacement inventory contains only the proposed 15:35 direct train, an itinerary arriving at 17:55, and an itinerary arriving at 18:05. Mara’s valid flexible Alderwick-to-Bellhaven ticket is accepted on the proposed train, which requires no transfers and provides a step-free route throughout. The 17:55 itinerary requires two transfers. The 18:05 itinerary has a seven-minute interchange; the applicable standard is the train operations coordinator’s published ten-minute minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs.\n\nThe finalized operating record places the proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled Bellhaven arrival six minutes after reference instant Q7. Reference instant Q7 in Mara's travel record is 18:30 on the service date, the latest time at which Mara may arrive at Bellhaven.\n\nPolicy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}, "method": "c2d", "provenance": {"source_id": "diverse-094", "source_is_synthetic": true, "source_sha256": "7c5946099125dac256dc280abea7b852aa82ad5f37dd617c38d101dd520d00f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the feasibility, constraint, ranking, and support-routing policy and retain Mara, the proposed 15:35 Alderwick-to-Bellhaven train, and the 17 September 2026 timing scope. The two evidence spans are complete factual sentences. The counterfactual changes only the proposed train’s arrival relative to marker H, implying 18:36 rather than the base context’s 18:19, without conflicting with any unchanged measurement or assertion. Neither constructed context contains a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A5 is one exhaustive-set fact rather than an unrelated bundle. A1 is a factual timing relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the proposed train's relation to the deadline changing. Policy evidence correctly preserves the state-originating constraints, minimum interchange requirement, ranking rule, and support fallback; observations such as the particular ticket and itinerary details need not be preserved as policy evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the proposed train meets the deadline, ticket-validity, transfer, and stair-avoidance requirements. A5 makes the listed alternatives exhaustive; A6 excludes the 17:55 option for exceeding Mara's transfer limit, and A7 excludes the 18:05 option for violating the applicable minimum interchange time. Thus no more suitable feasible alternative remains.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 entails that the proposed train arrives after Mara's 18:30 deadline. It therefore violates an explicit constraint and cannot receive the true outcome, irrespective of the alternatives.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled arrival time at Bellhaven is no later than Mara's 18:30 latest-arrival time."}, {"id": "A2", "statement": "Mara's valid flexible Alderwick-to-Bellhaven ticket is valid for travel on the proposed 15:35 direct Alderwick-to-Bellhaven train."}, {"id": "A3", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train requires zero transfers."}, {"id": "A4", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven itinerary permits Mara to travel without using stairs."}, {"id": "A5", "statement": "The complete set of available replacement itineraries for Mara's canceled 15:10 Alderwick-to-Bellhaven train consists of the proposed 15:35 direct train, the itinerary arriving at 17:55, and the itinerary arriving at 18:05."}, {"id": "A6", "statement": "The Alderwick-to-Bellhaven itinerary arriving at 17:55 requires two transfers."}, {"id": "A7", "statement": "The seven-minute interchange of the Alderwick-to-Bellhaven itinerary arriving at 18:05 is shorter than the applicable ten-minute minimum interchange time."}], "base_state_json": "\"At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 11 minutes before marker H. For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time. Mara’s 15:10 service was canceled, and her valid flexible Alderwick-to-Bellhaven ticket is valid on the proposed replacement. That service is direct, requires zero transfers, and provides a step-free itinerary throughout. The complete replacement inventory contains exactly three itineraries: the proposed 15:35 direct train, an option arriving at 17:55, and an option arriving at 18:05. The 17:55 option requires two transfers. The 18:05 option has a seven-minute interchange, which is shorter than the applicable minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs. The applicable standard is the train operations coordinator’s published ten-minute minimum. Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 11 minutes before marker H."}, {"path": [], "text": "For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time."}], "policy_evidence": [{"path": [], "text": "She must arrive by 18:30, use no more than one transfer, and avoid stairs."}, {"path": [], "text": "the train operations coordinator’s published ten-minute minimum"}, {"path": [], "text": "Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}], "rules": [{"justification": "The proposed train is ticket-valid, arrives by Mara's deadline, has no transfers, and avoids stairs. The exhaustive alternatives are infeasible: the 17:55 itinerary exceeds the one-transfer limit, and the 18:05 itinerary has an interchange below the applicable minimum. Therefore the proposed train is the most suitable feasible option.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Refutation of A1 entails that the proposed train arrives after Mara's 18:30 latest-arrival time. It therefore violates her arrival constraint and must not be used, regardless of the unchanged deficiencies of the other available itineraries.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 11 minutes before marker H.", "negative_left": "At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 6 minutes after marker H.", "negative_right": "For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time.", "right": "For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time."}, "verifier_independent_model": false}, "family": "scale-diverse-094-003", "id": "scale-diverse-094-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not use the proposed train; route Mara to support because it is infeasible, violates her constraints, or a more suitable compliant option exists.", "true": "Yes — rebook Mara onto the proposed 15:35 direct train because it is the most suitable feasible option under her constraints and the stated policy."}, "instructions": "Decide whether the agent should rebook Mara onto the proposed 15:35 direct train rather than route her to support. Answer yes or no using the stated feasibility and ranking policy.", "type": "noul"}}, "state": "At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 11 minutes before marker H. For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time. Mara’s 15:10 service was canceled, and her valid flexible Alderwick-to-Bellhaven ticket is valid on the proposed replacement. That service is direct, requires zero transfers, and provides a step-free itinerary throughout. The complete replacement inventory contains exactly three itineraries: the proposed 15:35 direct train, an option arriving at 17:55, and an option arriving at 18:05. The 17:55 option requires two transfers. The 18:05 option has a seven-minute interchange, which is shorter than the applicable minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs. The applicable standard is the train operations coordinator’s published ten-minute minimum. Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}, "method": "c2d", "provenance": {"source_id": "diverse-094", "source_is_synthetic": true, "source_sha256": "7c5946099125dac256dc280abea7b852aa82ad5f37dd617c38d101dd520d00f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the feasibility, constraint, ranking, and support-routing policy and retain Mara, the proposed 15:35 Alderwick-to-Bellhaven train, and the 17 September 2026 timing scope. The two evidence spans are complete factual sentences. The counterfactual changes only the proposed train’s arrival relative to marker H, implying 18:36 rather than the base context’s 18:19, without conflicting with any unchanged measurement or assertion. Neither constructed context contains a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A5 is one exhaustive-set fact rather than an unrelated bundle. A1 is a factual timing relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the proposed train's relation to the deadline changing. Policy evidence correctly preserves the state-originating constraints, minimum interchange requirement, ranking rule, and support fallback; observations such as the particular ticket and itinerary details need not be preserved as policy evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the proposed train meets the deadline, ticket-validity, transfer, and stair-avoidance requirements. A5 makes the listed alternatives exhaustive; A6 excludes the 17:55 option for exceeding Mara's transfer limit, and A7 excludes the 18:05 option for violating the applicable minimum interchange time. Thus no more suitable feasible alternative remains.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 entails that the proposed train arrives after Mara's 18:30 deadline. It therefore violates an explicit constraint and cannot receive the true outcome, irrespective of the alternatives.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled arrival time at Bellhaven is no later than Mara's 18:30 latest-arrival time."}, {"id": "A2", "statement": "Mara's valid flexible Alderwick-to-Bellhaven ticket is valid for travel on the proposed 15:35 direct Alderwick-to-Bellhaven train."}, {"id": "A3", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train requires zero transfers."}, {"id": "A4", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven itinerary permits Mara to travel without using stairs."}, {"id": "A5", "statement": "The complete set of available replacement itineraries for Mara's canceled 15:10 Alderwick-to-Bellhaven train consists of the proposed 15:35 direct train, the itinerary arriving at 17:55, and the itinerary arriving at 18:05."}, {"id": "A6", "statement": "The Alderwick-to-Bellhaven itinerary arriving at 17:55 requires two transfers."}, {"id": "A7", "statement": "The seven-minute interchange of the Alderwick-to-Bellhaven itinerary arriving at 18:05 is shorter than the applicable ten-minute minimum interchange time."}], "base_state_json": "\"At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 11 minutes before marker H. For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time. Mara’s 15:10 service was canceled, and her valid flexible Alderwick-to-Bellhaven ticket is valid on the proposed replacement. That service is direct, requires zero transfers, and provides a step-free itinerary throughout. The complete replacement inventory contains exactly three itineraries: the proposed 15:35 direct train, an option arriving at 17:55, and an option arriving at 18:05. The 17:55 option requires two transfers. The 18:05 option has a seven-minute interchange, which is shorter than the applicable minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs. The applicable standard is the train operations coordinator’s published ten-minute minimum. Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 11 minutes before marker H."}, {"path": [], "text": "For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time."}], "policy_evidence": [{"path": [], "text": "She must arrive by 18:30, use no more than one transfer, and avoid stairs."}, {"path": [], "text": "the train operations coordinator’s published ten-minute minimum"}, {"path": [], "text": "Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}], "rules": [{"justification": "The proposed train is ticket-valid, arrives by Mara's deadline, has no transfers, and avoids stairs. The exhaustive alternatives are infeasible: the 17:55 itinerary exceeds the one-transfer limit, and the 18:05 itinerary has an interchange below the applicable minimum. Therefore the proposed train is the most suitable feasible option.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Refutation of A1 entails that the proposed train arrives after Mara's 18:30 latest-arrival time. It therefore violates her arrival constraint and must not be used, regardless of the unchanged deficiencies of the other available itineraries.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 11 minutes before marker H.", "negative_left": "At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 6 minutes after marker H.", "negative_right": "For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time.", "right": "For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time."}, "verifier_independent_model": false}, "family": "scale-diverse-094-003", "id": "scale-diverse-094-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not use the proposed train; route Mara to support because it is infeasible, violates her constraints, or a more suitable compliant option exists.", "true": "Yes — rebook Mara onto the proposed 15:35 direct train because it is the most suitable feasible option under her constraints and the stated policy."}, "instructions": "Decide whether the agent should rebook Mara onto the proposed 15:35 direct train rather than route her to support. Answer yes or no using the stated feasibility and ranking policy.", "type": "noul"}}, "state": "At the 14:42 operational handoff on 17 September 2026, the timetable entry for the proposed 15:35 direct Alderwick-to-Bellhaven train placed its scheduled Bellhaven arrival 6 minutes after marker H. For Mara's Alderwick-to-Bellhaven replacement journey on 17 September 2026, marker H denotes 18:30 local time at Bellhaven, exactly her latest permissible arrival time. Mara’s 15:10 service was canceled, and her valid flexible Alderwick-to-Bellhaven ticket is valid on the proposed replacement. That service is direct, requires zero transfers, and provides a step-free itinerary throughout. The complete replacement inventory contains exactly three itineraries: the proposed 15:35 direct train, an option arriving at 17:55, and an option arriving at 18:05. The 17:55 option requires two transfers. The 18:05 option has a seven-minute interchange, which is shorter than the applicable minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs. The applicable standard is the train operations coordinator’s published ten-minute minimum. Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}, "method": "c2d", "provenance": {"source_id": "diverse-094", "source_is_synthetic": true, "source_sha256": "7c5946099125dac256dc280abea7b852aa82ad5f37dd617c38d101dd520d00f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing ranking and support-routing policy, while the unchanged questions preserve the decision criteria and instructions. Mara, the Alderwick-to-Bellhaven route, the proposed 15:35 direct train, her 18:30 deadline, and the rebook-versus-support decision remain bound correctly. The two evidence spans are complete factual sentences rather than instructions or policy definitions. Changing AB-742’s arrival from 18:24 to 18:39 is coherent with the remaining facts and creates no duplicate contradictory measurement. Neither context contains an answer code, gold label, proposition identifier, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A5 is one exhaustive-set fact rather than an unrelated bundle. A1 is a factual timing relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the proposed train's relation to the deadline changing. Policy evidence correctly preserves the state-originating constraints, minimum interchange requirement, ranking rule, and support fallback; observations such as the particular ticket and itinerary details need not be preserved as policy evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the proposed train meets the deadline, ticket-validity, transfer, and stair-avoidance requirements. A5 makes the listed alternatives exhaustive; A6 excludes the 17:55 option for exceeding Mara's transfer limit, and A7 excludes the 18:05 option for violating the applicable minimum interchange time. Thus no more suitable feasible alternative remains.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 entails that the proposed train arrives after Mara's 18:30 deadline. It therefore violates an explicit constraint and cannot receive the true outcome, irrespective of the alternatives.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled arrival time at Bellhaven is no later than Mara's 18:30 latest-arrival time."}, {"id": "A2", "statement": "Mara's valid flexible Alderwick-to-Bellhaven ticket is valid for travel on the proposed 15:35 direct Alderwick-to-Bellhaven train."}, {"id": "A3", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train requires zero transfers."}, {"id": "A4", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven itinerary permits Mara to travel without using stairs."}, {"id": "A5", "statement": "The complete set of available replacement itineraries for Mara's canceled 15:10 Alderwick-to-Bellhaven train consists of the proposed 15:35 direct train, the itinerary arriving at 17:55, and the itinerary arriving at 18:05."}, {"id": "A6", "statement": "The Alderwick-to-Bellhaven itinerary arriving at 17:55 requires two transfers."}, {"id": "A7", "statement": "The seven-minute interchange of the Alderwick-to-Bellhaven itinerary arriving at 18:05 is shorter than the applicable ten-minute minimum interchange time."}], "base_state_json": "\"Field note: Mara’s 15:10 Alderwick-to-Bellhaven train was canceled. Her valid flexible Alderwick-to-Bellhaven ticket is valid on the proposed 15:35 direct replacement, which requires zero transfers and provides a step-free route throughout. The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train. The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:24, while Mara's recorded latest-arrival time is 18:30. The completed replacement search found exactly three available itineraries: the proposed service, one arriving at 17:55, and one arriving at 18:05. The 17:55 itinerary requires two transfers. The 18:05 itinerary’s seven-minute interchange is shorter than the applicable ten-minute minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs. The applicable threshold is the train operations coordinator’s published ten-minute minimum. Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train."}, {"path": [], "text": "The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:24, while Mara's recorded latest-arrival time is 18:30."}], "policy_evidence": [{"path": [], "text": "She must arrive by 18:30, use no more than one transfer, and avoid stairs."}, {"path": [], "text": "the train operations coordinator’s published ten-minute minimum"}, {"path": [], "text": "Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}], "rules": [{"justification": "The proposed train is ticket-valid, arrives by Mara's deadline, has no transfers, and avoids stairs. The exhaustive alternatives are infeasible: the 17:55 itinerary exceeds the one-transfer limit, and the 18:05 itinerary has an interchange below the applicable minimum. Therefore the proposed train is the most suitable feasible option.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Refutation of A1 entails that the proposed train arrives after Mara's 18:30 latest-arrival time. It therefore violates her arrival constraint and must not be used, regardless of the unchanged deficiencies of the other available itineraries.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train.", "negative_left": "The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train.", "negative_right": "The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:39, while Mara's recorded latest-arrival time is 18:30.", "right": "The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:24, while Mara's recorded latest-arrival time is 18:30."}, "verifier_independent_model": false}, "family": "scale-diverse-094-004", "id": "scale-diverse-094-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not use the proposed train; route Mara to support because it is infeasible, violates her constraints, or a more suitable compliant option exists.", "true": "Yes — rebook Mara onto the proposed 15:35 direct train because it is the most suitable feasible option under her constraints and the stated policy."}, "instructions": "Decide whether the agent should rebook Mara onto the proposed 15:35 direct train rather than route her to support. Answer yes or no using the stated feasibility and ranking policy.", "type": "noul"}}, "state": "Field note: Mara’s 15:10 Alderwick-to-Bellhaven train was canceled. Her valid flexible Alderwick-to-Bellhaven ticket is valid on the proposed 15:35 direct replacement, which requires zero transfers and provides a step-free route throughout. The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train. The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:24, while Mara's recorded latest-arrival time is 18:30. The completed replacement search found exactly three available itineraries: the proposed service, one arriving at 17:55, and one arriving at 18:05. The 17:55 itinerary requires two transfers. The 18:05 itinerary’s seven-minute interchange is shorter than the applicable ten-minute minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs. The applicable threshold is the train operations coordinator’s published ten-minute minimum. Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}, "method": "c2d", "provenance": {"source_id": "diverse-094", "source_is_synthetic": true, "source_sha256": "7c5946099125dac256dc280abea7b852aa82ad5f37dd617c38d101dd520d00f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing ranking and support-routing policy, while the unchanged questions preserve the decision criteria and instructions. Mara, the Alderwick-to-Bellhaven route, the proposed 15:35 direct train, her 18:30 deadline, and the rebook-versus-support decision remain bound correctly. The two evidence spans are complete factual sentences rather than instructions or policy definitions. Changing AB-742’s arrival from 18:24 to 18:39 is coherent with the remaining facts and creates no duplicate contradictory measurement. Neither context contains an answer code, gold label, proposition identifier, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A5 is one exhaustive-set fact rather than an unrelated bundle. A1 is a factual timing relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the proposed train's relation to the deadline changing. Policy evidence correctly preserves the state-originating constraints, minimum interchange requirement, ranking rule, and support fallback; observations such as the particular ticket and itinerary details need not be preserved as policy evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the proposed train meets the deadline, ticket-validity, transfer, and stair-avoidance requirements. A5 makes the listed alternatives exhaustive; A6 excludes the 17:55 option for exceeding Mara's transfer limit, and A7 excludes the 18:05 option for violating the applicable minimum interchange time. Thus no more suitable feasible alternative remains.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 entails that the proposed train arrives after Mara's 18:30 deadline. It therefore violates an explicit constraint and cannot receive the true outcome, irrespective of the alternatives.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train's scheduled arrival time at Bellhaven is no later than Mara's 18:30 latest-arrival time."}, {"id": "A2", "statement": "Mara's valid flexible Alderwick-to-Bellhaven ticket is valid for travel on the proposed 15:35 direct Alderwick-to-Bellhaven train."}, {"id": "A3", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven train requires zero transfers."}, {"id": "A4", "statement": "The proposed 15:35 direct Alderwick-to-Bellhaven itinerary permits Mara to travel without using stairs."}, {"id": "A5", "statement": "The complete set of available replacement itineraries for Mara's canceled 15:10 Alderwick-to-Bellhaven train consists of the proposed 15:35 direct train, the itinerary arriving at 17:55, and the itinerary arriving at 18:05."}, {"id": "A6", "statement": "The Alderwick-to-Bellhaven itinerary arriving at 17:55 requires two transfers."}, {"id": "A7", "statement": "The seven-minute interchange of the Alderwick-to-Bellhaven itinerary arriving at 18:05 is shorter than the applicable ten-minute minimum interchange time."}], "base_state_json": "\"Field note: Mara’s 15:10 Alderwick-to-Bellhaven train was canceled. Her valid flexible Alderwick-to-Bellhaven ticket is valid on the proposed 15:35 direct replacement, which requires zero transfers and provides a step-free route throughout. The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train. The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:24, while Mara's recorded latest-arrival time is 18:30. The completed replacement search found exactly three available itineraries: the proposed service, one arriving at 17:55, and one arriving at 18:05. The 17:55 itinerary requires two transfers. The 18:05 itinerary’s seven-minute interchange is shorter than the applicable ten-minute minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs. The applicable threshold is the train operations coordinator’s published ten-minute minimum. Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train."}, {"path": [], "text": "The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:24, while Mara's recorded latest-arrival time is 18:30."}], "policy_evidence": [{"path": [], "text": "She must arrive by 18:30, use no more than one transfer, and avoid stairs."}, {"path": [], "text": "the train operations coordinator’s published ten-minute minimum"}, {"path": [], "text": "Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}], "rules": [{"justification": "The proposed train is ticket-valid, arrives by Mara's deadline, has no transfers, and avoids stairs. The exhaustive alternatives are infeasible: the 17:55 itinerary exceeds the one-transfer limit, and the 18:05 itinerary has an interchange below the applicable minimum. Therefore the proposed train is the most suitable feasible option.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Refutation of A1 entails that the proposed train arrives after Mara's 18:30 latest-arrival time. It therefore violates her arrival constraint and must not be used, regardless of the unchanged deficiencies of the other available itineraries.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train.", "negative_left": "The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train.", "negative_right": "The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:39, while Mara's recorded latest-arrival time is 18:30.", "right": "The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:24, while Mara's recorded latest-arrival time is 18:30."}, "verifier_independent_model": false}, "family": "scale-diverse-094-004", "id": "scale-diverse-094-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not use the proposed train; route Mara to support because it is infeasible, violates her constraints, or a more suitable compliant option exists.", "true": "Yes — rebook Mara onto the proposed 15:35 direct train because it is the most suitable feasible option under her constraints and the stated policy."}, "instructions": "Decide whether the agent should rebook Mara onto the proposed 15:35 direct train rather than route her to support. Answer yes or no using the stated feasibility and ranking policy.", "type": "noul"}}, "state": "Field note: Mara’s 15:10 Alderwick-to-Bellhaven train was canceled. Her valid flexible Alderwick-to-Bellhaven ticket is valid on the proposed 15:35 direct replacement, which requires zero transfers and provides a step-free route throughout. The finalized 12 October 2026 timetable assigns the unique service identifier AB-742 to the proposed 15:35 direct Alderwick-to-Bellhaven train. The finalized 12 October 2026 timetable schedules service AB-742 to arrive at Bellhaven at 18:39, while Mara's recorded latest-arrival time is 18:30. The completed replacement search found exactly three available itineraries: the proposed service, one arriving at 17:55, and one arriving at 18:05. The 17:55 itinerary requires two transfers. The 18:05 itinerary’s seven-minute interchange is shorter than the applicable ten-minute minimum. She must arrive by 18:30, use no more than one transfer, and avoid stairs. The applicable threshold is the train operations coordinator’s published ten-minute minimum. Policy ranks feasible, constraint-compliant options by earlier arrival, then fewer transfers; otherwise route the traveler to support."}, "method": "c2d", "provenance": {"source_id": "diverse-094", "source_is_synthetic": true, "source_sha256": "7c5946099125dac256dc280abea7b852aa82ad5f37dd617c38d101dd520d00f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric, instructions, request, Birch entity, Lanton-to-Merrow journey, and confirmed constraint interpretation. The two focus-evidence spans are complete factual sentences. Changing Birch’s arrival from 18:24 to 18:47 is coherent with the unchanged statement that Alder arrives at 18:21 earlier than Birch, and neither context contains an answer, code, rule table, proposition identifier, label rationale, or output instruction beyond the governing rubric.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.\",\"1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.\",\"2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority.\"],\"instructions\":\"Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.\",\"type\":\"score\"}},\"state\":{\"context\":\"After the cancellation, the operations coordinator recorded Alder and Birch as operating replacement routes from Lanton to Merrow. The traveler’s flexible ticket with the cancellation endorsement is valid on both. Each route requires one transfer. Alder has a 13-minute interchange and is scheduled to arrive at 18:21, earlier than Birch. Birch’s 14-minute connection exceeds the published 12-minute minimum. The traveler confirmed that the agent’s stated transfer limit and arrival cutoff were the only hard constraints. They also confirmed having no preference or priority that should override earliest-arrival ordering among eligible routes.\",\"evidence\":[\"For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow.\",\"Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:24 that day.\",\"For that journey, the same confirmed paraphrase specified one as the maximum permissible number of transfers.\"],\"request\":\"Using the rubric, score Birch’s suitability for this traveler.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow."}, {"path": ["state", "evidence", "1"], "text": "Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:24 that day."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow.", "negative_left": "For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow.", "negative_right": "Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:47 that day.", "right": "Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:24 that day."}, "verifier_independent_model": false}, "family": "scale-diverse-095-001", "id": "scale-diverse-095-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "After the cancellation, the operations coordinator recorded Alder and Birch as operating replacement routes from Lanton to Merrow. The traveler’s flexible ticket with the cancellation endorsement is valid on both. Each route requires one transfer. Alder has a 13-minute interchange and is scheduled to arrive at 18:21, earlier than Birch. Birch’s 14-minute connection exceeds the published 12-minute minimum. The traveler confirmed that the agent’s stated transfer limit and arrival cutoff were the only hard constraints. They also confirmed having no preference or priority that should override earliest-arrival ordering among eligible routes.", "evidence": ["For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow.", "Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:24 that day.", "For that journey, the same confirmed paraphrase specified one as the maximum permissible number of transfers."], "request": "Using the rubric, score Birch’s suitability for this traveler."}}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric, instructions, request, Birch entity, Lanton-to-Merrow journey, and confirmed constraint interpretation. The two focus-evidence spans are complete factual sentences. Changing Birch’s arrival from 18:24 to 18:47 is coherent with the unchanged statement that Alder arrives at 18:21 earlier than Birch, and neither context contains an answer, code, rule table, proposition identifier, label rationale, or output instruction beyond the governing rubric.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.\",\"1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.\",\"2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority.\"],\"instructions\":\"Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.\",\"type\":\"score\"}},\"state\":{\"context\":\"After the cancellation, the operations coordinator recorded Alder and Birch as operating replacement routes from Lanton to Merrow. The traveler’s flexible ticket with the cancellation endorsement is valid on both. Each route requires one transfer. Alder has a 13-minute interchange and is scheduled to arrive at 18:21, earlier than Birch. Birch’s 14-minute connection exceeds the published 12-minute minimum. The traveler confirmed that the agent’s stated transfer limit and arrival cutoff were the only hard constraints. They also confirmed having no preference or priority that should override earliest-arrival ordering among eligible routes.\",\"evidence\":[\"For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow.\",\"Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:24 that day.\",\"For that journey, the same confirmed paraphrase specified one as the maximum permissible number of transfers.\"],\"request\":\"Using the rubric, score Birch’s suitability for this traveler.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow."}, {"path": ["state", "evidence", "1"], "text": "Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:24 that day."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow.", "negative_left": "For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow.", "negative_right": "Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:47 that day.", "right": "Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:24 that day."}, "verifier_independent_model": false}, "family": "scale-diverse-095-001", "id": "scale-diverse-095-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "After the cancellation, the operations coordinator recorded Alder and Birch as operating replacement routes from Lanton to Merrow. The traveler’s flexible ticket with the cancellation endorsement is valid on both. Each route requires one transfer. Alder has a 13-minute interchange and is scheduled to arrive at 18:21, earlier than Birch. Birch’s 14-minute connection exceeds the published 12-minute minimum. The traveler confirmed that the agent’s stated transfer limit and arrival cutoff were the only hard constraints. They also confirmed having no preference or priority that should override earliest-arrival ordering among eligible routes.", "evidence": ["For the replacement journey from Lanton to Merrow on 14 October 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 that day as the latest permissible arrival at Merrow.", "Birch’s schedule for the replacement journey from Lanton to Merrow on 14 October 2026 lists its arrival at Merrow as 18:47 that day.", "For that journey, the same confirmed paraphrase specified one as the maximum permissible number of transfers."], "request": "Using the rubric, score Birch’s suitability for this traveler."}}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original suitability rubric and its earliest-arrival-then-transfer ordering without adding exceptions or default rules. The request remains about scoring Birch for the same traveler and Lanton-to-Merrow replacement journey. The two focus spans are complete factual sentences. The counterfactual changes only Birch’s scheduled arrival from 18:24 to 18:36, which is consistent with Alder arriving at 18:18 and creates no duplicate conflicting measurement. Neither context states a score, answer code, proposition identifier, rule table, label rationale, or output instruction beyond the unchanged request to use the rubric.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"A cancellation has left a traveler needing an endorsed replacement from Lanton to Merrow. Operations records identify Birch and Alder as operating replacement routes. The traveler’s flexible ticket bearing the cancellation endorsement is valid on both.\",\"evidence\":[\"At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow.\",\"The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:24.\",\"Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"Alder has a 15-minute interchange, also exceeding the published 12-minute minimum.\",\"Birch and Alder each require one transfer. Alder is scheduled to arrive at Merrow at 18:18, earlier than Birch.\",\"The traveler confirmed that the agent’s paraphrased maximum was one transfer. They also confirmed having no hard constraints beyond that limit and the latest arrival time, and no priority overriding earliest-arrival ordering among eligible routes.\"],\"request\":\"Using the rubric, score Birch’s suitability for this traveler.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow."}, {"path": ["evidence", "1"], "text": "The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:24."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow.", "negative_left": "At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow.", "negative_right": "The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:36.", "right": "The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:24."}, "verifier_independent_model": false}, "family": "scale-diverse-095-002", "id": "scale-diverse-095-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "A cancellation has left a traveler needing an endorsed replacement from Lanton to Merrow. Operations records identify Birch and Alder as operating replacement routes. The traveler’s flexible ticket bearing the cancellation endorsement is valid on both.", "evidence": ["At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow.", "The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:24.", "Birch’s 14-minute connection exceeds the published 12-minute minimum.", "Alder has a 15-minute interchange, also exceeding the published 12-minute minimum.", "Birch and Alder each require one transfer. Alder is scheduled to arrive at Merrow at 18:18, earlier than Birch.", "The traveler confirmed that the agent’s paraphrased maximum was one transfer. They also confirmed having no hard constraints beyond that limit and the latest arrival time, and no priority overriding earliest-arrival ordering among eligible routes."], "request": "Using the rubric, score Birch’s suitability for this traveler."}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original suitability rubric and its earliest-arrival-then-transfer ordering without adding exceptions or default rules. The request remains about scoring Birch for the same traveler and Lanton-to-Merrow replacement journey. The two focus spans are complete factual sentences. The counterfactual changes only Birch’s scheduled arrival from 18:24 to 18:36, which is consistent with Alder arriving at 18:18 and creates no duplicate conflicting measurement. Neither context states a score, answer code, proposition identifier, rule table, label rationale, or output instruction beyond the unchanged request to use the rubric.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"A cancellation has left a traveler needing an endorsed replacement from Lanton to Merrow. Operations records identify Birch and Alder as operating replacement routes. The traveler’s flexible ticket bearing the cancellation endorsement is valid on both.\",\"evidence\":[\"At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow.\",\"The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:24.\",\"Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"Alder has a 15-minute interchange, also exceeding the published 12-minute minimum.\",\"Birch and Alder each require one transfer. Alder is scheduled to arrive at Merrow at 18:18, earlier than Birch.\",\"The traveler confirmed that the agent’s paraphrased maximum was one transfer. They also confirmed having no hard constraints beyond that limit and the latest arrival time, and no priority overriding earliest-arrival ordering among eligible routes.\"],\"request\":\"Using the rubric, score Birch’s suitability for this traveler.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow."}, {"path": ["evidence", "1"], "text": "The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:24."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow.", "negative_left": "At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow.", "negative_right": "The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:36.", "right": "The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:24."}, "verifier_independent_model": false}, "family": "scale-diverse-095-002", "id": "scale-diverse-095-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "A cancellation has left a traveler needing an endorsed replacement from Lanton to Merrow. Operations records identify Birch and Alder as operating replacement routes. The traveler’s flexible ticket bearing the cancellation endorsement is valid on both.", "evidence": ["At 09:12 on 6 October 2026, the traveler confirmed that the rebooking agent’s paraphrase set 18:30 on 8 October 2026 as the latest permissible arrival time at Merrow for the replacement journey from Lanton to Merrow.", "The finalized timetable for Birch’s replacement journey from Lanton to Merrow on 8 October 2026 lists its scheduled arrival at Merrow as 18:36.", "Birch’s 14-minute connection exceeds the published 12-minute minimum.", "Alder has a 15-minute interchange, also exceeding the published 12-minute minimum.", "Birch and Alder each require one transfer. Alder is scheduled to arrive at Merrow at 18:18, earlier than Birch.", "The traveler confirmed that the agent’s paraphrased maximum was one transfer. They also confirmed having no hard constraints beyond that limit and the latest arrival time, and no priority overriding earliest-arrival ordering among eligible routes."], "request": "Using the rubric, score Birch’s suitability for this traveler."}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the full scoring policy, while both contexts retain the request to score Birch for the same traveler and Lanton-to-Merrow journey. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Birch’s arrival from 18:17 to 18:43; Alder’s 18:12 arrival remains earlier in either case, so no duplicate measurement or unchanged assertion conflicts with the change. Neither context states a score, answer code, proposition identifier, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"A cancellation has triggered an operational handoff for the traveler’s replacement journey from Lanton to Merrow.\",\"evidence\":[\"In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow.\",\"The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:17.\",\"The same confirmation set one transfer as the maximum.\",\"The traveler confirmed these were the only hard constraints and that no other preference should displace the rubric’s earliest-arrival ordering.\",\"The dispatcher confirms Birch and Alder are operating replacement routes from Lanton to Merrow, each with one transfer.\",\"The traveler’s flexible ticket with the cancellation endorsement is valid on both routes.\",\"Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"Alder has a 13-minute interchange and is scheduled to reach Merrow at 18:12.\",\"The operations comparison records Alder as arriving earlier than Birch.\"],\"request\":\"Using the rubric, score Birch’s suitability for this traveler.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow."}, {"path": ["evidence", "1"], "text": "The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:17."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow.", "negative_left": "In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow.", "negative_right": "The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:43.", "right": "The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:17."}, "verifier_independent_model": false}, "family": "scale-diverse-095-003", "id": "scale-diverse-095-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "A cancellation has triggered an operational handoff for the traveler’s replacement journey from Lanton to Merrow.", "evidence": ["In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow.", "The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:17.", "The same confirmation set one transfer as the maximum.", "The traveler confirmed these were the only hard constraints and that no other preference should displace the rubric’s earliest-arrival ordering.", "The dispatcher confirms Birch and Alder are operating replacement routes from Lanton to Merrow, each with one transfer.", "The traveler’s flexible ticket with the cancellation endorsement is valid on both routes.", "Birch’s 14-minute connection exceeds the published 12-minute minimum.", "Alder has a 13-minute interchange and is scheduled to reach Merrow at 18:12.", "The operations comparison records Alder as arriving earlier than Birch."], "request": "Using the rubric, score Birch’s suitability for this traveler."}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the full scoring policy, while both contexts retain the request to score Birch for the same traveler and Lanton-to-Merrow journey. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Birch’s arrival from 18:17 to 18:43; Alder’s 18:12 arrival remains earlier in either case, so no duplicate measurement or unchanged assertion conflicts with the change. Neither context states a score, answer code, proposition identifier, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"A cancellation has triggered an operational handoff for the traveler’s replacement journey from Lanton to Merrow.\",\"evidence\":[\"In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow.\",\"The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:17.\",\"The same confirmation set one transfer as the maximum.\",\"The traveler confirmed these were the only hard constraints and that no other preference should displace the rubric’s earliest-arrival ordering.\",\"The dispatcher confirms Birch and Alder are operating replacement routes from Lanton to Merrow, each with one transfer.\",\"The traveler’s flexible ticket with the cancellation endorsement is valid on both routes.\",\"Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"Alder has a 13-minute interchange and is scheduled to reach Merrow at 18:12.\",\"The operations comparison records Alder as arriving earlier than Birch.\"],\"request\":\"Using the rubric, score Birch’s suitability for this traveler.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow."}, {"path": ["evidence", "1"], "text": "The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:17."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow.", "negative_left": "In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow.", "negative_right": "The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:43.", "right": "The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:17."}, "verifier_independent_model": false}, "family": "scale-diverse-095-003", "id": "scale-diverse-095-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "A cancellation has triggered an operational handoff for the traveler’s replacement journey from Lanton to Merrow.", "evidence": ["In the operational handoff for the replacement journey from Lanton to Merrow on 17 September 2026, the traveler confirmed the rebooking agent’s paraphrase specifying 18:30 as the latest permissible arrival time at Merrow.", "The finalized schedule for Birch’s replacement journey from Lanton to Merrow on 17 September 2026 records its arrival at Merrow at 18:43.", "The same confirmation set one transfer as the maximum.", "The traveler confirmed these were the only hard constraints and that no other preference should displace the rubric’s earliest-arrival ordering.", "The dispatcher confirms Birch and Alder are operating replacement routes from Lanton to Merrow, each with one transfer.", "The traveler’s flexible ticket with the cancellation endorsement is valid on both routes.", "Birch’s 14-minute connection exceeds the published 12-minute minimum.", "Alder has a 13-minute interchange and is scheduled to reach Merrow at 18:12.", "The operations comparison records Alder as arriving earlier than Birch."], "request": "Using the rubric, score Birch’s suitability for this traveler."}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all scoring rules and mandatory constraints, while both contexts retain the state-originating eight-minute station minimum without alteration. Option A, the Alderport-to-Bellhaven route, the September 17, 2026/today time binding, wheelchair assistance, no-bus requirement, and 16:00 deadline remain aligned. The two evidence spans are complete factual sentences. The counterfactual changes only H-47’s arrival value from 15:53 to 16:11 and creates no duplicate or contradictory arrival assertion. Neither context embeds a score, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"arrival_meets_deadline": "supported", "connection_at_least_minimum": "supported", "connection_at_most_minimum": "supported", "option_a_has_no_bus": "supported", "wheelchair_boarding_confirmed": "supported"}, "full_context_fact_states": {"base": {"arrival_meets_deadline": "supported", "connection_at_least_minimum": "supported", "connection_at_most_minimum": "supported", "option_a_has_no_bus": "supported", "wheelchair_boarding_confirmed": "supported"}, "counterfactual": {"arrival_meets_deadline": "refuted", "connection_at_least_minimum": "supported", "connection_at_most_minimum": "supported", "option_a_has_no_bus": "supported", "wheelchair_boarding_confirmed": "supported"}, "remove_left": {"arrival_meets_deadline": "unknown"}, "remove_right": {"arrival_meets_deadline": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"arrival_meets_deadline": "unknown"}, "negative_pair": {"arrival_meets_deadline": "refuted"}, "negative_sentence": {"arrival_meets_deadline": "unknown"}, "positive_pair": {"arrival_meets_deadline": "supported"}, "right": {"arrival_meets_deadline": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus is the factual relation between Option A’s arrival and the deadline. The base and counter assignments are realizable with only that arrival relation changing. Policy evidence preserves the state-originating deadline/no-bus requirements and the station-specific eight-minute minimum; the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The supported lower and upper bounds jointly entail an exactly eight-minute connection. Together with arrival by 16:00, no bus segment, and confirmed wheelchair boarding, all mandatory conditions are satisfied and the exact-minimum connection makes level 1 sufficient. The exact upper bound also excludes the longer-than-eight-minute condition for level 2.", "rule_index": 0, "sound": true}, {"reason": "Refuting arrival_meets_deadline entails that Option A arrives after the mandatory 16:00 deadline. Violation of either mandatory traveler constraint is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "connection_at_least_minimum", "statement": "Option A's timetable gap at Carden today from arrival on platform 2 to departure from platform 3 is at least eight minutes."}, {"id": "connection_at_most_minimum", "statement": "Option A's timetable gap at Carden today from arrival on platform 2 to departure from platform 3 is no more than eight minutes."}, {"id": "arrival_meets_deadline", "statement": "Option A's scheduled arrival at Bellhaven today is no later than the rail traveler's mandatory Bellhaven arrival deadline of 16:00 today."}, {"id": "option_a_has_no_bus", "statement": "No transport segment of Option A from Alderport to Bellhaven today is a bus segment."}, {"id": "wheelchair_boarding_confirmed", "statement": "Wheelchair boarding assistance is confirmed for the rail traveler on Option A from Alderport to Bellhaven today."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I use a wheelchair, cannot take a bus, and must arrive by 16:00.\"},{\"speaker\":\"operations handoff\",\"text\":\"Option A is the rebooked Alderport-to-Bellhaven itinerary for today. Every transport segment is by train. At Carden, its first train is timetabled to arrive on platform 2 at 14:22, and its onward train is timetabled to depart from platform 3 at 14:30.\"},{\"speaker\":\"operations handoff\",\"text\":\"For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time.\"},{\"speaker\":\"timetable controller\",\"text\":\"Timetable field H-47 contains 15:53 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00.\"},{\"speaker\":\"Option A booking record\",\"text\":\"Station policy sets an eight-minute minimum connection between those platforms; assisted boarding is confirmed.\"}]", "base_states": [{"atom_id": "connection_at_least_minimum", "state": "supported"}, {"atom_id": "connection_at_most_minimum", "state": "supported"}, {"atom_id": "arrival_meets_deadline", "state": "supported"}, {"atom_id": "option_a_has_no_bus", "state": "supported"}, {"atom_id": "wheelchair_boarding_confirmed", "state": "supported"}], "counter_states": [{"atom_id": "connection_at_least_minimum", "state": "supported"}, {"atom_id": "connection_at_most_minimum", "state": "supported"}, {"atom_id": "arrival_meets_deadline", "state": "refuted"}, {"atom_id": "option_a_has_no_bus", "state": "supported"}, {"atom_id": "wheelchair_boarding_confirmed", "state": "supported"}], "focus_atom": "arrival_meets_deadline", "focus_evidence": [{"path": ["2", "text"], "text": "For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time."}, {"path": ["3", "text"], "text": "Timetable field H-47 contains 15:53 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00."}], "policy_evidence": [{"path": ["0", "text"], "text": "I use a wheelchair, cannot take a bus, and must arrive by 16:00."}, {"path": ["1", "text"], "text": "Station policy sets an eight-minute minimum connection between those platforms; assisted boarding is confirmed."}], "rules": [{"justification": "The two connection bounds establish a gap of exactly eight minutes. Option A also meets both mandatory constraints and has confirmed wheelchair boarding, so it is acceptable but fragile.", "target": "1", "when": [{"atom_id": "connection_at_least_minimum", "state": "supported"}, {"atom_id": "connection_at_most_minimum", "state": "supported"}, {"atom_id": "arrival_meets_deadline", "state": "supported"}, {"atom_id": "option_a_has_no_bus", "state": "supported"}, {"atom_id": "wheelchair_boarding_confirmed", "state": "supported"}]}, {"justification": "An arrival later than the mandatory 16:00 deadline violates a mandatory traveler constraint, which is sufficient for an unsuitable rating.", "target": "0", "when": [{"atom_id": "arrival_meets_deadline", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time.", "negative_left": "For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time.", "negative_right": "Timetable field H-47 contains 16:11 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00.", "right": "Timetable field H-47 contains 15:53 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00."}, "verifier_independent_model": false}, "family": "scale-diverse-096-003", "id": "scale-diverse-096-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: violates either mandatory traveler constraint, lacks confirmed wheelchair boarding, or has a connection shorter than eight minutes.", "1 — Acceptable but fragile: satisfies all mandatory constraints and confirmed assistance, but either the connection exactly equals the eight-minute minimum or arrival is less than five minutes before 16:00.", "2 — Strongly suitable: satisfies all mandatory constraints and confirmed assistance, with a connection longer than eight minutes and arrival at least five minutes before 16:00."], "instructions": "Rate Option A’s suitability using only the explicit evidence. Apply these rules: a connection is feasible when its timetable gap is at least the stated station minimum, including equality; the 16:00 arrival deadline and no-bus constraint are mandatory. Select the single level whose definition fits best.", "type": "score"}}, "state": [{"speaker": "rail traveler", "text": "I use a wheelchair, cannot take a bus, and must arrive by 16:00."}, {"speaker": "operations handoff", "text": "Option A is the rebooked Alderport-to-Bellhaven itinerary for today. Every transport segment is by train. At Carden, its first train is timetabled to arrive on platform 2 at 14:22, and its onward train is timetabled to depart from platform 3 at 14:30."}, {"speaker": "operations handoff", "text": "For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time."}, {"speaker": "timetable controller", "text": "Timetable field H-47 contains 15:53 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00."}, {"speaker": "Option A booking record", "text": "Station policy sets an eight-minute minimum connection between those platforms; assisted boarding is confirmed."}]}, "method": "c2d", "provenance": {"source_id": "diverse-096", "source_is_synthetic": true, "source_sha256": "d0ff1d31d5fa023be5fe937290fddf2410923bf977198d111954e1fcce408f0d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all scoring rules and mandatory constraints, while both contexts retain the state-originating eight-minute station minimum without alteration. Option A, the Alderport-to-Bellhaven route, the September 17, 2026/today time binding, wheelchair assistance, no-bus requirement, and 16:00 deadline remain aligned. The two evidence spans are complete factual sentences. The counterfactual changes only H-47’s arrival value from 15:53 to 16:11 and creates no duplicate or contradictory arrival assertion. Neither context embeds a score, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"arrival_meets_deadline": "refuted", "connection_at_least_minimum": "supported", "connection_at_most_minimum": "supported", "option_a_has_no_bus": "supported", "wheelchair_boarding_confirmed": "supported"}, "full_context_fact_states": {"base": {"arrival_meets_deadline": "supported", "connection_at_least_minimum": "supported", "connection_at_most_minimum": "supported", "option_a_has_no_bus": "supported", "wheelchair_boarding_confirmed": "supported"}, "counterfactual": {"arrival_meets_deadline": "refuted", "connection_at_least_minimum": "supported", "connection_at_most_minimum": "supported", "option_a_has_no_bus": "supported", "wheelchair_boarding_confirmed": "supported"}, "remove_left": {"arrival_meets_deadline": "unknown"}, "remove_right": {"arrival_meets_deadline": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"arrival_meets_deadline": "unknown"}, "negative_pair": {"arrival_meets_deadline": "refuted"}, "negative_sentence": {"arrival_meets_deadline": "unknown"}, "positive_pair": {"arrival_meets_deadline": "supported"}, "right": {"arrival_meets_deadline": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus is the factual relation between Option A’s arrival and the deadline. The base and counter assignments are realizable with only that arrival relation changing. Policy evidence preserves the state-originating deadline/no-bus requirements and the station-specific eight-minute minimum; the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The supported lower and upper bounds jointly entail an exactly eight-minute connection. Together with arrival by 16:00, no bus segment, and confirmed wheelchair boarding, all mandatory conditions are satisfied and the exact-minimum connection makes level 1 sufficient. The exact upper bound also excludes the longer-than-eight-minute condition for level 2.", "rule_index": 0, "sound": true}, {"reason": "Refuting arrival_meets_deadline entails that Option A arrives after the mandatory 16:00 deadline. Violation of either mandatory traveler constraint is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "connection_at_least_minimum", "statement": "Option A's timetable gap at Carden today from arrival on platform 2 to departure from platform 3 is at least eight minutes."}, {"id": "connection_at_most_minimum", "statement": "Option A's timetable gap at Carden today from arrival on platform 2 to departure from platform 3 is no more than eight minutes."}, {"id": "arrival_meets_deadline", "statement": "Option A's scheduled arrival at Bellhaven today is no later than the rail traveler's mandatory Bellhaven arrival deadline of 16:00 today."}, {"id": "option_a_has_no_bus", "statement": "No transport segment of Option A from Alderport to Bellhaven today is a bus segment."}, {"id": "wheelchair_boarding_confirmed", "statement": "Wheelchair boarding assistance is confirmed for the rail traveler on Option A from Alderport to Bellhaven today."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I use a wheelchair, cannot take a bus, and must arrive by 16:00.\"},{\"speaker\":\"operations handoff\",\"text\":\"Option A is the rebooked Alderport-to-Bellhaven itinerary for today. Every transport segment is by train. At Carden, its first train is timetabled to arrive on platform 2 at 14:22, and its onward train is timetabled to depart from platform 3 at 14:30.\"},{\"speaker\":\"operations handoff\",\"text\":\"For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time.\"},{\"speaker\":\"timetable controller\",\"text\":\"Timetable field H-47 contains 15:53 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00.\"},{\"speaker\":\"Option A booking record\",\"text\":\"Station policy sets an eight-minute minimum connection between those platforms; assisted boarding is confirmed.\"}]", "base_states": [{"atom_id": "connection_at_least_minimum", "state": "supported"}, {"atom_id": "connection_at_most_minimum", "state": "supported"}, {"atom_id": "arrival_meets_deadline", "state": "supported"}, {"atom_id": "option_a_has_no_bus", "state": "supported"}, {"atom_id": "wheelchair_boarding_confirmed", "state": "supported"}], "counter_states": [{"atom_id": "connection_at_least_minimum", "state": "supported"}, {"atom_id": "connection_at_most_minimum", "state": "supported"}, {"atom_id": "arrival_meets_deadline", "state": "refuted"}, {"atom_id": "option_a_has_no_bus", "state": "supported"}, {"atom_id": "wheelchair_boarding_confirmed", "state": "supported"}], "focus_atom": "arrival_meets_deadline", "focus_evidence": [{"path": ["2", "text"], "text": "For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time."}, {"path": ["3", "text"], "text": "Timetable field H-47 contains 15:53 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00."}], "policy_evidence": [{"path": ["0", "text"], "text": "I use a wheelchair, cannot take a bus, and must arrive by 16:00."}, {"path": ["1", "text"], "text": "Station policy sets an eight-minute minimum connection between those platforms; assisted boarding is confirmed."}], "rules": [{"justification": "The two connection bounds establish a gap of exactly eight minutes. Option A also meets both mandatory constraints and has confirmed wheelchair boarding, so it is acceptable but fragile.", "target": "1", "when": [{"atom_id": "connection_at_least_minimum", "state": "supported"}, {"atom_id": "connection_at_most_minimum", "state": "supported"}, {"atom_id": "arrival_meets_deadline", "state": "supported"}, {"atom_id": "option_a_has_no_bus", "state": "supported"}, {"atom_id": "wheelchair_boarding_confirmed", "state": "supported"}]}, {"justification": "An arrival later than the mandatory 16:00 deadline violates a mandatory traveler constraint, which is sufficient for an unsuitable rating.", "target": "0", "when": [{"atom_id": "arrival_meets_deadline", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time.", "negative_left": "For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time.", "negative_right": "Timetable field H-47 contains 16:11 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00.", "right": "Timetable field H-47 contains 15:53 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00."}, "verifier_independent_model": false}, "family": "scale-diverse-096-003", "id": "scale-diverse-096-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: violates either mandatory traveler constraint, lacks confirmed wheelchair boarding, or has a connection shorter than eight minutes.", "1 — Acceptable but fragile: satisfies all mandatory constraints and confirmed assistance, but either the connection exactly equals the eight-minute minimum or arrival is less than five minutes before 16:00.", "2 — Strongly suitable: satisfies all mandatory constraints and confirmed assistance, with a connection longer than eight minutes and arrival at least five minutes before 16:00."], "instructions": "Rate Option A’s suitability using only the explicit evidence. Apply these rules: a connection is feasible when its timetable gap is at least the stated station minimum, including equality; the 16:00 arrival deadline and no-bus constraint are mandatory. Select the single level whose definition fits best.", "type": "score"}}, "state": [{"speaker": "rail traveler", "text": "I use a wheelchair, cannot take a bus, and must arrive by 16:00."}, {"speaker": "operations handoff", "text": "Option A is the rebooked Alderport-to-Bellhaven itinerary for today. Every transport segment is by train. At Carden, its first train is timetabled to arrive on platform 2 at 14:22, and its onward train is timetabled to depart from platform 3 at 14:30."}, {"speaker": "operations handoff", "text": "For Option A on September 17, 2026 (today), the operational handoff designates timetable field H-47 as its scheduled Bellhaven arrival time."}, {"speaker": "timetable controller", "text": "Timetable field H-47 contains 16:11 on September 17, 2026, and the rail traveler's mandatory Bellhaven arrival deadline that day is 16:00."}, {"speaker": "Option A booking record", "text": "Station policy sets an eight-minute minimum connection between those platforms; assisted boarding is confirmed."}]}, "method": "c2d", "provenance": {"source_id": "diverse-096", "source_is_synthetic": true, "source_sha256": "d0ff1d31d5fa023be5fe937290fddf2410923bf977198d111954e1fcce408f0d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Monday Mara itinerary, route, fixed times, required rest, ferry vehicle restriction, replacement option, and lodging window without adding policy exceptions or decision defaults. The only counterfactual change is the compact-car catalog code from HC-417 to HC-592; this does not contradict any unchanged measurement or assertion. The two evidence spans are complete factual sentences, and neither context contains a score, gold answer, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"On Monday, Mara’s Harbor City file limits the reservation’s canonical vehicle class to either SUV or compact car. At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417. At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-417. The coordinator reported that the ferry does not accept SUVs; it accepts compact cars, and that a compact car is available if Mara changes her reservation before pickup. Every listed driving time is confirmed for both the reserved vehicle and the available replacement compact car. Mara plans to collect the vehicle at 08:00, drive two hours, and reach Pine Gate at 10:00. She will take the required 30-minute rest from 10:00–10:30, then attend the fixed 10:30–11:30 garden visit. A 90-minute drive brings her to Seaside Ferry for 13:00 boarding. The ferry ride runs 13:00–13:45, followed by a one-hour drive to Island Lodge, arriving at 14:45. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417."}, {"path": [], "text": "At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-417."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417.", "negative_left": "At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417.", "negative_right": "At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-592.", "right": "At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-417."}, "verifier_independent_model": false}, "family": "scale-diverse-107-001", "id": "scale-diverse-107-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "On Monday, Mara’s Harbor City file limits the reservation’s canonical vehicle class to either SUV or compact car. At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417. At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-417. The coordinator reported that the ferry does not accept SUVs; it accepts compact cars, and that a compact car is available if Mara changes her reservation before pickup. Every listed driving time is confirmed for both the reserved vehicle and the available replacement compact car. Mara plans to collect the vehicle at 08:00, drive two hours, and reach Pine Gate at 10:00. She will take the required 30-minute rest from 10:00–10:30, then attend the fixed 10:30–11:30 garden visit. A 90-minute drive brings her to Seaside Ferry for 13:00 boarding. The ferry ride runs 13:00–13:45, followed by a one-hour drive to Island Lodge, arriving at 14:45. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Monday Mara itinerary, route, fixed times, required rest, ferry vehicle restriction, replacement option, and lodging window without adding policy exceptions or decision defaults. The only counterfactual change is the compact-car catalog code from HC-417 to HC-592; this does not contradict any unchanged measurement or assertion. The two evidence spans are complete factual sentences, and neither context contains a score, gold answer, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"On Monday, Mara’s Harbor City file limits the reservation’s canonical vehicle class to either SUV or compact car. At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417. At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-417. The coordinator reported that the ferry does not accept SUVs; it accepts compact cars, and that a compact car is available if Mara changes her reservation before pickup. Every listed driving time is confirmed for both the reserved vehicle and the available replacement compact car. Mara plans to collect the vehicle at 08:00, drive two hours, and reach Pine Gate at 10:00. She will take the required 30-minute rest from 10:00–10:30, then attend the fixed 10:30–11:30 garden visit. A 90-minute drive brings her to Seaside Ferry for 13:00 boarding. The ferry ride runs 13:00–13:45, followed by a one-hour drive to Island Lodge, arriving at 14:45. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417."}, {"path": [], "text": "At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-417."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417.", "negative_left": "At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417.", "negative_right": "At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-592.", "right": "At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-417."}, "verifier_independent_model": false}, "family": "scale-diverse-107-001", "id": "scale-diverse-107-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "On Monday, Mara’s Harbor City file limits the reservation’s canonical vehicle class to either SUV or compact car. At 07:41, the canonical vehicle-class field on Mara's Monday Harbor City rental reservation contained the code HC-417. At 07:42, the rental catalog's canonical entry for a compact car contained the code HC-592. The coordinator reported that the ferry does not accept SUVs; it accepts compact cars, and that a compact car is available if Mara changes her reservation before pickup. Every listed driving time is confirmed for both the reserved vehicle and the available replacement compact car. Mara plans to collect the vehicle at 08:00, drive two hours, and reach Pine Gate at 10:00. She will take the required 30-minute rest from 10:00–10:30, then attend the fixed 10:30–11:30 garden visit. A 90-minute drive brings her to Seaside Ferry for 13:00 boarding. The ferry ride runs 13:00–13:45, followed by a one-hour drive to Island Lodge, arriving at 14:45. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original feasibility criteria and scope without adding exceptions, priorities, or missing-evidence rules. Mara, her Monday reservation, the vehicle-class identification path, schedule, ferry restriction, rest, and lodging bindings remain fixed; only the catalog’s compact-class code changes. The two evidence spans are complete factual sentences. The counterfactual is coherent because HC-C17 and HC-K42 are not asserted to identify the same class, so there is no contradictory duplicate measurement or assertion. Neither context includes a score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"The audited reservation file says Mara’s canonical vehicle class is one of SUV or compact car. At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17. At 07:12 on Monday, the audited rental catalog listed HC-C17 as its canonical vehicle-class code for a compact car. A replacement compact is available if Mara changes the reservation before pickup. The operating notes state that the ferry does not accept SUVs; it accepts compact cars. Mara’s Monday schedule has Harbor City pickup at 08:00, Pine Gate arrival at 10:00, and rest from 10:00–10:30; she must take the required 30-minute rest and attend the fixed 10:30–11:30 garden visit. She reaches Seaside Ferry for 13:00 boarding, rides from 13:00–13:45, and reaches Island Lodge at 14:45. Every listed driving time—two hours, 1.5 hours, and one hour—is confirmed for both the reserved vehicle and replacement compact. Check-in is 15:00–20:00, and Mara’s Monday room is confirmed.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17."}, {"path": [], "text": "At 07:12 on Monday, the audited rental catalog listed HC-C17 as its canonical vehicle-class code for a compact car."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17.", "negative_left": "At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17.", "negative_right": "At 07:12 on Monday, the audited rental catalog listed HC-K42 as its canonical vehicle-class code for a compact car.", "right": "At 07:12 on Monday, the audited rental catalog listed HC-C17 as its canonical vehicle-class code for a compact car."}, "verifier_independent_model": false}, "family": "scale-diverse-107-002", "id": "scale-diverse-107-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "The audited reservation file says Mara’s canonical vehicle class is one of SUV or compact car. At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17. At 07:12 on Monday, the audited rental catalog listed HC-C17 as its canonical vehicle-class code for a compact car. A replacement compact is available if Mara changes the reservation before pickup. The operating notes state that the ferry does not accept SUVs; it accepts compact cars. Mara’s Monday schedule has Harbor City pickup at 08:00, Pine Gate arrival at 10:00, and rest from 10:00–10:30; she must take the required 30-minute rest and attend the fixed 10:30–11:30 garden visit. She reaches Seaside Ferry for 13:00 boarding, rides from 13:00–13:45, and reaches Island Lodge at 14:45. Every listed driving time—two hours, 1.5 hours, and one hour—is confirmed for both the reserved vehicle and replacement compact. Check-in is 15:00–20:00, and Mara’s Monday room is confirmed."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original feasibility criteria and scope without adding exceptions, priorities, or missing-evidence rules. Mara, her Monday reservation, the vehicle-class identification path, schedule, ferry restriction, rest, and lodging bindings remain fixed; only the catalog’s compact-class code changes. The two evidence spans are complete factual sentences. The counterfactual is coherent because HC-C17 and HC-K42 are not asserted to identify the same class, so there is no contradictory duplicate measurement or assertion. Neither context includes a score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"The audited reservation file says Mara’s canonical vehicle class is one of SUV or compact car. At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17. At 07:12 on Monday, the audited rental catalog listed HC-C17 as its canonical vehicle-class code for a compact car. A replacement compact is available if Mara changes the reservation before pickup. The operating notes state that the ferry does not accept SUVs; it accepts compact cars. Mara’s Monday schedule has Harbor City pickup at 08:00, Pine Gate arrival at 10:00, and rest from 10:00–10:30; she must take the required 30-minute rest and attend the fixed 10:30–11:30 garden visit. She reaches Seaside Ferry for 13:00 boarding, rides from 13:00–13:45, and reaches Island Lodge at 14:45. Every listed driving time—two hours, 1.5 hours, and one hour—is confirmed for both the reserved vehicle and replacement compact. Check-in is 15:00–20:00, and Mara’s Monday room is confirmed.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17."}, {"path": [], "text": "At 07:12 on Monday, the audited rental catalog listed HC-C17 as its canonical vehicle-class code for a compact car."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17.", "negative_left": "At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17.", "negative_right": "At 07:12 on Monday, the audited rental catalog listed HC-K42 as its canonical vehicle-class code for a compact car.", "right": "At 07:12 on Monday, the audited rental catalog listed HC-C17 as its canonical vehicle-class code for a compact car."}, "verifier_independent_model": false}, "family": "scale-diverse-107-002", "id": "scale-diverse-107-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "The audited reservation file says Mara’s canonical vehicle class is one of SUV or compact car. At 07:12 on Monday, the audited Harbor City rental record for Mara's Monday reservation listed its canonical vehicle-class code as HC-C17. At 07:12 on Monday, the audited rental catalog listed HC-K42 as its canonical vehicle-class code for a compact car. A replacement compact is available if Mara changes the reservation before pickup. The operating notes state that the ferry does not accept SUVs; it accepts compact cars. Mara’s Monday schedule has Harbor City pickup at 08:00, Pine Gate arrival at 10:00, and rest from 10:00–10:30; she must take the required 30-minute rest and attend the fixed 10:30–11:30 garden visit. She reaches Seaside Ferry for 13:00 boarding, rides from 13:00–13:45, and reaches Island Lodge at 14:45. Every listed driving time—two hours, 1.5 hours, and one hour—is confirmed for both the reserved vehicle and replacement compact. Check-in is 15:00–20:00, and Mara’s Monday room is confirmed."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rest, visit, boarding, check-in, ferry vehicle, replacement availability, and confirmed-timing constraints, while preserving Mara’s Monday Harbor City itinerary and times. The counterfactual changes only the reservation’s canonical code from HC-482 to HC-917; because the catalog sentence maps only HC-482 to compact, the altered context leaves HC-917’s class unresolved but does not create a contradiction or invent a missing-evidence default. The two evidence spans are complete factual sentences, and neither context contains a score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Operational handoff: Mara’s Monday Harbor City reservation is classified as either an SUV or a compact car. The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-482. The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car. A replacement compact remains available if she changes the reservation before pickup. All listed driving times are confirmed for both the reserved vehicle and the replacement. Pickup is at 08:00. Mara reaches Pine Gate at 10:00, rests 10:00–10:30, attends the garden 10:30–11:30, reaches Seaside Ferry at 13:00, rides 13:00–13:45, and reaches Island Lodge at 14:45. Her Monday room is confirmed. Governing notes state: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-482."}, {"path": [], "text": "The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-482.", "negative_left": "The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-917.", "negative_right": "The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car.", "right": "The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car."}, "verifier_independent_model": false}, "family": "scale-diverse-107-003", "id": "scale-diverse-107-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Operational handoff: Mara’s Monday Harbor City reservation is classified as either an SUV or a compact car. The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-482. The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car. A replacement compact remains available if she changes the reservation before pickup. All listed driving times are confirmed for both the reserved vehicle and the replacement. Pickup is at 08:00. Mara reaches Pine Gate at 10:00, rests 10:00–10:30, attends the garden 10:30–11:30, reaches Seaside Ferry at 13:00, rides 13:00–13:45, and reaches Island Lodge at 14:45. Her Monday room is confirmed. Governing notes state: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rest, visit, boarding, check-in, ferry vehicle, replacement availability, and confirmed-timing constraints, while preserving Mara’s Monday Harbor City itinerary and times. The counterfactual changes only the reservation’s canonical code from HC-482 to HC-917; because the catalog sentence maps only HC-482 to compact, the altered context leaves HC-917’s class unresolved but does not create a contradiction or invent a missing-evidence default. The two evidence spans are complete factual sentences, and neither context contains a score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Operational handoff: Mara’s Monday Harbor City reservation is classified as either an SUV or a compact car. The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-482. The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car. A replacement compact remains available if she changes the reservation before pickup. All listed driving times are confirmed for both the reserved vehicle and the replacement. Pickup is at 08:00. Mara reaches Pine Gate at 10:00, rests 10:00–10:30, attends the garden 10:30–11:30, reaches Seaside Ferry at 13:00, rides 13:00–13:45, and reaches Island Lodge at 14:45. Her Monday room is confirmed. Governing notes state: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-482."}, {"path": [], "text": "The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-482.", "negative_left": "The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-917.", "negative_right": "The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car.", "right": "The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car."}, "verifier_independent_model": false}, "family": "scale-diverse-107-003", "id": "scale-diverse-107-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Operational handoff: Mara’s Monday Harbor City reservation is classified as either an SUV or a compact car. The 07:12 operational handoff record for Mara's Monday Harbor City rental reservation lists its canonical vehicle-class code as HC-917. The rental catalog in effect for Mara's Monday Harbor City pickup lists HC-482 as the canonical vehicle-class code for a compact car. A replacement compact remains available if she changes the reservation before pickup. All listed driving times are confirmed for both the reserved vehicle and the replacement. Pickup is at 08:00. Mara reaches Pine Gate at 10:00, rests 10:00–10:30, attends the garden 10:30–11:30, reaches Seaside Ferry at 13:00, rides 13:00–13:45, and reaches Island Lodge at 14:45. Her Monday room is confirmed. Governing notes state: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the governing itinerary constraints, ferry vehicle rule, replacement availability, confirmed travel times, required rest, fixed visit, and lodging facts, while the unchanged questions object preserves the scoring policy. The same traveler, reservation, route, locations, and Monday time bindings are maintained. The two evidence spans are complete factual sentences. The counterfactual changes only the reservation’s recorded code from HC-17 to HC-42; because it does not assert a conflicting class mapping for HC-42, it remains coherent, though the reserved class becomes unresolved from the supplied catalog facts. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Monday’s rental audit says Mara’s Harbor City reservation has a canonical vehicle class in the catalog set comprising only SUV and compact car. At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-17. At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car. Before pickup, the coordinator confirmed that a compact car is available if Mara changes her reservation and that every listed driving time is confirmed for both the reserved vehicle and the replacement compact. Mara plans to collect the vehicle Monday at 08:00 and reach Pine Gate at 10:00, where she will take the required 30-minute rest from 10:00 to 10:30 and attend the fixed 10:30–11:30 garden visit. She will reach Seaside Ferry for 13:00 boarding; the ferry does not accept SUVs; it accepts compact cars. The ride runs 13:00–13:45, followed by the confirmed one-hour drive to Island Lodge, arriving at 14:45. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-17."}, {"path": [], "text": "At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-17.", "negative_left": "At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-42.", "negative_right": "At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car.", "right": "At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car."}, "verifier_independent_model": false}, "family": "scale-diverse-107-005", "id": "scale-diverse-107-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Monday’s rental audit says Mara’s Harbor City reservation has a canonical vehicle class in the catalog set comprising only SUV and compact car. At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-17. At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car. Before pickup, the coordinator confirmed that a compact car is available if Mara changes her reservation and that every listed driving time is confirmed for both the reserved vehicle and the replacement compact. Mara plans to collect the vehicle Monday at 08:00 and reach Pine Gate at 10:00, where she will take the required 30-minute rest from 10:00 to 10:30 and attend the fixed 10:30–11:30 garden visit. She will reach Seaside Ferry for 13:00 boarding; the ferry does not accept SUVs; it accepts compact cars. The ride runs 13:00–13:45, followed by the confirmed one-hour drive to Island Lodge, arriving at 14:45. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the governing itinerary constraints, ferry vehicle rule, replacement availability, confirmed travel times, required rest, fixed visit, and lodging facts, while the unchanged questions object preserves the scoring policy. The same traveler, reservation, route, locations, and Monday time bindings are maintained. The two evidence spans are complete factual sentences. The counterfactual changes only the reservation’s recorded code from HC-17 to HC-42; because it does not assert a conflicting class mapping for HC-42, it remains coherent, though the reserved class becomes unresolved from the supplied catalog facts. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Monday’s rental audit says Mara’s Harbor City reservation has a canonical vehicle class in the catalog set comprising only SUV and compact car. At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-17. At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car. Before pickup, the coordinator confirmed that a compact car is available if Mara changes her reservation and that every listed driving time is confirmed for both the reserved vehicle and the replacement compact. Mara plans to collect the vehicle Monday at 08:00 and reach Pine Gate at 10:00, where she will take the required 30-minute rest from 10:00 to 10:30 and attend the fixed 10:30–11:30 garden visit. She will reach Seaside Ferry for 13:00 boarding; the ferry does not accept SUVs; it accepts compact cars. The ride runs 13:00–13:45, followed by the confirmed one-hour drive to Island Lodge, arriving at 14:45. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-17."}, {"path": [], "text": "At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-17.", "negative_left": "At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-42.", "negative_right": "At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car.", "right": "At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car."}, "verifier_independent_model": false}, "family": "scale-diverse-107-005", "id": "scale-diverse-107-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Monday’s rental audit says Mara’s Harbor City reservation has a canonical vehicle class in the catalog set comprising only SUV and compact car. At 07:15 on Monday, Mara's Harbor City rental reservation recorded the canonical vehicle-class code HC-42. At 07:20 on Monday, the Harbor City rental catalog recorded HC-17 as the canonical vehicle-class code for a compact car. Before pickup, the coordinator confirmed that a compact car is available if Mara changes her reservation and that every listed driving time is confirmed for both the reserved vehicle and the replacement compact. Mara plans to collect the vehicle Monday at 08:00 and reach Pine Gate at 10:00, where she will take the required 30-minute rest from 10:00 to 10:30 and attend the fixed 10:30–11:30 garden visit. She will reach Seaside Ferry for 13:00 boarding; the ferry does not accept SUVs; it accepts compact cars. The ride runs 13:00–13:45, followed by the confirmed one-hour drive to Island Lodge, arriving at 14:45. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same scoring policy, Friday self-drive scope, conditional plan-selection rule, and relevant entity/path/time bindings. The two focus spans are complete factual sentences. Changing the nonstop museum-to-winery duration from 64 to 71 minutes coherently changes arrival from 12:10 to 12:17 while leaving the 11:06 departure and 12:15 fixed start unchanged; no duplicate measurement contradicts that change. The lodging assertion can still coexist with a missed winery booking, and neither context embeds a score, answer code, rule table, proposition identifier, output instruction, or explicit gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"Friday’s rental record confirms an automatic compact and shows that no AWD vehicle is confirmed.\",\"The compact is permitted on every road in the Valley plan, and the museum log records entry at 10:15, within its 10:00–11:00 window.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\",\"On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time.\",\"The Friday Valley plan schedules 64 minutes from the museum departure to the winery arrival.\",\"Keeping the Friday Valley route, reservation, and vehicle unchanged, that nonstop leg yields the earliest obtainable winery arrival.\",\"The lodging log records arrival at 15:05, ten minutes before check-in closes at 15:15; every continuous driving segment is at most two hours.\",\"A room for Friday night and a parking space for the Valley plan are both confirmed.\"],\"request\":\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "4"], "text": "On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time."}, {"path": ["evidence", "5"], "text": "The Friday Valley plan schedules 64 minutes from the museum departure to the winery arrival."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time.", "negative_left": "On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time.", "negative_right": "The Friday Valley plan schedules 71 minutes from the museum departure to the winery arrival.", "right": "The Friday Valley plan schedules 64 minutes from the museum departure to the winery arrival."}, "verifier_independent_model": false}, "family": "scale-diverse-108-001", "id": "scale-diverse-108-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "Friday’s rental record confirms an automatic compact and shows that no AWD vehicle is confirmed.", "The compact is permitted on every road in the Valley plan, and the museum log records entry at 10:15, within its 10:00–11:00 window.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock.", "On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time.", "The Friday Valley plan schedules 64 minutes from the museum departure to the winery arrival.", "Keeping the Friday Valley route, reservation, and vehicle unchanged, that nonstop leg yields the earliest obtainable winery arrival.", "The lodging log records arrival at 15:05, ten minutes before check-in closes at 15:15; every continuous driving segment is at most two hours.", "A room for Friday night and a parking space for the Valley plan are both confirmed."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same scoring policy, Friday self-drive scope, conditional plan-selection rule, and relevant entity/path/time bindings. The two focus spans are complete factual sentences. Changing the nonstop museum-to-winery duration from 64 to 71 minutes coherently changes arrival from 12:10 to 12:17 while leaving the 11:06 departure and 12:15 fixed start unchanged; no duplicate measurement contradicts that change. The lodging assertion can still coexist with a missed winery booking, and neither context embeds a score, answer code, rule table, proposition identifier, output instruction, or explicit gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"Friday’s rental record confirms an automatic compact and shows that no AWD vehicle is confirmed.\",\"The compact is permitted on every road in the Valley plan, and the museum log records entry at 10:15, within its 10:00–11:00 window.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\",\"On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time.\",\"The Friday Valley plan schedules 64 minutes from the museum departure to the winery arrival.\",\"Keeping the Friday Valley route, reservation, and vehicle unchanged, that nonstop leg yields the earliest obtainable winery arrival.\",\"The lodging log records arrival at 15:05, ten minutes before check-in closes at 15:15; every continuous driving segment is at most two hours.\",\"A room for Friday night and a parking space for the Valley plan are both confirmed.\"],\"request\":\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "4"], "text": "On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time."}, {"path": ["evidence", "5"], "text": "The Friday Valley plan schedules 64 minutes from the museum departure to the winery arrival."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time.", "negative_left": "On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time.", "negative_right": "The Friday Valley plan schedules 71 minutes from the museum departure to the winery arrival.", "right": "The Friday Valley plan schedules 64 minutes from the museum departure to the winery arrival."}, "verifier_independent_model": false}, "family": "scale-diverse-108-001", "id": "scale-diverse-108-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "Friday’s rental record confirms an automatic compact and shows that no AWD vehicle is confirmed.", "The compact is permitted on every road in the Valley plan, and the museum log records entry at 10:15, within its 10:00–11:00 window.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock.", "On Friday, the Valley plan has the traveler leave the museum at 11:06 and proceed directly, without stops, to the winery, where the booked visit has a fixed 12:15 start time.", "The Friday Valley plan schedules 71 minutes from the museum departure to the winery arrival.", "Keeping the Friday Valley route, reservation, and vehicle unchanged, that nonstop leg yields the earliest obtainable winery arrival.", "The lodging log records arrival at 15:05, ten minutes before check-in closes at 15:15; every continuous driving segment is at most two hours.", "A room for Friday night and a parking space for the Valley plan are both confirmed."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler’s selection condition, Friday/Valley itinerary bindings, request scope, vehicle and lodging requirements, timing-window concepts, and the unchanged rest policy. The two focus spans are complete factual sentences. Changing the museum-to-winery elapsed time from 46 to 52 minutes coherently changes the implied arrival from 12:12 to 12:18 relative to the fixed 12:15 start; this creates no contradictory duplicate measurement within either context. Neither context states a score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"The rental record confirms no AWD vehicle and confirms one automatic compact for Friday. The compact is permitted on every road in the Valley plan.\",\"Valley’s museum entry is logged at 10:41, within its 10:00–11:00 entry window.\",\"The route audit records every continuous driving segment as no longer than two hours.\",\"On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday.\",\"The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 46 minutes.\",\"The logged winery arrival is the earliest obtainable without changing the Valley route, reservation, or vehicle.\",\"The lodging record shows arrival ten minutes before the 15:15 check-in closing time. A room for Friday night and a parking space are both confirmed.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\"],\"request\":\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "4"], "text": "On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday."}, {"path": ["evidence", "5"], "text": "The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 46 minutes."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday.", "negative_left": "On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday.", "negative_right": "The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 52 minutes.", "right": "The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 46 minutes."}, "verifier_independent_model": false}, "family": "scale-diverse-108-002", "id": "scale-diverse-108-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "The rental record confirms no AWD vehicle and confirms one automatic compact for Friday. The compact is permitted on every road in the Valley plan.", "Valley’s museum entry is logged at 10:41, within its 10:00–11:00 entry window.", "The route audit records every continuous driving segment as no longer than two hours.", "On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday.", "The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 46 minutes.", "The logged winery arrival is the earliest obtainable without changing the Valley route, reservation, or vehicle.", "The lodging record shows arrival ten minutes before the 15:15 check-in closing time. A room for Friday night and a parking space are both confirmed.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler’s selection condition, Friday/Valley itinerary bindings, request scope, vehicle and lodging requirements, timing-window concepts, and the unchanged rest policy. The two focus spans are complete factual sentences. Changing the museum-to-winery elapsed time from 46 to 52 minutes coherently changes the implied arrival from 12:12 to 12:18 relative to the fixed 12:15 start; this creates no contradictory duplicate measurement within either context. Neither context states a score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"The rental record confirms no AWD vehicle and confirms one automatic compact for Friday. The compact is permitted on every road in the Valley plan.\",\"Valley’s museum entry is logged at 10:41, within its 10:00–11:00 entry window.\",\"The route audit records every continuous driving segment as no longer than two hours.\",\"On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday.\",\"The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 46 minutes.\",\"The logged winery arrival is the earliest obtainable without changing the Valley route, reservation, or vehicle.\",\"The lodging record shows arrival ten minutes before the 15:15 check-in closing time. A room for Friday night and a parking space are both confirmed.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\"],\"request\":\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "4"], "text": "On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday."}, {"path": ["evidence", "5"], "text": "The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 46 minutes."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday.", "negative_left": "On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday.", "negative_right": "The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 52 minutes.", "right": "The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 46 minutes."}, "verifier_independent_model": false}, "family": "scale-diverse-108-002", "id": "scale-diverse-108-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "The rental record confirms no AWD vehicle and confirms one automatic compact for Friday. The compact is permitted on every road in the Valley plan.", "Valley’s museum entry is logged at 10:41, within its 10:00–11:00 entry window.", "The route audit records every continuous driving segment as no longer than two hours.", "On Friday, the Valley plan schedules departure from the museum at 11:26, and the booked winery visit has a fixed start time of 12:15 that Friday.", "The recorded elapsed time from the Friday Valley plan’s museum departure to its winery arrival is 52 minutes.", "The logged winery arrival is the earliest obtainable without changing the Valley route, reservation, or vehicle.", "The lodging record shows arrival ten minutes before the 15:15 check-in closing time. A room for Friday night and a parking space are both confirmed.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scoring criteria, instructions, Friday itinerary scope, conditional plan-selection rule, and relevant entity/path/time bindings. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the winery arrival from seven minutes before marker W to seven minutes after it; because W remains the 12:15 Friday fixed start, this change is coherent with the unchanged dispatch statement and creates no duplicate or contradictory measurement. Neither context embeds a score, gold answer, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.\",\"1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.\",\"2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.\",\"3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.\",\"4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes.\"],\"instructions\":\"First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.\",\"type\":\"score\"}},\"state\":{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"The fleet handoff confirms no AWD vehicle; the confirmed rental is an automatic compact permitted on every road in the Valley plan.\",\"The Valley museum arrival is 10:05 Friday, within its 10:00–11:00 entry window.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\",\"The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes before reference marker W.\",\"Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start.\",\"Dispatch confirms that the logged winery arrival is the earliest obtainable without changing the Valley route, reservation, or vehicle.\",\"Every continuous driving segment in the Valley plan is two hours or less.\",\"The plan reaches the lodging at 15:00 Friday, and check-in closes at 15:15.\",\"The required Friday-night room and a parking space for the Valley plan are both confirmed.\"],\"request\":\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["state", "evidence", "4"], "text": "The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes before reference marker W."}, {"path": ["state", "evidence", "5"], "text": "Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes before reference marker W.", "negative_left": "The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes after reference marker W.", "negative_right": "Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start.", "right": "Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start."}, "verifier_independent_model": false}, "family": "scale-diverse-108-003", "id": "scale-diverse-108-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "The fleet handoff confirms no AWD vehicle; the confirmed rental is an automatic compact permitted on every road in the Valley plan.", "The Valley museum arrival is 10:05 Friday, within its 10:00–11:00 entry window.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock.", "The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes before reference marker W.", "Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start.", "Dispatch confirms that the logged winery arrival is the earliest obtainable without changing the Valley route, reservation, or vehicle.", "Every continuous driving segment in the Valley plan is two hours or less.", "The plan reaches the lodging at 15:00 Friday, and check-in closes at 15:15.", "The required Friday-night room and a parking space for the Valley plan are both confirmed."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scoring criteria, instructions, Friday itinerary scope, conditional plan-selection rule, and relevant entity/path/time bindings. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the winery arrival from seven minutes before marker W to seven minutes after it; because W remains the 12:15 Friday fixed start, this change is coherent with the unchanged dispatch statement and creates no duplicate or contradictory measurement. Neither context embeds a score, gold answer, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.\",\"1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.\",\"2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.\",\"3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.\",\"4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes.\"],\"instructions\":\"First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.\",\"type\":\"score\"}},\"state\":{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"The fleet handoff confirms no AWD vehicle; the confirmed rental is an automatic compact permitted on every road in the Valley plan.\",\"The Valley museum arrival is 10:05 Friday, within its 10:00–11:00 entry window.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\",\"The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes before reference marker W.\",\"Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start.\",\"Dispatch confirms that the logged winery arrival is the earliest obtainable without changing the Valley route, reservation, or vehicle.\",\"Every continuous driving segment in the Valley plan is two hours or less.\",\"The plan reaches the lodging at 15:00 Friday, and check-in closes at 15:15.\",\"The required Friday-night room and a parking space for the Valley plan are both confirmed.\"],\"request\":\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["state", "evidence", "4"], "text": "The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes before reference marker W."}, {"path": ["state", "evidence", "5"], "text": "Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes before reference marker W.", "negative_left": "The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes after reference marker W.", "negative_right": "Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start.", "right": "Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start."}, "verifier_independent_model": false}, "family": "scale-diverse-108-003", "id": "scale-diverse-108-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "The fleet handoff confirms no AWD vehicle; the confirmed rental is an automatic compact permitted on every road in the Valley plan.", "The Valley museum arrival is 10:05 Friday, within its 10:00–11:00 entry window.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock.", "The Friday Valley plan’s operational handoff log records its winery arrival as seven minutes after reference marker W.", "Reference marker W denotes 12:15 on Friday, the same instant as the booked winery visit’s fixed start.", "Dispatch confirms that the logged winery arrival is the earliest obtainable without changing the Valley route, reservation, or vehicle.", "Every continuous driving segment in the Valley plan is two hours or less.", "The plan reaches the lodging at 15:00 Friday, and check-in closes at 15:15.", "The required Friday-night room and a parking space for the Valley plan are both confirmed."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the unchanged scoring question and retain the traveler’s conditional selection rule and rest policy from the original state. The Friday, Valley-plan, vehicle, route, reservation, and scoring bindings remain intact. The two focus spans are complete factual sentences. Changing the direct drive from 47 to 52 minutes coherently changes the winery arrival from 12:13 to 12:18 relative to the fixed 12:15 start, without creating a duplicate or contradictory measurement. Neither context states a score, answer code, proposition ID, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"For Friday, no AWD is confirmed; the rental coordinator confirms only an automatic compact, which is permitted on every road in the Valley plan.\",\"Valley reaches the museum at 10:20 Friday, within its 10:00–11:00 entry window; the depot-to-museum drive lasts 1 hour 45 minutes.\",\"The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday.\",\"The plan records 47 minutes from that museum departure until arrival at the winery, with no intervening stop.\",\"That direct museum-to-winery run is the earliest obtainable without changing the Valley route, winery reservation, or confirmed vehicle.\",\"After the booked visit, the winery-to-lodging drive lasts exactly 2 hours. Lodging arrival is 15:00, and check-in closes at 15:15; every continuous driving segment is at most 2 hours.\",\"A room and a parking space are both confirmed for the Friday Valley plan and required Friday night.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\"],\"request\":\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "3"], "text": "The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday."}, {"path": ["evidence", "4"], "text": "The plan records 47 minutes from that museum departure until arrival at the winery, with no intervening stop."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday.", "negative_left": "The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday.", "negative_right": "The plan records 52 minutes from that museum departure until arrival at the winery, with no intervening stop.", "right": "The plan records 47 minutes from that museum departure until arrival at the winery, with no intervening stop."}, "verifier_independent_model": false}, "family": "scale-diverse-108-005", "id": "scale-diverse-108-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "For Friday, no AWD is confirmed; the rental coordinator confirms only an automatic compact, which is permitted on every road in the Valley plan.", "Valley reaches the museum at 10:20 Friday, within its 10:00–11:00 entry window; the depot-to-museum drive lasts 1 hour 45 minutes.", "The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday.", "The plan records 47 minutes from that museum departure until arrival at the winery, with no intervening stop.", "That direct museum-to-winery run is the earliest obtainable without changing the Valley route, winery reservation, or confirmed vehicle.", "After the booked visit, the winery-to-lodging drive lasts exactly 2 hours. Lodging arrival is 15:00, and check-in closes at 15:15; every continuous driving segment is at most 2 hours.", "A room and a parking space are both confirmed for the Friday Valley plan and required Friday night.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the unchanged scoring question and retain the traveler’s conditional selection rule and rest policy from the original state. The Friday, Valley-plan, vehicle, route, reservation, and scoring bindings remain intact. The two focus spans are complete factual sentences. Changing the direct drive from 47 to 52 minutes coherently changes the winery arrival from 12:13 to 12:18 relative to the fixed 12:15 start, without creating a duplicate or contradictory measurement. Neither context states a score, answer code, proposition ID, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"For Friday, no AWD is confirmed; the rental coordinator confirms only an automatic compact, which is permitted on every road in the Valley plan.\",\"Valley reaches the museum at 10:20 Friday, within its 10:00–11:00 entry window; the depot-to-museum drive lasts 1 hour 45 minutes.\",\"The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday.\",\"The plan records 47 minutes from that museum departure until arrival at the winery, with no intervening stop.\",\"That direct museum-to-winery run is the earliest obtainable without changing the Valley route, winery reservation, or confirmed vehicle.\",\"After the booked visit, the winery-to-lodging drive lasts exactly 2 hours. Lodging arrival is 15:00, and check-in closes at 15:15; every continuous driving segment is at most 2 hours.\",\"A room and a parking space are both confirmed for the Friday Valley plan and required Friday night.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\"],\"request\":\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "3"], "text": "The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday."}, {"path": ["evidence", "4"], "text": "The plan records 47 minutes from that museum departure until arrival at the winery, with no intervening stop."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday.", "negative_left": "The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday.", "negative_right": "The plan records 52 minutes from that museum departure until arrival at the winery, with no intervening stop.", "right": "The plan records 47 minutes from that museum departure until arrival at the winery, with no intervening stop."}, "verifier_independent_model": false}, "family": "scale-diverse-108-005", "id": "scale-diverse-108-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "For Friday, no AWD is confirmed; the rental coordinator confirms only an automatic compact, which is permitted on every road in the Valley plan.", "Valley reaches the museum at 10:20 Friday, within its 10:00–11:00 entry window; the depot-to-museum drive lasts 1 hour 45 minutes.", "The Friday Valley plan records an 11:26 Friday departure from the museum and lists the booked winery visit’s fixed start time as 12:15 that Friday.", "The plan records 52 minutes from that museum departure until arrival at the winery, with no intervening stop.", "That direct museum-to-winery run is the earliest obtainable without changing the Valley route, winery reservation, or confirmed vehicle.", "After the booked visit, the winery-to-lodging drive lasts exactly 2 hours. Lodging arrival is 15:00, and check-in closes at 15:15; every continuous driving segment is at most 2 hours.", "A room and a parking space are both confirmed for the Friday Valley plan and required Friday night.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the airport-to-hotel routing scope, Elena’s party and luggage requirements, the 18:40 deadline, and the mandatory ramp and wheelchair-securement policy. The two evidence spans are complete factual sentences. The counterfactual changes only the accessible van’s documented passenger capacity from three to two, which is coherent with the unchanged three-traveler manifest and creates no duplicate conflicting measurement within either context. Neither context includes a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At fictional Northstar Airport, dispatch staff reviewed Elena’s airport-to-hotel booking. At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party. At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of three travelers for the 18:35 accessible van. The booking listed four suitcases and Elena’s rigid wheelchair and requested pickup by 18:40. The van’s specification documented capacity for all four suitcases and the rigid wheelchair, together with a ramp and wheelchair securement. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan’s specification recorded no ramp. The accessible minibus was scheduled for 18:50, after the requested deadline. The airport assistance attendant’s assignment ended at the curb and did not provide transportation to Elena’s hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party."}, {"path": [], "text": "At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of three travelers for the 18:35 accessible van."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party.", "negative_left": "At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party.", "negative_right": "At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of two travelers for the 18:35 accessible van.", "right": "At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of three travelers for the 18:35 accessible van."}, "verifier_independent_model": false}, "family": "scale-diverse-109-001", "id": "scale-diverse-109-001-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At fictional Northstar Airport, dispatch staff reviewed Elena’s airport-to-hotel booking. At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party. At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of three travelers for the 18:35 accessible van. The booking listed four suitcases and Elena’s rigid wheelchair and requested pickup by 18:40. The van’s specification documented capacity for all four suitcases and the rigid wheelchair, together with a ramp and wheelchair securement. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan’s specification recorded no ramp. The accessible minibus was scheduled for 18:50, after the requested deadline. The airport assistance attendant’s assignment ended at the curb and did not provide transportation to Elena’s hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the airport-to-hotel routing scope, Elena’s party and luggage requirements, the 18:40 deadline, and the mandatory ramp and wheelchair-securement policy. The two evidence spans are complete factual sentences. The counterfactual changes only the accessible van’s documented passenger capacity from three to two, which is coherent with the unchanged three-traveler manifest and creates no duplicate conflicting measurement within either context. Neither context includes a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At fictional Northstar Airport, dispatch staff reviewed Elena’s airport-to-hotel booking. At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party. At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of three travelers for the 18:35 accessible van. The booking listed four suitcases and Elena’s rigid wheelchair and requested pickup by 18:40. The van’s specification documented capacity for all four suitcases and the rigid wheelchair, together with a ramp and wheelchair securement. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan’s specification recorded no ramp. The accessible minibus was scheduled for 18:50, after the requested deadline. The airport assistance attendant’s assignment ended at the curb and did not provide transportation to Elena’s hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party."}, {"path": [], "text": "At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of three travelers for the 18:35 accessible van."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party.", "negative_left": "At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party.", "negative_right": "At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of two travelers for the 18:35 accessible van.", "right": "At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of three travelers for the 18:35 accessible van."}, "verifier_independent_model": false}, "family": "scale-diverse-109-001", "id": "scale-diverse-109-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At fictional Northstar Airport, dispatch staff reviewed Elena’s airport-to-hotel booking. At 17:50 on 14 June 2026, the finalized transfer manifest identified Elena, Omar, and Priya as the only travelers in Elena’s entire airport-to-hotel party. At 18:00 on 14 June 2026, the vehicle capacity certificate recorded a maximum passenger capacity of two travelers for the 18:35 accessible van. The booking listed four suitcases and Elena’s rigid wheelchair and requested pickup by 18:40. The van’s specification documented capacity for all four suitcases and the rigid wheelchair, together with a ramp and wheelchair securement. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan’s specification recorded no ramp. The accessible minibus was scheduled for 18:50, after the requested deadline. The airport assistance attendant’s assignment ended at the curb and did not provide transportation to Elena’s hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same airport-to-hotel scope, Elena/party bindings, timing requirement, whole-party and luggage capacity requirements, and documented ramp and securement policy. The only counterfactual change is the accessible van’s passenger capacity from five places to two; this is coherent with the unchanged three-traveler count and creates no duplicate contradictory measurement. The two evidence spans are complete factual sentences, and neither context contains a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Dispatch handoff — Northstar Airport, 17 September 2026. The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party. Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents five passenger places in the 18:35 accessible van. The luggage manifest assigns four suitcases to Elena’s transfer, and VC-742 rates the van’s luggage bay for six suitcases. The vehicle file documents a dedicated bay approved for Elena’s rigid wheelchair, a boarding ramp, and four-point wheelchair securement. Its booked pickup is 18:35; Elena’s latest acceptable hotel pickup is 18:40. The 18:25 standard sedan’s inspection record explicitly reports no ramp. The accessible minibus is scheduled for 18:50, after the deadline. Airport assistance ends at the curb after baggage claim and does not include travel to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party."}, {"path": [], "text": "Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents five passenger places in the 18:35 accessible van."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party.", "negative_left": "The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party.", "negative_right": "Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents two passenger places in the 18:35 accessible van.", "right": "Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents five passenger places in the 18:35 accessible van."}, "verifier_independent_model": false}, "family": "scale-diverse-109-003", "id": "scale-diverse-109-003-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Dispatch handoff — Northstar Airport, 17 September 2026. The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party. Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents five passenger places in the 18:35 accessible van. The luggage manifest assigns four suitcases to Elena’s transfer, and VC-742 rates the van’s luggage bay for six suitcases. The vehicle file documents a dedicated bay approved for Elena’s rigid wheelchair, a boarding ramp, and four-point wheelchair securement. Its booked pickup is 18:35; Elena’s latest acceptable hotel pickup is 18:40. The 18:25 standard sedan’s inspection record explicitly reports no ramp. The accessible minibus is scheduled for 18:50, after the deadline. Airport assistance ends at the curb after baggage claim and does not include travel to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same airport-to-hotel scope, Elena/party bindings, timing requirement, whole-party and luggage capacity requirements, and documented ramp and securement policy. The only counterfactual change is the accessible van’s passenger capacity from five places to two; this is coherent with the unchanged three-traveler count and creates no duplicate contradictory measurement. The two evidence spans are complete factual sentences, and neither context contains a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Dispatch handoff — Northstar Airport, 17 September 2026. The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party. Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents five passenger places in the 18:35 accessible van. The luggage manifest assigns four suitcases to Elena’s transfer, and VC-742 rates the van’s luggage bay for six suitcases. The vehicle file documents a dedicated bay approved for Elena’s rigid wheelchair, a boarding ramp, and four-point wheelchair securement. Its booked pickup is 18:35; Elena’s latest acceptable hotel pickup is 18:40. The 18:25 standard sedan’s inspection record explicitly reports no ramp. The accessible minibus is scheduled for 18:50, after the deadline. Airport assistance ends at the curb after baggage claim and does not include travel to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party."}, {"path": [], "text": "Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents five passenger places in the 18:35 accessible van."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party.", "negative_left": "The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party.", "negative_right": "Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents two passenger places in the 18:35 accessible van.", "right": "Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents five passenger places in the 18:35 accessible van."}, "verifier_independent_model": false}, "family": "scale-diverse-109-003", "id": "scale-diverse-109-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Dispatch handoff — Northstar Airport, 17 September 2026. The operational handoff recorded at 17:55 on 17 September 2026 lists exactly three travelers in Elena’s entire Northstar Airport-to-hotel party. Capacity record VC-742, valid for Elena’s airport-to-hotel transfer on 17 September 2026, documents two passenger places in the 18:35 accessible van. The luggage manifest assigns four suitcases to Elena’s transfer, and VC-742 rates the van’s luggage bay for six suitcases. The vehicle file documents a dedicated bay approved for Elena’s rigid wheelchair, a boarding ramp, and four-point wheelchair securement. Its booked pickup is 18:35; Elena’s latest acceptable hotel pickup is 18:40. The 18:25 standard sedan’s inspection record explicitly reports no ramp. The accessible minibus is scheduled for 18:50, after the deadline. Airport assistance ends at the curb after baggage claim and does not include travel to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the governing suitability policy through the unchanged questions and the repeated context policy, while retaining Elena’s airport-to-hotel transfer, party, accessibility, location, and timing bindings. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the accessible van’s passenger capacity from five to two without creating a duplicate or conflicting capacity assertion. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Field note, Northstar Airport, 17 September 2026. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena. The vehicle record issued on 17 September 2026 documents an exact passenger capacity of five travelers for the 18:35 accessible van. Elena’s manifest also records four suitcases and a rigid wheelchair. The van’s record documents room for four suitcases and one rigid wheelchair, together with a ramp and wheelchair securement. Its pickup is scheduled for 18:35, before Elena’s requested 18:40 deadline. The 18:25 standard sedan’s record explicitly states that it has no ramp. The accessible minibus is scheduled for 18:50, after the deadline. The airport assistance attendant’s service ends at the curb and does not provide transportation from Northstar Airport to Elena’s hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena."}, {"path": [], "text": "The vehicle record issued on 17 September 2026 documents an exact passenger capacity of five travelers for the 18:35 accessible van."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena.", "negative_left": "Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena.", "negative_right": "The vehicle record issued on 17 September 2026 documents an exact passenger capacity of two travelers for the 18:35 accessible van.", "right": "The vehicle record issued on 17 September 2026 documents an exact passenger capacity of five travelers for the 18:35 accessible van."}, "verifier_independent_model": false}, "family": "scale-diverse-109-004", "id": "scale-diverse-109-004-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Field note, Northstar Airport, 17 September 2026. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena. The vehicle record issued on 17 September 2026 documents an exact passenger capacity of five travelers for the 18:35 accessible van. Elena’s manifest also records four suitcases and a rigid wheelchair. The van’s record documents room for four suitcases and one rigid wheelchair, together with a ramp and wheelchair securement. Its pickup is scheduled for 18:35, before Elena’s requested 18:40 deadline. The 18:25 standard sedan’s record explicitly states that it has no ramp. The accessible minibus is scheduled for 18:50, after the deadline. The airport assistance attendant’s service ends at the curb and does not provide transportation from Northstar Airport to Elena’s hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the governing suitability policy through the unchanged questions and the repeated context policy, while retaining Elena’s airport-to-hotel transfer, party, accessibility, location, and timing bindings. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the accessible van’s passenger capacity from five to two without creating a duplicate or conflicting capacity assertion. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Field note, Northstar Airport, 17 September 2026. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena. The vehicle record issued on 17 September 2026 documents an exact passenger capacity of five travelers for the 18:35 accessible van. Elena’s manifest also records four suitcases and a rigid wheelchair. The van’s record documents room for four suitcases and one rigid wheelchair, together with a ramp and wheelchair securement. Its pickup is scheduled for 18:35, before Elena’s requested 18:40 deadline. The 18:25 standard sedan’s record explicitly states that it has no ramp. The accessible minibus is scheduled for 18:50, after the deadline. The airport assistance attendant’s service ends at the curb and does not provide transportation from Northstar Airport to Elena’s hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena."}, {"path": [], "text": "The vehicle record issued on 17 September 2026 documents an exact passenger capacity of five travelers for the 18:35 accessible van."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena.", "negative_left": "Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena.", "negative_right": "The vehicle record issued on 17 September 2026 documents an exact passenger capacity of two travelers for the 18:35 accessible van.", "right": "The vehicle record issued on 17 September 2026 documents an exact passenger capacity of five travelers for the 18:35 accessible van."}, "verifier_independent_model": false}, "family": "scale-diverse-109-004", "id": "scale-diverse-109-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Field note, Northstar Airport, 17 September 2026. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena’s finalized airport-to-hotel manifest for 17 September 2026 lists exactly three travelers in her entire party, including Elena. The vehicle record issued on 17 September 2026 documents an exact passenger capacity of two travelers for the 18:35 accessible van. Elena’s manifest also records four suitcases and a rigid wheelchair. The van’s record documents room for four suitcases and one rigid wheelchair, together with a ramp and wheelchair securement. Its pickup is scheduled for 18:35, before Elena’s requested 18:40 deadline. The 18:25 standard sedan’s record explicitly states that it has no ramp. The accessible minibus is scheduled for 18:50, after the deadline. The airport assistance attendant’s service ends at the curb and does not provide transportation from Northstar Airport to Elena’s hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same dispatch eligibility, capacity, assistance-scope, and ranking policies, as well as the same airport, routing choices, curb-ready time, and request. The two focus spans are complete factual sentences. The counterfactual coherently changes only Access B’s luggage capacity from seven bags to five; no unchanged evidence supplies a conflicting capacity. Neither context explicitly states a selected option, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"At 18:30, the three-person party confirmed that reaching its hotel requires travel beyond the curb. The passenger must remain in her rigid power chair, so one occupied-chair securement is required and transfer into the 18:41 standard van is impossible.\",\"At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party.\",\"At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as seven bags.\",\"At 18:37, vehicle records listed Access A with a ramp, two securements, four passenger places, a luggage bay approved for the manifest, and an 18:50 pickup. Access B was listed with a lift, two securements, four passenger places, and an 18:42 pickup.\",\"At 18:38, the 19:00 hotel shuttle was confirmed to have a two-person and four-bag limit.\",\"At 18:40, the assistance attendant delivered the party to the curb and confirmed that the attendant's service ends there.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party."}, {"path": ["evidence", "2"], "text": "At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as seven bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party.", "negative_left": "At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party.", "negative_right": "At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as five bags.", "right": "At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as seven bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-001", "id": "scale-diverse-110-001-base", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["At 18:30, the three-person party confirmed that reaching its hotel requires travel beyond the curb. The passenger must remain in her rigid power chair, so one occupied-chair securement is required and transfer into the 18:41 standard van is impossible.", "At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party.", "At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as seven bags.", "At 18:37, vehicle records listed Access A with a ramp, two securements, four passenger places, a luggage bay approved for the manifest, and an 18:50 pickup. Access B was listed with a lift, two securements, four passenger places, and an 18:42 pickup.", "At 18:38, the 19:00 hotel shuttle was confirmed to have a two-person and four-bag limit.", "At 18:40, the assistance attendant delivered the party to the curb and confirmed that the attendant's service ends there."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_b"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same dispatch eligibility, capacity, assistance-scope, and ranking policies, as well as the same airport, routing choices, curb-ready time, and request. The two focus spans are complete factual sentences. The counterfactual coherently changes only Access B’s luggage capacity from seven bags to five; no unchanged evidence supplies a conflicting capacity. Neither context explicitly states a selected option, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"At 18:30, the three-person party confirmed that reaching its hotel requires travel beyond the curb. The passenger must remain in her rigid power chair, so one occupied-chair securement is required and transfer into the 18:41 standard van is impossible.\",\"At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party.\",\"At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as seven bags.\",\"At 18:37, vehicle records listed Access A with a ramp, two securements, four passenger places, a luggage bay approved for the manifest, and an 18:50 pickup. Access B was listed with a lift, two securements, four passenger places, and an 18:42 pickup.\",\"At 18:38, the 19:00 hotel shuttle was confirmed to have a two-person and four-bag limit.\",\"At 18:40, the assistance attendant delivered the party to the curb and confirmed that the attendant's service ends there.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party."}, {"path": ["evidence", "2"], "text": "At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as seven bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party.", "negative_left": "At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party.", "negative_right": "At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as five bags.", "right": "At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as seven bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-001", "id": "scale-diverse-110-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["At 18:30, the three-person party confirmed that reaching its hotel requires travel beyond the curb. The passenger must remain in her rigid power chair, so one occupied-chair securement is required and transfer into the 18:41 standard van is impossible.", "At 18:35 on 17 September 2026, the Northstar Airport arrival manifest recorded six bags for the arriving party.", "At 18:36 on 17 September 2026, Northstar Airport's vehicle inventory recorded Access B's luggage capacity as five bags.", "At 18:37, vehicle records listed Access A with a ramp, two securements, four passenger places, a luggage bay approved for the manifest, and an 18:50 pickup. Access B was listed with a lift, two securements, four passenger places, and an 18:42 pickup.", "At 18:38, the 19:00 hotel shuttle was confirmed to have a two-person and four-bag limit.", "At 18:40, the assistance attendant delivered the party to the curb and confirmed that the attendant's service ends there."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_a"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same dispatch eligibility, capacity, assistance-scope, and pickup-ranking policies, as well as the same party, routing request, curb-ready time, candidate vehicles, and hotel-bound path. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Access B’s luggage capacity from nine bags to six, which is consistent with the unchanged seven-bag manifest and does not conflict with the separate Access B record, which states no luggage figure. Neither context supplies a gold answer, code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The three-person party needs road transport from the curb to its hotel; the attendant’s assignment ends at the curb, where the party will be ready at 18:40.\",\"At 18:40, the arriving party's finalized baggage manifest lists seven bags.\",\"For its 18:42 pickup, Access B has a certified luggage capacity of nine bags.\",\"The passenger must remain seated in her rigid power chair, which cannot be folded, and the party requires one occupied-chair securement.\",\"Access A’s certified record documents a ramp, two occupied-chair securements, capacity for four passengers and eight bags, and an 18:50 pickup.\",\"Access B’s certified record documents a lift, two occupied-chair securements, capacity for four passengers, and an 18:42 pickup.\",\"The standard van is listed for 18:41 and requires the passenger to leave the chair.\",\"The 19:00 hotel shuttle is limited to two passengers and four bags.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:40, the arriving party's finalized baggage manifest lists seven bags."}, {"path": ["evidence", "2"], "text": "For its 18:42 pickup, Access B has a certified luggage capacity of nine bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:40, the arriving party's finalized baggage manifest lists seven bags.", "negative_left": "At 18:40, the arriving party's finalized baggage manifest lists seven bags.", "negative_right": "For its 18:42 pickup, Access B has a certified luggage capacity of six bags.", "right": "For its 18:42 pickup, Access B has a certified luggage capacity of nine bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-002", "id": "scale-diverse-110-002-base", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The three-person party needs road transport from the curb to its hotel; the attendant’s assignment ends at the curb, where the party will be ready at 18:40.", "At 18:40, the arriving party's finalized baggage manifest lists seven bags.", "For its 18:42 pickup, Access B has a certified luggage capacity of nine bags.", "The passenger must remain seated in her rigid power chair, which cannot be folded, and the party requires one occupied-chair securement.", "Access A’s certified record documents a ramp, two occupied-chair securements, capacity for four passengers and eight bags, and an 18:50 pickup.", "Access B’s certified record documents a lift, two occupied-chair securements, capacity for four passengers, and an 18:42 pickup.", "The standard van is listed for 18:41 and requires the passenger to leave the chair.", "The 19:00 hotel shuttle is limited to two passengers and four bags."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_b"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same dispatch eligibility, capacity, assistance-scope, and pickup-ranking policies, as well as the same party, routing request, curb-ready time, candidate vehicles, and hotel-bound path. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Access B’s luggage capacity from nine bags to six, which is consistent with the unchanged seven-bag manifest and does not conflict with the separate Access B record, which states no luggage figure. Neither context supplies a gold answer, code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The three-person party needs road transport from the curb to its hotel; the attendant’s assignment ends at the curb, where the party will be ready at 18:40.\",\"At 18:40, the arriving party's finalized baggage manifest lists seven bags.\",\"For its 18:42 pickup, Access B has a certified luggage capacity of nine bags.\",\"The passenger must remain seated in her rigid power chair, which cannot be folded, and the party requires one occupied-chair securement.\",\"Access A’s certified record documents a ramp, two occupied-chair securements, capacity for four passengers and eight bags, and an 18:50 pickup.\",\"Access B’s certified record documents a lift, two occupied-chair securements, capacity for four passengers, and an 18:42 pickup.\",\"The standard van is listed for 18:41 and requires the passenger to leave the chair.\",\"The 19:00 hotel shuttle is limited to two passengers and four bags.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:40, the arriving party's finalized baggage manifest lists seven bags."}, {"path": ["evidence", "2"], "text": "For its 18:42 pickup, Access B has a certified luggage capacity of nine bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:40, the arriving party's finalized baggage manifest lists seven bags.", "negative_left": "At 18:40, the arriving party's finalized baggage manifest lists seven bags.", "negative_right": "For its 18:42 pickup, Access B has a certified luggage capacity of six bags.", "right": "For its 18:42 pickup, Access B has a certified luggage capacity of nine bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-002", "id": "scale-diverse-110-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The three-person party needs road transport from the curb to its hotel; the attendant’s assignment ends at the curb, where the party will be ready at 18:40.", "At 18:40, the arriving party's finalized baggage manifest lists seven bags.", "For its 18:42 pickup, Access B has a certified luggage capacity of six bags.", "The passenger must remain seated in her rigid power chair, which cannot be folded, and the party requires one occupied-chair securement.", "Access A’s certified record documents a ramp, two occupied-chair securements, capacity for four passengers and eight bags, and an 18:50 pickup.", "Access B’s certified record documents a lift, two occupied-chair securements, capacity for four passengers, and an 18:42 pickup.", "The standard van is listed for 18:41 and requires the passenger to leave the chair.", "The 19:00 hotel shuttle is limited to two passengers and four bags."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_a"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both inputs preserve the questions object verbatim and retain the original governing policy, request, airport, routing entities, pickup times, curb-ready time, and decision scope. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Access B’s luggage capacity from seven to four bags, with no conflicting duplicate capacity stated within that context, so it remains coherent with the unchanged facts. Neither context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or output instruction beyond the preserved question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"accessible_access_a\":\"Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.\",\"accessible_access_b\":\"Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.\",\"assistance_staff_only\":\"Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.\",\"hotel_shuttle\":\"Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.\",\"none_of_above\":\"Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.\",\"standard_van_dispatch\":\"Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded.\"},\"instructions\":\"Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.\",\"type\":\"choice\"}},\"state\":{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The operational handoff says the party needs transport from the curb to its hotel. Three travelers are present, and one occupied-chair securement is required for the passenger's rigid power chair; she cannot leave or fold it. The attendant's curb-ready time is 18:40.\",\"The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags.\",\"Access A's dispatch sheet documents a ramp, four occupied-chair securements, room for four passengers and six bags, and an 18:50 pickup.\",\"Access B's inspection documents a lift and two occupied-chair securements; its dispatch sheet lists room for four passengers and an 18:42 pickup.\",\"The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of seven bags.\",\"The listed 19:00 hotel shuttle is limited to two passengers and four bags.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags."}, {"path": ["state", "evidence", "4"], "text": "The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of seven bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags.", "negative_left": "The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags.", "negative_right": "The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of four bags.", "right": "The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of seven bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-003", "id": "scale-diverse-110-003-base", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The operational handoff says the party needs transport from the curb to its hotel. Three travelers are present, and one occupied-chair securement is required for the passenger's rigid power chair; she cannot leave or fold it. The attendant's curb-ready time is 18:40.", "The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags.", "Access A's dispatch sheet documents a ramp, four occupied-chair securements, room for four passengers and six bags, and an 18:50 pickup.", "Access B's inspection documents a lift and two occupied-chair securements; its dispatch sheet lists room for four passengers and an 18:42 pickup.", "The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of seven bags.", "The listed 19:00 hotel shuttle is limited to two passengers and four bags."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_b"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both inputs preserve the questions object verbatim and retain the original governing policy, request, airport, routing entities, pickup times, curb-ready time, and decision scope. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Access B’s luggage capacity from seven to four bags, with no conflicting duplicate capacity stated within that context, so it remains coherent with the unchanged facts. Neither context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or output instruction beyond the preserved question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"accessible_access_a\":\"Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.\",\"accessible_access_b\":\"Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.\",\"assistance_staff_only\":\"Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.\",\"hotel_shuttle\":\"Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.\",\"none_of_above\":\"Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.\",\"standard_van_dispatch\":\"Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded.\"},\"instructions\":\"Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.\",\"type\":\"choice\"}},\"state\":{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The operational handoff says the party needs transport from the curb to its hotel. Three travelers are present, and one occupied-chair securement is required for the passenger's rigid power chair; she cannot leave or fold it. The attendant's curb-ready time is 18:40.\",\"The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags.\",\"Access A's dispatch sheet documents a ramp, four occupied-chair securements, room for four passengers and six bags, and an 18:50 pickup.\",\"Access B's inspection documents a lift and two occupied-chair securements; its dispatch sheet lists room for four passengers and an 18:42 pickup.\",\"The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of seven bags.\",\"The listed 19:00 hotel shuttle is limited to two passengers and four bags.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags."}, {"path": ["state", "evidence", "4"], "text": "The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of seven bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags.", "negative_left": "The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags.", "negative_right": "The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of four bags.", "right": "The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of seven bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-003", "id": "scale-diverse-110-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The operational handoff says the party needs transport from the curb to its hotel. Three travelers are present, and one occupied-chair securement is required for the passenger's rigid power chair; she cannot leave or fold it. The attendant's curb-ready time is 18:40.", "The Northstar Airport curb-handoff record dated 17 September 2026 lists the arriving party with five bags.", "Access A's dispatch sheet documents a ramp, four occupied-chair securements, room for four passengers and six bags, and an 18:50 pickup.", "Access B's inspection documents a lift and two occupied-chair securements; its dispatch sheet lists room for four passengers and an 18:42 pickup.", "The Access B vehicle inspection completed at 18:20 on 17 September 2026 records a maximum luggage capacity of four bags.", "The listed 19:00 hotel shuttle is limited to two passengers and four bags."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_a"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing dispatch, accessibility, assistance-scope, and ranking policy, while the unchanged questions object preserves the detailed choice criteria. The request, entities, curb-ready time, and relevant evidence paths remain bound to the original question. The two focus spans are complete factual sentences. The counterfactual coherently changes only Access B’s luggage capacity from seven bags to four, without conflicting with any other measurement or assertion. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The attendant will bring the three-person party to the curb by 18:40, but their request includes a road ride from there to the hotel. The passenger must remain in her rigid, nonfolding power chair, requiring one occupied-chair position.\",\"The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags.\",\"Fleet records list Access A with a documented ramp, two occupied-chair restraints, room for four passengers and six bags, and an 18:50 pickup.\",\"The dispatch manifest dated 17 September 2026 lists Access B with space for seven bags.\",\"The same manifest records Access B's documented lift, two occupied-chair restraints, four-passenger limit, and 18:42 pickup.\",\"The 19:00 hotel shuttle is limited to two passengers and four bags.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags."}, {"path": ["evidence", "3"], "text": "The dispatch manifest dated 17 September 2026 lists Access B with space for seven bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags.", "negative_left": "The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags.", "negative_right": "The dispatch manifest dated 17 September 2026 lists Access B with space for four bags.", "right": "The dispatch manifest dated 17 September 2026 lists Access B with space for seven bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-004", "id": "scale-diverse-110-004-base", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The attendant will bring the three-person party to the curb by 18:40, but their request includes a road ride from there to the hotel. The passenger must remain in her rigid, nonfolding power chair, requiring one occupied-chair position.", "The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags.", "Fleet records list Access A with a documented ramp, two occupied-chair restraints, room for four passengers and six bags, and an 18:50 pickup.", "The dispatch manifest dated 17 September 2026 lists Access B with space for seven bags.", "The same manifest records Access B's documented lift, two occupied-chair restraints, four-passenger limit, and 18:42 pickup.", "The 19:00 hotel shuttle is limited to two passengers and four bags."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_b"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing dispatch, accessibility, assistance-scope, and ranking policy, while the unchanged questions object preserves the detailed choice criteria. The request, entities, curb-ready time, and relevant evidence paths remain bound to the original question. The two focus spans are complete factual sentences. The counterfactual coherently changes only Access B’s luggage capacity from seven bags to four, without conflicting with any other measurement or assertion. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The attendant will bring the three-person party to the curb by 18:40, but their request includes a road ride from there to the hotel. The passenger must remain in her rigid, nonfolding power chair, requiring one occupied-chair position.\",\"The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags.\",\"Fleet records list Access A with a documented ramp, two occupied-chair restraints, room for four passengers and six bags, and an 18:50 pickup.\",\"The dispatch manifest dated 17 September 2026 lists Access B with space for seven bags.\",\"The same manifest records Access B's documented lift, two occupied-chair restraints, four-passenger limit, and 18:42 pickup.\",\"The 19:00 hotel shuttle is limited to two passengers and four bags.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags."}, {"path": ["evidence", "3"], "text": "The dispatch manifest dated 17 September 2026 lists Access B with space for seven bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags.", "negative_left": "The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags.", "negative_right": "The dispatch manifest dated 17 September 2026 lists Access B with space for four bags.", "right": "The dispatch manifest dated 17 September 2026 lists Access B with space for seven bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-004", "id": "scale-diverse-110-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The attendant will bring the three-person party to the curb by 18:40, but their request includes a road ride from there to the hotel. The passenger must remain in her rigid, nonfolding power chair, requiring one occupied-chair position.", "The arriving party recorded at Northstar Airport's curb desk at 18:40 has five bags.", "Fleet records list Access A with a documented ramp, two occupied-chair restraints, room for four passengers and six bags, and an 18:50 pickup.", "The dispatch manifest dated 17 September 2026 lists Access B with space for four bags.", "The same manifest records Access B's documented lift, two occupied-chair restraints, four-passenger limit, and 18:42 pickup.", "The 19:00 hotel shuttle is limited to two passengers and four bags."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_a"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy, request, airport, routing entities, and curb-ready timing, while the unchanged questions object preserves the full option criteria. The two focus spans are complete factual sentences. The counterfactual coherently changes Access B’s luggage capacity from eight bags to six bags without creating a duplicate or contradictory measurement within that context. Neither context contains a gold answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"At 18:34, three travelers requested onward road transport to their hotel. One passenger occupied a rigid power chair and said she could not transfer out of it.\",\"At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer.\",\"At 18:36, Access B's vehicle manifest recorded space for eight bags.\",\"At 18:37, Access A was logged with a ramp, four occupied-chair restraints, room for four passengers and nine bags, and an 18:50 pickup.\",\"At 18:38, Access B was logged with a lift, two occupied-chair restraints, room for four passengers, and an 18:42 pickup.\",\"At 18:39, the 19:00 hotel shuttle was confirmed to have limits of two passengers and four bags.\",\"The assistance attendant completed aircraft-to-curb help at 18:40, leaving the party ready for the requested road journey.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer."}, {"path": ["evidence", "2"], "text": "At 18:36, Access B's vehicle manifest recorded space for eight bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer.", "negative_left": "At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer.", "negative_right": "At 18:36, Access B's vehicle manifest recorded space for six bags.", "right": "At 18:36, Access B's vehicle manifest recorded space for eight bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-005", "id": "scale-diverse-110-005-base", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["At 18:34, three travelers requested onward road transport to their hotel. One passenger occupied a rigid power chair and said she could not transfer out of it.", "At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer.", "At 18:36, Access B's vehicle manifest recorded space for eight bags.", "At 18:37, Access A was logged with a ramp, four occupied-chair restraints, room for four passengers and nine bags, and an 18:50 pickup.", "At 18:38, Access B was logged with a lift, two occupied-chair restraints, room for four passengers, and an 18:42 pickup.", "At 18:39, the 19:00 hotel shuttle was confirmed to have limits of two passengers and four bags.", "The assistance attendant completed aircraft-to-curb help at 18:40, leaving the party ready for the requested road journey."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_b"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy, request, airport, routing entities, and curb-ready timing, while the unchanged questions object preserves the full option criteria. The two focus spans are complete factual sentences. The counterfactual coherently changes Access B’s luggage capacity from eight bags to six bags without creating a duplicate or contradictory measurement within that context. Neither context contains a gold answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"At 18:34, three travelers requested onward road transport to their hotel. One passenger occupied a rigid power chair and said she could not transfer out of it.\",\"At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer.\",\"At 18:36, Access B's vehicle manifest recorded space for eight bags.\",\"At 18:37, Access A was logged with a ramp, four occupied-chair restraints, room for four passengers and nine bags, and an 18:50 pickup.\",\"At 18:38, Access B was logged with a lift, two occupied-chair restraints, room for four passengers, and an 18:42 pickup.\",\"At 18:39, the 19:00 hotel shuttle was confirmed to have limits of two passengers and four bags.\",\"The assistance attendant completed aircraft-to-curb help at 18:40, leaving the party ready for the requested road journey.\"],\"request\":\"Which single routing choice is suitable and ranks highest under the stated policy?\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer."}, {"path": ["evidence", "2"], "text": "At 18:36, Access B's vehicle manifest recorded space for eight bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer.", "negative_left": "At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer.", "negative_right": "At 18:36, Access B's vehicle manifest recorded space for six bags.", "right": "At 18:36, Access B's vehicle manifest recorded space for eight bags."}, "verifier_independent_model": false}, "family": "scale-diverse-110-005", "id": "scale-diverse-110-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["At 18:34, three travelers requested onward road transport to their hotel. One passenger occupied a rigid power chair and said she could not transfer out of it.", "At 18:35 at Northstar Airport, the arriving party presented seven bags for the road transfer.", "At 18:36, Access B's vehicle manifest recorded space for six bags.", "At 18:37, Access A was logged with a ramp, four occupied-chair restraints, room for four passengers and nine bags, and an 18:50 pickup.", "At 18:38, Access B was logged with a lift, two occupied-chair restraints, room for four passengers, and an 18:42 pickup.", "At 18:39, the 19:00 hotel shuttle was confirmed to have limits of two passengers and four bags.", "The assistance attendant completed aircraft-to-curb help at 18:40, leaving the party ready for the requested road journey."], "request": "Which single routing choice is suitable and ranks highest under the stated policy?"}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_a"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original accessibility-routing and suitability/ranking policy, as well as the question’s focus on LiftVan B, this party, and pickup by 18:35. The two focus-evidence spans are complete factual sentences. The counterfactual changes only LiftVan B’s suitcase capacity from eight to four while retaining the six-suitcase count, creating no contradictory duplicate measurement. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"arriving passenger\",\"text\":\"I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35.\"},{\"speaker\":\"dispatch log\",\"text\":\"At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B.\"},{\"speaker\":\"trip manifest\",\"text\":\"At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of eight suitcases for that party's transfer.\"},{\"speaker\":\"transfer dispatcher\",\"text\":\"The only logged transfer options were sedan S at 18:18, LiftVan B at 18:28, and LiftVan C at 18:34. LiftVan B was assigned as the party’s accessible vehicle; its file explicitly documented a ramp and wheelchair tie-downs. Sedan S’s file had no explicit documented evidence of boarding equipment.\"},{\"speaker\":\"capacity record\",\"text\":\"The party count was four travelers, including one rigid-wheelchair user. LiftVan B’s documented limits were four travelers and one rigid wheelchair.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"Policy routes non-transfer passengers to accessible dispatch.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"},{\"speaker\":\"airport assistance booking\",\"text\":\"Boarding assistance was booked at the curb for 18:25, before LiftVan B’s 18:28 pickup.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B."}, {"path": ["2", "text"], "text": "At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of eight suitcases for that party's transfer."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B.", "negative_left": "At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B.", "negative_right": "At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of four suitcases for that party's transfer.", "right": "At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of eight suitcases for that party's transfer."}, "verifier_independent_model": false}, "family": "scale-diverse-111-001", "id": "scale-diverse-111-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "arriving passenger", "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"speaker": "dispatch log", "text": "At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B."}, {"speaker": "trip manifest", "text": "At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of eight suitcases for that party's transfer."}, {"speaker": "transfer dispatcher", "text": "The only logged transfer options were sedan S at 18:18, LiftVan B at 18:28, and LiftVan C at 18:34. LiftVan B was assigned as the party’s accessible vehicle; its file explicitly documented a ramp and wheelchair tie-downs. Sedan S’s file had no explicit documented evidence of boarding equipment."}, {"speaker": "capacity record", "text": "The party count was four travelers, including one rigid-wheelchair user. LiftVan B’s documented limits were four travelers and one rigid wheelchair."}, {"speaker": "airport assistance attendant", "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"speaker": "airport assistance attendant", "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}, {"speaker": "airport assistance booking", "text": "Boarding assistance was booked at the curb for 18:25, before LiftVan B’s 18:28 pickup."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original accessibility-routing and suitability/ranking policy, as well as the question’s focus on LiftVan B, this party, and pickup by 18:35. The two focus-evidence spans are complete factual sentences. The counterfactual changes only LiftVan B’s suitcase capacity from eight to four while retaining the six-suitcase count, creating no contradictory duplicate measurement. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"arriving passenger\",\"text\":\"I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35.\"},{\"speaker\":\"dispatch log\",\"text\":\"At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B.\"},{\"speaker\":\"trip manifest\",\"text\":\"At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of eight suitcases for that party's transfer.\"},{\"speaker\":\"transfer dispatcher\",\"text\":\"The only logged transfer options were sedan S at 18:18, LiftVan B at 18:28, and LiftVan C at 18:34. LiftVan B was assigned as the party’s accessible vehicle; its file explicitly documented a ramp and wheelchair tie-downs. Sedan S’s file had no explicit documented evidence of boarding equipment.\"},{\"speaker\":\"capacity record\",\"text\":\"The party count was four travelers, including one rigid-wheelchair user. LiftVan B’s documented limits were four travelers and one rigid wheelchair.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"Policy routes non-transfer passengers to accessible dispatch.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"},{\"speaker\":\"airport assistance booking\",\"text\":\"Boarding assistance was booked at the curb for 18:25, before LiftVan B’s 18:28 pickup.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B."}, {"path": ["2", "text"], "text": "At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of eight suitcases for that party's transfer."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B.", "negative_left": "At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B.", "negative_right": "At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of four suitcases for that party's transfer.", "right": "At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of eight suitcases for that party's transfer."}, "verifier_independent_model": false}, "family": "scale-diverse-111-001", "id": "scale-diverse-111-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "arriving passenger", "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"speaker": "dispatch log", "text": "At 17:50 on 14 June 2026, the dispatcher recorded exactly six suitcases for the party assigned to LiftVan B."}, {"speaker": "trip manifest", "text": "At 17:55 on 14 June 2026, LiftVan B's trip manifest recorded a suitcase capacity of four suitcases for that party's transfer."}, {"speaker": "transfer dispatcher", "text": "The only logged transfer options were sedan S at 18:18, LiftVan B at 18:28, and LiftVan C at 18:34. LiftVan B was assigned as the party’s accessible vehicle; its file explicitly documented a ramp and wheelchair tie-downs. Sedan S’s file had no explicit documented evidence of boarding equipment."}, {"speaker": "capacity record", "text": "The party count was four travelers, including one rigid-wheelchair user. LiftVan B’s documented limits were four travelers and one rigid wheelchair."}, {"speaker": "airport assistance attendant", "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"speaker": "airport assistance attendant", "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}, {"speaker": "airport assistance booking", "text": "Boarding assistance was booked at the curb for 18:25, before LiftVan B’s 18:28 pickup."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing suitability and ranking policy and retain the question’s focus on LiftVan B for the same party, pickup deadline, and relevant time. The two evidence spans are complete factual sentences. The counterfactual coherently changes only LiftVan B’s suitcase capacity from eight to six while the manifest remains seven, without creating a contradictory duplicate measurement. Neither constructed context embeds a gold answer, output code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"arriving passenger\",\"text\":\"I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35.\"},{\"speaker\":\"manifest clerk\",\"text\":\"At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases.\"},{\"speaker\":\"certification officer\",\"text\":\"LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly eight suitcases.\"},{\"speaker\":\"transfer dispatcher\",\"text\":\"The reconciliation log lists exactly three transfer options: sedan S, LiftVan B, and LiftVan C. Sedan S is due at 18:18 and lacks explicit documented evidence of boarding equipment. LiftVan B is due at 18:28 and is recorded as accessible for this non-transfer passenger, with a verified ramp and tie-downs and certified room for four travelers and one rigid wheelchair. LiftVan C is due at 18:34. The party has four travelers and one rigid wheelchair.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"Boarding assistance is booked at the curb for 18:25. Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases."}, {"path": ["2", "text"], "text": "LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly eight suitcases."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases.", "negative_left": "At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases.", "negative_right": "LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly six suitcases.", "right": "LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly eight suitcases."}, "verifier_independent_model": false}, "family": "scale-diverse-111-002", "id": "scale-diverse-111-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "arriving passenger", "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"speaker": "manifest clerk", "text": "At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases."}, {"speaker": "certification officer", "text": "LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly eight suitcases."}, {"speaker": "transfer dispatcher", "text": "The reconciliation log lists exactly three transfer options: sedan S, LiftVan B, and LiftVan C. Sedan S is due at 18:18 and lacks explicit documented evidence of boarding equipment. LiftVan B is due at 18:28 and is recorded as accessible for this non-transfer passenger, with a verified ramp and tie-downs and certified room for four travelers and one rigid wheelchair. LiftVan C is due at 18:34. The party has four travelers and one rigid wheelchair."}, {"speaker": "airport assistance attendant", "text": "Boarding assistance is booked at the curb for 18:25. Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing suitability and ranking policy and retain the question’s focus on LiftVan B for the same party, pickup deadline, and relevant time. The two evidence spans are complete factual sentences. The counterfactual coherently changes only LiftVan B’s suitcase capacity from eight to six while the manifest remains seven, without creating a contradictory duplicate measurement. Neither constructed context embeds a gold answer, output code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"arriving passenger\",\"text\":\"I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35.\"},{\"speaker\":\"manifest clerk\",\"text\":\"At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases.\"},{\"speaker\":\"certification officer\",\"text\":\"LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly eight suitcases.\"},{\"speaker\":\"transfer dispatcher\",\"text\":\"The reconciliation log lists exactly three transfer options: sedan S, LiftVan B, and LiftVan C. Sedan S is due at 18:18 and lacks explicit documented evidence of boarding equipment. LiftVan B is due at 18:28 and is recorded as accessible for this non-transfer passenger, with a verified ramp and tie-downs and certified room for four travelers and one rigid wheelchair. LiftVan C is due at 18:34. The party has four travelers and one rigid wheelchair.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"Boarding assistance is booked at the curb for 18:25. Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases."}, {"path": ["2", "text"], "text": "LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly eight suitcases."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases.", "negative_left": "At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases.", "negative_right": "LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly six suitcases.", "right": "LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly eight suitcases."}, "verifier_independent_model": false}, "family": "scale-diverse-111-002", "id": "scale-diverse-111-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "arriving passenger", "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"speaker": "manifest clerk", "text": "At 18:00 on 17 September 2026, the finalized manifest for the party booked on LiftVan B listed exactly seven suitcases."}, {"speaker": "certification officer", "text": "LiftVan B's capacity certificate, valid at 18:00 on 17 September 2026, lists a suitcase capacity of exactly six suitcases."}, {"speaker": "transfer dispatcher", "text": "The reconciliation log lists exactly three transfer options: sedan S, LiftVan B, and LiftVan C. Sedan S is due at 18:18 and lacks explicit documented evidence of boarding equipment. LiftVan B is due at 18:28 and is recorded as accessible for this non-transfer passenger, with a verified ramp and tie-downs and certified room for four travelers and one rigid wheelchair. LiftVan C is due at 18:34. The party has four travelers and one rigid wheelchair."}, {"speaker": "airport assistance attendant", "text": "Boarding assistance is booked at the curb for 18:25. Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same suitability requirements, accessible-dispatch rule, earliest-ETA ranking rule, LiftVan B decision target, party/chair scope, and 18:35 pickup deadline. The focus evidence consists of two complete factual sentences. The sole counterfactual change lowers LiftVan B’s documented suitcase capacity from eight to four while the party still has six suitcases; this is a coherent capacity shortfall, not a contradictory duplicate measurement. Neither constructed context contains an answer code, explicit yes/no conclusion, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"arriving passenger\",\"text\":\"I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35. Our party has four travelers and one rigid wheelchair.\"},{\"speaker\":\"transfer dispatcher\",\"text\":\"The final transfer log lists exactly sedan S at 18:18, LiftVan B at 18:28, and LiftVan C at 18:34. LiftVan B is the accessible vehicle assigned to this non-transfer passenger; sedan S lacks explicit documented evidence of boarding equipment.\"},{\"speaker\":\"handoff coordinator\",\"text\":\"At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases.\"},{\"speaker\":\"fleet records officer\",\"text\":\"At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was eight suitcases.\"},{\"speaker\":\"fleet technician\",\"text\":\"LiftVan B's inspection file explicitly documents a boarding ramp and wheelchair tie-downs; its rated capacity is four travelers and one rigid wheelchair.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"Curb boarding assistance is booked for the party at 18:25, before LiftVan B's logged pickup. Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases."}, {"path": ["3", "text"], "text": "At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was eight suitcases."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases.", "negative_left": "At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases.", "negative_right": "At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was four suitcases.", "right": "At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was eight suitcases."}, "verifier_independent_model": false}, "family": "scale-diverse-111-003", "id": "scale-diverse-111-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "arriving passenger", "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35. Our party has four travelers and one rigid wheelchair."}, {"speaker": "transfer dispatcher", "text": "The final transfer log lists exactly sedan S at 18:18, LiftVan B at 18:28, and LiftVan C at 18:34. LiftVan B is the accessible vehicle assigned to this non-transfer passenger; sedan S lacks explicit documented evidence of boarding equipment."}, {"speaker": "handoff coordinator", "text": "At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases."}, {"speaker": "fleet records officer", "text": "At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was eight suitcases."}, {"speaker": "fleet technician", "text": "LiftVan B's inspection file explicitly documents a boarding ramp and wheelchair tie-downs; its rated capacity is four travelers and one rigid wheelchair."}, {"speaker": "airport assistance attendant", "text": "Curb boarding assistance is booked for the party at 18:25, before LiftVan B's logged pickup. Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same suitability requirements, accessible-dispatch rule, earliest-ETA ranking rule, LiftVan B decision target, party/chair scope, and 18:35 pickup deadline. The focus evidence consists of two complete factual sentences. The sole counterfactual change lowers LiftVan B’s documented suitcase capacity from eight to four while the party still has six suitcases; this is a coherent capacity shortfall, not a contradictory duplicate measurement. Neither constructed context contains an answer code, explicit yes/no conclusion, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"arriving passenger\",\"text\":\"I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35. Our party has four travelers and one rigid wheelchair.\"},{\"speaker\":\"transfer dispatcher\",\"text\":\"The final transfer log lists exactly sedan S at 18:18, LiftVan B at 18:28, and LiftVan C at 18:34. LiftVan B is the accessible vehicle assigned to this non-transfer passenger; sedan S lacks explicit documented evidence of boarding equipment.\"},{\"speaker\":\"handoff coordinator\",\"text\":\"At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases.\"},{\"speaker\":\"fleet records officer\",\"text\":\"At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was eight suitcases.\"},{\"speaker\":\"fleet technician\",\"text\":\"LiftVan B's inspection file explicitly documents a boarding ramp and wheelchair tie-downs; its rated capacity is four travelers and one rigid wheelchair.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"Curb boarding assistance is booked for the party at 18:25, before LiftVan B's logged pickup. Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases."}, {"path": ["3", "text"], "text": "At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was eight suitcases."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases.", "negative_left": "At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases.", "negative_right": "At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was four suitcases.", "right": "At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was eight suitcases."}, "verifier_independent_model": false}, "family": "scale-diverse-111-003", "id": "scale-diverse-111-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "arriving passenger", "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35. Our party has four travelers and one rigid wheelchair."}, {"speaker": "transfer dispatcher", "text": "The final transfer log lists exactly sedan S at 18:18, LiftVan B at 18:28, and LiftVan C at 18:34. LiftVan B is the accessible vehicle assigned to this non-transfer passenger; sedan S lacks explicit documented evidence of boarding equipment."}, {"speaker": "handoff coordinator", "text": "At the 17:42 handoff on 14 June 2026, the party assigned to LiftVan B had six suitcases."}, {"speaker": "fleet records officer", "text": "At the 17:42 handoff on 14 June 2026, LiftVan B's documented suitcase capacity was four suitcases."}, {"speaker": "fleet technician", "text": "LiftVan B's inspection file explicitly documents a boarding ramp and wheelchair tie-downs; its rated capacity is four travelers and one rigid wheelchair."}, {"speaker": "airport assistance attendant", "text": "Curb boarding assistance is booked for the party at 18:25, before LiftVan B's logged pickup. Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing dispatch and suitability policy, while the unchanged questions preserve the decision criteria and instructions. The request remains bound to selecting LiftVan B for the same party and 18:35 pickup deadline. The two focus spans are complete factual sentences. The counterfactual coherently changes LiftVan B’s luggage capacity from 9 to 6 against a manifest count of 7, without creating duplicate contradictory measurements. Neither context contains an answer, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"passenger note\",\"text\":\"The party consists of four travelers, including one person using a rigid wheelchair. I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35.\"},{\"speaker\":\"dispatch manifest\",\"text\":\"At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B.\"},{\"speaker\":\"inspection record\",\"text\":\"At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 9 suitcases.\"},{\"speaker\":\"dispatcher field note\",\"text\":\"The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C, with pickup ETAs of 18:18, 18:28, and 18:34 respectively. The accessibility register identifies LiftVan B as an accessible vehicle for this non-transfer passenger. Its file explicitly documents a ramp, wheelchair tie-downs, room for four travelers, and capacity for one rigid wheelchair. Sedan S has no explicit documented evidence of boarding equipment.\"},{\"speaker\":\"assistance booking\",\"text\":\"Curbside boarding assistance for this party is booked for 18:25, before LiftVan B's scheduled arrival.\"},{\"speaker\":\"policy\",\"text\":\"Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B."}, {"path": ["2", "text"], "text": "At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 9 suitcases."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B.", "negative_left": "At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B.", "negative_right": "At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 6 suitcases.", "right": "At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 9 suitcases."}, "verifier_independent_model": false}, "family": "scale-diverse-111-004", "id": "scale-diverse-111-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "passenger note", "text": "The party consists of four travelers, including one person using a rigid wheelchair. I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"speaker": "dispatch manifest", "text": "At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B."}, {"speaker": "inspection record", "text": "At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 9 suitcases."}, {"speaker": "dispatcher field note", "text": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C, with pickup ETAs of 18:18, 18:28, and 18:34 respectively. The accessibility register identifies LiftVan B as an accessible vehicle for this non-transfer passenger. Its file explicitly documents a ramp, wheelchair tie-downs, room for four travelers, and capacity for one rigid wheelchair. Sedan S has no explicit documented evidence of boarding equipment."}, {"speaker": "assistance booking", "text": "Curbside boarding assistance for this party is booked for 18:25, before LiftVan B's scheduled arrival."}, {"speaker": "policy", "text": "Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing dispatch and suitability policy, while the unchanged questions preserve the decision criteria and instructions. The request remains bound to selecting LiftVan B for the same party and 18:35 pickup deadline. The two focus spans are complete factual sentences. The counterfactual coherently changes LiftVan B’s luggage capacity from 9 to 6 against a manifest count of 7, without creating duplicate contradictory measurements. Neither context contains an answer, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"passenger note\",\"text\":\"The party consists of four travelers, including one person using a rigid wheelchair. I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35.\"},{\"speaker\":\"dispatch manifest\",\"text\":\"At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B.\"},{\"speaker\":\"inspection record\",\"text\":\"At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 9 suitcases.\"},{\"speaker\":\"dispatcher field note\",\"text\":\"The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C, with pickup ETAs of 18:18, 18:28, and 18:34 respectively. The accessibility register identifies LiftVan B as an accessible vehicle for this non-transfer passenger. Its file explicitly documents a ramp, wheelchair tie-downs, room for four travelers, and capacity for one rigid wheelchair. Sedan S has no explicit documented evidence of boarding equipment.\"},{\"speaker\":\"assistance booking\",\"text\":\"Curbside boarding assistance for this party is booked for 18:25, before LiftVan B's scheduled arrival.\"},{\"speaker\":\"policy\",\"text\":\"Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B."}, {"path": ["2", "text"], "text": "At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 9 suitcases."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B.", "negative_left": "At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B.", "negative_right": "At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 6 suitcases.", "right": "At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 9 suitcases."}, "verifier_independent_model": false}, "family": "scale-diverse-111-004", "id": "scale-diverse-111-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "passenger note", "text": "The party consists of four travelers, including one person using a rigid wheelchair. I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"speaker": "dispatch manifest", "text": "At 18:10 on 17 September 2026, the dispatch manifest recorded 7 suitcases for the party assigned to LiftVan B."}, {"speaker": "inspection record", "text": "At 18:10 on 17 September 2026, LiftVan B's inspected luggage compartment was rated to carry 6 suitcases."}, {"speaker": "dispatcher field note", "text": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C, with pickup ETAs of 18:18, 18:28, and 18:34 respectively. The accessibility register identifies LiftVan B as an accessible vehicle for this non-transfer passenger. Its file explicitly documents a ramp, wheelchair tie-downs, room for four travelers, and capacity for one rigid wheelchair. Sedan S has no explicit documented evidence of boarding equipment."}, {"speaker": "assistance booking", "text": "Curbside boarding assistance for this party is booked for 18:25, before LiftVan B's scheduled arrival."}, {"speaker": "policy", "text": "Policy routes non-transfer passengers to accessible dispatch. Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same accessibility-routing and earliest-pickup policy and preserve Mira, LiftVan A, the airport-to-hotel path, and the transfer-specific decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only LiftVan B’s pickup from 14:35 to 13:45, which is after the 13:30 party-ready time and creates no duplicate or contradictory measurement. Neither context states the requested yes/no answer, an answer code, a rule table, proposition identifiers, label rationale, or an output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Mira cannot transfer out of her powered wheelchair during the airport-to-hotel transfer. Her party consists of exactly one passenger traveling in an occupied wheelchair—Mira—plus exactly two companions, and they carry exactly three suitcases. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch. Dispatch specifications show that each LiftVan has a wheelchair ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. Mira’s party-ready time is 13:30 local time. For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time. For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 14:35 local time. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time."}, {"path": [], "text": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 14:35 local time."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time.", "negative_left": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time.", "negative_right": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 13:45 local time.", "right": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 14:35 local time."}, "verifier_independent_model": false}, "family": "scale-diverse-112-002", "id": "scale-diverse-112-002-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Mira cannot transfer out of her powered wheelchair during the airport-to-hotel transfer. Her party consists of exactly one passenger traveling in an occupied wheelchair—Mira—plus exactly two companions, and they carry exactly three suitcases. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch. Dispatch specifications show that each LiftVan has a wheelchair ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. Mira’s party-ready time is 13:30 local time. For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time. For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 14:35 local time. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same accessibility-routing and earliest-pickup policy and preserve Mira, LiftVan A, the airport-to-hotel path, and the transfer-specific decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only LiftVan B’s pickup from 14:35 to 13:45, which is after the 13:30 party-ready time and creates no duplicate or contradictory measurement. Neither context states the requested yes/no answer, an answer code, a rule table, proposition identifiers, label rationale, or an output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Mira cannot transfer out of her powered wheelchair during the airport-to-hotel transfer. Her party consists of exactly one passenger traveling in an occupied wheelchair—Mira—plus exactly two companions, and they carry exactly three suitcases. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch. Dispatch specifications show that each LiftVan has a wheelchair ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. Mira’s party-ready time is 13:30 local time. For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time. For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 14:35 local time. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time."}, {"path": [], "text": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 14:35 local time."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time.", "negative_left": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time.", "negative_right": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 13:45 local time.", "right": "For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 14:35 local time."}, "verifier_independent_model": false}, "family": "scale-diverse-112-002", "id": "scale-diverse-112-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Mira cannot transfer out of her powered wheelchair during the airport-to-hotel transfer. Her party consists of exactly one passenger traveling in an occupied wheelchair—Mira—plus exactly two companions, and they carry exactly three suitcases. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch. Dispatch specifications show that each LiftVan has a wheelchair ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. Mira’s party-ready time is 13:30 local time. For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan A's confirmed pickup time is 14:10 local time. For Mira's airport-to-hotel transfer on 18 September 2026, LiftVan B's confirmed pickup time is 13:45 local time. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same accessibility-dispatch and earliest-pickup policy, the same decision target (LiftVan A), passenger, airport-to-hotel route, vehicle set, and transfer date. The two evidence spans are complete factual log sentences. The counterfactual changes only LiftVan B’s pickup from 14:47 to 13:58; this remains consistent with the 13:50 readiness time and creates no duplicate or contradictory measurement. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Operational handoff for Mira’s airport-to-hotel transfer, 18 September 2026: Mira cannot transfer out of her powered wheelchair. Her party has exactly one passenger traveling in an occupied wheelchair, exactly two companions, and exactly three suitcases, and will be ready at 13:50 UTC. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch and has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. LiftVan B likewise has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026. The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:47 UTC on 18 September 2026. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026."}, {"path": [], "text": "The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:47 UTC on 18 September 2026."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026.", "negative_left": "The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026.", "negative_right": "The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 13:58 UTC on 18 September 2026.", "right": "The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:47 UTC on 18 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-112-003", "id": "scale-diverse-112-003-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Operational handoff for Mira’s airport-to-hotel transfer, 18 September 2026: Mira cannot transfer out of her powered wheelchair. Her party has exactly one passenger traveling in an occupied wheelchair, exactly two companions, and exactly three suitcases, and will be ready at 13:50 UTC. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch and has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. LiftVan B likewise has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026. The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:47 UTC on 18 September 2026. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same accessibility-dispatch and earliest-pickup policy, the same decision target (LiftVan A), passenger, airport-to-hotel route, vehicle set, and transfer date. The two evidence spans are complete factual log sentences. The counterfactual changes only LiftVan B’s pickup from 14:47 to 13:58; this remains consistent with the 13:50 readiness time and creates no duplicate or contradictory measurement. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Operational handoff for Mira’s airport-to-hotel transfer, 18 September 2026: Mira cannot transfer out of her powered wheelchair. Her party has exactly one passenger traveling in an occupied wheelchair, exactly two companions, and exactly three suitcases, and will be ready at 13:50 UTC. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch and has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. LiftVan B likewise has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026. The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:47 UTC on 18 September 2026. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026."}, {"path": [], "text": "The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:47 UTC on 18 September 2026."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026.", "negative_left": "The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026.", "negative_right": "The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 13:58 UTC on 18 September 2026.", "right": "The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:47 UTC on 18 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-112-003", "id": "scale-diverse-112-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Operational handoff for Mira’s airport-to-hotel transfer, 18 September 2026: Mira cannot transfer out of her powered wheelchair. Her party has exactly one passenger traveling in an occupied wheelchair, exactly two companions, and exactly three suitcases, and will be ready at 13:50 UTC. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch and has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. LiftVan B likewise has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. The operational handoff log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026. The operational handoff log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 13:58 UTC on 18 September 2026. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing dispatch and earliest-pickup policy, while the unchanged questions preserve the decision criteria and instructions. Mira, LiftVan A, the airport-to-hotel route, and the relevant pickup-ranking scope remain bound; the changed trip date, readiness time, and LiftVan B pickup are permissible case-observation changes. The two evidence spans are complete factual sentences. The counterfactual changes only LiftVan B's pickup from 14:29 to 14:03, which is coherent with the 14:00 readiness time and does not create duplicate or contradictory measurements. Neither context contains a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Transfer field note: Mira cannot transfer out of her powered wheelchair for the airport-to-hotel trip. Her party consists of Mira, the sole passenger traveling in an occupied wheelchair, and exactly two companions. They have exactly three suitcases and are ready at 14:00 UTC on 18 September 2026. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has neither a wheelchair ramp nor wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch; it has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. LiftVan B also has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026. The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:29 UTC on 18 September 2026. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026."}, {"path": [], "text": "The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:29 UTC on 18 September 2026."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026.", "negative_left": "The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026.", "negative_right": "The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:03 UTC on 18 September 2026.", "right": "The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:29 UTC on 18 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-112-004", "id": "scale-diverse-112-004-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Transfer field note: Mira cannot transfer out of her powered wheelchair for the airport-to-hotel trip. Her party consists of Mira, the sole passenger traveling in an occupied wheelchair, and exactly two companions. They have exactly three suitcases and are ready at 14:00 UTC on 18 September 2026. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has neither a wheelchair ramp nor wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch; it has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. LiftVan B also has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026. The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:29 UTC on 18 September 2026. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing dispatch and earliest-pickup policy, while the unchanged questions preserve the decision criteria and instructions. Mira, LiftVan A, the airport-to-hotel route, and the relevant pickup-ranking scope remain bound; the changed trip date, readiness time, and LiftVan B pickup are permissible case-observation changes. The two evidence spans are complete factual sentences. The counterfactual changes only LiftVan B's pickup from 14:29 to 14:03, which is coherent with the 14:00 readiness time and does not create duplicate or contradictory measurements. Neither context contains a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Transfer field note: Mira cannot transfer out of her powered wheelchair for the airport-to-hotel trip. Her party consists of Mira, the sole passenger traveling in an occupied wheelchair, and exactly two companions. They have exactly three suitcases and are ready at 14:00 UTC on 18 September 2026. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has neither a wheelchair ramp nor wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch; it has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. LiftVan B also has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026. The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:29 UTC on 18 September 2026. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026."}, {"path": [], "text": "The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:29 UTC on 18 September 2026."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026.", "negative_left": "The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026.", "negative_right": "The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:03 UTC on 18 September 2026.", "right": "The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:29 UTC on 18 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-112-004", "id": "scale-diverse-112-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Transfer field note: Mira cannot transfer out of her powered wheelchair for the airport-to-hotel trip. Her party consists of Mira, the sole passenger traveling in an occupied wheelchair, and exactly two companions. They have exactly three suitcases and are ready at 14:00 UTC on 18 September 2026. The only vehicles under consideration are Standard Sedan S, LiftVan A, and LiftVan B. Sedan S has neither a wheelchair ramp nor wheelchair securement. LiftVan A is offered through accessible-vehicle dispatch; it has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. LiftVan B also has a ramp, exactly one occupied-wheelchair position, exactly three additional passenger seats, and room for exactly three suitcases. The dispatch log records LiftVan A's pickup for Mira's airport-to-hotel transfer at 14:12 UTC on 18 September 2026. The dispatch log records LiftVan B's pickup for Mira's airport-to-hotel transfer at 14:03 UTC on 18 September 2026. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mira’s conditional routing policy, the assistance policy, the same trip, entities, date, and dispatch request. The two focus spans are complete factual sentences; the counterfactual changes only S4’s verified suitcase capacity from five to three, which coherently contrasts with the unchanged inventory of four suitcases and creates no duplicate or contradictory measurement. Neither context contains a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"On 17 September 2026, Northstar Airport operations reviewed Mira’s requested trip to the Lumen Hotel. The records below were entered in chronological order before a vehicle-routing decision.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel.\",\"At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as five.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\",\"The manifest lists Mira and one companion, while S4 has verified seating for four passengers.\",\"The inventory records one backpack, and S4 is verified to carry one backpack.\",\"S4’s trunk is verified to accommodate Mira’s single folded manual wheelchair.\",\"S4’s confirmed pickup time at Northstar Airport is 18 minutes.\",\"Mira neither requests nor requires help reaching, boarding, or transferring into the dispatched vehicle.\",\"Mira will transfer independently into a passenger seat and will not remain in her wheelchair during the trip.\",\"Mira requires neither a ramp nor wheelchair securement for this transfer.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "2"], "text": "At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as five."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel.", "negative_left": "At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel.", "negative_right": "At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as three.", "right": "At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as five."}, "verifier_independent_model": false}, "family": "scale-diverse-113-001", "id": "scale-diverse-113-001-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "On 17 September 2026, Northstar Airport operations reviewed Mira’s requested trip to the Lumen Hotel. The records below were entered in chronological order before a vehicle-routing decision.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel.", "At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as five.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.", "The manifest lists Mira and one companion, while S4 has verified seating for four passengers.", "The inventory records one backpack, and S4 is verified to carry one backpack.", "S4’s trunk is verified to accommodate Mira’s single folded manual wheelchair.", "S4’s confirmed pickup time at Northstar Airport is 18 minutes.", "Mira neither requests nor requires help reaching, boarding, or transferring into the dispatched vehicle.", "Mira will transfer independently into a passenger seat and will not remain in her wheelchair during the trip.", "Mira requires neither a ramp nor wheelchair securement for this transfer."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mira’s conditional routing policy, the assistance policy, the same trip, entities, date, and dispatch request. The two focus spans are complete factual sentences; the counterfactual changes only S4’s verified suitcase capacity from five to three, which coherently contrasts with the unchanged inventory of four suitcases and creates no duplicate or contradictory measurement. Neither context contains a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"On 17 September 2026, Northstar Airport operations reviewed Mira’s requested trip to the Lumen Hotel. The records below were entered in chronological order before a vehicle-routing decision.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel.\",\"At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as five.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\",\"The manifest lists Mira and one companion, while S4 has verified seating for four passengers.\",\"The inventory records one backpack, and S4 is verified to carry one backpack.\",\"S4’s trunk is verified to accommodate Mira’s single folded manual wheelchair.\",\"S4’s confirmed pickup time at Northstar Airport is 18 minutes.\",\"Mira neither requests nor requires help reaching, boarding, or transferring into the dispatched vehicle.\",\"Mira will transfer independently into a passenger seat and will not remain in her wheelchair during the trip.\",\"Mira requires neither a ramp nor wheelchair securement for this transfer.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "2"], "text": "At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as five."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel.", "negative_left": "At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel.", "negative_right": "At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as three.", "right": "At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as five."}, "verifier_independent_model": false}, "family": "scale-diverse-113-001", "id": "scale-diverse-113-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "On 17 September 2026, Northstar Airport operations reviewed Mira’s requested trip to the Lumen Hotel. The records below were entered in chronological order before a vehicle-routing decision.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "At 08:10 on 17 September 2026, the completed transfer inventory recorded that Mira's party carried four medium suitcases from Northstar Airport to the Lumen Hotel.", "At 08:15 on 17 September 2026, the completed verification record listed Sedan S4's medium-suitcase capacity for that transfer as three.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.", "The manifest lists Mira and one companion, while S4 has verified seating for four passengers.", "The inventory records one backpack, and S4 is verified to carry one backpack.", "S4’s trunk is verified to accommodate Mira’s single folded manual wheelchair.", "S4’s confirmed pickup time at Northstar Airport is 18 minutes.", "Mira neither requests nor requires help reaching, boarding, or transferring into the dispatched vehicle.", "Mira will transfer independently into a passenger seat and will not remain in her wheelchair during the trip.", "Mira requires neither a ramp nor wheelchair securement for this transfer."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mira’s governing conditional request and the airport-assistance policy, while the unchanged questions preserve all routing criteria and instructions. The traveler, route, vehicle, transfer, and relevant operational-time bindings remain consistent. The two focus spans are complete factual sentences. The counterfactual coherently changes Sedan S4’s measured suitcase capacity from nine to six while retaining the seven-suitcase count, without creating a duplicate conflicting measurement within either context. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"The dispatch desk prepared an operational handoff for Mira’s trip from Northstar Airport to the Lumen Hotel. The manifest identifies Mira and one companion, one backpack, and one folding manual wheelchair.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases.\",\"For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly nine medium suitcases.\",\"Sedan S4 is verified for four passengers, one backpack, and one folded manual wheelchair in its trunk. Its pickup time is 18 minutes.\",\"Mira confirms that she can independently reach, board, and transfer into Sedan S4. She will ride in a standard passenger seat and does not request or require boarding, transfer, or reaching assistance, a ramp, or wheelchair securement.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "1"], "text": "At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases."}, {"path": ["evidence", "2"], "text": "For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly nine medium suitcases."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases.", "negative_left": "At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases.", "negative_right": "For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly six medium suitcases.", "right": "For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly nine medium suitcases."}, "verifier_independent_model": false}, "family": "scale-diverse-113-003", "id": "scale-diverse-113-003-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "The dispatch desk prepared an operational handoff for Mira’s trip from Northstar Airport to the Lumen Hotel. The manifest identifies Mira and one companion, one backpack, and one folding manual wheelchair.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases.", "For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly nine medium suitcases.", "Sedan S4 is verified for four passengers, one backpack, and one folded manual wheelchair in its trunk. Its pickup time is 18 minutes.", "Mira confirms that she can independently reach, board, and transfer into Sedan S4. She will ride in a standard passenger seat and does not request or require boarding, transfer, or reaching assistance, a ramp, or wheelchair securement.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mira’s governing conditional request and the airport-assistance policy, while the unchanged questions preserve all routing criteria and instructions. The traveler, route, vehicle, transfer, and relevant operational-time bindings remain consistent. The two focus spans are complete factual sentences. The counterfactual coherently changes Sedan S4’s measured suitcase capacity from nine to six while retaining the seven-suitcase count, without creating a duplicate conflicting measurement within either context. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"The dispatch desk prepared an operational handoff for Mira’s trip from Northstar Airport to the Lumen Hotel. The manifest identifies Mira and one companion, one backpack, and one folding manual wheelchair.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases.\",\"For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly nine medium suitcases.\",\"Sedan S4 is verified for four passengers, one backpack, and one folded manual wheelchair in its trunk. Its pickup time is 18 minutes.\",\"Mira confirms that she can independently reach, board, and transfer into Sedan S4. She will ride in a standard passenger seat and does not request or require boarding, transfer, or reaching assistance, a ramp, or wheelchair securement.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "1"], "text": "At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases."}, {"path": ["evidence", "2"], "text": "For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly nine medium suitcases."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases.", "negative_left": "At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases.", "negative_right": "For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly six medium suitcases.", "right": "For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly nine medium suitcases."}, "verifier_independent_model": false}, "family": "scale-diverse-113-003", "id": "scale-diverse-113-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "The dispatch desk prepared an operational handoff for Mira’s trip from Northstar Airport to the Lumen Hotel. The manifest identifies Mira and one companion, one backpack, and one folding manual wheelchair.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "At the 2026-09-17 14:05 UTC operational handoff for Mira's party's transfer from Northstar Airport to the Lumen Hotel, the baggage count recorded seven medium suitcases.", "For the same transfer, a load test completed at 2026-09-17 13:40 UTC verified that Sedan S4 can carry exactly six medium suitcases.", "Sedan S4 is verified for four passengers, one backpack, and one folded manual wheelchair in its trunk. Its pickup time is 18 minutes.", "Mira confirms that she can independently reach, board, and transfer into Sedan S4. She will ride in a standard passenger seat and does not request or require boarding, transfer, or reaching assistance, a ramp, or wheelchair securement.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the routing criteria and instructions, while both contexts retain Mira’s conditional sedan policy and the airport-assistance policy from the original state. Mira, the Northstar Airport–Lumen Hotel route, transfer request, pickup condition, and relevant service requirements remain bound consistently. The two focus-evidence spans are complete factual sentences concerning the luggage tally and tested sedan capacity. The counterfactual changes only the tested suitcase capacity from eight to six; this is coherent with the unchanged seven-suitcase tally and does not conflict with the context’s separate statements about passenger, backpack, and folded-wheelchair capacity. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or instruction specifying which route to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"Dispatch field note for Mira’s Northstar Airport transfer to the Lumen Hotel on 17 September 2026. Mira is traveling with one companion. The party carries one backpack and Mira’s folding manual wheelchair. Sedan S4’s completed load test verified room for four passengers, one backpack, and the folded chair in its trunk. Its confirmed airport pickup time is 18 minutes. Mira can reach and enter S4 independently, does not request or require help reaching, boarding, or transferring into it, and will ride in a regular seat after folding the chair. She requires neither a ramp nor wheelchair securement.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel.\",\"At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry eight medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "2"], "text": "At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry eight medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry six medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry eight medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "scale-diverse-113-004", "id": "scale-diverse-113-004-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Dispatch field note for Mira’s Northstar Airport transfer to the Lumen Hotel on 17 September 2026. Mira is traveling with one companion. The party carries one backpack and Mira’s folding manual wheelchair. Sedan S4’s completed load test verified room for four passengers, one backpack, and the folded chair in its trunk. Its confirmed airport pickup time is 18 minutes. Mira can reach and enter S4 independently, does not request or require help reaching, boarding, or transferring into it, and will ride in a regular seat after folding the chair. She requires neither a ramp nor wheelchair securement.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel.", "At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry eight medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the routing criteria and instructions, while both contexts retain Mira’s conditional sedan policy and the airport-assistance policy from the original state. Mira, the Northstar Airport–Lumen Hotel route, transfer request, pickup condition, and relevant service requirements remain bound consistently. The two focus-evidence spans are complete factual sentences concerning the luggage tally and tested sedan capacity. The counterfactual changes only the tested suitcase capacity from eight to six; this is coherent with the unchanged seven-suitcase tally and does not conflict with the context’s separate statements about passenger, backpack, and folded-wheelchair capacity. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or instruction specifying which route to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"Dispatch field note for Mira’s Northstar Airport transfer to the Lumen Hotel on 17 September 2026. Mira is traveling with one companion. The party carries one backpack and Mira’s folding manual wheelchair. Sedan S4’s completed load test verified room for four passengers, one backpack, and the folded chair in its trunk. Its confirmed airport pickup time is 18 minutes. Mira can reach and enter S4 independently, does not request or require help reaching, boarding, or transferring into it, and will ride in a regular seat after folding the chair. She requires neither a ramp nor wheelchair securement.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel.\",\"At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry eight medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "2"], "text": "At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry eight medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry six medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry eight medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "scale-diverse-113-004", "id": "scale-diverse-113-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Dispatch field note for Mira’s Northstar Airport transfer to the Lumen Hotel on 17 September 2026. Mira is traveling with one companion. The party carries one backpack and Mira’s folding manual wheelchair. Sedan S4’s completed load test verified room for four passengers, one backpack, and the folded chair in its trunk. Its confirmed airport pickup time is 18 minutes. Mira can reach and enter S4 independently, does not request or require help reaching, boarding, or transferring into it, and will ride in a regular seat after folding the chair. She requires neither a ramp nor wheelchair securement.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "At 14:10 on 17 September 2026, the checked-luggage tally recorded seven medium suitcases carried by Mira's party for the transfer from Northstar Airport to the Lumen Hotel.", "At 14:12 on 17 September 2026, the completed load test verified that Sedan S4 could carry six medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing driver/attendant boarding-assistance policy, while the unchanged questions preserve all scoring criteria and instructions. The transfer, vehicle, flight, handoff time, passenger configuration, and relevant capacity path remain bound to the original request. The two evidence spans are complete factual sentences. Changing AV-7’s certified maximum load from 276 kilograms to 248 kilograms is coherent with the unchanged 253-kilogram measured load; it creates an over-capacity condition rather than a contradictory duplicate measurement. Neither context contains an answer label, score code, rule table, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the policy-related atoms, which describe concrete permission or prohibition facts rather than making the final classification. The focus a1 is a factual capacity relationship. The base and counter assignments are realizable with only a1 changing because wheelchair-bay capacity can remain sufficient while total passenger capacity fails to accommodate the whole party. Policy evidence correctly preserves the substantive driver/attendant rule originating in the original state; all other governing criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a1 establishes that AV-7 does not accommodate the required passenger configuration. Missing required passenger capacity is independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes sufficient wheelchair, passenger, and luggage capacity; documented and inspected ramp and securement features; post-arrival availability; and exactly one required boarding task. The passenger cannot perform that task, the driver is prohibited from doing it, and a permitted, able, available attendant can do it. This satisfies level 1 and excludes level 0 and level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "AV-7's verified maximum passenger load accommodates the transfer party of one occupied power-wheelchair bay plus two fixed-seat companions."}, {"id": "a2", "statement": "AV-7's verified capacity of one wheelchair bay accommodates the transfer's one occupied power wheelchair."}, {"id": "a3", "statement": "AV-7's verified capacity of four suitcase spaces accommodates the transfer's three checked suitcases."}, {"id": "a4", "statement": "AV-7's signed manifest documents a ramp."}, {"id": "a5", "statement": "AV-7's ramp passed its inspection on the day of the transfer."}, {"id": "a6", "statement": "AV-7's signed manifest documents certified four-point tie-downs."}, {"id": "a7", "statement": "AV-7's tie-downs passed their inspection on the day of the transfer."}, {"id": "a8", "statement": "AV-7 is available at 18:25 after flight QX214's 18:00 arrival."}, {"id": "a9", "statement": "The wheelchair user cannot independently move the occupied chair up AV-7's ramp."}, {"id": "a10", "statement": "Pushing the occupied wheelchair up AV-7's ramp is the only boarding-assistance task required by the wheelchair user."}, {"id": "a11", "statement": "Dispatch policy prohibits AV-7's driver from pushing the occupied wheelchair."}, {"id": "a12", "statement": "Dispatch policy permits an airport assistance attendant to push the occupied wheelchair."}, {"id": "a13", "statement": "The airport assistance attendant is able to push the occupied wheelchair aboard AV-7."}, {"id": "a14", "statement": "The airport assistance attendant can meet AV-7 for its boarding at 18:20."}], "base_state_json": "[{\"speaker\":\"verification officer\",\"text\":\"At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 276 kilograms.\"},{\"speaker\":\"handoff recorder\",\"text\":\"At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms.\"},{\"speaker\":\"transfer dispatcher\",\"text\":\"QX214 arrives at 18:00, and AV-7 is available afterward at 18:25. The signed manifest lists one wheelchair bay, four suitcase spaces, a ramp, and certified four-point tie-downs. The transfer has one occupied power wheelchair assigned to that bay and three checked suitcases.\"},{\"speaker\":\"vehicle inspector\",\"text\":\"AV-7's ramp and tie-downs both passed inspection today.\"},{\"speaker\":\"wheelchair user\",\"text\":\"I cannot independently move my occupied chair up AV-7's ramp. Pushing it up the ramp is the only boarding-assistance task I require.\"},{\"speaker\":\"dispatch policy officer\",\"text\":\"Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"I can meet AV-7 at 18:20 for its boarding and am able to push the occupied wheelchair aboard.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 276 kilograms."}, {"path": ["1", "text"], "text": "At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms."}], "policy_evidence": [{"path": ["4", "text"], "text": "Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it."}], "rules": [{"justification": "AV-7 lacks the required passenger capacity, which independently satisfies the not-suitable criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}]}, {"justification": "The available accessible vehicle has documented, operational accessibility features and sufficient wheelchair, passenger, and luggage capacities, including the exact passenger-capacity case. The passenger requires one boarding task that the driver is prohibited from performing, while a permitted and able airport attendant is available to perform it.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 276 kilograms.", "negative_left": "At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 248 kilograms.", "negative_right": "At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms.", "right": "At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms."}, "verifier_independent_model": false}, "family": "scale-diverse-114-003", "id": "scale-diverse-114-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not suitable: Reject the proposed transfer or seek another accessible vehicle because a required ramp, securement feature, passenger or luggage capacity, or necessary boarding assistance is missing or unavailable.", "1 — Conditionally ready with assistance: Route to accessible-vehicle dispatch and coordinate an airport attendant because the vehicle’s verified features and capacities meet all needs, including exact-capacity cases, but the latest update creates a boarding task the driver is prohibited from performing.", "2 — Fully ready for direct accessible dispatch: The accessible vehicle’s verified features and capacities meet all needs, and the passenger can board independently or the driver is permitted and able to provide every required boarding task, so no attendant dependency remains."], "instructions": "Using the latest information, select the readiness and routing level for this transfer. Treat exact passenger or luggage capacity as acceptable, but require documented accessibility features and compliance with the stated boarding-assistance policy.", "type": "score"}}, "state": [{"speaker": "verification officer", "text": "At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 276 kilograms."}, {"speaker": "handoff recorder", "text": "At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms."}, {"speaker": "transfer dispatcher", "text": "QX214 arrives at 18:00, and AV-7 is available afterward at 18:25. The signed manifest lists one wheelchair bay, four suitcase spaces, a ramp, and certified four-point tie-downs. The transfer has one occupied power wheelchair assigned to that bay and three checked suitcases."}, {"speaker": "vehicle inspector", "text": "AV-7's ramp and tie-downs both passed inspection today."}, {"speaker": "wheelchair user", "text": "I cannot independently move my occupied chair up AV-7's ramp. Pushing it up the ramp is the only boarding-assistance task I require."}, {"speaker": "dispatch policy officer", "text": "Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it."}, {"speaker": "airport assistance attendant", "text": "I can meet AV-7 at 18:20 for its boarding and am able to push the occupied wheelchair aboard."}]}, "method": "c2d", "provenance": {"source_id": "diverse-114", "source_is_synthetic": true, "source_sha256": "b56354508ccf5f9113502827336a33bc1795f4d0a49c1d7d10d47ee939f8fcc2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing driver/attendant boarding-assistance policy, while the unchanged questions preserve all scoring criteria and instructions. The transfer, vehicle, flight, handoff time, passenger configuration, and relevant capacity path remain bound to the original request. The two evidence spans are complete factual sentences. Changing AV-7’s certified maximum load from 276 kilograms to 248 kilograms is coherent with the unchanged 253-kilogram measured load; it creates an over-capacity condition rather than a contradictory duplicate measurement. Neither context contains an answer label, score code, rule table, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the policy-related atoms, which describe concrete permission or prohibition facts rather than making the final classification. The focus a1 is a factual capacity relationship. The base and counter assignments are realizable with only a1 changing because wheelchair-bay capacity can remain sufficient while total passenger capacity fails to accommodate the whole party. Policy evidence correctly preserves the substantive driver/attendant rule originating in the original state; all other governing criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a1 establishes that AV-7 does not accommodate the required passenger configuration. Missing required passenger capacity is independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes sufficient wheelchair, passenger, and luggage capacity; documented and inspected ramp and securement features; post-arrival availability; and exactly one required boarding task. The passenger cannot perform that task, the driver is prohibited from doing it, and a permitted, able, available attendant can do it. This satisfies level 1 and excludes level 0 and level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "AV-7's verified maximum passenger load accommodates the transfer party of one occupied power-wheelchair bay plus two fixed-seat companions."}, {"id": "a2", "statement": "AV-7's verified capacity of one wheelchair bay accommodates the transfer's one occupied power wheelchair."}, {"id": "a3", "statement": "AV-7's verified capacity of four suitcase spaces accommodates the transfer's three checked suitcases."}, {"id": "a4", "statement": "AV-7's signed manifest documents a ramp."}, {"id": "a5", "statement": "AV-7's ramp passed its inspection on the day of the transfer."}, {"id": "a6", "statement": "AV-7's signed manifest documents certified four-point tie-downs."}, {"id": "a7", "statement": "AV-7's tie-downs passed their inspection on the day of the transfer."}, {"id": "a8", "statement": "AV-7 is available at 18:25 after flight QX214's 18:00 arrival."}, {"id": "a9", "statement": "The wheelchair user cannot independently move the occupied chair up AV-7's ramp."}, {"id": "a10", "statement": "Pushing the occupied wheelchair up AV-7's ramp is the only boarding-assistance task required by the wheelchair user."}, {"id": "a11", "statement": "Dispatch policy prohibits AV-7's driver from pushing the occupied wheelchair."}, {"id": "a12", "statement": "Dispatch policy permits an airport assistance attendant to push the occupied wheelchair."}, {"id": "a13", "statement": "The airport assistance attendant is able to push the occupied wheelchair aboard AV-7."}, {"id": "a14", "statement": "The airport assistance attendant can meet AV-7 for its boarding at 18:20."}], "base_state_json": "[{\"speaker\":\"verification officer\",\"text\":\"At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 276 kilograms.\"},{\"speaker\":\"handoff recorder\",\"text\":\"At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms.\"},{\"speaker\":\"transfer dispatcher\",\"text\":\"QX214 arrives at 18:00, and AV-7 is available afterward at 18:25. The signed manifest lists one wheelchair bay, four suitcase spaces, a ramp, and certified four-point tie-downs. The transfer has one occupied power wheelchair assigned to that bay and three checked suitcases.\"},{\"speaker\":\"vehicle inspector\",\"text\":\"AV-7's ramp and tie-downs both passed inspection today.\"},{\"speaker\":\"wheelchair user\",\"text\":\"I cannot independently move my occupied chair up AV-7's ramp. Pushing it up the ramp is the only boarding-assistance task I require.\"},{\"speaker\":\"dispatch policy officer\",\"text\":\"Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"I can meet AV-7 at 18:20 for its boarding and am able to push the occupied wheelchair aboard.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 276 kilograms."}, {"path": ["1", "text"], "text": "At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms."}], "policy_evidence": [{"path": ["4", "text"], "text": "Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it."}], "rules": [{"justification": "AV-7 lacks the required passenger capacity, which independently satisfies the not-suitable criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}]}, {"justification": "The available accessible vehicle has documented, operational accessibility features and sufficient wheelchair, passenger, and luggage capacities, including the exact passenger-capacity case. The passenger requires one boarding task that the driver is prohibited from performing, while a permitted and able airport attendant is available to perform it.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 276 kilograms.", "negative_left": "At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 248 kilograms.", "negative_right": "At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms.", "right": "At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms."}, "verifier_independent_model": false}, "family": "scale-diverse-114-003", "id": "scale-diverse-114-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not suitable: Reject the proposed transfer or seek another accessible vehicle because a required ramp, securement feature, passenger or luggage capacity, or necessary boarding assistance is missing or unavailable.", "1 — Conditionally ready with assistance: Route to accessible-vehicle dispatch and coordinate an airport attendant because the vehicle’s verified features and capacities meet all needs, including exact-capacity cases, but the latest update creates a boarding task the driver is prohibited from performing.", "2 — Fully ready for direct accessible dispatch: The accessible vehicle’s verified features and capacities meet all needs, and the passenger can board independently or the driver is permitted and able to provide every required boarding task, so no attendant dependency remains."], "instructions": "Using the latest information, select the readiness and routing level for this transfer. Treat exact passenger or luggage capacity as acceptable, but require documented accessibility features and compliance with the stated boarding-assistance policy.", "type": "score"}}, "state": [{"speaker": "verification officer", "text": "At the 18:25 transfer handoff, AV-7's current verification certificate records a maximum combined passenger-and-mobility-device load of 248 kilograms."}, {"speaker": "handoff recorder", "text": "At the 18:25 transfer handoff, the measured combined mass of the occupied power wheelchair and the two fixed-seat companions is 253 kilograms."}, {"speaker": "transfer dispatcher", "text": "QX214 arrives at 18:00, and AV-7 is available afterward at 18:25. The signed manifest lists one wheelchair bay, four suitcase spaces, a ramp, and certified four-point tie-downs. The transfer has one occupied power wheelchair assigned to that bay and three checked suitcases."}, {"speaker": "vehicle inspector", "text": "AV-7's ramp and tie-downs both passed inspection today."}, {"speaker": "wheelchair user", "text": "I cannot independently move my occupied chair up AV-7's ramp. Pushing it up the ramp is the only boarding-assistance task I require."}, {"speaker": "dispatch policy officer", "text": "Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it."}, {"speaker": "airport assistance attendant", "text": "I can meet AV-7 at 18:20 for its boarding and am able to push the occupied wheelchair aboard."}]}, "method": "c2d", "provenance": {"source_id": "diverse-114", "source_is_synthetic": true, "source_sha256": "b56354508ccf5f9113502827336a33bc1795f4d0a49c1d7d10d47ee939f8fcc2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question criteria and preserve the original dispatch policy that drivers may not push occupied chairs while attendants may do so. The same flight, vehicle, transfer time, party, luggage, and boarding-assistance path remain bound to the question. The two evidence spans are complete factual sentences. The counterfactual changes only AV-7’s verified maximum load from five to four allocation units; this coherently creates a shortfall against the unchanged five-unit requirement rather than a contradictory duplicate measurement. Neither context contains an answer code, gold label, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the policy-related atoms, which describe concrete permission or prohibition facts rather than making the final classification. The focus a1 is a factual capacity relationship. The base and counter assignments are realizable with only a1 changing because wheelchair-bay capacity can remain sufficient while total passenger capacity fails to accommodate the whole party. Policy evidence correctly preserves the substantive driver/attendant rule originating in the original state; all other governing criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a1 establishes that AV-7 does not accommodate the required passenger configuration. Missing required passenger capacity is independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes sufficient wheelchair, passenger, and luggage capacity; documented and inspected ramp and securement features; post-arrival availability; and exactly one required boarding task. The passenger cannot perform that task, the driver is prohibited from doing it, and a permitted, able, available attendant can do it. This satisfies level 1 and excludes level 0 and level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "AV-7's verified maximum passenger load accommodates the transfer party of one occupied power-wheelchair bay plus two fixed-seat companions."}, {"id": "a2", "statement": "AV-7's verified capacity of one wheelchair bay accommodates the transfer's one occupied power wheelchair."}, {"id": "a3", "statement": "AV-7's verified capacity of four suitcase spaces accommodates the transfer's three checked suitcases."}, {"id": "a4", "statement": "AV-7's signed manifest documents a ramp."}, {"id": "a5", "statement": "AV-7's ramp passed its inspection on the day of the transfer."}, {"id": "a6", "statement": "AV-7's signed manifest documents certified four-point tie-downs."}, {"id": "a7", "statement": "AV-7's tie-downs passed their inspection on the day of the transfer."}, {"id": "a8", "statement": "AV-7 is available at 18:25 after flight QX214's 18:00 arrival."}, {"id": "a9", "statement": "The wheelchair user cannot independently move the occupied chair up AV-7's ramp."}, {"id": "a10", "statement": "Pushing the occupied wheelchair up AV-7's ramp is the only boarding-assistance task required by the wheelchair user."}, {"id": "a11", "statement": "Dispatch policy prohibits AV-7's driver from pushing the occupied wheelchair."}, {"id": "a12", "statement": "Dispatch policy permits an airport assistance attendant to push the occupied wheelchair."}, {"id": "a13", "statement": "The airport assistance attendant is able to push the occupied wheelchair aboard AV-7."}, {"id": "a14", "statement": "The airport assistance attendant can meet AV-7 for its boarding at 18:20."}], "base_state_json": "[{\"speaker\":\"operations coordinator\",\"text\":\"Flight QX214 arrives at 18:00, and AV-7 is available at 18:25. The party has three checked suitcases.\"},{\"speaker\":\"vehicle records clerk\",\"text\":\"AV-7's signed manifest documents a ramp, one wheelchair bay, four suitcase spaces, and certified four-point tie-downs. The ramp and tie-downs both passed inspection on the transfer date.\"},{\"speaker\":\"capacity auditor\",\"text\":\"For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is five allocation units.\"},{\"speaker\":\"load auditor\",\"text\":\"Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each.\"},{\"speaker\":\"wheelchair user\",\"text\":\"I cannot independently move my occupied chair up AV-7's ramp. Pushing it up that ramp is my only required boarding-assistance task.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"I am able to push the occupied wheelchair aboard AV-7 and can meet the vehicle for boarding at 18:20.\"},{\"speaker\":\"dispatch policy record\",\"text\":\"Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is five allocation units."}, {"path": ["3", "text"], "text": "Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each."}], "policy_evidence": [{"path": ["4", "text"], "text": "Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it."}], "rules": [{"justification": "AV-7 lacks the required passenger capacity, which independently satisfies the not-suitable criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}]}, {"justification": "The available accessible vehicle has documented, operational accessibility features and sufficient wheelchair, passenger, and luggage capacities, including the exact passenger-capacity case. The passenger requires one boarding task that the driver is prohibited from performing, while a permitted and able airport attendant is available to perform it.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is five allocation units.", "negative_left": "For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is four allocation units.", "negative_right": "Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each.", "right": "Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each."}, "verifier_independent_model": false}, "family": "scale-diverse-114-004", "id": "scale-diverse-114-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Not suitable: Reject the proposed transfer or seek another accessible vehicle because a required ramp, securement feature, passenger or luggage capacity, or necessary boarding assistance is missing or unavailable.", "1 — Conditionally ready with assistance: Route to accessible-vehicle dispatch and coordinate an airport attendant because the vehicle’s verified features and capacities meet all needs, including exact-capacity cases, but the latest update creates a boarding task the driver is prohibited from performing.", "2 — Fully ready for direct accessible dispatch: The accessible vehicle’s verified features and capacities meet all needs, and the passenger can board independently or the driver is permitted and able to provide every required boarding task, so no attendant dependency remains."], "instructions": "Using the latest information, select the readiness and routing level for this transfer. Treat exact passenger or luggage capacity as acceptable, but require documented accessibility features and compliance with the stated boarding-assistance policy.", "type": "score"}}, "state": [{"speaker": "operations coordinator", "text": "Flight QX214 arrives at 18:00, and AV-7 is available at 18:25. The party has three checked suitcases."}, {"speaker": "vehicle records clerk", "text": "AV-7's signed manifest documents a ramp, one wheelchair bay, four suitcase spaces, and certified four-point tie-downs. The ramp and tie-downs both passed inspection on the transfer date."}, {"speaker": "capacity auditor", "text": "For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is five allocation units."}, {"speaker": "load auditor", "text": "Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each."}, {"speaker": "wheelchair user", "text": "I cannot independently move my occupied chair up AV-7's ramp. Pushing it up that ramp is my only required boarding-assistance task."}, {"speaker": "airport assistance attendant", "text": "I am able to push the occupied wheelchair aboard AV-7 and can meet the vehicle for boarding at 18:20."}, {"speaker": "dispatch policy record", "text": "Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-114", "source_is_synthetic": true, "source_sha256": "b56354508ccf5f9113502827336a33bc1795f4d0a49c1d7d10d47ee939f8fcc2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question criteria and preserve the original dispatch policy that drivers may not push occupied chairs while attendants may do so. The same flight, vehicle, transfer time, party, luggage, and boarding-assistance path remain bound to the question. The two evidence spans are complete factual sentences. The counterfactual changes only AV-7’s verified maximum load from five to four allocation units; this coherently creates a shortfall against the unchanged five-unit requirement rather than a contradictory duplicate measurement. Neither context contains an answer code, gold label, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the policy-related atoms, which describe concrete permission or prohibition facts rather than making the final classification. The focus a1 is a factual capacity relationship. The base and counter assignments are realizable with only a1 changing because wheelchair-bay capacity can remain sufficient while total passenger capacity fails to accommodate the whole party. Policy evidence correctly preserves the substantive driver/attendant rule originating in the original state; all other governing criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a1 establishes that AV-7 does not accommodate the required passenger configuration. Missing required passenger capacity is independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes sufficient wheelchair, passenger, and luggage capacity; documented and inspected ramp and securement features; post-arrival availability; and exactly one required boarding task. The passenger cannot perform that task, the driver is prohibited from doing it, and a permitted, able, available attendant can do it. This satisfies level 1 and excludes level 0 and level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "AV-7's verified maximum passenger load accommodates the transfer party of one occupied power-wheelchair bay plus two fixed-seat companions."}, {"id": "a2", "statement": "AV-7's verified capacity of one wheelchair bay accommodates the transfer's one occupied power wheelchair."}, {"id": "a3", "statement": "AV-7's verified capacity of four suitcase spaces accommodates the transfer's three checked suitcases."}, {"id": "a4", "statement": "AV-7's signed manifest documents a ramp."}, {"id": "a5", "statement": "AV-7's ramp passed its inspection on the day of the transfer."}, {"id": "a6", "statement": "AV-7's signed manifest documents certified four-point tie-downs."}, {"id": "a7", "statement": "AV-7's tie-downs passed their inspection on the day of the transfer."}, {"id": "a8", "statement": "AV-7 is available at 18:25 after flight QX214's 18:00 arrival."}, {"id": "a9", "statement": "The wheelchair user cannot independently move the occupied chair up AV-7's ramp."}, {"id": "a10", "statement": "Pushing the occupied wheelchair up AV-7's ramp is the only boarding-assistance task required by the wheelchair user."}, {"id": "a11", "statement": "Dispatch policy prohibits AV-7's driver from pushing the occupied wheelchair."}, {"id": "a12", "statement": "Dispatch policy permits an airport assistance attendant to push the occupied wheelchair."}, {"id": "a13", "statement": "The airport assistance attendant is able to push the occupied wheelchair aboard AV-7."}, {"id": "a14", "statement": "The airport assistance attendant can meet AV-7 for its boarding at 18:20."}], "base_state_json": "[{\"speaker\":\"operations coordinator\",\"text\":\"Flight QX214 arrives at 18:00, and AV-7 is available at 18:25. The party has three checked suitcases.\"},{\"speaker\":\"vehicle records clerk\",\"text\":\"AV-7's signed manifest documents a ramp, one wheelchair bay, four suitcase spaces, and certified four-point tie-downs. The ramp and tie-downs both passed inspection on the transfer date.\"},{\"speaker\":\"capacity auditor\",\"text\":\"For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is five allocation units.\"},{\"speaker\":\"load auditor\",\"text\":\"Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each.\"},{\"speaker\":\"wheelchair user\",\"text\":\"I cannot independently move my occupied chair up AV-7's ramp. Pushing it up that ramp is my only required boarding-assistance task.\"},{\"speaker\":\"airport assistance attendant\",\"text\":\"I am able to push the occupied wheelchair aboard AV-7 and can meet the vehicle for boarding at 18:20.\"},{\"speaker\":\"dispatch policy record\",\"text\":\"Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is five allocation units."}, {"path": ["3", "text"], "text": "Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each."}], "policy_evidence": [{"path": ["4", "text"], "text": "Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it."}], "rules": [{"justification": "AV-7 lacks the required passenger capacity, which independently satisfies the not-suitable criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}]}, {"justification": "The available accessible vehicle has documented, operational accessibility features and sufficient wheelchair, passenger, and luggage capacities, including the exact passenger-capacity case. The passenger requires one boarding task that the driver is prohibited from performing, while a permitted and able airport attendant is available to perform it.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is five allocation units.", "negative_left": "For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is four allocation units.", "negative_right": "Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each.", "right": "Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each."}, "verifier_independent_model": false}, "family": "scale-diverse-114-004", "id": "scale-diverse-114-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not suitable: Reject the proposed transfer or seek another accessible vehicle because a required ramp, securement feature, passenger or luggage capacity, or necessary boarding assistance is missing or unavailable.", "1 — Conditionally ready with assistance: Route to accessible-vehicle dispatch and coordinate an airport attendant because the vehicle’s verified features and capacities meet all needs, including exact-capacity cases, but the latest update creates a boarding task the driver is prohibited from performing.", "2 — Fully ready for direct accessible dispatch: The accessible vehicle’s verified features and capacities meet all needs, and the passenger can board independently or the driver is permitted and able to provide every required boarding task, so no attendant dependency remains."], "instructions": "Using the latest information, select the readiness and routing level for this transfer. Treat exact passenger or luggage capacity as acceptable, but require documented accessibility features and compliance with the stated boarding-assistance policy.", "type": "score"}}, "state": [{"speaker": "operations coordinator", "text": "Flight QX214 arrives at 18:00, and AV-7 is available at 18:25. The party has three checked suitcases."}, {"speaker": "vehicle records clerk", "text": "AV-7's signed manifest documents a ramp, one wheelchair bay, four suitcase spaces, and certified four-point tie-downs. The ramp and tie-downs both passed inspection on the transfer date."}, {"speaker": "capacity auditor", "text": "For the 18:25 transfer on 17 September 2026, AV-7's verified maximum passenger load is four allocation units."}, {"speaker": "load auditor", "text": "Under the load accounting used for that transfer, its one occupied power-wheelchair bay requires three allocation units and its two fixed-seat companions require one allocation unit each."}, {"speaker": "wheelchair user", "text": "I cannot independently move my occupied chair up AV-7's ramp. Pushing it up that ramp is my only required boarding-assistance task."}, {"speaker": "airport assistance attendant", "text": "I am able to push the occupied wheelchair aboard AV-7 and can meet the vehicle for boarding at 18:20."}, {"speaker": "dispatch policy record", "text": "Dispatch policy forbids drivers from pushing occupied chairs; attendants may do it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-114", "source_is_synthetic": true, "source_sha256": "b56354508ccf5f9113502827336a33bc1795f4d0a49c1d7d10d47ee939f8fcc2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy and remain within the original single-owner CI-triage scope. The two evidence spans are complete factual sentences. The counterfactual coherently makes CMP-617 the sole cache-invalidation component; because CMP-482 is distinct and the architecture audit limits CMP-482 to either that cache component or an application source component, the unchanged assertions remain consistent. Neither context includes an answer code, explicit classification, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail. At 09:19 UTC on 12 August 2026, the component registry identified CMP-482 as the sole build cache-invalidation component. At 09:26 UTC, the change audit found that pull request PR-204 introduced the defect in the component named by IR-731. Causal analysis confirmed that this defect was the single primary cause of the compile-job failure. The architecture audit constrained the incident-record component to exactly one of two possibilities: the component registry's build cache-invalidation component or an application source component. Job records ruled out any test assertion failure, test setup failure, or flakiness. They also confirmed that no CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail."}, {"path": [], "text": "At 09:19 UTC on 12 August 2026, the component registry identified CMP-482 as the sole build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail.", "negative_left": "At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail.", "negative_right": "At 09:19 UTC on 12 August 2026, the component registry identified CMP-617, which is distinct from CMP-482, as the sole build cache-invalidation component.", "right": "At 09:19 UTC on 12 August 2026, the component registry identified CMP-482 as the sole build cache-invalidation component."}, "verifier_independent_model": false}, "family": "scale-diverse-127-001", "id": "scale-diverse-127-001-base", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail. At 09:19 UTC on 12 August 2026, the component registry identified CMP-482 as the sole build cache-invalidation component. At 09:26 UTC, the change audit found that pull request PR-204 introduced the defect in the component named by IR-731. Causal analysis confirmed that this defect was the single primary cause of the compile-job failure. The architecture audit constrained the incident-record component to exactly one of two possibilities: the component registry's build cache-invalidation component or an application source component. Job records ruled out any test assertion failure, test setup failure, or flakiness. They also confirmed that no CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy and remain within the original single-owner CI-triage scope. The two evidence spans are complete factual sentences. The counterfactual coherently makes CMP-617 the sole cache-invalidation component; because CMP-482 is distinct and the architecture audit limits CMP-482 to either that cache component or an application source component, the unchanged assertions remain consistent. Neither context includes an answer code, explicit classification, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail. At 09:19 UTC on 12 August 2026, the component registry identified CMP-482 as the sole build cache-invalidation component. At 09:26 UTC, the change audit found that pull request PR-204 introduced the defect in the component named by IR-731. Causal analysis confirmed that this defect was the single primary cause of the compile-job failure. The architecture audit constrained the incident-record component to exactly one of two possibilities: the component registry's build cache-invalidation component or an application source component. Job records ruled out any test assertion failure, test setup failure, or flakiness. They also confirmed that no CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail."}, {"path": [], "text": "At 09:19 UTC on 12 August 2026, the component registry identified CMP-482 as the sole build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail.", "negative_left": "At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail.", "negative_right": "At 09:19 UTC on 12 August 2026, the component registry identified CMP-617, which is distinct from CMP-482, as the sole build cache-invalidation component.", "right": "At 09:19 UTC on 12 August 2026, the component registry identified CMP-482 as the sole build cache-invalidation component."}, "verifier_independent_model": false}, "family": "scale-diverse-127-001", "id": "scale-diverse-127-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "At 09:14 UTC on 12 August 2026, incident record IR-731 identified component CMP-482 as the sole component containing the defect that caused compile job CJ-906 to fail. At 09:19 UTC on 12 August 2026, the component registry identified CMP-617, which is distinct from CMP-482, as the sole build cache-invalidation component. At 09:26 UTC, the change audit found that pull request PR-204 introduced the defect in the component named by IR-731. Causal analysis confirmed that this defect was the single primary cause of the compile-job failure. The architecture audit constrained the incident-record component to exactly one of two possibilities: the component registry's build cache-invalidation component or an application source component. Job records ruled out any test assertion failure, test setup failure, or flakiness. They also confirmed that no CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "application_owner"}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy and continue to concern the single primary owner of the pull request’s compile-job failure. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the registry mapping: component 741 is distinct from cache-invalidation component 925, and the unchanged two-possibility assertion therefore remains consistent without duplicate contradictory measurements. Neither context includes a gold answer, answer code, rule table, proposition identifier, output instruction, or impermissible label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"Evidence reconciliation note: The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741. The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 741 to the build cache-invalidation component. Change analysis concluded that the pull request introduced a defect in the incident record’s affected component and that this defect was the single primary cause of the compile-job failure. The affected component is exactly one of two possibilities: the component registry’s build cache-invalidation component or an application source component. Investigators excluded a test assertion failure as a cause. They also excluded test setup failure and flakiness. No CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741."}, {"path": [], "text": "The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 741 to the build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741.", "negative_left": "The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741.", "negative_right": "The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 925 to the build cache-invalidation component and records that the component bearing number 925 is distinct from the component bearing number 741.", "right": "The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 741 to the build cache-invalidation component."}, "verifier_independent_model": false}, "family": "scale-diverse-127-002", "id": "scale-diverse-127-002-base", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "Evidence reconciliation note: The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741. The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 741 to the build cache-invalidation component. Change analysis concluded that the pull request introduced a defect in the incident record’s affected component and that this defect was the single primary cause of the compile-job failure. The affected component is exactly one of two possibilities: the component registry’s build cache-invalidation component or an application source component. Investigators excluded a test assertion failure as a cause. They also excluded test setup failure and flakiness. No CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy and continue to concern the single primary owner of the pull request’s compile-job failure. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the registry mapping: component 741 is distinct from cache-invalidation component 925, and the unchanged two-possibility assertion therefore remains consistent without duplicate contradictory measurements. Neither context includes a gold answer, answer code, rule table, proposition identifier, output instruction, or impermissible label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"Evidence reconciliation note: The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741. The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 741 to the build cache-invalidation component. Change analysis concluded that the pull request introduced a defect in the incident record’s affected component and that this defect was the single primary cause of the compile-job failure. The affected component is exactly one of two possibilities: the component registry’s build cache-invalidation component or an application source component. Investigators excluded a test assertion failure as a cause. They also excluded test setup failure and flakiness. No CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741."}, {"path": [], "text": "The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 741 to the build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741.", "negative_left": "The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741.", "negative_right": "The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 925 to the build cache-invalidation component and records that the component bearing number 925 is distinct from the component bearing number 741.", "right": "The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 741 to the build cache-invalidation component."}, "verifier_independent_model": false}, "family": "scale-diverse-127-002", "id": "scale-diverse-127-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "Evidence reconciliation note: The signed incident record IR-2011, finalized at 14:32 UTC on 17 September 2026, identifies its affected component solely by the component registry's unique component number 741. The component registry snapshot taken at 14:30 UTC on 17 September 2026 assigns unique component number 925 to the build cache-invalidation component and records that the component bearing number 925 is distinct from the component bearing number 741. Change analysis concluded that the pull request introduced a defect in the incident record’s affected component and that this defect was the single primary cause of the compile-job failure. The affected component is exactly one of two possibilities: the component registry’s build cache-invalidation component or an application source component. Investigators excluded a test assertion failure as a cause. They also excluded test setup failure and flakiness. No CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "application_owner"}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, while the unchanged questions object preserves all decision instructions and criteria. The request remains about routing the single primary cause of the CI failure, and the counterfactual changes only the registry component identity while preserving the incident, component, and handoff-time bindings. The two evidence spans are complete factual sentences. In the counterfactual, CMP-4821 is distinct from the registry’s cache-invalidation component CMP-7754 and can coherently be the stipulated application source component, so there are no contradictory duplicate identities or measurements. Neither context includes an answer code, explicit output instruction, rule table, proposition identifier, or gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821. At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-4821. Investigators established that the pull request introduced a defect in the incident record’s component. That defect was the single primary cause of the compile-job failure. The handoff stipulated that the incident record’s component was exactly one of two possibilities: the component registry’s build cache-invalidation component or an application source component. No test assertion failure caused the compile-job failure. Neither a test setup failure nor flakiness caused the failure. No CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821."}, {"path": [], "text": "At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-4821."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821.", "negative_left": "At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821.", "negative_right": "At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-7754, which was distinct from component CMP-4821.", "right": "At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-4821."}, "verifier_independent_model": false}, "family": "scale-diverse-127-003", "id": "scale-diverse-127-003-base", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821. At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-4821. Investigators established that the pull request introduced a defect in the incident record’s component. That defect was the single primary cause of the compile-job failure. The handoff stipulated that the incident record’s component was exactly one of two possibilities: the component registry’s build cache-invalidation component or an application source component. No test assertion failure caused the compile-job failure. Neither a test setup failure nor flakiness caused the failure. No CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, while the unchanged questions object preserves all decision instructions and criteria. The request remains about routing the single primary cause of the CI failure, and the counterfactual changes only the registry component identity while preserving the incident, component, and handoff-time bindings. The two evidence spans are complete factual sentences. In the counterfactual, CMP-4821 is distinct from the registry’s cache-invalidation component CMP-7754 and can coherently be the stipulated application source component, so there are no contradictory duplicate identities or measurements. Neither context includes an answer code, explicit output instruction, rule table, proposition identifier, or gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821. At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-4821. Investigators established that the pull request introduced a defect in the incident record’s component. That defect was the single primary cause of the compile-job failure. The handoff stipulated that the incident record’s component was exactly one of two possibilities: the component registry’s build cache-invalidation component or an application source component. No test assertion failure caused the compile-job failure. Neither a test setup failure nor flakiness caused the failure. No CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821."}, {"path": [], "text": "At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-4821."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821.", "negative_left": "At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821.", "negative_right": "At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-7754, which was distinct from component CMP-4821.", "right": "At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-4821."}, "verifier_independent_model": false}, "family": "scale-diverse-127-003", "id": "scale-diverse-127-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "At the 2026-09-17 14:35 UTC operational handoff, incident record IR-734 identified exactly one component, component CMP-4821. At the 2026-09-17 14:35 UTC operational handoff, component registry CR-96 identified exactly one build cache-invalidation component, component CMP-7754, which was distinct from component CMP-4821. Investigators established that the pull request introduced a defect in the incident record’s component. That defect was the single primary cause of the compile-job failure. The handoff stipulated that the incident record’s component was exactly one of two possibilities: the component registry’s build cache-invalidation component or an application source component. No test assertion failure caused the compile-job failure. Neither a test setup failure nor flakiness caused the failure. No CI worker fault, network fault, or hosted-service fault caused the failure. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "application_owner"}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, while the unchanged questions preserve all choice criteria and instructions. The same compile-job failure, pull request, incident component, serial scheme, and timestamps remain bound across the contexts; only the registry serial changes. The two evidence spans are complete factual sentences. Under the stated immutable, one-to-one serial scheme, CS-731 and CS-948 identify distinct components, so the counterfactual is coherent with the ledger’s exclusive two-component classification and does not duplicate or contradict an unchanged measurement. Neither context contains an answer code, gold label, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"Field note: compile-job review. The audit established that the pull request introduced a defect in the component identified by IR-4826, and that defect was the single primary cause of the failure. The classification ledger restricts that incident component to exactly one of the component registry's build cache-invalidation component or an application source component. At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component. At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component. Job records explicitly ruled out a test assertion failure, test setup failure or flakiness, CI worker fault, network fault, and hosted-service fault. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component."}, {"path": [], "text": "At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component.", "negative_left": "At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component.", "negative_right": "At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-948 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component.", "right": "At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component."}, "verifier_independent_model": false}, "family": "scale-diverse-127-004", "id": "scale-diverse-127-004-base", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "Field note: compile-job review. The audit established that the pull request introduced a defect in the component identified by IR-4826, and that defect was the single primary cause of the failure. The classification ledger restricts that incident component to exactly one of the component registry's build cache-invalidation component or an application source component. At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component. At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component. Job records explicitly ruled out a test assertion failure, test setup failure or flakiness, CI worker fault, network fault, and hosted-service fault. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, while the unchanged questions preserve all choice criteria and instructions. The same compile-job failure, pull request, incident component, serial scheme, and timestamps remain bound across the contexts; only the registry serial changes. The two evidence spans are complete factual sentences. Under the stated immutable, one-to-one serial scheme, CS-731 and CS-948 identify distinct components, so the counterfactual is coherent with the ledger’s exclusive two-component classification and does not duplicate or contradict an unchanged measurement. Neither context contains an answer code, gold label, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"Field note: compile-job review. The audit established that the pull request introduced a defect in the component identified by IR-4826, and that defect was the single primary cause of the failure. The classification ledger restricts that incident component to exactly one of the component registry's build cache-invalidation component or an application source component. At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component. At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component. Job records explicitly ruled out a test assertion failure, test setup failure or flakiness, CI worker fault, network fault, and hosted-service fault. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component."}, {"path": [], "text": "At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component.", "negative_left": "At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component.", "negative_right": "At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-948 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component.", "right": "At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component."}, "verifier_independent_model": false}, "family": "scale-diverse-127-004", "id": "scale-diverse-127-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "Field note: compile-job review. The audit established that the pull request introduced a defect in the component identified by IR-4826, and that defect was the single primary cause of the failure. The classification ledger restricts that incident component to exactly one of the component registry's build cache-invalidation component or an application source component. At 2026-08-14T09:12:00Z, incident record IR-4826 listed its identified component under serial CS-731 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component. At 2026-08-14T09:16:00Z, the component registry listed its build cache-invalidation component under serial CS-948 in Component Serial Scheme S9, which assigns each component exactly one immutable serial and assigns each serial to at most one component. Job records explicitly ruled out a test assertion failure, test setup failure or flakiness, CI worker fault, network fault, and hosted-service fault. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "application_owner"}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the flaky-exemption threshold and leave the repository readiness rule unchanged, while preserving the decision’s bindings to change 8f31, latest required run 442, the same-environment base-branch search, and the stated timestamps. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the complete identifier set found in frozen search set S; excluding ConfigParserTest.rejectsEmptyOverride does not contradict run 442’s failure because S consists of base-branch runs, whereas run 442 is the required change run. Neither constructed context contains a gold decision, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"At 14:10 UTC on 17 September 2026, the required-run ledger identified run 442 as the latest required run for change 8f31.\"},{\"speaker\":\"Test historian\",\"text\":\"The flaky-exemption history comprised 20 base-branch runs; ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of them.\"},{\"speaker\":\"Policy note\",\"text\":\"policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Search custodian\",\"text\":\"At 14:20 UTC, reproduction-search set S was frozen. It contains every recorded base-branch run executed in the same environment as run 442, every run in S is a base-branch run, and every run in S was executed in that environment.\"},{\"speaker\":\"Build recorder\",\"text\":\"At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Evidence archivist\",\"text\":\"At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {ConfigParserTest.rejectsEmptyOverride, HeaderCodecTest.preservesQuotedValue}.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["4", "text"], "text": "At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride."}, {"path": ["5", "text"], "text": "At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {ConfigParserTest.rejectsEmptyOverride, HeaderCodecTest.preservesQuotedValue}."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride.", "negative_left": "At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride.", "negative_right": "At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {HeaderCodecTest.preservesQuotedValue, TokenReaderTest.acceptsFinalDelimiter}.", "right": "At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {ConfigParserTest.rejectsEmptyOverride, HeaderCodecTest.preservesQuotedValue}."}, "verifier_independent_model": false}, "family": "scale-diverse-129-001", "id": "scale-diverse-129-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "At 14:10 UTC on 17 September 2026, the required-run ledger identified run 442 as the latest required run for change 8f31."}, {"speaker": "Test historian", "text": "The flaky-exemption history comprised 20 base-branch runs; ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of them."}, {"speaker": "Policy note", "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Search custodian", "text": "At 14:20 UTC, reproduction-search set S was frozen. It contains every recorded base-branch run executed in the same environment as run 442, every run in S is a base-branch run, and every run in S was executed in that environment."}, {"speaker": "Build recorder", "text": "At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Evidence archivist", "text": "At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {ConfigParserTest.rejectsEmptyOverride, HeaderCodecTest.preservesQuotedValue}."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the flaky-exemption threshold and leave the repository readiness rule unchanged, while preserving the decision’s bindings to change 8f31, latest required run 442, the same-environment base-branch search, and the stated timestamps. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the complete identifier set found in frozen search set S; excluding ConfigParserTest.rejectsEmptyOverride does not contradict run 442’s failure because S consists of base-branch runs, whereas run 442 is the required change run. Neither constructed context contains a gold decision, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"At 14:10 UTC on 17 September 2026, the required-run ledger identified run 442 as the latest required run for change 8f31.\"},{\"speaker\":\"Test historian\",\"text\":\"The flaky-exemption history comprised 20 base-branch runs; ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of them.\"},{\"speaker\":\"Policy note\",\"text\":\"policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Search custodian\",\"text\":\"At 14:20 UTC, reproduction-search set S was frozen. It contains every recorded base-branch run executed in the same environment as run 442, every run in S is a base-branch run, and every run in S was executed in that environment.\"},{\"speaker\":\"Build recorder\",\"text\":\"At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Evidence archivist\",\"text\":\"At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {ConfigParserTest.rejectsEmptyOverride, HeaderCodecTest.preservesQuotedValue}.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["4", "text"], "text": "At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride."}, {"path": ["5", "text"], "text": "At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {ConfigParserTest.rejectsEmptyOverride, HeaderCodecTest.preservesQuotedValue}."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride.", "negative_left": "At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride.", "negative_right": "At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {HeaderCodecTest.preservesQuotedValue, TokenReaderTest.acceptsFinalDelimiter}.", "right": "At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {ConfigParserTest.rejectsEmptyOverride, HeaderCodecTest.preservesQuotedValue}."}, "verifier_independent_model": false}, "family": "scale-diverse-129-001", "id": "scale-diverse-129-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "At 14:10 UTC on 17 September 2026, the required-run ledger identified run 442 as the latest required run for change 8f31."}, {"speaker": "Test historian", "text": "The flaky-exemption history comprised 20 base-branch runs; ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of them."}, {"speaker": "Policy note", "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Search custodian", "text": "At 14:20 UTC, reproduction-search set S was frozen. It contains every recorded base-branch run executed in the same environment as run 442, every run in S is a base-branch run, and every run in S was executed in that environment."}, {"speaker": "Build recorder", "text": "At 14:32 UTC on 17 September 2026, run 442 completed with exactly one failure record, whose test-case identifier was ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Evidence archivist", "text": "At 14:40 UTC on 17 September 2026, the complete set of test-case identifiers in failure records across all runs in reproduction-search set S was exactly {HeaderCodecTest.preservesQuotedValue, TokenReaderTest.acceptsFinalDelimiter}."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy conditions without adding exceptions or defaults, and they preserve the bindings to change 8f31, latest required run 442, the 20-run exemption history, and same-environment base-branch reproduction. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the finalized reproduction-search result; it remains compatible with the separate statement that the test failed once among the 20 exemption-history runs because those runs are not asserted to be identical to same-environment set S. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"For change 8f31, run 442 superseded all earlier CI activity and is the latest required run.\"},{\"speaker\":\"Failure-ledger auditor\",\"text\":\"The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442.\"},{\"speaker\":\"Test-history analyst\",\"text\":\"Across the 20 base-branch runs used for exemption history, ConfigParserTest.rejectsEmptyOverride failed in exactly one run. policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Reproduction auditor\",\"text\":\"Set S was finalized from the execution archive and contains every recorded base-branch run executed in the same environment as run 442. Every member of S is a base-branch run, and every member used that same environment. The finalized membership and failure ledger for reproduction-search set S records member run 386 with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["1", "text"], "text": "The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442."}, {"path": ["3", "text"], "text": "The finalized membership and failure ledger for reproduction-search set S records member run 386 with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442.", "negative_left": "The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442.", "negative_right": "The finalized membership and failure ledger for reproduction-search set S records no member run with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "right": "The finalized membership and failure ledger for reproduction-search set S records member run 386 with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "scale-diverse-129-002", "id": "scale-diverse-129-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "For change 8f31, run 442 superseded all earlier CI activity and is the latest required run."}, {"speaker": "Failure-ledger auditor", "text": "The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442."}, {"speaker": "Test-history analyst", "text": "Across the 20 base-branch runs used for exemption history, ConfigParserTest.rejectsEmptyOverride failed in exactly one run. policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Reproduction auditor", "text": "Set S was finalized from the execution archive and contains every recorded base-branch run executed in the same environment as run 442. Every member of S is a base-branch run, and every member used that same environment. The finalized membership and failure ledger for reproduction-search set S records member run 386 with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy conditions without adding exceptions or defaults, and they preserve the bindings to change 8f31, latest required run 442, the 20-run exemption history, and same-environment base-branch reproduction. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the finalized reproduction-search result; it remains compatible with the separate statement that the test failed once among the 20 exemption-history runs because those runs are not asserted to be identical to same-environment set S. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"For change 8f31, run 442 superseded all earlier CI activity and is the latest required run.\"},{\"speaker\":\"Failure-ledger auditor\",\"text\":\"The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442.\"},{\"speaker\":\"Test-history analyst\",\"text\":\"Across the 20 base-branch runs used for exemption history, ConfigParserTest.rejectsEmptyOverride failed in exactly one run. policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Reproduction auditor\",\"text\":\"Set S was finalized from the execution archive and contains every recorded base-branch run executed in the same environment as run 442. Every member of S is a base-branch run, and every member used that same environment. The finalized membership and failure ledger for reproduction-search set S records member run 386 with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["1", "text"], "text": "The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442."}, {"path": ["3", "text"], "text": "The finalized membership and failure ledger for reproduction-search set S records member run 386 with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442.", "negative_left": "The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442.", "negative_right": "The finalized membership and failure ledger for reproduction-search set S records no member run with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "right": "The finalized membership and failure ledger for reproduction-search set S records member run 386 with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "scale-diverse-129-002", "id": "scale-diverse-129-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "For change 8f31, run 442 superseded all earlier CI activity and is the latest required run."}, {"speaker": "Failure-ledger auditor", "text": "The failure ledger finalized at 2026-09-17T14:20:00Z records ConfigParserTest.rejectsEmptyOverride as the only failed test-case identifier in run 442."}, {"speaker": "Test-history analyst", "text": "Across the 20 base-branch runs used for exemption history, ConfigParserTest.rejectsEmptyOverride failed in exactly one run. policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Reproduction auditor", "text": "Set S was finalized from the execution archive and contains every recorded base-branch run executed in the same environment as run 442. Every member of S is a base-branch run, and every member used that same environment. The finalized membership and failure ledger for reproduction-search set S records no member run with a failure whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged question preserves the governing readiness rule, while both contexts retain the stated flaky-exemption threshold and the same-environment base-branch reproduction scope. The change, run, test, snapshot, and search-set bindings remain fixed. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes the complete inventory from containing a matching reproduction to containing none, without conflicting with the unchanged search-set description. Neither context includes an answer label, output instruction, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"For change 8f31, run 442 is the latest required run as of the operational handoff.\"},{\"speaker\":\"Test history analyst\",\"text\":\"ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history.\"},{\"speaker\":\"Repository policy\",\"text\":\"policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"CI recorder\",\"text\":\"At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Search custodian\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in that same environment.\"},{\"speaker\":\"Inventory recorder\",\"text\":\"At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S lists run 731 with a failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["3", "text"], "text": "At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, {"path": ["5", "text"], "text": "At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S lists run 731 with a failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "negative_left": "At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "negative_right": "At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S contains no failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "right": "At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S lists run 731 with a failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "scale-diverse-129-003", "id": "scale-diverse-129-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "For change 8f31, run 442 is the latest required run as of the operational handoff."}, {"speaker": "Test history analyst", "text": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"speaker": "Repository policy", "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "CI recorder", "text": "At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Search custodian", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in that same environment."}, {"speaker": "Inventory recorder", "text": "At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S lists run 731 with a failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged question preserves the governing readiness rule, while both contexts retain the stated flaky-exemption threshold and the same-environment base-branch reproduction scope. The change, run, test, snapshot, and search-set bindings remain fixed. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes the complete inventory from containing a matching reproduction to containing none, without conflicting with the unchanged search-set description. Neither context includes an answer label, output instruction, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"For change 8f31, run 442 is the latest required run as of the operational handoff.\"},{\"speaker\":\"Test history analyst\",\"text\":\"ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history.\"},{\"speaker\":\"Repository policy\",\"text\":\"policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"CI recorder\",\"text\":\"At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Search custodian\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in that same environment.\"},{\"speaker\":\"Inventory recorder\",\"text\":\"At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S lists run 731 with a failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["3", "text"], "text": "At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, {"path": ["5", "text"], "text": "At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S lists run 731 with a failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "negative_left": "At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "negative_right": "At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S contains no failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "right": "At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S lists run 731 with a failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "scale-diverse-129-003", "id": "scale-diverse-129-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "For change 8f31, run 442 is the latest required run as of the operational handoff."}, {"speaker": "Test history analyst", "text": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"speaker": "Repository policy", "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "CI recorder", "text": "At the 2026-09-17T14:20:00Z operational handoff snapshot, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Search custodian", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in that same environment."}, {"speaker": "Inventory recorder", "text": "At the 2026-09-17T14:20:00Z operational handoff snapshot, the complete failure-record inventory for reproduction-search set S contains no failure record whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the repository-policy decision rule, while both contexts retain the original-state flaky-exemption threshold of at least 2 failures among the 20 history runs. Both contexts preserve change 8f31, latest required run 442, the identified test failure, same-environment base-branch reproduction scope, and the stated timestamp. The two focus-evidence spans are complete factual sentences. The counterfactual coherently replaces the recorded matching reproduction with a complete-inventory assertion that no such reproduction exists; this does not contradict the separate 20-run history, which is not asserted to be identical to same-environment set S. Neither context contains an answer, answer code, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"Run 442 is the latest required run for change 8f31.\"},{\"speaker\":\"CI recorder\",\"text\":\"As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"History auditor\",\"text\":\"ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history.\"},{\"speaker\":\"Reproduction auditor\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in that same environment.\"},{\"speaker\":\"Evidence recorder\",\"text\":\"As of 2026-09-17T14:20:00Z, run B-907 belongs to reproduction-search set S and has a failure record with test-case identifier ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Policy recorder\",\"text\":\"policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["1", "text"], "text": "As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, {"path": ["4", "text"], "text": "As of 2026-09-17T14:20:00Z, run B-907 belongs to reproduction-search set S and has a failure record with test-case identifier ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "negative_left": "As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "negative_right": "As of 2026-09-17T14:20:00Z, the complete failure-record inventory for all runs in reproduction-search set S contains no record with test-case identifier ConfigParserTest.rejectsEmptyOverride.", "right": "As of 2026-09-17T14:20:00Z, run B-907 belongs to reproduction-search set S and has a failure record with test-case identifier ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "scale-diverse-129-004", "id": "scale-diverse-129-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "Run 442 is the latest required run for change 8f31."}, {"speaker": "CI recorder", "text": "As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "History auditor", "text": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"speaker": "Reproduction auditor", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in that same environment."}, {"speaker": "Evidence recorder", "text": "As of 2026-09-17T14:20:00Z, run B-907 belongs to reproduction-search set S and has a failure record with test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Policy recorder", "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the repository-policy decision rule, while both contexts retain the original-state flaky-exemption threshold of at least 2 failures among the 20 history runs. Both contexts preserve change 8f31, latest required run 442, the identified test failure, same-environment base-branch reproduction scope, and the stated timestamp. The two focus-evidence spans are complete factual sentences. The counterfactual coherently replaces the recorded matching reproduction with a complete-inventory assertion that no such reproduction exists; this does not contradict the separate 20-run history, which is not asserted to be identical to same-environment set S. Neither context contains an answer, answer code, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"Run 442 is the latest required run for change 8f31.\"},{\"speaker\":\"CI recorder\",\"text\":\"As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"History auditor\",\"text\":\"ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history.\"},{\"speaker\":\"Reproduction auditor\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in that same environment.\"},{\"speaker\":\"Evidence recorder\",\"text\":\"As of 2026-09-17T14:20:00Z, run B-907 belongs to reproduction-search set S and has a failure record with test-case identifier ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Policy recorder\",\"text\":\"policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["1", "text"], "text": "As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, {"path": ["4", "text"], "text": "As of 2026-09-17T14:20:00Z, run B-907 belongs to reproduction-search set S and has a failure record with test-case identifier ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "negative_left": "As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride.", "negative_right": "As of 2026-09-17T14:20:00Z, the complete failure-record inventory for all runs in reproduction-search set S contains no record with test-case identifier ConfigParserTest.rejectsEmptyOverride.", "right": "As of 2026-09-17T14:20:00Z, run B-907 belongs to reproduction-search set S and has a failure record with test-case identifier ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "scale-diverse-129-004", "id": "scale-diverse-129-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "Run 442 is the latest required run for change 8f31."}, {"speaker": "CI recorder", "text": "As of 2026-09-17T14:20:00Z, run 442 has exactly one failure record, whose test-case identifier is ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "History auditor", "text": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"speaker": "Reproduction auditor", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in that same environment."}, {"speaker": "Evidence recorder", "text": "As of 2026-09-17T14:20:00Z, the complete failure-record inventory for all runs in reproduction-search set S contains no record with test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Policy recorder", "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scoring rubric, mandatory runner-health-artifact policy, merge-routing request, and claimed infrastructure attribution for the parser timeout. The two focus spans are complete factual sentences. The counterfactual coherently relocates the parser timeout to CR-9175 while leaving RH-17 scoped to CR-6842; CR-6842 may have had a different failed job, so reporting no parser timeout there does not contradict the artifact's report of an infrastructure-caused job failure. Neither context states a selected score or instructs the classifier which level to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including permitted universal claims over explicit evidence sets. The focus concerns whether RH-17 pertains to the classified CI run, not a policy conclusion. The base and counter assignments are jointly realizable while changing only that relationship: RH-17 can concern the classified run in the base and a different failed job in the counter. The policy evidence correctly preserves the substantive repository rule from the original state; the remaining rubric is automatically retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the submitted runner-health artifact pertains to the classified run and attributes its failure to infrastructure; all other mandatory evidence is present and directly supportive; and all material competing causes are excluded. This is sufficient for level 2.", "rule_index": 0, "sound": true}, {"reason": "RH-17 is the sole submitted runner-health artifact but is established not to pertain to the classified run. Therefore the required runner-health evidence for that run is missing, which mandates level 0 regardless of other suggestive evidence.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Runner-health artifact RH-17 pertains to the CI run containing the parser timeout being classified."}, {"id": "a2", "statement": "Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause."}, {"id": "a3", "statement": "Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17."}, {"id": "a4", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present."}, {"id": "a5", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause."}, {"id": "a6", "statement": "The available evidence rules out every material competing cause of the parser timeout being classified."}, {"id": "a7", "statement": "The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.\",\"1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.\",\"2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue.\"],\"instructions\":\"Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.\",\"type\":\"score\"}},\"state\":{\"context\":\"A reviewer is reconciling the sealed evidence for classification case PC-604 before deciding whether to merge a parser refactor whose timeout is claimed to have an infrastructure cause.\",\"evidence\":[\"The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job.\",\"The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-6842 at 2026-08-11T14:37:22Z, and the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded.\",\"Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient.\"],\"request\":\"Rate how completely the claimed infrastructure cause is verified for merge routing.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job."}, {"path": ["state", "evidence", "1"], "text": "The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-6842 at 2026-08-11T14:37:22Z, and the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."}], "rules": [{"justification": "RH-17 is present, pertains to the run being classified, and directly attributes that run's failure to infrastructure. All other mandatory evidence is present and directly supportive, and every material competing cause is ruled out.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "RH-17 is the sole submitted runner-health artifact, but it does not pertain to the run being classified. The mandatory runner-health artifact for the relevant infrastructure attribution is therefore missing, which controls the result.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job.", "negative_left": "The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job.", "negative_right": "The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-9175 at 2026-08-11T16:09:41Z and records that CI run CR-6842 contained no parser timeout, while the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded.", "right": "The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-6842 at 2026-08-11T14:37:22Z, and the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded."}, "verifier_independent_model": false}, "family": "scale-diverse-131-002", "id": "scale-diverse-131-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"context": "A reviewer is reconciling the sealed evidence for classification case PC-604 before deciding whether to merge a parser refactor whose timeout is claimed to have an infrastructure cause.", "evidence": ["The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job.", "The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-6842 at 2026-08-11T14:37:22Z, and the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded.", "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."], "request": "Rate how completely the claimed infrastructure cause is verified for merge routing."}}}, "method": "c2d", "provenance": {"source_id": "diverse-131", "source_is_synthetic": true, "source_sha256": "69e57342498529b071bbca7c99a798ef59cafacc702cdd17ca2f15328e12e0aa", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scoring rubric, mandatory runner-health-artifact policy, merge-routing request, and claimed infrastructure attribution for the parser timeout. The two focus spans are complete factual sentences. The counterfactual coherently relocates the parser timeout to CR-9175 while leaving RH-17 scoped to CR-6842; CR-6842 may have had a different failed job, so reporting no parser timeout there does not contradict the artifact's report of an infrastructure-caused job failure. Neither context states a selected score or instructs the classifier which level to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including permitted universal claims over explicit evidence sets. The focus concerns whether RH-17 pertains to the classified CI run, not a policy conclusion. The base and counter assignments are jointly realizable while changing only that relationship: RH-17 can concern the classified run in the base and a different failed job in the counter. The policy evidence correctly preserves the substantive repository rule from the original state; the remaining rubric is automatically retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the submitted runner-health artifact pertains to the classified run and attributes its failure to infrastructure; all other mandatory evidence is present and directly supportive; and all material competing causes are excluded. This is sufficient for level 2.", "rule_index": 0, "sound": true}, {"reason": "RH-17 is the sole submitted runner-health artifact but is established not to pertain to the classified run. Therefore the required runner-health evidence for that run is missing, which mandates level 0 regardless of other suggestive evidence.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Runner-health artifact RH-17 pertains to the CI run containing the parser timeout being classified."}, {"id": "a2", "statement": "Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause."}, {"id": "a3", "statement": "Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17."}, {"id": "a4", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present."}, {"id": "a5", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause."}, {"id": "a6", "statement": "The available evidence rules out every material competing cause of the parser timeout being classified."}, {"id": "a7", "statement": "The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.\",\"1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.\",\"2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue.\"],\"instructions\":\"Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.\",\"type\":\"score\"}},\"state\":{\"context\":\"A reviewer is reconciling the sealed evidence for classification case PC-604 before deciding whether to merge a parser refactor whose timeout is claimed to have an infrastructure cause.\",\"evidence\":[\"The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job.\",\"The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-6842 at 2026-08-11T14:37:22Z, and the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded.\",\"Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient.\"],\"request\":\"Rate how completely the claimed infrastructure cause is verified for merge routing.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job."}, {"path": ["state", "evidence", "1"], "text": "The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-6842 at 2026-08-11T14:37:22Z, and the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."}], "rules": [{"justification": "RH-17 is present, pertains to the run being classified, and directly attributes that run's failure to infrastructure. All other mandatory evidence is present and directly supportive, and every material competing cause is ruled out.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "RH-17 is the sole submitted runner-health artifact, but it does not pertain to the run being classified. The mandatory runner-health artifact for the relevant infrastructure attribution is therefore missing, which controls the result.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job.", "negative_left": "The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job.", "negative_right": "The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-9175 at 2026-08-11T16:09:41Z and records that CI run CR-6842 contained no parser timeout, while the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded.", "right": "The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-6842 at 2026-08-11T14:37:22Z, and the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded."}, "verifier_independent_model": false}, "family": "scale-diverse-131-002", "id": "scale-diverse-131-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"context": "A reviewer is reconciling the sealed evidence for classification case PC-604 before deciding whether to merge a parser refactor whose timeout is claimed to have an infrastructure cause.", "evidence": ["The sealed evidence register for classification case PC-604 records RH-17 as the sole runner-health artifact submitted for the claimed infrastructure cause, and RH-17's complete scope names CI run CR-6842 and reports that an infrastructure malfunction caused that run's failed job.", "The synchronized CI ledger records the parser timeout being classified as occurring only in CI run CR-9175 at 2026-08-11T16:09:41Z and records that CI run CR-6842 contained no parser timeout, while the case audit records that every other mandatory evidence item is present and directly supports the claimed infrastructure cause and that all material competing causes were excluded.", "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."], "request": "Rate how completely the claimed infrastructure cause is verified for merge routing."}}}, "method": "c2d", "provenance": {"source_id": "diverse-131", "source_is_synthetic": true, "source_sha256": "69e57342498529b071bbca7c99a798ef59cafacc702cdd17ca2f15328e12e0aa", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both inputs retain the questions verbatim and preserve the original repository requirement that runner-health evidence is mandatory for infrastructure attribution. The request remains bound to the claimed infrastructure cause for the CI-4821 parser timeout and merge routing. The two focus spans are complete factual sentences. In the counterfactual, RH-17 pertains solely to CI-7394 rather than CI-4821; this creates an evidentiary mismatch without contradicting the unchanged statements, since the submission register may list RH-17 for the claim even though its embedded metadata shows it concerns another run. Neither context includes a gold score, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including permitted universal claims over explicit evidence sets. The focus concerns whether RH-17 pertains to the classified CI run, not a policy conclusion. The base and counter assignments are jointly realizable while changing only that relationship: RH-17 can concern the classified run in the base and a different failed job in the counter. The policy evidence correctly preserves the substantive repository rule from the original state; the remaining rubric is automatically retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the submitted runner-health artifact pertains to the classified run and attributes its failure to infrastructure; all other mandatory evidence is present and directly supportive; and all material competing causes are excluded. This is sufficient for level 2.", "rule_index": 0, "sound": true}, {"reason": "RH-17 is the sole submitted runner-health artifact but is established not to pertain to the classified run. Therefore the required runner-health evidence for that run is missing, which mandates level 0 regardless of other suggestive evidence.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Runner-health artifact RH-17 pertains to the CI run containing the parser timeout being classified."}, {"id": "a2", "statement": "Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause."}, {"id": "a3", "statement": "Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17."}, {"id": "a4", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present."}, {"id": "a5", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause."}, {"id": "a6", "statement": "The available evidence rules out every material competing cause of the parser timeout being classified."}, {"id": "a7", "statement": "The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.\",\"1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.\",\"2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue.\"],\"instructions\":\"Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.\",\"type\":\"score\"}},\"state\":{\"context\":\"The release desk prepared an operational handoff for review of the claimed infrastructure cause.\",\"evidence\":[\"The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified.\",\"The metadata embedded in runner-health artifact RH-17 identifies CI run CI-4821 as the sole CI run to which RH-17 pertains.\",\"The evidence submission register lists exactly one runner-health artifact for the claimed infrastructure cause: RH-17.\",\"RH-17’s signed diagnostic conclusion reports that an infrastructure malfunction caused the failed job recorded in the artifact.\",\"The completed mandatory-evidence checklist confirms that every required item other than RH-17 is present and directly supports the claimed infrastructure cause.\",\"The available logs, test analysis, and environmental checks rule out every material competing cause of the parser timeout being classified.\",\"Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient.\"],\"request\":\"Rate how completely the claimed infrastructure cause is verified for merge routing.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified."}, {"path": ["state", "evidence", "1"], "text": "The metadata embedded in runner-health artifact RH-17 identifies CI run CI-4821 as the sole CI run to which RH-17 pertains."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."}], "rules": [{"justification": "RH-17 is present, pertains to the run being classified, and directly attributes that run's failure to infrastructure. All other mandatory evidence is present and directly supportive, and every material competing cause is ruled out.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "RH-17 is the sole submitted runner-health artifact, but it does not pertain to the run being classified. The mandatory runner-health artifact for the relevant infrastructure attribution is therefore missing, which controls the result.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified.", "negative_left": "The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified.", "negative_right": "The metadata embedded in runner-health artifact RH-17 identifies CI run CI-7394 as the sole CI run to which RH-17 pertains.", "right": "The metadata embedded in runner-health artifact RH-17 identifies CI run CI-4821 as the sole CI run to which RH-17 pertains."}, "verifier_independent_model": false}, "family": "scale-diverse-131-003", "id": "scale-diverse-131-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"context": "The release desk prepared an operational handoff for review of the claimed infrastructure cause.", "evidence": ["The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified.", "The metadata embedded in runner-health artifact RH-17 identifies CI run CI-4821 as the sole CI run to which RH-17 pertains.", "The evidence submission register lists exactly one runner-health artifact for the claimed infrastructure cause: RH-17.", "RH-17’s signed diagnostic conclusion reports that an infrastructure malfunction caused the failed job recorded in the artifact.", "The completed mandatory-evidence checklist confirms that every required item other than RH-17 is present and directly supports the claimed infrastructure cause.", "The available logs, test analysis, and environmental checks rule out every material competing cause of the parser timeout being classified.", "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."], "request": "Rate how completely the claimed infrastructure cause is verified for merge routing."}}}, "method": "c2d", "provenance": {"source_id": "diverse-131", "source_is_synthetic": true, "source_sha256": "69e57342498529b071bbca7c99a798ef59cafacc702cdd17ca2f15328e12e0aa", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both inputs retain the questions verbatim and preserve the original repository requirement that runner-health evidence is mandatory for infrastructure attribution. The request remains bound to the claimed infrastructure cause for the CI-4821 parser timeout and merge routing. The two focus spans are complete factual sentences. In the counterfactual, RH-17 pertains solely to CI-7394 rather than CI-4821; this creates an evidentiary mismatch without contradicting the unchanged statements, since the submission register may list RH-17 for the claim even though its embedded metadata shows it concerns another run. Neither context includes a gold score, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including permitted universal claims over explicit evidence sets. The focus concerns whether RH-17 pertains to the classified CI run, not a policy conclusion. The base and counter assignments are jointly realizable while changing only that relationship: RH-17 can concern the classified run in the base and a different failed job in the counter. The policy evidence correctly preserves the substantive repository rule from the original state; the remaining rubric is automatically retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the submitted runner-health artifact pertains to the classified run and attributes its failure to infrastructure; all other mandatory evidence is present and directly supportive; and all material competing causes are excluded. This is sufficient for level 2.", "rule_index": 0, "sound": true}, {"reason": "RH-17 is the sole submitted runner-health artifact but is established not to pertain to the classified run. Therefore the required runner-health evidence for that run is missing, which mandates level 0 regardless of other suggestive evidence.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Runner-health artifact RH-17 pertains to the CI run containing the parser timeout being classified."}, {"id": "a2", "statement": "Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause."}, {"id": "a3", "statement": "Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17."}, {"id": "a4", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present."}, {"id": "a5", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause."}, {"id": "a6", "statement": "The available evidence rules out every material competing cause of the parser timeout being classified."}, {"id": "a7", "statement": "The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.\",\"1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.\",\"2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue.\"],\"instructions\":\"Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.\",\"type\":\"score\"}},\"state\":{\"context\":\"The release desk prepared an operational handoff for review of the claimed infrastructure cause.\",\"evidence\":[\"The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified.\",\"The metadata embedded in runner-health artifact RH-17 identifies CI run CI-4821 as the sole CI run to which RH-17 pertains.\",\"The evidence submission register lists exactly one runner-health artifact for the claimed infrastructure cause: RH-17.\",\"RH-17’s signed diagnostic conclusion reports that an infrastructure malfunction caused the failed job recorded in the artifact.\",\"The completed mandatory-evidence checklist confirms that every required item other than RH-17 is present and directly supports the claimed infrastructure cause.\",\"The available logs, test analysis, and environmental checks rule out every material competing cause of the parser timeout being classified.\",\"Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient.\"],\"request\":\"Rate how completely the claimed infrastructure cause is verified for merge routing.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified."}, {"path": ["state", "evidence", "1"], "text": "The metadata embedded in runner-health artifact RH-17 identifies CI run CI-4821 as the sole CI run to which RH-17 pertains."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."}], "rules": [{"justification": "RH-17 is present, pertains to the run being classified, and directly attributes that run's failure to infrastructure. All other mandatory evidence is present and directly supportive, and every material competing cause is ruled out.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "RH-17 is the sole submitted runner-health artifact, but it does not pertain to the run being classified. The mandatory runner-health artifact for the relevant infrastructure attribution is therefore missing, which controls the result.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified.", "negative_left": "The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified.", "negative_right": "The metadata embedded in runner-health artifact RH-17 identifies CI run CI-7394 as the sole CI run to which RH-17 pertains.", "right": "The metadata embedded in runner-health artifact RH-17 identifies CI run CI-4821 as the sole CI run to which RH-17 pertains."}, "verifier_independent_model": false}, "family": "scale-diverse-131-003", "id": "scale-diverse-131-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"context": "The release desk prepared an operational handoff for review of the claimed infrastructure cause.", "evidence": ["The operational handoff record dated 2026-08-14 identifies CI run CI-4821 as the run containing the parser timeout being classified.", "The metadata embedded in runner-health artifact RH-17 identifies CI run CI-7394 as the sole CI run to which RH-17 pertains.", "The evidence submission register lists exactly one runner-health artifact for the claimed infrastructure cause: RH-17.", "RH-17’s signed diagnostic conclusion reports that an infrastructure malfunction caused the failed job recorded in the artifact.", "The completed mandatory-evidence checklist confirms that every required item other than RH-17 is present and directly supports the claimed infrastructure cause.", "The available logs, test analysis, and environmental checks rule out every material competing cause of the parser timeout being classified.", "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."], "request": "Rate how completely the claimed infrastructure cause is verified for merge routing."}}}, "method": "c2d", "provenance": {"source_id": "diverse-131", "source_is_synthetic": true, "source_sha256": "69e57342498529b071bbca7c99a798ef59cafacc702cdd17ca2f15328e12e0aa", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing merge policy verbatim and preserve the question’s PR, required-test scope, current-run timestamp, and merge-readiness decision path. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the historical main-run signature from SIG-731 to SIG-944, without creating a duplicate or contradiction, and neither context embeds a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Maintainer\",\"text\":\"PR 482 modifies only ParserCache internals. Its review record confirms that it makes no public-interface change.\"},{\"speaker\":\"Scope reviewer\",\"text\":\"The affected-component map shows that Windows cli_integration exercises ParserCache and is a required test within PR 482's affected scope.\"},{\"speaker\":\"CI recorder\",\"text\":\"PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731.\"},{\"speaker\":\"Test engineer\",\"text\":\"The run is recorded as a test failure, and its signature is categorized as non-assertion. The required-test ledger lists this run as the only current in-scope failure for PR 482.\"},{\"speaker\":\"CI historian\",\"text\":\"The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-731.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731."}, {"path": ["4", "text"], "text": "The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-731."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731.", "negative_left": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731.", "negative_right": "The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-944, which is different from SIG-731.", "right": "The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-731."}, "verifier_independent_model": false}, "family": "scale-diverse-132-001", "id": "scale-diverse-132-001-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Maintainer", "text": "PR 482 modifies only ParserCache internals. Its review record confirms that it makes no public-interface change."}, {"speaker": "Scope reviewer", "text": "The affected-component map shows that Windows cli_integration exercises ParserCache and is a required test within PR 482's affected scope."}, {"speaker": "CI recorder", "text": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731."}, {"speaker": "Test engineer", "text": "The run is recorded as a test failure, and its signature is categorized as non-assertion. The required-test ledger lists this run as the only current in-scope failure for PR 482."}, {"speaker": "CI historian", "text": "The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-731."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing merge policy verbatim and preserve the question’s PR, required-test scope, current-run timestamp, and merge-readiness decision path. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the historical main-run signature from SIG-731 to SIG-944, without creating a duplicate or contradiction, and neither context embeds a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Maintainer\",\"text\":\"PR 482 modifies only ParserCache internals. Its review record confirms that it makes no public-interface change.\"},{\"speaker\":\"Scope reviewer\",\"text\":\"The affected-component map shows that Windows cli_integration exercises ParserCache and is a required test within PR 482's affected scope.\"},{\"speaker\":\"CI recorder\",\"text\":\"PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731.\"},{\"speaker\":\"Test engineer\",\"text\":\"The run is recorded as a test failure, and its signature is categorized as non-assertion. The required-test ledger lists this run as the only current in-scope failure for PR 482.\"},{\"speaker\":\"CI historian\",\"text\":\"The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-731.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731."}, {"path": ["4", "text"], "text": "The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-731."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731.", "negative_left": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731.", "negative_right": "The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-944, which is different from SIG-731.", "right": "The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-731."}, "verifier_independent_model": false}, "family": "scale-diverse-132-001", "id": "scale-diverse-132-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Maintainer", "text": "PR 482 modifies only ParserCache internals. Its review record confirms that it makes no public-interface change."}, {"speaker": "Scope reviewer", "text": "The affected-component map shows that Windows cli_integration exercises ParserCache and is a required test within PR 482's affected scope."}, {"speaker": "CI recorder", "text": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:20:00Z with failure signature SIG-731."}, {"speaker": "Test engineer", "text": "The run is recorded as a test failure, and its signature is categorized as non-assertion. The required-test ledger lists this run as the only current in-scope failure for PR 482."}, {"speaker": "CI historian", "text": "The sole Windows cli_integration run on main completed before 2026-09-17T14:20:00Z finished at 2026-09-17T09:45:00Z with failure signature SIG-944, which is different from SIG-731."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same scope rule and known-flake policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the PR, test, current-run, main-history, and timing bindings relevant to the merge-readiness question. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the sole prior main-run signature from exit-137 to exit-143, without creating a duplicate or contradictory measurement. Neither context includes a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Run record\",\"text\":\"As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`.\"},{\"speaker\":\"Main history\",\"text\":\"The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-137|stderr-empty|phase-cleanup`.\"},{\"speaker\":\"Release coordinator\",\"text\":\"The required-test inventory marks the current Windows cli_integration run as failed. It is the only current required-test failure for PR 482; every other current required test passed.\"},{\"speaker\":\"Maintainer\",\"text\":\"The affected-test map places Windows cli_integration within PR 482's scope because the CLI exercises the changed ParserCache internals. The patch makes no public-interface change.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["0", "text"], "text": "As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`."}, {"path": ["1", "text"], "text": "The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-137|stderr-empty|phase-cleanup`."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`.", "negative_left": "As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`.", "negative_right": "The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-143|stderr-empty|phase-cleanup`.", "right": "The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-137|stderr-empty|phase-cleanup`."}, "verifier_independent_model": false}, "family": "scale-diverse-132-002", "id": "scale-diverse-132-002-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Run record", "text": "As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`."}, {"speaker": "Main history", "text": "The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-137|stderr-empty|phase-cleanup`."}, {"speaker": "Release coordinator", "text": "The required-test inventory marks the current Windows cli_integration run as failed. It is the only current required-test failure for PR 482; every other current required test passed."}, {"speaker": "Maintainer", "text": "The affected-test map places Windows cli_integration within PR 482's scope because the CLI exercises the changed ParserCache internals. The patch makes no public-interface change."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same scope rule and known-flake policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the PR, test, current-run, main-history, and timing bindings relevant to the merge-readiness question. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the sole prior main-run signature from exit-137 to exit-143, without creating a duplicate or contradictory measurement. Neither context includes a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Run record\",\"text\":\"As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`.\"},{\"speaker\":\"Main history\",\"text\":\"The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-137|stderr-empty|phase-cleanup`.\"},{\"speaker\":\"Release coordinator\",\"text\":\"The required-test inventory marks the current Windows cli_integration run as failed. It is the only current required-test failure for PR 482; every other current required test passed.\"},{\"speaker\":\"Maintainer\",\"text\":\"The affected-test map places Windows cli_integration within PR 482's scope because the CLI exercises the changed ParserCache internals. The patch makes no public-interface change.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["0", "text"], "text": "As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`."}, {"path": ["1", "text"], "text": "The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-137|stderr-empty|phase-cleanup`."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`.", "negative_left": "As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`.", "negative_right": "The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-143|stderr-empty|phase-cleanup`.", "right": "The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-137|stderr-empty|phase-cleanup`."}, "verifier_independent_model": false}, "family": "scale-diverse-132-002", "id": "scale-diverse-132-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Run record", "text": "As of 2026-09-17T15:00:00Z, PR 482's current Windows cli_integration run had completed at 2026-09-17T14:42:10Z with the non-assertion failure signature `exit-137|stderr-empty|phase-cleanup`."}, {"speaker": "Main history", "text": "The only Windows cli_integration run on main completed before 2026-09-17T14:42:10Z had completed at 2026-09-16T21:08:34Z with the failure signature `exit-143|stderr-empty|phase-cleanup`."}, {"speaker": "Release coordinator", "text": "The required-test inventory marks the current Windows cli_integration run as failed. It is the only current required-test failure for PR 482; every other current required test passed."}, {"speaker": "Maintainer", "text": "The affected-test map places Windows cli_integration within PR 482's scope because the CLI exercises the changed ParserCache internals. The patch makes no public-interface change."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original governing merge policy and retain the question’s PR 482, current merge-readiness, Windows cli_integration, scope, and known-flake path bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes one historical main-run signature so the current signature is no longer present in the complete prior ledger, without contradicting unchanged facts. Neither context embeds a gold score, answer code, rule table, proposition identifier, label rationale, or output instruction; the conditional-merging language is governing policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"Operational handoff for PR 482: the current Windows cli_integration run is recorded as a test failure. This required test is within the affected scope because the changed ParserCache internals are exercised by the CLI. The complete current required-test matrix shows that this run is the only in-scope failure.\"},{\"speaker\":\"Evidence recorder\",\"text\":\"PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \\\"7E4C-91B2\\\".\"},{\"speaker\":\"Evidence recorder\",\"text\":\"The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \\\"2A10-6D55\\\", \\\"7E4C-91B2\\\", and \\\"C803-44FE\\\", respectively.\"},{\"speaker\":\"Test engineer\",\"text\":\"The diagnostic artifact classifies the current run's failure signature as non-assertion.\"},{\"speaker\":\"Developer\",\"text\":\"The change is confined to ParserCache internals and makes no public-interface change.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \"7E4C-91B2\"."}, {"path": ["2", "text"], "text": "The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \"2A10-6D55\", \"7E4C-91B2\", and \"C803-44FE\", respectively."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \"7E4C-91B2\".", "negative_left": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \"7E4C-91B2\".", "negative_right": "The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \"2A10-6D55\", \"B671-03AC\", and \"C803-44FE\", respectively.", "right": "The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \"2A10-6D55\", \"7E4C-91B2\", and \"C803-44FE\", respectively."}, "verifier_independent_model": false}, "family": "scale-diverse-132-003", "id": "scale-diverse-132-003-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Release coordinator", "text": "Operational handoff for PR 482: the current Windows cli_integration run is recorded as a test failure. This required test is within the affected scope because the changed ParserCache internals are exercised by the CLI. The complete current required-test matrix shows that this run is the only in-scope failure."}, {"speaker": "Evidence recorder", "text": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \"7E4C-91B2\"."}, {"speaker": "Evidence recorder", "text": "The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \"2A10-6D55\", \"7E4C-91B2\", and \"C803-44FE\", respectively."}, {"speaker": "Test engineer", "text": "The diagnostic artifact classifies the current run's failure signature as non-assertion."}, {"speaker": "Developer", "text": "The change is confined to ParserCache internals and makes no public-interface change."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original governing merge policy and retain the question’s PR 482, current merge-readiness, Windows cli_integration, scope, and known-flake path bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes one historical main-run signature so the current signature is no longer present in the complete prior ledger, without contradicting unchanged facts. Neither context embeds a gold score, answer code, rule table, proposition identifier, label rationale, or output instruction; the conditional-merging language is governing policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"Operational handoff for PR 482: the current Windows cli_integration run is recorded as a test failure. This required test is within the affected scope because the changed ParserCache internals are exercised by the CLI. The complete current required-test matrix shows that this run is the only in-scope failure.\"},{\"speaker\":\"Evidence recorder\",\"text\":\"PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \\\"7E4C-91B2\\\".\"},{\"speaker\":\"Evidence recorder\",\"text\":\"The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \\\"2A10-6D55\\\", \\\"7E4C-91B2\\\", and \\\"C803-44FE\\\", respectively.\"},{\"speaker\":\"Test engineer\",\"text\":\"The diagnostic artifact classifies the current run's failure signature as non-assertion.\"},{\"speaker\":\"Developer\",\"text\":\"The change is confined to ParserCache internals and makes no public-interface change.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \"7E4C-91B2\"."}, {"path": ["2", "text"], "text": "The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \"2A10-6D55\", \"7E4C-91B2\", and \"C803-44FE\", respectively."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \"7E4C-91B2\".", "negative_left": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \"7E4C-91B2\".", "negative_right": "The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \"2A10-6D55\", \"B671-03AC\", and \"C803-44FE\", respectively.", "right": "The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \"2A10-6D55\", \"7E4C-91B2\", and \"C803-44FE\", respectively."}, "verifier_independent_model": false}, "family": "scale-diverse-132-003", "id": "scale-diverse-132-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Release coordinator", "text": "Operational handoff for PR 482: the current Windows cli_integration run is recorded as a test failure. This required test is within the affected scope because the changed ParserCache internals are exercised by the CLI. The complete current required-test matrix shows that this run is the only in-scope failure."}, {"speaker": "Evidence recorder", "text": "PR 482's current Windows cli_integration run completed at 2026-09-17T14:32:08Z with the failure-signature byte string \"7E4C-91B2\"."}, {"speaker": "Evidence recorder", "text": "The complete Windows cli_integration execution ledger for main lists exactly three runs completed before 2026-09-17T14:32:08Z, with failure-signature byte strings \"2A10-6D55\", \"B671-03AC\", and \"C803-44FE\", respectively."}, {"speaker": "Test engineer", "text": "The diagnostic artifact classifies the current run's failure signature as non-assertion."}, {"speaker": "Developer", "text": "The change is confined to ParserCache internals and makes no public-interface change."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same incident threshold, deployment condition, routing rule, suppression condition, request, region, service path, and 14:20 UTC assessment time. The two focus spans are complete factual measurement sentences. The base durations total 10 minutes 25 seconds, while the counterfactual durations total 9 minutes 25 seconds; this single change is coherent with the unchanged point-in-time latency and error measurements and creates no duplicate contradictory measurement. Neither context states an answer option, code, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "supported", "p95_level": "supported", "search_exposure": "refuted"}, "full_context_fact_states": {"base": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "supported", "p95_level": "supported", "search_exposure": "refuted"}, "counterfactual": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "refuted", "p95_level": "supported", "search_exposure": "refuted"}, "remove_left": {"p95_duration": "unknown"}, "remove_right": {"p95_duration": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"p95_duration": "unknown"}, "negative_pair": {"p95_duration": "refuted"}, "negative_sentence": {"p95_duration": "unknown"}, "positive_pair": {"p95_duration": "supported"}, "right": {"p95_duration": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal dependency health remains a single quantified relationship rather than a bundle. The focus atom is a factual duration relation, not a policy conclusion. The base and counter assignments are realizable by varying only how long p95 remained above 800 ms while keeping the other observations fixed. Policy evidence preserves the substantive state-originated threshold, rollback condition, and routing rule; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails the Catalog outcome: p95 exceeds 800 ms for at least 10 minutes, errors exceed 2%, the first slow server span is in Catalog API, and relevant dependencies are healthy. Refuted Search exposure excludes Search rollback, and healthy dependencies exclude database routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the at-least-10-minute duration means the p95 incident threshold is unmet, which is sufficient for suppression under the stated rubric. The remaining conditions also exclude Catalog confirmation, Search rollback, database routing, and the none-of-above outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "p95_level", "statement": "During the elevated-latency interval immediately preceding the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1, p95 latency exceeded 800 ms."}, {"id": "p95_duration", "statement": "The elapsed time from the observed start of the p95-above-800-ms interval to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1 was at least 10 minutes."}, {"id": "error_rate", "statement": "At the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1, the observed error rate exceeded 2%."}, {"id": "first_slow_span", "statement": "In the sampled traces for the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC, the first slow server span was in Catalog API."}, {"id": "dependency_health", "statement": "Every dependency relevant to the Catalog API traces for the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC was healthy."}, {"id": "search_exposure", "statement": "At the time of the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC, search_v2 exposure in us-east-1 exceeded 20%."}, {"id": "evidence_contradiction", "statement": "The evidence used for the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1 contained a contradiction."}], "base_state_json": "{\"context\":\"At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%.\",\"evidence\":[\"A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1.\",\"The same synchronized telemetry ledger measured 3 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1.\",\"At assessment, telemetry showed p95 at 1.4 seconds and the error rate at 3.1%.\",\"Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%.\",\"Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet.\",\"The sampled traces consistently identified Catalog API as the first slow server span.\",\"Health checks marked every dependency relevant to those Catalog API traces healthy, including Search and the shared database.\",\"The reconciled telemetry, trace, deployment, and health records were mutually consistent and contained no contradiction.\"],\"request\":\"Choose the correct incident assessment and primary routing destination.\"}", "base_states": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "supported"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}], "counter_states": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "refuted"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}], "focus_atom": "p95_duration", "focus_evidence": [{"path": ["evidence", "0"], "text": "A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1."}, {"path": ["evidence", "1"], "text": "The same synchronized telemetry ledger measured 3 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1."}], "policy_evidence": [{"path": ["context"], "text": "At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%."}, {"path": ["evidence", "3"], "text": "Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%."}, {"path": ["evidence", "4"], "text": "Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet."}, {"path": ["request"], "text": "Choose the correct incident assessment and primary routing destination."}], "rules": [{"justification": "Both runbook thresholds are satisfied, the first slow server span is in Catalog API, and every relevant dependency is healthy. Refuted above-20% search_v2 exposure excludes the Search rollback condition, while the noncontradictory evidence satisfies the explicit Catalog rubric.", "target": "catalog_confirmed_incident", "when": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "supported"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}]}, {"justification": "Although the observed p95 level exceeded 800 ms and the error rate exceeded 2%, the p95 elevation did not last the required 10 minutes. The runbook’s p95 incident threshold is therefore unmet, requiring suppression; the refuted search exposure also excludes rollback.", "target": "suppress_nonincident", "when": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "refuted"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}]}]}, "verified_pair": {"left": "A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1.", "negative_left": "A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1.", "negative_right": "The same synchronized telemetry ledger measured 2 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1.", "right": "The same synchronized telemetry ledger measured 3 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1."}, "verifier_independent_model": false}, "family": "scale-diverse-133-002", "id": "scale-diverse-133-002-base", "input": {"questions": {"decision": {"criteria": {"catalog_confirmed_incident": "Confirm a real incident and route to Catalog when both alert thresholds are met, the first slow server span is in Catalog, and dependencies are healthy.", "database_dependency": "Route to the database team only when the shared database is degraded and trace latency originates in database subspans.", "none_of_above": "Use only if the evidence is contradictory or does not satisfy any of the four explicit rubrics.", "search_rollback": "Route to Search and request rollback only when the affected region has search_v2 exposure above 20% and latency rose after its deployment.", "suppress_nonincident": "Suppress the alert only when p95 or error-rate evidence fails to satisfy the runbook’s incident threshold."}, "instructions": "Apply the stated incident threshold and conditional routing rules. Select exactly one option; deployment timing alone does not override the deployment record’s rollback condition.", "type": "choice"}}, "state": {"context": "At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%.", "evidence": ["A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1.", "The same synchronized telemetry ledger measured 3 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1.", "At assessment, telemetry showed p95 at 1.4 seconds and the error rate at 3.1%.", "Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%.", "Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet.", "The sampled traces consistently identified Catalog API as the first slow server span.", "Health checks marked every dependency relevant to those Catalog API traces healthy, including Search and the shared database.", "The reconciled telemetry, trace, deployment, and health records were mutually consistent and contained no contradiction."], "request": "Choose the correct incident assessment and primary routing destination."}}, "method": "c2d", "provenance": {"source_id": "diverse-133", "source_is_synthetic": true, "source_sha256": "2a2cee811f9eba56284805a2dcace79cda1f9b59bd311a632a7cbbd4d02dbead", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "catalog_confirmed_incident"}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same incident threshold, deployment condition, routing rule, suppression condition, request, region, service path, and 14:20 UTC assessment time. The two focus spans are complete factual measurement sentences. The base durations total 10 minutes 25 seconds, while the counterfactual durations total 9 minutes 25 seconds; this single change is coherent with the unchanged point-in-time latency and error measurements and creates no duplicate contradictory measurement. Neither context states an answer option, code, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "refuted", "p95_level": "supported", "search_exposure": "refuted"}, "full_context_fact_states": {"base": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "supported", "p95_level": "supported", "search_exposure": "refuted"}, "counterfactual": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "refuted", "p95_level": "supported", "search_exposure": "refuted"}, "remove_left": {"p95_duration": "unknown"}, "remove_right": {"p95_duration": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"p95_duration": "unknown"}, "negative_pair": {"p95_duration": "refuted"}, "negative_sentence": {"p95_duration": "unknown"}, "positive_pair": {"p95_duration": "supported"}, "right": {"p95_duration": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal dependency health remains a single quantified relationship rather than a bundle. The focus atom is a factual duration relation, not a policy conclusion. The base and counter assignments are realizable by varying only how long p95 remained above 800 ms while keeping the other observations fixed. Policy evidence preserves the substantive state-originated threshold, rollback condition, and routing rule; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails the Catalog outcome: p95 exceeds 800 ms for at least 10 minutes, errors exceed 2%, the first slow server span is in Catalog API, and relevant dependencies are healthy. Refuted Search exposure excludes Search rollback, and healthy dependencies exclude database routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the at-least-10-minute duration means the p95 incident threshold is unmet, which is sufficient for suppression under the stated rubric. The remaining conditions also exclude Catalog confirmation, Search rollback, database routing, and the none-of-above outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "p95_level", "statement": "During the elevated-latency interval immediately preceding the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1, p95 latency exceeded 800 ms."}, {"id": "p95_duration", "statement": "The elapsed time from the observed start of the p95-above-800-ms interval to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1 was at least 10 minutes."}, {"id": "error_rate", "statement": "At the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1, the observed error rate exceeded 2%."}, {"id": "first_slow_span", "statement": "In the sampled traces for the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC, the first slow server span was in Catalog API."}, {"id": "dependency_health", "statement": "Every dependency relevant to the Catalog API traces for the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC was healthy."}, {"id": "search_exposure", "statement": "At the time of the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC, search_v2 exposure in us-east-1 exceeded 20%."}, {"id": "evidence_contradiction", "statement": "The evidence used for the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1 contained a contradiction."}], "base_state_json": "{\"context\":\"At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%.\",\"evidence\":[\"A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1.\",\"The same synchronized telemetry ledger measured 3 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1.\",\"At assessment, telemetry showed p95 at 1.4 seconds and the error rate at 3.1%.\",\"Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%.\",\"Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet.\",\"The sampled traces consistently identified Catalog API as the first slow server span.\",\"Health checks marked every dependency relevant to those Catalog API traces healthy, including Search and the shared database.\",\"The reconciled telemetry, trace, deployment, and health records were mutually consistent and contained no contradiction.\"],\"request\":\"Choose the correct incident assessment and primary routing destination.\"}", "base_states": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "supported"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}], "counter_states": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "refuted"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}], "focus_atom": "p95_duration", "focus_evidence": [{"path": ["evidence", "0"], "text": "A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1."}, {"path": ["evidence", "1"], "text": "The same synchronized telemetry ledger measured 3 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1."}], "policy_evidence": [{"path": ["context"], "text": "At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%."}, {"path": ["evidence", "3"], "text": "Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%."}, {"path": ["evidence", "4"], "text": "Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet."}, {"path": ["request"], "text": "Choose the correct incident assessment and primary routing destination."}], "rules": [{"justification": "Both runbook thresholds are satisfied, the first slow server span is in Catalog API, and every relevant dependency is healthy. Refuted above-20% search_v2 exposure excludes the Search rollback condition, while the noncontradictory evidence satisfies the explicit Catalog rubric.", "target": "catalog_confirmed_incident", "when": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "supported"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}]}, {"justification": "Although the observed p95 level exceeded 800 ms and the error rate exceeded 2%, the p95 elevation did not last the required 10 minutes. The runbook’s p95 incident threshold is therefore unmet, requiring suppression; the refuted search exposure also excludes rollback.", "target": "suppress_nonincident", "when": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "refuted"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}]}]}, "verified_pair": {"left": "A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1.", "negative_left": "A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1.", "negative_right": "The same synchronized telemetry ledger measured 2 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1.", "right": "The same synchronized telemetry ledger measured 3 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1."}, "verifier_independent_model": false}, "family": "scale-diverse-133-002", "id": "scale-diverse-133-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"catalog_confirmed_incident": "Confirm a real incident and route to Catalog when both alert thresholds are met, the first slow server span is in Catalog, and dependencies are healthy.", "database_dependency": "Route to the database team only when the shared database is degraded and trace latency originates in database subspans.", "none_of_above": "Use only if the evidence is contradictory or does not satisfy any of the four explicit rubrics.", "search_rollback": "Route to Search and request rollback only when the affected region has search_v2 exposure above 20% and latency rose after its deployment.", "suppress_nonincident": "Suppress the alert only when p95 or error-rate evidence fails to satisfy the runbook’s incident threshold."}, "instructions": "Apply the stated incident threshold and conditional routing rules. Select exactly one option; deployment timing alone does not override the deployment record’s rollback condition.", "type": "choice"}}, "state": {"context": "At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%.", "evidence": ["A synchronized telemetry ledger measured 7 minutes 15 seconds from the observed start of the p95-above-800-ms interval to checkpoint Q in the ShopFront product-page alert timeline for us-east-1.", "The same synchronized telemetry ledger measured 2 minutes 10 seconds from checkpoint Q to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1.", "At assessment, telemetry showed p95 at 1.4 seconds and the error rate at 3.1%.", "Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%.", "Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet.", "The sampled traces consistently identified Catalog API as the first slow server span.", "Health checks marked every dependency relevant to those Catalog API traces healthy, including Search and the shared database.", "The reconciled telemetry, trace, deployment, and health records were mutually consistent and contained no contradiction."], "request": "Choose the correct incident assessment and primary routing destination."}}, "method": "c2d", "provenance": {"source_id": "diverse-133", "source_is_synthetic": true, "source_sha256": "2a2cee811f9eba56284805a2dcace79cda1f9b59bd311a632a7cbbd4d02dbead", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "suppress_nonincident"}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all choice criteria and instructions, while both contexts retain the original incident threshold, Search rollback condition, and routing/suppression rule. The ShopFront product-page alert, us-east-1 region, 14:20 UTC assessment, request scope, and evidence paths remain bound consistently. The two focus-evidence spans are complete factual sentences. Changing the recorded threshold interval from 12 minutes to 7 minutes is a coherent single factual alteration with no duplicate contradictory duration, and neither context embeds an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "supported", "p95_level": "supported", "search_exposure": "refuted"}, "full_context_fact_states": {"base": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "supported", "p95_level": "supported", "search_exposure": "refuted"}, "counterfactual": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "refuted", "p95_level": "supported", "search_exposure": "refuted"}, "remove_left": {"p95_duration": "unknown"}, "remove_right": {"p95_duration": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"p95_duration": "unknown"}, "negative_pair": {"p95_duration": "refuted"}, "negative_sentence": {"p95_duration": "unknown"}, "positive_pair": {"p95_duration": "supported"}, "right": {"p95_duration": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal dependency health remains a single quantified relationship rather than a bundle. The focus atom is a factual duration relation, not a policy conclusion. The base and counter assignments are realizable by varying only how long p95 remained above 800 ms while keeping the other observations fixed. Policy evidence preserves the substantive state-originated threshold, rollback condition, and routing rule; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails the Catalog outcome: p95 exceeds 800 ms for at least 10 minutes, errors exceed 2%, the first slow server span is in Catalog API, and relevant dependencies are healthy. Refuted Search exposure excludes Search rollback, and healthy dependencies exclude database routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the at-least-10-minute duration means the p95 incident threshold is unmet, which is sufficient for suppression under the stated rubric. The remaining conditions also exclude Catalog confirmation, Search rollback, database routing, and the none-of-above outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "p95_level", "statement": "During the elevated-latency interval immediately preceding the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1, p95 latency exceeded 800 ms."}, {"id": "p95_duration", "statement": "The elapsed time from the observed start of the p95-above-800-ms interval to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1 was at least 10 minutes."}, {"id": "error_rate", "statement": "At the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1, the observed error rate exceeded 2%."}, {"id": "first_slow_span", "statement": "In the sampled traces for the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC, the first slow server span was in Catalog API."}, {"id": "dependency_health", "statement": "Every dependency relevant to the Catalog API traces for the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC was healthy."}, {"id": "search_exposure", "statement": "At the time of the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC, search_v2 exposure in us-east-1 exceeded 20%."}, {"id": "evidence_contradiction", "statement": "The evidence used for the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1 contained a contradiction."}], "base_state_json": "{\"context\":\"At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%.\",\"evidence\":[\"On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 12 minutes before checkpoint C-47.\",\"Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1.\",\"At the assessment, p95 measured 1.4 seconds and the observed error rate was 3.1%.\",\"Sampled traces placed the first slow server span in Catalog API. Every dependency relevant to those traces was healthy.\",\"The assembled metric, trace, dependency, and deployment records were mutually consistent.\",\"Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%.\",\"Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet.\"],\"request\":\"Choose the correct incident assessment and primary routing destination.\"}", "base_states": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "supported"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}], "counter_states": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "refuted"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}], "focus_atom": "p95_duration", "focus_evidence": [{"path": ["evidence", "0"], "text": "On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 12 minutes before checkpoint C-47."}, {"path": ["evidence", "1"], "text": "Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1."}], "policy_evidence": [{"path": ["context"], "text": "At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%."}, {"path": ["evidence", "3"], "text": "Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%."}, {"path": ["evidence", "4"], "text": "Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet."}, {"path": ["request"], "text": "Choose the correct incident assessment and primary routing destination."}], "rules": [{"justification": "Both runbook thresholds are satisfied, the first slow server span is in Catalog API, and every relevant dependency is healthy. Refuted above-20% search_v2 exposure excludes the Search rollback condition, while the noncontradictory evidence satisfies the explicit Catalog rubric.", "target": "catalog_confirmed_incident", "when": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "supported"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}]}, {"justification": "Although the observed p95 level exceeded 800 ms and the error rate exceeded 2%, the p95 elevation did not last the required 10 minutes. The runbook’s p95 incident threshold is therefore unmet, requiring suppression; the refuted search exposure also excludes rollback.", "target": "suppress_nonincident", "when": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "refuted"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 12 minutes before checkpoint C-47.", "negative_left": "On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 7 minutes before checkpoint C-47.", "negative_right": "Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1.", "right": "Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1."}, "verifier_independent_model": false}, "family": "scale-diverse-133-005", "id": "scale-diverse-133-005-base", "input": {"questions": {"decision": {"criteria": {"catalog_confirmed_incident": "Confirm a real incident and route to Catalog when both alert thresholds are met, the first slow server span is in Catalog, and dependencies are healthy.", "database_dependency": "Route to the database team only when the shared database is degraded and trace latency originates in database subspans.", "none_of_above": "Use only if the evidence is contradictory or does not satisfy any of the four explicit rubrics.", "search_rollback": "Route to Search and request rollback only when the affected region has search_v2 exposure above 20% and latency rose after its deployment.", "suppress_nonincident": "Suppress the alert only when p95 or error-rate evidence fails to satisfy the runbook’s incident threshold."}, "instructions": "Apply the stated incident threshold and conditional routing rules. Select exactly one option; deployment timing alone does not override the deployment record’s rollback condition.", "type": "choice"}}, "state": {"context": "At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%.", "evidence": ["On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 12 minutes before checkpoint C-47.", "Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1.", "At the assessment, p95 measured 1.4 seconds and the observed error rate was 3.1%.", "Sampled traces placed the first slow server span in Catalog API. Every dependency relevant to those traces was healthy.", "The assembled metric, trace, dependency, and deployment records were mutually consistent.", "Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%.", "Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet."], "request": "Choose the correct incident assessment and primary routing destination."}}, "method": "c2d", "provenance": {"source_id": "diverse-133", "source_is_synthetic": true, "source_sha256": "2a2cee811f9eba56284805a2dcace79cda1f9b59bd311a632a7cbbd4d02dbead", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "catalog_confirmed_incident"}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all choice criteria and instructions, while both contexts retain the original incident threshold, Search rollback condition, and routing/suppression rule. The ShopFront product-page alert, us-east-1 region, 14:20 UTC assessment, request scope, and evidence paths remain bound consistently. The two focus-evidence spans are complete factual sentences. Changing the recorded threshold interval from 12 minutes to 7 minutes is a coherent single factual alteration with no duplicate contradictory duration, and neither context embeds an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "refuted", "p95_level": "supported", "search_exposure": "refuted"}, "full_context_fact_states": {"base": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "supported", "p95_level": "supported", "search_exposure": "refuted"}, "counterfactual": {"dependency_health": "supported", "error_rate": "supported", "evidence_contradiction": "refuted", "first_slow_span": "supported", "p95_duration": "refuted", "p95_level": "supported", "search_exposure": "refuted"}, "remove_left": {"p95_duration": "unknown"}, "remove_right": {"p95_duration": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"p95_duration": "unknown"}, "negative_pair": {"p95_duration": "refuted"}, "negative_sentence": {"p95_duration": "unknown"}, "positive_pair": {"p95_duration": "supported"}, "right": {"p95_duration": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal dependency health remains a single quantified relationship rather than a bundle. The focus atom is a factual duration relation, not a policy conclusion. The base and counter assignments are realizable by varying only how long p95 remained above 800 ms while keeping the other observations fixed. Policy evidence preserves the substantive state-originated threshold, rollback condition, and routing rule; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails the Catalog outcome: p95 exceeds 800 ms for at least 10 minutes, errors exceed 2%, the first slow server span is in Catalog API, and relevant dependencies are healthy. Refuted Search exposure excludes Search rollback, and healthy dependencies exclude database routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the at-least-10-minute duration means the p95 incident threshold is unmet, which is sufficient for suppression under the stated rubric. The remaining conditions also exclude Catalog confirmation, Search rollback, database routing, and the none-of-above outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "p95_level", "statement": "During the elevated-latency interval immediately preceding the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1, p95 latency exceeded 800 ms."}, {"id": "p95_duration", "statement": "The elapsed time from the observed start of the p95-above-800-ms interval to the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1 was at least 10 minutes."}, {"id": "error_rate", "statement": "At the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1, the observed error rate exceeded 2%."}, {"id": "first_slow_span", "statement": "In the sampled traces for the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC, the first slow server span was in Catalog API."}, {"id": "dependency_health", "statement": "Every dependency relevant to the Catalog API traces for the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC was healthy."}, {"id": "search_exposure", "statement": "At the time of the ShopFront product-page alert assessed in us-east-1 at 14:20 UTC, search_v2 exposure in us-east-1 exceeded 20%."}, {"id": "evidence_contradiction", "statement": "The evidence used for the 14:20 UTC assessment of the ShopFront product-page alert in us-east-1 contained a contradiction."}], "base_state_json": "{\"context\":\"At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%.\",\"evidence\":[\"On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 12 minutes before checkpoint C-47.\",\"Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1.\",\"At the assessment, p95 measured 1.4 seconds and the observed error rate was 3.1%.\",\"Sampled traces placed the first slow server span in Catalog API. Every dependency relevant to those traces was healthy.\",\"The assembled metric, trace, dependency, and deployment records were mutually consistent.\",\"Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%.\",\"Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet.\"],\"request\":\"Choose the correct incident assessment and primary routing destination.\"}", "base_states": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "supported"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}], "counter_states": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "refuted"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}], "focus_atom": "p95_duration", "focus_evidence": [{"path": ["evidence", "0"], "text": "On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 12 minutes before checkpoint C-47."}, {"path": ["evidence", "1"], "text": "Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1."}], "policy_evidence": [{"path": ["context"], "text": "At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%."}, {"path": ["evidence", "3"], "text": "Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%."}, {"path": ["evidence", "4"], "text": "Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet."}, {"path": ["request"], "text": "Choose the correct incident assessment and primary routing destination."}], "rules": [{"justification": "Both runbook thresholds are satisfied, the first slow server span is in Catalog API, and every relevant dependency is healthy. Refuted above-20% search_v2 exposure excludes the Search rollback condition, while the noncontradictory evidence satisfies the explicit Catalog rubric.", "target": "catalog_confirmed_incident", "when": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "supported"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}]}, {"justification": "Although the observed p95 level exceeded 800 ms and the error rate exceeded 2%, the p95 elevation did not last the required 10 minutes. The runbook’s p95 incident threshold is therefore unmet, requiring suppression; the refuted search exposure also excludes rollback.", "target": "suppress_nonincident", "when": [{"atom_id": "p95_level", "state": "supported"}, {"atom_id": "p95_duration", "state": "refuted"}, {"atom_id": "error_rate", "state": "supported"}, {"atom_id": "first_slow_span", "state": "supported"}, {"atom_id": "dependency_health", "state": "supported"}, {"atom_id": "search_exposure", "state": "refuted"}, {"atom_id": "evidence_contradiction", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 12 minutes before checkpoint C-47.", "negative_left": "On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 7 minutes before checkpoint C-47.", "negative_right": "Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1.", "right": "Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1."}, "verifier_independent_model": false}, "family": "scale-diverse-133-005", "id": "scale-diverse-133-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"catalog_confirmed_incident": "Confirm a real incident and route to Catalog when both alert thresholds are met, the first slow server span is in Catalog, and dependencies are healthy.", "database_dependency": "Route to the database team only when the shared database is degraded and trace latency originates in database subspans.", "none_of_above": "Use only if the evidence is contradictory or does not satisfy any of the four explicit rubrics.", "search_rollback": "Route to Search and request rollback only when the affected region has search_v2 exposure above 20% and latency rose after its deployment.", "suppress_nonincident": "Suppress the alert only when p95 or error-rate evidence fails to satisfy the runbook’s incident threshold."}, "instructions": "Apply the stated incident threshold and conditional routing rules. Select exactly one option; deployment timing alone does not override the deployment record’s rollback condition.", "type": "choice"}}, "state": {"context": "At 14:20 UTC, an on-call engineer assesses a ShopFront product-page alert in us-east-1. The runbook confirms an incident when p95 exceeds 800 ms for 10 minutes and errors exceed 2%.", "evidence": ["On 17 September 2026, the monitoring system recorded the observed start of the p95-above-800-ms interval for the ShopFront product-page alert in us-east-1 as exactly 7 minutes before checkpoint C-47.", "Checkpoint C-47 occurred at 14:20 UTC on 17 September 2026 and coincided exactly with the assessment of the ShopFront product-page alert in us-east-1.", "At the assessment, p95 measured 1.4 seconds and the observed error rate was 3.1%.", "Sampled traces placed the first slow server span in Catalog API. Every dependency relevant to those traces was healthy.", "The assembled metric, trace, dependency, and deployment records were mutually consistent.", "Search deployed at 14:05. Its record says to roll back only if latency rises where search_v2 exposure exceeds 20%; exposure in us-east-1 is 0%.", "Routing rule: route to the service containing the first slow server span unless its dependency is degraded; suppress only when incident thresholds are unmet."], "request": "Choose the correct incident assessment and primary routing destination."}}, "method": "c2d", "provenance": {"source_id": "diverse-133", "source_is_synthetic": true, "source_sha256": "2a2cee811f9eba56284805a2dcace79cda1f9b59bd311a632a7cbbd4d02dbead", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "suppress_nonincident"}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the checkout-api alert-assessment scope, the 14:12 review time, user-visible HTTP 5xx measurement path, complete-bucket rule, and higher-band boundary policy without adding exceptions or defaults. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes the bucket timestamps while preserving eleven uniquely identified observations; the gap between 14:05 and 14:07 makes the series nonconsecutive but does not contradict any unchanged assertion. Neither context states a classification, answer code, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal bucket-rate atoms remain atomic. The focus a1 is factual rather than a policy conclusion. Both assignments are realizable with only a1 changing: one context can supply at least 10 qualifying consecutive buckets, while another can explicitly supply fewer than 10, with all supplied rates still between 5% and 20% and the latest rate at least 1%. The retained question already preserves the main rubric, and policy_evidence appropriately preserves the state-origin runbook rule concerning complete consecutive buckets and boundary handling.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 guarantees at least 10 consecutive complete buckets, while a2 and a3 place every supplied complete bucket in the inclusive-5%-to-exclusive-20% band. This is sufficient for SEV-2 and excludes SEV-1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that fewer than 10 consecutive complete buckets are supplied, and a4 excludes Noise by establishing a latest rate of at least 1%. Under the exhaustive rubric, these conditions are sufficient for Observe.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 14:12, the supplied evidence contains a sequence of at least 10 consecutive complete one-minute buckets for checkout-api."}, {"id": "a2", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5%."}, {"id": "a3", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate below 20%."}, {"id": "a4", "statement": "At 14:12, the latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}], "base_state_json": "[{\"speaker\":\"Monitoring analyst\",\"text\":\"At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present.\"},{\"speaker\":\"Incident reviewer\",\"text\":\"Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:06, 14:07, 14:08, 14:09, 14:10, and 14:11, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%.\"},{\"speaker\":\"On-call engineer\",\"text\":\"At the review time, C11 was the latest available user-visible result for checkout-api. Traffic remained normal, and the alert was ready for assessment.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present."}, {"path": ["1", "text"], "text": "Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:06, 14:07, 14:08, 14:09, 14:10, and 14:11, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}], "rules": [{"justification": "At least 10 consecutive complete buckets exist, and every supplied complete bucket is in the inclusive-5%-to-exclusive-20% band required for SEV-2.", "target": "sev2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "The latest rate is at least 1%, while the evidence does not contain 10 consecutive complete buckets, which is sufficient for Observe.", "target": "observe", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present.", "negative_left": "At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present.", "negative_right": "Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:07, 14:08, 14:09, 14:10, 14:11, and 14:12, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%.", "right": "Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:06, 14:07, 14:08, 14:09, 14:10, and 14:11, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%."}, "verifier_independent_model": false}, "family": "scale-diverse-134-001", "id": "scale-diverse-134-001-base", "input": {"questions": {"decision": {"criteria": {"noise": "Classify as Noise because the latest user-visible error rate is below 1%.", "none_of_above": "Choose None of the above because the supplied evidence cannot be classified under the stated rubric.", "observe": "Classify as Observe because at least 1% errors are present but fewer than 10 consecutive complete buckets exist.", "sev1": "Classify as SEV-1 because at least 10 consecutive complete buckets are at least 20%.", "sev2": "Classify as SEV-2 because at least 10 consecutive complete buckets are at least 5% but below 20%.", "sev3": "Classify as SEV-3 because at least 10 consecutive complete buckets are at least 1% but below 5%."}, "instructions": "Rate the alert under this exhaustive rubric: Noise if the latest rate is below 1%; Observe if it is at least 1% but fewer than 10 consecutive complete buckets exist; SEV-3 if 10 or more consecutive buckets are at least 1% but below 5%; SEV-2 if 10 or more are at least 5% but below 20%; SEV-1 if 10 or more are at least 20%; choose None of the above only if the supplied evidence cannot fit any rubric.", "type": "choice"}}, "state": [{"speaker": "Monitoring analyst", "text": "At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present."}, {"speaker": "Incident reviewer", "text": "Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:06, 14:07, 14:08, 14:09, 14:10, and 14:11, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%."}, {"speaker": "On-call engineer", "text": "At the review time, C11 was the latest available user-visible result for checkout-api. Traffic remained normal, and the alert was ready for assessment."}, {"speaker": "Incident coordinator", "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}]}, "method": "c2d", "provenance": {"source_id": "diverse-134", "source_is_synthetic": true, "source_sha256": "ad1d4e1484cdb6a3a2e5104e486414ba4c97d04018b99270162dffeb80a0d8f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sev2"}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the checkout-api alert-assessment scope, the 14:12 review time, user-visible HTTP 5xx measurement path, complete-bucket rule, and higher-band boundary policy without adding exceptions or defaults. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes the bucket timestamps while preserving eleven uniquely identified observations; the gap between 14:05 and 14:07 makes the series nonconsecutive but does not contradict any unchanged assertion. Neither context states a classification, answer code, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal bucket-rate atoms remain atomic. The focus a1 is factual rather than a policy conclusion. Both assignments are realizable with only a1 changing: one context can supply at least 10 qualifying consecutive buckets, while another can explicitly supply fewer than 10, with all supplied rates still between 5% and 20% and the latest rate at least 1%. The retained question already preserves the main rubric, and policy_evidence appropriately preserves the state-origin runbook rule concerning complete consecutive buckets and boundary handling.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 guarantees at least 10 consecutive complete buckets, while a2 and a3 place every supplied complete bucket in the inclusive-5%-to-exclusive-20% band. This is sufficient for SEV-2 and excludes SEV-1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that fewer than 10 consecutive complete buckets are supplied, and a4 excludes Noise by establishing a latest rate of at least 1%. Under the exhaustive rubric, these conditions are sufficient for Observe.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 14:12, the supplied evidence contains a sequence of at least 10 consecutive complete one-minute buckets for checkout-api."}, {"id": "a2", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5%."}, {"id": "a3", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate below 20%."}, {"id": "a4", "statement": "At 14:12, the latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}], "base_state_json": "[{\"speaker\":\"Monitoring analyst\",\"text\":\"At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present.\"},{\"speaker\":\"Incident reviewer\",\"text\":\"Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:06, 14:07, 14:08, 14:09, 14:10, and 14:11, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%.\"},{\"speaker\":\"On-call engineer\",\"text\":\"At the review time, C11 was the latest available user-visible result for checkout-api. Traffic remained normal, and the alert was ready for assessment.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present."}, {"path": ["1", "text"], "text": "Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:06, 14:07, 14:08, 14:09, 14:10, and 14:11, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}], "rules": [{"justification": "At least 10 consecutive complete buckets exist, and every supplied complete bucket is in the inclusive-5%-to-exclusive-20% band required for SEV-2.", "target": "sev2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "The latest rate is at least 1%, while the evidence does not contain 10 consecutive complete buckets, which is sufficient for Observe.", "target": "observe", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present.", "negative_left": "At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present.", "negative_right": "Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:07, 14:08, 14:09, 14:10, 14:11, and 14:12, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%.", "right": "Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:06, 14:07, 14:08, 14:09, 14:10, and 14:11, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%."}, "verifier_independent_model": false}, "family": "scale-diverse-134-001", "id": "scale-diverse-134-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"noise": "Classify as Noise because the latest user-visible error rate is below 1%.", "none_of_above": "Choose None of the above because the supplied evidence cannot be classified under the stated rubric.", "observe": "Classify as Observe because at least 1% errors are present but fewer than 10 consecutive complete buckets exist.", "sev1": "Classify as SEV-1 because at least 10 consecutive complete buckets are at least 20%.", "sev2": "Classify as SEV-2 because at least 10 consecutive complete buckets are at least 5% but below 20%.", "sev3": "Classify as SEV-3 because at least 10 consecutive complete buckets are at least 1% but below 5%."}, "instructions": "Rate the alert under this exhaustive rubric: Noise if the latest rate is below 1%; Observe if it is at least 1% but fewer than 10 consecutive complete buckets exist; SEV-3 if 10 or more consecutive buckets are at least 1% but below 5%; SEV-2 if 10 or more are at least 5% but below 20%; SEV-1 if 10 or more are at least 20%; choose None of the above only if the supplied evidence cannot fit any rubric.", "type": "choice"}}, "state": [{"speaker": "Monitoring analyst", "text": "At 14:12, the supplied evidence for checkout-api consists exclusively of eleven records designated C1 through C11, with each record containing exactly one observation and no other checkout-api data present."}, {"speaker": "Incident reviewer", "text": "Records C1 through C11 are complete one-minute buckets ending respectively at 14:01, 14:02, 14:03, 14:04, 14:05, 14:07, 14:08, 14:09, 14:10, 14:11, and 14:12, with respective user-visible HTTP 5xx rates of 6.1%, 7.4%, 8.2%, 9.5%, 10.3%, 11.6%, 12.8%, 13.1%, 14.7%, 15.2%, and 16.4%."}, {"speaker": "On-call engineer", "text": "At the review time, C11 was the latest available user-visible result for checkout-api. Traffic remained normal, and the alert was ready for assessment."}, {"speaker": "Incident coordinator", "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}]}, "method": "c2d", "provenance": {"source_id": "diverse-134", "source_is_synthetic": true, "source_sha256": "ad1d4e1484cdb6a3a2e5104e486414ba4c97d04018b99270162dffeb80a0d8f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "observe"}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full classification rubric, while both contexts retain the original-state runbook rules about consecutive complete buckets and boundary handling. Both contexts remain bound to checkout-api, the user-visible HTTP 5xx metric, and the 14:12 assessment snapshot. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes only Q36 from complete to incomplete; assigning a rate to an incomplete bucket does not contradict its incompleteness, and no duplicate assertion gives Q36 conflicting completeness states within that context. Neither context embeds a classification answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal bucket-rate atoms remain atomic. The focus a1 is factual rather than a policy conclusion. Both assignments are realizable with only a1 changing: one context can supply at least 10 qualifying consecutive buckets, while another can explicitly supply fewer than 10, with all supplied rates still between 5% and 20% and the latest rate at least 1%. The retained question already preserves the main rubric, and policy_evidence appropriately preserves the state-origin runbook rule concerning complete consecutive buckets and boundary handling.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 guarantees at least 10 consecutive complete buckets, while a2 and a3 place every supplied complete bucket in the inclusive-5%-to-exclusive-20% band. This is sufficient for SEV-2 and excludes SEV-1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that fewer than 10 consecutive complete buckets are supplied, and a4 excludes Noise by establishing a latest rate of at least 1%. Under the exhaustive rubric, these conditions are sufficient for Observe.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 14:12, the supplied evidence contains a sequence of at least 10 consecutive complete one-minute buckets for checkout-api."}, {"id": "a2", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5%."}, {"id": "a3", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate below 20%."}, {"id": "a4", "statement": "At 14:12, the latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}], "base_state_json": "[{\"speaker\":\"Metrics archivist\",\"text\":\"The 14:12 snapshot was exported from the user-visible edge metrics view; dependency telemetry and synthetic-check series were excluded.\"},{\"speaker\":\"On-call engineer\",\"text\":\"At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%.\"},{\"speaker\":\"Incident reviewer\",\"text\":\"At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in the supplied evidence are complete.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%."}, {"path": ["2", "text"], "text": "At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in the supplied evidence are complete."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}], "rules": [{"justification": "At least 10 consecutive complete buckets exist, and every supplied complete bucket is in the inclusive-5%-to-exclusive-20% band required for SEV-2.", "target": "sev2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "The latest rate is at least 1%, while the evidence does not contain 10 consecutive complete buckets, which is sufficient for Observe.", "target": "observe", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%.", "negative_left": "At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%.", "negative_right": "At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q37, Q38, Q39, and Q40 in the supplied evidence are complete, while bucket record Q36 is incomplete.", "right": "At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in the supplied evidence are complete."}, "verifier_independent_model": false}, "family": "scale-diverse-134-005", "id": "scale-diverse-134-005-base", "input": {"questions": {"decision": {"criteria": {"noise": "Classify as Noise because the latest user-visible error rate is below 1%.", "none_of_above": "Choose None of the above because the supplied evidence cannot be classified under the stated rubric.", "observe": "Classify as Observe because at least 1% errors are present but fewer than 10 consecutive complete buckets exist.", "sev1": "Classify as SEV-1 because at least 10 consecutive complete buckets are at least 20%.", "sev2": "Classify as SEV-2 because at least 10 consecutive complete buckets are at least 5% but below 20%.", "sev3": "Classify as SEV-3 because at least 10 consecutive complete buckets are at least 1% but below 5%."}, "instructions": "Rate the alert under this exhaustive rubric: Noise if the latest rate is below 1%; Observe if it is at least 1% but fewer than 10 consecutive complete buckets exist; SEV-3 if 10 or more consecutive buckets are at least 1% but below 5%; SEV-2 if 10 or more are at least 5% but below 20%; SEV-1 if 10 or more are at least 20%; choose None of the above only if the supplied evidence cannot fit any rubric.", "type": "choice"}}, "state": [{"speaker": "Metrics archivist", "text": "The 14:12 snapshot was exported from the user-visible edge metrics view; dependency telemetry and synthetic-check series were excluded."}, {"speaker": "On-call engineer", "text": "At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%."}, {"speaker": "Incident reviewer", "text": "At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in the supplied evidence are complete."}, {"speaker": "Incident coordinator", "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}]}, "method": "c2d", "provenance": {"source_id": "diverse-134", "source_is_synthetic": true, "source_sha256": "ad1d4e1484cdb6a3a2e5104e486414ba4c97d04018b99270162dffeb80a0d8f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sev2"}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full classification rubric, while both contexts retain the original-state runbook rules about consecutive complete buckets and boundary handling. Both contexts remain bound to checkout-api, the user-visible HTTP 5xx metric, and the 14:12 assessment snapshot. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes only Q36 from complete to incomplete; assigning a rate to an incomplete bucket does not contradict its incompleteness, and no duplicate assertion gives Q36 conflicting completeness states within that context. Neither context embeds a classification answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal bucket-rate atoms remain atomic. The focus a1 is factual rather than a policy conclusion. Both assignments are realizable with only a1 changing: one context can supply at least 10 qualifying consecutive buckets, while another can explicitly supply fewer than 10, with all supplied rates still between 5% and 20% and the latest rate at least 1%. The retained question already preserves the main rubric, and policy_evidence appropriately preserves the state-origin runbook rule concerning complete consecutive buckets and boundary handling.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 guarantees at least 10 consecutive complete buckets, while a2 and a3 place every supplied complete bucket in the inclusive-5%-to-exclusive-20% band. This is sufficient for SEV-2 and excludes SEV-1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that fewer than 10 consecutive complete buckets are supplied, and a4 excludes Noise by establishing a latest rate of at least 1%. Under the exhaustive rubric, these conditions are sufficient for Observe.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 14:12, the supplied evidence contains a sequence of at least 10 consecutive complete one-minute buckets for checkout-api."}, {"id": "a2", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5%."}, {"id": "a3", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate below 20%."}, {"id": "a4", "statement": "At 14:12, the latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}], "base_state_json": "[{\"speaker\":\"Metrics archivist\",\"text\":\"The 14:12 snapshot was exported from the user-visible edge metrics view; dependency telemetry and synthetic-check series were excluded.\"},{\"speaker\":\"On-call engineer\",\"text\":\"At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%.\"},{\"speaker\":\"Incident reviewer\",\"text\":\"At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in the supplied evidence are complete.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%."}, {"path": ["2", "text"], "text": "At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in the supplied evidence are complete."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}], "rules": [{"justification": "At least 10 consecutive complete buckets exist, and every supplied complete bucket is in the inclusive-5%-to-exclusive-20% band required for SEV-2.", "target": "sev2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "The latest rate is at least 1%, while the evidence does not contain 10 consecutive complete buckets, which is sufficient for Observe.", "target": "observe", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%.", "negative_left": "At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%.", "negative_right": "At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q37, Q38, Q39, and Q40 in the supplied evidence are complete, while bucket record Q36 is incomplete.", "right": "At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in the supplied evidence are complete."}, "verifier_independent_model": false}, "family": "scale-diverse-134-005", "id": "scale-diverse-134-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"noise": "Classify as Noise because the latest user-visible error rate is below 1%.", "none_of_above": "Choose None of the above because the supplied evidence cannot be classified under the stated rubric.", "observe": "Classify as Observe because at least 1% errors are present but fewer than 10 consecutive complete buckets exist.", "sev1": "Classify as SEV-1 because at least 10 consecutive complete buckets are at least 20%.", "sev2": "Classify as SEV-2 because at least 10 consecutive complete buckets are at least 5% but below 20%.", "sev3": "Classify as SEV-3 because at least 10 consecutive complete buckets are at least 1% but below 5%."}, "instructions": "Rate the alert under this exhaustive rubric: Noise if the latest rate is below 1%; Observe if it is at least 1% but fewer than 10 consecutive complete buckets exist; SEV-3 if 10 or more consecutive buckets are at least 1% but below 5%; SEV-2 if 10 or more are at least 5% but below 20%; SEV-1 if 10 or more are at least 20%; choose None of the above only if the supplied evidence cannot fit any rubric.", "type": "choice"}}, "state": [{"speaker": "Metrics archivist", "text": "The 14:12 snapshot was exported from the user-visible edge metrics view; dependency telemetry and synthetic-check series were excluded."}, {"speaker": "On-call engineer", "text": "At 14:12, the checkout-api evidence consists solely of one-minute bucket records Q31, Q32, Q33, Q34, Q35, Q36, Q37, Q38, Q39, and Q40 in consecutive earliest-to-latest order, with a user-visible HTTP 5xx rate of 8.4% for each record except the latest record Q40, whose rate is 12.6%."}, {"speaker": "Incident reviewer", "text": "At 14:12, checkout-api bucket records Q31, Q32, Q33, Q34, Q35, Q37, Q38, Q39, and Q40 in the supplied evidence are complete, while bucket record Q36 is incomplete."}, {"speaker": "Incident coordinator", "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}]}, "method": "c2d", "provenance": {"source_id": "diverse-134", "source_is_synthetic": true, "source_sha256": "ad1d4e1484cdb6a3a2e5104e486414ba4c97d04018b99270162dffeb80a0d8f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "observe"}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same runbook rule, request scope, Catalog-routing decision, and alert timeline. The two focus-evidence spans are complete factual sentences. The counterfactual coherently swaps the slot occupants without conflicting with the ledger’s statement that R-17 maps through slot 63 to exactly one of the two services. Neither context contains an explicit gold answer, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual assignment of R-17 to Pricing rather than a policy conclusion. The policy evidence preserves the substantive runbook rule originating in the original state; instructions and criteria in the retained questions object need not be repeated. The base and counter assignments are jointly realizable with only a8 changing: in the base R-17 is assigned to Pricing, while in the counter it is not assigned to Pricing and, given a6 and a7, is assigned uniquely to Catalog. Both rules include sufficient evidence and the necessary competing-service exclusion.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that at least three traces identify R-17 on the earliest failing span, an independent failure metric agrees on R-17 during the alert interval, and R-17 is uniquely assigned to Pricing. Under the cited runbook, the alert must therefore be routed to Pricing, which entails the false target.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same qualifying trace-and-metric agreement. Because R-17 is uniquely assigned to a service in {Catalog, Pricing} and its assignment to Pricing is refuted, it must be assigned to Catalog. The runbook therefore entails routing to Catalog, satisfying the true target.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of sampled traces for the Catalog API alert under review whose earliest failing span carries service-resource identifier R-17 is at least three."}, {"id": "a2", "statement": "Metric M-4 is independent of the sampled traces for the Catalog API alert under review."}, {"id": "a3", "statement": "Metric M-4 indicates a failure during the six-minute interval of the Catalog API alert under review."}, {"id": "a4", "statement": "Metric M-4 carries service-resource identifier R-17."}, {"id": "a5", "statement": "Exactly one service-resource identifier appears on the earliest failing span in at least three sampled traces for the Catalog API alert under review."}, {"id": "a6", "statement": "The service assigned service-resource identifier R-17 is a member of the set consisting of the Catalog service and the Pricing service."}, {"id": "a7", "statement": "Exactly one service is assigned service-resource identifier R-17."}, {"id": "a8", "statement": "Service-resource identifier R-17 is assigned to the Pricing service."}], "base_state_json": "\"At 2026-09-17T10:08:00Z, monitoring opened the Catalog API alert after failures persisted for six minutes. At 2026-09-17T10:12:00Z, reviewers examined the sampled traces: four had an earliest failing span carrying service-resource identifier R-17, and no other identifier appeared on the earliest failing span in at least three sampled traces. At 2026-09-17T10:13:00Z, metric M-4, collected independently of those sampled traces, indicated a failure during the same six-minute alert interval and carried service-resource identifier R-17. At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service. At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Pricing service, while ledger slot 64 was occupied by the Catalog service. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service."}, {"path": [], "text": "At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Pricing service, while ledger slot 64 was occupied by the Catalog service."}], "policy_evidence": [{"path": [], "text": "The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}], "rules": [{"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, R-17 is the sole trace-qualified identifier, and R-17 is uniquely assigned to Pricing. The runbook therefore routes the alert to Pricing, not Catalog.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, and R-17 is the sole trace-qualified identifier. Because R-17 has exactly one assignment, that assignment is within {Catalog, Pricing}, and it is explicitly not Pricing, R-17 is assigned to Catalog. The runbook therefore routes the alert to Catalog.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service.", "negative_left": "At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service.", "negative_right": "At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Catalog service, while ledger slot 64 was occupied by the Pricing service.", "right": "At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Pricing service, while ledger slot 64 was occupied by the Catalog service."}, "verifier_independent_model": false}, "family": "scale-diverse-135-001", "id": "scale-diverse-135-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the alert to the Catalog service owner; route it to the Pricing service owner.", "true": "Yes — route the alert to the Catalog service owner."}, "instructions": "Decide whether the alert should be routed to the Catalog service owner under the supplied runbook.", "type": "noul"}}, "state": "At 2026-09-17T10:08:00Z, monitoring opened the Catalog API alert after failures persisted for six minutes. At 2026-09-17T10:12:00Z, reviewers examined the sampled traces: four had an earliest failing span carrying service-resource identifier R-17, and no other identifier appeared on the earliest failing span in at least three sampled traces. At 2026-09-17T10:13:00Z, metric M-4, collected independently of those sampled traces, indicated a failure during the same six-minute alert interval and carried service-resource identifier R-17. At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service. At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Pricing service, while ledger slot 64 was occupied by the Catalog service. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}, "method": "c2d", "provenance": {"source_id": "diverse-135", "source_is_synthetic": true, "source_sha256": "7c01851c92ec31a0d400586f2e516fdb19e7f2a02cf1d4c8fb1140fb2791f3a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same runbook rule, request scope, Catalog-routing decision, and alert timeline. The two focus-evidence spans are complete factual sentences. The counterfactual coherently swaps the slot occupants without conflicting with the ledger’s statement that R-17 maps through slot 63 to exactly one of the two services. Neither context contains an explicit gold answer, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual assignment of R-17 to Pricing rather than a policy conclusion. The policy evidence preserves the substantive runbook rule originating in the original state; instructions and criteria in the retained questions object need not be repeated. The base and counter assignments are jointly realizable with only a8 changing: in the base R-17 is assigned to Pricing, while in the counter it is not assigned to Pricing and, given a6 and a7, is assigned uniquely to Catalog. Both rules include sufficient evidence and the necessary competing-service exclusion.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that at least three traces identify R-17 on the earliest failing span, an independent failure metric agrees on R-17 during the alert interval, and R-17 is uniquely assigned to Pricing. Under the cited runbook, the alert must therefore be routed to Pricing, which entails the false target.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same qualifying trace-and-metric agreement. Because R-17 is uniquely assigned to a service in {Catalog, Pricing} and its assignment to Pricing is refuted, it must be assigned to Catalog. The runbook therefore entails routing to Catalog, satisfying the true target.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of sampled traces for the Catalog API alert under review whose earliest failing span carries service-resource identifier R-17 is at least three."}, {"id": "a2", "statement": "Metric M-4 is independent of the sampled traces for the Catalog API alert under review."}, {"id": "a3", "statement": "Metric M-4 indicates a failure during the six-minute interval of the Catalog API alert under review."}, {"id": "a4", "statement": "Metric M-4 carries service-resource identifier R-17."}, {"id": "a5", "statement": "Exactly one service-resource identifier appears on the earliest failing span in at least three sampled traces for the Catalog API alert under review."}, {"id": "a6", "statement": "The service assigned service-resource identifier R-17 is a member of the set consisting of the Catalog service and the Pricing service."}, {"id": "a7", "statement": "Exactly one service is assigned service-resource identifier R-17."}, {"id": "a8", "statement": "Service-resource identifier R-17 is assigned to the Pricing service."}], "base_state_json": "\"At 2026-09-17T10:08:00Z, monitoring opened the Catalog API alert after failures persisted for six minutes. At 2026-09-17T10:12:00Z, reviewers examined the sampled traces: four had an earliest failing span carrying service-resource identifier R-17, and no other identifier appeared on the earliest failing span in at least three sampled traces. At 2026-09-17T10:13:00Z, metric M-4, collected independently of those sampled traces, indicated a failure during the same six-minute alert interval and carried service-resource identifier R-17. At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service. At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Pricing service, while ledger slot 64 was occupied by the Catalog service. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service."}, {"path": [], "text": "At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Pricing service, while ledger slot 64 was occupied by the Catalog service."}], "policy_evidence": [{"path": [], "text": "The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}], "rules": [{"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, R-17 is the sole trace-qualified identifier, and R-17 is uniquely assigned to Pricing. The runbook therefore routes the alert to Pricing, not Catalog.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, and R-17 is the sole trace-qualified identifier. Because R-17 has exactly one assignment, that assignment is within {Catalog, Pricing}, and it is explicitly not Pricing, R-17 is assigned to Catalog. The runbook therefore routes the alert to Catalog.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service.", "negative_left": "At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service.", "negative_right": "At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Catalog service, while ledger slot 64 was occupied by the Pricing service.", "right": "At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Pricing service, while ledger slot 64 was occupied by the Catalog service."}, "verifier_independent_model": false}, "family": "scale-diverse-135-001", "id": "scale-diverse-135-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the alert to the Catalog service owner; route it to the Pricing service owner.", "true": "Yes — route the alert to the Catalog service owner."}, "instructions": "Decide whether the alert should be routed to the Catalog service owner under the supplied runbook.", "type": "noul"}}, "state": "At 2026-09-17T10:08:00Z, monitoring opened the Catalog API alert after failures persisted for six minutes. At 2026-09-17T10:12:00Z, reviewers examined the sampled traces: four had an earliest failing span carrying service-resource identifier R-17, and no other identifier appeared on the earliest failing span in at least three sampled traces. At 2026-09-17T10:13:00Z, metric M-4, collected independently of those sampled traces, indicated a failure during the same six-minute alert interval and carried service-resource identifier R-17. At 2026-09-17T10:14:00Z, the complete assignment ledger linked service-resource identifier R-17 to exactly one service through ledger slot 63, and the service occupying that slot was either the Catalog service or the Pricing service. At 2026-09-17T10:14:00Z, ledger slot 63 was occupied by the Catalog service, while ledger slot 64 was occupied by the Pricing service. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}, "method": "c2d", "provenance": {"source_id": "diverse-135", "source_is_synthetic": true, "source_sha256": "7c01851c92ec31a0d400586f2e516fdb19e7f2a02cf1d4c8fb1140fb2791f3a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original runbook rule, while the unchanged questions preserve the decision instructions and criteria. The Catalog API routing decision and six-minute alert interval remain bound to the original request. The two evidence spans are complete factual sentences. Changing Q-62’s ledger assignment from Catalog to Pricing coherently reverses R-17’s assignment under the stated bijection without creating duplicate or contradictory measurements. Neither generated context states the required output, an answer code, or a label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual assignment of R-17 to Pricing rather than a policy conclusion. The policy evidence preserves the substantive runbook rule originating in the original state; instructions and criteria in the retained questions object need not be repeated. The base and counter assignments are jointly realizable with only a8 changing: in the base R-17 is assigned to Pricing, while in the counter it is not assigned to Pricing and, given a6 and a7, is assigned uniquely to Catalog. Both rules include sufficient evidence and the necessary competing-service exclusion.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that at least three traces identify R-17 on the earliest failing span, an independent failure metric agrees on R-17 during the alert interval, and R-17 is uniquely assigned to Pricing. Under the cited runbook, the alert must therefore be routed to Pricing, which entails the false target.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same qualifying trace-and-metric agreement. Because R-17 is uniquely assigned to a service in {Catalog, Pricing} and its assignment to Pricing is refuted, it must be assigned to Catalog. The runbook therefore entails routing to Catalog, satisfying the true target.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of sampled traces for the Catalog API alert under review whose earliest failing span carries service-resource identifier R-17 is at least three."}, {"id": "a2", "statement": "Metric M-4 is independent of the sampled traces for the Catalog API alert under review."}, {"id": "a3", "statement": "Metric M-4 indicates a failure during the six-minute interval of the Catalog API alert under review."}, {"id": "a4", "statement": "Metric M-4 carries service-resource identifier R-17."}, {"id": "a5", "statement": "Exactly one service-resource identifier appears on the earliest failing span in at least three sampled traces for the Catalog API alert under review."}, {"id": "a6", "statement": "The service assigned service-resource identifier R-17 is a member of the set consisting of the Catalog service and the Pricing service."}, {"id": "a7", "statement": "Exactly one service is assigned service-resource identifier R-17."}, {"id": "a8", "statement": "Service-resource identifier R-17 is assigned to the Pricing service."}], "base_state_json": "\"Field note: The engineer is deciding whether the Catalog API alert should be routed to the Catalog service owner. The trace audit covered twelve sampled traces. In four traces, the earliest failing span carried service-resource identifier R-17; every other identifier appeared on the earliest failing span in fewer than three traces. Metric M-4 came from host telemetry collected independently of those sampled traces. During the alert’s six-minute interval, M-4 registered a failure and carried R-17. At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service. At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Catalog service. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service."}, {"path": [], "text": "At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Catalog service."}], "policy_evidence": [{"path": [], "text": "The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}], "rules": [{"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, R-17 is the sole trace-qualified identifier, and R-17 is uniquely assigned to Pricing. The runbook therefore routes the alert to Pricing, not Catalog.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, and R-17 is the sole trace-qualified identifier. Because R-17 has exactly one assignment, that assignment is within {Catalog, Pricing}, and it is explicitly not Pricing, R-17 is assigned to Catalog. The runbook therefore routes the alert to Catalog.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service.", "negative_left": "At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service.", "negative_right": "At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Pricing service.", "right": "At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Catalog service."}, "verifier_independent_model": false}, "family": "scale-diverse-135-004", "id": "scale-diverse-135-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the alert to the Catalog service owner; route it to the Pricing service owner.", "true": "Yes — route the alert to the Catalog service owner."}, "instructions": "Decide whether the alert should be routed to the Catalog service owner under the supplied runbook.", "type": "noul"}}, "state": "Field note: The engineer is deciding whether the Catalog API alert should be routed to the Catalog service owner. The trace audit covered twelve sampled traces. In four traces, the earliest failing span carried service-resource identifier R-17; every other identifier appeared on the earliest failing span in fewer than three traces. Metric M-4 came from host telemetry collected independently of those sampled traces. During the alert’s six-minute interval, M-4 registered a failure and carried R-17. At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service. At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Catalog service. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}, "method": "c2d", "provenance": {"source_id": "diverse-135", "source_is_synthetic": true, "source_sha256": "7c01851c92ec31a0d400586f2e516fdb19e7f2a02cf1d4c8fb1140fb2791f3a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original runbook rule, while the unchanged questions preserve the decision instructions and criteria. The Catalog API routing decision and six-minute alert interval remain bound to the original request. The two evidence spans are complete factual sentences. Changing Q-62’s ledger assignment from Catalog to Pricing coherently reverses R-17’s assignment under the stated bijection without creating duplicate or contradictory measurements. Neither generated context states the required output, an answer code, or a label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual assignment of R-17 to Pricing rather than a policy conclusion. The policy evidence preserves the substantive runbook rule originating in the original state; instructions and criteria in the retained questions object need not be repeated. The base and counter assignments are jointly realizable with only a8 changing: in the base R-17 is assigned to Pricing, while in the counter it is not assigned to Pricing and, given a6 and a7, is assigned uniquely to Catalog. Both rules include sufficient evidence and the necessary competing-service exclusion.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that at least three traces identify R-17 on the earliest failing span, an independent failure metric agrees on R-17 during the alert interval, and R-17 is uniquely assigned to Pricing. Under the cited runbook, the alert must therefore be routed to Pricing, which entails the false target.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same qualifying trace-and-metric agreement. Because R-17 is uniquely assigned to a service in {Catalog, Pricing} and its assignment to Pricing is refuted, it must be assigned to Catalog. The runbook therefore entails routing to Catalog, satisfying the true target.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of sampled traces for the Catalog API alert under review whose earliest failing span carries service-resource identifier R-17 is at least three."}, {"id": "a2", "statement": "Metric M-4 is independent of the sampled traces for the Catalog API alert under review."}, {"id": "a3", "statement": "Metric M-4 indicates a failure during the six-minute interval of the Catalog API alert under review."}, {"id": "a4", "statement": "Metric M-4 carries service-resource identifier R-17."}, {"id": "a5", "statement": "Exactly one service-resource identifier appears on the earliest failing span in at least three sampled traces for the Catalog API alert under review."}, {"id": "a6", "statement": "The service assigned service-resource identifier R-17 is a member of the set consisting of the Catalog service and the Pricing service."}, {"id": "a7", "statement": "Exactly one service is assigned service-resource identifier R-17."}, {"id": "a8", "statement": "Service-resource identifier R-17 is assigned to the Pricing service."}], "base_state_json": "\"Field note: The engineer is deciding whether the Catalog API alert should be routed to the Catalog service owner. The trace audit covered twelve sampled traces. In four traces, the earliest failing span carried service-resource identifier R-17; every other identifier appeared on the earliest failing span in fewer than three traces. Metric M-4 came from host telemetry collected independently of those sampled traces. During the alert’s six-minute interval, M-4 registered a failure and carried R-17. At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service. At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Catalog service. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service."}, {"path": [], "text": "At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Catalog service."}], "policy_evidence": [{"path": [], "text": "The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}], "rules": [{"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, R-17 is the sole trace-qualified identifier, and R-17 is uniquely assigned to Pricing. The runbook therefore routes the alert to Pricing, not Catalog.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, and R-17 is the sole trace-qualified identifier. Because R-17 has exactly one assignment, that assignment is within {Catalog, Pricing}, and it is explicitly not Pricing, R-17 is assigned to Catalog. The runbook therefore routes the alert to Catalog.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service.", "negative_left": "At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service.", "negative_right": "At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Pricing service.", "right": "At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Catalog service."}, "verifier_independent_model": false}, "family": "scale-diverse-135-004", "id": "scale-diverse-135-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the alert to the Catalog service owner; route it to the Pricing service owner.", "true": "Yes — route the alert to the Catalog service owner."}, "instructions": "Decide whether the alert should be routed to the Catalog service owner under the supplied runbook.", "type": "noul"}}, "state": "Field note: The engineer is deciding whether the Catalog API alert should be routed to the Catalog service owner. The trace audit covered twelve sampled traces. In four traces, the earliest failing span carried service-resource identifier R-17; every other identifier appeared on the earliest failing span in fewer than three traces. Metric M-4 came from host telemetry collected independently of those sampled traces. During the alert’s six-minute interval, M-4 registered a failure and carried R-17. At 09:40 UTC on 17 September 2026, the service-assignment ledger recorded a bijection between identifiers R-17 and Q-62 and the Catalog service and Pricing service. At 09:40 UTC on 17 September 2026, the ledger entry for Q-62 named the Pricing service. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}, "method": "c2d", "provenance": {"source_id": "diverse-135", "source_is_synthetic": true, "source_sha256": "7c01851c92ec31a0d400586f2e516fdb19e7f2a02cf1d4c8fb1140fb2791f3a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same routing runbook, request scope, ShopWave alert entity, and relevant alert interval. The two focus spans are complete factual sentences. The counterfactual changes only CR-8046’s registered service from Payments to Fulfillment, remains consistent with the unchanged SC-17-to-CR-8046 mapping, and creates no duplicate or contradictory assertion. Neither context includes an answer code, proposition ID, rule table, explicit output instruction, or embedded gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"The operational handoff records trace T1 for the ShopWave production alert under assessment as completed.\",\"The completed trace identifies exactly one service as a slow component, and the service identifier recorded for that sole component is SC-17.\",\"In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046.\",\"In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Payments.\",\"The Payments dependency dashboard was checked for the alert interval; it contains no data for that interval and does not confirm Payments degradation.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046."}, {"path": ["evidence", "3"], "text": "In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Payments."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046.", "negative_left": "In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046.", "negative_right": "In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Fulfillment and which is distinct from the Payments service.", "right": "In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Payments."}, "verifier_independent_model": false}, "family": "scale-diverse-136-003", "id": "scale-diverse-136-003-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["The operational handoff records trace T1 for the ShopWave production alert under assessment as completed.", "The completed trace identifies exactly one service as a slow component, and the service identifier recorded for that sole component is SC-17.", "In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046.", "In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Payments.", "The Payments dependency dashboard was checked for the alert interval; it contains no data for that interval and does not confirm Payments degradation."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same routing runbook, request scope, ShopWave alert entity, and relevant alert interval. The two focus spans are complete factual sentences. The counterfactual changes only CR-8046’s registered service from Payments to Fulfillment, remains consistent with the unchanged SC-17-to-CR-8046 mapping, and creates no duplicate or contradictory assertion. Neither context includes an answer code, proposition ID, rule table, explicit output instruction, or embedded gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"The operational handoff records trace T1 for the ShopWave production alert under assessment as completed.\",\"The completed trace identifies exactly one service as a slow component, and the service identifier recorded for that sole component is SC-17.\",\"In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046.\",\"In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Payments.\",\"The Payments dependency dashboard was checked for the alert interval; it contains no data for that interval and does not confirm Payments degradation.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046."}, {"path": ["evidence", "3"], "text": "In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Payments."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046.", "negative_left": "In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046.", "negative_right": "In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Fulfillment and which is distinct from the Payments service.", "right": "In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Payments."}, "verifier_independent_model": false}, "family": "scale-diverse-136-003", "id": "scale-diverse-136-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["The operational handoff records trace T1 for the ShopWave production alert under assessment as completed.", "The completed trace identifies exactly one service as a slow component, and the service identifier recorded for that sole component is SC-17.", "In the authoritative service-registry snapshot attached to the ShopWave alert handoff at 2026-09-17T14:32:00Z, service identifier SC-17 denotes exactly the service represented by catalog record CR-8046.", "In that snapshot, catalog record CR-8046 represents exactly one service, whose registered name is Fulfillment and which is distinct from the Payments service.", "The Payments dependency dashboard was checked for the alert interval; it contains no data for that interval and does not confirm Payments degradation."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing runbook, request scope, ShopWave alert entity, Payments routing target, and assessment framing; the unchanged questions preserve the remaining instructions and decision criteria. The two focus-evidence spans are complete factual sentences, and the counterfactual coherently changes SC-17’s exclusive registry assignment from Payments to distinct Fulfillment without conflicting with the unchanged trace, dashboard, time-window, or request-path facts. Neither context embeds a gold answer, output code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval.\",\"The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Payments service.\",\"The incident log fixes the assessment window at 03:15–03:23 UTC and notes that T1 was collected from the affected production request path.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval."}, {"path": ["evidence", "1"], "text": "The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval.", "negative_left": "At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval.", "negative_right": "The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Fulfillment service, which it recorded as distinct from the Payments service.", "right": "The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Payments service."}, "verifier_independent_model": false}, "family": "scale-diverse-136-004", "id": "scale-diverse-136-004-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval.", "The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Payments service.", "The incident log fixes the assessment window at 03:15–03:23 UTC and notes that T1 was collected from the affected production request path."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing runbook, request scope, ShopWave alert entity, Payments routing target, and assessment framing; the unchanged questions preserve the remaining instructions and decision criteria. The two focus-evidence spans are complete factual sentences, and the counterfactual coherently changes SC-17’s exclusive registry assignment from Payments to distinct Fulfillment without conflicting with the unchanged trace, dashboard, time-window, or request-path facts. Neither context embeds a gold answer, output code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval.\",\"The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Payments service.\",\"The incident log fixes the assessment window at 03:15–03:23 UTC and notes that T1 was collected from the affected production request path.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval."}, {"path": ["evidence", "1"], "text": "The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval.", "negative_left": "At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval.", "negative_right": "The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Fulfillment service, which it recorded as distinct from the Payments service.", "right": "The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Payments service."}, "verifier_independent_model": false}, "family": "scale-diverse-136-004", "id": "scale-diverse-136-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["At 03:18 UTC on 14 May 2026, completed trace T1 for the ShopWave production alert identified exactly one slow component, recorded its identifier as SC-17, and the dependency dashboard showed no confirmed Payments degradation during the traced interval.", "The signed ShopWave service registry effective at 03:18 UTC on 14 May 2026 assigned the identifier recorded for T1's sole slow component exclusively to the Fulfillment service, which it recorded as distinct from the Payments service.", "The incident log fixes the assessment window at 03:15–03:23 UTC and notes that T1 was collected from the affected production request path."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both inputs preserve the questions verbatim and retain the original routing runbook, request, alert, entities, and relevant interval. The two focus spans are complete factual sentences. The counterfactual changes only the crosswalk association from CR-482 to the distinct CR-739; this coherently breaks the link between SC-17 and Payments without contradicting the unchanged registry assignment or dashboard evidence. Neither context contains a gold answer, answer code, proposition identifier, rule table, label rationale, or classifier-output instruction beyond the preserved question instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.\",\"true\":\"The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments.\"},\"instructions\":\"Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.\",\"type\":\"noul\"}},\"state\":{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"At 09:04 UTC on 14 March 2026, trace T1 for the ShopWave production alert finished processing and was marked completed.\",\"At 09:06 UTC, the finalized T1 results identified exactly one service as a slow component.\",\"At 09:07 UTC, the result record gave that sole slow component the service identifier SC-17.\",\"At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482.\",\"At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-482.\",\"At 09:20 UTC, the Payments dependency dashboard's interval report explicitly stated that it did not confirm Payments degradation during the relevant alert interval.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["state", "evidence", "3"], "text": "At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482."}, {"path": ["state", "evidence", "4"], "text": "At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-482."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482.", "negative_left": "At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482.", "negative_right": "At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-739, a record distinct from CR-482.", "right": "At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-482."}, "verifier_independent_model": false}, "family": "scale-diverse-136-005", "id": "scale-diverse-136-005-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["At 09:04 UTC on 14 March 2026, trace T1 for the ShopWave production alert finished processing and was marked completed.", "At 09:06 UTC, the finalized T1 results identified exactly one service as a slow component.", "At 09:07 UTC, the result record gave that sole slow component the service identifier SC-17.", "At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482.", "At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-482.", "At 09:20 UTC, the Payments dependency dashboard's interval report explicitly stated that it did not confirm Payments degradation during the relevant alert interval."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both inputs preserve the questions verbatim and retain the original routing runbook, request, alert, entities, and relevant interval. The two focus spans are complete factual sentences. The counterfactual changes only the crosswalk association from CR-482 to the distinct CR-739; this coherently breaks the link between SC-17 and Payments without contradicting the unchanged registry assignment or dashboard evidence. Neither context contains a gold answer, answer code, proposition identifier, rule table, label rationale, or classifier-output instruction beyond the preserved question instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.\",\"true\":\"The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments.\"},\"instructions\":\"Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.\",\"type\":\"noul\"}},\"state\":{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"At 09:04 UTC on 14 March 2026, trace T1 for the ShopWave production alert finished processing and was marked completed.\",\"At 09:06 UTC, the finalized T1 results identified exactly one service as a slow component.\",\"At 09:07 UTC, the result record gave that sole slow component the service identifier SC-17.\",\"At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482.\",\"At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-482.\",\"At 09:20 UTC, the Payments dependency dashboard's interval report explicitly stated that it did not confirm Payments degradation during the relevant alert interval.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["state", "evidence", "3"], "text": "At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482."}, {"path": ["state", "evidence", "4"], "text": "At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-482."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482.", "negative_left": "At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482.", "negative_right": "At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-739, a record distinct from CR-482.", "right": "At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-482."}, "verifier_independent_model": false}, "family": "scale-diverse-136-005", "id": "scale-diverse-136-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["At 09:04 UTC on 14 March 2026, trace T1 for the ShopWave production alert finished processing and was marked completed.", "At 09:06 UTC, the finalized T1 results identified exactly one service as a slow component.", "At 09:07 UTC, the result record gave that sole slow component the service identifier SC-17.", "At 09:10 UTC on 14 March 2026, ShopWave's service registry assigned the Payments service exclusively to canonical record CR-482.", "At 09:15 UTC on 14 March 2026, ShopWave's identifier crosswalk listed SC-17 exclusively as an alternate identifier for the service in canonical record CR-739, a record distinct from CR-482.", "At 09:20 UTC, the Payments dependency dashboard's interval report explicitly stated that it did not confirm Payments degradation during the relevant alert interval."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the production-only scope, synthetic/beta exclusion, inclusive-threshold rule, Checkout API entity, and fixed 14:20–14:30 UTC window without modifying the governing rubric. The two evidence spans are complete factual sentences; changing ap-south from 6% to 4% is coherent with the unchanged eu-west and other-region measurements, and neither context contains an impact-level answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual regional-failure relationship; A2 is a permissible universally quantified relation and A3 is a single region-count relation. A3 is factual rather than a policy conclusion. The base and counter assignments are realizable while changing only A3: the base can have two regions between 5% and 25%, while the counter can have only eu-west in that range, with no region at 25% or above in either case. The policy evidence correctly preserves the relevant state-originating scope, exclusions, threshold inclusivity, and escalation binding; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 alone entails the level-3 condition that at least 5% fail in two or more regions for the full 10-minute window. No exclusion of lower outcomes is needed because the ordered rubric classifies this directly as major broad impact.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes at least one qualifying region at 5% or more for 10 minutes. Refuted A3 entails that fewer than two regions qualify, so eu-west is exactly one qualifying region. A2 excludes every region from reaching the competing 25% level-3 threshold. Together these conditions entail level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During 14:20–14:30 UTC, at least 5% of in-scope production customer requests failed in eu-west for the full 10-minute window."}, {"id": "A2", "statement": "During 14:20–14:30 UTC, no region had at least 25% of in-scope production customer requests fail for the full 10-minute window."}, {"id": "A3", "statement": "During 14:20–14:30 UTC, at least two distinct regions each had at least 5% of in-scope production customer requests fail for the full 10-minute window."}], "base_state_json": "[{\"speaker\":\"Incident coordinator\",\"text\":\"Operational handoff for the Checkout API covers the fixed 14:20–14:30 UTC assessment window. The regional telemetry export is complete, and request outcomes were evaluated continuously across that entire interval rather than from a point sample.\"},{\"speaker\":\"On-call engineer\",\"text\":\"For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south.\"},{\"speaker\":\"Site reliability engineer\",\"text\":\"For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 6% failure rate in ap-south throughout the full window.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner.\"},{\"speaker\":\"Service owner\",\"text\":\"The handoff confirms that the dashboard’s scope filter was active for the complete window. The incident record is ready for impact-level assignment under the stated rubric.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south."}, {"path": ["2", "text"], "text": "For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 6% failure rate in ap-south throughout the full window."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}], "rules": [{"justification": "At least two regions each meeting the inclusive 5% threshold for the full 10-minute window is sufficient for major broad impact.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Eu-west supplies at least one region meeting 5% for 10 minutes; refutation of A3 limits the count of such regions to one, and A2 excludes the competing level-3 condition of any region reaching 25%.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south.", "negative_left": "For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south.", "negative_right": "For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 4% failure rate in ap-south throughout the full window.", "right": "For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 6% failure rate in ap-south throughout the full window."}, "verifier_independent_model": false}, "family": "scale-diverse-137-003", "id": "scale-diverse-137-003-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified production impact: only excluded synthetic or beta traffic is affected, or in-scope production failures remain below 1% in every region during the window.", "1 — Minor verified impact: at least 1% but less than 5% of in-scope production requests fail in one region for 10 minutes, with no other region at 1% or higher.", "2 — Significant regional impact: at least 5% but less than 25% of in-scope production requests fail in exactly one region for at least 10 minutes; escalate to the owning service team.", "3 — Major broad impact: at least 25% of in-scope production requests fail in any region for at least 10 minutes, or at least 5% fail in two or more regions for that duration."], "instructions": "Assign the alert an impact level using the ordered rubric. Assess only the stated 10-minute window and apply the runbook’s scope and exceptions.", "type": "score"}}, "state": [{"speaker": "Incident coordinator", "text": "Operational handoff for the Checkout API covers the fixed 14:20–14:30 UTC assessment window. The regional telemetry export is complete, and request outcomes were evaluated continuously across that entire interval rather than from a point sample."}, {"speaker": "On-call engineer", "text": "For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south."}, {"speaker": "Site reliability engineer", "text": "For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 6% failure rate in ap-south throughout the full window."}, {"speaker": "Incident coordinator", "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}, {"speaker": "Service owner", "text": "The handoff confirms that the dashboard’s scope filter was active for the complete window. The incident record is ready for impact-level assignment under the stated rubric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-137", "source_is_synthetic": true, "source_sha256": "6ad980dc01b48352a342f771b5dfe60d2cbfd6e9d0710110e93512959974536f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the production-only scope, synthetic/beta exclusion, inclusive-threshold rule, Checkout API entity, and fixed 14:20–14:30 UTC window without modifying the governing rubric. The two evidence spans are complete factual sentences; changing ap-south from 6% to 4% is coherent with the unchanged eu-west and other-region measurements, and neither context contains an impact-level answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual regional-failure relationship; A2 is a permissible universally quantified relation and A3 is a single region-count relation. A3 is factual rather than a policy conclusion. The base and counter assignments are realizable while changing only A3: the base can have two regions between 5% and 25%, while the counter can have only eu-west in that range, with no region at 25% or above in either case. The policy evidence correctly preserves the relevant state-originating scope, exclusions, threshold inclusivity, and escalation binding; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 alone entails the level-3 condition that at least 5% fail in two or more regions for the full 10-minute window. No exclusion of lower outcomes is needed because the ordered rubric classifies this directly as major broad impact.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes at least one qualifying region at 5% or more for 10 minutes. Refuted A3 entails that fewer than two regions qualify, so eu-west is exactly one qualifying region. A2 excludes every region from reaching the competing 25% level-3 threshold. Together these conditions entail level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During 14:20–14:30 UTC, at least 5% of in-scope production customer requests failed in eu-west for the full 10-minute window."}, {"id": "A2", "statement": "During 14:20–14:30 UTC, no region had at least 25% of in-scope production customer requests fail for the full 10-minute window."}, {"id": "A3", "statement": "During 14:20–14:30 UTC, at least two distinct regions each had at least 5% of in-scope production customer requests fail for the full 10-minute window."}], "base_state_json": "[{\"speaker\":\"Incident coordinator\",\"text\":\"Operational handoff for the Checkout API covers the fixed 14:20–14:30 UTC assessment window. The regional telemetry export is complete, and request outcomes were evaluated continuously across that entire interval rather than from a point sample.\"},{\"speaker\":\"On-call engineer\",\"text\":\"For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south.\"},{\"speaker\":\"Site reliability engineer\",\"text\":\"For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 6% failure rate in ap-south throughout the full window.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner.\"},{\"speaker\":\"Service owner\",\"text\":\"The handoff confirms that the dashboard’s scope filter was active for the complete window. The incident record is ready for impact-level assignment under the stated rubric.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south."}, {"path": ["2", "text"], "text": "For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 6% failure rate in ap-south throughout the full window."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}], "rules": [{"justification": "At least two regions each meeting the inclusive 5% threshold for the full 10-minute window is sufficient for major broad impact.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Eu-west supplies at least one region meeting 5% for 10 minutes; refutation of A3 limits the count of such regions to one, and A2 excludes the competing level-3 condition of any region reaching 25%.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south.", "negative_left": "For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south.", "negative_right": "For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 4% failure rate in ap-south throughout the full window.", "right": "For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 6% failure rate in ap-south throughout the full window."}, "verifier_independent_model": false}, "family": "scale-diverse-137-003", "id": "scale-diverse-137-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified production impact: only excluded synthetic or beta traffic is affected, or in-scope production failures remain below 1% in every region during the window.", "1 — Minor verified impact: at least 1% but less than 5% of in-scope production requests fail in one region for 10 minutes, with no other region at 1% or higher.", "2 — Significant regional impact: at least 5% but less than 25% of in-scope production requests fail in exactly one region for at least 10 minutes; escalate to the owning service team.", "3 — Major broad impact: at least 25% of in-scope production requests fail in any region for at least 10 minutes, or at least 5% fail in two or more regions for that duration."], "instructions": "Assign the alert an impact level using the ordered rubric. Assess only the stated 10-minute window and apply the runbook’s scope and exceptions.", "type": "score"}}, "state": [{"speaker": "Incident coordinator", "text": "Operational handoff for the Checkout API covers the fixed 14:20–14:30 UTC assessment window. The regional telemetry export is complete, and request outcomes were evaluated continuously across that entire interval rather than from a point sample."}, {"speaker": "On-call engineer", "text": "For in-scope production customer requests during 14:20–14:30 UTC, the exhaustive operational-handoff telemetry recorded an 8% failure rate in eu-west throughout the full window and rates below 5% throughout the full window in every region other than eu-west and ap-south."}, {"speaker": "Site reliability engineer", "text": "For in-scope production customer requests during 14:20–14:30 UTC, the operational-handoff telemetry recorded a 4% failure rate in ap-south throughout the full window."}, {"speaker": "Incident coordinator", "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}, {"speaker": "Service owner", "text": "The handoff confirms that the dashboard’s scope filter was active for the complete window. The incident record is ready for impact-level assignment under the stated rubric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-137", "source_is_synthetic": true, "source_sha256": "6ad980dc01b48352a342f771b5dfe60d2cbfd6e9d0710110e93512959974536f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and preserve the original runbook policy verbatim. They remain bound to the same production-traffic scope, regional assessment, and 14:20–14:30 UTC window. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the us-central failure rate from 9% to 4% without conflicting with the three-region count or other measurements. Neither context contains an answer code, gold label, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual regional-failure relationship; A2 is a permissible universally quantified relation and A3 is a single region-count relation. A3 is factual rather than a policy conclusion. The base and counter assignments are realizable while changing only A3: the base can have two regions between 5% and 25%, while the counter can have only eu-west in that range, with no region at 25% or above in either case. The policy evidence correctly preserves the relevant state-originating scope, exclusions, threshold inclusivity, and escalation binding; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 alone entails the level-3 condition that at least 5% fail in two or more regions for the full 10-minute window. No exclusion of lower outcomes is needed because the ordered rubric classifies this directly as major broad impact.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes at least one qualifying region at 5% or more for 10 minutes. Refuted A3 entails that fewer than two regions qualify, so eu-west is exactly one qualifying region. A2 excludes every region from reaching the competing 25% level-3 threshold. Together these conditions entail level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During 14:20–14:30 UTC, at least 5% of in-scope production customer requests failed in eu-west for the full 10-minute window."}, {"id": "A2", "statement": "During 14:20–14:30 UTC, no region had at least 25% of in-scope production customer requests fail for the full 10-minute window."}, {"id": "A3", "statement": "During 14:20–14:30 UTC, at least two distinct regions each had at least 5% of in-scope production customer requests fail for the full 10-minute window."}], "base_state_json": "[{\"speaker\":\"Regional traffic audit\",\"text\":\"During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast.\"},{\"speaker\":\"Failure-rate audit\",\"text\":\"During 14:20–14:30 UTC, exactly 9% of in-scope production customer requests in us-central failed throughout the full interval.\"},{\"speaker\":\"Validation note\",\"text\":\"The regional request logs were complete for the stated window, and each reported share was calculated over production customer requests observed continuously throughout that interval.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast."}, {"path": ["1", "text"], "text": "During 14:20–14:30 UTC, exactly 9% of in-scope production customer requests in us-central failed throughout the full interval."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}], "rules": [{"justification": "At least two regions each meeting the inclusive 5% threshold for the full 10-minute window is sufficient for major broad impact.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Eu-west supplies at least one region meeting 5% for 10 minutes; refutation of A3 limits the count of such regions to one, and A2 excludes the competing level-3 condition of any region reaching 25%.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast.", "negative_left": "During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast.", "negative_right": "During 14:20–14:30 UTC, exactly 4% of in-scope production customer requests in us-central failed throughout the full interval.", "right": "During 14:20–14:30 UTC, exactly 9% of in-scope production customer requests in us-central failed throughout the full interval."}, "verifier_independent_model": false}, "family": "scale-diverse-137-004", "id": "scale-diverse-137-004-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified production impact: only excluded synthetic or beta traffic is affected, or in-scope production failures remain below 1% in every region during the window.", "1 — Minor verified impact: at least 1% but less than 5% of in-scope production requests fail in one region for 10 minutes, with no other region at 1% or higher.", "2 — Significant regional impact: at least 5% but less than 25% of in-scope production requests fail in exactly one region for at least 10 minutes; escalate to the owning service team.", "3 — Major broad impact: at least 25% of in-scope production requests fail in any region for at least 10 minutes, or at least 5% fail in two or more regions for that duration."], "instructions": "Assign the alert an impact level using the ordered rubric. Assess only the stated 10-minute window and apply the runbook’s scope and exceptions.", "type": "score"}}, "state": [{"speaker": "Regional traffic audit", "text": "During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast."}, {"speaker": "Failure-rate audit", "text": "During 14:20–14:30 UTC, exactly 9% of in-scope production customer requests in us-central failed throughout the full interval."}, {"speaker": "Validation note", "text": "The regional request logs were complete for the stated window, and each reported share was calculated over production customer requests observed continuously throughout that interval."}, {"speaker": "Incident coordinator", "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}]}, "method": "c2d", "provenance": {"source_id": "diverse-137", "source_is_synthetic": true, "source_sha256": "6ad980dc01b48352a342f771b5dfe60d2cbfd6e9d0710110e93512959974536f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and preserve the original runbook policy verbatim. They remain bound to the same production-traffic scope, regional assessment, and 14:20–14:30 UTC window. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the us-central failure rate from 9% to 4% without conflicting with the three-region count or other measurements. Neither context contains an answer code, gold label, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual regional-failure relationship; A2 is a permissible universally quantified relation and A3 is a single region-count relation. A3 is factual rather than a policy conclusion. The base and counter assignments are realizable while changing only A3: the base can have two regions between 5% and 25%, while the counter can have only eu-west in that range, with no region at 25% or above in either case. The policy evidence correctly preserves the relevant state-originating scope, exclusions, threshold inclusivity, and escalation binding; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 alone entails the level-3 condition that at least 5% fail in two or more regions for the full 10-minute window. No exclusion of lower outcomes is needed because the ordered rubric classifies this directly as major broad impact.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes at least one qualifying region at 5% or more for 10 minutes. Refuted A3 entails that fewer than two regions qualify, so eu-west is exactly one qualifying region. A2 excludes every region from reaching the competing 25% level-3 threshold. Together these conditions entail level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During 14:20–14:30 UTC, at least 5% of in-scope production customer requests failed in eu-west for the full 10-minute window."}, {"id": "A2", "statement": "During 14:20–14:30 UTC, no region had at least 25% of in-scope production customer requests fail for the full 10-minute window."}, {"id": "A3", "statement": "During 14:20–14:30 UTC, at least two distinct regions each had at least 5% of in-scope production customer requests fail for the full 10-minute window."}], "base_state_json": "[{\"speaker\":\"Regional traffic audit\",\"text\":\"During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast.\"},{\"speaker\":\"Failure-rate audit\",\"text\":\"During 14:20–14:30 UTC, exactly 9% of in-scope production customer requests in us-central failed throughout the full interval.\"},{\"speaker\":\"Validation note\",\"text\":\"The regional request logs were complete for the stated window, and each reported share was calculated over production customer requests observed continuously throughout that interval.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast."}, {"path": ["1", "text"], "text": "During 14:20–14:30 UTC, exactly 9% of in-scope production customer requests in us-central failed throughout the full interval."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}], "rules": [{"justification": "At least two regions each meeting the inclusive 5% threshold for the full 10-minute window is sufficient for major broad impact.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Eu-west supplies at least one region meeting 5% for 10 minutes; refutation of A3 limits the count of such regions to one, and A2 excludes the competing level-3 condition of any region reaching 25%.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast.", "negative_left": "During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast.", "negative_right": "During 14:20–14:30 UTC, exactly 4% of in-scope production customer requests in us-central failed throughout the full interval.", "right": "During 14:20–14:30 UTC, exactly 9% of in-scope production customer requests in us-central failed throughout the full interval."}, "verifier_independent_model": false}, "family": "scale-diverse-137-004", "id": "scale-diverse-137-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified production impact: only excluded synthetic or beta traffic is affected, or in-scope production failures remain below 1% in every region during the window.", "1 — Minor verified impact: at least 1% but less than 5% of in-scope production requests fail in one region for 10 minutes, with no other region at 1% or higher.", "2 — Significant regional impact: at least 5% but less than 25% of in-scope production requests fail in exactly one region for at least 10 minutes; escalate to the owning service team.", "3 — Major broad impact: at least 25% of in-scope production requests fail in any region for at least 10 minutes, or at least 5% fail in two or more regions for that duration."], "instructions": "Assign the alert an impact level using the ordered rubric. Assess only the stated 10-minute window and apply the runbook’s scope and exceptions.", "type": "score"}}, "state": [{"speaker": "Regional traffic audit", "text": "During 14:20–14:30 UTC, in-scope production customer requests occurred in exactly three distinct regions—eu-west, us-central, and ap-southeast—and the shares failing throughout the full interval were exactly 8% in eu-west and 3% in ap-southeast."}, {"speaker": "Failure-rate audit", "text": "During 14:20–14:30 UTC, exactly 4% of in-scope production customer requests in us-central failed throughout the full interval."}, {"speaker": "Validation note", "text": "The regional request logs were complete for the stated window, and each reported share was calculated over production customer requests observed continuously throughout that interval."}, {"speaker": "Incident coordinator", "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}]}, "method": "c2d", "provenance": {"source_id": "diverse-137", "source_is_synthetic": true, "source_sha256": "6ad980dc01b48352a342f771b5dfe60d2cbfd6e9d0710110e93512959974536f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Checkout API alert, the 14-minute interval beginning at 14:20 UTC, customer-facing failures, purchase completion, and the applicable runbook policy; the unchanged questions preserve all scoring criteria and instructions. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only Harbor's availability, and the resulting state—both members of the complete candidate set being unavailable—is coherent with the remaining assertions. Neither context includes a score, answer code, proposition identifier, rule table, output instruction, or explicit classifier label; the major-incident language is the preserved governing runbook policy rather than answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom concerns the factual availability of an alternative purchase procedure rather than policy. The base and counter assignments differ only on that focus and are both realizable: the service owner's determination can remain accurate while changing consistently with whether a workaround exists. Policy evidence preserves the only substantive state-originated runbook rule needed for interpretation; criteria and instructions in the questions object need not and must not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than 20% of requests fail and the ordinary purchase path is blocked, but an alternative completion procedure exists. Together with the accurate service-owner determination, this establishes a confirmed workaround. That satisfies level 2's degraded-purchases-with-workaround branch and excludes level 3, which requires no workaround.", "rule_index": 0, "sound": true}, {"reason": "More than 20% of customer-facing requests fail, the ordinary purchase path is blocked, and refutation of the existential workaround atom entails that no alternative purchase-completion procedure exists. These conditions are sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, more than 20% of customer-facing Checkout API requests failed."}, {"id": "a2", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Checkout API failures blocked the ordinary purchase-completion path for affected customers."}, {"id": "a3", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, at least one alternative procedure was available to affected customers for completing purchases despite the Checkout API failures."}, {"id": "a4", "statement": "The Checkout service owner's determination of whether any alternative purchase-completion procedure was available to affected customers during the 14-minute Checkout API alert interval beginning at 14:20 UTC was accurate."}], "base_state_json": "\"Field note: Monitoring measured a 42% failure rate among customer-facing Checkout API requests during the 14-minute alert interval beginning at 14:20 UTC. For affected customers, those failures prevented the ordinary purchase-completion path from completing during the same interval. For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor available to affected customers and procedure Juniper unavailable. For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate. The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor available to affected customers and procedure Juniper unavailable."}, {"path": [], "text": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate."}], "policy_evidence": [{"path": [], "text": "The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation."}], "rules": [{"justification": "The ordinary purchase path is blocked, but an alternative purchase-completion procedure is available. Because the service owner's workaround-availability determination is accurate, customers have a confirmed workaround, satisfying level 2 and excluding level 3's no-workaround requirement.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "More than 20% of customer-facing requests fail, the ordinary purchase-completion path is blocked, and refutation of the existential workaround atom establishes that no alternative purchase-completion procedure is available. These facts satisfy level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor available to affected customers and procedure Juniper unavailable.", "negative_left": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor unavailable to affected customers and procedure Juniper unavailable.", "negative_right": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate.", "right": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate."}, "verifier_independent_model": false}, "family": "scale-diverse-138-004", "id": "scale-diverse-138-004-base", "input": {"questions": {"decision": {"criteria": ["0 — No incident: the alert is false or expected, customer-facing error rates remain at baseline, and completed purchases are unaffected.", "1 — Minor operational issue: a real fault exists, but it is confined to internal processing or affects fewer than 5% of customer requests without blocking purchases.", "2 — Moderate user impact: 5% to 20% of customer requests fail, or purchases are degraded but customers have a confirmed workaround; service-team response is required without major-incident escalation.", "3 — Major user impact: more than 20% of customer-facing requests fail and purchases are blocked without a workaround; notify the incident coordinator and immediately escalate to the responsible service or dependency team."], "instructions": "Assess the alert's user-impact level using the evidence and runbook. Select exactly one ordered level, resolving references such as “the former” and “the latter” to the named services.", "type": "score"}}, "state": "Field note: Monitoring measured a 42% failure rate among customer-facing Checkout API requests during the 14-minute alert interval beginning at 14:20 UTC. For affected customers, those failures prevented the ordinary purchase-completion path from completing during the same interval. For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor available to affected customers and procedure Juniper unavailable. For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate. The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation."}, "method": "c2d", "provenance": {"source_id": "diverse-138", "source_is_synthetic": true, "source_sha256": "d4407ec5f95dddf32844f514d87988ba4151765246a49398f43049b05769c0c4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Checkout API alert, the 14-minute interval beginning at 14:20 UTC, customer-facing failures, purchase completion, and the applicable runbook policy; the unchanged questions preserve all scoring criteria and instructions. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only Harbor's availability, and the resulting state—both members of the complete candidate set being unavailable—is coherent with the remaining assertions. Neither context includes a score, answer code, proposition identifier, rule table, output instruction, or explicit classifier label; the major-incident language is the preserved governing runbook policy rather than answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom concerns the factual availability of an alternative purchase procedure rather than policy. The base and counter assignments differ only on that focus and are both realizable: the service owner's determination can remain accurate while changing consistently with whether a workaround exists. Policy evidence preserves the only substantive state-originated runbook rule needed for interpretation; criteria and instructions in the questions object need not and must not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than 20% of requests fail and the ordinary purchase path is blocked, but an alternative completion procedure exists. Together with the accurate service-owner determination, this establishes a confirmed workaround. That satisfies level 2's degraded-purchases-with-workaround branch and excludes level 3, which requires no workaround.", "rule_index": 0, "sound": true}, {"reason": "More than 20% of customer-facing requests fail, the ordinary purchase path is blocked, and refutation of the existential workaround atom entails that no alternative purchase-completion procedure exists. These conditions are sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, more than 20% of customer-facing Checkout API requests failed."}, {"id": "a2", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Checkout API failures blocked the ordinary purchase-completion path for affected customers."}, {"id": "a3", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, at least one alternative procedure was available to affected customers for completing purchases despite the Checkout API failures."}, {"id": "a4", "statement": "The Checkout service owner's determination of whether any alternative purchase-completion procedure was available to affected customers during the 14-minute Checkout API alert interval beginning at 14:20 UTC was accurate."}], "base_state_json": "\"Field note: Monitoring measured a 42% failure rate among customer-facing Checkout API requests during the 14-minute alert interval beginning at 14:20 UTC. For affected customers, those failures prevented the ordinary purchase-completion path from completing during the same interval. For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor available to affected customers and procedure Juniper unavailable. For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate. The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor available to affected customers and procedure Juniper unavailable."}, {"path": [], "text": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate."}], "policy_evidence": [{"path": [], "text": "The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation."}], "rules": [{"justification": "The ordinary purchase path is blocked, but an alternative purchase-completion procedure is available. Because the service owner's workaround-availability determination is accurate, customers have a confirmed workaround, satisfying level 2 and excluding level 3's no-workaround requirement.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "More than 20% of customer-facing requests fail, the ordinary purchase-completion path is blocked, and refutation of the existential workaround atom establishes that no alternative purchase-completion procedure is available. These facts satisfy level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor available to affected customers and procedure Juniper unavailable.", "negative_left": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor unavailable to affected customers and procedure Juniper unavailable.", "negative_right": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate.", "right": "For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate."}, "verifier_independent_model": false}, "family": "scale-diverse-138-004", "id": "scale-diverse-138-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No incident: the alert is false or expected, customer-facing error rates remain at baseline, and completed purchases are unaffected.", "1 — Minor operational issue: a real fault exists, but it is confined to internal processing or affects fewer than 5% of customer requests without blocking purchases.", "2 — Moderate user impact: 5% to 20% of customer requests fail, or purchases are degraded but customers have a confirmed workaround; service-team response is required without major-incident escalation.", "3 — Major user impact: more than 20% of customer-facing requests fail and purchases are blocked without a workaround; notify the incident coordinator and immediately escalate to the responsible service or dependency team."], "instructions": "Assess the alert's user-impact level using the evidence and runbook. Select exactly one ordered level, resolving references such as “the former” and “the latter” to the named services.", "type": "score"}}, "state": "Field note: Monitoring measured a 42% failure rate among customer-facing Checkout API requests during the 14-minute alert interval beginning at 14:20 UTC. For affected customers, those failures prevented the ordinary purchase-completion path from completing during the same interval. For the 14-minute Checkout API alert interval beginning at 14:20 UTC, the Checkout service owner marked procedure Harbor unavailable to affected customers and procedure Juniper unavailable. For the 14-minute Checkout API alert interval beginning at 14:20 UTC, Harbor and Juniper were the complete set of candidate procedures for completing purchases despite the Checkout API failures, and the Checkout service owner's availability markings were accurate. The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation."}, "method": "c2d", "provenance": {"source_id": "diverse-138", "source_is_synthetic": true, "source_sha256": "d4407ec5f95dddf32844f514d87988ba4151765246a49398f43049b05769c0c4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing supersession, precedence, threshold, and transformation requirements without adding exceptions or defaults. The batch, newest-report path, readiness-disposition scope, and timestamps remain fixed. The two evidence spans are complete factual sentences. The counterfactual changes only the report’s SHA-256 value, creating a coherent mismatch with the unchanged immutable batch hash rather than a contradictory duplicate measurement. Neither context contains a gold disposition, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Submission coordinator\",\"text\":\"The readiness log for batch C-184 shows an earlier preliminary entry at 13:40. It confirms that the report timestamped 14:20 is the newest report submitted for the batch’s readiness disposition.\"},{\"speaker\":\"Validation operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 is marked completed.\"},{\"speaker\":\"Integrity recorder\",\"text\":\"The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49.\"},{\"speaker\":\"Batch custodian\",\"text\":\"The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49.\"},{\"speaker\":\"Quality analyst\",\"text\":\"The newest 14:20 report measures required-field completeness at exactly 98.0%, the duplicate rate at 0.9%, and the schema-error rate at 0.0%.\"},{\"speaker\":\"Transformation auditor\",\"text\":\"The audit marks every transformation used to generate the newest 14:20 report for batch C-184 as supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}, {"path": ["3", "text"], "text": "The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49.", "negative_left": "The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 4d8c2a719f30e6b45a12d7c893eb041f6c25a809de73b14f90c8625a1e47bd36.", "negative_right": "The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49.", "right": "The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}, "verifier_independent_model": false}, "family": "scale-diverse-139-001", "id": "scale-diverse-139-001-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Submission coordinator", "text": "The readiness log for batch C-184 shows an earlier preliminary entry at 13:40. It confirms that the report timestamped 14:20 is the newest report submitted for the batch’s readiness disposition."}, {"speaker": "Validation operator", "text": "The newest 14:20 report submitted for batch C-184 is marked completed."}, {"speaker": "Integrity recorder", "text": "The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}, {"speaker": "Batch custodian", "text": "The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}, {"speaker": "Quality analyst", "text": "The newest 14:20 report measures required-field completeness at exactly 98.0%, the duplicate rate at 0.9%, and the schema-error rate at 0.0%."}, {"speaker": "Transformation auditor", "text": "The audit marks every transformation used to generate the newest 14:20 report for batch C-184 as supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing supersession, precedence, threshold, and transformation requirements without adding exceptions or defaults. The batch, newest-report path, readiness-disposition scope, and timestamps remain fixed. The two evidence spans are complete factual sentences. The counterfactual changes only the report’s SHA-256 value, creating a coherent mismatch with the unchanged immutable batch hash rather than a contradictory duplicate measurement. Neither context contains a gold disposition, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Submission coordinator\",\"text\":\"The readiness log for batch C-184 shows an earlier preliminary entry at 13:40. It confirms that the report timestamped 14:20 is the newest report submitted for the batch’s readiness disposition.\"},{\"speaker\":\"Validation operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 is marked completed.\"},{\"speaker\":\"Integrity recorder\",\"text\":\"The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49.\"},{\"speaker\":\"Batch custodian\",\"text\":\"The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49.\"},{\"speaker\":\"Quality analyst\",\"text\":\"The newest 14:20 report measures required-field completeness at exactly 98.0%, the duplicate rate at 0.9%, and the schema-error rate at 0.0%.\"},{\"speaker\":\"Transformation auditor\",\"text\":\"The audit marks every transformation used to generate the newest 14:20 report for batch C-184 as supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}, {"path": ["3", "text"], "text": "The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49.", "negative_left": "The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 4d8c2a719f30e6b45a12d7c893eb041f6c25a809de73b14f90c8625a1e47bd36.", "negative_right": "The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49.", "right": "The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}, "verifier_independent_model": false}, "family": "scale-diverse-139-001", "id": "scale-diverse-139-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Submission coordinator", "text": "The readiness log for batch C-184 shows an earlier preliminary entry at 13:40. It confirms that the report timestamped 14:20 is the newest report submitted for the batch’s readiness disposition."}, {"speaker": "Validation operator", "text": "The newest 14:20 report submitted for batch C-184 is marked completed."}, {"speaker": "Integrity recorder", "text": "The SHA-256 field in the newest report timestamped 14:20 and submitted for batch C-184 contains 4d8c2a719f30e6b45a12d7c893eb041f6c25a809de73b14f90c8625a1e47bd36."}, {"speaker": "Batch custodian", "text": "The immutable SHA-256 value stored for batch C-184 is 7b21d18f3c946ea0528d4f1a6b9c03e74125df8a60bc37e914fa62d0ce538b49."}, {"speaker": "Quality analyst", "text": "The newest 14:20 report measures required-field completeness at exactly 98.0%, the duplicate rate at 0.9%, and the schema-error rate at 0.0%."}, {"speaker": "Transformation auditor", "text": "The audit marks every transformation used to generate the newest 14:20 report for batch C-184 as supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full disposition criteria and precedence rules, while both contexts retain the original-state governing policy without modification. The batch, newest-report path, and 14:20 temporal binding are preserved; changing the observed full hash value is permissible. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the report’s recorded hash so that it differs from the immutable batch hash, rather than asserting two conflicting values for the same field. Neither context contains an answer label, code, proposition identifier, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Records custodian\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51.\"},{\"speaker\":\"Submission monitor\",\"text\":\"The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51.\"},{\"speaker\":\"Readiness coordinator\",\"text\":\"During the operational handoff, the submission register identifies the report timestamped 14:20 as the newest report submitted for the readiness disposition of batch C-184. The earlier worksheet remains archived as historical material.\"},{\"speaker\":\"Validation analyst\",\"text\":\"The newest 14:20 report is completed. It measures required-field completeness at exactly 98.0%, the duplicate rate at 0.9%, and the schema-error rate at 0.0%.\"},{\"speaker\":\"Transformation auditor\",\"text\":\"Every transformation used to generate the newest 14:20 report for batch C-184 is supported. The audit entry covers each conversion and validation step used in its production.\"},{\"speaker\":\"Operations lead\",\"text\":\"The readiness packet is queued for disposition review, with the report, submission register, and transformation audit attached as the observed evidence.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}, {"path": ["1", "text"], "text": "The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51.", "negative_left": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51.", "negative_right": "The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 9b6e1d74a3c8502f67d49e18bc05a3264f7d91ea583c6b20d4a8f15e72c39b60.", "right": "The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}, "verifier_independent_model": false}, "family": "scale-diverse-139-003", "id": "scale-diverse-139-003-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Records custodian", "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}, {"speaker": "Submission monitor", "text": "The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}, {"speaker": "Readiness coordinator", "text": "During the operational handoff, the submission register identifies the report timestamped 14:20 as the newest report submitted for the readiness disposition of batch C-184. The earlier worksheet remains archived as historical material."}, {"speaker": "Validation analyst", "text": "The newest 14:20 report is completed. It measures required-field completeness at exactly 98.0%, the duplicate rate at 0.9%, and the schema-error rate at 0.0%."}, {"speaker": "Transformation auditor", "text": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported. The audit entry covers each conversion and validation step used in its production."}, {"speaker": "Operations lead", "text": "The readiness packet is queued for disposition review, with the report, submission register, and transformation audit attached as the observed evidence."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full disposition criteria and precedence rules, while both contexts retain the original-state governing policy without modification. The batch, newest-report path, and 14:20 temporal binding are preserved; changing the observed full hash value is permissible. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the report’s recorded hash so that it differs from the immutable batch hash, rather than asserting two conflicting values for the same field. Neither context contains an answer label, code, proposition identifier, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Records custodian\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51.\"},{\"speaker\":\"Submission monitor\",\"text\":\"The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51.\"},{\"speaker\":\"Readiness coordinator\",\"text\":\"During the operational handoff, the submission register identifies the report timestamped 14:20 as the newest report submitted for the readiness disposition of batch C-184. The earlier worksheet remains archived as historical material.\"},{\"speaker\":\"Validation analyst\",\"text\":\"The newest 14:20 report is completed. It measures required-field completeness at exactly 98.0%, the duplicate rate at 0.9%, and the schema-error rate at 0.0%.\"},{\"speaker\":\"Transformation auditor\",\"text\":\"Every transformation used to generate the newest 14:20 report for batch C-184 is supported. The audit entry covers each conversion and validation step used in its production.\"},{\"speaker\":\"Operations lead\",\"text\":\"The readiness packet is queued for disposition review, with the report, submission register, and transformation audit attached as the observed evidence.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}, {"path": ["1", "text"], "text": "The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51.", "negative_left": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51.", "negative_right": "The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 9b6e1d74a3c8502f67d49e18bc05a3264f7d91ea583c6b20d4a8f15e72c39b60.", "right": "The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}, "verifier_independent_model": false}, "family": "scale-diverse-139-003", "id": "scale-diverse-139-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Records custodian", "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2e6b1d4a8058e7f926c3d14b609af521ec84d76b305e19a4f867c2db51."}, {"speaker": "Submission monitor", "text": "The SHA-256 field in the newest 14:20 report submitted for batch C-184 records 9b6e1d74a3c8502f67d49e18bc05a3264f7d91ea583c6b20d4a8f15e72c39b60."}, {"speaker": "Readiness coordinator", "text": "During the operational handoff, the submission register identifies the report timestamped 14:20 as the newest report submitted for the readiness disposition of batch C-184. The earlier worksheet remains archived as historical material."}, {"speaker": "Validation analyst", "text": "The newest 14:20 report is completed. It measures required-field completeness at exactly 98.0%, the duplicate rate at 0.9%, and the schema-error rate at 0.0%."}, {"speaker": "Transformation auditor", "text": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported. The audit entry covers each conversion and validation step used in its production."}, {"speaker": "Operations lead", "text": "The readiness packet is queued for disposition review, with the report, submission register, and transformation audit attached as the observed evidence."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full disposition criteria and precedence, while the original-state policy statement is also retained in both contexts. The batch, report time, and decision scope remain bound to C-184 and the 14:20 report; the two evidence spans are complete factual sentences. The counterfactual coherently changes the immutable-register hash so that it differs from the report-recorded hash, which is a meaningful mismatch rather than a contradictory duplicate measurement. Neither context includes an answer code, gold disposition, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Readiness coordinator\",\"text\":\"For the readiness disposition of batch C-184, staff logged the following submission and immutable-record details.\"},{\"speaker\":\"Validation register\",\"text\":\"The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef.\"},{\"speaker\":\"Immutable batch register\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"The 14:20 report is completed. It measures required-field completeness at exactly 98.0%, duplicate rate at 0.9%, and schema-error rate at 0.0%. The transformation audit confirms that every transformation used to generate the report is supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}, {"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef.", "negative_left": "The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef.", "negative_right": "The immutable SHA-256 hash of batch C-184 is fedcba9876543210fedcba9876543210fedcba9876543210fedcba9876543210.", "right": "The immutable SHA-256 hash of batch C-184 is 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}, "verifier_independent_model": false}, "family": "scale-diverse-139-004", "id": "scale-diverse-139-004-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Readiness coordinator", "text": "For the readiness disposition of batch C-184, staff logged the following submission and immutable-record details."}, {"speaker": "Validation register", "text": "The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}, {"speaker": "Immutable batch register", "text": "The immutable SHA-256 hash of batch C-184 is 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}, {"speaker": "Audit reviewer", "text": "The 14:20 report is completed. It measures required-field completeness at exactly 98.0%, duplicate rate at 0.9%, and schema-error rate at 0.0%. The transformation audit confirms that every transformation used to generate the report is supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full disposition criteria and precedence, while the original-state policy statement is also retained in both contexts. The batch, report time, and decision scope remain bound to C-184 and the 14:20 report; the two evidence spans are complete factual sentences. The counterfactual coherently changes the immutable-register hash so that it differs from the report-recorded hash, which is a meaningful mismatch rather than a contradictory duplicate measurement. Neither context includes an answer code, gold disposition, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Readiness coordinator\",\"text\":\"For the readiness disposition of batch C-184, staff logged the following submission and immutable-record details.\"},{\"speaker\":\"Validation register\",\"text\":\"The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef.\"},{\"speaker\":\"Immutable batch register\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"The 14:20 report is completed. It measures required-field completeness at exactly 98.0%, duplicate rate at 0.9%, and schema-error rate at 0.0%. The transformation audit confirms that every transformation used to generate the report is supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}, {"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef.", "negative_left": "The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef.", "negative_right": "The immutable SHA-256 hash of batch C-184 is fedcba9876543210fedcba9876543210fedcba9876543210fedcba9876543210.", "right": "The immutable SHA-256 hash of batch C-184 is 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}, "verifier_independent_model": false}, "family": "scale-diverse-139-004", "id": "scale-diverse-139-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Readiness coordinator", "text": "For the readiness disposition of batch C-184, staff logged the following submission and immutable-record details."}, {"speaker": "Validation register", "text": "The report timestamped 14:20 UTC on 18 May 2031 is the newest report submitted for batch C-184 and records SHA-256 hash 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef."}, {"speaker": "Immutable batch register", "text": "The immutable SHA-256 hash of batch C-184 is fedcba9876543210fedcba9876543210fedcba9876543210fedcba9876543210."}, {"speaker": "Audit reviewer", "text": "The 14:20 report is completed. It measures required-field completeness at exactly 98.0%, duplicate rate at 0.9%, and schema-error rate at 0.0%. The transformation audit confirms that every transformation used to generate the report is supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same acceptance policy, request, batch entity, measurement definition, and relevant timeline. The two focus spans are complete factual sentences. The sole material change—from 49,800 to 49,700 valid instances out of 50,000—is coherent and creates no duplicate contradictory measurement; the earlier statement that reports had not yet arrived is also consistent with their later arrival. Neither context includes a gold answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy or final classification. The focus atom is the factual completeness-threshold relation. The base and counter assignments are jointly realizable: a supplied post-transformation result can either meet or fall below 99.5% while the duplicate rate and reconciliation remain supported. The policy evidence preserves the substantive state-originating thresholds, reconciliation requirement, and mandatory-report status; question-originating instructions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction demonstrates both numerical thresholds and both mandatory evidence requirements. Under the unchanged question, satisfying every acceptance condition is sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 entails that post-transformation completeness is below 99.5%. An unmet threshold is sufficient for a false decision regardless of the supported remaining conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The post-transformation critical-field completeness measured for customer batch CUS-184 is at least 99.5%."}, {"id": "a2", "statement": "A post-transformation critical-field completeness result for customer batch CUS-184 is supplied in the validation package."}, {"id": "a3", "statement": "The duplicate rate measured for customer batch CUS-184 is below 0.5%."}, {"id": "a4", "statement": "An input-to-output row reconciliation for customer batch CUS-184 is supplied in the validation package."}], "base_state_json": "{\"context\":\"Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied.\",\"evidence\":[\"At 09:00 UTC on 14 August 2026, Mira noted that the post-transformation completeness result and row reconciliation had not yet arrived.\",\"Application administrator confirms both missing reports are mandatory acceptance evidence.\",\"At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values.\",\"At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,800 required field instances.\",\"At 09:35 UTC, the completed post-transformation critical-field completeness result was indexed in the validation package.\",\"At 09:40 UTC, the duplicate report recorded a 0.3% duplicate rate for CUS-184.\",\"At 09:45 UTC, a signed input-to-output row reconciliation for CUS-184 was added to the validation package; its transformed, rejected, and loaded row counts balanced to the input total.\"],\"request\":\"Should Mira mark CUS-184 ready for acceptance now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values."}, {"path": ["evidence", "3"], "text": "At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,800 required field instances."}], "policy_evidence": [{"path": ["context"], "text": "Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied."}, {"path": ["evidence", "4"], "text": "Application administrator confirms both missing reports are mandatory acceptance evidence."}], "rules": [{"justification": "Every acceptance threshold is met, the mandatory post-transformation completeness result is supplied, and the required input-to-output reconciliation is supplied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The post-transformation critical-field completeness threshold is explicitly unmet, so the batch cannot be marked ready even though the other threshold and mandatory evidence requirements are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values.", "negative_left": "At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values.", "negative_right": "At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,700 required field instances.", "right": "At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,800 required field instances."}, "verifier_independent_model": false}, "family": "scale-diverse-141-001", "id": "scale-diverse-141-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark the batch ready because at least one threshold is unmet or required acceptance evidence is missing.", "true": "Yes — mark the batch ready because all acceptance thresholds and mandatory evidence requirements are demonstrated."}, "instructions": "Answer yes only if every stated acceptance condition is supported by the supplied evidence. If required evidence is missing, answer no even when available metrics appear to pass.", "type": "noul"}}, "state": {"context": "Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied.", "evidence": ["At 09:00 UTC on 14 August 2026, Mira noted that the post-transformation completeness result and row reconciliation had not yet arrived.", "Application administrator confirms both missing reports are mandatory acceptance evidence.", "At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values.", "At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,800 required field instances.", "At 09:35 UTC, the completed post-transformation critical-field completeness result was indexed in the validation package.", "At 09:40 UTC, the duplicate report recorded a 0.3% duplicate rate for CUS-184.", "At 09:45 UTC, a signed input-to-output row reconciliation for CUS-184 was added to the validation package; its transformed, rejected, and loaded row counts balanced to the input total."], "request": "Should Mira mark CUS-184 ready for acceptance now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-141", "source_is_synthetic": true, "source_sha256": "4338a888e7923294000aeef40f5b4f1e474f07015e64d45afe6ac50e7d87dce7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same acceptance policy, request, batch entity, measurement definition, and relevant timeline. The two focus spans are complete factual sentences. The sole material change—from 49,800 to 49,700 valid instances out of 50,000—is coherent and creates no duplicate contradictory measurement; the earlier statement that reports had not yet arrived is also consistent with their later arrival. Neither context includes a gold answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy or final classification. The focus atom is the factual completeness-threshold relation. The base and counter assignments are jointly realizable: a supplied post-transformation result can either meet or fall below 99.5% while the duplicate rate and reconciliation remain supported. The policy evidence preserves the substantive state-originating thresholds, reconciliation requirement, and mandatory-report status; question-originating instructions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction demonstrates both numerical thresholds and both mandatory evidence requirements. Under the unchanged question, satisfying every acceptance condition is sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 entails that post-transformation completeness is below 99.5%. An unmet threshold is sufficient for a false decision regardless of the supported remaining conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The post-transformation critical-field completeness measured for customer batch CUS-184 is at least 99.5%."}, {"id": "a2", "statement": "A post-transformation critical-field completeness result for customer batch CUS-184 is supplied in the validation package."}, {"id": "a3", "statement": "The duplicate rate measured for customer batch CUS-184 is below 0.5%."}, {"id": "a4", "statement": "An input-to-output row reconciliation for customer batch CUS-184 is supplied in the validation package."}], "base_state_json": "{\"context\":\"Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied.\",\"evidence\":[\"At 09:00 UTC on 14 August 2026, Mira noted that the post-transformation completeness result and row reconciliation had not yet arrived.\",\"Application administrator confirms both missing reports are mandatory acceptance evidence.\",\"At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values.\",\"At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,800 required field instances.\",\"At 09:35 UTC, the completed post-transformation critical-field completeness result was indexed in the validation package.\",\"At 09:40 UTC, the duplicate report recorded a 0.3% duplicate rate for CUS-184.\",\"At 09:45 UTC, a signed input-to-output row reconciliation for CUS-184 was added to the validation package; its transformed, rejected, and loaded row counts balanced to the input total.\"],\"request\":\"Should Mira mark CUS-184 ready for acceptance now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values."}, {"path": ["evidence", "3"], "text": "At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,800 required field instances."}], "policy_evidence": [{"path": ["context"], "text": "Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied."}, {"path": ["evidence", "4"], "text": "Application administrator confirms both missing reports are mandatory acceptance evidence."}], "rules": [{"justification": "Every acceptance threshold is met, the mandatory post-transformation completeness result is supplied, and the required input-to-output reconciliation is supplied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The post-transformation critical-field completeness threshold is explicitly unmet, so the batch cannot be marked ready even though the other threshold and mandatory evidence requirements are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values.", "negative_left": "At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values.", "negative_right": "At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,700 required field instances.", "right": "At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,800 required field instances."}, "verifier_independent_model": false}, "family": "scale-diverse-141-001", "id": "scale-diverse-141-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark the batch ready because at least one threshold is unmet or required acceptance evidence is missing.", "true": "Yes — mark the batch ready because all acceptance thresholds and mandatory evidence requirements are demonstrated."}, "instructions": "Answer yes only if every stated acceptance condition is supported by the supplied evidence. If required evidence is missing, answer no even when available metrics appear to pass.", "type": "noul"}}, "state": {"context": "Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied.", "evidence": ["At 09:00 UTC on 14 August 2026, Mira noted that the post-transformation completeness result and row reconciliation had not yet arrived.", "Application administrator confirms both missing reports are mandatory acceptance evidence.", "At 09:10 UTC on 14 August 2026, the validator fixed the measurement universe for customer batch CUS-184's post-transformation critical-field completeness at 50,000 required field instances, with completeness calculated as the percentage of those instances containing valid values.", "At 09:25 UTC on 14 August 2026, the completed validation count for that measurement recorded valid values in 49,700 required field instances.", "At 09:35 UTC, the completed post-transformation critical-field completeness result was indexed in the validation package.", "At 09:40 UTC, the duplicate report recorded a 0.3% duplicate rate for CUS-184.", "At 09:45 UTC, a signed input-to-output row reconciliation for CUS-184 was added to the validation package; its transformed, rejected, and loaded row counts balanced to the input total."], "request": "Should Mira mark CUS-184 ready for acceptance now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-141", "source_is_synthetic": true, "source_sha256": "4338a888e7923294000aeef40f5b4f1e474f07015e64d45afe6ac50e7d87dce7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both generated contexts retain the governing acceptance policy and the original administrator statement that the two reports are mandatory; the unchanged questions object preserves the decision instructions and criteria. The batch, decision request, evidence roles, record identifiers, and times remain bound consistently. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the completeness measurement from 99.7% to 99.2%, without creating a duplicate or contradictory measurement within that context. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy or final classification. The focus atom is the factual completeness-threshold relation. The base and counter assignments are jointly realizable: a supplied post-transformation result can either meet or fall below 99.5% while the duplicate rate and reconciliation remain supported. The policy evidence preserves the substantive state-originating thresholds, reconciliation requirement, and mandatory-report status; question-originating instructions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction demonstrates both numerical thresholds and both mandatory evidence requirements. Under the unchanged question, satisfying every acceptance condition is sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 entails that post-transformation completeness is below 99.5%. An unmet threshold is sufficient for a false decision regardless of the supported remaining conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The post-transformation critical-field completeness measured for customer batch CUS-184 is at least 99.5%."}, {"id": "a2", "statement": "A post-transformation critical-field completeness result for customer batch CUS-184 is supplied in the validation package."}, {"id": "a3", "statement": "The duplicate rate measured for customer batch CUS-184 is below 0.5%."}, {"id": "a4", "statement": "An input-to-output row reconciliation for customer batch CUS-184 is supplied in the validation package."}], "base_state_json": "{\"context\":\"Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied.\",\"handoff_note\":\"Mira is completing the operational handoff using the two timed records retained together in the review file. An intake note predating those records preserved the administrator’s original wording below.\",\"policy_record\":\"Application administrator confirms both missing reports are mandatory acceptance evidence.\",\"evidence\":[\"At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.7%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418.\",\"The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184.\"],\"request\":\"Should Mira mark CUS-184 ready for acceptance now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.7%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418."}, {"path": ["evidence", "1"], "text": "The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184."}], "policy_evidence": [{"path": ["context"], "text": "Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied."}, {"path": ["evidence", "4"], "text": "Application administrator confirms both missing reports are mandatory acceptance evidence."}], "rules": [{"justification": "Every acceptance threshold is met, the mandatory post-transformation completeness result is supplied, and the required input-to-output reconciliation is supplied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The post-transformation critical-field completeness threshold is explicitly unmet, so the batch cannot be marked ready even though the other threshold and mandatory evidence requirements are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.7%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418.", "negative_left": "At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.2%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418.", "negative_right": "The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184.", "right": "The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184."}, "verifier_independent_model": false}, "family": "scale-diverse-141-003", "id": "scale-diverse-141-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark the batch ready because at least one threshold is unmet or required acceptance evidence is missing.", "true": "Yes — mark the batch ready because all acceptance thresholds and mandatory evidence requirements are demonstrated."}, "instructions": "Answer yes only if every stated acceptance condition is supported by the supplied evidence. If required evidence is missing, answer no even when available metrics appear to pass.", "type": "noul"}}, "state": {"context": "Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied.", "evidence": ["At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.7%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418.", "The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184."], "handoff_note": "Mira is completing the operational handoff using the two timed records retained together in the review file. An intake note predating those records preserved the administrator’s original wording below.", "policy_record": "Application administrator confirms both missing reports are mandatory acceptance evidence.", "request": "Should Mira mark CUS-184 ready for acceptance now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-141", "source_is_synthetic": true, "source_sha256": "4338a888e7923294000aeef40f5b4f1e474f07015e64d45afe6ac50e7d87dce7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both generated contexts retain the governing acceptance policy and the original administrator statement that the two reports are mandatory; the unchanged questions object preserves the decision instructions and criteria. The batch, decision request, evidence roles, record identifiers, and times remain bound consistently. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the completeness measurement from 99.7% to 99.2%, without creating a duplicate or contradictory measurement within that context. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy or final classification. The focus atom is the factual completeness-threshold relation. The base and counter assignments are jointly realizable: a supplied post-transformation result can either meet or fall below 99.5% while the duplicate rate and reconciliation remain supported. The policy evidence preserves the substantive state-originating thresholds, reconciliation requirement, and mandatory-report status; question-originating instructions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction demonstrates both numerical thresholds and both mandatory evidence requirements. Under the unchanged question, satisfying every acceptance condition is sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 entails that post-transformation completeness is below 99.5%. An unmet threshold is sufficient for a false decision regardless of the supported remaining conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The post-transformation critical-field completeness measured for customer batch CUS-184 is at least 99.5%."}, {"id": "a2", "statement": "A post-transformation critical-field completeness result for customer batch CUS-184 is supplied in the validation package."}, {"id": "a3", "statement": "The duplicate rate measured for customer batch CUS-184 is below 0.5%."}, {"id": "a4", "statement": "An input-to-output row reconciliation for customer batch CUS-184 is supplied in the validation package."}], "base_state_json": "{\"context\":\"Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied.\",\"handoff_note\":\"Mira is completing the operational handoff using the two timed records retained together in the review file. An intake note predating those records preserved the administrator’s original wording below.\",\"policy_record\":\"Application administrator confirms both missing reports are mandatory acceptance evidence.\",\"evidence\":[\"At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.7%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418.\",\"The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184.\"],\"request\":\"Should Mira mark CUS-184 ready for acceptance now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.7%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418."}, {"path": ["evidence", "1"], "text": "The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184."}], "policy_evidence": [{"path": ["context"], "text": "Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied."}, {"path": ["evidence", "4"], "text": "Application administrator confirms both missing reports are mandatory acceptance evidence."}], "rules": [{"justification": "Every acceptance threshold is met, the mandatory post-transformation completeness result is supplied, and the required input-to-output reconciliation is supplied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The post-transformation critical-field completeness threshold is explicitly unmet, so the batch cannot be marked ready even though the other threshold and mandatory evidence requirements are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.7%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418.", "negative_left": "At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.2%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418.", "negative_right": "The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184.", "right": "The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184."}, "verifier_independent_model": false}, "family": "scale-diverse-141-003", "id": "scale-diverse-141-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark the batch ready because at least one threshold is unmet or required acceptance evidence is missing.", "true": "Yes — mark the batch ready because all acceptance thresholds and mandatory evidence requirements are demonstrated."}, "instructions": "Answer yes only if every stated acceptance condition is supported by the supplied evidence. If required evidence is missing, answer no even when available metrics appear to pass.", "type": "noul"}}, "state": {"context": "Data quality analyst Mira is reviewing customer batch CUS-184 for the fictional Northstar CRM. Policy permits acceptance only when post-transformation critical-field completeness is at least 99.5%, duplicate rate is below 0.5%, and an input-to-output row reconciliation is supplied.", "evidence": ["At 14:20 UTC on 8 September 2026, validation-package record V-731 for customer batch CUS-184 listed metric slot M-17 as 99.2%, recorded the batch's duplicate rate as 0.32%, and included input-to-output row reconciliation RR-418.", "The signed operational-handoff manifest issued at 14:25 UTC on 8 September 2026 identifies metric slot M-17 in record V-731 as the measured post-transformation critical-field completeness result for customer batch CUS-184."], "handoff_note": "Mira is completing the operational handoff using the two timed records retained together in the review file. An intake note predating those records preserved the administrator’s original wording below.", "policy_record": "Application administrator confirms both missing reports are mandatory acceptance evidence.", "request": "Should Mira mark CUS-184 ready for acceptance now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-141", "source_is_synthetic": true, "source_sha256": "4338a888e7923294000aeef40f5b4f1e474f07015e64d45afe6ac50e7d87dce7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the zero-duplicate requirement, the 99.95% scoped ship_country completeness threshold, and the history-tag exception, while retaining the same batch, transformed-output path, audit scope, and acceptance question. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only the missing-value count from 73 to 127, with no conflicting duplicate count or measurement elsewhere. Neither context contains an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"At 13:50 UTC on 16 September 2026, the approved customer_id transformation for batch B-142 finished. Its reconciliation report recorded exactly zero duplicate customer_id values in the transformed output.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Audit recorder\",\"text\":\"At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history.\"},{\"speaker\":\"Audit recorder\",\"text\":\"At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 73 records in batch B-142 not tagged import_mode=history had no observed ship_country value.\"},{\"speaker\":\"Release manager\",\"text\":\"At 14:15 UTC, the validation package confirmed that Q7 examined the transformed batch and that no other acceptance requirements were pending. Should batch B-142 be accepted under the stated rubric?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history."}, {"path": ["4", "text"], "text": "At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 73 records in batch B-142 not tagged import_mode=history had no observed ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history.", "negative_left": "At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history.", "negative_right": "At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 127 records in batch B-142 not tagged import_mode=history had no observed ship_country value.", "right": "At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 73 records in batch B-142 not tagged import_mode=history had no observed ship_country value."}, "verifier_independent_model": false}, "family": "scale-diverse-142-001", "id": "scale-diverse-142-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "At 13:50 UTC on 16 September 2026, the approved customer_id transformation for batch B-142 finished. Its reconciliation report recorded exactly zero duplicate customer_id values in the transformed output."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Audit recorder", "text": "At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history."}, {"speaker": "Audit recorder", "text": "At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 73 records in batch B-142 not tagged import_mode=history had no observed ship_country value."}, {"speaker": "Release manager", "text": "At 14:15 UTC, the validation package confirmed that Q7 examined the transformed batch and that no other acceptance requirements were pending. Should batch B-142 be accepted under the stated rubric?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the zero-duplicate requirement, the 99.95% scoped ship_country completeness threshold, and the history-tag exception, while retaining the same batch, transformed-output path, audit scope, and acceptance question. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only the missing-value count from 73 to 127, with no conflicting duplicate count or measurement elsewhere. Neither context contains an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"At 13:50 UTC on 16 September 2026, the approved customer_id transformation for batch B-142 finished. Its reconciliation report recorded exactly zero duplicate customer_id values in the transformed output.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Audit recorder\",\"text\":\"At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history.\"},{\"speaker\":\"Audit recorder\",\"text\":\"At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 73 records in batch B-142 not tagged import_mode=history had no observed ship_country value.\"},{\"speaker\":\"Release manager\",\"text\":\"At 14:15 UTC, the validation package confirmed that Q7 examined the transformed batch and that no other acceptance requirements were pending. Should batch B-142 be accepted under the stated rubric?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history."}, {"path": ["4", "text"], "text": "At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 73 records in batch B-142 not tagged import_mode=history had no observed ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history.", "negative_left": "At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history.", "negative_right": "At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 127 records in batch B-142 not tagged import_mode=history had no observed ship_country value.", "right": "At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 73 records in batch B-142 not tagged import_mode=history had no observed ship_country value."}, "verifier_independent_model": false}, "family": "scale-diverse-142-001", "id": "scale-diverse-142-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "At 13:50 UTC on 16 September 2026, the approved customer_id transformation for batch B-142 finished. Its reconciliation report recorded exactly zero duplicate customer_id values in the transformed output."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Audit recorder", "text": "At 14:00 UTC on 16 September 2026, audit run Q7 counted exactly 200,000 records in batch B-142 that were not tagged import_mode=history."}, {"speaker": "Audit recorder", "text": "At 14:05 UTC on 16 September 2026, audit run Q7 found that exactly 127 records in batch B-142 not tagged import_mode=history had no observed ship_country value."}, {"speaker": "Release manager", "text": "At 14:15 UTC, the validation package confirmed that Q7 examined the transformed batch and that no other acceptance requirements were pending. Should batch B-142 be accepted under the stated rubric?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance policy: zero post-transformation duplicate IDs, at least 99.95% ship_country completeness among non-history records, and the history-record scope exception. They preserve the decision about batch B-142 and consistently bind the evidence to one finalized snapshot. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the observed complete-record count from 239,880 to 239,879 out of 240,000, which is internally coherent and does not conflict with another measurement. Neither context contains a gold answer, output instruction, answer code, proposition ID, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Reconciliation lead\",\"text\":\"The team reconciled batch B-142 against the finalized audit snapshot dated 2026-08-31T23:00:00Z. The snapshot is the controlling evidence for the acceptance review, and all cited checks were completed against that same frozen version.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Audit evidence register\",\"text\":\"In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history.\"},{\"speaker\":\"Audit evidence register\",\"text\":\"In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,880 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value.\"},{\"speaker\":\"Transformation operator\",\"text\":\"The approved customer_id deduplication was run before validation. Its final report recorded zero duplicate customer_id values remaining in B-142 after transformation.\"},{\"speaker\":\"Review chair\",\"text\":\"No other acceptance requirements are under review. Should batch B-142 be accepted under the stated rubric and documented scope exception?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["2", "text"], "text": "In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history."}, {"path": ["3", "text"], "text": "In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,880 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history.", "negative_left": "In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history.", "negative_right": "In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,879 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value.", "right": "In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,880 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value."}, "verifier_independent_model": false}, "family": "scale-diverse-142-002", "id": "scale-diverse-142-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Reconciliation lead", "text": "The team reconciled batch B-142 against the finalized audit snapshot dated 2026-08-31T23:00:00Z. The snapshot is the controlling evidence for the acceptance review, and all cited checks were completed against that same frozen version."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Audit evidence register", "text": "In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history."}, {"speaker": "Audit evidence register", "text": "In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,880 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value."}, {"speaker": "Transformation operator", "text": "The approved customer_id deduplication was run before validation. Its final report recorded zero duplicate customer_id values remaining in B-142 after transformation."}, {"speaker": "Review chair", "text": "No other acceptance requirements are under review. Should batch B-142 be accepted under the stated rubric and documented scope exception?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance policy: zero post-transformation duplicate IDs, at least 99.95% ship_country completeness among non-history records, and the history-record scope exception. They preserve the decision about batch B-142 and consistently bind the evidence to one finalized snapshot. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the observed complete-record count from 239,880 to 239,879 out of 240,000, which is internally coherent and does not conflict with another measurement. Neither context contains a gold answer, output instruction, answer code, proposition ID, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Reconciliation lead\",\"text\":\"The team reconciled batch B-142 against the finalized audit snapshot dated 2026-08-31T23:00:00Z. The snapshot is the controlling evidence for the acceptance review, and all cited checks were completed against that same frozen version.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Audit evidence register\",\"text\":\"In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history.\"},{\"speaker\":\"Audit evidence register\",\"text\":\"In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,880 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value.\"},{\"speaker\":\"Transformation operator\",\"text\":\"The approved customer_id deduplication was run before validation. Its final report recorded zero duplicate customer_id values remaining in B-142 after transformation.\"},{\"speaker\":\"Review chair\",\"text\":\"No other acceptance requirements are under review. Should batch B-142 be accepted under the stated rubric and documented scope exception?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["2", "text"], "text": "In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history."}, {"path": ["3", "text"], "text": "In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,880 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history.", "negative_left": "In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history.", "negative_right": "In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,879 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value.", "right": "In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,880 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value."}, "verifier_independent_model": false}, "family": "scale-diverse-142-002", "id": "scale-diverse-142-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Reconciliation lead", "text": "The team reconciled batch B-142 against the finalized audit snapshot dated 2026-08-31T23:00:00Z. The snapshot is the controlling evidence for the acceptance review, and all cited checks were completed against that same frozen version."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Audit evidence register", "text": "In the finalized 2026-08-31T23:00:00Z audit snapshot, batch B-142 contained exactly 240,000 records not tagged import_mode=history."}, {"speaker": "Audit evidence register", "text": "In the finalized 2026-08-31T23:00:00Z audit snapshot, exactly 239,879 records in batch B-142 were not tagged import_mode=history and had an observed ship_country value."}, {"speaker": "Transformation operator", "text": "The approved customer_id deduplication was run before validation. Its final report recorded zero duplicate customer_id values remaining in B-142 after transformation."}, {"speaker": "Review chair", "text": "No other acceptance requirements are under review. Should batch B-142 be accepted under the stated rubric and documented scope exception?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing zero-duplicate requirement, the 99.95% scoped-completeness threshold, and the history-record exemption; the unchanged questions object preserves the decision instructions and criteria. The batch entity, acceptance-review path, and stated handoff time remain fixed, while the permitted observation changes from four to five missing in-scope values. The two evidence spans are complete factual sentences. The counterfactual introduces no conflicting duplicate count, population count, or other assertion, and neither context states the decision, an answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Validation coordinator\",\"text\":\"At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"At the 2026-08-14T16:00:00Z validation handoff, exactly four records in batch B-142 not tagged import_mode=history lacked an observed ship_country value.\"},{\"speaker\":\"Data engineer\",\"text\":\"The completed transformation report for B-142 records a post-transformation duplicate customer_id count of zero.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Operations lead\",\"text\":\"No remaining validation findings concern the acceptance review. Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied. Should B-142 be accepted?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["0", "text"], "text": "At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history."}, {"path": ["1", "text"], "text": "At the 2026-08-14T16:00:00Z validation handoff, exactly four records in batch B-142 not tagged import_mode=history lacked an observed ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history.", "negative_left": "At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history.", "negative_right": "At the 2026-08-14T16:00:00Z validation handoff, exactly five records in batch B-142 not tagged import_mode=history lacked an observed ship_country value.", "right": "At the 2026-08-14T16:00:00Z validation handoff, exactly four records in batch B-142 not tagged import_mode=history lacked an observed ship_country value."}, "verifier_independent_model": false}, "family": "scale-diverse-142-003", "id": "scale-diverse-142-003-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Validation coordinator", "text": "At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history."}, {"speaker": "Quality reviewer", "text": "At the 2026-08-14T16:00:00Z validation handoff, exactly four records in batch B-142 not tagged import_mode=history lacked an observed ship_country value."}, {"speaker": "Data engineer", "text": "The completed transformation report for B-142 records a post-transformation duplicate customer_id count of zero."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Operations lead", "text": "No remaining validation findings concern the acceptance review. Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied. Should B-142 be accepted?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing zero-duplicate requirement, the 99.95% scoped-completeness threshold, and the history-record exemption; the unchanged questions object preserves the decision instructions and criteria. The batch entity, acceptance-review path, and stated handoff time remain fixed, while the permitted observation changes from four to five missing in-scope values. The two evidence spans are complete factual sentences. The counterfactual introduces no conflicting duplicate count, population count, or other assertion, and neither context states the decision, an answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Validation coordinator\",\"text\":\"At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"At the 2026-08-14T16:00:00Z validation handoff, exactly four records in batch B-142 not tagged import_mode=history lacked an observed ship_country value.\"},{\"speaker\":\"Data engineer\",\"text\":\"The completed transformation report for B-142 records a post-transformation duplicate customer_id count of zero.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Operations lead\",\"text\":\"No remaining validation findings concern the acceptance review. Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied. Should B-142 be accepted?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["0", "text"], "text": "At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history."}, {"path": ["1", "text"], "text": "At the 2026-08-14T16:00:00Z validation handoff, exactly four records in batch B-142 not tagged import_mode=history lacked an observed ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history.", "negative_left": "At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history.", "negative_right": "At the 2026-08-14T16:00:00Z validation handoff, exactly five records in batch B-142 not tagged import_mode=history lacked an observed ship_country value.", "right": "At the 2026-08-14T16:00:00Z validation handoff, exactly four records in batch B-142 not tagged import_mode=history lacked an observed ship_country value."}, "verifier_independent_model": false}, "family": "scale-diverse-142-003", "id": "scale-diverse-142-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Validation coordinator", "text": "At the 2026-08-14T16:00:00Z validation handoff, batch B-142 contained exactly 8,000 records not tagged import_mode=history."}, {"speaker": "Quality reviewer", "text": "At the 2026-08-14T16:00:00Z validation handoff, exactly five records in batch B-142 not tagged import_mode=history lacked an observed ship_country value."}, {"speaker": "Data engineer", "text": "The completed transformation report for B-142 records a post-transformation duplicate customer_id count of zero."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Operations lead", "text": "No remaining validation findings concern the acceptance review. Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied. Should B-142 be accepted?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the zero-duplicate requirement, the 99.95% scoped-completeness requirement, and the history-record exception, while preserving batch B-142 and the post-transformation review scope. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the observed ship_country count from 39,982 to 39,978 out of 40,000 without conflicting duplicate measurements or assertions. Neither context embeds an answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"For release review, the approved deduplication transform was completed on batch B-142 before the 2026-08-31 validation snapshot was frozen. The post-transformation customer_id scan found zero duplicate values.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Snapshot auditor\",\"text\":\"The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history.\"},{\"speaker\":\"Snapshot auditor\",\"text\":\"In the 2026-08-31 validation snapshot, exactly 39,982 of those records have an observed ship_country value.\"},{\"speaker\":\"Release coordinator\",\"text\":\"The snapshot and scan are the signed validation artifacts for this batch. The batch identifier was reconciled across both artifacts, and no later rerun or manual correction is part of the review record.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history."}, {"path": ["4", "text"], "text": "In the 2026-08-31 validation snapshot, exactly 39,982 of those records have an observed ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history.", "negative_left": "The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history.", "negative_right": "In the 2026-08-31 validation snapshot, exactly 39,978 of those records have an observed ship_country value.", "right": "In the 2026-08-31 validation snapshot, exactly 39,982 of those records have an observed ship_country value."}, "verifier_independent_model": false}, "family": "scale-diverse-142-004", "id": "scale-diverse-142-004-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "For release review, the approved deduplication transform was completed on batch B-142 before the 2026-08-31 validation snapshot was frozen. The post-transformation customer_id scan found zero duplicate values."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Snapshot auditor", "text": "The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history."}, {"speaker": "Snapshot auditor", "text": "In the 2026-08-31 validation snapshot, exactly 39,982 of those records have an observed ship_country value."}, {"speaker": "Release coordinator", "text": "The snapshot and scan are the signed validation artifacts for this batch. The batch identifier was reconciled across both artifacts, and no later rerun or manual correction is part of the review record."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the zero-duplicate requirement, the 99.95% scoped-completeness requirement, and the history-record exception, while preserving batch B-142 and the post-transformation review scope. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the observed ship_country count from 39,982 to 39,978 out of 40,000 without conflicting duplicate measurements or assertions. Neither context embeds an answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"For release review, the approved deduplication transform was completed on batch B-142 before the 2026-08-31 validation snapshot was frozen. The post-transformation customer_id scan found zero duplicate values.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Snapshot auditor\",\"text\":\"The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history.\"},{\"speaker\":\"Snapshot auditor\",\"text\":\"In the 2026-08-31 validation snapshot, exactly 39,982 of those records have an observed ship_country value.\"},{\"speaker\":\"Release coordinator\",\"text\":\"The snapshot and scan are the signed validation artifacts for this batch. The batch identifier was reconciled across both artifacts, and no later rerun or manual correction is part of the review record.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history."}, {"path": ["4", "text"], "text": "In the 2026-08-31 validation snapshot, exactly 39,982 of those records have an observed ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history.", "negative_left": "The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history.", "negative_right": "In the 2026-08-31 validation snapshot, exactly 39,978 of those records have an observed ship_country value.", "right": "In the 2026-08-31 validation snapshot, exactly 39,982 of those records have an observed ship_country value."}, "verifier_independent_model": false}, "family": "scale-diverse-142-004", "id": "scale-diverse-142-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "For release review, the approved deduplication transform was completed on batch B-142 before the 2026-08-31 validation snapshot was frozen. The post-transformation customer_id scan found zero duplicate values."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Snapshot auditor", "text": "The 2026-08-31 validation snapshot contains exactly 40,000 records in batch B-142 that are not tagged import_mode=history."}, {"speaker": "Snapshot auditor", "text": "In the 2026-08-31 validation snapshot, exactly 39,978 of those records have an observed ship_country value."}, {"speaker": "Release coordinator", "text": "The snapshot and scan are the signed validation artifacts for this batch. The batch identifier was reconciled across both artifacts, and no later rerun or manual correction is part of the review record."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same transformation rule and request a post-transformation completeness rating for the same customer batch and required-cell scope. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes only C-106’s postal_country status, leaving five of the six blank region_code cells derivable without conflicting with the stated pre-transformation blank counts. Neither context includes a rating, answer code, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"At 08:35 UTC on 14 May 2026, the analyst froze the customer-batch snapshot before any transformation. The product schema requires customer_name, email, and region_code, giving 300 required cells. A validation run on that snapshot found exactly 12 required cells blank, of which exactly six were required region_code cells. At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106. At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, C-91, and C-106. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” After the rule was scheduled for application, the application administrator requested the post-transformation completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106."}, {"path": [], "text": "At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, C-91, and C-106."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106.", "negative_left": "At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106.", "negative_right": "At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, and C-91, while the postal_country cell in customer-batch record C-106 was blank.", "right": "At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, C-91, and C-106."}, "verifier_independent_model": false}, "family": "scale-diverse-143-001", "id": "scale-diverse-143-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "At 08:35 UTC on 14 May 2026, the analyst froze the customer-batch snapshot before any transformation. The product schema requires customer_name, email, and region_code, giving 300 required cells. A validation run on that snapshot found exactly 12 required cells blank, of which exactly six were required region_code cells. At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106. At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, C-91, and C-106. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” After the rule was scheduled for application, the application administrator requested the post-transformation completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same transformation rule and request a post-transformation completeness rating for the same customer batch and required-cell scope. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes only C-106’s postal_country status, leaving five of the six blank region_code cells derivable without conflicting with the stated pre-transformation blank counts. Neither context includes a rating, answer code, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"At 08:35 UTC on 14 May 2026, the analyst froze the customer-batch snapshot before any transformation. The product schema requires customer_name, email, and region_code, giving 300 required cells. A validation run on that snapshot found exactly 12 required cells blank, of which exactly six were required region_code cells. At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106. At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, C-91, and C-106. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” After the rule was scheduled for application, the application administrator requested the post-transformation completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106."}, {"path": [], "text": "At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, C-91, and C-106."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106.", "negative_left": "At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106.", "negative_right": "At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, and C-91, while the postal_country cell in customer-batch record C-106 was blank.", "right": "At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, C-91, and C-106."}, "verifier_independent_model": false}, "family": "scale-diverse-143-001", "id": "scale-diverse-143-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "At 08:35 UTC on 14 May 2026, the analyst froze the customer-batch snapshot before any transformation. The product schema requires customer_name, email, and region_code, giving 300 required cells. A validation run on that snapshot found exactly 12 required cells blank, of which exactly six were required region_code cells. At 08:40 UTC on 14 May 2026, immediately before the stated transformation, the only customer-batch records whose required region_code cells were blank were records C-17, C-44, C-58, C-73, C-91, and C-106. At 08:40 UTC on 14 May 2026, postal_country values were present in customer-batch records C-17, C-44, C-58, C-73, and C-91, while the postal_country cell in customer-batch record C-106 was blank. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” After the rule was scheduled for application, the application administrator requested the post-transformation completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same derivation rule and the same request to assess post-transformation completeness for customer batch CB-41 at the specified snapshot using 300 required cells. The two evidence spans are complete factual sentences. The counterfactual coherently changes one postal_country identifier from R83 to R88 while leaving the six blank region_code records and all counts consistent, thereby changing only the overlap relevant to derivation. Neither context supplies a completeness result, score code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Reconciliation notes for customer batch CB-41 refer to the source snapshot captured at 2026-08-12T14:30:00Z, before any transformation. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation ledger recorded exactly 12 required cells as blank at that time. In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83. In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R83, and R95. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The analyst must assess completeness after applying that rule.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83."}, {"path": [], "text": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R83, and R95."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83.", "negative_left": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83.", "negative_right": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R88, and R95.", "right": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R83, and R95."}, "verifier_independent_model": false}, "family": "scale-diverse-143-002", "id": "scale-diverse-143-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Reconciliation notes for customer batch CB-41 refer to the source snapshot captured at 2026-08-12T14:30:00Z, before any transformation. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation ledger recorded exactly 12 required cells as blank at that time. In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83. In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R83, and R95. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The analyst must assess completeness after applying that rule."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same derivation rule and the same request to assess post-transformation completeness for customer batch CB-41 at the specified snapshot using 300 required cells. The two evidence spans are complete factual sentences. The counterfactual coherently changes one postal_country identifier from R83 to R88 while leaving the six blank region_code records and all counts consistent, thereby changing only the overlap relevant to derivation. Neither context supplies a completeness result, score code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Reconciliation notes for customer batch CB-41 refer to the source snapshot captured at 2026-08-12T14:30:00Z, before any transformation. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation ledger recorded exactly 12 required cells as blank at that time. In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83. In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R83, and R95. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The analyst must assess completeness after applying that rule.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83."}, {"path": [], "text": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R83, and R95."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83.", "negative_left": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83.", "negative_right": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R88, and R95.", "right": "In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R83, and R95."}, "verifier_independent_model": false}, "family": "scale-diverse-143-002", "id": "scale-diverse-143-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Reconciliation notes for customer batch CB-41 refer to the source snapshot captured at 2026-08-12T14:30:00Z, before any transformation. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation ledger recorded exactly 12 required cells as blank at that time. In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the required region_code cell was blank for exactly the records identified as R17, R28, R46, R59, R71, and R83. In customer batch CB-41 at the 2026-08-12T14:30:00Z pre-transformation snapshot, the records with present postal_country values had exactly the identifiers R09, R17, R28, R34, R46, R52, R59, R71, R88, and R95. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The analyst must assess completeness after applying that rule."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing derivation rule and the unchanged questions preserve the scoring thresholds and instructions. The same customer batch, required-cell total, pre-transformation snapshot, transformation path, and post-transformation completeness request are maintained. The two evidence spans are complete factual sentences. The counterfactual coherently changes one postal_country-bearing identifier from C-963 to C-927, reducing the overlap with blank region_code records without creating contradictory counts or duplicate assertions. Neither context supplies a completeness percentage, score, answer code, rule table, proposition ID, label rationale, or directed answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Operational handoff: Use the pre-transformation validation snapshot for the customer batch. The product schema requires customer_name, email, and region_code, giving 300 required cells. The snapshot reports exactly 12 required cells blank before any derivation is applied. In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963. In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-963, and C-988. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Apply only this rule when calculating the percentage of required cells populated after transformation, then select the matching completeness level.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963."}, {"path": [], "text": "In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-963, and C-988."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963.", "negative_left": "In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963.", "negative_right": "In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-927, and C-988.", "right": "In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-963, and C-988."}, "verifier_independent_model": false}, "family": "scale-diverse-143-003", "id": "scale-diverse-143-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Operational handoff: Use the pre-transformation validation snapshot for the customer batch. The product schema requires customer_name, email, and region_code, giving 300 required cells. The snapshot reports exactly 12 required cells blank before any derivation is applied. In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963. In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-963, and C-988. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Apply only this rule when calculating the percentage of required cells populated after transformation, then select the matching completeness level."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing derivation rule and the unchanged questions preserve the scoring thresholds and instructions. The same customer batch, required-cell total, pre-transformation snapshot, transformation path, and post-transformation completeness request are maintained. The two evidence spans are complete factual sentences. The counterfactual coherently changes one postal_country-bearing identifier from C-963 to C-927, reducing the overlap with blank region_code records without creating contradictory counts or duplicate assertions. Neither context supplies a completeness percentage, score, answer code, rule table, proposition ID, label rationale, or directed answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Operational handoff: Use the pre-transformation validation snapshot for the customer batch. The product schema requires customer_name, email, and region_code, giving 300 required cells. The snapshot reports exactly 12 required cells blank before any derivation is applied. In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963. In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-963, and C-988. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Apply only this rule when calculating the percentage of required cells populated after transformation, then select the matching completeness level.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963."}, {"path": [], "text": "In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-963, and C-988."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963.", "negative_left": "In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963.", "negative_right": "In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-927, and C-988.", "right": "In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-963, and C-988."}, "verifier_independent_model": false}, "family": "scale-diverse-143-003", "id": "scale-diverse-143-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Operational handoff: Use the pre-transformation validation snapshot for the customer batch. The product schema requires customer_name, email, and region_code, giving 300 required cells. The snapshot reports exactly 12 required cells blank before any derivation is applied. In the customer batch before the stated transformation, the exactly six blank required region_code cells are attached to record identifiers C-417, C-582, C-639, C-744, C-851, and C-963. In the customer batch before the stated transformation, present postal_country values are attached to exactly record identifiers C-205, C-417, C-582, C-639, C-744, C-851, C-927, and C-988. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Apply only this rule when calculating the percentage of required cells populated after transformation, then select the matching completeness level."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same Northstar batch-import request, schema constraints, governing readiness criteria, entity, process stage, and ownership scope. The two focus spans are complete factual sentences; the counterfactual changes only the batch identifier inventory from one duplicate to none, and the unchanged statement about every other difference being an occurrence-count increase remains vacuously coherent when no such other differences exist. Neither context embeds a readiness label, prescribed answer, answer code, rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A17": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A17": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A17": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified atoms. A1 is a factual comparison of identifier occurrence counts rather than a policy conclusion. Both assignments are realizable while changing only A1: the base can contain pipeline-added duplicate rows, while the counter can retain the same formatting differences without any count increase. The policy evidence correctly preserves the source-state schema requirements needed to interpret source conformance; all readiness criteria, thresholds, ordering, actions, and owners are already retained in the unchanged questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The source satisfies all stated schema requirements, retains all source rows and key values, and has unique source identifiers. With the reviewed pipeline as the sole intervening process, an increased occurrence count for an external_customer_id entails a pipeline-created duplicate. Source-attributable identifier conflicts and unresolved application-owned rules are excluded, so target 3, the data-engineer hold, is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 together with A14 excludes every non-formatting source-to-batch difference. A6 and A7 separately exclude record loss and key changes. The remaining differences are reversible, non-key, documented, value-preserving formatting differences affecting at most 500 of 50,000 rows, exactly 1%. Source schema violations and unresolved product-owned requirements are also excluded, so target 1, acceptance with monitoring, is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For at least one external_customer_id value, its occurrence count in the post-pipeline Northstar batch is greater than its occurrence count in the submitted 50,000-row Northstar source extract."}, {"id": "A2", "statement": "Every row in the submitted 50,000-row Northstar source extract has a non-null external_customer_id."}, {"id": "A3", "statement": "No two rows in the submitted 50,000-row Northstar source extract have the same external_customer_id."}, {"id": "A4", "statement": "Every account_status value in the submitted 50,000-row Northstar source extract is either active or inactive."}, {"id": "A5", "statement": "The reviewed pipeline is the only process between the submitted 50,000-row Northstar source extract and the post-pipeline Northstar batch."}, {"id": "A6", "statement": "Every row in the submitted 50,000-row Northstar source extract has a corresponding row in the post-pipeline Northstar batch."}, {"id": "A7", "statement": "Every source row and its corresponding post-pipeline batch row have the same external_customer_id value."}, {"id": "A8", "statement": "At least one row in the post-pipeline Northstar batch has an email-address or postal-code formatting difference from its corresponding source row."}, {"id": "A9", "statement": "At most 500 of the 50,000 submitted rows are affected by email-address or postal-code formatting differences in the post-pipeline Northstar batch."}, {"id": "A10", "statement": "Every email-address or postal-code formatting difference between a source row and its corresponding post-pipeline batch row is reversible."}, {"id": "A11", "statement": "Every email-address or postal-code formatting difference between a source row and its corresponding post-pipeline batch row affects a non-key field."}, {"id": "A12", "statement": "Every email-address or postal-code formatting difference between a source row and its corresponding post-pipeline batch row is covered by documented normalization."}, {"id": "A13", "statement": "For every email-address or postal-code formatting difference covered by documented normalization, the normalization preserves the source value."}, {"id": "A14", "statement": "Every difference between the submitted 50,000-row Northstar source extract and the post-pipeline Northstar batch that is not an email-address or postal-code formatting difference is an increase in the occurrence count of an external_customer_id."}, {"id": "A15", "statement": "No product-owned mapping remains unresolved for the submitted 50,000-row Northstar batch."}, {"id": "A16", "statement": "No allowed value remains unresolved for the submitted 50,000-row Northstar batch."}, {"id": "A17", "statement": "No display rule remains unresolved for the submitted 50,000-row Northstar batch."}], "base_state_json": "{\"context\":\"Mara reconciles the submitted Northstar extract against the resulting batch before production import.\",\"evidence\":[\"Schema requires external_customer_id to be non-null and unique; account_status must be active or inactive.\",\"The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.\",\"The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with NS-27184 occurring twice and each other identifier occurring once.\",\"Audit confirms every source row has a populated external_customer_id and an active or inactive account_status.\",\"The reviewed pipeline was the sole intervening process; every source row corresponds to a batch row bearing the same external_customer_id.\",\"Exactly 89 batch rows differ from source in email or postal-code formatting; all such differences are reversible, non-key, covered by documented normalization, and preserve the source value.\",\"Every other source-to-batch difference is an external_customer_id occurrence-count increase.\",\"No product-owned mapping, allowed value, or display rule remains unresolved for this batch.\"],\"request\":\"Score batch-import readiness and identify the action and owner prescribed by the selected level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}, {"atom_id": "A17", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}, {"atom_id": "A17", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "1"], "text": "The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once."}, {"path": ["evidence", "2"], "text": "The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with NS-27184 occurring twice and each other identifier occurring once."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Schema requires external_customer_id to be non-null and unique; account_status must be active or inactive."}], "rules": [{"justification": "The source separately satisfies every stated schema requirement. An external_customer_id occurrence-count increase appears only after the sole intervening pipeline, while source records and key values remain intact, so pipeline logic created a duplicate. The source-operator rejection conditions and application-owned unresolved conditions are excluded, making the prescribed result the data-engineer hold and rerun after pipeline correction.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}, {"atom_id": "A17", "state": "supported"}]}, {"justification": "A14 limits every non-formatting difference to an external_customer_id occurrence-count increase, and A1's explicit refutation eliminates all such increases. The remaining evidenced differences are therefore only reversible, non-key formatting differences, affect at most 500 of 50,000 rows, and have documented value-preserving normalization. Record loss, key changes, source schema violations, and unresolved product-owned requirements are excluded, so acceptance with monitoring applies.", "target": "1", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}, {"atom_id": "A17", "state": "supported"}]}]}, "verified_pair": {"left": "The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "negative_left": "The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "negative_right": "The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "right": "The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with NS-27184 occurring twice and each other identifier occurring once."}, "verifier_independent_model": false}, "family": "scale-diverse-144-002", "id": "scale-diverse-144-002-base", "input": {"questions": {"decision": {"criteria": ["Accept: No schema violations, conflicting required identifiers, unexplained record loss, or material transformation defects are evidenced.", "Accept with monitoring: Only reversible, non-key formatting differences affect at most 1% of rows, and documented normalization preserves the values.", "Conditional hold for application administrator: Valid source keys are intact, but product-owned mappings, allowed values, or display rules remain unresolved for at most 1% of rows.", "Hold for data engineer: The source extract conforms to schema, but pipeline logic creates duplicates, drops records, or changes key values; rerun after pipeline correction.", "Reject and route to source-system operator: At least 1% of rows contain conflicting required identifiers attributable to the source, and the configured import would overwrite identities or otherwise cause irreversible ambiguity; require a corrected extract before import."], "instructions": "Choose exactly one ordered readiness level. Apply the concrete thresholds below, distinguish schema-threatening evidence from harmless formatting facts, and do not rely solely on the validation summary’s recommendation.", "type": "score"}}, "state": {"context": "Mara reconciles the submitted Northstar extract against the resulting batch before production import.", "evidence": ["Schema requires external_customer_id to be non-null and unique; account_status must be active or inactive.", "The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with NS-27184 occurring twice and each other identifier occurring once.", "Audit confirms every source row has a populated external_customer_id and an active or inactive account_status.", "The reviewed pipeline was the sole intervening process; every source row corresponds to a batch row bearing the same external_customer_id.", "Exactly 89 batch rows differ from source in email or postal-code formatting; all such differences are reversible, non-key, covered by documented normalization, and preserve the source value.", "Every other source-to-batch difference is an external_customer_id occurrence-count increase.", "No product-owned mapping, allowed value, or display rule remains unresolved for this batch."], "request": "Score batch-import readiness and identify the action and owner prescribed by the selected level."}}, "method": "c2d", "provenance": {"source_id": "diverse-144", "source_is_synthetic": true, "source_sha256": "3ceabfed7faf0763ec5f5e1db35ca23557a098bee61f614bba650d8ffcdb8157", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same Northstar batch-import request, schema constraints, governing readiness criteria, entity, process stage, and ownership scope. The two focus spans are complete factual sentences; the counterfactual changes only the batch identifier inventory from one duplicate to none, and the unchanged statement about every other difference being an occurrence-count increase remains vacuously coherent when no such other differences exist. Neither context embeds a readiness label, prescribed answer, answer code, rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A17": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A17": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A17": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified atoms. A1 is a factual comparison of identifier occurrence counts rather than a policy conclusion. Both assignments are realizable while changing only A1: the base can contain pipeline-added duplicate rows, while the counter can retain the same formatting differences without any count increase. The policy evidence correctly preserves the source-state schema requirements needed to interpret source conformance; all readiness criteria, thresholds, ordering, actions, and owners are already retained in the unchanged questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The source satisfies all stated schema requirements, retains all source rows and key values, and has unique source identifiers. With the reviewed pipeline as the sole intervening process, an increased occurrence count for an external_customer_id entails a pipeline-created duplicate. Source-attributable identifier conflicts and unresolved application-owned rules are excluded, so target 3, the data-engineer hold, is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 together with A14 excludes every non-formatting source-to-batch difference. A6 and A7 separately exclude record loss and key changes. The remaining differences are reversible, non-key, documented, value-preserving formatting differences affecting at most 500 of 50,000 rows, exactly 1%. Source schema violations and unresolved product-owned requirements are also excluded, so target 1, acceptance with monitoring, is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For at least one external_customer_id value, its occurrence count in the post-pipeline Northstar batch is greater than its occurrence count in the submitted 50,000-row Northstar source extract."}, {"id": "A2", "statement": "Every row in the submitted 50,000-row Northstar source extract has a non-null external_customer_id."}, {"id": "A3", "statement": "No two rows in the submitted 50,000-row Northstar source extract have the same external_customer_id."}, {"id": "A4", "statement": "Every account_status value in the submitted 50,000-row Northstar source extract is either active or inactive."}, {"id": "A5", "statement": "The reviewed pipeline is the only process between the submitted 50,000-row Northstar source extract and the post-pipeline Northstar batch."}, {"id": "A6", "statement": "Every row in the submitted 50,000-row Northstar source extract has a corresponding row in the post-pipeline Northstar batch."}, {"id": "A7", "statement": "Every source row and its corresponding post-pipeline batch row have the same external_customer_id value."}, {"id": "A8", "statement": "At least one row in the post-pipeline Northstar batch has an email-address or postal-code formatting difference from its corresponding source row."}, {"id": "A9", "statement": "At most 500 of the 50,000 submitted rows are affected by email-address or postal-code formatting differences in the post-pipeline Northstar batch."}, {"id": "A10", "statement": "Every email-address or postal-code formatting difference between a source row and its corresponding post-pipeline batch row is reversible."}, {"id": "A11", "statement": "Every email-address or postal-code formatting difference between a source row and its corresponding post-pipeline batch row affects a non-key field."}, {"id": "A12", "statement": "Every email-address or postal-code formatting difference between a source row and its corresponding post-pipeline batch row is covered by documented normalization."}, {"id": "A13", "statement": "For every email-address or postal-code formatting difference covered by documented normalization, the normalization preserves the source value."}, {"id": "A14", "statement": "Every difference between the submitted 50,000-row Northstar source extract and the post-pipeline Northstar batch that is not an email-address or postal-code formatting difference is an increase in the occurrence count of an external_customer_id."}, {"id": "A15", "statement": "No product-owned mapping remains unresolved for the submitted 50,000-row Northstar batch."}, {"id": "A16", "statement": "No allowed value remains unresolved for the submitted 50,000-row Northstar batch."}, {"id": "A17", "statement": "No display rule remains unresolved for the submitted 50,000-row Northstar batch."}], "base_state_json": "{\"context\":\"Mara reconciles the submitted Northstar extract against the resulting batch before production import.\",\"evidence\":[\"Schema requires external_customer_id to be non-null and unique; account_status must be active or inactive.\",\"The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.\",\"The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with NS-27184 occurring twice and each other identifier occurring once.\",\"Audit confirms every source row has a populated external_customer_id and an active or inactive account_status.\",\"The reviewed pipeline was the sole intervening process; every source row corresponds to a batch row bearing the same external_customer_id.\",\"Exactly 89 batch rows differ from source in email or postal-code formatting; all such differences are reversible, non-key, covered by documented normalization, and preserve the source value.\",\"Every other source-to-batch difference is an external_customer_id occurrence-count increase.\",\"No product-owned mapping, allowed value, or display rule remains unresolved for this batch.\"],\"request\":\"Score batch-import readiness and identify the action and owner prescribed by the selected level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}, {"atom_id": "A17", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}, {"atom_id": "A17", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "1"], "text": "The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once."}, {"path": ["evidence", "2"], "text": "The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with NS-27184 occurring twice and each other identifier occurring once."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Schema requires external_customer_id to be non-null and unique; account_status must be active or inactive."}], "rules": [{"justification": "The source separately satisfies every stated schema requirement. An external_customer_id occurrence-count increase appears only after the sole intervening pipeline, while source records and key values remain intact, so pipeline logic created a duplicate. The source-operator rejection conditions and application-owned unresolved conditions are excluded, making the prescribed result the data-engineer hold and rerun after pipeline correction.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}, {"atom_id": "A17", "state": "supported"}]}, {"justification": "A14 limits every non-formatting difference to an external_customer_id occurrence-count increase, and A1's explicit refutation eliminates all such increases. The remaining evidenced differences are therefore only reversible, non-key formatting differences, affect at most 500 of 50,000 rows, and have documented value-preserving normalization. Record loss, key changes, source schema violations, and unresolved product-owned requirements are excluded, so acceptance with monitoring applies.", "target": "1", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}, {"atom_id": "A17", "state": "supported"}]}]}, "verified_pair": {"left": "The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "negative_left": "The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "negative_right": "The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "right": "The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with NS-27184 occurring twice and each other identifier occurring once."}, "verifier_independent_model": false}, "family": "scale-diverse-144-002", "id": "scale-diverse-144-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Accept: No schema violations, conflicting required identifiers, unexplained record loss, or material transformation defects are evidenced.", "Accept with monitoring: Only reversible, non-key formatting differences affect at most 1% of rows, and documented normalization preserves the values.", "Conditional hold for application administrator: Valid source keys are intact, but product-owned mappings, allowed values, or display rules remain unresolved for at most 1% of rows.", "Hold for data engineer: The source extract conforms to schema, but pipeline logic creates duplicates, drops records, or changes key values; rerun after pipeline correction.", "Reject and route to source-system operator: At least 1% of rows contain conflicting required identifiers attributable to the source, and the configured import would overwrite identities or otherwise cause irreversible ambiguity; require a corrected extract before import."], "instructions": "Choose exactly one ordered readiness level. Apply the concrete thresholds below, distinguish schema-threatening evidence from harmless formatting facts, and do not rely solely on the validation summary’s recommendation.", "type": "score"}}, "state": {"context": "Mara reconciles the submitted Northstar extract against the resulting batch before production import.", "evidence": ["Schema requires external_customer_id to be non-null and unique; account_status must be active or inactive.", "The complete external_customer_id inventory of the submitted 50,000-row Northstar source extract, finalized at 2026-08-11T09:00:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "The complete external_customer_id inventory of the post-pipeline Northstar batch, finalized at 2026-08-11T09:30:00Z, contains exactly NS-00001 through NS-50000, with each identifier occurring once.", "Audit confirms every source row has a populated external_customer_id and an active or inactive account_status.", "The reviewed pipeline was the sole intervening process; every source row corresponds to a batch row bearing the same external_customer_id.", "Exactly 89 batch rows differ from source in email or postal-code formatting; all such differences are reversible, non-key, covered by documented normalization, and preserve the source value.", "Every other source-to-batch difference is an external_customer_id occurrence-count increase.", "No product-owned mapping, allowed value, or display rule remains unresolved for this batch."], "request": "Score batch-import readiness and identify the action and owner prescribed by the selected level."}}, "method": "c2d", "provenance": {"source_id": "diverse-144", "source_is_synthetic": true, "source_sha256": "3ceabfed7faf0763ec5f5e1db35ca23557a098bee61f614bba650d8ffcdb8157", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the release-gate requirements, authoritative-evidence rule, and gate-owner assignments while retaining the Northstar Notes 6.2 RC3 release-readiness scope. The two evidence spans are complete factual sentences. The counterfactual changes DN-417’s rating from Sev-2 to Sev-3 without conflicting with the unchanged defect count, status, rollback failure, or other gate results. Neither context includes an answer label, code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed. In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-2 and Sev-4, respectively. At 16:35 UTC, the compatibility tester’s signed report recorded passes for every required Windows 11 and macOS 15 compatibility check. At 16:48 UTC, the deployment specialist’s signed packaging report recorded that the required installation gate and required signing gate passed. At 17:02 UTC, that report recorded that the required rollback gate failed because the updater service remained disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed."}, {"path": [], "text": "In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-2 and Sev-4, respectively."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed.", "negative_left": "At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed.", "negative_right": "In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-3 and Sev-4, respectively.", "right": "In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-2 and Sev-4, respectively."}, "verifier_independent_model": false}, "family": "scale-diverse-145-001", "id": "scale-diverse-145-001-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed. In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-2 and Sev-4, respectively. At 16:35 UTC, the compatibility tester’s signed report recorded passes for every required Windows 11 and macOS 15 compatibility check. At 16:48 UTC, the deployment specialist’s signed packaging report recorded that the required installation gate and required signing gate passed. At 17:02 UTC, that report recorded that the required rollback gate failed because the updater service remained disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the release-gate requirements, authoritative-evidence rule, and gate-owner assignments while retaining the Northstar Notes 6.2 RC3 release-readiness scope. The two evidence spans are complete factual sentences. The counterfactual changes DN-417’s rating from Sev-2 to Sev-3 without conflicting with the unchanged defect count, status, rollback failure, or other gate results. Neither context includes an answer label, code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed. In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-2 and Sev-4, respectively. At 16:35 UTC, the compatibility tester’s signed report recorded passes for every required Windows 11 and macOS 15 compatibility check. At 16:48 UTC, the deployment specialist’s signed packaging report recorded that the required installation gate and required signing gate passed. At 17:02 UTC, that report recorded that the required rollback gate failed because the updater service remained disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed."}, {"path": [], "text": "In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-2 and Sev-4, respectively."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed.", "negative_left": "At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed.", "negative_right": "In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-3 and Sev-4, respectively.", "right": "In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-2 and Sev-4, respectively."}, "verifier_independent_model": false}, "family": "scale-diverse-145-001", "id": "scale-diverse-145-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "At 16:20 UTC on 12 September 2026, signed defect list NSN62-RC3-47 was designated authoritative for Northstar Notes 6.2 RC3 and contained exactly defect DN-417, marked open, and defect DN-583, marked closed. In signed defect list NSN62-RC3-47, the recorded ratings for defects DN-417 and DN-583 were Sev-3 and Sev-4, respectively. At 16:35 UTC, the compatibility tester’s signed report recorded passes for every required Windows 11 and macOS 15 compatibility check. At 16:48 UTC, the deployment specialist’s signed packaging report recorded that the required installation gate and required signing gate passed. At 17:02 UTC, that report recorded that the required rollback gate failed because the updater service remained disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same release-readiness requirements, evidence-precedence rule, and gate-owner assignments while preserving the Northstar Notes 6.2 RC3 decision scope. The two focus spans are complete factual sentences. The counterfactual changes only the worksheet’s defect counts, placing zero qualifying defects in authoritative report NN62-D47 and three in nonauthoritative NN62-D91; this is coherent with the authority designation and the unchanged rollback failure. Neither context includes a classification label, answer code, rule table, proposition ID, output instruction, or explicit gold-answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Evidence reconciliation for Northstar Notes 6.2 RC3 was completed after the signed materials were checked. The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3. The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 3 defects in NN62-D47 and 0 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2. The compatibility tester’s signed report records that every required Windows 11 and macOS 15 compatibility check passes. The deployment specialist’s signed packaging report records that the required installation gate and required signing gate pass. It also records that the required rollback gate fails because restoration leaves the updater service disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3."}, {"path": [], "text": "The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 3 defects in NN62-D47 and 0 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3.", "negative_left": "The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3.", "negative_right": "The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 0 defects in NN62-D47 and 3 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2.", "right": "The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 3 defects in NN62-D47 and 0 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2."}, "verifier_independent_model": false}, "family": "scale-diverse-145-002", "id": "scale-diverse-145-002-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Evidence reconciliation for Northstar Notes 6.2 RC3 was completed after the signed materials were checked. The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3. The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 3 defects in NN62-D47 and 0 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2. The compatibility tester’s signed report records that every required Windows 11 and macOS 15 compatibility check passes. The deployment specialist’s signed packaging report records that the required installation gate and required signing gate pass. It also records that the required rollback gate fails because restoration leaves the updater service disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same release-readiness requirements, evidence-precedence rule, and gate-owner assignments while preserving the Northstar Notes 6.2 RC3 decision scope. The two focus spans are complete factual sentences. The counterfactual changes only the worksheet’s defect counts, placing zero qualifying defects in authoritative report NN62-D47 and three in nonauthoritative NN62-D91; this is coherent with the authority designation and the unchanged rollback failure. Neither context includes a classification label, answer code, rule table, proposition ID, output instruction, or explicit gold-answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Evidence reconciliation for Northstar Notes 6.2 RC3 was completed after the signed materials were checked. The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3. The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 3 defects in NN62-D47 and 0 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2. The compatibility tester’s signed report records that every required Windows 11 and macOS 15 compatibility check passes. The deployment specialist’s signed packaging report records that the required installation gate and required signing gate pass. It also records that the required rollback gate fails because restoration leaves the updater service disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3."}, {"path": [], "text": "The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 3 defects in NN62-D47 and 0 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3.", "negative_left": "The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3.", "negative_right": "The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 0 defects in NN62-D47 and 3 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2.", "right": "The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 3 defects in NN62-D47 and 0 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2."}, "verifier_independent_model": false}, "family": "scale-diverse-145-002", "id": "scale-diverse-145-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Evidence reconciliation for Northstar Notes 6.2 RC3 was completed after the signed materials were checked. The release audit record finalized at 14:20 UTC on 16 September 2026 designates signed defect report NN62-D47 as the authoritative signed defect list for Northstar Notes 6.2 RC3. The reconciliation worksheet completed at 14:12 UTC on 16 September 2026 records 0 defects in NN62-D47 and 3 defects in NN62-D91 that are both open and rated Sev-1 or Sev-2. The compatibility tester’s signed report records that every required Windows 11 and macOS 15 compatibility check passes. The deployment specialist’s signed packaging report records that the required installation gate and required signing gate pass. It also records that the required rollback gate fails because restoration leaves the updater service disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same release-gate requirements, authority rule, ownership assignments, product/version, and readiness-classification scope. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only NN-731’s severity from Sev-2 to Sev-3, without conflicting with the unchanged defect count, status, or other gate observations. Neither context embeds a classification answer, code, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Operational handoff for Northstar Notes 6.2 RC3, prepared from the final signed records. The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed. On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-2 and defect NN-846 carried the sole severity rating Sev-4. The compatibility team’s signed matrix records that every required Windows 11 and macOS 15 compatibility check passes. Packaging records show that the required installation gate passes and the required signing gate passes. The required rollback gate fails because the updater service remains disabled after rollback. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed."}, {"path": [], "text": "On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-2 and defect NN-846 carried the sole severity rating Sev-4."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed.", "negative_left": "The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed.", "negative_right": "On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-3 and defect NN-846 carried the sole severity rating Sev-4.", "right": "On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-2 and defect NN-846 carried the sole severity rating Sev-4."}, "verifier_independent_model": false}, "family": "scale-diverse-145-003", "id": "scale-diverse-145-003-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Operational handoff for Northstar Notes 6.2 RC3, prepared from the final signed records. The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed. On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-2 and defect NN-846 carried the sole severity rating Sev-4. The compatibility team’s signed matrix records that every required Windows 11 and macOS 15 compatibility check passes. Packaging records show that the required installation gate passes and the required signing gate passes. The required rollback gate fails because the updater service remains disabled after rollback. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same release-gate requirements, authority rule, ownership assignments, product/version, and readiness-classification scope. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only NN-731’s severity from Sev-2 to Sev-3, without conflicting with the unchanged defect count, status, or other gate observations. Neither context embeds a classification answer, code, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Operational handoff for Northstar Notes 6.2 RC3, prepared from the final signed records. The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed. On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-2 and defect NN-846 carried the sole severity rating Sev-4. The compatibility team’s signed matrix records that every required Windows 11 and macOS 15 compatibility check passes. Packaging records show that the required installation gate passes and the required signing gate passes. The required rollback gate fails because the updater service remains disabled after rollback. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed."}, {"path": [], "text": "On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-2 and defect NN-846 carried the sole severity rating Sev-4."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed.", "negative_left": "The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed.", "negative_right": "On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-3 and defect NN-846 carried the sole severity rating Sev-4.", "right": "On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-2 and defect NN-846 carried the sole severity rating Sev-4."}, "verifier_independent_model": false}, "family": "scale-diverse-145-003", "id": "scale-diverse-145-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Operational handoff for Northstar Notes 6.2 RC3, prepared from the final signed records. The authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026 contained exactly two entries: defect NN-731 marked open and defect NN-846 marked closed. On the authoritative defect list for Northstar Notes 6.2 RC3 signed at 16:40 UTC on 12 September 2026, defect NN-731 carried the sole severity rating Sev-3 and defect NN-846 carried the sole severity rating Sev-4. The compatibility team’s signed matrix records that every required Windows 11 and macOS 15 compatibility check passes. Packaging records show that the required installation gate passes and the required signing gate passes. The required rollback gate fails because the updater service remains disabled after rollback. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing release-gate, evidence-precedence, and ownership policies from the original state, while the unchanged questions preserve the classification criteria. They remain bound to Northstar Notes 6.2 RC3 and the same release-readiness decision scope. The two evidence spans are complete factual sentences. The counterfactual coherently changes NN-842’s sole severity rating from Sev-2 to Sev-3 without conflicting with any unchanged assertion; the rollback gate still fails. Neither context contains a gold label, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Field note — Northstar Notes 6.2 RC3. At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842. The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-2. The compatibility tester’s signed report records a pass for every required Windows 11 and macOS 15 compatibility check. The deployment specialist’s signed packaging report marks both the required installation gate and the required signing gate as passed. It records that the required rollback test left the updater service disabled and marks the rollback gate failed. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842."}, {"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-2."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842.", "negative_left": "At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842.", "negative_right": "The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-3.", "right": "The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-2."}, "verifier_independent_model": false}, "family": "scale-diverse-145-004", "id": "scale-diverse-145-004-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Field note — Northstar Notes 6.2 RC3. At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842. The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-2. The compatibility tester’s signed report records a pass for every required Windows 11 and macOS 15 compatibility check. The deployment specialist’s signed packaging report marks both the required installation gate and the required signing gate as passed. It records that the required rollback test left the updater service disabled and marks the rollback gate failed. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing release-gate, evidence-precedence, and ownership policies from the original state, while the unchanged questions preserve the classification criteria. They remain bound to Northstar Notes 6.2 RC3 and the same release-readiness decision scope. The two evidence spans are complete factual sentences. The counterfactual coherently changes NN-842’s sole severity rating from Sev-2 to Sev-3 without conflicting with any unchanged assertion; the rollback gate still fails. Neither context contains a gold label, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Field note — Northstar Notes 6.2 RC3. At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842. The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-2. The compatibility tester’s signed report records a pass for every required Windows 11 and macOS 15 compatibility check. The deployment specialist’s signed packaging report marks both the required installation gate and the required signing gate as passed. It records that the required rollback test left the updater service disabled and marks the rollback gate failed. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842."}, {"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-2."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842.", "negative_left": "At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842.", "negative_right": "The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-3.", "right": "The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-2."}, "verifier_independent_model": false}, "family": "scale-diverse-145-004", "id": "scale-diverse-145-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Field note — Northstar Notes 6.2 RC3. At 14:20 UTC on 8 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 classifies exactly one defect as open: NN-842. The authoritative signed defect list for Northstar Notes 6.2 RC3 assigns NN-842 exactly one severity rating: Sev-3. The compatibility tester’s signed report records a pass for every required Windows 11 and macOS 15 compatibility check. The deployment specialist’s signed packaging report marks both the required installation gate and the required signing gate as passed. It records that the required rollback test left the updater service disabled and marks the rollback gate failed. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory-check and missing-evidence policies, the QuillDesk 8.2 RC3/Windows 11 ARM64 scope, the deployment ownership relevant to the requested action, and the same readiness request. The two cited evidence spans are complete factual sentences. Changing R-731 from occupied to unoccupied coherently changes the bound ARM64 installation-result observation while leaving the other three ARM64 checks unchanged, with no duplicate contradictory measurement. Neither context includes an option code, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified platform atoms remain atomic. The focus atom concerns whether an ARM64 installation result is recorded, not a policy conclusion. Base and counter assignments can differ only on that fact: separate launch, uninstall, and signature records can exist without a recorded installation result. Policy evidence adequately preserves the relevant state-originated rules and role binding; governing criteria already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all four mandatory recorded checks for ARM64 and every other supported platform, while refuting product-code failure, open blocking defects, and other failed gates. It therefore sufficiently entails option A and excludes the competing failure and hold outcomes.", "rule_index": 0, "sound": true}, {"reason": "ARM64 is supported and is explicitly missing its mandatory recorded installation result. Product-code failure, blocking defects, other failed gates, and recorded ARM64 incompatibility are refuted, so the case is undetermined rather than failed and option C is sufficient. Deployment is the state-identified installer owner.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64."}, {"id": "a2", "statement": "QuillDesk 8.2 RC3 has a recorded launch result for Windows 11 ARM64."}, {"id": "a3", "statement": "QuillDesk 8.2 RC3 has a recorded uninstall result for Windows 11 ARM64."}, {"id": "a4", "statement": "QuillDesk 8.2 RC3 has a recorded signature-check result for Windows 11 ARM64."}, {"id": "a5", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded installation result."}, {"id": "a6", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded launch result."}, {"id": "a7", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded uninstall result."}, {"id": "a8", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded signature-check result."}, {"id": "a9", "statement": "A product-code failure is recorded for QuillDesk 8.2 RC3."}, {"id": "a10", "statement": "An open blocking defect remains for QuillDesk 8.2 RC3."}, {"id": "a11", "statement": "A recorded failed gate other than missing mandatory evidence remains for QuillDesk 8.2 RC3."}, {"id": "a12", "statement": "An incompatibility between QuillDesk 8.2 RC3 and Windows 11 ARM64 is recorded."}, {"id": "a13", "statement": "Windows 11 ARM64 is listed as a supported platform for QuillDesk 8.2 RC3."}], "base_state_json": "{\"request\":\"Under the stated gate policy, determine RC3 readiness and the proper next action.\",\"context\":\"the deployment specialist owns installers and rollback validation.\",\"handoff\":{\"scope\":\"The support matrix lists Windows 11 ARM64 for QuillDesk 8.2 RC3.\",\"arm64_checks\":\"The frozen register contains recorded launch, uninstall, and signature-check results for that platform.\",\"other_platforms\":\"For every other supported QuillDesk 8.2 RC3 platform, the register contains recorded installation, launch, uninstall, and signature-check results.\",\"exceptions\":\"Release review found no product-code failure, no open blocking defect, no other recorded failed gate, and no recorded incompatibility with Windows 11 ARM64.\",\"register_binding\":\"In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64.\",\"register_state\":\"Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is occupied.\"},\"policy\":[\"Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform.\",\"Missing mandatory evidence means readiness is undetermined, not failed.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["handoff", "register_binding"], "text": "In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64."}, {"path": ["handoff", "register_state"], "text": "Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is occupied."}], "policy_evidence": [{"path": ["request"], "text": "Under the stated gate policy, determine RC3 readiness and the proper next action."}, {"path": ["context"], "text": "the deployment specialist owns installers and rollback validation."}, {"path": ["evidence", "3"], "text": "Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform."}, {"path": ["evidence", "3"], "text": "Missing mandatory evidence means readiness is undetermined, not failed."}], "rules": [{"justification": "Windows 11 ARM64 and every other supported platform have each mandatory recorded check, while product-code failure, open blocking defect, another recorded failed gate, and ARM64 incompatibility are excluded.", "target": "A — Pass the release gate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Windows 11 ARM64 is supported but explicitly lacks its mandatory recorded installation result. Recorded failure outcomes and ARM64 incompatibility are excluded, so the policy treats readiness as undetermined and assigns collection of the missing installer evidence to deployment.", "target": "C — Hold as undetermined; request ARM64 evidence from deployment", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64.", "negative_left": "In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64.", "negative_right": "Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is unoccupied.", "right": "Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is occupied."}, "verifier_independent_model": false}, "family": "scale-diverse-146-003", "id": "scale-diverse-146-003-base", "input": {"questions": {"decision": {"criteria": {"A — Pass the release gate": "Choose only when every supported platform has all required recorded checks and no blocking defect or failed gate remains.", "B — Fail and route to desktop development": "Choose only when evidence records a product-code failure or an open blocking defect requiring a developer fix.", "C — Hold as undetermined; request ARM64 evidence from deployment": "Choose when a supported platform lacks mandatory packaging or installation evidence; ask the deployment specialist to run and record the missing checks.", "D — Fail and remove ARM64 from documentation": "Choose only when ARM64 has a recorded incompatibility and an approved scope change requires the compatibility tester and release manager to remove it from the support matrix."}, "instructions": "Select exactly one option. Apply the gate policy literally: distinguish a recorded failure from missing mandatory evidence and route evidence collection to the role that owns installers.", "type": "choice"}}, "state": {"context": "the deployment specialist owns installers and rollback validation.", "handoff": {"arm64_checks": "The frozen register contains recorded launch, uninstall, and signature-check results for that platform.", "exceptions": "Release review found no product-code failure, no open blocking defect, no other recorded failed gate, and no recorded incompatibility with Windows 11 ARM64.", "other_platforms": "For every other supported QuillDesk 8.2 RC3 platform, the register contains recorded installation, launch, uninstall, and signature-check results.", "register_binding": "In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64.", "register_state": "Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is occupied.", "scope": "The support matrix lists Windows 11 ARM64 for QuillDesk 8.2 RC3."}, "policy": ["Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform.", "Missing mandatory evidence means readiness is undetermined, not failed."], "request": "Under the stated gate policy, determine RC3 readiness and the proper next action."}}, "method": "c2d", "provenance": {"source_id": "diverse-146", "source_is_synthetic": true, "source_sha256": "b23b175fc407eb93f3fc39e956cded332d303bd5cf7d8099ec76fac590199f87", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A — Pass the release gate"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory-check and missing-evidence policies, the QuillDesk 8.2 RC3/Windows 11 ARM64 scope, the deployment ownership relevant to the requested action, and the same readiness request. The two cited evidence spans are complete factual sentences. Changing R-731 from occupied to unoccupied coherently changes the bound ARM64 installation-result observation while leaving the other three ARM64 checks unchanged, with no duplicate contradictory measurement. Neither context includes an option code, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified platform atoms remain atomic. The focus atom concerns whether an ARM64 installation result is recorded, not a policy conclusion. Base and counter assignments can differ only on that fact: separate launch, uninstall, and signature records can exist without a recorded installation result. Policy evidence adequately preserves the relevant state-originated rules and role binding; governing criteria already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all four mandatory recorded checks for ARM64 and every other supported platform, while refuting product-code failure, open blocking defects, and other failed gates. It therefore sufficiently entails option A and excludes the competing failure and hold outcomes.", "rule_index": 0, "sound": true}, {"reason": "ARM64 is supported and is explicitly missing its mandatory recorded installation result. Product-code failure, blocking defects, other failed gates, and recorded ARM64 incompatibility are refuted, so the case is undetermined rather than failed and option C is sufficient. Deployment is the state-identified installer owner.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64."}, {"id": "a2", "statement": "QuillDesk 8.2 RC3 has a recorded launch result for Windows 11 ARM64."}, {"id": "a3", "statement": "QuillDesk 8.2 RC3 has a recorded uninstall result for Windows 11 ARM64."}, {"id": "a4", "statement": "QuillDesk 8.2 RC3 has a recorded signature-check result for Windows 11 ARM64."}, {"id": "a5", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded installation result."}, {"id": "a6", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded launch result."}, {"id": "a7", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded uninstall result."}, {"id": "a8", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded signature-check result."}, {"id": "a9", "statement": "A product-code failure is recorded for QuillDesk 8.2 RC3."}, {"id": "a10", "statement": "An open blocking defect remains for QuillDesk 8.2 RC3."}, {"id": "a11", "statement": "A recorded failed gate other than missing mandatory evidence remains for QuillDesk 8.2 RC3."}, {"id": "a12", "statement": "An incompatibility between QuillDesk 8.2 RC3 and Windows 11 ARM64 is recorded."}, {"id": "a13", "statement": "Windows 11 ARM64 is listed as a supported platform for QuillDesk 8.2 RC3."}], "base_state_json": "{\"request\":\"Under the stated gate policy, determine RC3 readiness and the proper next action.\",\"context\":\"the deployment specialist owns installers and rollback validation.\",\"handoff\":{\"scope\":\"The support matrix lists Windows 11 ARM64 for QuillDesk 8.2 RC3.\",\"arm64_checks\":\"The frozen register contains recorded launch, uninstall, and signature-check results for that platform.\",\"other_platforms\":\"For every other supported QuillDesk 8.2 RC3 platform, the register contains recorded installation, launch, uninstall, and signature-check results.\",\"exceptions\":\"Release review found no product-code failure, no open blocking defect, no other recorded failed gate, and no recorded incompatibility with Windows 11 ARM64.\",\"register_binding\":\"In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64.\",\"register_state\":\"Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is occupied.\"},\"policy\":[\"Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform.\",\"Missing mandatory evidence means readiness is undetermined, not failed.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["handoff", "register_binding"], "text": "In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64."}, {"path": ["handoff", "register_state"], "text": "Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is occupied."}], "policy_evidence": [{"path": ["request"], "text": "Under the stated gate policy, determine RC3 readiness and the proper next action."}, {"path": ["context"], "text": "the deployment specialist owns installers and rollback validation."}, {"path": ["evidence", "3"], "text": "Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform."}, {"path": ["evidence", "3"], "text": "Missing mandatory evidence means readiness is undetermined, not failed."}], "rules": [{"justification": "Windows 11 ARM64 and every other supported platform have each mandatory recorded check, while product-code failure, open blocking defect, another recorded failed gate, and ARM64 incompatibility are excluded.", "target": "A — Pass the release gate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Windows 11 ARM64 is supported but explicitly lacks its mandatory recorded installation result. Recorded failure outcomes and ARM64 incompatibility are excluded, so the policy treats readiness as undetermined and assigns collection of the missing installer evidence to deployment.", "target": "C — Hold as undetermined; request ARM64 evidence from deployment", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64.", "negative_left": "In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64.", "negative_right": "Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is unoccupied.", "right": "Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is occupied."}, "verifier_independent_model": false}, "family": "scale-diverse-146-003", "id": "scale-diverse-146-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Pass the release gate": "Choose only when every supported platform has all required recorded checks and no blocking defect or failed gate remains.", "B — Fail and route to desktop development": "Choose only when evidence records a product-code failure or an open blocking defect requiring a developer fix.", "C — Hold as undetermined; request ARM64 evidence from deployment": "Choose when a supported platform lacks mandatory packaging or installation evidence; ask the deployment specialist to run and record the missing checks.", "D — Fail and remove ARM64 from documentation": "Choose only when ARM64 has a recorded incompatibility and an approved scope change requires the compatibility tester and release manager to remove it from the support matrix."}, "instructions": "Select exactly one option. Apply the gate policy literally: distinguish a recorded failure from missing mandatory evidence and route evidence collection to the role that owns installers.", "type": "choice"}}, "state": {"context": "the deployment specialist owns installers and rollback validation.", "handoff": {"arm64_checks": "The frozen register contains recorded launch, uninstall, and signature-check results for that platform.", "exceptions": "Release review found no product-code failure, no open blocking defect, no other recorded failed gate, and no recorded incompatibility with Windows 11 ARM64.", "other_platforms": "For every other supported QuillDesk 8.2 RC3 platform, the register contains recorded installation, launch, uninstall, and signature-check results.", "register_binding": "In the complete release-evidence register snapshot frozen at 2026-09-16T14:37:22Z, slot R-731 is occupied if and only if QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64.", "register_state": "Slot R-731 in the release-evidence register snapshot frozen at 2026-09-16T14:37:22Z is unoccupied.", "scope": "The support matrix lists Windows 11 ARM64 for QuillDesk 8.2 RC3."}, "policy": ["Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform.", "Missing mandatory evidence means readiness is undetermined, not failed."], "request": "Under the stated gate policy, determine RC3 readiness and the proper next action."}}, "method": "c2d", "provenance": {"source_id": "diverse-146", "source_is_synthetic": true, "source_sha256": "b23b175fc407eb93f3fc39e956cded332d303bd5cf7d8099ec76fac590199f87", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Hold as undetermined; request ARM64 evidence from deployment"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full decision criteria and instructions, while both contexts retain the governing gate and missing-evidence policies and installer ownership. The request remains bound to QuillDesk 8.2 RC3, Windows 11 ARM64, the same review period, target, and archive. The two focus-evidence spans are complete factual chronology sentences. The counterfactual coherently changes the exhaustive archive's T-583 installation entry from present to absent without conflicting with the separately listed launch, uninstall, and signature results. Neither context contains an answer choice, output instruction, proposition identifier, rule table, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified platform atoms remain atomic. The focus atom concerns whether an ARM64 installation result is recorded, not a policy conclusion. Base and counter assignments can differ only on that fact: separate launch, uninstall, and signature records can exist without a recorded installation result. Policy evidence adequately preserves the relevant state-originated rules and role binding; governing criteria already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all four mandatory recorded checks for ARM64 and every other supported platform, while refuting product-code failure, open blocking defects, and other failed gates. It therefore sufficiently entails option A and excludes the competing failure and hold outcomes.", "rule_index": 0, "sound": true}, {"reason": "ARM64 is supported and is explicitly missing its mandatory recorded installation result. Product-code failure, blocking defects, other failed gates, and recorded ARM64 incompatibility are refuted, so the case is undetermined rather than failed and option C is sufficient. Deployment is the state-identified installer owner.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64."}, {"id": "a2", "statement": "QuillDesk 8.2 RC3 has a recorded launch result for Windows 11 ARM64."}, {"id": "a3", "statement": "QuillDesk 8.2 RC3 has a recorded uninstall result for Windows 11 ARM64."}, {"id": "a4", "statement": "QuillDesk 8.2 RC3 has a recorded signature-check result for Windows 11 ARM64."}, {"id": "a5", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded installation result."}, {"id": "a6", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded launch result."}, {"id": "a7", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded uninstall result."}, {"id": "a8", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded signature-check result."}, {"id": "a9", "statement": "A product-code failure is recorded for QuillDesk 8.2 RC3."}, {"id": "a10", "statement": "An open blocking defect remains for QuillDesk 8.2 RC3."}, {"id": "a11", "statement": "A recorded failed gate other than missing mandatory evidence remains for QuillDesk 8.2 RC3."}, {"id": "a12", "statement": "An incompatibility between QuillDesk 8.2 RC3 and Windows 11 ARM64 is recorded."}, {"id": "a13", "statement": "Windows 11 ARM64 is listed as a supported platform for QuillDesk 8.2 RC3."}], "base_state_json": "{\"context\":\"the deployment specialist owns installers and rollback validation.\",\"chronology\":[\"At 18:00 UTC on 14 September 2026, the support matrix listed Windows 11 ARM64 for QuillDesk 8.2 RC3.\",\"At 18:10 UTC, every supported RC3 platform other than Windows 11 ARM64 had separately recorded installation, launch, uninstall, and signature-check results.\",\"At 18:30 UTC, the Windows 11 ARM64 records contained launch, uninstall, and signature-check results for RC3.\",\"At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64.\",\"At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained an installation-result entry for target T-583.\",\"At review close, RC3 had no recorded product-code failure, no recorded ARM64 incompatibility, no open blocking defect, and no recorded failed gate.\"],\"evidence\":[\"Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform.\",\"Missing mandatory evidence means readiness is undetermined, not failed.\"],\"request\":\"Under the stated gate policy, determine RC3 readiness and the proper next action.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["chronology", "3"], "text": "At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64."}, {"path": ["chronology", "4"], "text": "At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained an installation-result entry for target T-583."}], "policy_evidence": [{"path": ["request"], "text": "Under the stated gate policy, determine RC3 readiness and the proper next action."}, {"path": ["context"], "text": "the deployment specialist owns installers and rollback validation."}, {"path": ["evidence", "3"], "text": "Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform."}, {"path": ["evidence", "3"], "text": "Missing mandatory evidence means readiness is undetermined, not failed."}], "rules": [{"justification": "Windows 11 ARM64 and every other supported platform have each mandatory recorded check, while product-code failure, open blocking defect, another recorded failed gate, and ARM64 incompatibility are excluded.", "target": "A — Pass the release gate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Windows 11 ARM64 is supported but explicitly lacks its mandatory recorded installation result. Recorded failure outcomes and ARM64 incompatibility are excluded, so the policy treats readiness as undetermined and assigns collection of the missing installer evidence to deployment.", "target": "C — Hold as undetermined; request ARM64 evidence from deployment", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64.", "negative_left": "At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64.", "negative_right": "At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained no installation-result entry for target T-583.", "right": "At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained an installation-result entry for target T-583."}, "verifier_independent_model": false}, "family": "scale-diverse-146-005", "id": "scale-diverse-146-005-base", "input": {"questions": {"decision": {"criteria": {"A — Pass the release gate": "Choose only when every supported platform has all required recorded checks and no blocking defect or failed gate remains.", "B — Fail and route to desktop development": "Choose only when evidence records a product-code failure or an open blocking defect requiring a developer fix.", "C — Hold as undetermined; request ARM64 evidence from deployment": "Choose when a supported platform lacks mandatory packaging or installation evidence; ask the deployment specialist to run and record the missing checks.", "D — Fail and remove ARM64 from documentation": "Choose only when ARM64 has a recorded incompatibility and an approved scope change requires the compatibility tester and release manager to remove it from the support matrix."}, "instructions": "Select exactly one option. Apply the gate policy literally: distinguish a recorded failure from missing mandatory evidence and route evidence collection to the role that owns installers.", "type": "choice"}}, "state": {"chronology": ["At 18:00 UTC on 14 September 2026, the support matrix listed Windows 11 ARM64 for QuillDesk 8.2 RC3.", "At 18:10 UTC, every supported RC3 platform other than Windows 11 ARM64 had separately recorded installation, launch, uninstall, and signature-check results.", "At 18:30 UTC, the Windows 11 ARM64 records contained launch, uninstall, and signature-check results for RC3.", "At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64.", "At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained an installation-result entry for target T-583.", "At review close, RC3 had no recorded product-code failure, no recorded ARM64 incompatibility, no open blocking defect, and no recorded failed gate."], "context": "the deployment specialist owns installers and rollback validation.", "evidence": ["Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform.", "Missing mandatory evidence means readiness is undetermined, not failed."], "request": "Under the stated gate policy, determine RC3 readiness and the proper next action."}}, "method": "c2d", "provenance": {"source_id": "diverse-146", "source_is_synthetic": true, "source_sha256": "b23b175fc407eb93f3fc39e956cded332d303bd5cf7d8099ec76fac590199f87", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A — Pass the release gate"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full decision criteria and instructions, while both contexts retain the governing gate and missing-evidence policies and installer ownership. The request remains bound to QuillDesk 8.2 RC3, Windows 11 ARM64, the same review period, target, and archive. The two focus-evidence spans are complete factual chronology sentences. The counterfactual coherently changes the exhaustive archive's T-583 installation entry from present to absent without conflicting with the separately listed launch, uninstall, and signature results. Neither context contains an answer choice, output instruction, proposition identifier, rule table, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified platform atoms remain atomic. The focus atom concerns whether an ARM64 installation result is recorded, not a policy conclusion. Base and counter assignments can differ only on that fact: separate launch, uninstall, and signature records can exist without a recorded installation result. Policy evidence adequately preserves the relevant state-originated rules and role binding; governing criteria already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all four mandatory recorded checks for ARM64 and every other supported platform, while refuting product-code failure, open blocking defects, and other failed gates. It therefore sufficiently entails option A and excludes the competing failure and hold outcomes.", "rule_index": 0, "sound": true}, {"reason": "ARM64 is supported and is explicitly missing its mandatory recorded installation result. Product-code failure, blocking defects, other failed gates, and recorded ARM64 incompatibility are refuted, so the case is undetermined rather than failed and option C is sufficient. Deployment is the state-identified installer owner.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "QuillDesk 8.2 RC3 has a recorded installation result for Windows 11 ARM64."}, {"id": "a2", "statement": "QuillDesk 8.2 RC3 has a recorded launch result for Windows 11 ARM64."}, {"id": "a3", "statement": "QuillDesk 8.2 RC3 has a recorded uninstall result for Windows 11 ARM64."}, {"id": "a4", "statement": "QuillDesk 8.2 RC3 has a recorded signature-check result for Windows 11 ARM64."}, {"id": "a5", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded installation result."}, {"id": "a6", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded launch result."}, {"id": "a7", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded uninstall result."}, {"id": "a8", "statement": "Every supported QuillDesk 8.2 RC3 platform other than Windows 11 ARM64 has a recorded signature-check result."}, {"id": "a9", "statement": "A product-code failure is recorded for QuillDesk 8.2 RC3."}, {"id": "a10", "statement": "An open blocking defect remains for QuillDesk 8.2 RC3."}, {"id": "a11", "statement": "A recorded failed gate other than missing mandatory evidence remains for QuillDesk 8.2 RC3."}, {"id": "a12", "statement": "An incompatibility between QuillDesk 8.2 RC3 and Windows 11 ARM64 is recorded."}, {"id": "a13", "statement": "Windows 11 ARM64 is listed as a supported platform for QuillDesk 8.2 RC3."}], "base_state_json": "{\"context\":\"the deployment specialist owns installers and rollback validation.\",\"chronology\":[\"At 18:00 UTC on 14 September 2026, the support matrix listed Windows 11 ARM64 for QuillDesk 8.2 RC3.\",\"At 18:10 UTC, every supported RC3 platform other than Windows 11 ARM64 had separately recorded installation, launch, uninstall, and signature-check results.\",\"At 18:30 UTC, the Windows 11 ARM64 records contained launch, uninstall, and signature-check results for RC3.\",\"At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64.\",\"At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained an installation-result entry for target T-583.\",\"At review close, RC3 had no recorded product-code failure, no recorded ARM64 incompatibility, no open blocking defect, and no recorded failed gate.\"],\"evidence\":[\"Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform.\",\"Missing mandatory evidence means readiness is undetermined, not failed.\"],\"request\":\"Under the stated gate policy, determine RC3 readiness and the proper next action.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["chronology", "3"], "text": "At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64."}, {"path": ["chronology", "4"], "text": "At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained an installation-result entry for target T-583."}], "policy_evidence": [{"path": ["request"], "text": "Under the stated gate policy, determine RC3 readiness and the proper next action."}, {"path": ["context"], "text": "the deployment specialist owns installers and rollback validation."}, {"path": ["evidence", "3"], "text": "Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform."}, {"path": ["evidence", "3"], "text": "Missing mandatory evidence means readiness is undetermined, not failed."}], "rules": [{"justification": "Windows 11 ARM64 and every other supported platform have each mandatory recorded check, while product-code failure, open blocking defect, another recorded failed gate, and ARM64 incompatibility are excluded.", "target": "A — Pass the release gate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Windows 11 ARM64 is supported but explicitly lacks its mandatory recorded installation result. Recorded failure outcomes and ARM64 incompatibility are excluded, so the policy treats readiness as undetermined and assigns collection of the missing installer evidence to deployment.", "target": "C — Hold as undetermined; request ARM64 evidence from deployment", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64.", "negative_left": "At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64.", "negative_right": "At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained no installation-result entry for target T-583.", "right": "At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained an installation-result entry for target T-583."}, "verifier_independent_model": false}, "family": "scale-diverse-146-005", "id": "scale-diverse-146-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Pass the release gate": "Choose only when every supported platform has all required recorded checks and no blocking defect or failed gate remains.", "B — Fail and route to desktop development": "Choose only when evidence records a product-code failure or an open blocking defect requiring a developer fix.", "C — Hold as undetermined; request ARM64 evidence from deployment": "Choose when a supported platform lacks mandatory packaging or installation evidence; ask the deployment specialist to run and record the missing checks.", "D — Fail and remove ARM64 from documentation": "Choose only when ARM64 has a recorded incompatibility and an approved scope change requires the compatibility tester and release manager to remove it from the support matrix."}, "instructions": "Select exactly one option. Apply the gate policy literally: distinguish a recorded failure from missing mandatory evidence and route evidence collection to the role that owns installers.", "type": "choice"}}, "state": {"chronology": ["At 18:00 UTC on 14 September 2026, the support matrix listed Windows 11 ARM64 for QuillDesk 8.2 RC3.", "At 18:10 UTC, every supported RC3 platform other than Windows 11 ARM64 had separately recorded installation, launch, uninstall, and signature-check results.", "At 18:30 UTC, the Windows 11 ARM64 records contained launch, uninstall, and signature-check results for RC3.", "At 18:40 UTC on 14 September 2026, target T-583 in evidence archive EA-942 denoted QuillDesk 8.2 RC3 on Windows 11 ARM64.", "At its final closure at 18:40 UTC on 14 September 2026, evidence archive EA-942 was the exhaustive record of installation results for its targets and contained no installation-result entry for target T-583.", "At review close, RC3 had no recorded product-code failure, no recorded ARM64 incompatibility, no open blocking defect, and no recorded failed gate."], "context": "the deployment specialist owns installers and rollback validation.", "evidence": ["Gate policy requires a recorded install, launch, uninstall, and signature check for every supported platform.", "Missing mandatory evidence means readiness is undetermined, not failed."], "request": "Under the stated gate policy, determine RC3 readiness and the proper next action."}}, "method": "c2d", "provenance": {"source_id": "diverse-146", "source_is_synthetic": true, "source_sha256": "b23b175fc407eb93f3fc39e956cded332d303bd5cf7d8099ec76fac590199f87", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Hold as undetermined; request ARM64 evidence from deployment"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The policy, request scope, and the blocker described by the deployment specialist remain unchanged in their bindings, and the two evidence spans are complete factual sentences. The counterfactual coherently changes defect D from a signing defect to an application-runtime defect while leaving defect S as the sole signing-related defect. However, the base context should be rejected: it identifies D as a signing defect while separately identifying S as the only signing-related defect, creating contradictory assertions about distinct defect identifiers. Neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"On 14 September 2026, the release manager convened the final deployment review for Northstar Notes 4.2 RC3. Functional testing had concluded, and the team was assessing the remaining blocker before release.\\n\\nAt 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned omission of the required signature from the Windows installer's updater executable.\\n\\nAt 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing.\\n\\nThe release manager then consulted the team-routing policy before assigning the blocker. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "At 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned omission of the required signature from the Windows installer's updater executable."}, {"path": [], "text": "At 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned omission of the required signature from the Windows installer's updater executable.", "negative_left": "At 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned an application-runtime failure but did not concern installer construction or signing.", "negative_right": "At 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing.", "right": "At 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing."}, "verifier_independent_model": false}, "family": "scale-diverse-148-001", "id": "scale-diverse-148-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "On 14 September 2026, the release manager convened the final deployment review for Northstar Notes 4.2 RC3. Functional testing had concluded, and the team was assessing the remaining blocker before release.\n\nAt 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned omission of the required signature from the Windows installer's updater executable.\n\nAt 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing.\n\nThe release manager then consulted the team-routing policy before assigning the blocker. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The policy, request scope, and the blocker described by the deployment specialist remain unchanged in their bindings, and the two evidence spans are complete factual sentences. The counterfactual coherently changes defect D from a signing defect to an application-runtime defect while leaving defect S as the sole signing-related defect. However, the base context should be rejected: it identifies D as a signing defect while separately identifying S as the only signing-related defect, creating contradictory assertions about distinct defect identifiers. Neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"On 14 September 2026, the release manager convened the final deployment review for Northstar Notes 4.2 RC3. Functional testing had concluded, and the team was assessing the remaining blocker before release.\\n\\nAt 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned omission of the required signature from the Windows installer's updater executable.\\n\\nAt 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing.\\n\\nThe release manager then consulted the team-routing policy before assigning the blocker. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "At 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned omission of the required signature from the Windows installer's updater executable."}, {"path": [], "text": "At 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned omission of the required signature from the Windows installer's updater executable.", "negative_left": "At 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned an application-runtime failure but did not concern installer construction or signing.", "negative_right": "At 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing.", "right": "At 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing."}, "verifier_independent_model": false}, "family": "scale-diverse-148-001", "id": "scale-diverse-148-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "On 14 September 2026, the release manager convened the final deployment review for Northstar Notes 4.2 RC3. Functional testing had concluded, and the team was assessing the remaining blocker before release.\n\nAt 08:17 UTC on 14 September 2026, the deployment specialist recorded that the sole defect underlying the Northstar Notes 4.2 RC3 blocker was defect D and that D concerned an application-runtime failure but did not concern installer construction or signing.\n\nAt 08:26 UTC on 14 September 2026, the completed release audit recorded defect S as the omission of the required signature from the Windows installer's updater executable and as the only Northstar Notes 4.2 RC3 defect concerning installer construction or signing.\n\nThe release manager then consulted the team-routing policy before assigning the blocker. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing routing policy is unchanged, and both contexts keep the decision bound to the single blocker D described by the deployment specialist for Northstar Notes 4.2 RC3. The two evidence spans are complete factual sentences, and the counterfactual changes one factual sentence so that D is an application-runtime defect while remaining consistent with S being the sole installer construction or signing defect. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction. However, the base context itself is internally inconsistent: it says S is the sole defect concerning installer construction or signing and later says D concerns installer signing, while D and S are separately designated and never identified as the same defect; therefore the overall constructed pair should be rejected despite the counterfactual context being coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"Operational handoff for the release-routing review:\\n\\nAt 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing.\\n\\nThe release audit defines defect S as the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3. The deployment specialist reported that the blocker remains open while its routing is decided.\\n\\nAt 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as concerning installer signing in Northstar Notes 4.2 RC3.\\n\\nPolicy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "At 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing."}, {"path": [], "text": "At 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as concerning installer signing in Northstar Notes 4.2 RC3."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing.", "negative_left": "At 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing.", "negative_right": "At 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as an application-runtime defect that did not concern installer construction or signing in Northstar Notes 4.2 RC3.", "right": "At 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as concerning installer signing in Northstar Notes 4.2 RC3."}, "verifier_independent_model": false}, "family": "scale-diverse-148-003", "id": "scale-diverse-148-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "Operational handoff for the release-routing review:\n\nAt 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing.\n\nThe release audit defines defect S as the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3. The deployment specialist reported that the blocker remains open while its routing is decided.\n\nAt 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as concerning installer signing in Northstar Notes 4.2 RC3.\n\nPolicy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing routing policy is unchanged, and both contexts keep the decision bound to the single blocker D described by the deployment specialist for Northstar Notes 4.2 RC3. The two evidence spans are complete factual sentences, and the counterfactual changes one factual sentence so that D is an application-runtime defect while remaining consistent with S being the sole installer construction or signing defect. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction. However, the base context itself is internally inconsistent: it says S is the sole defect concerning installer construction or signing and later says D concerns installer signing, while D and S are separately designated and never identified as the same defect; therefore the overall constructed pair should be rejected despite the counterfactual context being coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"Operational handoff for the release-routing review:\\n\\nAt 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing.\\n\\nThe release audit defines defect S as the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3. The deployment specialist reported that the blocker remains open while its routing is decided.\\n\\nAt 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as concerning installer signing in Northstar Notes 4.2 RC3.\\n\\nPolicy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "At 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing."}, {"path": [], "text": "At 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as concerning installer signing in Northstar Notes 4.2 RC3."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing.", "negative_left": "At 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing.", "negative_right": "At 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as an application-runtime defect that did not concern installer construction or signing in Northstar Notes 4.2 RC3.", "right": "At 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as concerning installer signing in Northstar Notes 4.2 RC3."}, "verifier_independent_model": false}, "family": "scale-diverse-148-003", "id": "scale-diverse-148-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "Operational handoff for the release-routing review:\n\nAt 14:35 UTC on 8 September 2026, the operational handoff for Northstar Notes 4.2 RC3 recorded exactly one blocker defect, designated D, and the release audit classified defect S as the sole defect concerning installer construction or signing.\n\nThe release audit defines defect S as the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3. The deployment specialist reported that the blocker remains open while its routing is decided.\n\nAt 14:41 UTC on 8 September 2026, the deployment specialist classified defect D as an application-runtime defect that did not concern installer construction or signing in Northstar Notes 4.2 RC3.\n\nPolicy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same release-gate policy and the request remains bound to Northstar Notes 6.2 RC3, its readiness, and the team responsible for any gate failure. The two focus spans are complete factual sentences. The counterfactual changes only the successful-rollback specification’s Build value from 6203 to 6204, creating a coherent observed-versus-required mismatch without contradictory duplicate observations. Neither context states a readiness label, answer code, classifier instruction, proposition ID, or label rationale; the responsibility assignment is ordinary evidence needed to answer the team portion of the request.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":\"Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\",\"evidence\":[\"At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\\\Software\\\\Northstar\\\\Notes\\\\Channel=\\\"RC3\\\" and HKCU\\\\Software\\\\Northstar\\\\Notes\\\\Build=\\\"6203\\\".\",\"At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\\\Software\\\\Northstar\\\\Notes\\\\Channel=\\\"RC3\\\" and HKCU\\\\Software\\\\Northstar\\\\Notes\\\\Build=\\\"6203\\\".\",\"The completed test report records that every supported-platform compatibility test passed and that every installer rollback other than the Windows MSI rollback succeeded.\",\"The defect review found no open release-blocking defect, and signature verification confirmed a valid signature on every release package.\",\"The release record documents the keyboard-shortcuts-table typo, classifies it as nonblocking, and notes that it still requires an approved release-note action.\",\"The responsibility register assigns the Northstar Notes deployment team to resolve any failure of the Windows MSI rollback gate.\"],\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\"."}, {"path": ["evidence", "1"], "text": "At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\"."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\".", "negative_left": "At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\".", "negative_right": "At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6204\".", "right": "At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\"."}, "verifier_independent_model": false}, "family": "scale-diverse-149-001", "id": "scale-diverse-149-001-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\".", "At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\".", "The completed test report records that every supported-platform compatibility test passed and that every installer rollback other than the Windows MSI rollback succeeded.", "The defect review found no open release-blocking defect, and signature verification confirmed a valid signature on every release package.", "The release record documents the keyboard-shortcuts-table typo, classifies it as nonblocking, and notes that it still requires an approved release-note action.", "The responsibility register assigns the Northstar Notes deployment team to resolve any failure of the Windows MSI rollback gate."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same release-gate policy and the request remains bound to Northstar Notes 6.2 RC3, its readiness, and the team responsible for any gate failure. The two focus spans are complete factual sentences. The counterfactual changes only the successful-rollback specification’s Build value from 6203 to 6204, creating a coherent observed-versus-required mismatch without contradictory duplicate observations. Neither context states a readiness label, answer code, classifier instruction, proposition ID, or label rationale; the responsibility assignment is ordinary evidence needed to answer the team portion of the request.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":\"Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\",\"evidence\":[\"At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\\\Software\\\\Northstar\\\\Notes\\\\Channel=\\\"RC3\\\" and HKCU\\\\Software\\\\Northstar\\\\Notes\\\\Build=\\\"6203\\\".\",\"At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\\\Software\\\\Northstar\\\\Notes\\\\Channel=\\\"RC3\\\" and HKCU\\\\Software\\\\Northstar\\\\Notes\\\\Build=\\\"6203\\\".\",\"The completed test report records that every supported-platform compatibility test passed and that every installer rollback other than the Windows MSI rollback succeeded.\",\"The defect review found no open release-blocking defect, and signature verification confirmed a valid signature on every release package.\",\"The release record documents the keyboard-shortcuts-table typo, classifies it as nonblocking, and notes that it still requires an approved release-note action.\",\"The responsibility register assigns the Northstar Notes deployment team to resolve any failure of the Windows MSI rollback gate.\"],\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\"."}, {"path": ["evidence", "1"], "text": "At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\"."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\".", "negative_left": "At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\".", "negative_right": "At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6204\".", "right": "At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\"."}, "verifier_independent_model": false}, "family": "scale-diverse-149-001", "id": "scale-diverse-149-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["At 14:37:22 UTC on 2026-08-19, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI consisted exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6203\".", "At 09:12:05 UTC on 2026-08-18, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified a complete post-rollback registration state consisting exclusively of the two registry values HKCU\\Software\\Northstar\\Notes\\Channel=\"RC3\" and HKCU\\Software\\Northstar\\Notes\\Build=\"6204\".", "The completed test report records that every supported-platform compatibility test passed and that every installer rollback other than the Windows MSI rollback succeeded.", "The defect review found no open release-blocking defect, and signature verification confirmed a valid signature on every release package.", "The release record documents the keyboard-shortcuts-table typo, classifies it as nonblocking, and notes that it still requires an approved release-note action.", "The responsibility register assigns the Northstar Notes deployment team to resolve any failure of the Windows MSI rollback gate."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing release policy and the same request concerning Northstar Notes 6.2 RC3 and the team responsible for any gate failure. The two focus spans are complete factual sentences. The counterfactual changes only the prescribed rollback-marker value from \"complete\" to \"verified\"; this creates a coherent observed-versus-required mismatch rather than contradictory duplicate measurements. Neither context contains a readiness answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":\"Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\",\"evidence\":[\"At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \\\"complete\\\").\",\"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \\\"complete\\\").\",\"Every supported-platform compatibility test passed, and every installer rollback other than the Windows MSI rollback succeeded.\",\"The defect register shows no open release-blocking defect, and validation confirms a valid signature on every release package.\",\"The release record documents the keyboard-shortcuts-table typo, classifies it as nonblocking, and notes that an approved release-note action is still required.\",\"The Northstar Notes deployment team is responsible for resolving any failure of the Windows MSI rollback gate.\"],\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\")."}, {"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\")."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\").", "negative_left": "At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\").", "negative_right": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"verified\").", "right": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\")."}, "verifier_independent_model": false}, "family": "scale-diverse-149-002", "id": "scale-diverse-149-002-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\").", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\").", "Every supported-platform compatibility test passed, and every installer rollback other than the Windows MSI rollback succeeded.", "The defect register shows no open release-blocking defect, and validation confirms a valid signature on every release package.", "The release record documents the keyboard-shortcuts-table typo, classifies it as nonblocking, and notes that an approved release-note action is still required.", "The Northstar Notes deployment team is responsible for resolving any failure of the Windows MSI rollback gate."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing release policy and the same request concerning Northstar Notes 6.2 RC3 and the team responsible for any gate failure. The two focus spans are complete factual sentences. The counterfactual changes only the prescribed rollback-marker value from \"complete\" to \"verified\"; this creates a coherent observed-versus-required mismatch rather than contradictory duplicate measurements. Neither context contains a readiness answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":\"Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\",\"evidence\":[\"At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \\\"complete\\\").\",\"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \\\"complete\\\").\",\"Every supported-platform compatibility test passed, and every installer rollback other than the Windows MSI rollback succeeded.\",\"The defect register shows no open release-blocking defect, and validation confirms a valid signature on every release package.\",\"The release record documents the keyboard-shortcuts-table typo, classifies it as nonblocking, and notes that an approved release-note action is still required.\",\"The Northstar Notes deployment team is responsible for resolving any failure of the Windows MSI rollback gate.\"],\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\")."}, {"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\")."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\").", "negative_left": "At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\").", "negative_right": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"verified\").", "right": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\")."}, "verifier_independent_model": false}, "family": "scale-diverse-149-002", "id": "scale-diverse-149-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["At 2026-09-14T18:42:11Z in rollback run NR-842, the complete observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"complete\").", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI, signed at 2026-09-13T09:16:00Z, prescribes exactly the tuple (product-registration keys: none; registered COM class identifiers: none; uninstall entry: absent; installer rollback marker: \"verified\").", "Every supported-platform compatibility test passed, and every installer rollback other than the Windows MSI rollback succeeded.", "The defect register shows no open release-blocking defect, and validation confirms a valid signature on every release package.", "The release record documents the keyboard-shortcuts-table typo, classifies it as nonblocking, and notes that an approved release-note action is still required.", "The Northstar Notes deployment team is responsible for resolving any failure of the Windows MSI rollback gate."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release-readiness criteria and remain bound to RC 4 and its gate decision. The two evidence spans are complete factual sentences. The counterfactual changes only record U-731 from passing to failing; this is coherent with the unchanged statements that the Ubuntu check failed without remediation and that the cleanup command was the only tested workaround, and it creates no duplicate or conflicting outcome. Neither context contains a readiness label, answer code, rule table, proposition identifier, output instruction, or explicit answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom a6 is factual rather than a policy classification. The base and counter assignments are jointly realizable while changing only whether the cleanup-command trial passes; a7 merely excludes other tested workarounds and need not assert that the cleanup command succeeds. Empty policy_evidence is correct because all governing readiness criteria and interpretive instructions occur in the retained questions object, while the original state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The Ubuntu required-platform check fails without remediation, the cleanup command does not produce a passing result, and no other tested workaround exists. Therefore a required-platform check fails without a tested workaround, which is sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes no open blocker/high-severity defects, passing core and required-platform checks through the tested Ubuntu workaround, verified installer signing/integrity, validated rollback, and exactly one open medium packaging-or-compatibility defect whose documentation approval remains pending. This is sufficient for level 1 and excludes the stated level-0 conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of open blocker defects for RC 4 is zero."}, {"id": "a2", "statement": "The number of open high-severity defects for RC 4 is zero."}, {"id": "a3", "statement": "Each of the 186 required smoke and regression tests for RC 4 passed."}, {"id": "a4", "statement": "Every required-platform check for RC 4 other than the supported-Ubuntu obsolete-tray-helper check passed."}, {"id": "a5", "statement": "Without remediation, the obsolete-tray-helper defect causes its required-platform check on the supported Ubuntu build for RC 4 to fail."}, {"id": "a6", "statement": "The cleanup-command trial for the obsolete-tray-helper defect produced a passing result for its required-platform check on the supported Ubuntu build for RC 4."}, {"id": "a7", "statement": "No tested workaround other than the cleanup command exists for the obsolete-tray-helper defect in RC 4."}, {"id": "a8", "statement": "The obsolete-tray-helper defect in RC 4 is open."}, {"id": "a9", "statement": "The obsolete-tray-helper defect in RC 4 has medium severity."}, {"id": "a10", "statement": "The obsolete-tray-helper defect in RC 4 is a compatibility defect."}, {"id": "a11", "statement": "Every open packaging or compatibility defect in RC 4 other than the obsolete-tray-helper defect has severity below medium."}, {"id": "a12", "statement": "Every RC 4 installer has the expected publisher identity."}, {"id": "a13", "statement": "Every RC 4 installer digest matches its build record."}, {"id": "a14", "statement": "The RC 4 removal package restored the prior version 7.3."}, {"id": "a15", "statement": "The RC 4 removal package retained every user file during restoration of version 7.3."}, {"id": "a16", "statement": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}], "base_state_json": "[{\"speaker\":\"Gate coordinator\",\"text\":\"At 13:30 UTC, RC 4 had zero open blocker defects and zero open high-severity defects. Each of its 186 required smoke and regression tests had passed.\"},{\"speaker\":\"Platform lead\",\"text\":\"Every required-platform check other than the supported-Ubuntu obsolete-tray-helper check passed. Without remediation, that defect caused its check to fail, and no tested workaround other than the cleanup command existed.\"},{\"speaker\":\"Test recorder\",\"text\":\"At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731.\"},{\"speaker\":\"Audit clerk\",\"text\":\"At 14:09 UTC on 12 September 2026, record U-731 was finalized with a passing outcome.\"},{\"speaker\":\"Triage lead\",\"text\":\"Triage kept the obsolete-tray-helper defect open as a medium-severity compatibility defect. Every other open packaging or compatibility defect in RC 4 was below medium severity.\"},{\"speaker\":\"Release auditor\",\"text\":\"Every RC 4 installer showed the expected publisher identity, and every digest matched its build record. The removal package restored prior version 7.3 and retained every user file. Final approval of the obsolete-tray-helper release-note documentation remained pending.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731."}, {"path": ["3", "text"], "text": "At 14:09 UTC on 12 September 2026, record U-731 was finalized with a passing outcome."}], "policy_evidence": [], "rules": [{"justification": "The required supported-Ubuntu platform check fails, the cleanup-command trial does not make it pass, and no other tested workaround exists; this satisfies the blocked criterion.", "target": "0", "when": [{"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "There are no blocker or high-severity defects; the core and required-platform gates pass, including the Ubuntu gate through a tested cleanup workaround; signing, integrity, and rollback are verified; and exactly one medium compatibility defect remains with documentation approval pending.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731.", "negative_left": "At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731.", "negative_right": "At 14:09 UTC on 12 September 2026, record U-731 was finalized with a failing outcome.", "right": "At 14:09 UTC on 12 September 2026, record U-731 was finalized with a passing outcome."}, "verifier_independent_model": false}, "family": "scale-diverse-150-001", "id": "scale-diverse-150-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Blocked: At least one blocker/high-severity defect is open, a required platform or core test fails without a tested workaround, installer signing/integrity is unverified, or rollback has not been validated.", "1 — Conditionally ready: No blocker/high-severity defects remain; core tests, required-platform checks, installer signing/integrity, and rollback pass, but exactly one medium packaging or compatibility defect remains with a tested workaround that still needs final documentation or release-operations approval.", "2 — Fully ready: All core gates pass, rollback and installer integrity are validated, no medium-or-higher packaging or compatibility defects remain, and all release notes and operational approvals are final."], "instructions": "Assign the release-readiness level. Interpret evidence by meaning rather than requiring exact rubric phrases; for example, restoring the prior version without losing files constitutes validated rollback, and matching publisher identity/digests constitutes signing and integrity verification.", "type": "score"}}, "state": [{"speaker": "Gate coordinator", "text": "At 13:30 UTC, RC 4 had zero open blocker defects and zero open high-severity defects. Each of its 186 required smoke and regression tests had passed."}, {"speaker": "Platform lead", "text": "Every required-platform check other than the supported-Ubuntu obsolete-tray-helper check passed. Without remediation, that defect caused its check to fail, and no tested workaround other than the cleanup command existed."}, {"speaker": "Test recorder", "text": "At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731."}, {"speaker": "Audit clerk", "text": "At 14:09 UTC on 12 September 2026, record U-731 was finalized with a passing outcome."}, {"speaker": "Triage lead", "text": "Triage kept the obsolete-tray-helper defect open as a medium-severity compatibility defect. Every other open packaging or compatibility defect in RC 4 was below medium severity."}, {"speaker": "Release auditor", "text": "Every RC 4 installer showed the expected publisher identity, and every digest matched its build record. The removal package restored prior version 7.3 and retained every user file. Final approval of the obsolete-tray-helper release-note documentation remained pending."}]}, "method": "c2d", "provenance": {"source_id": "diverse-150", "source_is_synthetic": true, "source_sha256": "debcb2db8f1984d1db79d80259af413a1a60045f28238c51df71d3deaccb8e61", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release-readiness criteria and remain bound to RC 4 and its gate decision. The two evidence spans are complete factual sentences. The counterfactual changes only record U-731 from passing to failing; this is coherent with the unchanged statements that the Ubuntu check failed without remediation and that the cleanup command was the only tested workaround, and it creates no duplicate or conflicting outcome. Neither context contains a readiness label, answer code, rule table, proposition identifier, output instruction, or explicit answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom a6 is factual rather than a policy classification. The base and counter assignments are jointly realizable while changing only whether the cleanup-command trial passes; a7 merely excludes other tested workarounds and need not assert that the cleanup command succeeds. Empty policy_evidence is correct because all governing readiness criteria and interpretive instructions occur in the retained questions object, while the original state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The Ubuntu required-platform check fails without remediation, the cleanup command does not produce a passing result, and no other tested workaround exists. Therefore a required-platform check fails without a tested workaround, which is sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes no open blocker/high-severity defects, passing core and required-platform checks through the tested Ubuntu workaround, verified installer signing/integrity, validated rollback, and exactly one open medium packaging-or-compatibility defect whose documentation approval remains pending. This is sufficient for level 1 and excludes the stated level-0 conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of open blocker defects for RC 4 is zero."}, {"id": "a2", "statement": "The number of open high-severity defects for RC 4 is zero."}, {"id": "a3", "statement": "Each of the 186 required smoke and regression tests for RC 4 passed."}, {"id": "a4", "statement": "Every required-platform check for RC 4 other than the supported-Ubuntu obsolete-tray-helper check passed."}, {"id": "a5", "statement": "Without remediation, the obsolete-tray-helper defect causes its required-platform check on the supported Ubuntu build for RC 4 to fail."}, {"id": "a6", "statement": "The cleanup-command trial for the obsolete-tray-helper defect produced a passing result for its required-platform check on the supported Ubuntu build for RC 4."}, {"id": "a7", "statement": "No tested workaround other than the cleanup command exists for the obsolete-tray-helper defect in RC 4."}, {"id": "a8", "statement": "The obsolete-tray-helper defect in RC 4 is open."}, {"id": "a9", "statement": "The obsolete-tray-helper defect in RC 4 has medium severity."}, {"id": "a10", "statement": "The obsolete-tray-helper defect in RC 4 is a compatibility defect."}, {"id": "a11", "statement": "Every open packaging or compatibility defect in RC 4 other than the obsolete-tray-helper defect has severity below medium."}, {"id": "a12", "statement": "Every RC 4 installer has the expected publisher identity."}, {"id": "a13", "statement": "Every RC 4 installer digest matches its build record."}, {"id": "a14", "statement": "The RC 4 removal package restored the prior version 7.3."}, {"id": "a15", "statement": "The RC 4 removal package retained every user file during restoration of version 7.3."}, {"id": "a16", "statement": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}], "base_state_json": "[{\"speaker\":\"Gate coordinator\",\"text\":\"At 13:30 UTC, RC 4 had zero open blocker defects and zero open high-severity defects. Each of its 186 required smoke and regression tests had passed.\"},{\"speaker\":\"Platform lead\",\"text\":\"Every required-platform check other than the supported-Ubuntu obsolete-tray-helper check passed. Without remediation, that defect caused its check to fail, and no tested workaround other than the cleanup command existed.\"},{\"speaker\":\"Test recorder\",\"text\":\"At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731.\"},{\"speaker\":\"Audit clerk\",\"text\":\"At 14:09 UTC on 12 September 2026, record U-731 was finalized with a passing outcome.\"},{\"speaker\":\"Triage lead\",\"text\":\"Triage kept the obsolete-tray-helper defect open as a medium-severity compatibility defect. Every other open packaging or compatibility defect in RC 4 was below medium severity.\"},{\"speaker\":\"Release auditor\",\"text\":\"Every RC 4 installer showed the expected publisher identity, and every digest matched its build record. The removal package restored prior version 7.3 and retained every user file. Final approval of the obsolete-tray-helper release-note documentation remained pending.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731."}, {"path": ["3", "text"], "text": "At 14:09 UTC on 12 September 2026, record U-731 was finalized with a passing outcome."}], "policy_evidence": [], "rules": [{"justification": "The required supported-Ubuntu platform check fails, the cleanup-command trial does not make it pass, and no other tested workaround exists; this satisfies the blocked criterion.", "target": "0", "when": [{"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "There are no blocker or high-severity defects; the core and required-platform gates pass, including the Ubuntu gate through a tested cleanup workaround; signing, integrity, and rollback are verified; and exactly one medium compatibility defect remains with documentation approval pending.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731.", "negative_left": "At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731.", "negative_right": "At 14:09 UTC on 12 September 2026, record U-731 was finalized with a failing outcome.", "right": "At 14:09 UTC on 12 September 2026, record U-731 was finalized with a passing outcome."}, "verifier_independent_model": false}, "family": "scale-diverse-150-001", "id": "scale-diverse-150-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Blocked: At least one blocker/high-severity defect is open, a required platform or core test fails without a tested workaround, installer signing/integrity is unverified, or rollback has not been validated.", "1 — Conditionally ready: No blocker/high-severity defects remain; core tests, required-platform checks, installer signing/integrity, and rollback pass, but exactly one medium packaging or compatibility defect remains with a tested workaround that still needs final documentation or release-operations approval.", "2 — Fully ready: All core gates pass, rollback and installer integrity are validated, no medium-or-higher packaging or compatibility defects remain, and all release notes and operational approvals are final."], "instructions": "Assign the release-readiness level. Interpret evidence by meaning rather than requiring exact rubric phrases; for example, restoring the prior version without losing files constitutes validated rollback, and matching publisher identity/digests constitutes signing and integrity verification.", "type": "score"}}, "state": [{"speaker": "Gate coordinator", "text": "At 13:30 UTC, RC 4 had zero open blocker defects and zero open high-severity defects. Each of its 186 required smoke and regression tests had passed."}, {"speaker": "Platform lead", "text": "Every required-platform check other than the supported-Ubuntu obsolete-tray-helper check passed. Without remediation, that defect caused its check to fail, and no tested workaround other than the cleanup command existed."}, {"speaker": "Test recorder", "text": "At 14:08 UTC on 12 September 2026, the cleanup-command trial for the obsolete-tray-helper defect ran its required-platform check exactly once on the supported Ubuntu build for RC 4 and stored the sole outcome in record U-731."}, {"speaker": "Audit clerk", "text": "At 14:09 UTC on 12 September 2026, record U-731 was finalized with a failing outcome."}, {"speaker": "Triage lead", "text": "Triage kept the obsolete-tray-helper defect open as a medium-severity compatibility defect. Every other open packaging or compatibility defect in RC 4 was below medium severity."}, {"speaker": "Release auditor", "text": "Every RC 4 installer showed the expected publisher identity, and every digest matched its build record. The removal package restored prior version 7.3 and retained every user file. Final approval of the obsolete-tray-helper release-note documentation remained pending."}]}, "method": "c2d", "provenance": {"source_id": "diverse-150", "source_is_synthetic": true, "source_sha256": "debcb2db8f1984d1db79d80259af413a1a60045f28238c51df71d3deaccb8e61", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the release-readiness policy and remain scoped to RC 4, its supported-Ubuntu required-platform check, and the identified compatibility defect. The two evidence spans are complete factual sentences. The counterfactual changes only H-27’s recorded check result from passed to failed; this does not contradict the statement that the cleanup command is the only tested workaround, because that statement does not assert that every application or trial succeeds. Neither context contains a readiness score, gold answer, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom a6 is factual rather than a policy classification. The base and counter assignments are jointly realizable while changing only whether the cleanup-command trial passes; a7 merely excludes other tested workarounds and need not assert that the cleanup command succeeds. Empty policy_evidence is correct because all governing readiness criteria and interpretive instructions occur in the retained questions object, while the original state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The Ubuntu required-platform check fails without remediation, the cleanup command does not produce a passing result, and no other tested workaround exists. Therefore a required-platform check fails without a tested workaround, which is sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes no open blocker/high-severity defects, passing core and required-platform checks through the tested Ubuntu workaround, verified installer signing/integrity, validated rollback, and exactly one open medium packaging-or-compatibility defect whose documentation approval remains pending. This is sufficient for level 1 and excludes the stated level-0 conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of open blocker defects for RC 4 is zero."}, {"id": "a2", "statement": "The number of open high-severity defects for RC 4 is zero."}, {"id": "a3", "statement": "Each of the 186 required smoke and regression tests for RC 4 passed."}, {"id": "a4", "statement": "Every required-platform check for RC 4 other than the supported-Ubuntu obsolete-tray-helper check passed."}, {"id": "a5", "statement": "Without remediation, the obsolete-tray-helper defect causes its required-platform check on the supported Ubuntu build for RC 4 to fail."}, {"id": "a6", "statement": "The cleanup-command trial for the obsolete-tray-helper defect produced a passing result for its required-platform check on the supported Ubuntu build for RC 4."}, {"id": "a7", "statement": "No tested workaround other than the cleanup command exists for the obsolete-tray-helper defect in RC 4."}, {"id": "a8", "statement": "The obsolete-tray-helper defect in RC 4 is open."}, {"id": "a9", "statement": "The obsolete-tray-helper defect in RC 4 has medium severity."}, {"id": "a10", "statement": "The obsolete-tray-helper defect in RC 4 is a compatibility defect."}, {"id": "a11", "statement": "Every open packaging or compatibility defect in RC 4 other than the obsolete-tray-helper defect has severity below medium."}, {"id": "a12", "statement": "Every RC 4 installer has the expected publisher identity."}, {"id": "a13", "statement": "Every RC 4 installer digest matches its build record."}, {"id": "a14", "statement": "The RC 4 removal package restored the prior version 7.3."}, {"id": "a15", "statement": "The RC 4 removal package retained every user file during restoration of version 7.3."}, {"id": "a16", "statement": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed. Every required-platform check except the supported-Ubuntu obsolete-tray-helper check passed; without remediation, that defect makes the excepted check fail.\"},{\"speaker\":\"Handoff operator\",\"text\":\"At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build.\"},{\"speaker\":\"Records analyst\",\"text\":\"The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as passed.\"},{\"speaker\":\"Release coordinator\",\"text\":\"No tested workaround other than the cleanup command exists for this defect. The defect is open, medium-severity, and a compatibility defect; every other open packaging or compatibility defect is below medium. Every RC 4 installer has the expected publisher identity, and every installer digest matches its build record. The removal package restored version 7.3 and retained every user file. Final approval of the defect's release-note documentation is pending.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build."}, {"path": ["2", "text"], "text": "The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as passed."}], "policy_evidence": [], "rules": [{"justification": "The required supported-Ubuntu platform check fails, the cleanup-command trial does not make it pass, and no other tested workaround exists; this satisfies the blocked criterion.", "target": "0", "when": [{"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "There are no blocker or high-severity defects; the core and required-platform gates pass, including the Ubuntu gate through a tested cleanup workaround; signing, integrity, and rollback are verified; and exactly one medium compatibility defect remains with documentation approval pending.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build.", "negative_left": "At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build.", "negative_right": "The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as failed.", "right": "The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as passed."}, "verifier_independent_model": false}, "family": "scale-diverse-150-003", "id": "scale-diverse-150-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Blocked: At least one blocker/high-severity defect is open, a required platform or core test fails without a tested workaround, installer signing/integrity is unverified, or rollback has not been validated.", "1 — Conditionally ready: No blocker/high-severity defects remain; core tests, required-platform checks, installer signing/integrity, and rollback pass, but exactly one medium packaging or compatibility defect remains with a tested workaround that still needs final documentation or release-operations approval.", "2 — Fully ready: All core gates pass, rollback and installer integrity are validated, no medium-or-higher packaging or compatibility defects remain, and all release notes and operational approvals are final."], "instructions": "Assign the release-readiness level. Interpret evidence by meaning rather than requiring exact rubric phrases; for example, restoring the prior version without losing files constitutes validated rollback, and matching publisher identity/digests constitutes signing and integrity verification.", "type": "score"}}, "state": [{"speaker": "Release manager", "text": "RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed. Every required-platform check except the supported-Ubuntu obsolete-tray-helper check passed; without remediation, that defect makes the excepted check fail."}, {"speaker": "Handoff operator", "text": "At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build."}, {"speaker": "Records analyst", "text": "The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as passed."}, {"speaker": "Release coordinator", "text": "No tested workaround other than the cleanup command exists for this defect. The defect is open, medium-severity, and a compatibility defect; every other open packaging or compatibility defect is below medium. Every RC 4 installer has the expected publisher identity, and every installer digest matches its build record. The removal package restored version 7.3 and retained every user file. Final approval of the defect's release-note documentation is pending."}]}, "method": "c2d", "provenance": {"source_id": "diverse-150", "source_is_synthetic": true, "source_sha256": "debcb2db8f1984d1db79d80259af413a1a60045f28238c51df71d3deaccb8e61", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the release-readiness policy and remain scoped to RC 4, its supported-Ubuntu required-platform check, and the identified compatibility defect. The two evidence spans are complete factual sentences. The counterfactual changes only H-27’s recorded check result from passed to failed; this does not contradict the statement that the cleanup command is the only tested workaround, because that statement does not assert that every application or trial succeeds. Neither context contains a readiness score, gold answer, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom a6 is factual rather than a policy classification. The base and counter assignments are jointly realizable while changing only whether the cleanup-command trial passes; a7 merely excludes other tested workarounds and need not assert that the cleanup command succeeds. Empty policy_evidence is correct because all governing readiness criteria and interpretive instructions occur in the retained questions object, while the original state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The Ubuntu required-platform check fails without remediation, the cleanup command does not produce a passing result, and no other tested workaround exists. Therefore a required-platform check fails without a tested workaround, which is sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes no open blocker/high-severity defects, passing core and required-platform checks through the tested Ubuntu workaround, verified installer signing/integrity, validated rollback, and exactly one open medium packaging-or-compatibility defect whose documentation approval remains pending. This is sufficient for level 1 and excludes the stated level-0 conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of open blocker defects for RC 4 is zero."}, {"id": "a2", "statement": "The number of open high-severity defects for RC 4 is zero."}, {"id": "a3", "statement": "Each of the 186 required smoke and regression tests for RC 4 passed."}, {"id": "a4", "statement": "Every required-platform check for RC 4 other than the supported-Ubuntu obsolete-tray-helper check passed."}, {"id": "a5", "statement": "Without remediation, the obsolete-tray-helper defect causes its required-platform check on the supported Ubuntu build for RC 4 to fail."}, {"id": "a6", "statement": "The cleanup-command trial for the obsolete-tray-helper defect produced a passing result for its required-platform check on the supported Ubuntu build for RC 4."}, {"id": "a7", "statement": "No tested workaround other than the cleanup command exists for the obsolete-tray-helper defect in RC 4."}, {"id": "a8", "statement": "The obsolete-tray-helper defect in RC 4 is open."}, {"id": "a9", "statement": "The obsolete-tray-helper defect in RC 4 has medium severity."}, {"id": "a10", "statement": "The obsolete-tray-helper defect in RC 4 is a compatibility defect."}, {"id": "a11", "statement": "Every open packaging or compatibility defect in RC 4 other than the obsolete-tray-helper defect has severity below medium."}, {"id": "a12", "statement": "Every RC 4 installer has the expected publisher identity."}, {"id": "a13", "statement": "Every RC 4 installer digest matches its build record."}, {"id": "a14", "statement": "The RC 4 removal package restored the prior version 7.3."}, {"id": "a15", "statement": "The RC 4 removal package retained every user file during restoration of version 7.3."}, {"id": "a16", "statement": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed. Every required-platform check except the supported-Ubuntu obsolete-tray-helper check passed; without remediation, that defect makes the excepted check fail.\"},{\"speaker\":\"Handoff operator\",\"text\":\"At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build.\"},{\"speaker\":\"Records analyst\",\"text\":\"The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as passed.\"},{\"speaker\":\"Release coordinator\",\"text\":\"No tested workaround other than the cleanup command exists for this defect. The defect is open, medium-severity, and a compatibility defect; every other open packaging or compatibility defect is below medium. Every RC 4 installer has the expected publisher identity, and every installer digest matches its build record. The removal package restored version 7.3 and retained every user file. Final approval of the defect's release-note documentation is pending.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build."}, {"path": ["2", "text"], "text": "The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as passed."}], "policy_evidence": [], "rules": [{"justification": "The required supported-Ubuntu platform check fails, the cleanup-command trial does not make it pass, and no other tested workaround exists; this satisfies the blocked criterion.", "target": "0", "when": [{"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "There are no blocker or high-severity defects; the core and required-platform gates pass, including the Ubuntu gate through a tested cleanup workaround; signing, integrity, and rollback are verified; and exactly one medium compatibility defect remains with documentation approval pending.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build.", "negative_left": "At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build.", "negative_right": "The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as failed.", "right": "The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as passed."}, "verifier_independent_model": false}, "family": "scale-diverse-150-003", "id": "scale-diverse-150-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Blocked: At least one blocker/high-severity defect is open, a required platform or core test fails without a tested workaround, installer signing/integrity is unverified, or rollback has not been validated.", "1 — Conditionally ready: No blocker/high-severity defects remain; core tests, required-platform checks, installer signing/integrity, and rollback pass, but exactly one medium packaging or compatibility defect remains with a tested workaround that still needs final documentation or release-operations approval.", "2 — Fully ready: All core gates pass, rollback and installer integrity are validated, no medium-or-higher packaging or compatibility defects remain, and all release notes and operational approvals are final."], "instructions": "Assign the release-readiness level. Interpret evidence by meaning rather than requiring exact rubric phrases; for example, restoring the prior version without losing files constitutes validated rollback, and matching publisher identity/digests constitutes signing and integrity verification.", "type": "score"}}, "state": [{"speaker": "Release manager", "text": "RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed. Every required-platform check except the supported-Ubuntu obsolete-tray-helper check passed; without remediation, that defect makes the excepted check fail."}, {"speaker": "Handoff operator", "text": "At 2026-09-15T14:32:00Z, operational-handoff trial H-27 applied the cleanup command for RC 4's obsolete-tray-helper defect and then ran the defect's required-platform check on the supported Ubuntu build."}, {"speaker": "Records analyst", "text": "The final result record for operational-handoff trial H-27 marks the check executed at 2026-09-15T14:32:00Z as failed."}, {"speaker": "Release coordinator", "text": "No tested workaround other than the cleanup command exists for this defect. The defect is open, medium-severity, and a compatibility defect; every other open packaging or compatibility defect is below medium. Every RC 4 installer has the expected publisher identity, and every installer digest matches its build record. The removal package restored version 7.3 and retained every user file. Final approval of the defect's release-note documentation is pending."}]}, "method": "c2d", "provenance": {"source_id": "diverse-150", "source_is_synthetic": true, "source_sha256": "debcb2db8f1984d1db79d80259af413a1a60045f28238c51df71d3deaccb8e61", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing return-eligibility and hazard-priority policies without adding exceptions, priorities, or missing-evidence defaults. The address, Zone R4, Tuesday recycling claim, reporting deadline, placement-time requirement, hazard assessment, request scope, and decision path remain fixed. The two focus spans are complete factual sentences describing the exhaustive placement-evidence set and its timing limitations. The counterfactual changes only the contents of PX-314 to provide continuous curbside evidence spanning 7:00 a.m.; this is coherent with the unchanged statement that PX-314 and PX-882 are the sole submitted records and creates no duplicate or contradictory measurement. Neither full context contains a gold option, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"Municipal address records place 18 Larch Court in Zone R4.\",\"The published route calendar lists recycling for Zone R4 on Tuesday.\",\"The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart.\",\"Record PX-314 begins at 7:06 a.m., record PX-882 is a still image captured at 7:13 a.m., and neither record contains imagery, metadata, or testimony concerning any time before 7:06 a.m.\",\"The missed-pickup report was received Wednesday at 8:20 a.m., within the 48-hour reporting deadline.\",\"The supervisor’s hazard checklist records a clear travel lane, contained waste without leakage, and no visible pests.\"],\"request\":\"Choose the correct routing and response classification.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart."}, {"path": ["evidence", "3"], "text": "Record PX-314 begins at 7:06 a.m., record PX-882 is a still image captured at 7:13 a.m., and neither record contains imagery, metadata, or testimony concerning any time before 7:06 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart.", "negative_left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart.", "negative_right": "Record PX-314 is an authenticated continuous video from 6:48 a.m. through 7:04 a.m. showing the blue cart stationary curbside throughout, and record PX-882 is a still image captured at 7:13 a.m.", "right": "Record PX-314 begins at 7:06 a.m., record PX-882 is a still image captured at 7:13 a.m., and neither record contains imagery, metadata, or testimony concerning any time before 7:06 a.m."}, "verifier_independent_model": false}, "family": "scale-diverse-151-001", "id": "scale-diverse-151-001-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["Municipal address records place 18 Larch Court in Zone R4.", "The published route calendar lists recycling for Zone R4 on Tuesday.", "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart.", "Record PX-314 begins at 7:06 a.m., record PX-882 is a still image captured at 7:13 a.m., and neither record contains imagery, metadata, or testimony concerning any time before 7:06 a.m.", "The missed-pickup report was received Wednesday at 8:20 a.m., within the 48-hour reporting deadline.", "The supervisor’s hazard checklist records a clear travel lane, contained waste without leakage, and no visible pests."], "request": "Choose the correct routing and response classification."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing return-eligibility and hazard-priority policies without adding exceptions, priorities, or missing-evidence defaults. The address, Zone R4, Tuesday recycling claim, reporting deadline, placement-time requirement, hazard assessment, request scope, and decision path remain fixed. The two focus spans are complete factual sentences describing the exhaustive placement-evidence set and its timing limitations. The counterfactual changes only the contents of PX-314 to provide continuous curbside evidence spanning 7:00 a.m.; this is coherent with the unchanged statement that PX-314 and PX-882 are the sole submitted records and creates no duplicate or contradictory measurement. Neither full context contains a gold option, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"Municipal address records place 18 Larch Court in Zone R4.\",\"The published route calendar lists recycling for Zone R4 on Tuesday.\",\"The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart.\",\"Record PX-314 begins at 7:06 a.m., record PX-882 is a still image captured at 7:13 a.m., and neither record contains imagery, metadata, or testimony concerning any time before 7:06 a.m.\",\"The missed-pickup report was received Wednesday at 8:20 a.m., within the 48-hour reporting deadline.\",\"The supervisor’s hazard checklist records a clear travel lane, contained waste without leakage, and no visible pests.\"],\"request\":\"Choose the correct routing and response classification.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart."}, {"path": ["evidence", "3"], "text": "Record PX-314 begins at 7:06 a.m., record PX-882 is a still image captured at 7:13 a.m., and neither record contains imagery, metadata, or testimony concerning any time before 7:06 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart.", "negative_left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart.", "negative_right": "Record PX-314 is an authenticated continuous video from 6:48 a.m. through 7:04 a.m. showing the blue cart stationary curbside throughout, and record PX-882 is a still image captured at 7:13 a.m.", "right": "Record PX-314 begins at 7:06 a.m., record PX-882 is a still image captured at 7:13 a.m., and neither record contains imagery, metadata, or testimony concerning any time before 7:06 a.m."}, "verifier_independent_model": false}, "family": "scale-diverse-151-001", "id": "scale-diverse-151-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["Municipal address records place 18 Larch Court in Zone R4.", "The published route calendar lists recycling for Zone R4 on Tuesday.", "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of records PX-314 and PX-882 concerning the blue cart.", "Record PX-314 is an authenticated continuous video from 6:48 a.m. through 7:04 a.m. showing the blue cart stationary curbside throughout, and record PX-882 is a still image captured at 7:13 a.m.", "The missed-pickup report was received Wednesday at 8:20 a.m., within the 48-hour reporting deadline.", "The supervisor’s hazard checklist records a clear travel lane, contained waste without leakage, and no visible pests."], "request": "Choose the correct routing and response classification."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same return-eligibility and hazard-priority policies, request the same routing and response classification, and preserve the claim’s address, Tuesday recycling stream, and relevant timing bindings. The two focus spans are complete factual sentences; the counterfactual changes only PX-41’s observed placement, remains consistent with the stated two-photo evidence set and PX-86, and includes no duplicate contradictory measurement or assertion. Neither context embeds an answer option, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The municipal address register places 18 Larch Court in Zone R4.\",\"The final collection calendar lists recycling as Zone R4’s scheduled waste stream for Tuesday.\",\"The missed-pickup report concerning Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline.\",\"The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86.\",\"Photograph PX-41 shows the blue cart inside the closed garage at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday.\",\"The site review documents no blocked travel lane or other obstruction associated with the claim.\",\"The same review records no leaking waste and no visible pests.\"],\"request\":\"Determine the routing and response classification for the missed-pickup claim.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86."}, {"path": ["evidence", "4"], "text": "Photograph PX-41 shows the blue cart inside the closed garage at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86.", "negative_left": "The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86.", "negative_right": "Photograph PX-41 shows the blue cart at the curb at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday.", "right": "Photograph PX-41 shows the blue cart inside the closed garage at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-151-002", "id": "scale-diverse-151-002-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The municipal address register places 18 Larch Court in Zone R4.", "The final collection calendar lists recycling as Zone R4’s scheduled waste stream for Tuesday.", "The missed-pickup report concerning Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline.", "The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86.", "Photograph PX-41 shows the blue cart inside the closed garage at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday.", "The site review documents no blocked travel lane or other obstruction associated with the claim.", "The same review records no leaking waste and no visible pests."], "request": "Determine the routing and response classification for the missed-pickup claim."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same return-eligibility and hazard-priority policies, request the same routing and response classification, and preserve the claim’s address, Tuesday recycling stream, and relevant timing bindings. The two focus spans are complete factual sentences; the counterfactual changes only PX-41’s observed placement, remains consistent with the stated two-photo evidence set and PX-86, and includes no duplicate contradictory measurement or assertion. Neither context embeds an answer option, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The municipal address register places 18 Larch Court in Zone R4.\",\"The final collection calendar lists recycling as Zone R4’s scheduled waste stream for Tuesday.\",\"The missed-pickup report concerning Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline.\",\"The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86.\",\"Photograph PX-41 shows the blue cart inside the closed garage at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday.\",\"The site review documents no blocked travel lane or other obstruction associated with the claim.\",\"The same review records no leaking waste and no visible pests.\"],\"request\":\"Determine the routing and response classification for the missed-pickup claim.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86."}, {"path": ["evidence", "4"], "text": "Photograph PX-41 shows the blue cart inside the closed garage at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86.", "negative_left": "The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86.", "negative_right": "Photograph PX-41 shows the blue cart at the curb at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday.", "right": "Photograph PX-41 shows the blue cart inside the closed garage at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-151-002", "id": "scale-diverse-151-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The municipal address register places 18 Larch Court in Zone R4.", "The final collection calendar lists recycling as Zone R4’s scheduled waste stream for Tuesday.", "The missed-pickup report concerning Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline.", "The complete set of items submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court consists of two single-frame photographs identified as PX-41 and PX-86.", "Photograph PX-41 shows the blue cart at the curb at its authenticated capture time of 6:54 a.m. on Tuesday, and photograph PX-86 shows the blue cart at the curb at its authenticated capture time of 7:13 a.m. that Tuesday.", "The site review documents no blocked travel lane or other obstruction associated with the claim.", "The same review records no leaking waste and no visible pests."], "request": "Determine the routing and response classification for the missed-pickup claim."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy, request, address, waste stream, claim day, and placement-time binding, while the unchanged questions preserve all decision criteria. The two focus spans are complete factual sentences. The counterfactual coherently changes the assessment so that one submitted item verifies timely curb placement while the other does not; these statements are not contradictory. Neither context embeds a gold answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The route register places 18 Larch Court in Zone R4 and lists recycling as that zone’s scheduled Tuesday collection stream.\",\"The missed-pickup report for Tuesday recycling at 18 Larch Court was received Wednesday at 8:20 a.m., within the 48-hour reporting deadline.\",\"The inspection record states that there was no obstruction or blocked travel lane, no leaking waste, and no visible pests.\",\"At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83.\",\"The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that neither item verifies the blue cart was curbside by 7:00 a.m. on Tuesday.\"],\"request\":\"Choose the correct routing and response classification.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83."}, {"path": ["evidence", "4"], "text": "The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that neither item verifies the blue cart was curbside by 7:00 a.m. on Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83.", "negative_left": "At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83.", "negative_right": "The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that photograph PX-47 verifies the blue cart was curbside by 7:00 a.m. on Tuesday, while sensor log SL-83 does not.", "right": "The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that neither item verifies the blue cart was curbside by 7:00 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-151-003", "id": "scale-diverse-151-003-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The route register places 18 Larch Court in Zone R4 and lists recycling as that zone’s scheduled Tuesday collection stream.", "The missed-pickup report for Tuesday recycling at 18 Larch Court was received Wednesday at 8:20 a.m., within the 48-hour reporting deadline.", "The inspection record states that there was no obstruction or blocked travel lane, no leaking waste, and no visible pests.", "At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83.", "The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that neither item verifies the blue cart was curbside by 7:00 a.m. on Tuesday."], "request": "Choose the correct routing and response classification."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy, request, address, waste stream, claim day, and placement-time binding, while the unchanged questions preserve all decision criteria. The two focus spans are complete factual sentences. The counterfactual coherently changes the assessment so that one submitted item verifies timely curb placement while the other does not; these statements are not contradictory. Neither context embeds a gold answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The route register places 18 Larch Court in Zone R4 and lists recycling as that zone’s scheduled Tuesday collection stream.\",\"The missed-pickup report for Tuesday recycling at 18 Larch Court was received Wednesday at 8:20 a.m., within the 48-hour reporting deadline.\",\"The inspection record states that there was no obstruction or blocked travel lane, no leaking waste, and no visible pests.\",\"At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83.\",\"The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that neither item verifies the blue cart was curbside by 7:00 a.m. on Tuesday.\"],\"request\":\"Choose the correct routing and response classification.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83."}, {"path": ["evidence", "4"], "text": "The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that neither item verifies the blue cart was curbside by 7:00 a.m. on Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83.", "negative_left": "At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83.", "negative_right": "The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that photograph PX-47 verifies the blue cart was curbside by 7:00 a.m. on Tuesday, while sensor log SL-83 does not.", "right": "The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that neither item verifies the blue cart was curbside by 7:00 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-151-003", "id": "scale-diverse-151-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The route register places 18 Larch Court in Zone R4 and lists recycling as that zone’s scheduled Tuesday collection stream.", "The missed-pickup report for Tuesday recycling at 18 Larch Court was received Wednesday at 8:20 a.m., within the 48-hour reporting deadline.", "The inspection record states that there was no obstruction or blocked travel lane, no leaking waste, and no visible pests.", "At 10:14 a.m. on Wednesday, the complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consisted solely of photograph PX-47 and sensor log SL-83.", "The finalized evidence assessment for photograph PX-47 and sensor log SL-83 records that photograph PX-47 verifies the blue cart was curbside by 7:00 a.m. on Tuesday, while sensor log SL-83 does not."], "request": "Choose the correct routing and response classification."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original sanitation routing and hazard-priority policy, while the unchanged questions object preserves the full choice criteria and instructions. The address, Tuesday recycling claim, request, and relevant evidence paths remain bound consistently. Both focus spans are complete factual sentences. The counterfactual coherently changes the review result for LC-41 without conflicting with the exactly-two-item inventory; LC-42 can remain non-verifying while LC-41 verifies placement. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The municipal address registry lists 18 Larch Court in Zone R4.\",\"The published collection calendar records recycling service for Zone R4 on Tuesday.\",\"The claim log records receipt of the Tuesday recycling report within the 48-hour reporting deadline.\",\"The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42.\",\"The completed item-review record marks both file LC-41 and file LC-42 as not verifying that the blue cart was curbside by 7:00 a.m.\",\"The intake record and supervisor review explicitly document no obstruction, no leaking waste, and no visible pests for this claim.\",\"The records clerk confirmed that the address entry, collection calendar, claim log, evidence inventory, item review, and hazard review all belong to the same Tuesday recycling claim.\"],\"request\":\"Choose the correct routing and response classification.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42."}, {"path": ["evidence", "4"], "text": "The completed item-review record marks both file LC-41 and file LC-42 as not verifying that the blue cart was curbside by 7:00 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42.", "negative_left": "The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42.", "negative_right": "The completed item-review record marks file LC-41 as verifying that the blue cart was curbside by 7:00 a.m. and file LC-42 as not verifying it.", "right": "The completed item-review record marks both file LC-41 and file LC-42 as not verifying that the blue cart was curbside by 7:00 a.m."}, "verifier_independent_model": false}, "family": "scale-diverse-151-004", "id": "scale-diverse-151-004-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The municipal address registry lists 18 Larch Court in Zone R4.", "The published collection calendar records recycling service for Zone R4 on Tuesday.", "The claim log records receipt of the Tuesday recycling report within the 48-hour reporting deadline.", "The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42.", "The completed item-review record marks both file LC-41 and file LC-42 as not verifying that the blue cart was curbside by 7:00 a.m.", "The intake record and supervisor review explicitly document no obstruction, no leaking waste, and no visible pests for this claim.", "The records clerk confirmed that the address entry, collection calendar, claim log, evidence inventory, item review, and hazard review all belong to the same Tuesday recycling claim."], "request": "Choose the correct routing and response classification."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original sanitation routing and hazard-priority policy, while the unchanged questions object preserves the full choice criteria and instructions. The address, Tuesday recycling claim, request, and relevant evidence paths remain bound consistently. Both focus spans are complete factual sentences. The counterfactual coherently changes the review result for LC-41 without conflicting with the exactly-two-item inventory; LC-42 can remain non-verifying while LC-41 verifies placement. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The municipal address registry lists 18 Larch Court in Zone R4.\",\"The published collection calendar records recycling service for Zone R4 on Tuesday.\",\"The claim log records receipt of the Tuesday recycling report within the 48-hour reporting deadline.\",\"The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42.\",\"The completed item-review record marks both file LC-41 and file LC-42 as not verifying that the blue cart was curbside by 7:00 a.m.\",\"The intake record and supervisor review explicitly document no obstruction, no leaking waste, and no visible pests for this claim.\",\"The records clerk confirmed that the address entry, collection calendar, claim log, evidence inventory, item review, and hazard review all belong to the same Tuesday recycling claim.\"],\"request\":\"Choose the correct routing and response classification.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42."}, {"path": ["evidence", "4"], "text": "The completed item-review record marks both file LC-41 and file LC-42 as not verifying that the blue cart was curbside by 7:00 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42.", "negative_left": "The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42.", "negative_right": "The completed item-review record marks file LC-41 as verifying that the blue cart was curbside by 7:00 a.m. and file LC-42 as not verifying it.", "right": "The completed item-review record marks both file LC-41 and file LC-42 as not verifying that the blue cart was curbside by 7:00 a.m."}, "verifier_independent_model": false}, "family": "scale-diverse-151-004", "id": "scale-diverse-151-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The municipal address registry lists 18 Larch Court in Zone R4.", "The published collection calendar records recycling service for Zone R4 on Tuesday.", "The claim log records receipt of the Tuesday recycling report within the 48-hour reporting deadline.", "The placement-evidence inventory for the Tuesday recycling claim at 18 Larch Court contains exactly two items, file LC-41 and file LC-42.", "The completed item-review record marks file LC-41 as verifying that the blue cart was curbside by 7:00 a.m. and file LC-42 as not verifying it.", "The intake record and supervisor review explicitly document no obstruction, no leaking waste, and no visible pests for this claim.", "The records clerk confirmed that the address entry, collection calendar, claim log, evidence inventory, item review, and hazard review all belong to the same Tuesday recycling claim."], "request": "Choose the correct routing and response classification."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same return-pickup scope, 5 cm preparation threshold, branch exception, service type, address, Tuesday event, and cart identity. The only counterfactual change is the complete factual set-out-time sentence, from 6:47 a.m. to 7:13 a.m.; this creates no duplicate or contradictory measurement within the counterfactual context. The two focus-evidence spans are complete factual sentences rather than policy instructions. Neither context includes an answer code, proposition ID, rule table, gold label, output instruction, or impermissible label rationale; the service-desk and preparation terminology is natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relations or closely defined factual alternatives, and the focus atom is the factual timeliness relation. The base and counter assignments differ only on timeliness and are jointly realizable: an otherwise identical unemptied organics cart can have been timely or late. Policy evidence correctly preserves the substantive return-scope and preparation-exception rule originating in the original state; the remaining governing criteria and priority rules are already retained in the questions object and need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes organics rather than recycling, an unemptied timely and accessible cart, a lid gap over 5 cm, and the explicit absence of every listed urgent-hazard category. The preparation exception therefore applies and uniquely supports routine organics-desk handling with no return pickup.", "rule_index": 0, "sound": true}, {"reason": "Refuting the no-later-than-deadline atom entails a late set-out, placing the cart outside the return-pickup scope. Under the instruction to apply scope before the preparation exception, the over-5-cm gap does not independently establish outcome D. Recycling and all urgent hazards are excluded, so none of A–D applies and E is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported cart at 18 Lark Street was an organics cart for the Tuesday service event."}, {"id": "a2", "statement": "Organics collection was scheduled at 18 Lark Street for the Tuesday service event."}, {"id": "a3", "statement": "The evidence identified recycling as the scheduled or reported material at 18 Lark Street for the Tuesday service event."}, {"id": "a4", "statement": "The reported cart at 18 Lark Street remained unemptied after the Tuesday collection pass."}, {"id": "a5", "statement": "The reported cart's Tuesday curb set-out time was no later than the applicable 7:00 a.m. set-out deadline."}, {"id": "a6", "statement": "The reported cart at 18 Lark Street was accessible to the collection crew during the Tuesday collection pass."}, {"id": "a7", "statement": "The reported cart's lid gap during the Tuesday collection pass was greater than 5 cm."}, {"id": "a8", "statement": "Waste from the reported cart was spilled during the relevant Tuesday service period."}, {"id": "a9", "statement": "The reported cart or its waste blocked travel during the relevant Tuesday service period."}, {"id": "a10", "statement": "The reported cart leaked during the relevant Tuesday service period."}, {"id": "a11", "statement": "Pests were present at the reported cart during the relevant Tuesday service period."}, {"id": "a12", "statement": "The reported cart presented another stated immediate sanitation hazard during the relevant Tuesday service period."}], "base_state_json": "[{\"speaker\":\"Service file\",\"text\":\"For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m.\"},{\"speaker\":\"Curbside record\",\"text\":\"The timestamped curbside record shows that cart serial ORG-4826 was set out at 6:47 a.m. for the Tuesday service event at 18 Lark Street.\"},{\"speaker\":\"Dispatch reconciliation\",\"text\":\"The route manifest and resident report identify ORG-4826 as the organics cart at 18 Lark Street. Organics, not recycling, was scheduled there for Tuesday.\"},{\"speaker\":\"Post-pass inspection\",\"text\":\"ORG-4826 was unobstructed and accessible to the crew during the Tuesday pass, remained unemptied afterward, and had a branch holding its lid 8 cm open.\"},{\"speaker\":\"Incident checklist\",\"text\":\"No waste was spilled, nothing blocked travel, the cart did not leak, no pests were present, and no other immediate sanitation hazard was observed during the relevant Tuesday service period.\"},{\"speaker\":\"Collection crew supervisor\",\"text\":\"Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m."}, {"path": ["1", "text"], "text": "The timestamped curbside record shows that cart serial ORG-4826 was set out at 6:47 a.m. for the Tuesday service event at 18 Lark Street."}], "policy_evidence": [{"path": ["3", "text"], "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}], "rules": [{"justification": "A timely, accessible, apparently missed organics cart is within return-pickup scope, but its lid gap over 5 cm triggers the preparation exception. With every stated urgent-hazard category explicitly absent and recycling explicitly excluded, the organics service desk handles it at routine priority with no return pickup.", "target": "D_organics_desk_routine_no_return", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "An explicitly late cart is outside the timely return-pickup scope, so neither organics return-pickup rubric applies and the scope-first instruction prevents use of the preparation exception as an independent routing basis. Recycling and every urgent-hazard category are explicitly excluded, leaving none of the four substantive rubrics uniquely satisfied.", "target": "E_none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m.", "negative_left": "For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m.", "negative_right": "The timestamped curbside record shows that cart serial ORG-4826 was set out at 7:13 a.m. for the Tuesday service event at 18 Lark Street.", "right": "The timestamped curbside record shows that cart serial ORG-4826 was set out at 6:47 a.m. for the Tuesday service event at 18 Lark Street."}, "verifier_independent_model": false}, "family": "scale-diverse-152-002", "id": "scale-diverse-152-002-base", "input": {"questions": {"decision": {"criteria": {"A_organics_return_routine": "Route to the organics collection crew for a routine return pickup because the cart was timely, accessible, properly prepared within the 5 cm lid limit, and apparently missed.", "B_organics_return_urgent": "Route to the organics collection crew for an urgent return pickup because an eligible missed cart also presents a stated spill, obstruction, leakage, pest, or immediate sanitation hazard.", "C_recycling_desk_routine": "Route to the recycling service desk at routine priority because the evidence identifies recycling, rather than organics, as the scheduled or reported material.", "D_organics_desk_routine_no_return": "Route to the organics service desk at routine priority with no return pickup because a protruding branch or lid gap over 5 cm triggers the preparation exception and no urgent hazard is present.", "E_none_of_above": "Use only if the evidence does not uniquely satisfy any of the four routing, return, and priority rubrics above."}, "instructions": "Choose the routing, return-pickup decision, and priority that match the supplied policy. Apply the return-pickup scope first, then its preparation exception. Urgent priority requires spilled waste, blocked travel, leakage, pests, or another stated immediate sanitation hazard; otherwise use routine priority.", "type": "choice"}}, "state": [{"speaker": "Service file", "text": "For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m."}, {"speaker": "Curbside record", "text": "The timestamped curbside record shows that cart serial ORG-4826 was set out at 6:47 a.m. for the Tuesday service event at 18 Lark Street."}, {"speaker": "Dispatch reconciliation", "text": "The route manifest and resident report identify ORG-4826 as the organics cart at 18 Lark Street. Organics, not recycling, was scheduled there for Tuesday."}, {"speaker": "Post-pass inspection", "text": "ORG-4826 was unobstructed and accessible to the crew during the Tuesday pass, remained unemptied afterward, and had a branch holding its lid 8 cm open."}, {"speaker": "Incident checklist", "text": "No waste was spilled, nothing blocked travel, the cart did not leak, no pests were present, and no other immediate sanitation hazard was observed during the relevant Tuesday service period."}, {"speaker": "Collection crew supervisor", "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-152", "source_is_synthetic": true, "source_sha256": "1dfc7c00860982758ea45c9c5b47d8291aac858825f6143cf430090701b0ce5f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_organics_desk_routine_no_return"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same return-pickup scope, 5 cm preparation threshold, branch exception, service type, address, Tuesday event, and cart identity. The only counterfactual change is the complete factual set-out-time sentence, from 6:47 a.m. to 7:13 a.m.; this creates no duplicate or contradictory measurement within the counterfactual context. The two focus-evidence spans are complete factual sentences rather than policy instructions. Neither context includes an answer code, proposition ID, rule table, gold label, output instruction, or impermissible label rationale; the service-desk and preparation terminology is natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relations or closely defined factual alternatives, and the focus atom is the factual timeliness relation. The base and counter assignments differ only on timeliness and are jointly realizable: an otherwise identical unemptied organics cart can have been timely or late. Policy evidence correctly preserves the substantive return-scope and preparation-exception rule originating in the original state; the remaining governing criteria and priority rules are already retained in the questions object and need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes organics rather than recycling, an unemptied timely and accessible cart, a lid gap over 5 cm, and the explicit absence of every listed urgent-hazard category. The preparation exception therefore applies and uniquely supports routine organics-desk handling with no return pickup.", "rule_index": 0, "sound": true}, {"reason": "Refuting the no-later-than-deadline atom entails a late set-out, placing the cart outside the return-pickup scope. Under the instruction to apply scope before the preparation exception, the over-5-cm gap does not independently establish outcome D. Recycling and all urgent hazards are excluded, so none of A–D applies and E is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported cart at 18 Lark Street was an organics cart for the Tuesday service event."}, {"id": "a2", "statement": "Organics collection was scheduled at 18 Lark Street for the Tuesday service event."}, {"id": "a3", "statement": "The evidence identified recycling as the scheduled or reported material at 18 Lark Street for the Tuesday service event."}, {"id": "a4", "statement": "The reported cart at 18 Lark Street remained unemptied after the Tuesday collection pass."}, {"id": "a5", "statement": "The reported cart's Tuesday curb set-out time was no later than the applicable 7:00 a.m. set-out deadline."}, {"id": "a6", "statement": "The reported cart at 18 Lark Street was accessible to the collection crew during the Tuesday collection pass."}, {"id": "a7", "statement": "The reported cart's lid gap during the Tuesday collection pass was greater than 5 cm."}, {"id": "a8", "statement": "Waste from the reported cart was spilled during the relevant Tuesday service period."}, {"id": "a9", "statement": "The reported cart or its waste blocked travel during the relevant Tuesday service period."}, {"id": "a10", "statement": "The reported cart leaked during the relevant Tuesday service period."}, {"id": "a11", "statement": "Pests were present at the reported cart during the relevant Tuesday service period."}, {"id": "a12", "statement": "The reported cart presented another stated immediate sanitation hazard during the relevant Tuesday service period."}], "base_state_json": "[{\"speaker\":\"Service file\",\"text\":\"For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m.\"},{\"speaker\":\"Curbside record\",\"text\":\"The timestamped curbside record shows that cart serial ORG-4826 was set out at 6:47 a.m. for the Tuesday service event at 18 Lark Street.\"},{\"speaker\":\"Dispatch reconciliation\",\"text\":\"The route manifest and resident report identify ORG-4826 as the organics cart at 18 Lark Street. Organics, not recycling, was scheduled there for Tuesday.\"},{\"speaker\":\"Post-pass inspection\",\"text\":\"ORG-4826 was unobstructed and accessible to the crew during the Tuesday pass, remained unemptied afterward, and had a branch holding its lid 8 cm open.\"},{\"speaker\":\"Incident checklist\",\"text\":\"No waste was spilled, nothing blocked travel, the cart did not leak, no pests were present, and no other immediate sanitation hazard was observed during the relevant Tuesday service period.\"},{\"speaker\":\"Collection crew supervisor\",\"text\":\"Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m."}, {"path": ["1", "text"], "text": "The timestamped curbside record shows that cart serial ORG-4826 was set out at 6:47 a.m. for the Tuesday service event at 18 Lark Street."}], "policy_evidence": [{"path": ["3", "text"], "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}], "rules": [{"justification": "A timely, accessible, apparently missed organics cart is within return-pickup scope, but its lid gap over 5 cm triggers the preparation exception. With every stated urgent-hazard category explicitly absent and recycling explicitly excluded, the organics service desk handles it at routine priority with no return pickup.", "target": "D_organics_desk_routine_no_return", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "An explicitly late cart is outside the timely return-pickup scope, so neither organics return-pickup rubric applies and the scope-first instruction prevents use of the preparation exception as an independent routing basis. Recycling and every urgent-hazard category are explicitly excluded, leaving none of the four substantive rubrics uniquely satisfied.", "target": "E_none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m.", "negative_left": "For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m.", "negative_right": "The timestamped curbside record shows that cart serial ORG-4826 was set out at 7:13 a.m. for the Tuesday service event at 18 Lark Street.", "right": "The timestamped curbside record shows that cart serial ORG-4826 was set out at 6:47 a.m. for the Tuesday service event at 18 Lark Street."}, "verifier_independent_model": false}, "family": "scale-diverse-152-002", "id": "scale-diverse-152-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_organics_return_routine": "Route to the organics collection crew for a routine return pickup because the cart was timely, accessible, properly prepared within the 5 cm lid limit, and apparently missed.", "B_organics_return_urgent": "Route to the organics collection crew for an urgent return pickup because an eligible missed cart also presents a stated spill, obstruction, leakage, pest, or immediate sanitation hazard.", "C_recycling_desk_routine": "Route to the recycling service desk at routine priority because the evidence identifies recycling, rather than organics, as the scheduled or reported material.", "D_organics_desk_routine_no_return": "Route to the organics service desk at routine priority with no return pickup because a protruding branch or lid gap over 5 cm triggers the preparation exception and no urgent hazard is present.", "E_none_of_above": "Use only if the evidence does not uniquely satisfy any of the four routing, return, and priority rubrics above."}, "instructions": "Choose the routing, return-pickup decision, and priority that match the supplied policy. Apply the return-pickup scope first, then its preparation exception. Urgent priority requires spilled waste, blocked travel, leakage, pests, or another stated immediate sanitation hazard; otherwise use routine priority.", "type": "choice"}}, "state": [{"speaker": "Service file", "text": "For the Tuesday service event at 18 Lark Street, the service file identifies cart serial ORG-4826 as the reported cart and lists its applicable curb set-out deadline as 7:00 a.m."}, {"speaker": "Curbside record", "text": "The timestamped curbside record shows that cart serial ORG-4826 was set out at 7:13 a.m. for the Tuesday service event at 18 Lark Street."}, {"speaker": "Dispatch reconciliation", "text": "The route manifest and resident report identify ORG-4826 as the organics cart at 18 Lark Street. Organics, not recycling, was scheduled there for Tuesday."}, {"speaker": "Post-pass inspection", "text": "ORG-4826 was unobstructed and accessible to the crew during the Tuesday pass, remained unemptied afterward, and had a branch holding its lid 8 cm open."}, {"speaker": "Incident checklist", "text": "No waste was spilled, nothing blocked travel, the cart did not leak, no pests were present, and no other immediate sanitation hazard was observed during the relevant Tuesday service period."}, {"speaker": "Collection crew supervisor", "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-152", "source_is_synthetic": true, "source_sha256": "1dfc7c00860982758ea45c9c5b47d8291aac858825f6143cf430090701b0ce5f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "E_none_of_above"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original-state governing preparation policy, while the unchanged questions preserve the full decision criteria and instructions. The same Tuesday service, 18 Lark Street address, organics cart, and routing decision remain in scope. The two focus-evidence spans are complete factual sentences. The counterfactual changes only event Q-47's time from 6:38 a.m. to 7:16 a.m.; no unchanged statement supplies a conflicting set-out time, so the result remains coherent. Neither context contains an explicit answer, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relations or closely defined factual alternatives, and the focus atom is the factual timeliness relation. The base and counter assignments differ only on timeliness and are jointly realizable: an otherwise identical unemptied organics cart can have been timely or late. Policy evidence correctly preserves the substantive return-scope and preparation-exception rule originating in the original state; the remaining governing criteria and priority rules are already retained in the questions object and need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes organics rather than recycling, an unemptied timely and accessible cart, a lid gap over 5 cm, and the explicit absence of every listed urgent-hazard category. The preparation exception therefore applies and uniquely supports routine organics-desk handling with no return pickup.", "rule_index": 0, "sound": true}, {"reason": "Refuting the no-later-than-deadline atom entails a late set-out, placing the cart outside the return-pickup scope. Under the instruction to apply scope before the preparation exception, the over-5-cm gap does not independently establish outcome D. Recycling and all urgent hazards are excluded, so none of A–D applies and E is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported cart at 18 Lark Street was an organics cart for the Tuesday service event."}, {"id": "a2", "statement": "Organics collection was scheduled at 18 Lark Street for the Tuesday service event."}, {"id": "a3", "statement": "The evidence identified recycling as the scheduled or reported material at 18 Lark Street for the Tuesday service event."}, {"id": "a4", "statement": "The reported cart at 18 Lark Street remained unemptied after the Tuesday collection pass."}, {"id": "a5", "statement": "The reported cart's Tuesday curb set-out time was no later than the applicable 7:00 a.m. set-out deadline."}, {"id": "a6", "statement": "The reported cart at 18 Lark Street was accessible to the collection crew during the Tuesday collection pass."}, {"id": "a7", "statement": "The reported cart's lid gap during the Tuesday collection pass was greater than 5 cm."}, {"id": "a8", "statement": "Waste from the reported cart was spilled during the relevant Tuesday service period."}, {"id": "a9", "statement": "The reported cart or its waste blocked travel during the relevant Tuesday service period."}, {"id": "a10", "statement": "The reported cart leaked during the relevant Tuesday service period."}, {"id": "a11", "statement": "Pests were present at the reported cart during the relevant Tuesday service period."}, {"id": "a12", "statement": "The reported cart presented another stated immediate sanitation hazard during the relevant Tuesday service period."}], "base_state_json": "[{\"speaker\":\"Municipal service record\",\"text\":\"The Tuesday report for 18 Lark Street concerned a green organics cart, and the route schedule confirms organics service at that address. The material was not recycling.\"},{\"speaker\":\"Set-out record\",\"text\":\"For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47.\"},{\"speaker\":\"Event log\",\"text\":\"Field event Q-47 occurred at 6:38 a.m. on Tuesday.\"},{\"speaker\":\"Post-pass inspection\",\"text\":\"The reported cart remained unemptied after the Tuesday collection pass. Crew access records show that it was unobstructed and reachable during that pass.\"},{\"speaker\":\"Photographic inspection\",\"text\":\"A branch held the reported cart's lid 8 cm open during the Tuesday collection pass.\"},{\"speaker\":\"Safety inspection\",\"text\":\"No waste was spilled, and neither the cart nor its waste blocked travel. The cart did not leak, pests were absent, and no other immediate sanitation hazard was observed during the relevant Tuesday service period.\"},{\"speaker\":\"Collection crew supervisor\",\"text\":\"Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47."}, {"path": ["2", "text"], "text": "Field event Q-47 occurred at 6:38 a.m. on Tuesday."}], "policy_evidence": [{"path": ["3", "text"], "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}], "rules": [{"justification": "A timely, accessible, apparently missed organics cart is within return-pickup scope, but its lid gap over 5 cm triggers the preparation exception. With every stated urgent-hazard category explicitly absent and recycling explicitly excluded, the organics service desk handles it at routine priority with no return pickup.", "target": "D_organics_desk_routine_no_return", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "An explicitly late cart is outside the timely return-pickup scope, so neither organics return-pickup rubric applies and the scope-first instruction prevents use of the preparation exception as an independent routing basis. Recycling and every urgent-hazard category are explicitly excluded, leaving none of the four substantive rubrics uniquely satisfied.", "target": "E_none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47.", "negative_left": "For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47.", "negative_right": "Field event Q-47 occurred at 7:16 a.m. on Tuesday.", "right": "Field event Q-47 occurred at 6:38 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-152-004", "id": "scale-diverse-152-004-base", "input": {"questions": {"decision": {"criteria": {"A_organics_return_routine": "Route to the organics collection crew for a routine return pickup because the cart was timely, accessible, properly prepared within the 5 cm lid limit, and apparently missed.", "B_organics_return_urgent": "Route to the organics collection crew for an urgent return pickup because an eligible missed cart also presents a stated spill, obstruction, leakage, pest, or immediate sanitation hazard.", "C_recycling_desk_routine": "Route to the recycling service desk at routine priority because the evidence identifies recycling, rather than organics, as the scheduled or reported material.", "D_organics_desk_routine_no_return": "Route to the organics service desk at routine priority with no return pickup because a protruding branch or lid gap over 5 cm triggers the preparation exception and no urgent hazard is present.", "E_none_of_above": "Use only if the evidence does not uniquely satisfy any of the four routing, return, and priority rubrics above."}, "instructions": "Choose the routing, return-pickup decision, and priority that match the supplied policy. Apply the return-pickup scope first, then its preparation exception. Urgent priority requires spilled waste, blocked travel, leakage, pests, or another stated immediate sanitation hazard; otherwise use routine priority.", "type": "choice"}}, "state": [{"speaker": "Municipal service record", "text": "The Tuesday report for 18 Lark Street concerned a green organics cart, and the route schedule confirms organics service at that address. The material was not recycling."}, {"speaker": "Set-out record", "text": "For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47."}, {"speaker": "Event log", "text": "Field event Q-47 occurred at 6:38 a.m. on Tuesday."}, {"speaker": "Post-pass inspection", "text": "The reported cart remained unemptied after the Tuesday collection pass. Crew access records show that it was unobstructed and reachable during that pass."}, {"speaker": "Photographic inspection", "text": "A branch held the reported cart's lid 8 cm open during the Tuesday collection pass."}, {"speaker": "Safety inspection", "text": "No waste was spilled, and neither the cart nor its waste blocked travel. The cart did not leak, pests were absent, and no other immediate sanitation hazard was observed during the relevant Tuesday service period."}, {"speaker": "Collection crew supervisor", "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-152", "source_is_synthetic": true, "source_sha256": "1dfc7c00860982758ea45c9c5b47d8291aac858825f6143cf430090701b0ce5f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_organics_desk_routine_no_return"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original-state governing preparation policy, while the unchanged questions preserve the full decision criteria and instructions. The same Tuesday service, 18 Lark Street address, organics cart, and routing decision remain in scope. The two focus-evidence spans are complete factual sentences. The counterfactual changes only event Q-47's time from 6:38 a.m. to 7:16 a.m.; no unchanged statement supplies a conflicting set-out time, so the result remains coherent. Neither context contains an explicit answer, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relations or closely defined factual alternatives, and the focus atom is the factual timeliness relation. The base and counter assignments differ only on timeliness and are jointly realizable: an otherwise identical unemptied organics cart can have been timely or late. Policy evidence correctly preserves the substantive return-scope and preparation-exception rule originating in the original state; the remaining governing criteria and priority rules are already retained in the questions object and need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes organics rather than recycling, an unemptied timely and accessible cart, a lid gap over 5 cm, and the explicit absence of every listed urgent-hazard category. The preparation exception therefore applies and uniquely supports routine organics-desk handling with no return pickup.", "rule_index": 0, "sound": true}, {"reason": "Refuting the no-later-than-deadline atom entails a late set-out, placing the cart outside the return-pickup scope. Under the instruction to apply scope before the preparation exception, the over-5-cm gap does not independently establish outcome D. Recycling and all urgent hazards are excluded, so none of A–D applies and E is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported cart at 18 Lark Street was an organics cart for the Tuesday service event."}, {"id": "a2", "statement": "Organics collection was scheduled at 18 Lark Street for the Tuesday service event."}, {"id": "a3", "statement": "The evidence identified recycling as the scheduled or reported material at 18 Lark Street for the Tuesday service event."}, {"id": "a4", "statement": "The reported cart at 18 Lark Street remained unemptied after the Tuesday collection pass."}, {"id": "a5", "statement": "The reported cart's Tuesday curb set-out time was no later than the applicable 7:00 a.m. set-out deadline."}, {"id": "a6", "statement": "The reported cart at 18 Lark Street was accessible to the collection crew during the Tuesday collection pass."}, {"id": "a7", "statement": "The reported cart's lid gap during the Tuesday collection pass was greater than 5 cm."}, {"id": "a8", "statement": "Waste from the reported cart was spilled during the relevant Tuesday service period."}, {"id": "a9", "statement": "The reported cart or its waste blocked travel during the relevant Tuesday service period."}, {"id": "a10", "statement": "The reported cart leaked during the relevant Tuesday service period."}, {"id": "a11", "statement": "Pests were present at the reported cart during the relevant Tuesday service period."}, {"id": "a12", "statement": "The reported cart presented another stated immediate sanitation hazard during the relevant Tuesday service period."}], "base_state_json": "[{\"speaker\":\"Municipal service record\",\"text\":\"The Tuesday report for 18 Lark Street concerned a green organics cart, and the route schedule confirms organics service at that address. The material was not recycling.\"},{\"speaker\":\"Set-out record\",\"text\":\"For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47.\"},{\"speaker\":\"Event log\",\"text\":\"Field event Q-47 occurred at 6:38 a.m. on Tuesday.\"},{\"speaker\":\"Post-pass inspection\",\"text\":\"The reported cart remained unemptied after the Tuesday collection pass. Crew access records show that it was unobstructed and reachable during that pass.\"},{\"speaker\":\"Photographic inspection\",\"text\":\"A branch held the reported cart's lid 8 cm open during the Tuesday collection pass.\"},{\"speaker\":\"Safety inspection\",\"text\":\"No waste was spilled, and neither the cart nor its waste blocked travel. The cart did not leak, pests were absent, and no other immediate sanitation hazard was observed during the relevant Tuesday service period.\"},{\"speaker\":\"Collection crew supervisor\",\"text\":\"Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47."}, {"path": ["2", "text"], "text": "Field event Q-47 occurred at 6:38 a.m. on Tuesday."}], "policy_evidence": [{"path": ["3", "text"], "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}], "rules": [{"justification": "A timely, accessible, apparently missed organics cart is within return-pickup scope, but its lid gap over 5 cm triggers the preparation exception. With every stated urgent-hazard category explicitly absent and recycling explicitly excluded, the organics service desk handles it at routine priority with no return pickup.", "target": "D_organics_desk_routine_no_return", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "An explicitly late cart is outside the timely return-pickup scope, so neither organics return-pickup rubric applies and the scope-first instruction prevents use of the preparation exception as an independent routing basis. Recycling and every urgent-hazard category are explicitly excluded, leaving none of the four substantive rubrics uniquely satisfied.", "target": "E_none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47.", "negative_left": "For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47.", "negative_right": "Field event Q-47 occurred at 7:16 a.m. on Tuesday.", "right": "Field event Q-47 occurred at 6:38 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-152-004", "id": "scale-diverse-152-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_organics_return_routine": "Route to the organics collection crew for a routine return pickup because the cart was timely, accessible, properly prepared within the 5 cm lid limit, and apparently missed.", "B_organics_return_urgent": "Route to the organics collection crew for an urgent return pickup because an eligible missed cart also presents a stated spill, obstruction, leakage, pest, or immediate sanitation hazard.", "C_recycling_desk_routine": "Route to the recycling service desk at routine priority because the evidence identifies recycling, rather than organics, as the scheduled or reported material.", "D_organics_desk_routine_no_return": "Route to the organics service desk at routine priority with no return pickup because a protruding branch or lid gap over 5 cm triggers the preparation exception and no urgent hazard is present.", "E_none_of_above": "Use only if the evidence does not uniquely satisfy any of the four routing, return, and priority rubrics above."}, "instructions": "Choose the routing, return-pickup decision, and priority that match the supplied policy. Apply the return-pickup scope first, then its preparation exception. Urgent priority requires spilled waste, blocked travel, leakage, pests, or another stated immediate sanitation hazard; otherwise use routine priority.", "type": "choice"}}, "state": [{"speaker": "Municipal service record", "text": "The Tuesday report for 18 Lark Street concerned a green organics cart, and the route schedule confirms organics service at that address. The material was not recycling."}, {"speaker": "Set-out record", "text": "For Tuesday service at 18 Lark Street, the applicable curb set-out deadline was 7:00 a.m., and the reported cart's set-out was logged as field event Q-47."}, {"speaker": "Event log", "text": "Field event Q-47 occurred at 7:16 a.m. on Tuesday."}, {"speaker": "Post-pass inspection", "text": "The reported cart remained unemptied after the Tuesday collection pass. Crew access records show that it was unobstructed and reachable during that pass."}, {"speaker": "Photographic inspection", "text": "A branch held the reported cart's lid 8 cm open during the Tuesday collection pass."}, {"speaker": "Safety inspection", "text": "No waste was spilled, and neither the cart nor its waste blocked travel. The cart did not leak, pests were absent, and no other immediate sanitation hazard was observed during the relevant Tuesday service period."}, {"speaker": "Collection crew supervisor", "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-152", "source_is_synthetic": true, "source_sha256": "1dfc7c00860982758ea45c9c5b47d8291aac858825f6143cf430090701b0ce5f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "E_none_of_above"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing and urgency policy and keep the same report, resident, address, Tuesday timeframe, routing path, and urgency decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only the reported cart’s manifest material from organics to recycling; together with the exclusive Cart A/Cart B assignments, inspection findings, and recorded priorities, this remains internally coherent without duplicate contradictory assertions. The contexts contain ordinary operational evidence and policy terminology, but no gold label, answer code, proposition identifier, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual identity of the reported cart. The policy evidence preserves the state-originating routing and urgency rules; question-originating criteria need not be repeated. The base and counter assignments can describe the same two carts and circumstances while changing only which cart is the reported cart, without violating the policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the reported cart as Cart A, establish a timely and properly placed missed collection, give organics as its scheduled material, and explicitly set the controlling priority to urgent. The leaking-waste condition also satisfies the policy's necessary urgent-condition restriction.", "rule_index": 0, "sound": true}, {"reason": "Membership in {Cart A, Cart B} together with refutation of Cart A identifies the reported cart as Cart B. The remaining conditions establish a timely and properly placed missed collection, scheduled recycling, routine priority, and absence of both policy-listed urgent conditions, which is sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane belongs to the candidate asset set {Cart A, Cart B}."}, {"id": "a2", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is Cart A."}, {"id": "a3", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was properly placed for collection."}, {"id": "a4", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was curbside no later than the 6:30 a.m. Tuesday cutoff."}, {"id": "a5", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane remained uncollected after the scheduled Tuesday collection."}, {"id": "a6", "statement": "The scheduled Tuesday material for Cart A at 18 Birch Lane is organics."}, {"id": "a7", "statement": "The priority field in the controlling dispatch record for Cart A’s Tuesday missed collection at 18 Birch Lane is urgent."}, {"id": "a8", "statement": "Waste in Cart A was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a9", "statement": "The scheduled Tuesday material for Cart B at 18 Birch Lane is recycling."}, {"id": "a10", "statement": "The priority field in the controlling dispatch record for Cart B’s Tuesday missed collection at 18 Birch Lane is routine."}, {"id": "a11", "statement": "Waste in Cart B was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a12", "statement": "Cart B blocked the roadway at 18 Birch Lane during the Tuesday collection period."}], "base_state_json": "\"At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B. At 6:12 a.m., a timestamped image showed the reported cart correctly positioned curbside for collection, ahead of the 6:30 a.m. Tuesday cutoff. For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned organics. After the scheduled collection, a 7:20 a.m. inspection confirmed that the reported cart remained uncollected. The collection-period inspection log recorded leaking waste from Cart A, no leaking waste from Cart B, and no roadway blockage by Cart B. The controlling Tuesday missed-collection dispatch records listed Cart A’s priority as urgent and Cart B’s as routine. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B."}, {"path": [], "text": "For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned organics."}], "policy_evidence": [{"path": [], "text": "Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}], "rules": [{"justification": "The report concerns Cart A, which was properly placed, timely, and missed; its scheduled material is organics and its controlling priority is urgent. Its leaking waste also satisfies the policy’s necessary condition for urgent status.", "target": "true", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The reported cart is one of Cart A and Cart B but is not Cart A, so it is Cart B. It was properly placed, timely, and missed; its scheduled material is recycling and its controlling priority is routine, while both policy-listed urgent conditions are explicitly absent.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B.", "negative_left": "At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B.", "negative_right": "For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned recycling.", "right": "For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned organics."}, "verifier_independent_model": false}, "family": "scale-diverse-153-001", "id": "scale-diverse-153-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route it to the organics crew as urgent; route it to the recycling return-pickup crew with routine priority.", "true": "Route the report to the organics return-pickup crew with urgent priority."}, "instructions": "Decide whether this report should be routed to the organics return-pickup crew as an urgent case.", "type": "noul"}}, "state": "At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B. At 6:12 a.m., a timestamped image showed the reported cart correctly positioned curbside for collection, ahead of the 6:30 a.m. Tuesday cutoff. For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned organics. After the scheduled collection, a 7:20 a.m. inspection confirmed that the reported cart remained uncollected. The collection-period inspection log recorded leaking waste from Cart A, no leaking waste from Cart B, and no roadway blockage by Cart B. The controlling Tuesday missed-collection dispatch records listed Cart A’s priority as urgent and Cart B’s as routine. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}, "method": "c2d", "provenance": {"source_id": "diverse-153", "source_is_synthetic": true, "source_sha256": "f8cc49d0970071d68e752b9fc8ad734adab047d8cd33245838da80f07b42cdfc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing and urgency policy and keep the same report, resident, address, Tuesday timeframe, routing path, and urgency decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only the reported cart’s manifest material from organics to recycling; together with the exclusive Cart A/Cart B assignments, inspection findings, and recorded priorities, this remains internally coherent without duplicate contradictory assertions. The contexts contain ordinary operational evidence and policy terminology, but no gold label, answer code, proposition identifier, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual identity of the reported cart. The policy evidence preserves the state-originating routing and urgency rules; question-originating criteria need not be repeated. The base and counter assignments can describe the same two carts and circumstances while changing only which cart is the reported cart, without violating the policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the reported cart as Cart A, establish a timely and properly placed missed collection, give organics as its scheduled material, and explicitly set the controlling priority to urgent. The leaking-waste condition also satisfies the policy's necessary urgent-condition restriction.", "rule_index": 0, "sound": true}, {"reason": "Membership in {Cart A, Cart B} together with refutation of Cart A identifies the reported cart as Cart B. The remaining conditions establish a timely and properly placed missed collection, scheduled recycling, routine priority, and absence of both policy-listed urgent conditions, which is sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane belongs to the candidate asset set {Cart A, Cart B}."}, {"id": "a2", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is Cart A."}, {"id": "a3", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was properly placed for collection."}, {"id": "a4", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was curbside no later than the 6:30 a.m. Tuesday cutoff."}, {"id": "a5", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane remained uncollected after the scheduled Tuesday collection."}, {"id": "a6", "statement": "The scheduled Tuesday material for Cart A at 18 Birch Lane is organics."}, {"id": "a7", "statement": "The priority field in the controlling dispatch record for Cart A’s Tuesday missed collection at 18 Birch Lane is urgent."}, {"id": "a8", "statement": "Waste in Cart A was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a9", "statement": "The scheduled Tuesday material for Cart B at 18 Birch Lane is recycling."}, {"id": "a10", "statement": "The priority field in the controlling dispatch record for Cart B’s Tuesday missed collection at 18 Birch Lane is routine."}, {"id": "a11", "statement": "Waste in Cart B was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a12", "statement": "Cart B blocked the roadway at 18 Birch Lane during the Tuesday collection period."}], "base_state_json": "\"At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B. At 6:12 a.m., a timestamped image showed the reported cart correctly positioned curbside for collection, ahead of the 6:30 a.m. Tuesday cutoff. For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned organics. After the scheduled collection, a 7:20 a.m. inspection confirmed that the reported cart remained uncollected. The collection-period inspection log recorded leaking waste from Cart A, no leaking waste from Cart B, and no roadway blockage by Cart B. The controlling Tuesday missed-collection dispatch records listed Cart A’s priority as urgent and Cart B’s as routine. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B."}, {"path": [], "text": "For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned organics."}], "policy_evidence": [{"path": [], "text": "Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}], "rules": [{"justification": "The report concerns Cart A, which was properly placed, timely, and missed; its scheduled material is organics and its controlling priority is urgent. Its leaking waste also satisfies the policy’s necessary condition for urgent status.", "target": "true", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The reported cart is one of Cart A and Cart B but is not Cart A, so it is Cart B. It was properly placed, timely, and missed; its scheduled material is recycling and its controlling priority is routine, while both policy-listed urgent conditions are explicitly absent.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B.", "negative_left": "At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B.", "negative_right": "For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned recycling.", "right": "For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned organics."}, "verifier_independent_model": false}, "family": "scale-diverse-153-001", "id": "scale-diverse-153-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route it to the organics crew as urgent; route it to the recycling return-pickup crew with routine priority.", "true": "Route the report to the organics return-pickup crew with urgent priority."}, "instructions": "Decide whether this report should be routed to the organics return-pickup crew as an urgent case.", "type": "noul"}}, "state": "At 5:55 a.m. on the Tuesday covered by Mara Chen’s missed-collection report at 18 Birch Lane, the reported cart was exactly one of two assets: Cart A or Cart B. At 6:12 a.m., a timestamped image showed the reported cart correctly positioned curbside for collection, ahead of the 6:30 a.m. Tuesday cutoff. For that Tuesday at 18 Birch Lane, the collection manifest assigned only organics to Cart A and only recycling to Cart B, while the reported cart’s manifest entry assigned recycling. After the scheduled collection, a 7:20 a.m. inspection confirmed that the reported cart remained uncollected. The collection-period inspection log recorded leaking waste from Cart A, no leaking waste from Cart B, and no roadway blockage by Cart B. The controlling Tuesday missed-collection dispatch records listed Cart A’s priority as urgent and Cart B’s as routine. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}, "method": "c2d", "provenance": {"source_id": "diverse-153", "source_is_synthetic": true, "source_sha256": "f8cc49d0970071d68e752b9fc8ad734adab047d8cd33245838da80f07b42cdfc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing and urgency policy verbatim and retain the report’s resident, address, Tuesday timing, container, and routing-decision scope. The two evidence spans are complete factual sentences. The counterfactual coherently changes the asset-tag mapping so the reported MC-417 cart is Cart B; its scheduled recycling assignment, routine dispatch status, and absence of leakage or roadway blockage are mutually consistent. Although the contexts contain facts that support a decision, neither embeds a gold label, answer code, rule table, proposition identifier, classifier instruction, or explicit final routing answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual identity of the reported cart. The policy evidence preserves the state-originating routing and urgency rules; question-originating criteria need not be repeated. The base and counter assignments can describe the same two carts and circumstances while changing only which cart is the reported cart, without violating the policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the reported cart as Cart A, establish a timely and properly placed missed collection, give organics as its scheduled material, and explicitly set the controlling priority to urgent. The leaking-waste condition also satisfies the policy's necessary urgent-condition restriction.", "rule_index": 0, "sound": true}, {"reason": "Membership in {Cart A, Cart B} together with refutation of Cart A identifies the reported cart as Cart B. The remaining conditions establish a timely and properly placed missed collection, scheduled recycling, routine priority, and absence of both policy-listed urgent conditions, which is sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane belongs to the candidate asset set {Cart A, Cart B}."}, {"id": "a2", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is Cart A."}, {"id": "a3", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was properly placed for collection."}, {"id": "a4", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was curbside no later than the 6:30 a.m. Tuesday cutoff."}, {"id": "a5", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane remained uncollected after the scheduled Tuesday collection."}, {"id": "a6", "statement": "The scheduled Tuesday material for Cart A at 18 Birch Lane is organics."}, {"id": "a7", "statement": "The priority field in the controlling dispatch record for Cart A’s Tuesday missed collection at 18 Birch Lane is urgent."}, {"id": "a8", "statement": "Waste in Cart A was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a9", "statement": "The scheduled Tuesday material for Cart B at 18 Birch Lane is recycling."}, {"id": "a10", "statement": "The priority field in the controlling dispatch record for Cart B’s Tuesday missed collection at 18 Birch Lane is routine."}, {"id": "a11", "statement": "Waste in Cart B was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a12", "statement": "Cart B blocked the roadway at 18 Birch Lane during the Tuesday collection period."}], "base_state_json": "\"A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417. The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-417 and Cart B with asset tag MC-862. A placement review found the reported cart properly positioned curbside before the 6:30 a.m. Tuesday cutoff and still uncollected after the scheduled collection. The address schedule assigns Tuesday organics to Cart A and Tuesday recycling to Cart B. The controlling dispatch records mark Cart A’s Tuesday missed collection urgent and Cart B’s routine. During the collection period, waste in Cart A was leaking; waste in Cart B was not leaking, and Cart B did not block the roadway. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417."}, {"path": [], "text": "The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-417 and Cart B with asset tag MC-862."}], "policy_evidence": [{"path": [], "text": "Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}], "rules": [{"justification": "The report concerns Cart A, which was properly placed, timely, and missed; its scheduled material is organics and its controlling priority is urgent. Its leaking waste also satisfies the policy’s necessary condition for urgent status.", "target": "true", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The reported cart is one of Cart A and Cart B but is not Cart A, so it is Cart B. It was properly placed, timely, and missed; its scheduled material is recycling and its controlling priority is routine, while both policy-listed urgent conditions are explicitly absent.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417.", "negative_left": "A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417.", "negative_right": "The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-862 and Cart B with asset tag MC-417.", "right": "The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-417 and Cart B with asset tag MC-862."}, "verifier_independent_model": false}, "family": "scale-diverse-153-002", "id": "scale-diverse-153-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route it to the organics crew as urgent; route it to the recycling return-pickup crew with routine priority.", "true": "Route the report to the organics return-pickup crew with urgent priority."}, "instructions": "Decide whether this report should be routed to the organics return-pickup crew as an urgent case.", "type": "noul"}}, "state": "A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417. The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-417 and Cart B with asset tag MC-862. A placement review found the reported cart properly positioned curbside before the 6:30 a.m. Tuesday cutoff and still uncollected after the scheduled collection. The address schedule assigns Tuesday organics to Cart A and Tuesday recycling to Cart B. The controlling dispatch records mark Cart A’s Tuesday missed collection urgent and Cart B’s routine. During the collection period, waste in Cart A was leaking; waste in Cart B was not leaking, and Cart B did not block the roadway. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}, "method": "c2d", "provenance": {"source_id": "diverse-153", "source_is_synthetic": true, "source_sha256": "f8cc49d0970071d68e752b9fc8ad734adab047d8cd33245838da80f07b42cdfc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing and urgency policy verbatim and retain the report’s resident, address, Tuesday timing, container, and routing-decision scope. The two evidence spans are complete factual sentences. The counterfactual coherently changes the asset-tag mapping so the reported MC-417 cart is Cart B; its scheduled recycling assignment, routine dispatch status, and absence of leakage or roadway blockage are mutually consistent. Although the contexts contain facts that support a decision, neither embeds a gold label, answer code, rule table, proposition identifier, classifier instruction, or explicit final routing answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual identity of the reported cart. The policy evidence preserves the state-originating routing and urgency rules; question-originating criteria need not be repeated. The base and counter assignments can describe the same two carts and circumstances while changing only which cart is the reported cart, without violating the policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the reported cart as Cart A, establish a timely and properly placed missed collection, give organics as its scheduled material, and explicitly set the controlling priority to urgent. The leaking-waste condition also satisfies the policy's necessary urgent-condition restriction.", "rule_index": 0, "sound": true}, {"reason": "Membership in {Cart A, Cart B} together with refutation of Cart A identifies the reported cart as Cart B. The remaining conditions establish a timely and properly placed missed collection, scheduled recycling, routine priority, and absence of both policy-listed urgent conditions, which is sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane belongs to the candidate asset set {Cart A, Cart B}."}, {"id": "a2", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is Cart A."}, {"id": "a3", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was properly placed for collection."}, {"id": "a4", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was curbside no later than the 6:30 a.m. Tuesday cutoff."}, {"id": "a5", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane remained uncollected after the scheduled Tuesday collection."}, {"id": "a6", "statement": "The scheduled Tuesday material for Cart A at 18 Birch Lane is organics."}, {"id": "a7", "statement": "The priority field in the controlling dispatch record for Cart A’s Tuesday missed collection at 18 Birch Lane is urgent."}, {"id": "a8", "statement": "Waste in Cart A was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a9", "statement": "The scheduled Tuesday material for Cart B at 18 Birch Lane is recycling."}, {"id": "a10", "statement": "The priority field in the controlling dispatch record for Cart B’s Tuesday missed collection at 18 Birch Lane is routine."}, {"id": "a11", "statement": "Waste in Cart B was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a12", "statement": "Cart B blocked the roadway at 18 Birch Lane during the Tuesday collection period."}], "base_state_json": "\"A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417. The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-417 and Cart B with asset tag MC-862. A placement review found the reported cart properly positioned curbside before the 6:30 a.m. Tuesday cutoff and still uncollected after the scheduled collection. The address schedule assigns Tuesday organics to Cart A and Tuesday recycling to Cart B. The controlling dispatch records mark Cart A’s Tuesday missed collection urgent and Cart B’s routine. During the collection period, waste in Cart A was leaking; waste in Cart B was not leaking, and Cart B did not block the roadway. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417."}, {"path": [], "text": "The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-417 and Cart B with asset tag MC-862."}], "policy_evidence": [{"path": [], "text": "Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}], "rules": [{"justification": "The report concerns Cart A, which was properly placed, timely, and missed; its scheduled material is organics and its controlling priority is urgent. Its leaking waste also satisfies the policy’s necessary condition for urgent status.", "target": "true", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The reported cart is one of Cart A and Cart B but is not Cart A, so it is Cart B. It was properly placed, timely, and missed; its scheduled material is recycling and its controlling priority is routine, while both policy-listed urgent conditions are explicitly absent.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417.", "negative_left": "A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417.", "negative_right": "The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-862 and Cart B with asset tag MC-417.", "right": "The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-417 and Cart B with asset tag MC-862."}, "verifier_independent_model": false}, "family": "scale-diverse-153-002", "id": "scale-diverse-153-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route it to the organics crew as urgent; route it to the recycling return-pickup crew with routine priority.", "true": "Route the report to the organics return-pickup crew with urgent priority."}, "instructions": "Decide whether this report should be routed to the organics return-pickup crew as an urgent case.", "type": "noul"}}, "state": "A reconciliation completed at 9:14 a.m. on Wednesday established that the cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was exactly one of Cart A and Cart B and bore asset tag MC-417. The asset register exported at 9:12 a.m. on Wednesday lists Cart A with asset tag MC-862 and Cart B with asset tag MC-417. A placement review found the reported cart properly positioned curbside before the 6:30 a.m. Tuesday cutoff and still uncollected after the scheduled collection. The address schedule assigns Tuesday organics to Cart A and Tuesday recycling to Cart B. The controlling dispatch records mark Cart A’s Tuesday missed collection urgent and Cart B’s routine. During the collection period, waste in Cart A was leaking; waste in Cart B was not leaking, and Cart B did not block the roadway. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}, "method": "c2d", "provenance": {"source_id": "diverse-153", "source_is_synthetic": true, "source_sha256": "f8cc49d0970071d68e752b9fc8ad734adab047d8cd33245838da80f07b42cdfc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing routing and urgency policy, while the unchanged questions object preserves the decision criteria. The report remains bound to Mara Chen, 18 Birch Lane, the Tuesday missed collection, and the same routing decision. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the registry association of serial QH-7316 from Cart A to Cart B; the unchanged schedule, condition, and dispatch facts remain internally consistent with that change. Neither context includes an answer code, proposition ID, rule table, output instruction, or explicit final routing answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual identity of the reported cart. The policy evidence preserves the state-originating routing and urgency rules; question-originating criteria need not be repeated. The base and counter assignments can describe the same two carts and circumstances while changing only which cart is the reported cart, without violating the policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the reported cart as Cart A, establish a timely and properly placed missed collection, give organics as its scheduled material, and explicitly set the controlling priority to urgent. The leaking-waste condition also satisfies the policy's necessary urgent-condition restriction.", "rule_index": 0, "sound": true}, {"reason": "Membership in {Cart A, Cart B} together with refutation of Cart A identifies the reported cart as Cart B. The remaining conditions establish a timely and properly placed missed collection, scheduled recycling, routine priority, and absence of both policy-listed urgent conditions, which is sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane belongs to the candidate asset set {Cart A, Cart B}."}, {"id": "a2", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is Cart A."}, {"id": "a3", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was properly placed for collection."}, {"id": "a4", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was curbside no later than the 6:30 a.m. Tuesday cutoff."}, {"id": "a5", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane remained uncollected after the scheduled Tuesday collection."}, {"id": "a6", "statement": "The scheduled Tuesday material for Cart A at 18 Birch Lane is organics."}, {"id": "a7", "statement": "The priority field in the controlling dispatch record for Cart A’s Tuesday missed collection at 18 Birch Lane is urgent."}, {"id": "a8", "statement": "Waste in Cart A was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a9", "statement": "The scheduled Tuesday material for Cart B at 18 Birch Lane is recycling."}, {"id": "a10", "statement": "The priority field in the controlling dispatch record for Cart B’s Tuesday missed collection at 18 Birch Lane is routine."}, {"id": "a11", "statement": "Waste in Cart B was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a12", "statement": "Cart B blocked the roadway at 18 Birch Lane during the Tuesday collection period."}], "base_state_json": "\"Operational handoff: At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316. The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart A. Intake reconciliation limits the reported asset to the candidate set consisting of Cart A and Cart B. Site records confirm that the reported cart was properly placed curbside by the 6:30 a.m. Tuesday cutoff and remained uncollected after the scheduled collection. The schedule assigns Tuesday organics to Cart A and Tuesday recycling to Cart B. Controlling dispatch records mark Cart A’s Tuesday missed collection urgent and Cart B’s routine. During the collection period, waste in Cart A was leaking; Cart B had no leaking waste and did not block the roadway. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316."}, {"path": [], "text": "The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart A."}], "policy_evidence": [{"path": [], "text": "Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}], "rules": [{"justification": "The report concerns Cart A, which was properly placed, timely, and missed; its scheduled material is organics and its controlling priority is urgent. Its leaking waste also satisfies the policy’s necessary condition for urgent status.", "target": "true", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The reported cart is one of Cart A and Cart B but is not Cart A, so it is Cart B. It was properly placed, timely, and missed; its scheduled material is recycling and its controlling priority is routine, while both policy-listed urgent conditions are explicitly absent.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316.", "negative_left": "At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316.", "negative_right": "The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart B.", "right": "The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart A."}, "verifier_independent_model": false}, "family": "scale-diverse-153-003", "id": "scale-diverse-153-003-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route it to the organics crew as urgent; route it to the recycling return-pickup crew with routine priority.", "true": "Route the report to the organics return-pickup crew with urgent priority."}, "instructions": "Decide whether this report should be routed to the organics return-pickup crew as an urgent case.", "type": "noul"}}, "state": "Operational handoff: At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316. The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart A. Intake reconciliation limits the reported asset to the candidate set consisting of Cart A and Cart B. Site records confirm that the reported cart was properly placed curbside by the 6:30 a.m. Tuesday cutoff and remained uncollected after the scheduled collection. The schedule assigns Tuesday organics to Cart A and Tuesday recycling to Cart B. Controlling dispatch records mark Cart A’s Tuesday missed collection urgent and Cart B’s routine. During the collection period, waste in Cart A was leaking; Cart B had no leaking waste and did not block the roadway. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}, "method": "c2d", "provenance": {"source_id": "diverse-153", "source_is_synthetic": true, "source_sha256": "f8cc49d0970071d68e752b9fc8ad734adab047d8cd33245838da80f07b42cdfc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing routing and urgency policy, while the unchanged questions object preserves the decision criteria. The report remains bound to Mara Chen, 18 Birch Lane, the Tuesday missed collection, and the same routing decision. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the registry association of serial QH-7316 from Cart A to Cart B; the unchanged schedule, condition, and dispatch facts remain internally consistent with that change. Neither context includes an answer code, proposition ID, rule table, output instruction, or explicit final routing answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual identity of the reported cart. The policy evidence preserves the state-originating routing and urgency rules; question-originating criteria need not be repeated. The base and counter assignments can describe the same two carts and circumstances while changing only which cart is the reported cart, without violating the policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the reported cart as Cart A, establish a timely and properly placed missed collection, give organics as its scheduled material, and explicitly set the controlling priority to urgent. The leaking-waste condition also satisfies the policy's necessary urgent-condition restriction.", "rule_index": 0, "sound": true}, {"reason": "Membership in {Cart A, Cart B} together with refutation of Cart A identifies the reported cart as Cart B. The remaining conditions establish a timely and properly placed missed collection, scheduled recycling, routine priority, and absence of both policy-listed urgent conditions, which is sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane belongs to the candidate asset set {Cart A, Cart B}."}, {"id": "a2", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is Cart A."}, {"id": "a3", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was properly placed for collection."}, {"id": "a4", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was curbside no later than the 6:30 a.m. Tuesday cutoff."}, {"id": "a5", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane remained uncollected after the scheduled Tuesday collection."}, {"id": "a6", "statement": "The scheduled Tuesday material for Cart A at 18 Birch Lane is organics."}, {"id": "a7", "statement": "The priority field in the controlling dispatch record for Cart A’s Tuesday missed collection at 18 Birch Lane is urgent."}, {"id": "a8", "statement": "Waste in Cart A was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a9", "statement": "The scheduled Tuesday material for Cart B at 18 Birch Lane is recycling."}, {"id": "a10", "statement": "The priority field in the controlling dispatch record for Cart B’s Tuesday missed collection at 18 Birch Lane is routine."}, {"id": "a11", "statement": "Waste in Cart B was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a12", "statement": "Cart B blocked the roadway at 18 Birch Lane during the Tuesday collection period."}], "base_state_json": "\"Operational handoff: At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316. The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart A. Intake reconciliation limits the reported asset to the candidate set consisting of Cart A and Cart B. Site records confirm that the reported cart was properly placed curbside by the 6:30 a.m. Tuesday cutoff and remained uncollected after the scheduled collection. The schedule assigns Tuesday organics to Cart A and Tuesday recycling to Cart B. Controlling dispatch records mark Cart A’s Tuesday missed collection urgent and Cart B’s routine. During the collection period, waste in Cart A was leaking; Cart B had no leaking waste and did not block the roadway. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316."}, {"path": [], "text": "The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart A."}], "policy_evidence": [{"path": [], "text": "Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}], "rules": [{"justification": "The report concerns Cart A, which was properly placed, timely, and missed; its scheduled material is organics and its controlling priority is urgent. Its leaking waste also satisfies the policy’s necessary condition for urgent status.", "target": "true", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The reported cart is one of Cart A and Cart B but is not Cart A, so it is Cart B. It was properly placed, timely, and missed; its scheduled material is recycling and its controlling priority is routine, while both policy-listed urgent conditions are explicitly absent.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316.", "negative_left": "At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316.", "negative_right": "The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart B.", "right": "The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart A."}, "verifier_independent_model": false}, "family": "scale-diverse-153-003", "id": "scale-diverse-153-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route it to the organics crew as urgent; route it to the recycling return-pickup crew with routine priority.", "true": "Route the report to the organics return-pickup crew with urgent priority."}, "instructions": "Decide whether this report should be routed to the organics return-pickup crew as an urgent case.", "type": "noul"}}, "state": "Operational handoff: At 7:18 a.m. on Tuesday, July 14, 2026, a handheld scan of the cart documented in Mara Chen’s missed-collection report at 18 Birch Lane returned factory serial QH-7316. The asset registry snapshot at 6:00 a.m. on Tuesday, July 14, 2026, records Cart A and Cart B as distinct carts and QH-7316 as the unique factory serial of Cart B. Intake reconciliation limits the reported asset to the candidate set consisting of Cart A and Cart B. Site records confirm that the reported cart was properly placed curbside by the 6:30 a.m. Tuesday cutoff and remained uncollected after the scheduled collection. The schedule assigns Tuesday organics to Cart A and Tuesday recycling to Cart B. Controlling dispatch records mark Cart A’s Tuesday missed collection urgent and Cart B’s routine. During the collection period, waste in Cart A was leaking; Cart B had no leaking waste and did not block the roadway. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}, "method": "c2d", "provenance": {"source_id": "diverse-153", "source_is_synthetic": true, "source_sha256": "f8cc49d0970071d68e752b9fc8ad734adab047d8cd33245838da80f07b42cdfc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, request, Tuesday timing, 18 Alder Court, Zone C, and organics-crew binding. The two focus spans are complete factual sentences. The counterfactual changes only the cart’s material designation from organics-only to landfill-waste-only; this conflicts with the scheduled material for eligibility purposes but does not contradict the unchanged image count, timestamps, placement, noncollection, or same-cart assertions. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction beyond the preserved governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"Do not route directly; send the report to the service desk for review or reject the return-pickup request.\",\"true\":\"Route directly to the Zone C organics crew for a return pickup.\"},\"instructions\":\"Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.\",\"type\":\"noul\"}},\"state\":{\"context\":\"Mara Chen submitted a Tuesday missed-pickup report for 18 Alder Court. The address registry assigns the property to Zone C. Her submission contains exactly two cart images, and each depicts a cart at 18 Alder Court at its stated capture time.\",\"evidence\":[\"At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material.\",\"The submitted placement image, captured at 5:55 a.m. Tuesday, shows compliant cart placement at 18 Alder Court at that time.\",\"The submitted noncollection image was captured at 7:18 p.m. Tuesday and shows the cart uncollected at 18 Alder Court at that time.\",\"Image comparison confirms that the cart in the placement image is the same cart in the noncollection image.\",\"At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore an organics-only material designation.\"],\"request\":\"Is this report ready to route to the Zone C organics crew for a return pickup?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material."}, {"path": ["state", "evidence", "4"], "text": "At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore an organics-only material designation."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material.", "negative_left": "At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material.", "negative_right": "At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore a landfill-waste-only material designation.", "right": "At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore an organics-only material designation."}, "verifier_independent_model": false}, "family": "scale-diverse-154-001", "id": "scale-diverse-154-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Mara Chen submitted a Tuesday missed-pickup report for 18 Alder Court. The address registry assigns the property to Zone C. Her submission contains exactly two cart images, and each depicts a cart at 18 Alder Court at its stated capture time.", "evidence": ["At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material.", "The submitted placement image, captured at 5:55 a.m. Tuesday, shows compliant cart placement at 18 Alder Court at that time.", "The submitted noncollection image was captured at 7:18 p.m. Tuesday and shows the cart uncollected at 18 Alder Court at that time.", "Image comparison confirms that the cart in the placement image is the same cart in the noncollection image.", "At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore an organics-only material designation."], "request": "Is this report ready to route to the Zone C organics crew for a return pickup?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, request, Tuesday timing, 18 Alder Court, Zone C, and organics-crew binding. The two focus spans are complete factual sentences. The counterfactual changes only the cart’s material designation from organics-only to landfill-waste-only; this conflicts with the scheduled material for eligibility purposes but does not contradict the unchanged image count, timestamps, placement, noncollection, or same-cart assertions. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction beyond the preserved governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"Do not route directly; send the report to the service desk for review or reject the return-pickup request.\",\"true\":\"Route directly to the Zone C organics crew for a return pickup.\"},\"instructions\":\"Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.\",\"type\":\"noul\"}},\"state\":{\"context\":\"Mara Chen submitted a Tuesday missed-pickup report for 18 Alder Court. The address registry assigns the property to Zone C. Her submission contains exactly two cart images, and each depicts a cart at 18 Alder Court at its stated capture time.\",\"evidence\":[\"At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material.\",\"The submitted placement image, captured at 5:55 a.m. Tuesday, shows compliant cart placement at 18 Alder Court at that time.\",\"The submitted noncollection image was captured at 7:18 p.m. Tuesday and shows the cart uncollected at 18 Alder Court at that time.\",\"Image comparison confirms that the cart in the placement image is the same cart in the noncollection image.\",\"At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore an organics-only material designation.\"],\"request\":\"Is this report ready to route to the Zone C organics crew for a return pickup?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material."}, {"path": ["state", "evidence", "4"], "text": "At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore an organics-only material designation."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material.", "negative_left": "At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material.", "negative_right": "At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore a landfill-waste-only material designation.", "right": "At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore an organics-only material designation."}, "verifier_independent_model": false}, "family": "scale-diverse-154-001", "id": "scale-diverse-154-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Mara Chen submitted a Tuesday missed-pickup report for 18 Alder Court. The address registry assigns the property to Zone C. Her submission contains exactly two cart images, and each depicts a cart at 18 Alder Court at its stated capture time.", "evidence": ["At 5:42 a.m. on the Tuesday of Mara Chen's report, the service schedule for 18 Alder Court listed organics as the sole scheduled material.", "The submitted placement image, captured at 5:55 a.m. Tuesday, shows compliant cart placement at 18 Alder Court at that time.", "The submitted noncollection image was captured at 7:18 p.m. Tuesday and shows the cart uncollected at 18 Alder Court at that time.", "Image comparison confirms that the cart in the placement image is the same cart in the noncollection image.", "At 8:16 p.m. that Tuesday, the cart depicted throughout the submitted images for Mara Chen's report bore a landfill-waste-only material designation."], "request": "Is this report ready to route to the Zone C organics crew for a return pickup?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing policy, request scope, and bindings to 18 Alder Court, Zone C, Tuesday, organics service, and return pickup. The two focus spans are complete factual sentences. The counterfactual changes only the observed cart designation from Organics to Residual Waste; this coherently creates a mismatch with the scheduled material rather than a contradictory duplicate measurement. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction beyond the preserved governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"Do not route directly; send the report to the service desk for review or reject the return-pickup request.\",\"true\":\"Route directly to the Zone C organics crew for a return pickup.\"},\"instructions\":\"Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.\",\"type\":\"noul\"}},\"state\":{\"context\":\"Operations handoff: Mara Chen filed Tuesday's missed-pickup report for 18 Alder Court and requested return collection. The address registry assigns the property to Zone C.\",\"evidence\":[\"The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.\",\"The material designation visible on the cart in each submitted image for Mara Chen's report is Organics.\",\"The submission manifest lists exactly two cart images, each taken at 18 Alder Court that Tuesday: a placement image timestamped 5:57 a.m. and a noncollection image timestamped 7:16 p.m.\",\"The placement image shows the cart positioned in full compliance with curb, clearance, lid, and access requirements at capture.\",\"Matching serial and damage marks establish that the evening image shows the same cart, which was still uncollected at capture.\"],\"request\":\"Is this report ready to route to the Zone C organics crew for a return pickup?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"path": ["state", "evidence", "1"], "text": "The material designation visible on the cart in each submitted image for Mara Chen's report is Organics."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "negative_left": "The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "negative_right": "The material designation visible on the cart in each submitted image for Mara Chen's report is Residual Waste.", "right": "The material designation visible on the cart in each submitted image for Mara Chen's report is Organics."}, "verifier_independent_model": false}, "family": "scale-diverse-154-003", "id": "scale-diverse-154-003-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Operations handoff: Mara Chen filed Tuesday's missed-pickup report for 18 Alder Court and requested return collection. The address registry assigns the property to Zone C.", "evidence": ["The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "The material designation visible on the cart in each submitted image for Mara Chen's report is Organics.", "The submission manifest lists exactly two cart images, each taken at 18 Alder Court that Tuesday: a placement image timestamped 5:57 a.m. and a noncollection image timestamped 7:16 p.m.", "The placement image shows the cart positioned in full compliance with curb, clearance, lid, and access requirements at capture.", "Matching serial and damage marks establish that the evening image shows the same cart, which was still uncollected at capture."], "request": "Is this report ready to route to the Zone C organics crew for a return pickup?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing policy, request scope, and bindings to 18 Alder Court, Zone C, Tuesday, organics service, and return pickup. The two focus spans are complete factual sentences. The counterfactual changes only the observed cart designation from Organics to Residual Waste; this coherently creates a mismatch with the scheduled material rather than a contradictory duplicate measurement. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction beyond the preserved governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"Do not route directly; send the report to the service desk for review or reject the return-pickup request.\",\"true\":\"Route directly to the Zone C organics crew for a return pickup.\"},\"instructions\":\"Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.\",\"type\":\"noul\"}},\"state\":{\"context\":\"Operations handoff: Mara Chen filed Tuesday's missed-pickup report for 18 Alder Court and requested return collection. The address registry assigns the property to Zone C.\",\"evidence\":[\"The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.\",\"The material designation visible on the cart in each submitted image for Mara Chen's report is Organics.\",\"The submission manifest lists exactly two cart images, each taken at 18 Alder Court that Tuesday: a placement image timestamped 5:57 a.m. and a noncollection image timestamped 7:16 p.m.\",\"The placement image shows the cart positioned in full compliance with curb, clearance, lid, and access requirements at capture.\",\"Matching serial and damage marks establish that the evening image shows the same cart, which was still uncollected at capture.\"],\"request\":\"Is this report ready to route to the Zone C organics crew for a return pickup?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"path": ["state", "evidence", "1"], "text": "The material designation visible on the cart in each submitted image for Mara Chen's report is Organics."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "negative_left": "The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "negative_right": "The material designation visible on the cart in each submitted image for Mara Chen's report is Residual Waste.", "right": "The material designation visible on the cart in each submitted image for Mara Chen's report is Organics."}, "verifier_independent_model": false}, "family": "scale-diverse-154-003", "id": "scale-diverse-154-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Operations handoff: Mara Chen filed Tuesday's missed-pickup report for 18 Alder Court and requested return collection. The address registry assigns the property to Zone C.", "evidence": ["The service handoff ledger records Organics as the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "The material designation visible on the cart in each submitted image for Mara Chen's report is Residual Waste.", "The submission manifest lists exactly two cart images, each taken at 18 Alder Court that Tuesday: a placement image timestamped 5:57 a.m. and a noncollection image timestamped 7:16 p.m.", "The placement image shows the cart positioned in full compliance with curb, clearance, lid, and access requirements at capture.", "Matching serial and damage marks establish that the evening image shows the same cart, which was still uncollected at capture."], "request": "Is this report ready to route to the Zone C organics crew for a return pickup?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same address, Zone C Wednesday trash route, gray-cart requirements, incident dates, missed-service facts, and nonhazardous contained-waste conditions. The only counterfactual change is the report-submission time, from 09:17 to 18:06 on July 9; it does not alter the governing 24-hour policy or create a contradictory duplicate assertion. The two focus-evidence spans are complete factual sentences, and neither context contains an answer label, code, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route control\",\"text\":\"The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026.\"},{\"speaker\":\"Service desk\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 09:17 UTC on Thursday, July 9, 2026.\"},{\"speaker\":\"Records coordinator\",\"text\":\"The address-zone file assigns 18 Alder Mews to Zone C. The collection schedule specifies Wednesday trash service using a gray cart. Inspection footage confirms that the resident submitted a gray cart and that it was curbside before 7:00 a.m. Wednesday. Crew records and the post-route footage show that the scheduled crew did not empty that cart.\"},{\"speaker\":\"Site inspector\",\"text\":\"The cart remains closed and all waste is fully contained. No waste is exposed, and neither the cart nor its contents are leaking or producing an odor. No animals are accessing or disturbing the waste, no pests are present, and there is no loose material outside the cart. The cart and waste create no obstruction and no acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 09:17 UTC on Thursday, July 9, 2026."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026.", "negative_left": "The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 18:06 UTC on Thursday, July 9, 2026.", "right": "The missed-collection report for 18 Alder Mews was submitted at 09:17 UTC on Thursday, July 9, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-155-003", "id": "scale-diverse-155-003-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route control", "text": "The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026."}, {"speaker": "Service desk", "text": "The missed-collection report for 18 Alder Mews was submitted at 09:17 UTC on Thursday, July 9, 2026."}, {"speaker": "Records coordinator", "text": "The address-zone file assigns 18 Alder Mews to Zone C. The collection schedule specifies Wednesday trash service using a gray cart. Inspection footage confirms that the resident submitted a gray cart and that it was curbside before 7:00 a.m. Wednesday. Crew records and the post-route footage show that the scheduled crew did not empty that cart."}, {"speaker": "Site inspector", "text": "The cart remains closed and all waste is fully contained. No waste is exposed, and neither the cart nor its contents are leaking or producing an odor. No animals are accessing or disturbing the waste, no pests are present, and there is no loose material outside the cart. The cart and waste create no obstruction and no acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same address, Zone C Wednesday trash route, gray-cart requirements, incident dates, missed-service facts, and nonhazardous contained-waste conditions. The only counterfactual change is the report-submission time, from 09:17 to 18:06 on July 9; it does not alter the governing 24-hour policy or create a contradictory duplicate assertion. The two focus-evidence spans are complete factual sentences, and neither context contains an answer label, code, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route control\",\"text\":\"The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026.\"},{\"speaker\":\"Service desk\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 09:17 UTC on Thursday, July 9, 2026.\"},{\"speaker\":\"Records coordinator\",\"text\":\"The address-zone file assigns 18 Alder Mews to Zone C. The collection schedule specifies Wednesday trash service using a gray cart. Inspection footage confirms that the resident submitted a gray cart and that it was curbside before 7:00 a.m. Wednesday. Crew records and the post-route footage show that the scheduled crew did not empty that cart.\"},{\"speaker\":\"Site inspector\",\"text\":\"The cart remains closed and all waste is fully contained. No waste is exposed, and neither the cart nor its contents are leaking or producing an odor. No animals are accessing or disturbing the waste, no pests are present, and there is no loose material outside the cart. The cart and waste create no obstruction and no acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 09:17 UTC on Thursday, July 9, 2026."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026.", "negative_left": "The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 18:06 UTC on Thursday, July 9, 2026.", "right": "The missed-collection report for 18 Alder Mews was submitted at 09:17 UTC on Thursday, July 9, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-155-003", "id": "scale-diverse-155-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route control", "text": "The operations log records closure of the relevant Zone C Wednesday trash route for 18 Alder Mews at 16:42 UTC on Wednesday, July 8, 2026."}, {"speaker": "Service desk", "text": "The missed-collection report for 18 Alder Mews was submitted at 18:06 UTC on Thursday, July 9, 2026."}, {"speaker": "Records coordinator", "text": "The address-zone file assigns 18 Alder Mews to Zone C. The collection schedule specifies Wednesday trash service using a gray cart. Inspection footage confirms that the resident submitted a gray cart and that it was curbside before 7:00 a.m. Wednesday. Crew records and the post-route footage show that the scheduled crew did not empty that cart."}, {"speaker": "Site inspector", "text": "The cart remains closed and all waste is fully contained. No waste is exposed, and neither the cart nor its contents are leaking or producing an odor. No animals are accessing or disturbing the waste, no pests are present, and there is no loose material outside the cart. The cart and waste create no obstruction and no acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing routing policy, eligibility conditions, thresholds, and requested decision scope. Both contexts retain the same address, zone, scheduled collection, cart type, route, and relevant date framework; the counterfactual changes only the report time from 9:55 a.m. to 11:05 a.m. The two focus-evidence spans are complete factual sentences. The revised report time does not contradict any unchanged assertion, and neither context contains a gold answer, output code, rule table, proposition identifier, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route audit\",\"text\":\"The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026.\"},{\"speaker\":\"Intake record\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 9:55 a.m. UTC+01:00 on 15 October 2026.\"},{\"speaker\":\"Records clerk\",\"text\":\"The address-zone file assigns 18 Alder Mews to Zone C. Zone C has scheduled trash collection on Wednesday, using gray carts. The container submitted at the address was a gray cart, recorded curbside at 6:50 a.m. Wednesday.\"},{\"speaker\":\"Evidence reviewer\",\"text\":\"Truck video and the crew log show that the scheduled crew passed 18 Alder Mews without emptying the submitted cart.\"},{\"speaker\":\"Site inspector\",\"text\":\"The cart lid is shut, and all remaining waste is fully contained. No waste is exposed; neither the waste nor the cart is leaking or producing an odor. No animals are accessing or disturbing the waste, no pests are present, and no loose material lies outside the cart. The waste and cart create no obstruction to the roadway, sidewalk access, or fire route, and there is no acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 9:55 a.m. UTC+01:00 on 15 October 2026."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026.", "negative_left": "The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 11:05 a.m. UTC+01:00 on 15 October 2026.", "right": "The missed-collection report for 18 Alder Mews was submitted at 9:55 a.m. UTC+01:00 on 15 October 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-155-004", "id": "scale-diverse-155-004-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route audit", "text": "The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026."}, {"speaker": "Intake record", "text": "The missed-collection report for 18 Alder Mews was submitted at 9:55 a.m. UTC+01:00 on 15 October 2026."}, {"speaker": "Records clerk", "text": "The address-zone file assigns 18 Alder Mews to Zone C. Zone C has scheduled trash collection on Wednesday, using gray carts. The container submitted at the address was a gray cart, recorded curbside at 6:50 a.m. Wednesday."}, {"speaker": "Evidence reviewer", "text": "Truck video and the crew log show that the scheduled crew passed 18 Alder Mews without emptying the submitted cart."}, {"speaker": "Site inspector", "text": "The cart lid is shut, and all remaining waste is fully contained. No waste is exposed; neither the waste nor the cart is leaking or producing an odor. No animals are accessing or disturbing the waste, no pests are present, and no loose material lies outside the cart. The waste and cart create no obstruction to the roadway, sidewalk access, or fire route, and there is no acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing routing policy, eligibility conditions, thresholds, and requested decision scope. Both contexts retain the same address, zone, scheduled collection, cart type, route, and relevant date framework; the counterfactual changes only the report time from 9:55 a.m. to 11:05 a.m. The two focus-evidence spans are complete factual sentences. The revised report time does not contradict any unchanged assertion, and neither context contains a gold answer, output code, rule table, proposition identifier, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route audit\",\"text\":\"The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026.\"},{\"speaker\":\"Intake record\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 9:55 a.m. UTC+01:00 on 15 October 2026.\"},{\"speaker\":\"Records clerk\",\"text\":\"The address-zone file assigns 18 Alder Mews to Zone C. Zone C has scheduled trash collection on Wednesday, using gray carts. The container submitted at the address was a gray cart, recorded curbside at 6:50 a.m. Wednesday.\"},{\"speaker\":\"Evidence reviewer\",\"text\":\"Truck video and the crew log show that the scheduled crew passed 18 Alder Mews without emptying the submitted cart.\"},{\"speaker\":\"Site inspector\",\"text\":\"The cart lid is shut, and all remaining waste is fully contained. No waste is exposed; neither the waste nor the cart is leaking or producing an odor. No animals are accessing or disturbing the waste, no pests are present, and no loose material lies outside the cart. The waste and cart create no obstruction to the roadway, sidewalk access, or fire route, and there is no acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 9:55 a.m. UTC+01:00 on 15 October 2026."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026.", "negative_left": "The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 11:05 a.m. UTC+01:00 on 15 October 2026.", "right": "The missed-collection report for 18 Alder Mews was submitted at 9:55 a.m. UTC+01:00 on 15 October 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-155-004", "id": "scale-diverse-155-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route audit", "text": "The relevant Zone C Wednesday trash route closed at 10:20 a.m. UTC+01:00 on 14 October 2026."}, {"speaker": "Intake record", "text": "The missed-collection report for 18 Alder Mews was submitted at 11:05 a.m. UTC+01:00 on 15 October 2026."}, {"speaker": "Records clerk", "text": "The address-zone file assigns 18 Alder Mews to Zone C. Zone C has scheduled trash collection on Wednesday, using gray carts. The container submitted at the address was a gray cart, recorded curbside at 6:50 a.m. Wednesday."}, {"speaker": "Evidence reviewer", "text": "Truck video and the crew log show that the scheduled crew passed 18 Alder Mews without emptying the submitted cart."}, {"speaker": "Site inspector", "text": "The cart lid is shut, and all remaining waste is fully contained. No waste is exposed; neither the waste nor the cart is leaking or producing an odor. No animals are accessing or disturbing the waste, no pests are present, and no loose material lies outside the cart. The waste and cart create no obstruction to the roadway, sidewalk access, or fire route, and there is no acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing instructions and score criteria, and neither context alters or omits any additional governing policy from the original state. Both contexts retain the same address, zone, scheduled material/day, cart, route event, and report path; the counterfactual coherently changes only the report-submission time from 3:55 p.m. to 5:05 p.m. The two focus-evidence spans are complete factual sentences. There are no conflicting duplicate facts, embedded answer codes, rule tables, proposition identifiers, label rationales, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route controller\",\"text\":\"The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026.\"},{\"speaker\":\"Records clerk\",\"text\":\"The address-zone file assigns 18 Alder Mews to Zone C. The collection schedule designates Wednesday for trash and specifies gray carts.\"},{\"speaker\":\"Resident\",\"text\":\"For Wednesday collection, I placed a gray cart curbside at 6:50 a.m. The lid remained shut, and all waste stayed fully inside.\"},{\"speaker\":\"Operations inspector\",\"text\":\"Service logs and vehicle footage show the scheduled crew passed without emptying the submitted cart. Inspection found no exposed waste, leakage, odor, animal access or disturbance, pests, or loose material outside the cart. Neither the cart nor the waste obstructed any roadway, sidewalk, access point, or fire route, and no acute dangerous spill was present.\"},{\"speaker\":\"Intake clerk\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 3:55 p.m. on Thursday, September 10, 2026.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026."}, {"path": ["4", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 3:55 p.m. on Thursday, September 10, 2026."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026.", "negative_left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 5:05 p.m. on Thursday, September 10, 2026.", "right": "The missed-collection report for 18 Alder Mews was submitted at 3:55 p.m. on Thursday, September 10, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-155-005", "id": "scale-diverse-155-005-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route controller", "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026."}, {"speaker": "Records clerk", "text": "The address-zone file assigns 18 Alder Mews to Zone C. The collection schedule designates Wednesday for trash and specifies gray carts."}, {"speaker": "Resident", "text": "For Wednesday collection, I placed a gray cart curbside at 6:50 a.m. The lid remained shut, and all waste stayed fully inside."}, {"speaker": "Operations inspector", "text": "Service logs and vehicle footage show the scheduled crew passed without emptying the submitted cart. Inspection found no exposed waste, leakage, odor, animal access or disturbance, pests, or loose material outside the cart. Neither the cart nor the waste obstructed any roadway, sidewalk, access point, or fire route, and no acute dangerous spill was present."}, {"speaker": "Intake clerk", "text": "The missed-collection report for 18 Alder Mews was submitted at 3:55 p.m. on Thursday, September 10, 2026."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing instructions and score criteria, and neither context alters or omits any additional governing policy from the original state. Both contexts retain the same address, zone, scheduled material/day, cart, route event, and report path; the counterfactual coherently changes only the report-submission time from 3:55 p.m. to 5:05 p.m. The two focus-evidence spans are complete factual sentences. There are no conflicting duplicate facts, embedded answer codes, rule tables, proposition identifiers, label rationales, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route controller\",\"text\":\"The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026.\"},{\"speaker\":\"Records clerk\",\"text\":\"The address-zone file assigns 18 Alder Mews to Zone C. The collection schedule designates Wednesday for trash and specifies gray carts.\"},{\"speaker\":\"Resident\",\"text\":\"For Wednesday collection, I placed a gray cart curbside at 6:50 a.m. The lid remained shut, and all waste stayed fully inside.\"},{\"speaker\":\"Operations inspector\",\"text\":\"Service logs and vehicle footage show the scheduled crew passed without emptying the submitted cart. Inspection found no exposed waste, leakage, odor, animal access or disturbance, pests, or loose material outside the cart. Neither the cart nor the waste obstructed any roadway, sidewalk, access point, or fire route, and no acute dangerous spill was present.\"},{\"speaker\":\"Intake clerk\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 3:55 p.m. on Thursday, September 10, 2026.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026."}, {"path": ["4", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 3:55 p.m. on Thursday, September 10, 2026."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026.", "negative_left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 5:05 p.m. on Thursday, September 10, 2026.", "right": "The missed-collection report for 18 Alder Mews was submitted at 3:55 p.m. on Thursday, September 10, 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-155-005", "id": "scale-diverse-155-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route controller", "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 4:40 p.m. on Wednesday, September 9, 2026."}, {"speaker": "Records clerk", "text": "The address-zone file assigns 18 Alder Mews to Zone C. The collection schedule designates Wednesday for trash and specifies gray carts."}, {"speaker": "Resident", "text": "For Wednesday collection, I placed a gray cart curbside at 6:50 a.m. The lid remained shut, and all waste stayed fully inside."}, {"speaker": "Operations inspector", "text": "Service logs and vehicle footage show the scheduled crew passed without emptying the submitted cart. Inspection found no exposed waste, leakage, odor, animal access or disturbance, pests, or loose material outside the cart. Neither the cart nor the waste obstructed any roadway, sidewalk, access point, or fire route, and no acute dangerous spill was present."}, {"speaker": "Intake clerk", "text": "The missed-collection report for 18 Alder Mews was submitted at 5:05 p.m. on Thursday, September 10, 2026."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same reporter, address, service type, assigned crew/zone, Tuesday report setting, and decision scope, while leaving the governing rubric unchanged. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the photo capture time from 2:06 p.m. to 1:32 p.m.; a photo showing an unemptied cart before the crew’s 1:47 p.m. completion is coherent and does not conflict with the unchanged facts. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a trash missed-collection report for 18 Alder Lane at 2:10 p.m. on Tuesday. The municipal service registry confirms that the address receives city collection, and Tuesday’s route schedule lists trash service there. Her report identifies the city-issued gray cart shown in the attached photo as the container needing collection. Dispatch records assign the Zone C trash crew to the address. Field notes confirm that the cart was at the curb and accessible to that crew when it completed service. The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday. The metadata of Maya Chen's attached photo records its capture time as 2:06 p.m. on the report Tuesday. The photo shows the identified gray cart still unemptied and is the report’s only evidence concerning whether it was emptied. The report documents no spill, leakage, pests, odor, roadway blockage, accessible-walkway obstruction, or other immediate sanitation or obstruction hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday."}, {"path": [], "text": "The metadata of Maya Chen's attached photo records its capture time as 2:06 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday.", "negative_left": "The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday.", "negative_right": "The metadata of Maya Chen's attached photo records its capture time as 1:32 p.m. on the report Tuesday.", "right": "The metadata of Maya Chen's attached photo records its capture time as 2:06 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-156-003", "id": "scale-diverse-156-003-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a trash missed-collection report for 18 Alder Lane at 2:10 p.m. on Tuesday. The municipal service registry confirms that the address receives city collection, and Tuesday’s route schedule lists trash service there. Her report identifies the city-issued gray cart shown in the attached photo as the container needing collection. Dispatch records assign the Zone C trash crew to the address. Field notes confirm that the cart was at the curb and accessible to that crew when it completed service. The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday. The metadata of Maya Chen's attached photo records its capture time as 2:06 p.m. on the report Tuesday. The photo shows the identified gray cart still unemptied and is the report’s only evidence concerning whether it was emptied. The report documents no spill, leakage, pests, odor, roadway blockage, accessible-walkway obstruction, or other immediate sanitation or obstruction hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same reporter, address, service type, assigned crew/zone, Tuesday report setting, and decision scope, while leaving the governing rubric unchanged. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the photo capture time from 2:06 p.m. to 1:32 p.m.; a photo showing an unemptied cart before the crew’s 1:47 p.m. completion is coherent and does not conflict with the unchanged facts. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a trash missed-collection report for 18 Alder Lane at 2:10 p.m. on Tuesday. The municipal service registry confirms that the address receives city collection, and Tuesday’s route schedule lists trash service there. Her report identifies the city-issued gray cart shown in the attached photo as the container needing collection. Dispatch records assign the Zone C trash crew to the address. Field notes confirm that the cart was at the curb and accessible to that crew when it completed service. The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday. The metadata of Maya Chen's attached photo records its capture time as 2:06 p.m. on the report Tuesday. The photo shows the identified gray cart still unemptied and is the report’s only evidence concerning whether it was emptied. The report documents no spill, leakage, pests, odor, roadway blockage, accessible-walkway obstruction, or other immediate sanitation or obstruction hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday."}, {"path": [], "text": "The metadata of Maya Chen's attached photo records its capture time as 2:06 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday.", "negative_left": "The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday.", "negative_right": "The metadata of Maya Chen's attached photo records its capture time as 1:32 p.m. on the report Tuesday.", "right": "The metadata of Maya Chen's attached photo records its capture time as 2:06 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-156-003", "id": "scale-diverse-156-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a trash missed-collection report for 18 Alder Lane at 2:10 p.m. on Tuesday. The municipal service registry confirms that the address receives city collection, and Tuesday’s route schedule lists trash service there. Her report identifies the city-issued gray cart shown in the attached photo as the container needing collection. Dispatch records assign the Zone C trash crew to the address. Field notes confirm that the cart was at the curb and accessible to that crew when it completed service. The assigned Zone C trash crew's service log records completion at 18 Alder Lane at 1:47 p.m. on the report Tuesday. The metadata of Maya Chen's attached photo records its capture time as 1:32 p.m. on the report Tuesday. The photo shows the identified gray cart still unemptied and is the report’s only evidence concerning whether it was emptied. The report documents no spill, leakage, pests, odor, roadway blockage, accessible-walkway obstruction, or other immediate sanitation or obstruction hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric and instructions, while both contexts retain the same reporter, address, waste stream, service day, crew, and relevant Tuesday time frame. The two evidence spans are complete factual sentences. The counterfactual changes only the photo capture time from 2:04 p.m. to 1:43 p.m.; a photo taken before the crew’s 1:51 p.m. completion is coherent with the remaining facts and creates no duplicate or contradictory measurement. Neither context includes an answer, score code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Field note: Municipal collection service covers 18 Alder Lane, with trash scheduled there on the report Tuesday. Maya Chen filed a report at 2:10 p.m. that Tuesday requesting trash collection. The Zone C trash crew was assigned to the address and completed service there that day. Access records confirm the reported city gray cart was available to that crew when service was completed. The attached photo depicts that same reported gray cart still unemptied and is the report’s only evidence of whether it was emptied. The metadata record for Maya Chen's attached photo lists its capture time as 2:04 p.m. on the report Tuesday. The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday. The report documents no immediate sanitation or obstruction hazard from the missed collection: no spill, pests, leakage, blocked road, or blocked accessible walkway.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The metadata record for Maya Chen's attached photo lists its capture time as 2:04 p.m. on the report Tuesday."}, {"path": [], "text": "The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The metadata record for Maya Chen's attached photo lists its capture time as 2:04 p.m. on the report Tuesday.", "negative_left": "The metadata record for Maya Chen's attached photo lists its capture time as 1:43 p.m. on the report Tuesday.", "negative_right": "The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday.", "right": "The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-156-004", "id": "scale-diverse-156-004-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Field note: Municipal collection service covers 18 Alder Lane, with trash scheduled there on the report Tuesday. Maya Chen filed a report at 2:10 p.m. that Tuesday requesting trash collection. The Zone C trash crew was assigned to the address and completed service there that day. Access records confirm the reported city gray cart was available to that crew when service was completed. The attached photo depicts that same reported gray cart still unemptied and is the report’s only evidence of whether it was emptied. The metadata record for Maya Chen's attached photo lists its capture time as 2:04 p.m. on the report Tuesday. The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday. The report documents no immediate sanitation or obstruction hazard from the missed collection: no spill, pests, leakage, blocked road, or blocked accessible walkway."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric and instructions, while both contexts retain the same reporter, address, waste stream, service day, crew, and relevant Tuesday time frame. The two evidence spans are complete factual sentences. The counterfactual changes only the photo capture time from 2:04 p.m. to 1:43 p.m.; a photo taken before the crew’s 1:51 p.m. completion is coherent with the remaining facts and creates no duplicate or contradictory measurement. Neither context includes an answer, score code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Field note: Municipal collection service covers 18 Alder Lane, with trash scheduled there on the report Tuesday. Maya Chen filed a report at 2:10 p.m. that Tuesday requesting trash collection. The Zone C trash crew was assigned to the address and completed service there that day. Access records confirm the reported city gray cart was available to that crew when service was completed. The attached photo depicts that same reported gray cart still unemptied and is the report’s only evidence of whether it was emptied. The metadata record for Maya Chen's attached photo lists its capture time as 2:04 p.m. on the report Tuesday. The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday. The report documents no immediate sanitation or obstruction hazard from the missed collection: no spill, pests, leakage, blocked road, or blocked accessible walkway.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The metadata record for Maya Chen's attached photo lists its capture time as 2:04 p.m. on the report Tuesday."}, {"path": [], "text": "The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The metadata record for Maya Chen's attached photo lists its capture time as 2:04 p.m. on the report Tuesday.", "negative_left": "The metadata record for Maya Chen's attached photo lists its capture time as 1:43 p.m. on the report Tuesday.", "negative_right": "The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday.", "right": "The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "scale-diverse-156-004", "id": "scale-diverse-156-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Field note: Municipal collection service covers 18 Alder Lane, with trash scheduled there on the report Tuesday. Maya Chen filed a report at 2:10 p.m. that Tuesday requesting trash collection. The Zone C trash crew was assigned to the address and completed service there that day. Access records confirm the reported city gray cart was available to that crew when service was completed. The attached photo depicts that same reported gray cart still unemptied and is the report’s only evidence of whether it was emptied. The metadata record for Maya Chen's attached photo lists its capture time as 1:43 p.m. on the report Tuesday. The service log for the assigned Zone C trash crew lists its completion time at 18 Alder Lane as 1:51 p.m. on the report Tuesday. The report documents no immediate sanitation or obstruction hazard from the missed collection: no spill, pests, leakage, blocked road, or blocked accessible walkway."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric, and neither context alters or invents policy. Both contexts retain Maya Chen, 18 Alder Lane, Tuesday trash service, the Zone C crew, and the same report scope and relevant time path. The two evidence spans are complete factual sentences. The counterfactual changes only the photo capture time from 1:23 p.m. to 10:16 a.m.; this places the sole evidence about the unemptied cart before the 11:47 a.m. crew completion without creating a duplicate or contradictory timestamp. Neither context contains a gold answer, score code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"At 2:10 p.m. Tuesday, Maya Chen filed a missed-collection report for 18 Alder Lane. Municipal service records list the address as active for collection and schedule trash service there on Tuesdays. Her report specifically requested trash collection. The routing log assigned the Zone C trash crew to service the address that day. On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m. On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 1:23 p.m. At the crew’s documented completion, the reported city-issued gray cart was positioned where the assigned crew could access it. The cart depicted in the attachment is the same gray cart identified in Maya’s report, and the image shows it unemptied. The evidence audit identifies that attachment as the report’s only evidence concerning whether the cart was emptied. The report documents no spill, pests, leakage, roadway blockage, walkway obstruction, or other immediate sanitation or obstruction hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m."}, {"path": [], "text": "On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 1:23 p.m."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m.", "negative_left": "On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m.", "negative_right": "On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 10:16 a.m.", "right": "On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 1:23 p.m."}, "verifier_independent_model": false}, "family": "scale-diverse-156-005", "id": "scale-diverse-156-005-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "At 2:10 p.m. Tuesday, Maya Chen filed a missed-collection report for 18 Alder Lane. Municipal service records list the address as active for collection and schedule trash service there on Tuesdays. Her report specifically requested trash collection. The routing log assigned the Zone C trash crew to service the address that day. On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m. On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 1:23 p.m. At the crew’s documented completion, the reported city-issued gray cart was positioned where the assigned crew could access it. The cart depicted in the attachment is the same gray cart identified in Maya’s report, and the image shows it unemptied. The evidence audit identifies that attachment as the report’s only evidence concerning whether the cart was emptied. The report documents no spill, pests, leakage, roadway blockage, walkway obstruction, or other immediate sanitation or obstruction hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric, and neither context alters or invents policy. Both contexts retain Maya Chen, 18 Alder Lane, Tuesday trash service, the Zone C crew, and the same report scope and relevant time path. The two evidence spans are complete factual sentences. The counterfactual changes only the photo capture time from 1:23 p.m. to 10:16 a.m.; this places the sole evidence about the unemptied cart before the 11:47 a.m. crew completion without creating a duplicate or contradictory timestamp. Neither context contains a gold answer, score code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"At 2:10 p.m. Tuesday, Maya Chen filed a missed-collection report for 18 Alder Lane. Municipal service records list the address as active for collection and schedule trash service there on Tuesdays. Her report specifically requested trash collection. The routing log assigned the Zone C trash crew to service the address that day. On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m. On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 1:23 p.m. At the crew’s documented completion, the reported city-issued gray cart was positioned where the assigned crew could access it. The cart depicted in the attachment is the same gray cart identified in Maya’s report, and the image shows it unemptied. The evidence audit identifies that attachment as the report’s only evidence concerning whether the cart was emptied. The report documents no spill, pests, leakage, roadway blockage, walkway obstruction, or other immediate sanitation or obstruction hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m."}, {"path": [], "text": "On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 1:23 p.m."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m.", "negative_left": "On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m.", "negative_right": "On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 10:16 a.m.", "right": "On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 1:23 p.m."}, "verifier_independent_model": false}, "family": "scale-diverse-156-005", "id": "scale-diverse-156-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "At 2:10 p.m. Tuesday, Maya Chen filed a missed-collection report for 18 Alder Lane. Municipal service records list the address as active for collection and schedule trash service there on Tuesdays. Her report specifically requested trash collection. The routing log assigned the Zone C trash crew to service the address that day. On the report Tuesday, the assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 11:47 a.m. On the report Tuesday, the recorded capture time of Maya Chen's attached photo was 10:16 a.m. At the crew’s documented completion, the reported city-issued gray cart was positioned where the assigned crew could access it. The cart depicted in the attachment is the same gray cart identified in Maya’s report, and the image shows it unemptied. The evidence audit identifies that attachment as the report’s only evidence concerning whether the cart was emptied. The report documents no spill, pests, leakage, roadway blockage, walkway obstruction, or other immediate sanitation or obstruction hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing booking policy without altering its exceptions, thresholds, readiness requirements, or routing rules. They preserve the request’s room, date, time, attendance, equipment, and decision scope. The two focus-evidence spans are complete factual sentences rather than policy or instructions. The counterfactual coherently changes only HU-827’s equipment classification to nonstandard; no unchanged assertion in that context says HU-827 is standard, and ownership and exemption from kitchen use do not require standard-equipment status. Neither context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Intake clerk\",\"text\":\"At 08:45 on September 18, 2026, the intake log recorded a request for the standard activity room on September 29, 2026, from 6–8 p.m., for a 38-person lecture. The room request form was attached. The requested equipment consists only of 40 chairs, the built-in projector, and hot-water urns; the chairs are standard furniture, and the projector is built-in audiovisual equipment.\"},{\"speaker\":\"Records clerk\",\"text\":\"At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827.\"},{\"speaker\":\"Asset custodian\",\"text\":\"At 09:35, the asset register confirmed that HU-314 and HU-827 are owned by the community center. The organizer confirmed that every requested urn would be used for packaged tea and that the request includes no kitchen use apart from the urn service.\"},{\"speaker\":\"Equipment inspector\",\"text\":\"At 10:05 on September 18, 2026, the equipment inspection recorded both unit HU-314 and unit HU-827 as standard equipment.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["1", "text"], "text": "At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827."}, {"path": ["3", "text"], "text": "At 10:05 on September 18, 2026, the equipment inspection recorded both unit HU-314 and unit HU-827 as standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827.", "negative_left": "At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827.", "negative_right": "At 10:05 on September 18, 2026, the equipment inspection recorded unit HU-314 as standard equipment and unit HU-827 as nonstandard equipment.", "right": "At 10:05 on September 18, 2026, the equipment inspection recorded both unit HU-314 and unit HU-827 as standard equipment."}, "verifier_independent_model": false}, "family": "scale-diverse-157-001", "id": "scale-diverse-157-001-base", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Intake clerk", "text": "At 08:45 on September 18, 2026, the intake log recorded a request for the standard activity room on September 29, 2026, from 6–8 p.m., for a 38-person lecture. The room request form was attached. The requested equipment consists only of 40 chairs, the built-in projector, and hot-water urns; the chairs are standard furniture, and the projector is built-in audiovisual equipment."}, {"speaker": "Records clerk", "text": "At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827."}, {"speaker": "Asset custodian", "text": "At 09:35, the asset register confirmed that HU-314 and HU-827 are owned by the community center. The organizer confirmed that every requested urn would be used for packaged tea and that the request includes no kitchen use apart from the urn service."}, {"speaker": "Equipment inspector", "text": "At 10:05 on September 18, 2026, the equipment inspection recorded both unit HU-314 and unit HU-827 as standard equipment."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_low"}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing booking policy without altering its exceptions, thresholds, readiness requirements, or routing rules. They preserve the request’s room, date, time, attendance, equipment, and decision scope. The two focus-evidence spans are complete factual sentences rather than policy or instructions. The counterfactual coherently changes only HU-827’s equipment classification to nonstandard; no unchanged assertion in that context says HU-827 is standard, and ownership and exemption from kitchen use do not require standard-equipment status. Neither context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Intake clerk\",\"text\":\"At 08:45 on September 18, 2026, the intake log recorded a request for the standard activity room on September 29, 2026, from 6–8 p.m., for a 38-person lecture. The room request form was attached. The requested equipment consists only of 40 chairs, the built-in projector, and hot-water urns; the chairs are standard furniture, and the projector is built-in audiovisual equipment.\"},{\"speaker\":\"Records clerk\",\"text\":\"At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827.\"},{\"speaker\":\"Asset custodian\",\"text\":\"At 09:35, the asset register confirmed that HU-314 and HU-827 are owned by the community center. The organizer confirmed that every requested urn would be used for packaged tea and that the request includes no kitchen use apart from the urn service.\"},{\"speaker\":\"Equipment inspector\",\"text\":\"At 10:05 on September 18, 2026, the equipment inspection recorded both unit HU-314 and unit HU-827 as standard equipment.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["1", "text"], "text": "At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827."}, {"path": ["3", "text"], "text": "At 10:05 on September 18, 2026, the equipment inspection recorded both unit HU-314 and unit HU-827 as standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827.", "negative_left": "At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827.", "negative_right": "At 10:05 on September 18, 2026, the equipment inspection recorded unit HU-314 as standard equipment and unit HU-827 as nonstandard equipment.", "right": "At 10:05 on September 18, 2026, the equipment inspection recorded both unit HU-314 and unit HU-827 as standard equipment."}, "verifier_independent_model": false}, "family": "scale-diverse-157-001", "id": "scale-diverse-157-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Intake clerk", "text": "At 08:45 on September 18, 2026, the intake log recorded a request for the standard activity room on September 29, 2026, from 6–8 p.m., for a 38-person lecture. The room request form was attached. The requested equipment consists only of 40 chairs, the built-in projector, and hot-water urns; the chairs are standard furniture, and the projector is built-in audiovisual equipment."}, {"speaker": "Records clerk", "text": "At 09:12 on September 18, 2026, the September 29, 2026 activity-room request listed exactly two hot-water urns, units HU-314 and HU-827."}, {"speaker": "Asset custodian", "text": "At 09:35, the asset register confirmed that HU-314 and HU-827 are owned by the community center. The organizer confirmed that every requested urn would be used for packaged tea and that the request includes no kitchen use apart from the urn service."}, {"speaker": "Equipment inspector", "text": "At 10:05 on September 18, 2026, the equipment inspection recorded unit HU-314 as standard equipment and unit HU-827 as nonstandard equipment."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_medium"}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the coordinator’s governing routing, readiness, exception, and complexity rules unchanged, and preserve the same activity-room request, date, time, attendance, equipment, and manifest bindings. The two focus-evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual changes only HU-846’s equipment classification from standard to nonstandard; no unchanged statement classifies that urn as standard, so it creates no contradictory duplicate assertion. Neither context includes a gold option, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"Operational handoff: the request is for the standard activity room on September 29, 2026, from 6–8 p.m., for 38 attendees. The complete equipment field lists 40 chairs, the built-in projector, and the hot-water urns documented in the final handoff manifest. The room request form is attached.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"},{\"speaker\":\"Site attendant\",\"text\":\"All requested chairs are standard furniture, and the requested projector is built-in audiovisual equipment. Each urn asset named in the final handoff manifest is owned by the community center and designated for packaged tea. No kitchen use is requested apart from that urn service.\"},{\"speaker\":\"Handoff manifest\",\"text\":\"The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846.\"},{\"speaker\":\"Equipment register\",\"text\":\"The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 and HU-846 as standard equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["3", "text"], "text": "The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846."}, {"path": ["4", "text"], "text": "The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 and HU-846 as standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846.", "negative_left": "The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846.", "negative_right": "The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 as standard equipment and HU-846 as nonstandard equipment.", "right": "The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 and HU-846 as standard equipment."}, "verifier_independent_model": false}, "family": "scale-diverse-157-003", "id": "scale-diverse-157-003-base", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Booking clerk", "text": "Operational handoff: the request is for the standard activity room on September 29, 2026, from 6–8 p.m., for 38 attendees. The complete equipment field lists 40 chairs, the built-in projector, and the hot-water urns documented in the final handoff manifest. The room request form is attached."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}, {"speaker": "Site attendant", "text": "All requested chairs are standard furniture, and the requested projector is built-in audiovisual equipment. Each urn asset named in the final handoff manifest is owned by the community center and designated for packaged tea. No kitchen use is requested apart from that urn service."}, {"speaker": "Handoff manifest", "text": "The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846."}, {"speaker": "Equipment register", "text": "The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 and HU-846 as standard equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_low"}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the coordinator’s governing routing, readiness, exception, and complexity rules unchanged, and preserve the same activity-room request, date, time, attendance, equipment, and manifest bindings. The two focus-evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual changes only HU-846’s equipment classification from standard to nonstandard; no unchanged statement classifies that urn as standard, so it creates no contradictory duplicate assertion. Neither context includes a gold option, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"Operational handoff: the request is for the standard activity room on September 29, 2026, from 6–8 p.m., for 38 attendees. The complete equipment field lists 40 chairs, the built-in projector, and the hot-water urns documented in the final handoff manifest. The room request form is attached.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"},{\"speaker\":\"Site attendant\",\"text\":\"All requested chairs are standard furniture, and the requested projector is built-in audiovisual equipment. Each urn asset named in the final handoff manifest is owned by the community center and designated for packaged tea. No kitchen use is requested apart from that urn service.\"},{\"speaker\":\"Handoff manifest\",\"text\":\"The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846.\"},{\"speaker\":\"Equipment register\",\"text\":\"The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 and HU-846 as standard equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["3", "text"], "text": "The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846."}, {"path": ["4", "text"], "text": "The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 and HU-846 as standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846.", "negative_left": "The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846.", "negative_right": "The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 as standard equipment and HU-846 as nonstandard equipment.", "right": "The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 and HU-846 as standard equipment."}, "verifier_independent_model": false}, "family": "scale-diverse-157-003", "id": "scale-diverse-157-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Booking clerk", "text": "Operational handoff: the request is for the standard activity room on September 29, 2026, from 6–8 p.m., for 38 attendees. The complete equipment field lists 40 chairs, the built-in projector, and the hot-water urns documented in the final handoff manifest. The room request form is attached."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}, {"speaker": "Site attendant", "text": "All requested chairs are standard furniture, and the requested projector is built-in audiovisual equipment. Each urn asset named in the final handoff manifest is owned by the community center and designated for packaged tea. No kitchen use is requested apart from that urn service."}, {"speaker": "Handoff manifest", "text": "The final handoff manifest for the September 29, 2026 activity-room request lists exactly two hot-water urns, identified as HU-731 and HU-846."}, {"speaker": "Equipment register", "text": "The community center's equipment register at 14:20 on September 28, 2026 categorizes HU-731 as standard equipment and HU-846 as nonstandard equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_medium"}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking policy verbatim, while the unchanged questions preserve the request instructions and choice criteria. The activity-room entity, September 29, 2026 date, 6–8 p.m. time, attendance, and requested equipment remain bound consistently. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only HU-682 from standard to nonstandard equipment; no unchanged assertion classifies that asset as standard, and center ownership remains compatible with nonstandard status. Neither context contains an answer code, explicit selected option, classifier instruction, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Intake clerk\",\"text\":\"The attached room request form seeks the standard activity room on September 29, 2026, from 6–8 p.m. for a 38-person lecture. Its complete equipment list consists of 40 chairs, the built-in projector, and hot-water-urn service.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"},{\"speaker\":\"Asset records clerk\",\"text\":\"The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682.\"},{\"speaker\":\"Equipment registrar\",\"text\":\"The community center’s equipment register, effective throughout September 29, 2026, classifies assets HU-417 and HU-682 as standard equipment.\"},{\"speaker\":\"Site attendant\",\"text\":\"Inspection records identify each requested chair as standard furniture and confirm the projector is built-in audiovisual equipment. The requested urns belong to the community center and will be used only for packaged tea. No kitchen use is planned beyond that urn service.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["2", "text"], "text": "The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682."}, {"path": ["3", "text"], "text": "The community center’s equipment register, effective throughout September 29, 2026, classifies assets HU-417 and HU-682 as standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682.", "negative_left": "The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682.", "negative_right": "The community center’s equipment register, effective throughout September 29, 2026, classifies asset HU-417 as standard equipment and asset HU-682 as nonstandard equipment.", "right": "The community center’s equipment register, effective throughout September 29, 2026, classifies assets HU-417 and HU-682 as standard equipment."}, "verifier_independent_model": false}, "family": "scale-diverse-157-004", "id": "scale-diverse-157-004-base", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Intake clerk", "text": "The attached room request form seeks the standard activity room on September 29, 2026, from 6–8 p.m. for a 38-person lecture. Its complete equipment list consists of 40 chairs, the built-in projector, and hot-water-urn service."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}, {"speaker": "Asset records clerk", "text": "The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682."}, {"speaker": "Equipment registrar", "text": "The community center’s equipment register, effective throughout September 29, 2026, classifies assets HU-417 and HU-682 as standard equipment."}, {"speaker": "Site attendant", "text": "Inspection records identify each requested chair as standard furniture and confirm the projector is built-in audiovisual equipment. The requested urns belong to the community center and will be used only for packaged tea. No kitchen use is planned beyond that urn service."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_low"}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking policy verbatim, while the unchanged questions preserve the request instructions and choice criteria. The activity-room entity, September 29, 2026 date, 6–8 p.m. time, attendance, and requested equipment remain bound consistently. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only HU-682 from standard to nonstandard equipment; no unchanged assertion classifies that asset as standard, and center ownership remains compatible with nonstandard status. Neither context contains an answer code, explicit selected option, classifier instruction, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Intake clerk\",\"text\":\"The attached room request form seeks the standard activity room on September 29, 2026, from 6–8 p.m. for a 38-person lecture. Its complete equipment list consists of 40 chairs, the built-in projector, and hot-water-urn service.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"},{\"speaker\":\"Asset records clerk\",\"text\":\"The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682.\"},{\"speaker\":\"Equipment registrar\",\"text\":\"The community center’s equipment register, effective throughout September 29, 2026, classifies assets HU-417 and HU-682 as standard equipment.\"},{\"speaker\":\"Site attendant\",\"text\":\"Inspection records identify each requested chair as standard furniture and confirm the projector is built-in audiovisual equipment. The requested urns belong to the community center and will be used only for packaged tea. No kitchen use is planned beyond that urn service.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["2", "text"], "text": "The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682."}, {"path": ["3", "text"], "text": "The community center’s equipment register, effective throughout September 29, 2026, classifies assets HU-417 and HU-682 as standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682.", "negative_left": "The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682.", "negative_right": "The community center’s equipment register, effective throughout September 29, 2026, classifies asset HU-417 as standard equipment and asset HU-682 as nonstandard equipment.", "right": "The community center’s equipment register, effective throughout September 29, 2026, classifies assets HU-417 and HU-682 as standard equipment."}, "verifier_independent_model": false}, "family": "scale-diverse-157-004", "id": "scale-diverse-157-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Intake clerk", "text": "The attached room request form seeks the standard activity room on September 29, 2026, from 6–8 p.m. for a 38-person lecture. Its complete equipment list consists of 40 chairs, the built-in projector, and hot-water-urn service."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}, {"speaker": "Asset records clerk", "text": "The September 29, 2026 activity-room request lists exactly two hot-water urns, assets HU-417 and HU-682."}, {"speaker": "Equipment registrar", "text": "The community center’s equipment register, effective throughout September 29, 2026, classifies asset HU-417 as standard equipment and asset HU-682 as nonstandard equipment."}, {"speaker": "Site attendant", "text": "Inspection records identify each requested chair as standard furniture and confirm the projector is built-in audiovisual equipment. The requested urns belong to the community center and will be used only for packaged tea. No kitchen use is planned beyond that urn service."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_medium"}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same center attendance policy and preserve Maya, Activity Studio B, October 18, and the attendance-estimate decision. The two focus-evidence spans are complete factual sentences describing the register scheme and Jordan’s recorded status; they are not governing-policy instructions. The counterfactual makes a single coherent observational change from Q7 to M4, consistent with the register’s mutually exclusive statuses and the unchanged payroll and roster facts. Neither context states a decision option, verified total, classifier label, proposition ID, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual set-membership, cardinality, or staff-role relationships rather than final classifications or policy conclusions. The focus atom is the factual question of whether Jordan is a center attendant. The base and counter assignments are jointly realizable and differ only on that focus: Jordan may be an attendant in the base and a non-staff participant in the counter. The policy evidence preserves the attendance-counting and exclusion rule from the original state; question-originating choice criteria need not be duplicated. Both rules include enough conditions to determine their exact totals, including the necessary competing staff-category exclusions in the 32-person rule.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 and a4 establish exactly 31 organizer-identified expected participants other than Jordan who are not facility staff and therefore are counted. a2 and a3 establish that Jordan is identified and expected but is excluded as a center attendant. Thus the verified estimate is exactly 31.", "rule_index": 0, "sound": true}, {"reason": "a1 and a4 establish 31 counted participants other than Jordan. a2 establishes Jordan is an organizer-identified expected participant, while the refutations of a3, a5, and a6 exclude every stated facility-staff category. Jordan is therefore also counted, yielding exactly 32.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly 31 people other than Jordan Lee are expected participants in Maya's October 18 event in Activity Studio B and are identified by Maya."}, {"id": "a2", "statement": "Jordan Lee is an expected participant in Maya's October 18 event in Activity Studio B and is identified by Maya."}, {"id": "a3", "statement": "Jordan Lee is a center attendant at the community center on October 18."}, {"id": "a4", "statement": "None of the 31 expected participants counted in a1 is facility staff at the community center on October 18."}, {"id": "a5", "statement": "Jordan Lee is a cleaner at the community center on October 18."}, {"id": "a6", "statement": "Jordan Lee is facility staff at the community center on October 18 in a role other than center attendant or cleaner."}], "base_state_json": "\"Maya’s finalized roster identifies Jordan Lee as an expected participant in her October 18 event in Activity Studio B, together with exactly 31 other identified expected participants. Employment records confirm that none of those 31 other participants is facility staff at the community center that day. A payroll reconciliation found no cleaner assignment for Jordan and no facility-staff position beyond the status recorded in the completed register. The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day. Jordan Lee's entry in the completed October 18 community-center register bears status code Q7. The facility booking coordinator must verify the attendance estimate. Center policy counts all expected event participants identified by the organizer but excludes center attendants, cleaners, and other facility staff from the reservation’s attendance estimate.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day."}, {"path": [], "text": "Jordan Lee's entry in the completed October 18 community-center register bears status code Q7."}], "policy_evidence": [{"path": [], "text": "The facility booking coordinator must verify the attendance estimate."}, {"path": [], "text": "Center policy counts all expected event participants identified by the organizer but excludes center attendants, cleaners, and other facility staff from the reservation’s attendance estimate."}], "rules": [{"justification": "The 31 identified participants other than Jordan are counted because none is facility staff. Jordan is excluded because Jordan is a center attendant, so the verified estimate is exactly 31.", "target": "exactly_31", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The 31 identified participants other than Jordan are counted. Jordan is also counted because Jordan is identified as an expected participant and is neither a center attendant, a cleaner, nor other facility staff, so the verified estimate is exactly 32.", "target": "exactly_32", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day.", "negative_left": "The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day.", "negative_right": "Jordan Lee's entry in the completed October 18 community-center register bears status code M4.", "right": "Jordan Lee's entry in the completed October 18 community-center register bears status code Q7."}, "verifier_independent_model": false}, "family": "scale-diverse-158-002", "id": "scale-diverse-158-002-base", "input": {"questions": {"decision": {"criteria": {"exactly_26": "The verified attendance estimate is exactly 26 people.", "exactly_28": "The verified attendance estimate is exactly 28 people.", "exactly_31": "The verified attendance estimate is exactly 31 people.", "exactly_32": "The verified attendance estimate is exactly 32 people.", "exactly_4": "The verified attendance estimate is exactly four people.", "none_of_above": "The verified attendance estimate is not any of the five listed totals."}, "instructions": "Select the option that exactly matches the verified attendance estimate under the stated center policy. Each substantive option applies only when the verified estimate equals its listed number; choose none_of_above if no listed number is exact.", "type": "choice"}}, "state": "Maya’s finalized roster identifies Jordan Lee as an expected participant in her October 18 event in Activity Studio B, together with exactly 31 other identified expected participants. Employment records confirm that none of those 31 other participants is facility staff at the community center that day. A payroll reconciliation found no cleaner assignment for Jordan and no facility-staff position beyond the status recorded in the completed register. The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day. Jordan Lee's entry in the completed October 18 community-center register bears status code Q7. The facility booking coordinator must verify the attendance estimate. Center policy counts all expected event participants identified by the organizer but excludes center attendants, cleaners, and other facility staff from the reservation’s attendance estimate."}, "method": "c2d", "provenance": {"source_id": "diverse-158", "source_is_synthetic": true, "source_sha256": "e1075a8056bdd8009837096e1f207634a30bf0ed596dcd439102ba1e2c1d47d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "exactly_31"}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same center attendance policy and preserve Maya, Activity Studio B, October 18, and the attendance-estimate decision. The two focus-evidence spans are complete factual sentences describing the register scheme and Jordan’s recorded status; they are not governing-policy instructions. The counterfactual makes a single coherent observational change from Q7 to M4, consistent with the register’s mutually exclusive statuses and the unchanged payroll and roster facts. Neither context states a decision option, verified total, classifier label, proposition ID, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual set-membership, cardinality, or staff-role relationships rather than final classifications or policy conclusions. The focus atom is the factual question of whether Jordan is a center attendant. The base and counter assignments are jointly realizable and differ only on that focus: Jordan may be an attendant in the base and a non-staff participant in the counter. The policy evidence preserves the attendance-counting and exclusion rule from the original state; question-originating choice criteria need not be duplicated. Both rules include enough conditions to determine their exact totals, including the necessary competing staff-category exclusions in the 32-person rule.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 and a4 establish exactly 31 organizer-identified expected participants other than Jordan who are not facility staff and therefore are counted. a2 and a3 establish that Jordan is identified and expected but is excluded as a center attendant. Thus the verified estimate is exactly 31.", "rule_index": 0, "sound": true}, {"reason": "a1 and a4 establish 31 counted participants other than Jordan. a2 establishes Jordan is an organizer-identified expected participant, while the refutations of a3, a5, and a6 exclude every stated facility-staff category. Jordan is therefore also counted, yielding exactly 32.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly 31 people other than Jordan Lee are expected participants in Maya's October 18 event in Activity Studio B and are identified by Maya."}, {"id": "a2", "statement": "Jordan Lee is an expected participant in Maya's October 18 event in Activity Studio B and is identified by Maya."}, {"id": "a3", "statement": "Jordan Lee is a center attendant at the community center on October 18."}, {"id": "a4", "statement": "None of the 31 expected participants counted in a1 is facility staff at the community center on October 18."}, {"id": "a5", "statement": "Jordan Lee is a cleaner at the community center on October 18."}, {"id": "a6", "statement": "Jordan Lee is facility staff at the community center on October 18 in a role other than center attendant or cleaner."}], "base_state_json": "\"Maya’s finalized roster identifies Jordan Lee as an expected participant in her October 18 event in Activity Studio B, together with exactly 31 other identified expected participants. Employment records confirm that none of those 31 other participants is facility staff at the community center that day. A payroll reconciliation found no cleaner assignment for Jordan and no facility-staff position beyond the status recorded in the completed register. The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day. Jordan Lee's entry in the completed October 18 community-center register bears status code Q7. The facility booking coordinator must verify the attendance estimate. Center policy counts all expected event participants identified by the organizer but excludes center attendants, cleaners, and other facility staff from the reservation’s attendance estimate.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day."}, {"path": [], "text": "Jordan Lee's entry in the completed October 18 community-center register bears status code Q7."}], "policy_evidence": [{"path": [], "text": "The facility booking coordinator must verify the attendance estimate."}, {"path": [], "text": "Center policy counts all expected event participants identified by the organizer but excludes center attendants, cleaners, and other facility staff from the reservation’s attendance estimate."}], "rules": [{"justification": "The 31 identified participants other than Jordan are counted because none is facility staff. Jordan is excluded because Jordan is a center attendant, so the verified estimate is exactly 31.", "target": "exactly_31", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The 31 identified participants other than Jordan are counted. Jordan is also counted because Jordan is identified as an expected participant and is neither a center attendant, a cleaner, nor other facility staff, so the verified estimate is exactly 32.", "target": "exactly_32", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day.", "negative_left": "The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day.", "negative_right": "Jordan Lee's entry in the completed October 18 community-center register bears status code M4.", "right": "Jordan Lee's entry in the completed October 18 community-center register bears status code Q7."}, "verifier_independent_model": false}, "family": "scale-diverse-158-002", "id": "scale-diverse-158-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"exactly_26": "The verified attendance estimate is exactly 26 people.", "exactly_28": "The verified attendance estimate is exactly 28 people.", "exactly_31": "The verified attendance estimate is exactly 31 people.", "exactly_32": "The verified attendance estimate is exactly 32 people.", "exactly_4": "The verified attendance estimate is exactly four people.", "none_of_above": "The verified attendance estimate is not any of the five listed totals."}, "instructions": "Select the option that exactly matches the verified attendance estimate under the stated center policy. Each substantive option applies only when the verified estimate equals its listed number; choose none_of_above if no listed number is exact.", "type": "choice"}}, "state": "Maya’s finalized roster identifies Jordan Lee as an expected participant in her October 18 event in Activity Studio B, together with exactly 31 other identified expected participants. Employment records confirm that none of those 31 other participants is facility staff at the community center that day. A payroll reconciliation found no cleaner assignment for Jordan and no facility-staff position beyond the status recorded in the completed register. The completed October 18 community-center register assigned every person exactly one of its exhaustive, mutually exclusive status codes, with code Q7 denoting center attendants and code M4 denoting people who held no facility-staff role that day. Jordan Lee's entry in the completed October 18 community-center register bears status code M4. The facility booking coordinator must verify the attendance estimate. Center policy counts all expected event participants identified by the organizer but excludes center attendants, cleaners, and other facility staff from the reservation’s attendance estimate."}, "method": "c2d", "provenance": {"source_id": "diverse-158", "source_is_synthetic": true, "source_sha256": "e1075a8056bdd8009837096e1f207634a30bf0ed596dcd439102ba1e2c1d47d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "exactly_32"}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing booking-readiness policy exactly and retain the same request scope, facility, teaching-kitchen reservation, October 14 date, 6–9 p.m. time, and 22-attendee binding. The two focus-evidence spans are complete factual sentences rather than instructions or embedded outputs. The counterfactual coherently changes F-17 from a Food Use Form to a Room Setup Form; because the register assigns each file exactly one type and no other attachments are present, this creates no contradictory duplicate classification or measurement. Neither context contains a gold answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship; a6 is a single negative relationship concerning the set of applicable forms rather than a bundled final classification. The focus a7 concerns whether a form is included, not the decision policy itself. The base and counter assignments can differ only in whether the Food Use Form is included while all other facts remain fixed. Policy evidence correctly cites the original state’s substantive readiness and kitchen-form rules; instructions and criteria in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the selected kitchen, all required reservation details, that no form other than the Food Use Form applies, and inclusion of that form. Therefore every stated readiness requirement is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the selected space is a kitchen while the required Food Use Form is absent. A missing required form is sufficient for a false decision regardless of the other satisfied details.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The selected space for the reservation request is the teaching kitchen."}, {"id": "a2", "statement": "The reservation request includes the event date October 14."}, {"id": "a3", "statement": "The reservation request includes the event time interval 6–9 p.m."}, {"id": "a4", "statement": "The reservation request includes an attendance estimate of 22 attendees."}, {"id": "a5", "statement": "The reservation request includes an equipment-needs specification."}, {"id": "a6", "statement": "No form other than the Food Use Form is required for this selected teaching-kitchen reservation."}, {"id": "a7", "statement": "The reservation request includes a Food Use Form."}], "base_state_json": "{\"context\":\"The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold.\",\"evidence\":[\"On August 20, staff checked the center’s current space-specific form matrix; its teaching-kitchen row listed the Food Use Form as the sole applicable form.\",\"At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present.\",\"The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Food Use Form.\"],\"request\":\"Is this request booking-ready for confirmation by the facility booking coordinator?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present."}, {"path": ["evidence", "2"], "text": "The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Food Use Form."}], "policy_evidence": [{"path": ["context"], "text": "The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold."}], "rules": [{"justification": "The request supplies every required reservation detail, and the required Food Use Form is included; no other form applies to the selected teaching kitchen.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The selected teaching kitchen requires a Food Use Form, and the request does not include that required form, so it is not booking-ready despite containing the other details.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present.", "negative_left": "At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present.", "negative_right": "The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Room Setup Form.", "right": "The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Food Use Form."}, "verifier_independent_model": false}, "family": "scale-diverse-159-001", "id": "scale-diverse-159-001-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks at least one required detail or applicable form, so it is not ready for coordinator confirmation.", "true": "The request contains all required reservation details and every form required for the selected space, so it is ready for coordinator confirmation."}, "instructions": "Answer yes only if all booking-readiness requirements in the context are satisfied. Answer no if any required detail or form is missing, even when capacity, equipment, and staffing are adequate.", "type": "noul"}}, "state": {"context": "The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold.", "evidence": ["On August 20, staff checked the center’s current space-specific form matrix; its teaching-kitchen row listed the Food Use Form as the sole applicable form.", "At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present.", "The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Food Use Form."], "request": "Is this request booking-ready for confirmation by the facility booking coordinator?"}}, "method": "c2d", "provenance": {"source_id": "diverse-159", "source_is_synthetic": true, "source_sha256": "bef4997ffa59657f3e9d777596ff439ee1a02aa66ae3676bb18f51233df96d1a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing booking-readiness policy exactly and retain the same request scope, facility, teaching-kitchen reservation, October 14 date, 6–9 p.m. time, and 22-attendee binding. The two focus-evidence spans are complete factual sentences rather than instructions or embedded outputs. The counterfactual coherently changes F-17 from a Food Use Form to a Room Setup Form; because the register assigns each file exactly one type and no other attachments are present, this creates no contradictory duplicate classification or measurement. Neither context contains a gold answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship; a6 is a single negative relationship concerning the set of applicable forms rather than a bundled final classification. The focus a7 concerns whether a form is included, not the decision policy itself. The base and counter assignments can differ only in whether the Food Use Form is included while all other facts remain fixed. Policy evidence correctly cites the original state’s substantive readiness and kitchen-form rules; instructions and criteria in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the selected kitchen, all required reservation details, that no form other than the Food Use Form applies, and inclusion of that form. Therefore every stated readiness requirement is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the selected space is a kitchen while the required Food Use Form is absent. A missing required form is sufficient for a false decision regardless of the other satisfied details.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The selected space for the reservation request is the teaching kitchen."}, {"id": "a2", "statement": "The reservation request includes the event date October 14."}, {"id": "a3", "statement": "The reservation request includes the event time interval 6–9 p.m."}, {"id": "a4", "statement": "The reservation request includes an attendance estimate of 22 attendees."}, {"id": "a5", "statement": "The reservation request includes an equipment-needs specification."}, {"id": "a6", "statement": "No form other than the Food Use Form is required for this selected teaching-kitchen reservation."}, {"id": "a7", "statement": "The reservation request includes a Food Use Form."}], "base_state_json": "{\"context\":\"The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold.\",\"evidence\":[\"On August 20, staff checked the center’s current space-specific form matrix; its teaching-kitchen row listed the Food Use Form as the sole applicable form.\",\"At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present.\",\"The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Food Use Form.\"],\"request\":\"Is this request booking-ready for confirmation by the facility booking coordinator?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present."}, {"path": ["evidence", "2"], "text": "The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Food Use Form."}], "policy_evidence": [{"path": ["context"], "text": "The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold."}], "rules": [{"justification": "The request supplies every required reservation detail, and the required Food Use Form is included; no other form applies to the selected teaching kitchen.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The selected teaching kitchen requires a Food Use Form, and the request does not include that required form, so it is not booking-ready despite containing the other details.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present.", "negative_left": "At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present.", "negative_right": "The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Room Setup Form.", "right": "The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Food Use Form."}, "verifier_independent_model": false}, "family": "scale-diverse-159-001", "id": "scale-diverse-159-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks at least one required detail or applicable form, so it is not ready for coordinator confirmation.", "true": "The request contains all required reservation details and every form required for the selected space, so it is ready for coordinator confirmation."}, "instructions": "Answer yes only if all booking-readiness requirements in the context are satisfied. Answer no if any required detail or form is missing, even when capacity, equipment, and staffing are adequate.", "type": "noul"}}, "state": {"context": "The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold.", "evidence": ["On August 20, staff checked the center’s current space-specific form matrix; its teaching-kitchen row listed the Food Use Form as the sole applicable form.", "At 4:10 p.m. on September 30, the Eastgate Community Center archived the complete reservation request for the teaching kitchen on October 14 from 6–9 p.m. for 22 attendees as files C-41 and F-17, with C-41 recording the equipment needs and no other pages or attachments present.", "The center's document register assigns each file exactly one form type and classifies C-41 as a reservation cover sheet and F-17 as a Room Setup Form."], "request": "Is this request booking-ready for confirmation by the facility booking coordinator?"}}, "method": "c2d", "provenance": {"source_id": "diverse-159", "source_is_synthetic": true, "source_sha256": "bef4997ffa59657f3e9d777596ff439ee1a02aa66ae3676bb18f51233df96d1a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original booking-readiness requirements without adding exceptions, priorities, or missing-evidence defaults, and they preserve the teaching-kitchen reservation’s date, time, equipment, form, and decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only the finalized attendance from 22 to 27; this coherently exceeds the unchanged capacity of 24 rather than contradicting another attendance measurement. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Roster update\",\"text\":\"At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 22 planned attendees for the October 24, 2026, dumpling workshop.\"},{\"speaker\":\"Facilities record\",\"text\":\"At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people.\"},{\"speaker\":\"Scheduling coordinator\",\"text\":\"Later that morning, the coordinator confirmed that the teaching kitchen was available on October 24, 2026, for the entire requested interval from 5:30 p.m. through 8:30 p.m.\"},{\"speaker\":\"Organizer\",\"text\":\"That afternoon, the organizer finalized the workshop's equipment request: all four induction burners and no other equipment. The organizer also uploaded a signed copy of the center's required food-preparation paperwork for activities involving food preparation onsite.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Before the reservation review concluded, the technician tested all four requested induction burners and recorded each one as operational for the October 24 workshop.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 22 planned attendees for the October 24, 2026, dumpling workshop."}, {"path": ["1", "text"], "text": "At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 22 planned attendees for the October 24, 2026, dumpling workshop.", "negative_left": "At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 27 planned attendees for the October 24, 2026, dumpling workshop.", "negative_right": "At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people.", "right": "At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people."}, "verifier_independent_model": false}, "family": "scale-diverse-160-001", "id": "scale-diverse-160-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Roster update", "text": "At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 22 planned attendees for the October 24, 2026, dumpling workshop."}, {"speaker": "Facilities record", "text": "At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people."}, {"speaker": "Scheduling coordinator", "text": "Later that morning, the coordinator confirmed that the teaching kitchen was available on October 24, 2026, for the entire requested interval from 5:30 p.m. through 8:30 p.m."}, {"speaker": "Organizer", "text": "That afternoon, the organizer finalized the workshop's equipment request: all four induction burners and no other equipment. The organizer also uploaded a signed copy of the center's required food-preparation paperwork for activities involving food preparation onsite."}, {"speaker": "Maintenance technician", "text": "Before the reservation review concluded, the technician tested all four requested induction burners and recorded each one as operational for the October 24 workshop."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original booking-readiness requirements without adding exceptions, priorities, or missing-evidence defaults, and they preserve the teaching-kitchen reservation’s date, time, equipment, form, and decision scope. The two evidence spans are complete factual sentences. The counterfactual changes only the finalized attendance from 22 to 27; this coherently exceeds the unchanged capacity of 24 rather than contradicting another attendance measurement. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Roster update\",\"text\":\"At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 22 planned attendees for the October 24, 2026, dumpling workshop.\"},{\"speaker\":\"Facilities record\",\"text\":\"At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people.\"},{\"speaker\":\"Scheduling coordinator\",\"text\":\"Later that morning, the coordinator confirmed that the teaching kitchen was available on October 24, 2026, for the entire requested interval from 5:30 p.m. through 8:30 p.m.\"},{\"speaker\":\"Organizer\",\"text\":\"That afternoon, the organizer finalized the workshop's equipment request: all four induction burners and no other equipment. The organizer also uploaded a signed copy of the center's required food-preparation paperwork for activities involving food preparation onsite.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Before the reservation review concluded, the technician tested all four requested induction burners and recorded each one as operational for the October 24 workshop.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 22 planned attendees for the October 24, 2026, dumpling workshop."}, {"path": ["1", "text"], "text": "At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 22 planned attendees for the October 24, 2026, dumpling workshop.", "negative_left": "At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 27 planned attendees for the October 24, 2026, dumpling workshop.", "negative_right": "At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people.", "right": "At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people."}, "verifier_independent_model": false}, "family": "scale-diverse-160-001", "id": "scale-diverse-160-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Roster update", "text": "At 9:00 a.m. on September 18, 2026, the organizer's finalized roster recorded 27 planned attendees for the October 24, 2026, dumpling workshop."}, {"speaker": "Facilities record", "text": "At 10:15 a.m. on September 18, 2026, the facilities register listed the teaching kitchen's maximum attendance capacity as 24 people."}, {"speaker": "Scheduling coordinator", "text": "Later that morning, the coordinator confirmed that the teaching kitchen was available on October 24, 2026, for the entire requested interval from 5:30 p.m. through 8:30 p.m."}, {"speaker": "Organizer", "text": "That afternoon, the organizer finalized the workshop's equipment request: all four induction burners and no other equipment. The organizer also uploaded a signed copy of the center's required food-preparation paperwork for activities involving food preparation onsite."}, {"speaker": "Maintenance technician", "text": "Before the reservation review concluded, the technician tested all four requested induction burners and recorded each one as operational for the October 24 workshop."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged question preserves all booking-readiness criteria, while both contexts retain the same reservation, facility, date, time, equipment, and form-confirmation scope. The two evidence spans are complete factual sentences; the counterfactual coherently changes only planned attendance from 22 to 27 while retaining the stated capacity of 24, and neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Reservation field note\",\"text\":\"The organizer requested the teaching kitchen for a dumpling workshop on October 24, 2026, from 5:30 p.m. through 8:30 p.m.\"},{\"speaker\":\"Facility scheduler\",\"text\":\"The teaching kitchen is available on October 24, 2026, and the entire requested interval from 5:30 p.m. through 8:30 p.m. is open.\"},{\"speaker\":\"Organizer\",\"text\":\"The workshop requires all four induction burners and no other equipment.\"},{\"speaker\":\"Equipment technician\",\"text\":\"Inspection records show that each of the four requested induction burners is operational.\"},{\"speaker\":\"Attendance record\",\"text\":\"The planned attendance for the October 24, 2026, dumpling workshop is 22 people.\"},{\"speaker\":\"Posted room notice\",\"text\":\"For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people.\"},{\"speaker\":\"Records clerk\",\"text\":\"The organizer’s uploaded sheet is signed. It is the center’s required onsite food-preparation form—the paperwork required when food will be prepared in the facility.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["4", "text"], "text": "The planned attendance for the October 24, 2026, dumpling workshop is 22 people."}, {"path": ["5", "text"], "text": "For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The planned attendance for the October 24, 2026, dumpling workshop is 22 people.", "negative_left": "The planned attendance for the October 24, 2026, dumpling workshop is 27 people.", "negative_right": "For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people.", "right": "For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people."}, "verifier_independent_model": false}, "family": "scale-diverse-160-004", "id": "scale-diverse-160-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Reservation field note", "text": "The organizer requested the teaching kitchen for a dumpling workshop on October 24, 2026, from 5:30 p.m. through 8:30 p.m."}, {"speaker": "Facility scheduler", "text": "The teaching kitchen is available on October 24, 2026, and the entire requested interval from 5:30 p.m. through 8:30 p.m. is open."}, {"speaker": "Organizer", "text": "The workshop requires all four induction burners and no other equipment."}, {"speaker": "Equipment technician", "text": "Inspection records show that each of the four requested induction burners is operational."}, {"speaker": "Attendance record", "text": "The planned attendance for the October 24, 2026, dumpling workshop is 22 people."}, {"speaker": "Posted room notice", "text": "For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people."}, {"speaker": "Records clerk", "text": "The organizer’s uploaded sheet is signed. It is the center’s required onsite food-preparation form—the paperwork required when food will be prepared in the facility."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged question preserves all booking-readiness criteria, while both contexts retain the same reservation, facility, date, time, equipment, and form-confirmation scope. The two evidence spans are complete factual sentences; the counterfactual coherently changes only planned attendance from 22 to 27 while retaining the stated capacity of 24, and neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Reservation field note\",\"text\":\"The organizer requested the teaching kitchen for a dumpling workshop on October 24, 2026, from 5:30 p.m. through 8:30 p.m.\"},{\"speaker\":\"Facility scheduler\",\"text\":\"The teaching kitchen is available on October 24, 2026, and the entire requested interval from 5:30 p.m. through 8:30 p.m. is open.\"},{\"speaker\":\"Organizer\",\"text\":\"The workshop requires all four induction burners and no other equipment.\"},{\"speaker\":\"Equipment technician\",\"text\":\"Inspection records show that each of the four requested induction burners is operational.\"},{\"speaker\":\"Attendance record\",\"text\":\"The planned attendance for the October 24, 2026, dumpling workshop is 22 people.\"},{\"speaker\":\"Posted room notice\",\"text\":\"For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people.\"},{\"speaker\":\"Records clerk\",\"text\":\"The organizer’s uploaded sheet is signed. It is the center’s required onsite food-preparation form—the paperwork required when food will be prepared in the facility.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["4", "text"], "text": "The planned attendance for the October 24, 2026, dumpling workshop is 22 people."}, {"path": ["5", "text"], "text": "For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The planned attendance for the October 24, 2026, dumpling workshop is 22 people.", "negative_left": "The planned attendance for the October 24, 2026, dumpling workshop is 27 people.", "negative_right": "For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people.", "right": "For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people."}, "verifier_independent_model": false}, "family": "scale-diverse-160-004", "id": "scale-diverse-160-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Reservation field note", "text": "The organizer requested the teaching kitchen for a dumpling workshop on October 24, 2026, from 5:30 p.m. through 8:30 p.m."}, {"speaker": "Facility scheduler", "text": "The teaching kitchen is available on October 24, 2026, and the entire requested interval from 5:30 p.m. through 8:30 p.m. is open."}, {"speaker": "Organizer", "text": "The workshop requires all four induction burners and no other equipment."}, {"speaker": "Equipment technician", "text": "Inspection records show that each of the four requested induction burners is operational."}, {"speaker": "Attendance record", "text": "The planned attendance for the October 24, 2026, dumpling workshop is 27 people."}, {"speaker": "Posted room notice", "text": "For the October 24, 2026, dumpling workshop, the teaching kitchen's posted maximum attendance capacity is 24 people."}, {"speaker": "Records clerk", "text": "The organizer’s uploaded sheet is signed. It is the center’s required onsite food-preparation form—the paperwork required when food will be prepared in the facility."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete scoring policy, and neither context alters or adds governing rules. Both contexts retain the same organizer, meeting, date, time, reservation scope, attendance, space, kitchen/layout, and staffing bindings; only Larkspur Hospitality LLC’s outside-vendor status changes. The two evidence spans are complete factual sentences. Treating Larkspur as an internal organizational unit that furnishes goods or services is coherent with the unchanged one-entry roster and creates no duplicate or contradictory measurement. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Field note: Priya Shah’s final reservation for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. records exactly 38 attendees. The booking is limited to Meeting Room B, with no other facility space included. Kitchen access is not included, and the existing room layout will remain unchanged without reconfiguration. The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC. Larkspur Hospitality LLC is a commercial vendor unaffiliated with Priya Shah’s neighborhood association. The operations schedule assigns exactly zero added staff hours; opening, routine support, ordinary cleaning, and closing will be handled entirely during regular shifts.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC."}, {"path": [], "text": "Larkspur Hospitality LLC is a commercial vendor unaffiliated with Priya Shah’s neighborhood association."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC.", "negative_left": "The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC.", "negative_right": "Larkspur Hospitality LLC is an internal unit of Priya Shah’s neighborhood association and is not a vendor.", "right": "Larkspur Hospitality LLC is a commercial vendor unaffiliated with Priya Shah’s neighborhood association."}, "verifier_independent_model": false}, "family": "scale-diverse-161-004", "id": "scale-diverse-161-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Field note: Priya Shah’s final reservation for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. records exactly 38 attendees. The booking is limited to Meeting Room B, with no other facility space included. Kitchen access is not included, and the existing room layout will remain unchanged without reconfiguration. The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC. Larkspur Hospitality LLC is a commercial vendor unaffiliated with Priya Shah’s neighborhood association. The operations schedule assigns exactly zero added staff hours; opening, routine support, ordinary cleaning, and closing will be handled entirely during regular shifts."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete scoring policy, and neither context alters or adds governing rules. Both contexts retain the same organizer, meeting, date, time, reservation scope, attendance, space, kitchen/layout, and staffing bindings; only Larkspur Hospitality LLC’s outside-vendor status changes. The two evidence spans are complete factual sentences. Treating Larkspur as an internal organizational unit that furnishes goods or services is coherent with the unchanged one-entry roster and creates no duplicate or contradictory measurement. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Field note: Priya Shah’s final reservation for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. records exactly 38 attendees. The booking is limited to Meeting Room B, with no other facility space included. Kitchen access is not included, and the existing room layout will remain unchanged without reconfiguration. The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC. Larkspur Hospitality LLC is a commercial vendor unaffiliated with Priya Shah’s neighborhood association. The operations schedule assigns exactly zero added staff hours; opening, routine support, ordinary cleaning, and closing will be handled entirely during regular shifts.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC."}, {"path": [], "text": "Larkspur Hospitality LLC is a commercial vendor unaffiliated with Priya Shah’s neighborhood association."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC.", "negative_left": "The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC.", "negative_right": "Larkspur Hospitality LLC is an internal unit of Priya Shah’s neighborhood association and is not a vendor.", "right": "Larkspur Hospitality LLC is a commercial vendor unaffiliated with Priya Shah’s neighborhood association."}, "verifier_independent_model": false}, "family": "scale-diverse-161-004", "id": "scale-diverse-161-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Field note: Priya Shah’s final reservation for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. records exactly 38 attendees. The booking is limited to Meeting Room B, with no other facility space included. Kitchen access is not included, and the existing room layout will remain unchanged without reconfiguration. The exhaustive roster of all persons and organizations scheduled by Priya Shah’s final request to furnish goods or services for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains exactly one entry, Larkspur Hospitality LLC. Larkspur Hospitality LLC is an internal unit of Priya Shah’s neighborhood association and is not a vendor. The operations schedule assigns exactly zero added staff hours; opening, routine support, ordinary cleaning, and closing will be handled entirely during regular shifts."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all scoring criteria and instructions, while both contexts retain the same organizer, meeting, date, time, request revision, attendance, room, kitchen, setup, and staffing bindings. The two evidence spans are complete factual sentences, and the counterfactual coherently changes only the final roster from listing one outside vendor to being empty without creating duplicate or contradictory facts. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. Revision N-417’s complete outside-vendor roster lists Juniper Table Catering LLC. The next morning, staff reviewed that revision. It specifies exactly 38 attendees and reserves only Meeting Room B, with no other facility space listed. The kitchen-use field is marked “No,” and the setup instructions retain the room’s standard arrangement without reconfiguration. The staffing worksheet records zero added attendant hours and zero added custodial hours; routine opening, closing, and cleaning remain within regular shifts.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m."}, {"path": [], "text": "Revision N-417’s complete outside-vendor roster lists Juniper Table Catering LLC."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m.", "negative_left": "At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m.", "negative_right": "Revision N-417’s complete outside-vendor roster is empty.", "right": "Revision N-417’s complete outside-vendor roster lists Juniper Table Catering LLC."}, "verifier_independent_model": false}, "family": "scale-diverse-161-005", "id": "scale-diverse-161-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. Revision N-417’s complete outside-vendor roster lists Juniper Table Catering LLC. The next morning, staff reviewed that revision. It specifies exactly 38 attendees and reserves only Meeting Room B, with no other facility space listed. The kitchen-use field is marked “No,” and the setup instructions retain the room’s standard arrangement without reconfiguration. The staffing worksheet records zero added attendant hours and zero added custodial hours; routine opening, closing, and cleaning remain within regular shifts."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all scoring criteria and instructions, while both contexts retain the same organizer, meeting, date, time, request revision, attendance, room, kitchen, setup, and staffing bindings. The two evidence spans are complete factual sentences, and the counterfactual coherently changes only the final roster from listing one outside vendor to being empty without creating duplicate or contradictory facts. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. Revision N-417’s complete outside-vendor roster lists Juniper Table Catering LLC. The next morning, staff reviewed that revision. It specifies exactly 38 attendees and reserves only Meeting Room B, with no other facility space listed. The kitchen-use field is marked “No,” and the setup instructions retain the room’s standard arrangement without reconfiguration. The staffing worksheet records zero added attendant hours and zero added custodial hours; routine opening, closing, and cleaning remain within regular shifts.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m."}, {"path": [], "text": "Revision N-417’s complete outside-vendor roster lists Juniper Table Catering LLC."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m.", "negative_left": "At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m.", "negative_right": "Revision N-417’s complete outside-vendor roster is empty.", "right": "Revision N-417’s complete outside-vendor roster lists Juniper Table Catering LLC."}, "verifier_independent_model": false}, "family": "scale-diverse-161-005", "id": "scale-diverse-161-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "At 4:37 p.m. on September 29, 2026, Priya Shah designated revision N-417 as the complete and exclusive final request for her neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. Revision N-417’s complete outside-vendor roster is empty. The next morning, staff reviewed that revision. It specifies exactly 38 attendees and reserves only Meeting Room B, with no other facility space listed. The kitchen-use field is marked “No,” and the setup instructions retain the room’s standard arrangement without reconfiguration. The staffing worksheet records zero added attendant hours and zero added custodial hours; routine opening, closing, and cleaning remain within regular shifts."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring rubric and instructions, while both contexts retain the original state’s governing policy without alteration. The booking entity, date, time, venue type, request scope, and decision target remain fixed. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the attendance estimate from 117 to 84; this does not conflict with the unchanged staffing, facility, or service facts. Neither context embeds a gold answer, output code, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\":\"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\",\"evidence\":[\"The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment.\",\"The attendance estimate filed under booking reference RCC-1847 records 117 attendees.\",\"The signed reservation form and facilities review are complete. The event will use the ordinary activity room without supervised audiovisual or sports equipment, and the schedule requires only the routine closing check beyond the recorded staffing assignment.\"],\"request\":\"Determine the booking’s operational-complexity level for routing and staffing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment."}, {"path": ["evidence", "1"], "text": "The attendance estimate filed under booking reference RCC-1847 records 117 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment.", "negative_left": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment.", "negative_right": "The attendance estimate filed under booking reference RCC-1847 records 84 attendees.", "right": "The attendance estimate filed under booking reference RCC-1847 records 117 attendees."}, "verifier_independent_model": false}, "family": "scale-diverse-162-004", "id": "scale-diverse-162-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment.", "The attendance estimate filed under booking reference RCC-1847 records 117 attendees.", "The signed reservation form and facilities review are complete. The event will use the ordinary activity room without supervised audiovisual or sports equipment, and the schedule requires only the routine closing check beyond the recorded staffing assignment."], "request": "Determine the booking’s operational-complexity level for routing and staffing."}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring rubric and instructions, while both contexts retain the original state’s governing policy without alteration. The booking entity, date, time, venue type, request scope, and decision target remain fixed. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the attendance estimate from 117 to 84; this does not conflict with the unchanged staffing, facility, or service facts. Neither context embeds a gold answer, output code, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\":\"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\",\"evidence\":[\"The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment.\",\"The attendance estimate filed under booking reference RCC-1847 records 117 attendees.\",\"The signed reservation form and facilities review are complete. The event will use the ordinary activity room without supervised audiovisual or sports equipment, and the schedule requires only the routine closing check beyond the recorded staffing assignment.\"],\"request\":\"Determine the booking’s operational-complexity level for routing and staffing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment."}, {"path": ["evidence", "1"], "text": "The attendance estimate filed under booking reference RCC-1847 records 117 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment.", "negative_left": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment.", "negative_right": "The attendance estimate filed under booking reference RCC-1847 records 84 attendees.", "right": "The attendance estimate filed under booking reference RCC-1847 records 117 attendees."}, "verifier_independent_model": false}, "family": "scale-diverse-162-004", "id": "scale-diverse-162-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is assigned the unique booking reference RCC-1847, involves no specialized spaces or kitchen access, and has exactly one additional staff assignment.", "The attendance estimate filed under booking reference RCC-1847 records 84 attendees.", "The signed reservation form and facilities review are complete. The event will use the ordinary activity room without supervised audiovisual or sports equipment, and the schedule requires only the routine closing check beyond the recorded staffing assignment."], "request": "Determine the booking’s operational-complexity level for routing and staffing."}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the unchanged questions object and retain the original state's governing policy and request. The booking entity, facility, event date, time, and decision scope remain fixed; only the attendance estimate changes from 117 to 94. The two focus-evidence spans are complete factual sentences, and neither context contains a conflicting duplicate attendance figure. The policy language mentions levels only as governing rubric content and does not disclose the correct level for either case or instruct the classifier what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\":\"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\",\"evidence\":[\"At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m.\",\"At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 117 attendees.\",\"At 10:05 a.m., the facilities review recorded that the ordinary activity room was the only space requested and that no specialized spaces were involved.\",\"At 10:18 a.m., the organizer confirmed that kitchen access was not requested and that no cooking or food service would occur.\",\"At 10:41 a.m., the staffing review recorded no additional attendant or custodial assignments; regular on-duty personnel would handle the routine opening and closing checks.\"],\"request\":\"Determine the booking’s operational-complexity level for routing and staffing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m."}, {"path": ["evidence", "1"], "text": "At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 117 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m.", "negative_left": "At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m.", "negative_right": "At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 94 attendees.", "right": "At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 117 attendees."}, "verifier_independent_model": false}, "family": "scale-diverse-162-005", "id": "scale-diverse-162-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m.", "At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 117 attendees.", "At 10:05 a.m., the facilities review recorded that the ordinary activity room was the only space requested and that no specialized spaces were involved.", "At 10:18 a.m., the organizer confirmed that kitchen access was not requested and that no cooking or food service would occur.", "At 10:41 a.m., the staffing review recorded no additional attendant or custodial assignments; regular on-duty personnel would handle the routine opening and closing checks."], "request": "Determine the booking’s operational-complexity level for routing and staffing."}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the unchanged questions object and retain the original state's governing policy and request. The booking entity, facility, event date, time, and decision scope remain fixed; only the attendance estimate changes from 117 to 94. The two focus-evidence spans are complete factual sentences, and neither context contains a conflicting duplicate attendance figure. The policy language mentions levels only as governing rubric content and does not disclose the correct level for either case or instruct the classifier what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\":\"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\",\"evidence\":[\"At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m.\",\"At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 117 attendees.\",\"At 10:05 a.m., the facilities review recorded that the ordinary activity room was the only space requested and that no specialized spaces were involved.\",\"At 10:18 a.m., the organizer confirmed that kitchen access was not requested and that no cooking or food service would occur.\",\"At 10:41 a.m., the staffing review recorded no additional attendant or custodial assignments; regular on-duty personnel would handle the routine opening and closing checks.\"],\"request\":\"Determine the booking’s operational-complexity level for routing and staffing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m."}, {"path": ["evidence", "1"], "text": "At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 117 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m.", "negative_left": "At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m.", "negative_right": "At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 94 attendees.", "right": "At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 117 attendees."}, "verifier_independent_model": false}, "family": "scale-diverse-162-005", "id": "scale-diverse-162-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["At 9:12 a.m. on September 3, 2026, the reservation system assigned booking reference RCC-5817 to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m.", "At 9:26 a.m. on September 3, 2026, the attendance estimate filed under booking reference RCC-5817 listed 94 attendees.", "At 10:05 a.m., the facilities review recorded that the ordinary activity room was the only space requested and that no specialized spaces were involved.", "At 10:18 a.m., the organizer confirmed that kitchen access was not requested and that no cooking or food service would occur.", "At 10:41 a.m., the staffing review recorded no additional attendant or custodial assignments; regular on-duty personnel would handle the routine opening and closing checks."], "request": "Determine the booking’s operational-complexity level for routing and staffing."}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same request, catalog-entry scope, date, document, reference number, and explicit-custodian routing framework without adding exceptions or defaults. The two evidence spans are complete factual sentences. The counterfactual consistently changes the responsible-unit text from “Aquatics Compliance Office” to “City Archives,” and no unchanged statement contradicts that change. Neither context contains an answer code, output instruction, rule table, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation, and a3 is a factual text-comparison relation rather than a policy conclusion. The base and counter assignments are jointly realizable by changing only the custodian text from a nonlisted responsible unit to “City Archives”; the remaining inequality atoms stay supported. Empty policy_evidence is correct because all governing routing rules are already retained in original_input.questions, and no substantive state-origin policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that the explicit custodian text does not differ from “City Archives,” i.e. it equals that name. With an explicit custodian present, the unchanged question requires the city_archives route.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish an explicit custodian naming a responsible unit while excluding exact matches to all four listed custodians. Therefore the custodian is a different unit and none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The scanned catalog entry under evaluation describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614."}, {"id": "a2", "statement": "The scanned catalog entry under evaluation has an explicit custodian field."}, {"id": "a3", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “City Archives”."}, {"id": "a4", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Municipal Clerk Service Desk”."}, {"id": "a5", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Facilities Scheduling Unit”."}, {"id": "a6", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Building Standards Records”."}, {"id": "a7", "statement": "The explicit custodian field of the scanned catalog entry under evaluation names a responsible unit."}], "base_state_json": "\"During evidence reconciliation, staff matched one scanned catalog entry to the request. For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426. The unit-name box of reconciliation card RC-7426 contains exactly the text “Aquatics Compliance Office”. A separate character-by-character check records that the text in that card box differs from each of the exact office names “Municipal Clerk Service Desk”, “Facilities Scheduling Unit”, and “Building Standards Records”. The reconciliation reviewer confirms that the box identifies a responsible municipal unit and that no abbreviation or alias was used. The scan is legible, and the custodian field is complete rather than truncated.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426."}, {"path": [], "text": "The unit-name box of reconciliation card RC-7426 contains exactly the text “Aquatics Compliance Office”."}], "policy_evidence": [], "rules": [{"justification": "Refuting the exact-text inequality establishes that the explicit custodian text equals “City Archives”; the policy therefore requires routing to City Archives.", "target": "city_archives", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The matching entry has an explicit custodian naming a responsible unit whose text differs from every listed custodian, so the policy requires none_of_above.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426.", "negative_left": "For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426.", "negative_right": "The unit-name box of reconciliation card RC-7426 contains exactly the text “City Archives”.", "right": "The unit-name box of reconciliation card RC-7426 contains exactly the text “Aquatics Compliance Office”."}, "verifier_independent_model": false}, "family": "scale-diverse-176-002", "id": "scale-diverse-176-002-base", "input": {"questions": {"decision": {"criteria": {"building_standards_records": "Route to Building Standards Records because the catalog explicitly names that office as custodian.", "city_archives": "Route to City Archives because the catalog explicitly names City Archives as custodian.", "clerk_service_desk": "Route to the Municipal Clerk Service Desk because the catalog explicitly names that desk as custodian.", "facilities_scheduling_unit": "Route to the Facilities Scheduling Unit because the catalog explicitly names that unit as custodian.", "none_of_above": "None of the listed routes fits because the catalog explicitly names a different responsible unit."}, "instructions": "Choose the correct routing result. Use the explicit custodian field when it is present; it overrides routing inferred from the document title. Select Clerk Service Desk only for entries naming that desk, Facilities Scheduling Unit only for entries naming that unit, Building Standards Records only for entries naming that office, and City Archives only for entries naming City Archives. Select none_of_above when the explicit custodian is a different unit.", "type": "choice"}}, "state": "During evidence reconciliation, staff matched one scanned catalog entry to the request. For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426. The unit-name box of reconciliation card RC-7426 contains exactly the text “Aquatics Compliance Office”. A separate character-by-character check records that the text in that card box differs from each of the exact office names “Municipal Clerk Service Desk”, “Facilities Scheduling Unit”, and “Building Standards Records”. The reconciliation reviewer confirms that the box identifies a responsible municipal unit and that no abbreviation or alias was used. The scan is legible, and the custodian field is complete rather than truncated."}, "method": "c2d", "provenance": {"source_id": "diverse-176", "source_is_synthetic": true, "source_sha256": "ba9d0c001db4226f311d86e7a27bafe2a25c689b1898d4887900a11d298f58f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same request, catalog-entry scope, date, document, reference number, and explicit-custodian routing framework without adding exceptions or defaults. The two evidence spans are complete factual sentences. The counterfactual consistently changes the responsible-unit text from “Aquatics Compliance Office” to “City Archives,” and no unchanged statement contradicts that change. Neither context contains an answer code, output instruction, rule table, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation, and a3 is a factual text-comparison relation rather than a policy conclusion. The base and counter assignments are jointly realizable by changing only the custodian text from a nonlisted responsible unit to “City Archives”; the remaining inequality atoms stay supported. Empty policy_evidence is correct because all governing routing rules are already retained in original_input.questions, and no substantive state-origin policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that the explicit custodian text does not differ from “City Archives,” i.e. it equals that name. With an explicit custodian present, the unchanged question requires the city_archives route.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish an explicit custodian naming a responsible unit while excluding exact matches to all four listed custodians. Therefore the custodian is a different unit and none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The scanned catalog entry under evaluation describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614."}, {"id": "a2", "statement": "The scanned catalog entry under evaluation has an explicit custodian field."}, {"id": "a3", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “City Archives”."}, {"id": "a4", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Municipal Clerk Service Desk”."}, {"id": "a5", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Facilities Scheduling Unit”."}, {"id": "a6", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Building Standards Records”."}, {"id": "a7", "statement": "The explicit custodian field of the scanned catalog entry under evaluation names a responsible unit."}], "base_state_json": "\"During evidence reconciliation, staff matched one scanned catalog entry to the request. For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426. The unit-name box of reconciliation card RC-7426 contains exactly the text “Aquatics Compliance Office”. A separate character-by-character check records that the text in that card box differs from each of the exact office names “Municipal Clerk Service Desk”, “Facilities Scheduling Unit”, and “Building Standards Records”. The reconciliation reviewer confirms that the box identifies a responsible municipal unit and that no abbreviation or alias was used. The scan is legible, and the custodian field is complete rather than truncated.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426."}, {"path": [], "text": "The unit-name box of reconciliation card RC-7426 contains exactly the text “Aquatics Compliance Office”."}], "policy_evidence": [], "rules": [{"justification": "Refuting the exact-text inequality establishes that the explicit custodian text equals “City Archives”; the policy therefore requires routing to City Archives.", "target": "city_archives", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The matching entry has an explicit custodian naming a responsible unit whose text differs from every listed custodian, so the policy requires none_of_above.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426.", "negative_left": "For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426.", "negative_right": "The unit-name box of reconciliation card RC-7426 contains exactly the text “City Archives”.", "right": "The unit-name box of reconciliation card RC-7426 contains exactly the text “Aquatics Compliance Office”."}, "verifier_independent_model": false}, "family": "scale-diverse-176-002", "id": "scale-diverse-176-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"building_standards_records": "Route to Building Standards Records because the catalog explicitly names that office as custodian.", "city_archives": "Route to City Archives because the catalog explicitly names City Archives as custodian.", "clerk_service_desk": "Route to the Municipal Clerk Service Desk because the catalog explicitly names that desk as custodian.", "facilities_scheduling_unit": "Route to the Facilities Scheduling Unit because the catalog explicitly names that unit as custodian.", "none_of_above": "None of the listed routes fits because the catalog explicitly names a different responsible unit."}, "instructions": "Choose the correct routing result. Use the explicit custodian field when it is present; it overrides routing inferred from the document title. Select Clerk Service Desk only for entries naming that desk, Facilities Scheduling Unit only for entries naming that unit, Building Standards Records only for entries naming that office, and City Archives only for entries naming City Archives. Select none_of_above when the explicit custodian is a different unit.", "type": "choice"}}, "state": "During evidence reconciliation, staff matched one scanned catalog entry to the request. For the scanned catalog entry under evaluation, which describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the complete text in its explicit custodian field is identical character for character to the responsible-unit name transcribed in the unit-name box of reconciliation card RC-7426. The unit-name box of reconciliation card RC-7426 contains exactly the text “City Archives”. A separate character-by-character check records that the text in that card box differs from each of the exact office names “Municipal Clerk Service Desk”, “Facilities Scheduling Unit”, and “Building Standards Records”. The reconciliation reviewer confirms that the box identifies a responsible municipal unit and that no abbreviation or alias was used. The scan is legible, and the custodian field is complete rather than truncated."}, "method": "c2d", "provenance": {"source_id": "diverse-176", "source_is_synthetic": true, "source_sha256": "ba9d0c001db4226f311d86e7a27bafe2a25c689b1898d4887900a11d298f58f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "city_archives"}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all routing rules, while neither generated context alters or omits any additional governing policy from the original state. Both contexts retain the same request scope, document identity, date, reference, custodian-field path, and relevant time binding. The two evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual consistently changes the ticket destination—and therefore the reproduced custodian text—from “Public Recreation Records Unit” to “City Archives” without conflicting with the unchanged exclusions or other facts. Neither context contains a gold label, answer code, rule table, proposition identifier, output instruction, or explicit label rationale; the office names are permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation, and a3 is a factual text-comparison relation rather than a policy conclusion. The base and counter assignments are jointly realizable by changing only the custodian text from a nonlisted responsible unit to “City Archives”; the remaining inequality atoms stay supported. Empty policy_evidence is correct because all governing routing rules are already retained in original_input.questions, and no substantive state-origin policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that the explicit custodian text does not differ from “City Archives,” i.e. it equals that name. With an explicit custodian present, the unchanged question requires the city_archives route.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish an explicit custodian naming a responsible unit while excluding exact matches to all four listed custodians. Therefore the custodian is a different unit and none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The scanned catalog entry under evaluation describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614."}, {"id": "a2", "statement": "The scanned catalog entry under evaluation has an explicit custodian field."}, {"id": "a3", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “City Archives”."}, {"id": "a4", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Municipal Clerk Service Desk”."}, {"id": "a5", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Facilities Scheduling Unit”."}, {"id": "a6", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Building Standards Records”."}, {"id": "a7", "statement": "The explicit custodian field of the scanned catalog entry under evaluation names a responsible unit."}], "base_state_json": "\"During records intake, staff isolated the scanned catalog entry responsive to the pending request. The scan is legible and contains a populated, labeled custodian field. For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729. The destination line on operational handoff ticket OH-4729 reads exactly “Public Recreation Records Unit”. The catalog field legend defines a populated custodian field as naming the unit responsible for the entry. A controlled-name comparison separately records that the field text is not exactly “Municipal Clerk Service Desk,” not exactly “Facilities Scheduling Unit,” and not exactly “Building Standards Records.” Intake personnel confirmed that the request concerns only this inspection summary, and quality control found no competing catalog entry, bundled material, abbreviation, or unreadable text affecting the handoff record.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729."}, {"path": [], "text": "The destination line on operational handoff ticket OH-4729 reads exactly “Public Recreation Records Unit”."}], "policy_evidence": [], "rules": [{"justification": "Refuting the exact-text inequality establishes that the explicit custodian text equals “City Archives”; the policy therefore requires routing to City Archives.", "target": "city_archives", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The matching entry has an explicit custodian naming a responsible unit whose text differs from every listed custodian, so the policy requires none_of_above.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729.", "negative_left": "For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729.", "negative_right": "The destination line on operational handoff ticket OH-4729 reads exactly “City Archives”.", "right": "The destination line on operational handoff ticket OH-4729 reads exactly “Public Recreation Records Unit”."}, "verifier_independent_model": false}, "family": "scale-diverse-176-003", "id": "scale-diverse-176-003-base", "input": {"questions": {"decision": {"criteria": {"building_standards_records": "Route to Building Standards Records because the catalog explicitly names that office as custodian.", "city_archives": "Route to City Archives because the catalog explicitly names City Archives as custodian.", "clerk_service_desk": "Route to the Municipal Clerk Service Desk because the catalog explicitly names that desk as custodian.", "facilities_scheduling_unit": "Route to the Facilities Scheduling Unit because the catalog explicitly names that unit as custodian.", "none_of_above": "None of the listed routes fits because the catalog explicitly names a different responsible unit."}, "instructions": "Choose the correct routing result. Use the explicit custodian field when it is present; it overrides routing inferred from the document title. Select Clerk Service Desk only for entries naming that desk, Facilities Scheduling Unit only for entries naming that unit, Building Standards Records only for entries naming that office, and City Archives only for entries naming City Archives. Select none_of_above when the explicit custodian is a different unit.", "type": "choice"}}, "state": "During records intake, staff isolated the scanned catalog entry responsive to the pending request. The scan is legible and contains a populated, labeled custodian field. For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729. The destination line on operational handoff ticket OH-4729 reads exactly “Public Recreation Records Unit”. The catalog field legend defines a populated custodian field as naming the unit responsible for the entry. A controlled-name comparison separately records that the field text is not exactly “Municipal Clerk Service Desk,” not exactly “Facilities Scheduling Unit,” and not exactly “Building Standards Records.” Intake personnel confirmed that the request concerns only this inspection summary, and quality control found no competing catalog entry, bundled material, abbreviation, or unreadable text affecting the handoff record."}, "method": "c2d", "provenance": {"source_id": "diverse-176", "source_is_synthetic": true, "source_sha256": "ba9d0c001db4226f311d86e7a27bafe2a25c689b1898d4887900a11d298f58f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all routing rules, while neither generated context alters or omits any additional governing policy from the original state. Both contexts retain the same request scope, document identity, date, reference, custodian-field path, and relevant time binding. The two evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual consistently changes the ticket destination—and therefore the reproduced custodian text—from “Public Recreation Records Unit” to “City Archives” without conflicting with the unchanged exclusions or other facts. Neither context contains a gold label, answer code, rule table, proposition identifier, output instruction, or explicit label rationale; the office names are permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation, and a3 is a factual text-comparison relation rather than a policy conclusion. The base and counter assignments are jointly realizable by changing only the custodian text from a nonlisted responsible unit to “City Archives”; the remaining inequality atoms stay supported. Empty policy_evidence is correct because all governing routing rules are already retained in original_input.questions, and no substantive state-origin policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that the explicit custodian text does not differ from “City Archives,” i.e. it equals that name. With an explicit custodian present, the unchanged question requires the city_archives route.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish an explicit custodian naming a responsible unit while excluding exact matches to all four listed custodians. Therefore the custodian is a different unit and none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The scanned catalog entry under evaluation describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614."}, {"id": "a2", "statement": "The scanned catalog entry under evaluation has an explicit custodian field."}, {"id": "a3", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “City Archives”."}, {"id": "a4", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Municipal Clerk Service Desk”."}, {"id": "a5", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Facilities Scheduling Unit”."}, {"id": "a6", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Building Standards Records”."}, {"id": "a7", "statement": "The explicit custodian field of the scanned catalog entry under evaluation names a responsible unit."}], "base_state_json": "\"During records intake, staff isolated the scanned catalog entry responsive to the pending request. The scan is legible and contains a populated, labeled custodian field. For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729. The destination line on operational handoff ticket OH-4729 reads exactly “Public Recreation Records Unit”. The catalog field legend defines a populated custodian field as naming the unit responsible for the entry. A controlled-name comparison separately records that the field text is not exactly “Municipal Clerk Service Desk,” not exactly “Facilities Scheduling Unit,” and not exactly “Building Standards Records.” Intake personnel confirmed that the request concerns only this inspection summary, and quality control found no competing catalog entry, bundled material, abbreviation, or unreadable text affecting the handoff record.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729."}, {"path": [], "text": "The destination line on operational handoff ticket OH-4729 reads exactly “Public Recreation Records Unit”."}], "policy_evidence": [], "rules": [{"justification": "Refuting the exact-text inequality establishes that the explicit custodian text equals “City Archives”; the policy therefore requires routing to City Archives.", "target": "city_archives", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The matching entry has an explicit custodian naming a responsible unit whose text differs from every listed custodian, so the policy requires none_of_above.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729.", "negative_left": "For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729.", "negative_right": "The destination line on operational handoff ticket OH-4729 reads exactly “City Archives”.", "right": "The destination line on operational handoff ticket OH-4729 reads exactly “Public Recreation Records Unit”."}, "verifier_independent_model": false}, "family": "scale-diverse-176-003", "id": "scale-diverse-176-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"building_standards_records": "Route to Building Standards Records because the catalog explicitly names that office as custodian.", "city_archives": "Route to City Archives because the catalog explicitly names City Archives as custodian.", "clerk_service_desk": "Route to the Municipal Clerk Service Desk because the catalog explicitly names that desk as custodian.", "facilities_scheduling_unit": "Route to the Facilities Scheduling Unit because the catalog explicitly names that unit as custodian.", "none_of_above": "None of the listed routes fits because the catalog explicitly names a different responsible unit."}, "instructions": "Choose the correct routing result. Use the explicit custodian field when it is present; it overrides routing inferred from the document title. Select Clerk Service Desk only for entries naming that desk, Facilities Scheduling Unit only for entries naming that unit, Building Standards Records only for entries naming that office, and City Archives only for entries naming City Archives. Select none_of_above when the explicit custodian is a different unit.", "type": "choice"}}, "state": "During records intake, staff isolated the scanned catalog entry responsive to the pending request. The scan is legible and contains a populated, labeled custodian field. For the scanned catalog entry under evaluation describing the June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, the explicit custodian field exactly reproduces the destination line on operational handoff ticket OH-4729. The destination line on operational handoff ticket OH-4729 reads exactly “City Archives”. The catalog field legend defines a populated custodian field as naming the unit responsible for the entry. A controlled-name comparison separately records that the field text is not exactly “Municipal Clerk Service Desk,” not exactly “Facilities Scheduling Unit,” and not exactly “Building Standards Records.” Intake personnel confirmed that the request concerns only this inspection summary, and quality control found no competing catalog entry, bundled material, abbreviation, or unreadable text affecting the handoff record."}, "method": "c2d", "provenance": {"source_id": "diverse-176", "source_is_synthetic": true, "source_sha256": "ba9d0c001db4226f311d86e7a27bafe2a25c689b1898d4887900a11d298f58f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "city_archives"}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all routing rules, while both contexts retain the same requester, document, date, reference, entity, and explicit-custodian decision path. The two evidence spans are complete factual sentences. The counterfactual changes only the verification-card unit to City Archives; because the catalog field is stated to match that card and the unchanged non-match assertions concern three other units, the resulting context is coherent. Neither context includes an answer code, classifier instruction, rule table, proposition identifier, or label rationale; the custodian names are permissible natural policy terminology and case evidence.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation, and a3 is a factual text-comparison relation rather than a policy conclusion. The base and counter assignments are jointly realizable by changing only the custodian text from a nonlisted responsible unit to “City Archives”; the remaining inequality atoms stay supported. Empty policy_evidence is correct because all governing routing rules are already retained in original_input.questions, and no substantive state-origin policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that the explicit custodian text does not differ from “City Archives,” i.e. it equals that name. With an explicit custodian present, the unchanged question requires the city_archives route.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish an explicit custodian naming a responsible unit while excluding exact matches to all four listed custodians. Therefore the custodian is a different unit and none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The scanned catalog entry under evaluation describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614."}, {"id": "a2", "statement": "The scanned catalog entry under evaluation has an explicit custodian field."}, {"id": "a3", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “City Archives”."}, {"id": "a4", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Municipal Clerk Service Desk”."}, {"id": "a5", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Facilities Scheduling Unit”."}, {"id": "a6", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Building Standards Records”."}, {"id": "a7", "statement": "The explicit custodian field of the scanned catalog entry under evaluation names a responsible unit."}], "base_state_json": "\"Field note: The requester identified the June 14, 2018 Riverside Splash Pad inspection summary, reference PSP-2018-0614. Staff checked the scanned catalog record rather than inferring custody from the document title. The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827. Verification card VC-4827 displays the exact responsible-unit name “Aquatic Sites Compliance Office”. A separate comparison check records that the field text does not match “Municipal Clerk Service Desk,” “Facilities Scheduling Unit,” or “Building Standards Records.” The card is current, legible, and designated for confirming the responsible unit recorded in the catalog field.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827."}, {"path": [], "text": "Verification card VC-4827 displays the exact responsible-unit name “Aquatic Sites Compliance Office”."}], "policy_evidence": [], "rules": [{"justification": "Refuting the exact-text inequality establishes that the explicit custodian text equals “City Archives”; the policy therefore requires routing to City Archives.", "target": "city_archives", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The matching entry has an explicit custodian naming a responsible unit whose text differs from every listed custodian, so the policy requires none_of_above.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827.", "negative_left": "The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827.", "negative_right": "Verification card VC-4827 displays the exact responsible-unit name “City Archives”.", "right": "Verification card VC-4827 displays the exact responsible-unit name “Aquatic Sites Compliance Office”."}, "verifier_independent_model": false}, "family": "scale-diverse-176-004", "id": "scale-diverse-176-004-base", "input": {"questions": {"decision": {"criteria": {"building_standards_records": "Route to Building Standards Records because the catalog explicitly names that office as custodian.", "city_archives": "Route to City Archives because the catalog explicitly names City Archives as custodian.", "clerk_service_desk": "Route to the Municipal Clerk Service Desk because the catalog explicitly names that desk as custodian.", "facilities_scheduling_unit": "Route to the Facilities Scheduling Unit because the catalog explicitly names that unit as custodian.", "none_of_above": "None of the listed routes fits because the catalog explicitly names a different responsible unit."}, "instructions": "Choose the correct routing result. Use the explicit custodian field when it is present; it overrides routing inferred from the document title. Select Clerk Service Desk only for entries naming that desk, Facilities Scheduling Unit only for entries naming that unit, Building Standards Records only for entries naming that office, and City Archives only for entries naming City Archives. Select none_of_above when the explicit custodian is a different unit.", "type": "choice"}}, "state": "Field note: The requester identified the June 14, 2018 Riverside Splash Pad inspection summary, reference PSP-2018-0614. Staff checked the scanned catalog record rather than inferring custody from the document title. The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827. Verification card VC-4827 displays the exact responsible-unit name “Aquatic Sites Compliance Office”. A separate comparison check records that the field text does not match “Municipal Clerk Service Desk,” “Facilities Scheduling Unit,” or “Building Standards Records.” The card is current, legible, and designated for confirming the responsible unit recorded in the catalog field."}, "method": "c2d", "provenance": {"source_id": "diverse-176", "source_is_synthetic": true, "source_sha256": "ba9d0c001db4226f311d86e7a27bafe2a25c689b1898d4887900a11d298f58f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all routing rules, while both contexts retain the same requester, document, date, reference, entity, and explicit-custodian decision path. The two evidence spans are complete factual sentences. The counterfactual changes only the verification-card unit to City Archives; because the catalog field is stated to match that card and the unchanged non-match assertions concern three other units, the resulting context is coherent. Neither context includes an answer code, classifier instruction, rule table, proposition identifier, or label rationale; the custodian names are permissible natural policy terminology and case evidence.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation, and a3 is a factual text-comparison relation rather than a policy conclusion. The base and counter assignments are jointly realizable by changing only the custodian text from a nonlisted responsible unit to “City Archives”; the remaining inequality atoms stay supported. Empty policy_evidence is correct because all governing routing rules are already retained in original_input.questions, and no substantive state-origin policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that the explicit custodian text does not differ from “City Archives,” i.e. it equals that name. With an explicit custodian present, the unchanged question requires the city_archives route.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish an explicit custodian naming a responsible unit while excluding exact matches to all four listed custodians. Therefore the custodian is a different unit and none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The scanned catalog entry under evaluation describes the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614."}, {"id": "a2", "statement": "The scanned catalog entry under evaluation has an explicit custodian field."}, {"id": "a3", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “City Archives”."}, {"id": "a4", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Municipal Clerk Service Desk”."}, {"id": "a5", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Facilities Scheduling Unit”."}, {"id": "a6", "statement": "The exact text in the explicit custodian field of the scanned catalog entry under evaluation differs from the exact office name “Building Standards Records”."}, {"id": "a7", "statement": "The explicit custodian field of the scanned catalog entry under evaluation names a responsible unit."}], "base_state_json": "\"Field note: The requester identified the June 14, 2018 Riverside Splash Pad inspection summary, reference PSP-2018-0614. Staff checked the scanned catalog record rather than inferring custody from the document title. The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827. Verification card VC-4827 displays the exact responsible-unit name “Aquatic Sites Compliance Office”. A separate comparison check records that the field text does not match “Municipal Clerk Service Desk,” “Facilities Scheduling Unit,” or “Building Standards Records.” The card is current, legible, and designated for confirming the responsible unit recorded in the catalog field.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827."}, {"path": [], "text": "Verification card VC-4827 displays the exact responsible-unit name “Aquatic Sites Compliance Office”."}], "policy_evidence": [], "rules": [{"justification": "Refuting the exact-text inequality establishes that the explicit custodian text equals “City Archives”; the policy therefore requires routing to City Archives.", "target": "city_archives", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The matching entry has an explicit custodian naming a responsible unit whose text differs from every listed custodian, so the policy requires none_of_above.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827.", "negative_left": "The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827.", "negative_right": "Verification card VC-4827 displays the exact responsible-unit name “City Archives”.", "right": "Verification card VC-4827 displays the exact responsible-unit name “Aquatic Sites Compliance Office”."}, "verifier_independent_model": false}, "family": "scale-diverse-176-004", "id": "scale-diverse-176-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"building_standards_records": "Route to Building Standards Records because the catalog explicitly names that office as custodian.", "city_archives": "Route to City Archives because the catalog explicitly names City Archives as custodian.", "clerk_service_desk": "Route to the Municipal Clerk Service Desk because the catalog explicitly names that desk as custodian.", "facilities_scheduling_unit": "Route to the Facilities Scheduling Unit because the catalog explicitly names that unit as custodian.", "none_of_above": "None of the listed routes fits because the catalog explicitly names a different responsible unit."}, "instructions": "Choose the correct routing result. Use the explicit custodian field when it is present; it overrides routing inferred from the document title. Select Clerk Service Desk only for entries naming that desk, Facilities Scheduling Unit only for entries naming that unit, Building Standards Records only for entries naming that office, and City Archives only for entries naming City Archives. Select none_of_above when the explicit custodian is a different unit.", "type": "choice"}}, "state": "Field note: The requester identified the June 14, 2018 Riverside Splash Pad inspection summary, reference PSP-2018-0614. Staff checked the scanned catalog record rather than inferring custody from the document title. The scanned catalog entry under evaluation, describing the requested June 14, 2018 Riverside Splash Pad inspection summary with reference PSP-2018-0614, has an explicit custodian field whose exact text matches the responsible-unit name on verification card VC-4827. Verification card VC-4827 displays the exact responsible-unit name “City Archives”. A separate comparison check records that the field text does not match “Municipal Clerk Service Desk,” “Facilities Scheduling Unit,” or “Building Standards Records.” The card is current, legible, and designated for confirming the responsible unit recorded in the catalog field."}, "method": "c2d", "provenance": {"source_id": "diverse-176", "source_is_synthetic": true, "source_sha256": "ba9d0c001db4226f311d86e7a27bafe2a25c689b1898d4887900a11d298f58f9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "city_archives"}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts reproduce the governing routing and status rubric without adding exceptions, priorities, or missing-evidence defaults. The request continues to concern whether the identified council-packet request qualifies as Ready—standard at the City Clerk Records Unit, and the focus paths identify two complete factual sentences. The counterfactual changes only P2’s cataloged meeting date from 2024-06-11 to 2024-06-18; this is coherent with P1 matching request R’s 2024-06-11 date, the stated complete two-packet archive set, and the unchanged one-box retrieval fact. Neither constructed context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or output directive.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\":\"For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\",\"evidence\":[\"Request R was submitted to the City Clerk Records Unit for a council meeting packet. The submitted form includes no packet number.\",\"The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11.\",\"During the archive review, staff confirmed that, as of 2026-09-17, the unit's complete set of archived council meeting packets was exactly {packet P1, packet P2}. The catalog entry for P1 marks its meeting date as matching the date stated in request R.\",\"As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-11.\",\"The retrieval log records that locating packets P1 and P2 spans one archive box.\"],\"request\":\"Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "1"], "text": "The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11."}, {"path": ["evidence", "3"], "text": "As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-11."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11.", "negative_left": "The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11.", "negative_right": "As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-18.", "right": "As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-11."}, "verifier_independent_model": false}, "family": "scale-diverse-177-001", "id": "scale-diverse-177-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.", "evidence": ["Request R was submitted to the City Clerk Records Unit for a council meeting packet. The submitted form includes no packet number.", "The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11.", "During the archive review, staff confirmed that, as of 2026-09-17, the unit's complete set of archived council meeting packets was exactly {packet P1, packet P2}. The catalog entry for P1 marks its meeting date as matching the date stated in request R.", "As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-11.", "The retrieval log records that locating packets P1 and P2 spans one archive box."], "request": "Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts reproduce the governing routing and status rubric without adding exceptions, priorities, or missing-evidence defaults. The request continues to concern whether the identified council-packet request qualifies as Ready—standard at the City Clerk Records Unit, and the focus paths identify two complete factual sentences. The counterfactual changes only P2’s cataloged meeting date from 2024-06-11 to 2024-06-18; this is coherent with P1 matching request R’s 2024-06-11 date, the stated complete two-packet archive set, and the unchanged one-box retrieval fact. Neither constructed context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or output directive.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\":\"For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\",\"evidence\":[\"Request R was submitted to the City Clerk Records Unit for a council meeting packet. The submitted form includes no packet number.\",\"The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11.\",\"During the archive review, staff confirmed that, as of 2026-09-17, the unit's complete set of archived council meeting packets was exactly {packet P1, packet P2}. The catalog entry for P1 marks its meeting date as matching the date stated in request R.\",\"As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-11.\",\"The retrieval log records that locating packets P1 and P2 spans one archive box.\"],\"request\":\"Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "1"], "text": "The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11."}, {"path": ["evidence", "3"], "text": "As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-11."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11.", "negative_left": "The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11.", "negative_right": "As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-18.", "right": "As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-11."}, "verifier_independent_model": false}, "family": "scale-diverse-177-001", "id": "scale-diverse-177-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.", "evidence": ["Request R was submitted to the City Clerk Records Unit for a council meeting packet. The submitted form includes no packet number.", "The City Clerk Records Unit's timestamped copy of request R records the requested meeting date as 2024-06-11.", "During the archive review, staff confirmed that, as of 2026-09-17, the unit's complete set of archived council meeting packets was exactly {packet P1, packet P2}. The catalog entry for P1 marks its meeting date as matching the date stated in request R.", "As of 2026-09-17, the City Clerk Records Unit's catalog entry for packet P2 records its meeting date as 2024-06-18.", "The retrieval log records that locating packets P1 and P2 spans one archive box."], "request": "Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing rubric verbatim and retain the same request scope, unit, entity, and classification question. The two focus-evidence spans are complete factual sentences. The counterfactual changes P2’s cataloged date from 2024-06-11 to 2024-06-18 without conflicting with the unchanged inventory, P1 date-match assertion, or archive-box count. Neither constructed context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or output directive; the status terms are permissible policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\":\"For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\",\"evidence\":[\"Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date.\",\"The intake form identifies the requested material type as a council meeting packet and contains no packet number.\",\"An inventory snapshot dated 2026-09-17 records the unit’s complete archived council meeting packet holdings as exactly P1 and P2.\",\"The catalog entry for P1 states that its meeting date matches the date entered in request R.\",\"As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-11 as the meeting date for archived council meeting packet P2.\",\"The archive pull log shows that retrieving P1 and P2 spans one archive box.\"],\"request\":\"Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "0"], "text": "Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date."}, {"path": ["evidence", "4"], "text": "As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-11 as the meeting date for archived council meeting packet P2."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date.", "negative_left": "Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date.", "negative_right": "As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-18 as the meeting date for archived council meeting packet P2.", "right": "As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-11 as the meeting date for archived council meeting packet P2."}, "verifier_independent_model": false}, "family": "scale-diverse-177-002", "id": "scale-diverse-177-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.", "evidence": ["Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date.", "The intake form identifies the requested material type as a council meeting packet and contains no packet number.", "An inventory snapshot dated 2026-09-17 records the unit’s complete archived council meeting packet holdings as exactly P1 and P2.", "The catalog entry for P1 states that its meeting date matches the date entered in request R.", "As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-11 as the meeting date for archived council meeting packet P2.", "The archive pull log shows that retrieving P1 and P2 spans one archive box."], "request": "Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing rubric verbatim and retain the same request scope, unit, entity, and classification question. The two focus-evidence spans are complete factual sentences. The counterfactual changes P2’s cataloged date from 2024-06-11 to 2024-06-18 without conflicting with the unchanged inventory, P1 date-match assertion, or archive-box count. Neither constructed context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or output directive; the status terms are permissible policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\":\"For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\",\"evidence\":[\"Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date.\",\"The intake form identifies the requested material type as a council meeting packet and contains no packet number.\",\"An inventory snapshot dated 2026-09-17 records the unit’s complete archived council meeting packet holdings as exactly P1 and P2.\",\"The catalog entry for P1 states that its meeting date matches the date entered in request R.\",\"As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-11 as the meeting date for archived council meeting packet P2.\",\"The archive pull log shows that retrieving P1 and P2 spans one archive box.\"],\"request\":\"Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "0"], "text": "Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date."}, {"path": ["evidence", "4"], "text": "As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-11 as the meeting date for archived council meeting packet P2."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date.", "negative_left": "Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date.", "negative_right": "As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-18 as the meeting date for archived council meeting packet P2.", "right": "As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-11 as the meeting date for archived council meeting packet P2."}, "verifier_independent_model": false}, "family": "scale-diverse-177-002", "id": "scale-diverse-177-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.", "evidence": ["Request R, submitted to the City Clerk Records Unit, specifies 2024-06-11 as the meeting date.", "The intake form identifies the requested material type as a council meeting packet and contains no packet number.", "An inventory snapshot dated 2026-09-17 records the unit’s complete archived council meeting packet holdings as exactly P1 and P2.", "The catalog entry for P1 states that its meeting date matches the date entered in request R.", "As of 2026-09-17, the City Clerk Records Unit catalog lists 2024-06-18 as the meeting date for archived council meeting packet P2.", "The archive pull log shows that retrieving P1 and P2 spans one archive box."], "request": "Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing rubric and request scope, while the unchanged questions preserve the decision criteria and instructions. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only P2’s cataloged date from 2024-06-11 to 2024-06-12; this does not conflict with P1’s matching date, the complete two-packet inventory, or the archive-box statement. Neither context contains a gold answer, output instruction, answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\":\"For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\",\"request\":\"Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard.\",\"handoff\":{\"summary\":\"Operations staff reviewed request R against the unit’s current archive inventory and catalog before retrieval.\",\"observations\":[\"During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11.\",\"Request R was submitted to the City Clerk Records Unit and identifies the requested material type as a council meeting packet.\",\"The handoff copy of request R contains no packet number.\",\"As of the handoff date, the unit’s complete set of archived council meeting packets was exactly packet P1 and packet P2.\",\"Packet P1’s cataloged meeting date matched the meeting-date field in request R.\",\"During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-11 as its meeting date.\",\"Archive staff confirmed that retrieving packets P1 and P2 together spans one archive box.\"]}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["handoff", "observations", "0"], "text": "During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11."}, {"path": ["handoff", "observations", "5"], "text": "During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-11 as its meeting date."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11.", "negative_left": "During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11.", "negative_right": "During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-12 as its meeting date.", "right": "During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-11 as its meeting date."}, "verifier_independent_model": false}, "family": "scale-diverse-177-003", "id": "scale-diverse-177-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.", "handoff": {"observations": ["During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11.", "Request R was submitted to the City Clerk Records Unit and identifies the requested material type as a council meeting packet.", "The handoff copy of request R contains no packet number.", "As of the handoff date, the unit’s complete set of archived council meeting packets was exactly packet P1 and packet P2.", "Packet P1’s cataloged meeting date matched the meeting-date field in request R.", "During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-11 as its meeting date.", "Archive staff confirmed that retrieving packets P1 and P2 together spans one archive box."], "summary": "Operations staff reviewed request R against the unit’s current archive inventory and catalog before retrieval."}, "request": "Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing rubric and request scope, while the unchanged questions preserve the decision criteria and instructions. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only P2’s cataloged date from 2024-06-11 to 2024-06-12; this does not conflict with P1’s matching date, the complete two-packet inventory, or the archive-box statement. Neither context contains a gold answer, output instruction, answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\":\"For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\",\"request\":\"Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard.\",\"handoff\":{\"summary\":\"Operations staff reviewed request R against the unit’s current archive inventory and catalog before retrieval.\",\"observations\":[\"During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11.\",\"Request R was submitted to the City Clerk Records Unit and identifies the requested material type as a council meeting packet.\",\"The handoff copy of request R contains no packet number.\",\"As of the handoff date, the unit’s complete set of archived council meeting packets was exactly packet P1 and packet P2.\",\"Packet P1’s cataloged meeting date matched the meeting-date field in request R.\",\"During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-11 as its meeting date.\",\"Archive staff confirmed that retrieving packets P1 and P2 together spans one archive box.\"]}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["handoff", "observations", "0"], "text": "During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11."}, {"path": ["handoff", "observations", "5"], "text": "During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-11 as its meeting date."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11.", "negative_left": "During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11.", "negative_right": "During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-12 as its meeting date.", "right": "During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-11 as its meeting date."}, "verifier_independent_model": false}, "family": "scale-diverse-177-003", "id": "scale-diverse-177-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.", "handoff": {"observations": ["During the City Clerk Records Unit handoff on 2026-09-17, the meeting-date field in request R contained 2024-06-11.", "Request R was submitted to the City Clerk Records Unit and identifies the requested material type as a council meeting packet.", "The handoff copy of request R contains no packet number.", "As of the handoff date, the unit’s complete set of archived council meeting packets was exactly packet P1 and packet P2.", "Packet P1’s cataloged meeting date matched the meeting-date field in request R.", "During the City Clerk Records Unit handoff on 2026-09-17, packet P2's catalog record listed 2024-06-12 as its meeting date.", "Archive staff confirmed that retrieving packets P1 and P2 together spans one archive box."], "summary": "Operations staff reviewed request R against the unit’s current archive inventory and catalog before retrieval."}, "request": "Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing rubric and preserve the request’s entity, decision path, and relevant meeting-date binding. The two focus spans are complete factual sentences. The counterfactual changes only P2’s meeting date from 2024-06-11 to 2024-06-18, which is coherent with the unchanged evidence and removes the duplicate-date condition without creating a contradictory measurement or assertion. Neither context contains an answer, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\":\"For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\",\"evidence\":[\"Request R specifies 2024-06-11 as its meeting date.\",\"The intake record shows that Request R was submitted to the City Clerk Records Unit for a council meeting packet.\",\"The archive inventory dated 2026-09-17 identifies the unit’s complete set of archived council meeting packets as packet P1 and packet P2.\",\"Packet P1’s catalog entry identifies its meeting date as the date specified in Request R.\",\"The City Clerk Records Unit catalog lists 2024-06-11 as packet P2's meeting date.\",\"The packet-number field on Request R is blank, and no packet number appears elsewhere on the submission.\",\"Staff report that retrieving packets P1 and P2 requires access to one archive box.\"],\"request\":\"Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "0"], "text": "Request R specifies 2024-06-11 as its meeting date."}, {"path": ["evidence", "4"], "text": "The City Clerk Records Unit catalog lists 2024-06-11 as packet P2's meeting date."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Request R specifies 2024-06-11 as its meeting date.", "negative_left": "Request R specifies 2024-06-11 as its meeting date.", "negative_right": "The City Clerk Records Unit catalog lists 2024-06-18 as packet P2's meeting date.", "right": "The City Clerk Records Unit catalog lists 2024-06-11 as packet P2's meeting date."}, "verifier_independent_model": false}, "family": "scale-diverse-177-004", "id": "scale-diverse-177-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.", "evidence": ["Request R specifies 2024-06-11 as its meeting date.", "The intake record shows that Request R was submitted to the City Clerk Records Unit for a council meeting packet.", "The archive inventory dated 2026-09-17 identifies the unit’s complete set of archived council meeting packets as packet P1 and packet P2.", "Packet P1’s catalog entry identifies its meeting date as the date specified in Request R.", "The City Clerk Records Unit catalog lists 2024-06-11 as packet P2's meeting date.", "The packet-number field on Request R is blank, and no packet number appears elsewhere on the submission.", "Staff report that retrieving packets P1 and P2 requires access to one archive box."], "request": "Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing rubric and preserve the request’s entity, decision path, and relevant meeting-date binding. The two focus spans are complete factual sentences. The counterfactual changes only P2’s meeting date from 2024-06-11 to 2024-06-18, which is coherent with the unchanged evidence and removes the duplicate-date condition without creating a contradictory measurement or assertion. Neither context contains an answer, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\":\"For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\",\"evidence\":[\"Request R specifies 2024-06-11 as its meeting date.\",\"The intake record shows that Request R was submitted to the City Clerk Records Unit for a council meeting packet.\",\"The archive inventory dated 2026-09-17 identifies the unit’s complete set of archived council meeting packets as packet P1 and packet P2.\",\"Packet P1’s catalog entry identifies its meeting date as the date specified in Request R.\",\"The City Clerk Records Unit catalog lists 2024-06-11 as packet P2's meeting date.\",\"The packet-number field on Request R is blank, and no packet number appears elsewhere on the submission.\",\"Staff report that retrieving packets P1 and P2 requires access to one archive box.\"],\"request\":\"Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "0"], "text": "Request R specifies 2024-06-11 as its meeting date."}, {"path": ["evidence", "4"], "text": "The City Clerk Records Unit catalog lists 2024-06-11 as packet P2's meeting date."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Request R specifies 2024-06-11 as its meeting date.", "negative_left": "Request R specifies 2024-06-11 as its meeting date.", "negative_right": "The City Clerk Records Unit catalog lists 2024-06-18 as packet P2's meeting date.", "right": "The City Clerk Records Unit catalog lists 2024-06-11 as packet P2's meeting date."}, "verifier_independent_model": false}, "family": "scale-diverse-177-004", "id": "scale-diverse-177-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.", "evidence": ["Request R specifies 2024-06-11 as its meeting date.", "The intake record shows that Request R was submitted to the City Clerk Records Unit for a council meeting packet.", "The archive inventory dated 2026-09-17 identifies the unit’s complete set of archived council meeting packets as packet P1 and packet P2.", "Packet P1’s catalog entry identifies its meeting date as the date specified in Request R.", "The City Clerk Records Unit catalog lists 2024-06-18 as packet P2's meeting date.", "The packet-number field on Request R is blank, and no packet number appears elsewhere on the submission.", "Staff report that retrieving packets P1 and P2 requires access to one archive box."], "request": "Decide whether the request is complete enough for the City Clerk Records Unit to classify as Ready—standard."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the routing, process-readiness, and difficulty policies without adding exceptions or defaults, while retaining the same requester reference, board, meeting date, requested records, fallback, and relevant time path. The two focus-evidence spans are complete factual sentences: one reports the recorded availability code and the other reports the configuration mapping. The counterfactual coherently changes only the availability observation from code 47 (UNCONFIRMED) to code 82 (CONFIRMED), with no duplicate conflicting measurement in either context. The domain-specific codes describe an availability field rather than classifier outputs, and neither context embeds a gold answer, proposition identifier, label rationale, or instruction directing the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Intake record\",\"text\":\"On May 15, 2018, the requester filed reference PB-2018-05-14, seeking the complete Parks Advisory Board meeting packet for its May 14, 2018 meeting, a date earlier than January 1, 2022. The filing permits the archived agenda and available attachments to be supplied if the complete packet cannot be found.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Intake review\",\"text\":\"The submission identifies the board, meeting date, exact requested item, and acceptable conditional fallback; the clerk marked the records request process-ready.\"},{\"speaker\":\"Configuration report\",\"text\":\"At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field.\"},{\"speaker\":\"Audit log\",\"text\":\"At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["4", "text"], "text": "At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14."}, {"path": ["3", "text"], "text": "At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.", "negative_left": "At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 82 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.", "negative_right": "At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field.", "right": "At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}, "verifier_independent_model": false}, "family": "scale-diverse-178-001", "id": "scale-diverse-178-001-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Intake record", "text": "On May 15, 2018, the requester filed reference PB-2018-05-14, seeking the complete Parks Advisory Board meeting packet for its May 14, 2018 meeting, a date earlier than January 1, 2022. The filing permits the archived agenda and available attachments to be supplied if the complete packet cannot be found."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Intake review", "text": "The submission identifies the board, meeting date, exact requested item, and acceptable conditional fallback; the clerk marked the records request process-ready."}, {"speaker": "Configuration report", "text": "At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}, {"speaker": "Audit log", "text": "At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the routing, process-readiness, and difficulty policies without adding exceptions or defaults, while retaining the same requester reference, board, meeting date, requested records, fallback, and relevant time path. The two focus-evidence spans are complete factual sentences: one reports the recorded availability code and the other reports the configuration mapping. The counterfactual coherently changes only the availability observation from code 47 (UNCONFIRMED) to code 82 (CONFIRMED), with no duplicate conflicting measurement in either context. The domain-specific codes describe an availability field rather than classifier outputs, and neither context embeds a gold answer, proposition identifier, label rationale, or instruction directing the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Intake record\",\"text\":\"On May 15, 2018, the requester filed reference PB-2018-05-14, seeking the complete Parks Advisory Board meeting packet for its May 14, 2018 meeting, a date earlier than January 1, 2022. The filing permits the archived agenda and available attachments to be supplied if the complete packet cannot be found.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Intake review\",\"text\":\"The submission identifies the board, meeting date, exact requested item, and acceptable conditional fallback; the clerk marked the records request process-ready.\"},{\"speaker\":\"Configuration report\",\"text\":\"At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field.\"},{\"speaker\":\"Audit log\",\"text\":\"At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["4", "text"], "text": "At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14."}, {"path": ["3", "text"], "text": "At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.", "negative_left": "At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 82 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.", "negative_right": "At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field.", "right": "At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}, "verifier_independent_model": false}, "family": "scale-diverse-178-001", "id": "scale-diverse-178-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Intake record", "text": "On May 15, 2018, the requester filed reference PB-2018-05-14, seeking the complete Parks Advisory Board meeting packet for its May 14, 2018 meeting, a date earlier than January 1, 2022. The filing permits the archived agenda and available attachments to be supplied if the complete packet cannot be found."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Intake review", "text": "The submission identifies the board, meeting date, exact requested item, and acceptable conditional fallback; the clerk marked the records request process-ready."}, {"speaker": "Configuration report", "text": "At 09:10 UTC on May 16, 2018, the Archive Unit's configuration report mapped code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}, {"speaker": "Audit log", "text": "At 09:14 UTC on May 16, 2018, the Archive Unit's audit log recorded code 82 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the routing, process-readiness, and difficulty policies without adding exceptions or defaults, and they retain the same request reference, board, meeting date, requested records, fallback, and decision scope. The two focus spans are complete factual sentences. The counterfactual changes only cell H27 from UNCONFIRMED to CONFIRMED; this is coherent with the reconciliation statement and all unchanged facts, although it changes which stated difficulty category applies. Neither context includes a gold answer, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records requester\",\"text\":\"For reference PB-2018-05-14, I request the complete Parks Advisory Board packet from its May 14, 2018 meeting. If that complete packet cannot be found, you may provide the archived agenda and available attachments.\"},{\"speaker\":\"Intake clerk\",\"text\":\"After confirming the required request details and fallback permission, intake marked the records request process-ready.\"},{\"speaker\":\"Reconciliation log\",\"text\":\"The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042.\"},{\"speaker\":\"Archive snapshot\",\"text\":\"Cell H27 in snapshot AR-6042 contains the value UNCONFIRMED.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["2", "text"], "text": "The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042."}, {"path": ["3", "text"], "text": "Cell H27 in snapshot AR-6042 contains the value UNCONFIRMED."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042.", "negative_left": "The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042.", "negative_right": "Cell H27 in snapshot AR-6042 contains the value CONFIRMED.", "right": "Cell H27 in snapshot AR-6042 contains the value UNCONFIRMED."}, "verifier_independent_model": false}, "family": "scale-diverse-178-002", "id": "scale-diverse-178-002-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records requester", "text": "For reference PB-2018-05-14, I request the complete Parks Advisory Board packet from its May 14, 2018 meeting. If that complete packet cannot be found, you may provide the archived agenda and available attachments."}, {"speaker": "Intake clerk", "text": "After confirming the required request details and fallback permission, intake marked the records request process-ready."}, {"speaker": "Reconciliation log", "text": "The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042."}, {"speaker": "Archive snapshot", "text": "Cell H27 in snapshot AR-6042 contains the value UNCONFIRMED."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the routing, process-readiness, and difficulty policies without adding exceptions or defaults, and they retain the same request reference, board, meeting date, requested records, fallback, and decision scope. The two focus spans are complete factual sentences. The counterfactual changes only cell H27 from UNCONFIRMED to CONFIRMED; this is coherent with the reconciliation statement and all unchanged facts, although it changes which stated difficulty category applies. Neither context includes a gold answer, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records requester\",\"text\":\"For reference PB-2018-05-14, I request the complete Parks Advisory Board packet from its May 14, 2018 meeting. If that complete packet cannot be found, you may provide the archived agenda and available attachments.\"},{\"speaker\":\"Intake clerk\",\"text\":\"After confirming the required request details and fallback permission, intake marked the records request process-ready.\"},{\"speaker\":\"Reconciliation log\",\"text\":\"The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042.\"},{\"speaker\":\"Archive snapshot\",\"text\":\"Cell H27 in snapshot AR-6042 contains the value UNCONFIRMED.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["2", "text"], "text": "The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042."}, {"path": ["3", "text"], "text": "Cell H27 in snapshot AR-6042 contains the value UNCONFIRMED."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042.", "negative_left": "The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042.", "negative_right": "Cell H27 in snapshot AR-6042 contains the value CONFIRMED.", "right": "Cell H27 in snapshot AR-6042 contains the value UNCONFIRMED."}, "verifier_independent_model": false}, "family": "scale-diverse-178-002", "id": "scale-diverse-178-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records requester", "text": "For reference PB-2018-05-14, I request the complete Parks Advisory Board packet from its May 14, 2018 meeting. If that complete packet cannot be found, you may provide the archived agenda and available attachments."}, {"speaker": "Intake clerk", "text": "After confirming the required request details and fallback permission, intake marked the records request process-ready."}, {"speaker": "Reconciliation log", "text": "The Archive Unit's 2026-08-19 reconciliation log states that the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 exactly matches cell H27 in snapshot AR-6042."}, {"speaker": "Archive snapshot", "text": "Cell H27 in snapshot AR-6042 contains the value CONFIRMED."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing, process-readiness, and difficulty policies, while the unchanged questions preserve all decision criteria. The same board, meeting date, reference, requested packet, fallback, routing path, and relevant timestamps are maintained. The two evidence spans are complete factual sentences describing a stored field value and its contemporaneous metadata mapping, rather than classifier instructions. Changing the field from code 47 (UNCONFIRMED) to code 82 (CONFIRMED) creates a coherent counterfactual without conflicting duplicate measurements. Neither context states the final yes/no answer, provides an answer code or proposition label, or instructs the classifier what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records intake specialist\",\"text\":\"Operational handoff for PB-2018-05-14: the requested record is the complete Parks Advisory Board meeting packet for the meeting held May 14, 2018. The requester permits the archived agenda and available attachments to be supplied only if that complete packet cannot be found. The intake checklist marks the request process-ready.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"},{\"speaker\":\"Routing coordinator\",\"text\":\"The May 14, 2018 meeting date is earlier than January 1, 2022, so the handoff is assigned to the Archive Unit.\"},{\"speaker\":\"Archive system log\",\"text\":\"The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.\"},{\"speaker\":\"Archive metadata record\",\"text\":\"The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["4", "text"], "text": "The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14."}, {"path": ["5", "text"], "text": "The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.", "negative_left": "The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 82 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.", "negative_right": "The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field.", "right": "The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}, "verifier_independent_model": false}, "family": "scale-diverse-178-003", "id": "scale-diverse-178-003-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records intake specialist", "text": "Operational handoff for PB-2018-05-14: the requested record is the complete Parks Advisory Board meeting packet for the meeting held May 14, 2018. The requester permits the archived agenda and available attachments to be supplied only if that complete packet cannot be found. The intake checklist marks the request process-ready."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}, {"speaker": "Routing coordinator", "text": "The May 14, 2018 meeting date is earlier than January 1, 2022, so the handoff is assigned to the Archive Unit."}, {"speaker": "Archive system log", "text": "The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14."}, {"speaker": "Archive metadata record", "text": "The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing, process-readiness, and difficulty policies, while the unchanged questions preserve all decision criteria. The same board, meeting date, reference, requested packet, fallback, routing path, and relevant timestamps are maintained. The two evidence spans are complete factual sentences describing a stored field value and its contemporaneous metadata mapping, rather than classifier instructions. Changing the field from code 47 (UNCONFIRMED) to code 82 (CONFIRMED) creates a coherent counterfactual without conflicting duplicate measurements. Neither context states the final yes/no answer, provides an answer code or proposition label, or instructs the classifier what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records intake specialist\",\"text\":\"Operational handoff for PB-2018-05-14: the requested record is the complete Parks Advisory Board meeting packet for the meeting held May 14, 2018. The requester permits the archived agenda and available attachments to be supplied only if that complete packet cannot be found. The intake checklist marks the request process-ready.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"},{\"speaker\":\"Routing coordinator\",\"text\":\"The May 14, 2018 meeting date is earlier than January 1, 2022, so the handoff is assigned to the Archive Unit.\"},{\"speaker\":\"Archive system log\",\"text\":\"The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.\"},{\"speaker\":\"Archive metadata record\",\"text\":\"The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["4", "text"], "text": "The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14."}, {"path": ["5", "text"], "text": "The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 47 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.", "negative_left": "The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 82 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14.", "negative_right": "The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field.", "right": "The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}, "verifier_independent_model": false}, "family": "scale-diverse-178-003", "id": "scale-diverse-178-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records intake specialist", "text": "Operational handoff for PB-2018-05-14: the requested record is the complete Parks Advisory Board meeting packet for the meeting held May 14, 2018. The requester permits the archived agenda and available attachments to be supplied only if that complete packet cannot be found. The intake checklist marks the request process-ready."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}, {"speaker": "Routing coordinator", "text": "The May 14, 2018 meeting date is earlier than January 1, 2022, so the handoff is assigned to the Archive Unit."}, {"speaker": "Archive system log", "text": "The Archive Unit operational-handoff record generated at 14:32 UTC on May 16, 2018, stores code 82 in the binary exact-item availability field for the complete packet requested under reference PB-2018-05-14."}, {"speaker": "Archive metadata record", "text": "The Archive Unit data dictionary issued at 09:00 UTC on May 1, 2018, maps code 47 to UNCONFIRMED and code 82 to CONFIRMED for the binary exact-item availability field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same fulfillment rubric, request scope, requester, records unit, notice, and relevant dates. The two evidence spans are complete factual sentences. In the counterfactual, the prepared complete PDF is EB-7316 while the completed check names the distinct globally unique artifact EB-9042; this breaks the same-artifact linkage needed to establish that the prepared copy was checked, but does not itself contradict the unchanged facts because the context does not state that EB-9042 is another complete prepared copy. Neither context embeds a score, gold answer, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified step and detail claims. The focus atom is the factual identity relation between two PDF artifacts, not a policy conclusion. Both assignments are realizable with only that identity changing: in the counter assignment, Nia can complete the sole recorded verification action on a different PDF while the sole complete prepared PDF remains unchecked. Empty policy_evidence is correct because the scoring rubric, latest-update priority, and other governing instructions are all in the questions object; the original state contributes case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails score 4: the responsible unit and sufficient identifying details are established; the material is located; the sole complete PDF is the artifact successfully checked against every verified detail; final verification is completed; all retrieval and other preparation steps are complete; and no later update supersedes this status.", "rule_index": 0, "sound": true}, {"reason": "The conjunction entails score 3. The material and its sole complete PDF have been located and prepared, but the only completed final-verification action checked a different artifact. Therefore the prepared complete copy has not received the required final verification, excluding score 4, while location of the material excludes the lower retrieval stages.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Public Works Records is the responsible records unit for Maya Chen's September 15, 2018 request for the archived Eastbrook footbridge closure notice."}, {"id": "a2", "statement": "The verified notice code PW-2018-0314 and verified posting date March 14, 2018 are sufficient to identify the notice requested by Maya Chen on September 15, 2018."}, {"id": "a3", "statement": "By September 17, 2018, staff had located the archived Eastbrook footbridge closure notice requested by Maya Chen."}, {"id": "a4", "statement": "By September 17, 2018, exactly one complete PDF copy had been prepared for Maya Chen's request."}, {"id": "a5", "statement": "For every verified detail of Maya Chen's request, Nia's September 17, 2018 final-check record records a successful comparison of its named PDF artifact against that detail."}, {"id": "a6", "statement": "The PDF artifact named in Nia's September 17, 2018 final-check record is the same PDF artifact as the complete PDF copy counted in the September 17, 2018 preparation record for Maya Chen's request."}, {"id": "a7", "statement": "Nia's September 17, 2018 check was a completed final-verification action for Maya Chen's request."}, {"id": "a8", "statement": "Exactly one final-verification action had been completed for Maya Chen's request by September 17, 2018."}, {"id": "a9", "statement": "Every required retrieval step for Maya Chen's request was complete by September 17, 2018."}, {"id": "a10", "statement": "Every required preparation step for Maya Chen's request other than final verification was complete by September 17, 2018."}, {"id": "a11", "statement": "No fulfillment-status update for Maya Chen's request is dated later than September 17, 2018."}], "base_state_json": "\"Public Works Records is the responsible records unit for Maya Chen’s September 15, 2018 request for the archived Eastbrook footbridge closure notice. The verified notice code PW-2018-0314 and posting date March 14, 2018 suffice to identify it. By September 17, staff had located the notice. The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316. Every required retrieval step was complete that day, as was every preparation step other than final verification. Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-7316. That final-check record reports a successful comparison of its named PDF against every verified request detail and marks Nia’s check as a completed final-verification action. Exactly one final-verification action had then been completed for the request. No fulfillment-status update is dated after September 17, 2018.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316."}, {"path": [], "text": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-7316."}], "policy_evidence": [], "rules": [{"justification": "The responsible unit and sufficient reference details are established, the requested notice has been located, and the sole complete PDF is the artifact successfully checked against every verified request detail. All required retrieval and non-verification preparation steps are complete, and the completed check supplies final verification of that PDF. With no later update, no retrieval or preparation step remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The requested notice has been located and its sole complete PDF has been prepared, but Nia's completed check concerned a different PDF artifact. Because exactly one final-verification action has been completed, there is no completed final verification of the prepared complete PDF. Final verification therefore remains unfinished, and no later update supersedes this located-but-pending stage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316.", "negative_left": "The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316.", "negative_right": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-9042.", "right": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-7316."}, "verifier_independent_model": false}, "family": "scale-diverse-179-002", "id": "scale-diverse-179-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unroutable: the requested municipal material or responsible records unit cannot be identified from the supplied details.", "1 — Routed but incomplete: the responsible unit is identified, but essential reference details are still missing or too broad for retrieval.", "2 — Complete and awaiting retrieval: sufficient details have been verified and the request is queued, but staff have not yet located the material.", "3 — Located, preparation pending: staff have found the material, but copying or final verification remains unfinished.", "4 — Ready for delivery: a complete copy has been prepared and checked against the verified request details, with no retrieval or preparation step remaining."], "instructions": "Rate the current fulfillment status using the ordered rubric. Apply the latest dated update; an earlier incomplete stage does not control after later work is documented.", "type": "score"}}, "state": "Public Works Records is the responsible records unit for Maya Chen’s September 15, 2018 request for the archived Eastbrook footbridge closure notice. The verified notice code PW-2018-0314 and posting date March 14, 2018 suffice to identify it. By September 17, staff had located the notice. The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316. Every required retrieval step was complete that day, as was every preparation step other than final verification. Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-7316. That final-check record reports a successful comparison of its named PDF against every verified request detail and marks Nia’s check as a completed final-verification action. Exactly one final-verification action had then been completed for the request. No fulfillment-status update is dated after September 17, 2018."}, "method": "c2d", "provenance": {"source_id": "diverse-179", "source_is_synthetic": true, "source_sha256": "6e7a9f82d359e340ca568b11d3b41a7dd157dd961cbbe35311953aaff2665a8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same fulfillment rubric, request scope, requester, records unit, notice, and relevant dates. The two evidence spans are complete factual sentences. In the counterfactual, the prepared complete PDF is EB-7316 while the completed check names the distinct globally unique artifact EB-9042; this breaks the same-artifact linkage needed to establish that the prepared copy was checked, but does not itself contradict the unchanged facts because the context does not state that EB-9042 is another complete prepared copy. Neither context embeds a score, gold answer, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified step and detail claims. The focus atom is the factual identity relation between two PDF artifacts, not a policy conclusion. Both assignments are realizable with only that identity changing: in the counter assignment, Nia can complete the sole recorded verification action on a different PDF while the sole complete prepared PDF remains unchecked. Empty policy_evidence is correct because the scoring rubric, latest-update priority, and other governing instructions are all in the questions object; the original state contributes case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails score 4: the responsible unit and sufficient identifying details are established; the material is located; the sole complete PDF is the artifact successfully checked against every verified detail; final verification is completed; all retrieval and other preparation steps are complete; and no later update supersedes this status.", "rule_index": 0, "sound": true}, {"reason": "The conjunction entails score 3. The material and its sole complete PDF have been located and prepared, but the only completed final-verification action checked a different artifact. Therefore the prepared complete copy has not received the required final verification, excluding score 4, while location of the material excludes the lower retrieval stages.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Public Works Records is the responsible records unit for Maya Chen's September 15, 2018 request for the archived Eastbrook footbridge closure notice."}, {"id": "a2", "statement": "The verified notice code PW-2018-0314 and verified posting date March 14, 2018 are sufficient to identify the notice requested by Maya Chen on September 15, 2018."}, {"id": "a3", "statement": "By September 17, 2018, staff had located the archived Eastbrook footbridge closure notice requested by Maya Chen."}, {"id": "a4", "statement": "By September 17, 2018, exactly one complete PDF copy had been prepared for Maya Chen's request."}, {"id": "a5", "statement": "For every verified detail of Maya Chen's request, Nia's September 17, 2018 final-check record records a successful comparison of its named PDF artifact against that detail."}, {"id": "a6", "statement": "The PDF artifact named in Nia's September 17, 2018 final-check record is the same PDF artifact as the complete PDF copy counted in the September 17, 2018 preparation record for Maya Chen's request."}, {"id": "a7", "statement": "Nia's September 17, 2018 check was a completed final-verification action for Maya Chen's request."}, {"id": "a8", "statement": "Exactly one final-verification action had been completed for Maya Chen's request by September 17, 2018."}, {"id": "a9", "statement": "Every required retrieval step for Maya Chen's request was complete by September 17, 2018."}, {"id": "a10", "statement": "Every required preparation step for Maya Chen's request other than final verification was complete by September 17, 2018."}, {"id": "a11", "statement": "No fulfillment-status update for Maya Chen's request is dated later than September 17, 2018."}], "base_state_json": "\"Public Works Records is the responsible records unit for Maya Chen’s September 15, 2018 request for the archived Eastbrook footbridge closure notice. The verified notice code PW-2018-0314 and posting date March 14, 2018 suffice to identify it. By September 17, staff had located the notice. The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316. Every required retrieval step was complete that day, as was every preparation step other than final verification. Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-7316. That final-check record reports a successful comparison of its named PDF against every verified request detail and marks Nia’s check as a completed final-verification action. Exactly one final-verification action had then been completed for the request. No fulfillment-status update is dated after September 17, 2018.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316."}, {"path": [], "text": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-7316."}], "policy_evidence": [], "rules": [{"justification": "The responsible unit and sufficient reference details are established, the requested notice has been located, and the sole complete PDF is the artifact successfully checked against every verified request detail. All required retrieval and non-verification preparation steps are complete, and the completed check supplies final verification of that PDF. With no later update, no retrieval or preparation step remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The requested notice has been located and its sole complete PDF has been prepared, but Nia's completed check concerned a different PDF artifact. Because exactly one final-verification action has been completed, there is no completed final verification of the prepared complete PDF. Final verification therefore remains unfinished, and no later update supersedes this located-but-pending stage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316.", "negative_left": "The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316.", "negative_right": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-9042.", "right": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-7316."}, "verifier_independent_model": false}, "family": "scale-diverse-179-002", "id": "scale-diverse-179-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unroutable: the requested municipal material or responsible records unit cannot be identified from the supplied details.", "1 — Routed but incomplete: the responsible unit is identified, but essential reference details are still missing or too broad for retrieval.", "2 — Complete and awaiting retrieval: sufficient details have been verified and the request is queued, but staff have not yet located the material.", "3 — Located, preparation pending: staff have found the material, but copying or final verification remains unfinished.", "4 — Ready for delivery: a complete copy has been prepared and checked against the verified request details, with no retrieval or preparation step remaining."], "instructions": "Rate the current fulfillment status using the ordered rubric. Apply the latest dated update; an earlier incomplete stage does not control after later work is documented.", "type": "score"}}, "state": "Public Works Records is the responsible records unit for Maya Chen’s September 15, 2018 request for the archived Eastbrook footbridge closure notice. The verified notice code PW-2018-0314 and posting date March 14, 2018 suffice to identify it. By September 17, staff had located the notice. The September 17, 2018 preparation ledger for Maya Chen's request counts exactly one complete PDF copy and assigns it the globally unique PDF-artifact identifier EB-7316. Every required retrieval step was complete that day, as was every preparation step other than final verification. Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF bearing the globally unique PDF-artifact identifier EB-9042. That final-check record reports a successful comparison of its named PDF against every verified request detail and marks Nia’s check as a completed final-verification action. Exactly one final-verification action had then been completed for the request. No fulfillment-status update is dated after September 17, 2018."}, "method": "c2d", "provenance": {"source_id": "diverse-179", "source_is_synthetic": true, "source_sha256": "6e7a9f82d359e340ca568b11d3b41a7dd157dd961cbbe35311953aaff2665a8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the full scoring policy, including the latest-update instruction, and neither context alters it. Both contexts remain bound to Maya Chen’s request, the Eastbrook footbridge notice, Public Works Records, and the relevant September 2018 timeline. The two evidence spans are complete factual sentences. In the counterfactual, the sole final-verification action concerns MC-29V rather than the prepared complete copy MC-74K; this breaks the artifact linkage without contradicting the statements that MC-74K has no other tag or that no other artifact bears MC-74K. Neither context states a score, answer code, classifier instruction, or explicit gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified step and detail claims. The focus atom is the factual identity relation between two PDF artifacts, not a policy conclusion. Both assignments are realizable with only that identity changing: in the counter assignment, Nia can complete the sole recorded verification action on a different PDF while the sole complete prepared PDF remains unchecked. Empty policy_evidence is correct because the scoring rubric, latest-update priority, and other governing instructions are all in the questions object; the original state contributes case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails score 4: the responsible unit and sufficient identifying details are established; the material is located; the sole complete PDF is the artifact successfully checked against every verified detail; final verification is completed; all retrieval and other preparation steps are complete; and no later update supersedes this status.", "rule_index": 0, "sound": true}, {"reason": "The conjunction entails score 3. The material and its sole complete PDF have been located and prepared, but the only completed final-verification action checked a different artifact. Therefore the prepared complete copy has not received the required final verification, excluding score 4, while location of the material excludes the lower retrieval stages.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Public Works Records is the responsible records unit for Maya Chen's September 15, 2018 request for the archived Eastbrook footbridge closure notice."}, {"id": "a2", "statement": "The verified notice code PW-2018-0314 and verified posting date March 14, 2018 are sufficient to identify the notice requested by Maya Chen on September 15, 2018."}, {"id": "a3", "statement": "By September 17, 2018, staff had located the archived Eastbrook footbridge closure notice requested by Maya Chen."}, {"id": "a4", "statement": "By September 17, 2018, exactly one complete PDF copy had been prepared for Maya Chen's request."}, {"id": "a5", "statement": "For every verified detail of Maya Chen's request, Nia's September 17, 2018 final-check record records a successful comparison of its named PDF artifact against that detail."}, {"id": "a6", "statement": "The PDF artifact named in Nia's September 17, 2018 final-check record is the same PDF artifact as the complete PDF copy counted in the September 17, 2018 preparation record for Maya Chen's request."}, {"id": "a7", "statement": "Nia's September 17, 2018 check was a completed final-verification action for Maya Chen's request."}, {"id": "a8", "statement": "Exactly one final-verification action had been completed for Maya Chen's request by September 17, 2018."}, {"id": "a9", "statement": "Every required retrieval step for Maya Chen's request was complete by September 17, 2018."}, {"id": "a10", "statement": "Every required preparation step for Maya Chen's request other than final verification was complete by September 17, 2018."}, {"id": "a11", "statement": "No fulfillment-status update for Maya Chen's request is dated later than September 17, 2018."}], "base_state_json": "\"Public Works Records is the responsible records unit for Maya Chen's September 15, 2018 request for the archived Eastbrook footbridge closure notice. The verified code PW-2018-0314 and posting date March 14, 2018 suffice to identify that notice. By September 17, staff had located it, and every required retrieval step was complete. The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K. It also records every required preparation step other than final verification as complete. Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-74K. That record documents a successful comparison of its named artifact against every verified request detail and a completed final-verification action. Exactly one final-verification action had been completed for the request by September 17. No fulfillment-status update is dated later than September 17, 2018.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K."}, {"path": [], "text": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-74K."}], "policy_evidence": [], "rules": [{"justification": "The responsible unit and sufficient reference details are established, the requested notice has been located, and the sole complete PDF is the artifact successfully checked against every verified request detail. All required retrieval and non-verification preparation steps are complete, and the completed check supplies final verification of that PDF. With no later update, no retrieval or preparation step remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The requested notice has been located and its sole complete PDF has been prepared, but Nia's completed check concerned a different PDF artifact. Because exactly one final-verification action has been completed, there is no completed final verification of the prepared complete PDF. Final verification therefore remains unfinished, and no later update supersedes this located-but-pending stage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K.", "negative_left": "The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K.", "negative_right": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-29V.", "right": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-74K."}, "verifier_independent_model": false}, "family": "scale-diverse-179-004", "id": "scale-diverse-179-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Unroutable: the requested municipal material or responsible records unit cannot be identified from the supplied details.", "1 — Routed but incomplete: the responsible unit is identified, but essential reference details are still missing or too broad for retrieval.", "2 — Complete and awaiting retrieval: sufficient details have been verified and the request is queued, but staff have not yet located the material.", "3 — Located, preparation pending: staff have found the material, but copying or final verification remains unfinished.", "4 — Ready for delivery: a complete copy has been prepared and checked against the verified request details, with no retrieval or preparation step remaining."], "instructions": "Rate the current fulfillment status using the ordered rubric. Apply the latest dated update; an earlier incomplete stage does not control after later work is documented.", "type": "score"}}, "state": "Public Works Records is the responsible records unit for Maya Chen's September 15, 2018 request for the archived Eastbrook footbridge closure notice. The verified code PW-2018-0314 and posting date March 14, 2018 suffice to identify that notice. By September 17, staff had located it, and every required retrieval step was complete. The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K. It also records every required preparation step other than final verification as complete. Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-74K. That record documents a successful comparison of its named artifact against every verified request detail and a completed final-verification action. Exactly one final-verification action had been completed for the request by September 17. No fulfillment-status update is dated later than September 17, 2018."}, "method": "c2d", "provenance": {"source_id": "diverse-179", "source_is_synthetic": true, "source_sha256": "6e7a9f82d359e340ca568b11d3b41a7dd157dd961cbbe35311953aaff2665a8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the full scoring policy, including the latest-update instruction, and neither context alters it. Both contexts remain bound to Maya Chen’s request, the Eastbrook footbridge notice, Public Works Records, and the relevant September 2018 timeline. The two evidence spans are complete factual sentences. In the counterfactual, the sole final-verification action concerns MC-29V rather than the prepared complete copy MC-74K; this breaks the artifact linkage without contradicting the statements that MC-74K has no other tag or that no other artifact bears MC-74K. Neither context states a score, answer code, classifier instruction, or explicit gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified step and detail claims. The focus atom is the factual identity relation between two PDF artifacts, not a policy conclusion. Both assignments are realizable with only that identity changing: in the counter assignment, Nia can complete the sole recorded verification action on a different PDF while the sole complete prepared PDF remains unchecked. Empty policy_evidence is correct because the scoring rubric, latest-update priority, and other governing instructions are all in the questions object; the original state contributes case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails score 4: the responsible unit and sufficient identifying details are established; the material is located; the sole complete PDF is the artifact successfully checked against every verified detail; final verification is completed; all retrieval and other preparation steps are complete; and no later update supersedes this status.", "rule_index": 0, "sound": true}, {"reason": "The conjunction entails score 3. The material and its sole complete PDF have been located and prepared, but the only completed final-verification action checked a different artifact. Therefore the prepared complete copy has not received the required final verification, excluding score 4, while location of the material excludes the lower retrieval stages.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Public Works Records is the responsible records unit for Maya Chen's September 15, 2018 request for the archived Eastbrook footbridge closure notice."}, {"id": "a2", "statement": "The verified notice code PW-2018-0314 and verified posting date March 14, 2018 are sufficient to identify the notice requested by Maya Chen on September 15, 2018."}, {"id": "a3", "statement": "By September 17, 2018, staff had located the archived Eastbrook footbridge closure notice requested by Maya Chen."}, {"id": "a4", "statement": "By September 17, 2018, exactly one complete PDF copy had been prepared for Maya Chen's request."}, {"id": "a5", "statement": "For every verified detail of Maya Chen's request, Nia's September 17, 2018 final-check record records a successful comparison of its named PDF artifact against that detail."}, {"id": "a6", "statement": "The PDF artifact named in Nia's September 17, 2018 final-check record is the same PDF artifact as the complete PDF copy counted in the September 17, 2018 preparation record for Maya Chen's request."}, {"id": "a7", "statement": "Nia's September 17, 2018 check was a completed final-verification action for Maya Chen's request."}, {"id": "a8", "statement": "Exactly one final-verification action had been completed for Maya Chen's request by September 17, 2018."}, {"id": "a9", "statement": "Every required retrieval step for Maya Chen's request was complete by September 17, 2018."}, {"id": "a10", "statement": "Every required preparation step for Maya Chen's request other than final verification was complete by September 17, 2018."}, {"id": "a11", "statement": "No fulfillment-status update for Maya Chen's request is dated later than September 17, 2018."}], "base_state_json": "\"Public Works Records is the responsible records unit for Maya Chen's September 15, 2018 request for the archived Eastbrook footbridge closure notice. The verified code PW-2018-0314 and posting date March 14, 2018 suffice to identify that notice. By September 17, staff had located it, and every required retrieval step was complete. The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K. It also records every required preparation step other than final verification as complete. Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-74K. That record documents a successful comparison of its named artifact against every verified request detail and a completed final-verification action. Exactly one final-verification action had been completed for the request by September 17. No fulfillment-status update is dated later than September 17, 2018.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K."}, {"path": [], "text": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-74K."}], "policy_evidence": [], "rules": [{"justification": "The responsible unit and sufficient reference details are established, the requested notice has been located, and the sole complete PDF is the artifact successfully checked against every verified request detail. All required retrieval and non-verification preparation steps are complete, and the completed check supplies final verification of that PDF. With no later update, no retrieval or preparation step remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The requested notice has been located and its sole complete PDF has been prepared, but Nia's completed check concerned a different PDF artifact. Because exactly one final-verification action has been completed, there is no completed final verification of the prepared complete PDF. Final verification therefore remains unfinished, and no later update supersedes this located-but-pending stage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K.", "negative_left": "The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K.", "negative_right": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-29V.", "right": "Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-74K."}, "verifier_independent_model": false}, "family": "scale-diverse-179-004", "id": "scale-diverse-179-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unroutable: the requested municipal material or responsible records unit cannot be identified from the supplied details.", "1 — Routed but incomplete: the responsible unit is identified, but essential reference details are still missing or too broad for retrieval.", "2 — Complete and awaiting retrieval: sufficient details have been verified and the request is queued, but staff have not yet located the material.", "3 — Located, preparation pending: staff have found the material, but copying or final verification remains unfinished.", "4 — Ready for delivery: a complete copy has been prepared and checked against the verified request details, with no retrieval or preparation step remaining."], "instructions": "Rate the current fulfillment status using the ordered rubric. Apply the latest dated update; an earlier incomplete stage does not control after later work is documented.", "type": "score"}}, "state": "Public Works Records is the responsible records unit for Maya Chen's September 15, 2018 request for the archived Eastbrook footbridge closure notice. The verified code PW-2018-0314 and posting date March 14, 2018 suffice to identify that notice. By September 17, staff had located it, and every required retrieval step was complete. The September 17, 2018 preparation record for Maya Chen's request counts one complete PDF copy bearing serial tag MC-74K, records that the copy bears no other serial tag, and records that no other PDF artifact for the request bears MC-74K. It also records every required preparation step other than final verification as complete. Nia's September 17, 2018 final-check record for Maya Chen's request names the PDF artifact bearing serial tag MC-29V. That record documents a successful comparison of its named artifact against every verified request detail and a completed final-verification action. Exactly one final-verification action had been completed for the request by September 17. No fulfillment-status update is dated later than September 17, 2018."}, "method": "c2d", "provenance": {"source_id": "diverse-179", "source_is_synthetic": true, "source_sha256": "6e7a9f82d359e340ca568b11d3b41a7dd157dd961cbbe35311953aaff2665a8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing external-product and appliance-origin policies without adding exceptions or defaults, and they preserve the decision scope and the 9:10 p.m. spill, bottle, and washer bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the bottle-liquid identifier from SP-417 to SP-862; given the stated identifier-set and uniqueness facts, this does not create a contradictory duplicate measurement. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"At 9:12 p.m., Inspector Nia limited the investigated possible origins of the 9:10 p.m. floor spill to exactly the uncapped detergent bottle on the shelf and the washer. At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample. At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-417 for the liquid in the uncapped detergent bottle on the shelf. At 9:31 p.m., the report stated that the spill identifier belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier; those two identifiers differ. The investigation protocol uniquely associates every identifier used with its possible origin in that set, and the spill has exactly one origin. Records classify the bottle as external to the washer and its liquid as a detergent product. Inspection found no active flooding or electrical signs at the washer. Logs show no recurring fault. Cycle notes do not confirm an unbalanced or overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample."}, {"path": [], "text": "At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-417 for the liquid in the uncapped detergent bottle on the shelf."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample.", "negative_left": "At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample.", "negative_right": "At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-862 for the liquid in the uncapped detergent bottle on the shelf.", "right": "At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-417 for the liquid in the uncapped detergent bottle on the shelf."}, "verifier_independent_model": false}, "family": "scale-diverse-181-001", "id": "scale-diverse-181-001-base", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "At 9:12 p.m., Inspector Nia limited the investigated possible origins of the 9:10 p.m. floor spill to exactly the uncapped detergent bottle on the shelf and the washer. At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample. At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-417 for the liquid in the uncapped detergent bottle on the shelf. At 9:31 p.m., the report stated that the spill identifier belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier; those two identifiers differ. The investigation protocol uniquely associates every identifier used with its possible origin in that set, and the spill has exactly one origin. Records classify the bottle as external to the washer and its liquid as a detergent product. Inspection found no active flooding or electrical signs at the washer. Logs show no recurring fault. Cycle notes do not confirm an unbalanced or overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing external-product and appliance-origin policies without adding exceptions or defaults, and they preserve the decision scope and the 9:10 p.m. spill, bottle, and washer bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the bottle-liquid identifier from SP-417 to SP-862; given the stated identifier-set and uniqueness facts, this does not create a contradictory duplicate measurement. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"At 9:12 p.m., Inspector Nia limited the investigated possible origins of the 9:10 p.m. floor spill to exactly the uncapped detergent bottle on the shelf and the washer. At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample. At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-417 for the liquid in the uncapped detergent bottle on the shelf. At 9:31 p.m., the report stated that the spill identifier belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier; those two identifiers differ. The investigation protocol uniquely associates every identifier used with its possible origin in that set, and the spill has exactly one origin. Records classify the bottle as external to the washer and its liquid as a detergent product. Inspection found no active flooding or electrical signs at the washer. Logs show no recurring fault. Cycle notes do not confirm an unbalanced or overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample."}, {"path": [], "text": "At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-417 for the liquid in the uncapped detergent bottle on the shelf."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample.", "negative_left": "At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample.", "negative_right": "At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-862 for the liquid in the uncapped detergent bottle on the shelf.", "right": "At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-417 for the liquid in the uncapped detergent bottle on the shelf."}, "verifier_independent_model": false}, "family": "scale-diverse-181-001", "id": "scale-diverse-181-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "At 9:12 p.m., Inspector Nia limited the investigated possible origins of the 9:10 p.m. floor spill to exactly the uncapped detergent bottle on the shelf and the washer. At 9:18 p.m., the laboratory analyzer recorded source-profile identifier SP-417 for the 9:10 p.m. floor-spill sample. At 9:26 p.m., the same laboratory analyzer recorded source-profile identifier SP-862 for the liquid in the uncapped detergent bottle on the shelf. At 9:31 p.m., the report stated that the spill identifier belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier; those two identifiers differ. The investigation protocol uniquely associates every identifier used with its possible origin in that set, and the spill has exactly one origin. Records classify the bottle as external to the washer and its liquid as a detergent product. Inspection found no active flooding or electrical signs at the washer. Logs show no recurring fault. Cycle notes do not confirm an unbalanced or overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "basic_inspection_hold_high"}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing spill policy without adding exceptions, priorities, or missing-evidence defaults; the fuller route/reuse/urgency criteria and selection instructions remain supplied by the unchanged original question. The decision scope and washer/spill/time bindings are preserved. The two focus spans are complete factual measurement sentences. The counterfactual coherently changes the bottle identifier from SP-641 to distinct SP-908; given the stated two-origin set, differing origin identifiers, uniqueness, and single-origin constraint, this creates no duplicate or contradictory measurement. Neither context includes an option code, gold answer, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Evidence reconciliation note: The investigation limited possible origins of the 9:10 p.m. spill to exactly the uncapped detergent bottle on the shelf and the washer. At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample. At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the liquid in the uncapped detergent bottle on the shelf. The spill identifier belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier; those two origin identifiers differ. Each source-profile identifier used in the investigation uniquely identifies its associated possible origin within that set, and the spill has exactly one origin. The bottle is external to the washer, and its liquid is a detergent product. Inspection found no active flooding or electrical signs at the washer. Washer logs show no recurring fault. Cycle notes confirm neither an unbalanced nor an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample."}, {"path": [], "text": "At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the liquid in the uncapped detergent bottle on the shelf."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample.", "negative_left": "At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample.", "negative_right": "At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-908 for the liquid in the uncapped detergent bottle on the shelf, and SP-641 and SP-908 are distinct identifiers.", "right": "At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the liquid in the uncapped detergent bottle on the shelf."}, "verifier_independent_model": false}, "family": "scale-diverse-181-002", "id": "scale-diverse-181-002-base", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Evidence reconciliation note: The investigation limited possible origins of the 9:10 p.m. spill to exactly the uncapped detergent bottle on the shelf and the washer. At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample. At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the liquid in the uncapped detergent bottle on the shelf. The spill identifier belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier; those two origin identifiers differ. Each source-profile identifier used in the investigation uniquely identifies its associated possible origin within that set, and the spill has exactly one origin. The bottle is external to the washer, and its liquid is a detergent product. Inspection found no active flooding or electrical signs at the washer. Washer logs show no recurring fault. Cycle notes confirm neither an unbalanced nor an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing spill policy without adding exceptions, priorities, or missing-evidence defaults; the fuller route/reuse/urgency criteria and selection instructions remain supplied by the unchanged original question. The decision scope and washer/spill/time bindings are preserved. The two focus spans are complete factual measurement sentences. The counterfactual coherently changes the bottle identifier from SP-641 to distinct SP-908; given the stated two-origin set, differing origin identifiers, uniqueness, and single-origin constraint, this creates no duplicate or contradictory measurement. Neither context includes an option code, gold answer, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Evidence reconciliation note: The investigation limited possible origins of the 9:10 p.m. spill to exactly the uncapped detergent bottle on the shelf and the washer. At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample. At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the liquid in the uncapped detergent bottle on the shelf. The spill identifier belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier; those two origin identifiers differ. Each source-profile identifier used in the investigation uniquely identifies its associated possible origin within that set, and the spill has exactly one origin. The bottle is external to the washer, and its liquid is a detergent product. Inspection found no active flooding or electrical signs at the washer. Washer logs show no recurring fault. Cycle notes confirm neither an unbalanced nor an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample."}, {"path": [], "text": "At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the liquid in the uncapped detergent bottle on the shelf."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample.", "negative_left": "At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample.", "negative_right": "At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-908 for the liquid in the uncapped detergent bottle on the shelf, and SP-641 and SP-908 are distinct identifiers.", "right": "At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the liquid in the uncapped detergent bottle on the shelf."}, "verifier_independent_model": false}, "family": "scale-diverse-181-002", "id": "scale-diverse-181-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Evidence reconciliation note: The investigation limited possible origins of the 9:10 p.m. spill to exactly the uncapped detergent bottle on the shelf and the washer. At 9:24 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-641 for the 9:10 p.m. floor-spill sample. At 9:26 p.m., calibrated analyzer Q7 recorded source-profile identifier SP-908 for the liquid in the uncapped detergent bottle on the shelf, and SP-641 and SP-908 are distinct identifiers. The spill identifier belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier; those two origin identifiers differ. Each source-profile identifier used in the investigation uniquely identifies its associated possible origin within that set, and the spill has exactly one origin. The bottle is external to the washer, and its liquid is a detergent product. Inspection found no active flooding or electrical signs at the washer. Washer logs show no recurring fault. Cycle notes confirm neither an unbalanced nor an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "basic_inspection_hold_high"}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original incident, washer, timing, decision scope, and policy statements without adding exceptions or defaults. The two evidence spans are complete factual sentences; changing only the bottle-liquid identifier coherently shifts the identifier relationship while remaining consistent with the two-origin set, uniqueness, and single-origin assertions, and neither context contains a gold answer, option code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Operational handoff: The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer. The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316. The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-7316. Chain-of-custody certification states that the spill identifier belongs to the set containing the bottle-liquid identifier and washer-moisture identifier. Those two possible-origin identifiers differ, and every identifier used in the investigation uniquely identifies its associated possible origin within that set. The incident record assigns the spill exactly one origin. The uncapped bottle is external to the washer, and its liquid is a detergent product. Site checks found no active flooding or electrical signs at the washer. Washer logs show no recurring fault. Cycle notes do not confirm an unbalanced load and do not confirm an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316."}, {"path": [], "text": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-7316."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316.", "negative_left": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316.", "negative_right": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-2849.", "right": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-7316."}, "verifier_independent_model": false}, "family": "scale-diverse-181-003", "id": "scale-diverse-181-003-base", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Operational handoff: The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer. The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316. The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-7316. Chain-of-custody certification states that the spill identifier belongs to the set containing the bottle-liquid identifier and washer-moisture identifier. Those two possible-origin identifiers differ, and every identifier used in the investigation uniquely identifies its associated possible origin within that set. The incident record assigns the spill exactly one origin. The uncapped bottle is external to the washer, and its liquid is a detergent product. Site checks found no active flooding or electrical signs at the washer. Washer logs show no recurring fault. Cycle notes do not confirm an unbalanced load and do not confirm an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original incident, washer, timing, decision scope, and policy statements without adding exceptions or defaults. The two evidence spans are complete factual sentences; changing only the bottle-liquid identifier coherently shifts the identifier relationship while remaining consistent with the two-origin set, uniqueness, and single-origin assertions, and neither context contains a gold answer, option code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Operational handoff: The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer. The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316. The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-7316. Chain-of-custody certification states that the spill identifier belongs to the set containing the bottle-liquid identifier and washer-moisture identifier. Those two possible-origin identifiers differ, and every identifier used in the investigation uniquely identifies its associated possible origin within that set. The incident record assigns the spill exactly one origin. The uncapped bottle is external to the washer, and its liquid is a detergent product. Site checks found no active flooding or electrical signs at the washer. Washer logs show no recurring fault. Cycle notes do not confirm an unbalanced load and do not confirm an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316."}, {"path": [], "text": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-7316."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316.", "negative_left": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316.", "negative_right": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-2849.", "right": "The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-7316."}, "verifier_independent_model": false}, "family": "scale-diverse-181-003", "id": "scale-diverse-181-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Operational handoff: The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer. The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the 9:10 p.m. floor-spill sample as QF-7316. The 9:18 p.m. laboratory handoff record lists the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf as QF-2849. Chain-of-custody certification states that the spill identifier belongs to the set containing the bottle-liquid identifier and washer-moisture identifier. Those two possible-origin identifiers differ, and every identifier used in the investigation uniquely identifies its associated possible origin within that set. The incident record assigns the spill exactly one origin. The uncapped bottle is external to the washer, and its liquid is a detergent product. Site checks found no active flooding or electrical signs at the washer. Washer logs show no recurring fault. Cycle notes do not confirm an unbalanced load and do not confirm an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "basic_inspection_hold_high"}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original-state governing policy, while the unchanged questions object preserves all decision criteria and instructions. The washer incident, source-identification path, and 9:10 p.m. sample binding remain intact. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the bottle-liquid identifier: because the spill identifier must equal one of two distinct, unique origin identifiers, it shifts the inferred origin from the bottle to the washer without creating duplicate contradictory measurements. Neither context contains a gold answer, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Field note: The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample. The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf. Investigators limited the complete possible-origin set to that bottle and the washer. The spill identifier belongs to the set containing the bottle-liquid and washer-moisture identifiers. Those two origin identifiers differ; each investigation identifier uniquely identifies its associated origin within that set. The spill has exactly one origin. The bottle is external to the washer, and its contents are a detergent product. Inspection found no active flooding or electrical signs. Washer logs show no recurring fault. Cycle notes confirm neither an unbalanced nor an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample."}, {"path": [], "text": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample.", "negative_left": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample.", "negative_right": "The 9:24 p.m. laboratory field note records identifier SP-4829 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf.", "right": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf."}, "verifier_independent_model": false}, "family": "scale-diverse-181-004", "id": "scale-diverse-181-004-base", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Field note: The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample. The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf. Investigators limited the complete possible-origin set to that bottle and the washer. The spill identifier belongs to the set containing the bottle-liquid and washer-moisture identifiers. Those two origin identifiers differ; each investigation identifier uniquely identifies its associated origin within that set. The spill has exactly one origin. The bottle is external to the washer, and its contents are a detergent product. Inspection found no active flooding or electrical signs. Washer logs show no recurring fault. Cycle notes confirm neither an unbalanced nor an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original-state governing policy, while the unchanged questions object preserves all decision criteria and instructions. The washer incident, source-identification path, and 9:10 p.m. sample binding remain intact. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the bottle-liquid identifier: because the spill identifier must equal one of two distinct, unique origin identifiers, it shifts the inferred origin from the bottle to the washer without creating duplicate contradictory measurements. Neither context contains a gold answer, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Field note: The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample. The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf. Investigators limited the complete possible-origin set to that bottle and the washer. The spill identifier belongs to the set containing the bottle-liquid and washer-moisture identifiers. Those two origin identifiers differ; each investigation identifier uniquely identifies its associated origin within that set. The spill has exactly one origin. The bottle is external to the washer, and its contents are a detergent product. Inspection found no active flooding or electrical signs. Washer logs show no recurring fault. Cycle notes confirm neither an unbalanced nor an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample."}, {"path": [], "text": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample.", "negative_left": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample.", "negative_right": "The 9:24 p.m. laboratory field note records identifier SP-4829 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf.", "right": "The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf."}, "verifier_independent_model": false}, "family": "scale-diverse-181-004", "id": "scale-diverse-181-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Field note: The 9:24 p.m. laboratory field note records identifier SP-7316 for the source profile measured from the 9:10 p.m. floor-spill sample. The 9:24 p.m. laboratory field note records identifier SP-4829 for the source profile measured from the liquid in the uncapped detergent bottle on the shelf. Investigators limited the complete possible-origin set to that bottle and the washer. The spill identifier belongs to the set containing the bottle-liquid and washer-moisture identifiers. Those two origin identifiers differ; each investigation identifier uniquely identifies its associated origin within that set. The spill has exactly one origin. The bottle is external to the washer, and its contents are a detergent product. Inspection found no active flooding or electrical signs. Washer logs show no recurring fault. Cycle notes confirm neither an unbalanced nor an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "basic_inspection_hold_high"}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the resident’s specified reuse condition and remain scoped to the same washer, inspector-completed balanced empty-spin test, and readiness decision interval. The two evidence spans are complete factual sentences describing the counter’s operation and readings, not policy instructions. The counterfactual changes only the completion reading from 73 to 75; this is coherent with a cumulative event counter that increments once per observed banging event, because two events can account for the two-count increase and no unchanged statement limits the run to one event. Neither context embeds a decision, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "refuted", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A15": "unknown"}, "remove_right": {"A15": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A15": "unknown"}, "negative_pair": {"A15": "refuted"}, "negative_sentence": {"A15": "unknown"}, "positive_pair": {"A15": "supported"}, "right": {"A15": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a factual relationship; interval-bounded and uniqueness propositions remain atomic despite quantifying over times or tests. A15 is a factual observation about the banging count, not a policy conclusion. The base and counter assignments differ only on A15 and are realizable: a spin may be balanced yet still produce banging. Policy evidence correctly preserves the resident’s state-originating, case-specific reuse condition, while the general governing readiness rule remains automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that inspector P completed the resident’s specified test in the relevant interval: an empty, balanced spin with zero banging. This successfully fulfills the express reuse condition, so readiness follows. The additional incident and uniqueness conditions are unnecessary but do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "A15 being refuted entails that test R had a nonzero banging count, so it did not satisfy the required no-banging test. A16 excludes any other completed inspector spin test in the relevant interval, preventing a competing successful test. Therefore the resident’s express condition remains unfulfilled and the false target follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During the period covered by the evidence through readiness decision D, incident E was the washer's only banging incident."}, {"id": "A2", "statement": "The washer contained exactly one item during incident E."}, {"id": "A3", "statement": "The item in the washer was off-center during incident E."}, {"id": "A4", "statement": "No water leaked from the washer during or after incident E through readiness decision D."}, {"id": "A5", "statement": "No odor was detected from the washer during or after incident E through readiness decision D."}, {"id": "A6", "statement": "No scraping was detected from the washer during or after incident E through readiness decision D."}, {"id": "A7", "statement": "Inspection of the washer after incident E detected no looseness."}, {"id": "A8", "statement": "Inspection of the washer after incident E detected no damage."}, {"id": "A9", "statement": "Person P was an inspector at the time of test run R."}, {"id": "A10", "statement": "Person P performed test run R on the washer."}, {"id": "A11", "statement": "Test run R was completed."}, {"id": "A12", "statement": "Test run R occurred after the resident stated the reuse condition and before readiness decision D."}, {"id": "A13", "statement": "The washer drum contained zero items during test run R."}, {"id": "A14", "statement": "The washer's spin was balanced during test run R."}, {"id": "A15", "statement": "The observed banging-event count during test run R was zero."}, {"id": "A16", "statement": "Test run R was the only washer spin test completed by any inspector after the resident stated the reuse condition and before readiness decision D."}], "base_state_json": "[{\"speaker\":\"Resident\",\"text\":\"The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging.\"},{\"speaker\":\"Incident handoff\",\"text\":\"Case records through readiness decision D list incident E as the washer's only banging incident; controlled test observations are logged separately rather than as incidents. E occurred with exactly one bath mat, visibly folded off-center. No water leakage, odor, or scraping was detected during or after E through D. Post-E inspection found neither looseness nor damage.\"},{\"speaker\":\"Test counter record\",\"text\":\"At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases.\"},{\"speaker\":\"Test counter record\",\"text\":\"At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73.\"},{\"speaker\":\"Completion handoff\",\"text\":\"Person P was the assigned inspector when R occurred and personally performed and completed it. R took place after the resident stated the reuse condition and before D. The drum contained zero items, and the spin was balanced. R was the only washer spin test completed by any inspector within that interval.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "supported"}], "focus_atom": "A15", "focus_evidence": [{"path": ["2", "text"], "text": "At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases."}, {"path": ["3", "text"], "text": "At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73."}], "policy_evidence": [{"path": ["0", "text"], "text": "The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging."}], "rules": [{"justification": "The incident satisfies every stated routine load-correction qualification, and inspector P completed an empty, balanced spin without banging after the resident imposed the condition. The resident's specified test was therefore completed successfully.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}]}, {"justification": "Although the incident satisfies every stated routine load-correction qualification, the only inspector-completed spin test in the relevant interval had a nonzero banging count. Thus no competing inspector-completed test could have fulfilled the resident's requirement for a balanced empty spin without banging.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases.", "negative_left": "At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases.", "negative_right": "At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 75.", "right": "At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73."}, "verifier_independent_model": false}, "family": "scale-diverse-183-003", "id": "scale-diverse-183-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one stated condition for the resident’s intended reuse remains unfulfilled, even if the incident otherwise appears routine.", "true": "Yes — the evidence establishes readiness for the resident’s intended reuse, including fulfillment of every stated condition."}, "instructions": "Answer whether the washer is ready for the resident’s intended reuse. Routing options are cleanup, load correction, basic inspection, and deferred maintenance. Urgency is ordered: 1 = routine, 2 = prompt, 3 = urgent, 4 = immediate shutdown. Under the household policy, an isolated banging event with an off-center single item, no leak, odor, scraping, looseness, or damage routes to load correction at urgency 1. However, when the resident expressly conditions reuse on a specified test, readiness is established only after that test is completed successfully. An inspection plan or other favorable evidence does not satisfy the condition.", "type": "noul"}}, "state": [{"speaker": "Resident", "text": "The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging."}, {"speaker": "Incident handoff", "text": "Case records through readiness decision D list incident E as the washer's only banging incident; controlled test observations are logged separately rather than as incidents. E occurred with exactly one bath mat, visibly folded off-center. No water leakage, odor, or scraping was detected during or after E through D. Post-E inspection found neither looseness nor damage."}, {"speaker": "Test counter record", "text": "At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases."}, {"speaker": "Test counter record", "text": "At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73."}, {"speaker": "Completion handoff", "text": "Person P was the assigned inspector when R occurred and personally performed and completed it. R took place after the resident stated the reuse condition and before D. The drum contained zero items, and the spin was balanced. R was the only washer spin test completed by any inspector within that interval."}]}, "method": "c2d", "provenance": {"source_id": "diverse-183", "source_is_synthetic": true, "source_sha256": "e3c0610d70e581b628bc4f9c3a5638fa1feb83c55d1a41d1e45a1f8fe0e3e6ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the resident’s specified reuse condition and remain scoped to the same washer, inspector-completed balanced empty-spin test, and readiness decision interval. The two evidence spans are complete factual sentences describing the counter’s operation and readings, not policy instructions. The counterfactual changes only the completion reading from 73 to 75; this is coherent with a cumulative event counter that increments once per observed banging event, because two events can account for the two-count increase and no unchanged statement limits the run to one event. Neither context embeds a decision, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "refuted", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "refuted", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A15": "unknown"}, "remove_right": {"A15": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A15": "unknown"}, "negative_pair": {"A15": "refuted"}, "negative_sentence": {"A15": "unknown"}, "positive_pair": {"A15": "supported"}, "right": {"A15": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a factual relationship; interval-bounded and uniqueness propositions remain atomic despite quantifying over times or tests. A15 is a factual observation about the banging count, not a policy conclusion. The base and counter assignments differ only on A15 and are realizable: a spin may be balanced yet still produce banging. Policy evidence correctly preserves the resident’s state-originating, case-specific reuse condition, while the general governing readiness rule remains automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that inspector P completed the resident’s specified test in the relevant interval: an empty, balanced spin with zero banging. This successfully fulfills the express reuse condition, so readiness follows. The additional incident and uniqueness conditions are unnecessary but do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "A15 being refuted entails that test R had a nonzero banging count, so it did not satisfy the required no-banging test. A16 excludes any other completed inspector spin test in the relevant interval, preventing a competing successful test. Therefore the resident’s express condition remains unfulfilled and the false target follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During the period covered by the evidence through readiness decision D, incident E was the washer's only banging incident."}, {"id": "A2", "statement": "The washer contained exactly one item during incident E."}, {"id": "A3", "statement": "The item in the washer was off-center during incident E."}, {"id": "A4", "statement": "No water leaked from the washer during or after incident E through readiness decision D."}, {"id": "A5", "statement": "No odor was detected from the washer during or after incident E through readiness decision D."}, {"id": "A6", "statement": "No scraping was detected from the washer during or after incident E through readiness decision D."}, {"id": "A7", "statement": "Inspection of the washer after incident E detected no looseness."}, {"id": "A8", "statement": "Inspection of the washer after incident E detected no damage."}, {"id": "A9", "statement": "Person P was an inspector at the time of test run R."}, {"id": "A10", "statement": "Person P performed test run R on the washer."}, {"id": "A11", "statement": "Test run R was completed."}, {"id": "A12", "statement": "Test run R occurred after the resident stated the reuse condition and before readiness decision D."}, {"id": "A13", "statement": "The washer drum contained zero items during test run R."}, {"id": "A14", "statement": "The washer's spin was balanced during test run R."}, {"id": "A15", "statement": "The observed banging-event count during test run R was zero."}, {"id": "A16", "statement": "Test run R was the only washer spin test completed by any inspector after the resident stated the reuse condition and before readiness decision D."}], "base_state_json": "[{\"speaker\":\"Resident\",\"text\":\"The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging.\"},{\"speaker\":\"Incident handoff\",\"text\":\"Case records through readiness decision D list incident E as the washer's only banging incident; controlled test observations are logged separately rather than as incidents. E occurred with exactly one bath mat, visibly folded off-center. No water leakage, odor, or scraping was detected during or after E through D. Post-E inspection found neither looseness nor damage.\"},{\"speaker\":\"Test counter record\",\"text\":\"At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases.\"},{\"speaker\":\"Test counter record\",\"text\":\"At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73.\"},{\"speaker\":\"Completion handoff\",\"text\":\"Person P was the assigned inspector when R occurred and personally performed and completed it. R took place after the resident stated the reuse condition and before D. The drum contained zero items, and the spin was balanced. R was the only washer spin test completed by any inspector within that interval.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "supported"}], "focus_atom": "A15", "focus_evidence": [{"path": ["2", "text"], "text": "At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases."}, {"path": ["3", "text"], "text": "At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73."}], "policy_evidence": [{"path": ["0", "text"], "text": "The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging."}], "rules": [{"justification": "The incident satisfies every stated routine load-correction qualification, and inspector P completed an empty, balanced spin without banging after the resident imposed the condition. The resident's specified test was therefore completed successfully.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}]}, {"justification": "Although the incident satisfies every stated routine load-correction qualification, the only inspector-completed spin test in the relevant interval had a nonzero banging count. Thus no competing inspector-completed test could have fulfilled the resident's requirement for a balanced empty spin without banging.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases.", "negative_left": "At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases.", "negative_right": "At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 75.", "right": "At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73."}, "verifier_independent_model": false}, "family": "scale-diverse-183-003", "id": "scale-diverse-183-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one stated condition for the resident’s intended reuse remains unfulfilled, even if the incident otherwise appears routine.", "true": "Yes — the evidence establishes readiness for the resident’s intended reuse, including fulfillment of every stated condition."}, "instructions": "Answer whether the washer is ready for the resident’s intended reuse. Routing options are cleanup, load correction, basic inspection, and deferred maintenance. Urgency is ordered: 1 = routine, 2 = prompt, 3 = urgent, 4 = immediate shutdown. Under the household policy, an isolated banging event with an off-center single item, no leak, odor, scraping, looseness, or damage routes to load correction at urgency 1. However, when the resident expressly conditions reuse on a specified test, readiness is established only after that test is completed successfully. An inspection plan or other favorable evidence does not satisfy the condition.", "type": "noul"}}, "state": [{"speaker": "Resident", "text": "The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging."}, {"speaker": "Incident handoff", "text": "Case records through readiness decision D list incident E as the washer's only banging incident; controlled test observations are logged separately rather than as incidents. E occurred with exactly one bath mat, visibly folded off-center. No water leakage, odor, or scraping was detected during or after E through D. Post-E inspection found neither looseness nor damage."}, {"speaker": "Test counter record", "text": "At the 14:00 start of test run R on 17 September 2026, the washer's cumulative banging-event counter read 73, and during R it increased by one if and only if a banging event was observed, with no resets or decreases."}, {"speaker": "Test counter record", "text": "At the 14:12 completion of test run R on 17 September 2026, the washer's cumulative banging-event counter read 75."}, {"speaker": "Completion handoff", "text": "Person P was the assigned inspector when R occurred and personally performed and completed it. R took place after the resident stated the reuse condition and before D. The drum contained zero items, and the spin was balanced. R was the only washer spin test completed by any inspector within that interval."}]}, "method": "c2d", "provenance": {"source_id": "diverse-183", "source_is_synthetic": true, "source_sha256": "e3c0610d70e581b628bc4f9c3a5638fa1feb83c55d1a41d1e45a1f8fe0e3e6ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same washer disposition policy, request, people, appliance, decision path, and relevant date/time framing while changing only the source analysis for pad P-17. Each evidence array contains two complete factual sentences. The counterfactual coherently turns P-17 into confirmation of washer-discharged water during the short rinse without contradicting the earlier finding that the reported-cycle puddle was detergent or the inspection finding of no damaged component. Neither context includes a disposition answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Devon is reviewing Inez's dated findings after Mara reported a puddle beside her washer following a spin cycle.\",\"evidence\":[\"On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17.\",\"At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as detergent from the uncapped bottle.\"],\"request\":\"Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "0"], "text": "On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17."}, {"path": ["evidence", "1"], "text": "At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as detergent from the uncapped bottle."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17.", "negative_left": "On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17.", "negative_right": "At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as water discharged by Mara's washer during the post-inspection short rinse.", "right": "At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as detergent from the uncapped bottle."}, "verifier_independent_model": false}, "family": "scale-diverse-185-001", "id": "scale-diverse-185-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Devon is reviewing Inez's dated findings after Mara reported a puddle beside her washer following a spin cycle.", "evidence": ["On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17.", "At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as detergent from the uncapped bottle."], "request": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same washer disposition policy, request, people, appliance, decision path, and relevant date/time framing while changing only the source analysis for pad P-17. Each evidence array contains two complete factual sentences. The counterfactual coherently turns P-17 into confirmation of washer-discharged water during the short rinse without contradicting the earlier finding that the reported-cycle puddle was detergent or the inspection finding of no damaged component. Neither context includes a disposition answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Devon is reviewing Inez's dated findings after Mara reported a puddle beside her washer following a spin cycle.\",\"evidence\":[\"On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17.\",\"At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as detergent from the uncapped bottle.\"],\"request\":\"Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "0"], "text": "On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17."}, {"path": ["evidence", "1"], "text": "At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as detergent from the uncapped bottle."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17.", "negative_left": "On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17.", "negative_right": "At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as water discharged by Mara's washer during the post-inspection short rinse.", "right": "At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as detergent from the uncapped bottle."}, "verifier_independent_model": false}, "family": "scale-diverse-185-001", "id": "scale-diverse-185-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Devon is reviewing Inez's dated findings after Mara reported a puddle beside her washer following a spin cycle.", "evidence": ["On 12 May 2026, the puddle beside Mara's washer after the 13:20 reported spin cycle was traced to detergent from an uncapped bottle; Inez's 14:00 inspection found no appliance-origin leak or damaged component; for her 14:30 post-inspection short rinse, sealed collector pads captured every drop of liquid entering the laundry nook from any route, and afterward only pad P-17 contained liquid; and throughout the reported cycle, inspection, and short rinse, the washer recorded zero error codes and exhibited no electrical symptom, burning symptom, uncontrolled heat, or repeated symptom other than the possible rinse leak represented by P-17.", "At 15:10 on 12 May 2026, source analysis identified the liquid on pad P-17 as water discharged by Mara's washer during the post-inspection short rinse."], "request": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same washer-disposition request and leave the original ordered criteria and instructions unchanged, without adding exceptions, priorities, or missing-evidence defaults. The two focus spans are complete factual sentences at the specified evidence paths. Changing the sampler result from zero to 17 milliliters is coherent with the tracer setup and with the unchanged wording that allows a possible leak during the post-inspection rinse; it does not contradict the earlier inspection finding. Neither context contains a disposition level, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Devon reviewed Mara's washer incident using Inez's inspection notes, instrument records, and the laboratory result.\",\"evidence\":[\"The small puddle beside the washer after the reported spin cycle was traced to detergent discharged from the uncapped bottle. Inez's inspection found no appliance-origin leak and no damaged washer component.\",\"During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook.\",\"Laboratory analysis of the sealed sampler after the rinse measured zero milliliters of water containing tracer Q.\",\"The washer recorded no error codes during either the reported cycle or the short rinse. No electrical symptom, burning symptom, or uncontrolled heat occurred during the reported cycle, inspection, or rinse.\",\"No washer symptom other than a possible leak during the short rinse repeated across the reported cycle, inspection, or rinse.\"],\"request\":\"Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "1"], "text": "During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook."}, {"path": ["evidence", "2"], "text": "Laboratory analysis of the sealed sampler after the rinse measured zero milliliters of water containing tracer Q."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook.", "negative_left": "During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook.", "negative_right": "Laboratory analysis of the sealed sampler after the rinse measured 17 milliliters of water containing tracer Q.", "right": "Laboratory analysis of the sealed sampler after the rinse measured zero milliliters of water containing tracer Q."}, "verifier_independent_model": false}, "family": "scale-diverse-185-002", "id": "scale-diverse-185-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Devon reviewed Mara's washer incident using Inez's inspection notes, instrument records, and the laboratory result.", "evidence": ["The small puddle beside the washer after the reported spin cycle was traced to detergent discharged from the uncapped bottle. Inez's inspection found no appliance-origin leak and no damaged washer component.", "During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook.", "Laboratory analysis of the sealed sampler after the rinse measured zero milliliters of water containing tracer Q.", "The washer recorded no error codes during either the reported cycle or the short rinse. No electrical symptom, burning symptom, or uncontrolled heat occurred during the reported cycle, inspection, or rinse.", "No washer symptom other than a possible leak during the short rinse repeated across the reported cycle, inspection, or rinse."], "request": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same washer-disposition request and leave the original ordered criteria and instructions unchanged, without adding exceptions, priorities, or missing-evidence defaults. The two focus spans are complete factual sentences at the specified evidence paths. Changing the sampler result from zero to 17 milliliters is coherent with the tracer setup and with the unchanged wording that allows a possible leak during the post-inspection rinse; it does not contradict the earlier inspection finding. Neither context contains a disposition level, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Devon reviewed Mara's washer incident using Inez's inspection notes, instrument records, and the laboratory result.\",\"evidence\":[\"The small puddle beside the washer after the reported spin cycle was traced to detergent discharged from the uncapped bottle. Inez's inspection found no appliance-origin leak and no damaged washer component.\",\"During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook.\",\"Laboratory analysis of the sealed sampler after the rinse measured zero milliliters of water containing tracer Q.\",\"The washer recorded no error codes during either the reported cycle or the short rinse. No electrical symptom, burning symptom, or uncontrolled heat occurred during the reported cycle, inspection, or rinse.\",\"No washer symptom other than a possible leak during the short rinse repeated across the reported cycle, inspection, or rinse.\"],\"request\":\"Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "1"], "text": "During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook."}, {"path": ["evidence", "2"], "text": "Laboratory analysis of the sealed sampler after the rinse measured zero milliliters of water containing tracer Q."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook.", "negative_left": "During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook.", "negative_right": "Laboratory analysis of the sealed sampler after the rinse measured 17 milliliters of water containing tracer Q.", "right": "Laboratory analysis of the sealed sampler after the rinse measured zero milliliters of water containing tracer Q."}, "verifier_independent_model": false}, "family": "scale-diverse-185-002", "id": "scale-diverse-185-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Devon reviewed Mara's washer incident using Inez's inspection notes, instrument records, and the laboratory result.", "evidence": ["The small puddle beside the washer after the reported spin cycle was traced to detergent discharged from the uncapped bottle. Inez's inspection found no appliance-origin leak and no damaged washer component.", "During Inez's post-inspection short rinse of Mara's washer from 14:12:00 to 14:16:00 on 8 September 2026, tracer Q was uniformly present in all water originating from the washer and absent from every other liquid, while a sealed sampler retained all water entering the laundry nook.", "Laboratory analysis of the sealed sampler after the rinse measured 17 milliliters of water containing tracer Q.", "The washer recorded no error codes during either the reported cycle or the short rinse. No electrical symptom, burning symptom, or uncontrolled heat occurred during the reported cycle, inspection, or rinse.", "No washer symptom other than a possible leak during the short rinse repeated across the reported cycle, inspection, or rinse."], "request": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all disposition criteria and instructions, while both contexts retain the request and Mara–washer decision binding. The two focus spans are complete factual sentences. The counterfactual changes only specimen R-62’s source; water emerging from the drain during the rinse is compatible with the washer having been intact and dry before the rinse, and it does not conflict with the zero-error or no-hazard observations. Neither context embeds a disposition answer, code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Evidence reconciliation note for Mara’s washer, compiled by household maintenance coordinator Devon after Inez’s inspection.\",\"evidence\":[\"Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62.\",\"Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Inez's 250-milliliter calibration beaker and never passed through Mara's washer.\",\"Chemical comparison established that the small puddle reported beside the machine after the spin cycle was detergent spilled from the uncapped bottle.\",\"Before the rinse, Inez found no physical evidence of an appliance-origin leak and identified no damaged washer component; the hoses, seal, underside, and floor beneath the machine were intact and dry.\",\"Controller exports covering both the reported cycle and the post-inspection rinse recorded zero error codes.\",\"Records and direct observations from the reported cycle, inspection, and rinse documented no electrical symptom, burning symptom, or uncontrolled heat.\",\"No washer symptom other than the investigated possible leak recurred during any of those events.\"],\"request\":\"Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "0"], "text": "Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62."}, {"path": ["evidence", "1"], "text": "Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Inez's 250-milliliter calibration beaker and never passed through Mara's washer."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62.", "negative_left": "Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62.", "negative_right": "Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Mara's washer's drain outlet.", "right": "Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Inez's 250-milliliter calibration beaker and never passed through Mara's washer."}, "verifier_independent_model": false}, "family": "scale-diverse-185-006", "id": "scale-diverse-185-006-base", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Evidence reconciliation note for Mara’s washer, compiled by household maintenance coordinator Devon after Inez’s inspection.", "evidence": ["Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62.", "Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Inez's 250-milliliter calibration beaker and never passed through Mara's washer.", "Chemical comparison established that the small puddle reported beside the machine after the spin cycle was detergent spilled from the uncapped bottle.", "Before the rinse, Inez found no physical evidence of an appliance-origin leak and identified no damaged washer component; the hoses, seal, underside, and floor beneath the machine were intact and dry.", "Controller exports covering both the reported cycle and the post-inspection rinse recorded zero error codes.", "Records and direct observations from the reported cycle, inspection, and rinse documented no electrical symptom, burning symptom, or uncontrolled heat.", "No washer symptom other than the investigated possible leak recurred during any of those events."], "request": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all disposition criteria and instructions, while both contexts retain the request and Mara–washer decision binding. The two focus spans are complete factual sentences. The counterfactual changes only specimen R-62’s source; water emerging from the drain during the rinse is compatible with the washer having been intact and dry before the rinse, and it does not conflict with the zero-error or no-hazard observations. Neither context embeds a disposition answer, code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Evidence reconciliation note for Mara’s washer, compiled by household maintenance coordinator Devon after Inez’s inspection.\",\"evidence\":[\"Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62.\",\"Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Inez's 250-milliliter calibration beaker and never passed through Mara's washer.\",\"Chemical comparison established that the small puddle reported beside the machine after the spin cycle was detergent spilled from the uncapped bottle.\",\"Before the rinse, Inez found no physical evidence of an appliance-origin leak and identified no damaged washer component; the hoses, seal, underside, and floor beneath the machine were intact and dry.\",\"Controller exports covering both the reported cycle and the post-inspection rinse recorded zero error codes.\",\"Records and direct observations from the reported cycle, inspection, and rinse documented no electrical symptom, burning symptom, or uncontrolled heat.\",\"No washer symptom other than the investigated possible leak recurred during any of those events.\"],\"request\":\"Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "0"], "text": "Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62."}, {"path": ["evidence", "1"], "text": "Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Inez's 250-milliliter calibration beaker and never passed through Mara's washer."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62.", "negative_left": "Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62.", "negative_right": "Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Mara's washer's drain outlet.", "right": "Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Inez's 250-milliliter calibration beaker and never passed through Mara's washer."}, "verifier_independent_model": false}, "family": "scale-diverse-185-006", "id": "scale-diverse-185-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Evidence reconciliation note for Mara’s washer, compiled by household maintenance coordinator Devon after Inez’s inspection.", "evidence": ["Synchronized floor-camera and moisture-grid records established that exactly one discrete quantity of water entered the laundry nook during Inez's 14:06–14:10 post-inspection short rinse of Mara's washer on September 12, 2026, and that quantity was catalogued as specimen R-62.", "Isotopic tracer analysis established that the water catalogued as specimen R-62 originated exclusively from Mara's washer's drain outlet.", "Chemical comparison established that the small puddle reported beside the machine after the spin cycle was detergent spilled from the uncapped bottle.", "Before the rinse, Inez found no physical evidence of an appliance-origin leak and identified no damaged washer component; the hoses, seal, underside, and floor beneath the machine were intact and dry.", "Controller exports covering both the reported cycle and the post-inspection rinse recorded zero error codes.", "Records and direct observations from the reported cycle, inspection, and rinse documented no electrical symptom, burning symptom, or uncontrolled heat.", "No washer symptom other than the investigated possible leak recurred during any of those events."], "request": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the washer incident, response-level decision scope, date/time frame, and governing response rubric without adding exceptions, priorities, or missing-evidence defaults. The two focus spans are complete factual sentences. The counterfactual changes the complete ledger from recording leakage at 08:47 to recording no leakage except at 08:19, which is consistent with the unchanged 08:19 report and does not create a duplicate-measurement contradiction. Neither context contains a gold response level, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or incident-assessment relationship rather than bundling unrelated requirements or encoding the final response level. A4 is a factual recurrence proposition, not a policy conclusion. The base and counter assignments can differ only in recurrence: two completed leakage occasions can occur without active leakage at decision time, while the counter can describe only one occasion; the remaining facts can stay fixed. Empty policy_evidence is correct because the governing rubric, evidence requirements, priorities, and exceptions all originate in the retained questions object, while the state contains only case observations that need not be preserved for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A4 establishes leakage on at least two distinct occasions, which is recurring leakage. Recurring leakage independently requires level 2, regardless of whether leakage is active at decision time.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a minor, contained incident; missing required rear-connection photographic evidence; and explicit absence of every enumerated level-2 indicator. Missing required evidence excludes level 0 and directs the case to level 1.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The reported washer incident was minor."}, {"id": "A2", "statement": "The reported washer incident was contained."}, {"id": "A3", "statement": "At the response-level decision time for the reported washer incident, the required incident-area photos lacked a view of the washer's rear connections."}, {"id": "A4", "statement": "From the start of the reported washer incident through the response-level decision time, the washer produced water leakage on at least two distinct occasions."}, {"id": "A5", "statement": "At the response-level decision time for the reported incident, the washer was actively leaking."}, {"id": "A6", "statement": "During the reported washer incident, water was near the electrical outlet serving the washer."}, {"id": "A7", "statement": "During the reported washer incident, the washer emitted smoke."}, {"id": "A8", "statement": "During the reported washer incident, the washer produced sparks."}, {"id": "A9", "statement": "During the reported washer incident, the washer emitted an electrical odor."}, {"id": "A10", "statement": "The washer's logs contained repeated fault entries associated with the reported incident."}, {"id": "A11", "statement": "The washer failed the basic inspection required before reuse after the reported incident."}], "base_state_json": "[{\"speaker\":\"Operations recorder\",\"text\":\"The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19.\"},{\"speaker\":\"Ledger custodian\",\"text\":\"The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records water leakage at 08:47.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"The incident was minor and remained contained to the laundry-room floor. At decision time, the required incident-area photos lacked a view of the washer's rear connections, and the washer was not actively leaking.\"},{\"speaker\":\"Safety inspector\",\"text\":\"Throughout the incident, there was no water near the outlet serving the washer, and the appliance emitted no smoke, sparks, or electrical odor. Review of the washer's logs found no repeated fault entries associated with the incident. The basic inspection required before reuse was completed, and the washer passed it.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["0", "text"], "text": "The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19."}, {"path": ["1", "text"], "text": "The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records water leakage at 08:47."}], "policy_evidence": [], "rules": [{"justification": "At least two distinct leakage occasions during the incident-to-decision interval establish recurring leakage, which independently requires the high-response route.", "target": "2", "when": [{"atom_id": "A4", "state": "supported"}]}, {"justification": "The incident is minor and contained, required rear-connection photographic evidence is missing, and every enumerated level-2 indicator is explicitly absent, so the moderate-response route applies.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19.", "negative_left": "The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19.", "negative_right": "The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records no water leakage except at 08:19.", "right": "The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records water leakage at 08:47."}, "verifier_independent_model": false}, "family": "scale-diverse-186-003", "id": "scale-diverse-186-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Low response: Cleanup or load correction; reuse is allowed now only when all required evidence is present, shows no ongoing leak or danger indicator, and supports a one-time spill or load-balance cause.", "1 — Moderate response: Basic inspection before reuse; apply when the incident is minor and contained with no level-2 indicator, but required evidence is missing or a simple load-related cause remains unconfirmed. Complete the inspection the same day before another cycle.", "2 — High response: Stop use and arrange urgent maintenance; apply for active or recurring leakage, water near an outlet, smoke, sparks, electrical odor, repeated fault logs, or failure of the basic inspection."], "instructions": "Assign one response-intensity level using the ordered rubric. Required evidence for immediate reuse is: incident-area photos including rear connections, the cycle log, and symptom notes covering recurrence after unloading. Missing required evidence prevents level 0. Determine the route, whether the washer may be used again, and the urgency.", "type": "score"}}, "state": [{"speaker": "Operations recorder", "text": "The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19."}, {"speaker": "Ledger custodian", "text": "The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records water leakage at 08:47."}, {"speaker": "Incident coordinator", "text": "The incident was minor and remained contained to the laundry-room floor. At decision time, the required incident-area photos lacked a view of the washer's rear connections, and the washer was not actively leaking."}, {"speaker": "Safety inspector", "text": "Throughout the incident, there was no water near the outlet serving the washer, and the appliance emitted no smoke, sparks, or electrical odor. Review of the washer's logs found no repeated fault entries associated with the incident. The basic inspection required before reuse was completed, and the washer passed it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-186", "source_is_synthetic": true, "source_sha256": "bedc0c805625693d340a70785d59687aa7dc865190604e3fc159604b114d855a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the washer incident, response-level decision scope, date/time frame, and governing response rubric without adding exceptions, priorities, or missing-evidence defaults. The two focus spans are complete factual sentences. The counterfactual changes the complete ledger from recording leakage at 08:47 to recording no leakage except at 08:19, which is consistent with the unchanged 08:19 report and does not create a duplicate-measurement contradiction. Neither context contains a gold response level, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or incident-assessment relationship rather than bundling unrelated requirements or encoding the final response level. A4 is a factual recurrence proposition, not a policy conclusion. The base and counter assignments can differ only in recurrence: two completed leakage occasions can occur without active leakage at decision time, while the counter can describe only one occasion; the remaining facts can stay fixed. Empty policy_evidence is correct because the governing rubric, evidence requirements, priorities, and exceptions all originate in the retained questions object, while the state contains only case observations that need not be preserved for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A4 establishes leakage on at least two distinct occasions, which is recurring leakage. Recurring leakage independently requires level 2, regardless of whether leakage is active at decision time.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a minor, contained incident; missing required rear-connection photographic evidence; and explicit absence of every enumerated level-2 indicator. Missing required evidence excludes level 0 and directs the case to level 1.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The reported washer incident was minor."}, {"id": "A2", "statement": "The reported washer incident was contained."}, {"id": "A3", "statement": "At the response-level decision time for the reported washer incident, the required incident-area photos lacked a view of the washer's rear connections."}, {"id": "A4", "statement": "From the start of the reported washer incident through the response-level decision time, the washer produced water leakage on at least two distinct occasions."}, {"id": "A5", "statement": "At the response-level decision time for the reported incident, the washer was actively leaking."}, {"id": "A6", "statement": "During the reported washer incident, water was near the electrical outlet serving the washer."}, {"id": "A7", "statement": "During the reported washer incident, the washer emitted smoke."}, {"id": "A8", "statement": "During the reported washer incident, the washer produced sparks."}, {"id": "A9", "statement": "During the reported washer incident, the washer emitted an electrical odor."}, {"id": "A10", "statement": "The washer's logs contained repeated fault entries associated with the reported incident."}, {"id": "A11", "statement": "The washer failed the basic inspection required before reuse after the reported incident."}], "base_state_json": "[{\"speaker\":\"Operations recorder\",\"text\":\"The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19.\"},{\"speaker\":\"Ledger custodian\",\"text\":\"The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records water leakage at 08:47.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"The incident was minor and remained contained to the laundry-room floor. At decision time, the required incident-area photos lacked a view of the washer's rear connections, and the washer was not actively leaking.\"},{\"speaker\":\"Safety inspector\",\"text\":\"Throughout the incident, there was no water near the outlet serving the washer, and the appliance emitted no smoke, sparks, or electrical odor. Review of the washer's logs found no repeated fault entries associated with the incident. The basic inspection required before reuse was completed, and the washer passed it.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["0", "text"], "text": "The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19."}, {"path": ["1", "text"], "text": "The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records water leakage at 08:47."}], "policy_evidence": [], "rules": [{"justification": "At least two distinct leakage occasions during the incident-to-decision interval establish recurring leakage, which independently requires the high-response route.", "target": "2", "when": [{"atom_id": "A4", "state": "supported"}]}, {"justification": "The incident is minor and contained, required rear-connection photographic evidence is missing, and every enumerated level-2 indicator is explicitly absent, so the moderate-response route applies.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19.", "negative_left": "The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19.", "negative_right": "The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records no water leakage except at 08:19.", "right": "The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records water leakage at 08:47."}, "verifier_independent_model": false}, "family": "scale-diverse-186-003", "id": "scale-diverse-186-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low response: Cleanup or load correction; reuse is allowed now only when all required evidence is present, shows no ongoing leak or danger indicator, and supports a one-time spill or load-balance cause.", "1 — Moderate response: Basic inspection before reuse; apply when the incident is minor and contained with no level-2 indicator, but required evidence is missing or a simple load-related cause remains unconfirmed. Complete the inspection the same day before another cycle.", "2 — High response: Stop use and arrange urgent maintenance; apply for active or recurring leakage, water near an outlet, smoke, sparks, electrical odor, repeated fault logs, or failure of the basic inspection."], "instructions": "Assign one response-intensity level using the ordered rubric. Required evidence for immediate reuse is: incident-area photos including rear connections, the cycle log, and symptom notes covering recurrence after unloading. Missing required evidence prevents level 0. Determine the route, whether the washer may be used again, and the urgency.", "type": "score"}}, "state": [{"speaker": "Operations recorder", "text": "The reported washer incident began at 08:12 on 14 May 2026, the response-level decision occurred at 09:05 that day, and the washer leaked water at 08:19."}, {"speaker": "Ledger custodian", "text": "The complete observation ledger for the washer from 08:12 through 09:05 on 14 May 2026 records no water leakage except at 08:19."}, {"speaker": "Incident coordinator", "text": "The incident was minor and remained contained to the laundry-room floor. At decision time, the required incident-area photos lacked a view of the washer's rear connections, and the washer was not actively leaking."}, {"speaker": "Safety inspector", "text": "Throughout the incident, there was no water near the outlet serving the washer, and the appliance emitted no smoke, sparks, or electrical odor. Review of the washer's logs found no repeated fault entries associated with the incident. The basic inspection required before reuse was completed, and the washer passed it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-186", "source_is_synthetic": true, "source_sha256": "bedc0c805625693d340a70785d59687aa7dc865190604e3fc159604b114d855a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the complete response rubric, required-evidence rule, scope, and scoring instructions; neither context alters or adds governing policy. Both contexts concern the same washer incident and decision interval, and the counterfactual changes only the probe activation observation from two distinct activations to one while remaining consistent with the unchanged probe specification and all other facts. The two focus-evidence spans are complete factual sentences. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or incident-assessment relationship rather than bundling unrelated requirements or encoding the final response level. A4 is a factual recurrence proposition, not a policy conclusion. The base and counter assignments can differ only in recurrence: two completed leakage occasions can occur without active leakage at decision time, while the counter can describe only one occasion; the remaining facts can stay fixed. Empty policy_evidence is correct because the governing rubric, evidence requirements, priorities, and exceptions all originate in the retained questions object, while the state contains only case observations that need not be preserved for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A4 establishes leakage on at least two distinct occasions, which is recurring leakage. Recurring leakage independently requires level 2, regardless of whether leakage is active at decision time.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a minor, contained incident; missing required rear-connection photographic evidence; and explicit absence of every enumerated level-2 indicator. Missing required evidence excludes level 0 and directs the case to level 1.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The reported washer incident was minor."}, {"id": "A2", "statement": "The reported washer incident was contained."}, {"id": "A3", "statement": "At the response-level decision time for the reported washer incident, the required incident-area photos lacked a view of the washer's rear connections."}, {"id": "A4", "statement": "From the start of the reported washer incident through the response-level decision time, the washer produced water leakage on at least two distinct occasions."}, {"id": "A5", "statement": "At the response-level decision time for the reported incident, the washer was actively leaking."}, {"id": "A6", "statement": "During the reported washer incident, water was near the electrical outlet serving the washer."}, {"id": "A7", "statement": "During the reported washer incident, the washer emitted smoke."}, {"id": "A8", "statement": "During the reported washer incident, the washer produced sparks."}, {"id": "A9", "statement": "During the reported washer incident, the washer emitted an electrical odor."}, {"id": "A10", "statement": "The washer's logs contained repeated fault entries associated with the reported incident."}, {"id": "A11", "statement": "The washer failed the basic inspection required before reuse after the reported incident."}], "base_state_json": "[{\"speaker\":\"Sensor record\",\"text\":\"The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 14:26:11 and 15:52:43 during that interval with no other activations.\"},{\"speaker\":\"Probe specification\",\"text\":\"Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event.\"},{\"speaker\":\"Incident reviewer\",\"text\":\"The incident was minor and contained. At the response-level decision time, the required incident-area photos lacked a view of the washer’s rear connections, and the washer was not actively leaking.\"},{\"speaker\":\"Safety inspector\",\"text\":\"During the incident, water was never near the electrical outlet serving the washer; the appliance emitted no smoke, sparks, or electrical odor. Its logs contained no repeated fault entries associated with the incident, and it passed the basic inspection required before reuse.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["0", "text"], "text": "The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 14:26:11 and 15:52:43 during that interval with no other activations."}, {"path": ["1", "text"], "text": "Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event."}], "policy_evidence": [], "rules": [{"justification": "At least two distinct leakage occasions during the incident-to-decision interval establish recurring leakage, which independently requires the high-response route.", "target": "2", "when": [{"atom_id": "A4", "state": "supported"}]}, {"justification": "The incident is minor and contained, required rear-connection photographic evidence is missing, and every enumerated level-2 indicator is explicitly absent, so the moderate-response route applies.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 14:26:11 and 15:52:43 during that interval with no other activations.", "negative_left": "The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 15:19:28 during that interval with no other activations.", "negative_right": "Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event.", "right": "Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event."}, "verifier_independent_model": false}, "family": "scale-diverse-186-006", "id": "scale-diverse-186-006-base", "input": {"questions": {"decision": {"criteria": ["0 — Low response: Cleanup or load correction; reuse is allowed now only when all required evidence is present, shows no ongoing leak or danger indicator, and supports a one-time spill or load-balance cause.", "1 — Moderate response: Basic inspection before reuse; apply when the incident is minor and contained with no level-2 indicator, but required evidence is missing or a simple load-related cause remains unconfirmed. Complete the inspection the same day before another cycle.", "2 — High response: Stop use and arrange urgent maintenance; apply for active or recurring leakage, water near an outlet, smoke, sparks, electrical odor, repeated fault logs, or failure of the basic inspection."], "instructions": "Assign one response-intensity level using the ordered rubric. Required evidence for immediate reuse is: incident-area photos including rear connections, the cycle log, and symptom notes covering recurrence after unloading. Missing required evidence prevents level 0. Determine the route, whether the washer may be used again, and the urgency.", "type": "score"}}, "state": [{"speaker": "Sensor record", "text": "The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 14:26:11 and 15:52:43 during that interval with no other activations."}, {"speaker": "Probe specification", "text": "Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event."}, {"speaker": "Incident reviewer", "text": "The incident was minor and contained. At the response-level decision time, the required incident-area photos lacked a view of the washer’s rear connections, and the washer was not actively leaking."}, {"speaker": "Safety inspector", "text": "During the incident, water was never near the electrical outlet serving the washer; the appliance emitted no smoke, sparks, or electrical odor. Its logs contained no repeated fault entries associated with the incident, and it passed the basic inspection required before reuse."}]}, "method": "c2d", "provenance": {"source_id": "diverse-186", "source_is_synthetic": true, "source_sha256": "bedc0c805625693d340a70785d59687aa7dc865190604e3fc159604b114d855a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the complete response rubric, required-evidence rule, scope, and scoring instructions; neither context alters or adds governing policy. Both contexts concern the same washer incident and decision interval, and the counterfactual changes only the probe activation observation from two distinct activations to one while remaining consistent with the unchanged probe specification and all other facts. The two focus-evidence spans are complete factual sentences. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or incident-assessment relationship rather than bundling unrelated requirements or encoding the final response level. A4 is a factual recurrence proposition, not a policy conclusion. The base and counter assignments can differ only in recurrence: two completed leakage occasions can occur without active leakage at decision time, while the counter can describe only one occasion; the remaining facts can stay fixed. Empty policy_evidence is correct because the governing rubric, evidence requirements, priorities, and exceptions all originate in the retained questions object, while the state contains only case observations that need not be preserved for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A4 establishes leakage on at least two distinct occasions, which is recurring leakage. Recurring leakage independently requires level 2, regardless of whether leakage is active at decision time.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a minor, contained incident; missing required rear-connection photographic evidence; and explicit absence of every enumerated level-2 indicator. Missing required evidence excludes level 0 and directs the case to level 1.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The reported washer incident was minor."}, {"id": "A2", "statement": "The reported washer incident was contained."}, {"id": "A3", "statement": "At the response-level decision time for the reported washer incident, the required incident-area photos lacked a view of the washer's rear connections."}, {"id": "A4", "statement": "From the start of the reported washer incident through the response-level decision time, the washer produced water leakage on at least two distinct occasions."}, {"id": "A5", "statement": "At the response-level decision time for the reported incident, the washer was actively leaking."}, {"id": "A6", "statement": "During the reported washer incident, water was near the electrical outlet serving the washer."}, {"id": "A7", "statement": "During the reported washer incident, the washer emitted smoke."}, {"id": "A8", "statement": "During the reported washer incident, the washer produced sparks."}, {"id": "A9", "statement": "During the reported washer incident, the washer emitted an electrical odor."}, {"id": "A10", "statement": "The washer's logs contained repeated fault entries associated with the reported incident."}, {"id": "A11", "statement": "The washer failed the basic inspection required before reuse after the reported incident."}], "base_state_json": "[{\"speaker\":\"Sensor record\",\"text\":\"The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 14:26:11 and 15:52:43 during that interval with no other activations.\"},{\"speaker\":\"Probe specification\",\"text\":\"Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event.\"},{\"speaker\":\"Incident reviewer\",\"text\":\"The incident was minor and contained. At the response-level decision time, the required incident-area photos lacked a view of the washer’s rear connections, and the washer was not actively leaking.\"},{\"speaker\":\"Safety inspector\",\"text\":\"During the incident, water was never near the electrical outlet serving the washer; the appliance emitted no smoke, sparks, or electrical odor. Its logs contained no repeated fault entries associated with the incident, and it passed the basic inspection required before reuse.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["0", "text"], "text": "The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 14:26:11 and 15:52:43 during that interval with no other activations."}, {"path": ["1", "text"], "text": "Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event."}], "policy_evidence": [], "rules": [{"justification": "At least two distinct leakage occasions during the incident-to-decision interval establish recurring leakage, which independently requires the high-response route.", "target": "2", "when": [{"atom_id": "A4", "state": "supported"}]}, {"justification": "The incident is minor and contained, required rear-connection photographic evidence is missing, and every enumerated level-2 indicator is explicitly absent, so the moderate-response route applies.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 14:26:11 and 15:52:43 during that interval with no other activations.", "negative_left": "The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 15:19:28 during that interval with no other activations.", "negative_right": "Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event.", "right": "Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event."}, "verifier_independent_model": false}, "family": "scale-diverse-186-006", "id": "scale-diverse-186-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low response: Cleanup or load correction; reuse is allowed now only when all required evidence is present, shows no ongoing leak or danger indicator, and supports a one-time spill or load-balance cause.", "1 — Moderate response: Basic inspection before reuse; apply when the incident is minor and contained with no level-2 indicator, but required evidence is missing or a simple load-related cause remains unconfirmed. Complete the inspection the same day before another cycle.", "2 — High response: Stop use and arrange urgent maintenance; apply for active or recurring leakage, water near an outlet, smoke, sparks, electrical odor, repeated fault logs, or failure of the basic inspection."], "instructions": "Assign one response-intensity level using the ordered rubric. Required evidence for immediate reuse is: incident-area photos including rear connections, the cycle log, and symptom notes covering recurrence after unloading. Missing required evidence prevents level 0. Determine the route, whether the washer may be used again, and the urgency.", "type": "score"}}, "state": [{"speaker": "Sensor record", "text": "The reported washer incident began at 14:08:00 on 12 September 2026, the response-level decision occurred at 16:35:00 that day, and continuously recorded probe P-47 activated at 15:19:28 during that interval with no other activations."}, {"speaker": "Probe specification", "text": "Throughout the interval from 14:08:00 through 16:35:00 on 12 September 2026, operational probe P-47 activated exactly once on each distinct occasion when the washer involved in the reported incident produced water leakage and never activated for any other event."}, {"speaker": "Incident reviewer", "text": "The incident was minor and contained. At the response-level decision time, the required incident-area photos lacked a view of the washer’s rear connections, and the washer was not actively leaking."}, {"speaker": "Safety inspector", "text": "During the incident, water was never near the electrical outlet serving the washer; the appliance emitted no smoke, sparks, or electrical odor. Its logs contained no repeated fault entries associated with the incident, and it passed the basic inspection required before reuse."}]}, "method": "c2d", "provenance": {"source_id": "diverse-186", "source_is_synthetic": true, "source_sha256": "bedc0c805625693d340a70785d59687aa7dc865190604e3fc159604b114d855a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric and required-evidence rules, and neither context omits or alters any additional policy from the original state. Both contexts retain the same washer, incident interval, response-level decision time, and requested decision path. The two focus-evidence spans are complete factual sentences. Changing the decision-time counter from 286 to 285 coherently changes the inferred number of leakage occasions from two to one under the unchanged counter behavior, without conflicting with the statements that the incident was contained and no leak was active at decision time. Neither full context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or incident-assessment relationship rather than bundling unrelated requirements or encoding the final response level. A4 is a factual recurrence proposition, not a policy conclusion. The base and counter assignments can differ only in recurrence: two completed leakage occasions can occur without active leakage at decision time, while the counter can describe only one occasion; the remaining facts can stay fixed. Empty policy_evidence is correct because the governing rubric, evidence requirements, priorities, and exceptions all originate in the retained questions object, while the state contains only case observations that need not be preserved for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A4 establishes leakage on at least two distinct occasions, which is recurring leakage. Recurring leakage independently requires level 2, regardless of whether leakage is active at decision time.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a minor, contained incident; missing required rear-connection photographic evidence; and explicit absence of every enumerated level-2 indicator. Missing required evidence excludes level 0 and directs the case to level 1.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The reported washer incident was minor."}, {"id": "A2", "statement": "The reported washer incident was contained."}, {"id": "A3", "statement": "At the response-level decision time for the reported washer incident, the required incident-area photos lacked a view of the washer's rear connections."}, {"id": "A4", "statement": "From the start of the reported washer incident through the response-level decision time, the washer produced water leakage on at least two distinct occasions."}, {"id": "A5", "statement": "At the response-level decision time for the reported incident, the washer was actively leaking."}, {"id": "A6", "statement": "During the reported washer incident, water was near the electrical outlet serving the washer."}, {"id": "A7", "statement": "During the reported washer incident, the washer emitted smoke."}, {"id": "A8", "statement": "During the reported washer incident, the washer produced sparks."}, {"id": "A9", "statement": "During the reported washer incident, the washer emitted an electrical odor."}, {"id": "A10", "statement": "The washer's logs contained repeated fault entries associated with the reported incident."}, {"id": "A11", "statement": "The washer failed the basic inspection required before reuse after the reported incident."}], "base_state_json": "[{\"speaker\":\"Device record custodian\",\"text\":\"Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason.\"},{\"speaker\":\"Decision-time recorder\",\"text\":\"At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 286.\"},{\"speaker\":\"Operational handoff\",\"text\":\"At the response-level decision time, the reported incident was minor and contained. The required incident-area photos lacked a view of the washer's rear connections, and the washer was not actively leaking.\"},{\"speaker\":\"Safety observer\",\"text\":\"During the incident, no water was near the electrical outlet serving the washer. The washer emitted no smoke, produced no sparks, and emitted no electrical odor.\"},{\"speaker\":\"Maintenance reviewer\",\"text\":\"The washer's logs contained no repeated fault entries associated with the incident. After the incident, the washer passed the basic inspection required before reuse.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["0", "text"], "text": "Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason."}, {"path": ["1", "text"], "text": "At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 286."}], "policy_evidence": [], "rules": [{"justification": "At least two distinct leakage occasions during the incident-to-decision interval establish recurring leakage, which independently requires the high-response route.", "target": "2", "when": [{"atom_id": "A4", "state": "supported"}]}, {"justification": "The incident is minor and contained, required rear-connection photographic evidence is missing, and every enumerated level-2 indicator is explicitly absent, so the moderate-response route applies.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason.", "negative_left": "Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason.", "negative_right": "At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 285.", "right": "At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 286."}, "verifier_independent_model": false}, "family": "scale-diverse-186-007", "id": "scale-diverse-186-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Low response: Cleanup or load correction; reuse is allowed now only when all required evidence is present, shows no ongoing leak or danger indicator, and supports a one-time spill or load-balance cause.", "1 — Moderate response: Basic inspection before reuse; apply when the incident is minor and contained with no level-2 indicator, but required evidence is missing or a simple load-related cause remains unconfirmed. Complete the inspection the same day before another cycle.", "2 — High response: Stop use and arrange urgent maintenance; apply for active or recurring leakage, water near an outlet, smoke, sparks, electrical odor, repeated fault logs, or failure of the basic inspection."], "instructions": "Assign one response-intensity level using the ordered rubric. Required evidence for immediate reuse is: incident-area photos including rear connections, the cycle log, and symptom notes covering recurrence after unloading. Missing required evidence prevents level 0. Determine the route, whether the washer may be used again, and the urgency.", "type": "score"}}, "state": [{"speaker": "Device record custodian", "text": "Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason."}, {"speaker": "Decision-time recorder", "text": "At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 286."}, {"speaker": "Operational handoff", "text": "At the response-level decision time, the reported incident was minor and contained. The required incident-area photos lacked a view of the washer's rear connections, and the washer was not actively leaking."}, {"speaker": "Safety observer", "text": "During the incident, no water was near the electrical outlet serving the washer. The washer emitted no smoke, produced no sparks, and emitted no electrical odor."}, {"speaker": "Maintenance reviewer", "text": "The washer's logs contained no repeated fault entries associated with the incident. After the incident, the washer passed the basic inspection required before reuse."}]}, "method": "c2d", "provenance": {"source_id": "diverse-186", "source_is_synthetic": true, "source_sha256": "bedc0c805625693d340a70785d59687aa7dc865190604e3fc159604b114d855a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric and required-evidence rules, and neither context omits or alters any additional policy from the original state. Both contexts retain the same washer, incident interval, response-level decision time, and requested decision path. The two focus-evidence spans are complete factual sentences. Changing the decision-time counter from 286 to 285 coherently changes the inferred number of leakage occasions from two to one under the unchanged counter behavior, without conflicting with the statements that the incident was contained and no leak was active at decision time. Neither full context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or incident-assessment relationship rather than bundling unrelated requirements or encoding the final response level. A4 is a factual recurrence proposition, not a policy conclusion. The base and counter assignments can differ only in recurrence: two completed leakage occasions can occur without active leakage at decision time, while the counter can describe only one occasion; the remaining facts can stay fixed. Empty policy_evidence is correct because the governing rubric, evidence requirements, priorities, and exceptions all originate in the retained questions object, while the state contains only case observations that need not be preserved for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A4 establishes leakage on at least two distinct occasions, which is recurring leakage. Recurring leakage independently requires level 2, regardless of whether leakage is active at decision time.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a minor, contained incident; missing required rear-connection photographic evidence; and explicit absence of every enumerated level-2 indicator. Missing required evidence excludes level 0 and directs the case to level 1.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The reported washer incident was minor."}, {"id": "A2", "statement": "The reported washer incident was contained."}, {"id": "A3", "statement": "At the response-level decision time for the reported washer incident, the required incident-area photos lacked a view of the washer's rear connections."}, {"id": "A4", "statement": "From the start of the reported washer incident through the response-level decision time, the washer produced water leakage on at least two distinct occasions."}, {"id": "A5", "statement": "At the response-level decision time for the reported incident, the washer was actively leaking."}, {"id": "A6", "statement": "During the reported washer incident, water was near the electrical outlet serving the washer."}, {"id": "A7", "statement": "During the reported washer incident, the washer emitted smoke."}, {"id": "A8", "statement": "During the reported washer incident, the washer produced sparks."}, {"id": "A9", "statement": "During the reported washer incident, the washer emitted an electrical odor."}, {"id": "A10", "statement": "The washer's logs contained repeated fault entries associated with the reported incident."}, {"id": "A11", "statement": "The washer failed the basic inspection required before reuse after the reported incident."}], "base_state_json": "[{\"speaker\":\"Device record custodian\",\"text\":\"Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason.\"},{\"speaker\":\"Decision-time recorder\",\"text\":\"At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 286.\"},{\"speaker\":\"Operational handoff\",\"text\":\"At the response-level decision time, the reported incident was minor and contained. The required incident-area photos lacked a view of the washer's rear connections, and the washer was not actively leaking.\"},{\"speaker\":\"Safety observer\",\"text\":\"During the incident, no water was near the electrical outlet serving the washer. The washer emitted no smoke, produced no sparks, and emitted no electrical odor.\"},{\"speaker\":\"Maintenance reviewer\",\"text\":\"The washer's logs contained no repeated fault entries associated with the incident. After the incident, the washer passed the basic inspection required before reuse.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["0", "text"], "text": "Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason."}, {"path": ["1", "text"], "text": "At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 286."}], "policy_evidence": [], "rules": [{"justification": "At least two distinct leakage occasions during the incident-to-decision interval establish recurring leakage, which independently requires the high-response route.", "target": "2", "when": [{"atom_id": "A4", "state": "supported"}]}, {"justification": "The incident is minor and contained, required rear-connection photographic evidence is missing, and every enumerated level-2 indicator is explicitly absent, so the moderate-response route applies.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason.", "negative_left": "Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason.", "negative_right": "At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 285.", "right": "At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 286."}, "verifier_independent_model": false}, "family": "scale-diverse-186-007", "id": "scale-diverse-186-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low response: Cleanup or load correction; reuse is allowed now only when all required evidence is present, shows no ongoing leak or danger indicator, and supports a one-time spill or load-balance cause.", "1 — Moderate response: Basic inspection before reuse; apply when the incident is minor and contained with no level-2 indicator, but required evidence is missing or a simple load-related cause remains unconfirmed. Complete the inspection the same day before another cycle.", "2 — High response: Stop use and arrange urgent maintenance; apply for active or recurring leakage, water near an outlet, smoke, sparks, electrical odor, repeated fault logs, or failure of the basic inspection."], "instructions": "Assign one response-intensity level using the ordered rubric. Required evidence for immediate reuse is: incident-area photos including rear connections, the cycle log, and symptom notes covering recurrence after unloading. Missing required evidence prevents level 0. Determine the route, whether the washer may be used again, and the urgency.", "type": "score"}}, "state": [{"speaker": "Device record custodian", "text": "Immediately before the reported washer incident began at 14:06:00 on 8 May 2026, the washer's sealed leakage counter read 284; throughout the interval ending at the response-level decision time, the counter neither reset nor decreased and advanced by exactly one unit for each distinct occasion of water leakage from the washer and for no other reason."}, {"speaker": "Decision-time recorder", "text": "At the response-level decision time of 14:44:00 on 8 May 2026, the washer's sealed leakage counter read 285."}, {"speaker": "Operational handoff", "text": "At the response-level decision time, the reported incident was minor and contained. The required incident-area photos lacked a view of the washer's rear connections, and the washer was not actively leaking."}, {"speaker": "Safety observer", "text": "During the incident, no water was near the electrical outlet serving the washer. The washer emitted no smoke, produced no sparks, and emitted no electrical odor."}, {"speaker": "Maintenance reviewer", "text": "The washer's logs contained no repeated fault entries associated with the incident. After the incident, the washer passed the basic inspection required before reuse."}]}, "method": "c2d", "provenance": {"source_id": "diverse-186", "source_is_synthetic": true, "source_sha256": "bedc0c805625693d340a70785d59687aa7dc865190604e3fc159604b114d855a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original recipe requirements, kitchen policy, role assignments, request scope, and relevant glaze interval bindings. The two focus spans are complete factual sentences; the counterfactual changes only the recorded removal time, yielding a coherent shorter interval without duplicate or contradictory measurements, and neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Final reconciliation confirmed that the rice for all four bowls was fully cooked and the tofu was browned on every side listed by the recipe. Every required ingredient was present, and inspection found every food item in the batch uncontaminated. All four bowls were portioned, and every portion designated for refrigerated storage was labeled.\",\"The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat.\",\"The glaze's recorded removal from heat occurred at 18:44:31 on 14 September 2026.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "1"], "text": "The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat."}, {"path": ["evidence", "2"], "text": "The glaze's recorded removal from heat occurred at 18:44:31 on 14 September 2026."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat.", "negative_left": "The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat.", "negative_right": "The glaze's recorded removal from heat occurred at 18:43:49 on 14 September 2026.", "right": "The glaze's recorded removal from heat occurred at 18:44:31 on 14 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-187-002", "id": "scale-diverse-187-002-base", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Final reconciliation confirmed that the rice for all four bowls was fully cooked and the tofu was browned on every side listed by the recipe. Every required ingredient was present, and inspection found every food item in the batch uncontaminated. All four bowls were portioned, and every portion designated for refrigerated storage was labeled.", "The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat.", "The glaze's recorded removal from heat occurred at 18:44:31 on 14 September 2026."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_5_ready"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original recipe requirements, kitchen policy, role assignments, request scope, and relevant glaze interval bindings. The two focus spans are complete factual sentences; the counterfactual changes only the recorded removal time, yielding a coherent shorter interval without duplicate or contradictory measurements, and neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Final reconciliation confirmed that the rice for all four bowls was fully cooked and the tofu was browned on every side listed by the recipe. Every required ingredient was present, and inspection found every food item in the batch uncontaminated. All four bowls were portioned, and every portion designated for refrigerated storage was labeled.\",\"The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat.\",\"The glaze's recorded removal from heat occurred at 18:44:31 on 14 September 2026.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "1"], "text": "The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat."}, {"path": ["evidence", "2"], "text": "The glaze's recorded removal from heat occurred at 18:44:31 on 14 September 2026."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat.", "negative_left": "The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat.", "negative_right": "The glaze's recorded removal from heat occurred at 18:43:49 on 14 September 2026.", "right": "The glaze's recorded removal from heat occurred at 18:44:31 on 14 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-187-002", "id": "scale-diverse-187-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Final reconciliation confirmed that the rice for all four bowls was fully cooked and the tofu was browned on every side listed by the recipe. Every required ingredient was present, and inspection found every food item in the batch uncontaminated. All four bowls were portioned, and every portion designated for refrigerated storage was labeled.", "The sole observed continuous bubbling interval for the glaze before serving began at 18:42:17 on 14 September 2026 and ended at the glaze's recorded removal from heat.", "The glaze's recorded removal from heat occurred at 18:43:49 on 14 September 2026."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_2_major_correction"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same recipe requirements, workflow roles, request, and timed-heating rule without adding exceptions or defaults. The two focus spans are complete factual sentences, and the sole counterfactual change shortens the observed bubbling interval from 2 minutes 17 seconds to 1 minute 28 seconds while remaining consistent with the unchanged serving time and other evidence. Neither context contains an explicit answer, output instruction, answer code, proposition ID, or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Final handoff inspection recorded fully cooked rice and tofu browned on every side listed by the recipe.\",\"During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14.\",\"The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:05:31, before the bowls were served at 19:06:00.\",\"Ingredient checks confirmed that every ingredient required for the four bowls is present. Safety inspection found every food item in the batch uncontaminated.\",\"All four teriyaki tofu bowls are portioned, and every portion designated for refrigerated storage from the batch is labeled.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "1"], "text": "During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14."}, {"path": ["evidence", "2"], "text": "The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:05:31, before the bowls were served at 19:06:00."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14.", "negative_left": "During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14.", "negative_right": "The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:04:42, before the bowls were served at 19:06:00.", "right": "The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:05:31, before the bowls were served at 19:06:00."}, "verifier_independent_model": false}, "family": "scale-diverse-187-003", "id": "scale-diverse-187-003-base", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Final handoff inspection recorded fully cooked rice and tofu browned on every side listed by the recipe.", "During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14.", "The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:05:31, before the bowls were served at 19:06:00.", "Ingredient checks confirmed that every ingredient required for the four bowls is present. Safety inspection found every food item in the batch uncontaminated.", "All four teriyaki tofu bowls are portioned, and every portion designated for refrigerated storage from the batch is labeled."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_5_ready"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same recipe requirements, workflow roles, request, and timed-heating rule without adding exceptions or defaults. The two focus spans are complete factual sentences, and the sole counterfactual change shortens the observed bubbling interval from 2 minutes 17 seconds to 1 minute 28 seconds while remaining consistent with the unchanged serving time and other evidence. Neither context contains an explicit answer, output instruction, answer code, proposition ID, or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Final handoff inspection recorded fully cooked rice and tofu browned on every side listed by the recipe.\",\"During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14.\",\"The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:05:31, before the bowls were served at 19:06:00.\",\"Ingredient checks confirmed that every ingredient required for the four bowls is present. Safety inspection found every food item in the batch uncontaminated.\",\"All four teriyaki tofu bowls are portioned, and every portion designated for refrigerated storage from the batch is labeled.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "1"], "text": "During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14."}, {"path": ["evidence", "2"], "text": "The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:05:31, before the bowls were served at 19:06:00."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14.", "negative_left": "During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14.", "negative_right": "The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:04:42, before the bowls were served at 19:06:00.", "right": "The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:05:31, before the bowls were served at 19:06:00."}, "verifier_independent_model": false}, "family": "scale-diverse-187-003", "id": "scale-diverse-187-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Final handoff inspection recorded fully cooked rice and tofu browned on every side listed by the recipe.", "During the operational handoff for the four teriyaki tofu bowls, the synchronized kitchen timer marked the start of the observed continuous bubbling interval for the glaze at 19:03:14.", "The synchronized kitchen timer marked the end of that observed continuous bubbling interval at 19:04:42, before the bowls were served at 19:06:00.", "Ingredient checks confirmed that every ingredient required for the four bowls is present. Safety inspection found every food item in the batch uncontaminated.", "All four teriyaki tofu bowls are portioned, and every portion designated for refrigerated storage from the batch is labeled."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_2_major_correction"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the inclusive cooking, portioning, cooling, and conditional-storage policy without adding exceptions or defaults, and they preserve the same food items, containers, two-hour cooling endpoint, and workflow-evaluation scope. The two evidence spans are complete factual sentences. The counterfactual coherently changes only Flint-42's storage location; it does not conflict with the complete container inventory or the transfer, shallowness, and labeling facts. Neither context embeds a workflow-level answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Home cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Evaluation observer\",\"text\":\"At the workflow evaluation time, the chicken measured exactly 74°C. Exactly two dinner portions of rice had been prepared, and each portion weighed exactly 350 g.\"},{\"speaker\":\"Cooling monitor\",\"text\":\"The leftover rice was at 60°C when cooling began. At 8:00 p.m., two hours after cooling started, it had reached exactly 21°C.\"},{\"speaker\":\"Inventory observer\",\"text\":\"At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42.\"},{\"speaker\":\"Storage observer\",\"text\":\"At the workflow evaluation time, containers Cedar-17 and Flint-42 were both inside the refrigerator.\"},{\"speaker\":\"Transfer observer\",\"text\":\"By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Cedar-17 and Flint-42 were each shallow and labeled.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["4", "text"], "text": "At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42."}, {"path": ["5", "text"], "text": "At the workflow evaluation time, containers Cedar-17 and Flint-42 were both inside the refrigerator."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42.", "negative_left": "At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42.", "negative_right": "At the workflow evaluation time, container Cedar-17 was inside the refrigerator and container Flint-42 was outside the refrigerator on the preparation counter.", "right": "At the workflow evaluation time, containers Cedar-17 and Flint-42 were both inside the refrigerator."}, "verifier_independent_model": false}, "family": "scale-diverse-188-001", "id": "scale-diverse-188-001-base", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Meal planner", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Home cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Evaluation observer", "text": "At the workflow evaluation time, the chicken measured exactly 74°C. Exactly two dinner portions of rice had been prepared, and each portion weighed exactly 350 g."}, {"speaker": "Cooling monitor", "text": "The leftover rice was at 60°C when cooling began. At 8:00 p.m., two hours after cooling started, it had reached exactly 21°C."}, {"speaker": "Inventory observer", "text": "At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42."}, {"speaker": "Storage observer", "text": "At the workflow evaluation time, containers Cedar-17 and Flint-42 were both inside the refrigerator."}, {"speaker": "Transfer observer", "text": "By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Cedar-17 and Flint-42 were each shallow and labeled."}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_3_COMPLETE"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the inclusive cooking, portioning, cooling, and conditional-storage policy without adding exceptions or defaults, and they preserve the same food items, containers, two-hour cooling endpoint, and workflow-evaluation scope. The two evidence spans are complete factual sentences. The counterfactual coherently changes only Flint-42's storage location; it does not conflict with the complete container inventory or the transfer, shallowness, and labeling facts. Neither context embeds a workflow-level answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Home cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Evaluation observer\",\"text\":\"At the workflow evaluation time, the chicken measured exactly 74°C. Exactly two dinner portions of rice had been prepared, and each portion weighed exactly 350 g.\"},{\"speaker\":\"Cooling monitor\",\"text\":\"The leftover rice was at 60°C when cooling began. At 8:00 p.m., two hours after cooling started, it had reached exactly 21°C.\"},{\"speaker\":\"Inventory observer\",\"text\":\"At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42.\"},{\"speaker\":\"Storage observer\",\"text\":\"At the workflow evaluation time, containers Cedar-17 and Flint-42 were both inside the refrigerator.\"},{\"speaker\":\"Transfer observer\",\"text\":\"By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Cedar-17 and Flint-42 were each shallow and labeled.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["4", "text"], "text": "At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42."}, {"path": ["5", "text"], "text": "At the workflow evaluation time, containers Cedar-17 and Flint-42 were both inside the refrigerator."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42.", "negative_left": "At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42.", "negative_right": "At the workflow evaluation time, container Cedar-17 was inside the refrigerator and container Flint-42 was outside the refrigerator on the preparation counter.", "right": "At the workflow evaluation time, containers Cedar-17 and Flint-42 were both inside the refrigerator."}, "verifier_independent_model": false}, "family": "scale-diverse-188-001", "id": "scale-diverse-188-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Meal planner", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Home cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Evaluation observer", "text": "At the workflow evaluation time, the chicken measured exactly 74°C. Exactly two dinner portions of rice had been prepared, and each portion weighed exactly 350 g."}, {"speaker": "Cooling monitor", "text": "The leftover rice was at 60°C when cooling began. At 8:00 p.m., two hours after cooling started, it had reached exactly 21°C."}, {"speaker": "Inventory observer", "text": "At the workflow evaluation time, the complete set of containers holding any of the leftover rice consisted of containers Cedar-17 and Flint-42."}, {"speaker": "Storage observer", "text": "At the workflow evaluation time, container Cedar-17 was inside the refrigerator and container Flint-42 was outside the refrigerator on the preparation counter."}, {"speaker": "Transfer observer", "text": "By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Cedar-17 and Flint-42 were each shallow and labeled."}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_2_READY_CORRECTION_DUE"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the inclusive cooking, portioning, cooling, and conditional storage requirements without adding exceptions or altered priorities. They evaluate the same leftover-rice workflow at the same evaluation-time binding; the counterfactual changes only D42’s location, which is coherent with the unchanged transfer, inventory, shallowness, and labeling facts. The two evidence spans are complete factual sentences, and neither context contains a workflow-level answer, code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Home cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Kitchen recorder\",\"text\":\"At the workflow evaluation time, the chicken thermometer read exactly 74°C. Exactly two dinner portions of rice had been prepared, and each portion weighed exactly 350 g.\"},{\"speaker\":\"Cooling and packing auditor\",\"text\":\"The leftover rice began cooling at 6:00 p.m. at 60°C and measured exactly 21°C at 8:00 p.m. By the workflow evaluation time, the cooling bowl was empty and all the leftover rice had been transferred into containers. Every container holding it was shallow and labeled.\"},{\"speaker\":\"Inventory witness\",\"text\":\"At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42.\"},{\"speaker\":\"Location witness\",\"text\":\"At the workflow evaluation time, containers C17 and D42 were both inside the refrigerator.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["4", "text"], "text": "At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42."}, {"path": ["5", "text"], "text": "At the workflow evaluation time, containers C17 and D42 were both inside the refrigerator."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42.", "negative_left": "At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42.", "negative_right": "At the workflow evaluation time, container C17 was inside the refrigerator, while container D42 was outside the refrigerator on the preparation counter.", "right": "At the workflow evaluation time, containers C17 and D42 were both inside the refrigerator."}, "verifier_independent_model": false}, "family": "scale-diverse-188-002", "id": "scale-diverse-188-002-base", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Meal planner", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Home cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Kitchen recorder", "text": "At the workflow evaluation time, the chicken thermometer read exactly 74°C. Exactly two dinner portions of rice had been prepared, and each portion weighed exactly 350 g."}, {"speaker": "Cooling and packing auditor", "text": "The leftover rice began cooling at 6:00 p.m. at 60°C and measured exactly 21°C at 8:00 p.m. By the workflow evaluation time, the cooling bowl was empty and all the leftover rice had been transferred into containers. Every container holding it was shallow and labeled."}, {"speaker": "Inventory witness", "text": "At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42."}, {"speaker": "Location witness", "text": "At the workflow evaluation time, containers C17 and D42 were both inside the refrigerator."}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_3_COMPLETE"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the inclusive cooking, portioning, cooling, and conditional storage requirements without adding exceptions or altered priorities. They evaluate the same leftover-rice workflow at the same evaluation-time binding; the counterfactual changes only D42’s location, which is coherent with the unchanged transfer, inventory, shallowness, and labeling facts. The two evidence spans are complete factual sentences, and neither context contains a workflow-level answer, code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Home cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Kitchen recorder\",\"text\":\"At the workflow evaluation time, the chicken thermometer read exactly 74°C. Exactly two dinner portions of rice had been prepared, and each portion weighed exactly 350 g.\"},{\"speaker\":\"Cooling and packing auditor\",\"text\":\"The leftover rice began cooling at 6:00 p.m. at 60°C and measured exactly 21°C at 8:00 p.m. By the workflow evaluation time, the cooling bowl was empty and all the leftover rice had been transferred into containers. Every container holding it was shallow and labeled.\"},{\"speaker\":\"Inventory witness\",\"text\":\"At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42.\"},{\"speaker\":\"Location witness\",\"text\":\"At the workflow evaluation time, containers C17 and D42 were both inside the refrigerator.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["4", "text"], "text": "At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42."}, {"path": ["5", "text"], "text": "At the workflow evaluation time, containers C17 and D42 were both inside the refrigerator."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42.", "negative_left": "At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42.", "negative_right": "At the workflow evaluation time, container C17 was inside the refrigerator, while container D42 was outside the refrigerator on the preparation counter.", "right": "At the workflow evaluation time, containers C17 and D42 were both inside the refrigerator."}, "verifier_independent_model": false}, "family": "scale-diverse-188-002", "id": "scale-diverse-188-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Meal planner", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Home cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Kitchen recorder", "text": "At the workflow evaluation time, the chicken thermometer read exactly 74°C. Exactly two dinner portions of rice had been prepared, and each portion weighed exactly 350 g."}, {"speaker": "Cooling and packing auditor", "text": "The leftover rice began cooling at 6:00 p.m. at 60°C and measured exactly 21°C at 8:00 p.m. By the workflow evaluation time, the cooling bowl was empty and all the leftover rice had been transferred into containers. Every container holding it was shallow and labeled."}, {"speaker": "Inventory witness", "text": "At the workflow evaluation time, the only containers holding the leftover rice were containers C17 and D42."}, {"speaker": "Location witness", "text": "At the workflow evaluation time, container C17 was inside the refrigerator, while container D42 was outside the refrigerator on the preparation counter."}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_2_READY_CORRECTION_DUE"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the cooking, portioning, cooling, inclusive-boundary, conditional-discard, and storage requirements without adding exceptions or altered priorities. The same leftover-rice workflow is evaluated at the same relevant cooling/evaluation period, and the question still requests the overall workflow level. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the locations of the two identified containers: H17 is outside the refrigerator while H24 remains inside, without conflicting with the container count or the statements that all leftovers are in shallow labeled containers. Neither context contains a gold level, answer code, rule table, proposition identifier, label rationale, or output directive.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Home cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Shift verifier\",\"text\":\"At the workflow evaluation time, the calibrated probe showed the chicken at 74°C. Exactly two dinner portions of rice had been prepared, and each weighed exactly 350 g.\"},{\"speaker\":\"Cooling log reviewer\",\"text\":\"The leftover rice began cooling at 6:00 p.m. from 60°C. At 8:00 p.m. it measured 21°C, reaching 21°C or lower within two hours.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Every container holding it was shallow and labeled.\"},{\"speaker\":\"Inventory recorder\",\"text\":\"At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24.\"},{\"speaker\":\"Location checker\",\"text\":\"At the workflow evaluation time, canisters H17 and H24 stood inside the refrigerator.\"},{\"speaker\":\"Supervisor\",\"text\":\"Which overall workflow level applies?\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["5", "text"], "text": "At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24."}, {"path": ["6", "text"], "text": "At the workflow evaluation time, canisters H17 and H24 stood inside the refrigerator."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24.", "negative_left": "At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24.", "negative_right": "At the workflow evaluation time, canister H17 stood on the staging cart outside the refrigerator, while canister H24 stood inside the refrigerator.", "right": "At the workflow evaluation time, canisters H17 and H24 stood inside the refrigerator."}, "verifier_independent_model": false}, "family": "scale-diverse-188-003", "id": "scale-diverse-188-003-base", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Meal planner", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Home cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Shift verifier", "text": "At the workflow evaluation time, the calibrated probe showed the chicken at 74°C. Exactly two dinner portions of rice had been prepared, and each weighed exactly 350 g."}, {"speaker": "Cooling log reviewer", "text": "The leftover rice began cooling at 6:00 p.m. from 60°C. At 8:00 p.m. it measured 21°C, reaching 21°C or lower within two hours."}, {"speaker": "Handoff recorder", "text": "By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Every container holding it was shallow and labeled."}, {"speaker": "Inventory recorder", "text": "At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24."}, {"speaker": "Location checker", "text": "At the workflow evaluation time, canisters H17 and H24 stood inside the refrigerator."}, {"speaker": "Supervisor", "text": "Which overall workflow level applies?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_3_COMPLETE"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the cooking, portioning, cooling, inclusive-boundary, conditional-discard, and storage requirements without adding exceptions or altered priorities. The same leftover-rice workflow is evaluated at the same relevant cooling/evaluation period, and the question still requests the overall workflow level. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the locations of the two identified containers: H17 is outside the refrigerator while H24 remains inside, without conflicting with the container count or the statements that all leftovers are in shallow labeled containers. Neither context contains a gold level, answer code, rule table, proposition identifier, label rationale, or output directive.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Home cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Shift verifier\",\"text\":\"At the workflow evaluation time, the calibrated probe showed the chicken at 74°C. Exactly two dinner portions of rice had been prepared, and each weighed exactly 350 g.\"},{\"speaker\":\"Cooling log reviewer\",\"text\":\"The leftover rice began cooling at 6:00 p.m. from 60°C. At 8:00 p.m. it measured 21°C, reaching 21°C or lower within two hours.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Every container holding it was shallow and labeled.\"},{\"speaker\":\"Inventory recorder\",\"text\":\"At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24.\"},{\"speaker\":\"Location checker\",\"text\":\"At the workflow evaluation time, canisters H17 and H24 stood inside the refrigerator.\"},{\"speaker\":\"Supervisor\",\"text\":\"Which overall workflow level applies?\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["5", "text"], "text": "At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24."}, {"path": ["6", "text"], "text": "At the workflow evaluation time, canisters H17 and H24 stood inside the refrigerator."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24.", "negative_left": "At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24.", "negative_right": "At the workflow evaluation time, canister H17 stood on the staging cart outside the refrigerator, while canister H24 stood inside the refrigerator.", "right": "At the workflow evaluation time, canisters H17 and H24 stood inside the refrigerator."}, "verifier_independent_model": false}, "family": "scale-diverse-188-003", "id": "scale-diverse-188-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Meal planner", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Home cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Shift verifier", "text": "At the workflow evaluation time, the calibrated probe showed the chicken at 74°C. Exactly two dinner portions of rice had been prepared, and each weighed exactly 350 g."}, {"speaker": "Cooling log reviewer", "text": "The leftover rice began cooling at 6:00 p.m. from 60°C. At 8:00 p.m. it measured 21°C, reaching 21°C or lower within two hours."}, {"speaker": "Handoff recorder", "text": "By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Every container holding it was shallow and labeled."}, {"speaker": "Inventory recorder", "text": "At the workflow evaluation time, the leftover rice was held in exactly two containers, canisters H17 and H24."}, {"speaker": "Location checker", "text": "At the workflow evaluation time, canister H17 stood on the staging cart outside the refrigerator, while canister H24 stood inside the refrigerator."}, {"speaker": "Supervisor", "text": "Which overall workflow level applies?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_2_READY_CORRECTION_DUE"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the recipe requirements, the one-low-piece kitchen rule, the current-meal readiness scope, and the latest-update timing. The two evidence spans are complete factual sentences. The counterfactual changes only the checked piece’s seal and therefore its membership in the finalized Q-17 batch; it does not create a duplicate or contradictory measurement. Neither context states a readiness label, answer code, proposition ID, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified completion or exclusion claims remain atomic rather than bundling a final classification. The focus atom is the factual membership relation between the thick piece and the meal’s chicken batch. Both assignments are realizable while changing only that membership fact: the 70°C observation can concern either a batch piece or a non-batch piece, with the other listed facts unchanged. The policy evidence correctly preserves the state-originating recipe threshold and batch-return rule; the readiness rubric and temporal instructions need not be repeated because they remain in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the 70°C piece belongs to the meal’s chicken batch. Because 70°C is below the required 74°C and the state’s kitchen rule sends the entire batch back for cooking when any piece is low, a corrective cooking step remains. The meal therefore cannot be at Level 3 or 4.", "rule_index": 0, "sound": true}, {"reason": "Refuting batch membership makes the 70°C observation inapplicable to this meal’s chicken batch. The remaining conditions establish that all batch pieces reached 74°C, the other required cooking steps and setup are complete, portioning is complete, and no other corrective step remains. Under the stated policy, the low-piece cooking correction is not triggered, so these conditions are sufficient for readiness.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The newly checked thick piece at 6:43 p.m. belongs to the chicken batch for this meal."}, {"id": "a2", "statement": "The newly checked thick piece had a measured temperature of 70°C at 6:43 p.m."}, {"id": "a3", "statement": "At the latest update, every chicken piece in the meal's batch other than the newly checked thick piece has reached 74°C."}, {"id": "a4", "statement": "At the latest update, the rice has completed all cooking steps required for serving."}, {"id": "a5", "statement": "At the latest update, the broccoli has completed all cooking steps required for serving."}, {"id": "a6", "statement": "At the latest update, the required portioning of the chicken with rice and broccoli is complete."}, {"id": "a7", "statement": "At the latest update, the meal's setup is complete."}, {"id": "a8", "statement": "At the latest update, no corrective step other than continued chicken cooking prompted by the newly checked thick piece remains."}], "base_state_json": "\"At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours. By the latest update, the rice and broccoli had each completed all cooking steps required for serving. The required portioning of the chicken with rice and broccoli was complete, as was the meal setup. Every Q-17 piece apart from the item represented by the new reading had reached 74°C. At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17. At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact violet seal marked Q-17. Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C. The latest issue review found no corrective task remaining other than continued chicken cooking prompted by the newly checked thick piece, if that piece fell within the batch governed by the rule.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17."}, {"path": [], "text": "At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact violet seal marked Q-17."}], "policy_evidence": [{"path": [], "text": "At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours."}, {"path": [], "text": "Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C."}], "rules": [{"justification": "The 70°C piece belongs to the meal's chicken batch, so one batch piece is below 74°C. The kitchen rule therefore requires the entire batch to return to cooking, leaving a corrective cooking step and placing the meal at Level 2.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "The 70°C observation does not apply to the meal's chicken batch; every piece that does belong to that batch has reached 74°C. Setup, the other cooking steps, and required portioning are complete, and no applicable corrective step remains, so the meal is at Level 3.", "target": "true", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17.", "negative_left": "At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17.", "negative_right": "At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact amber seal marked M-28 instead of a Q-17 seal.", "right": "At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact violet seal marked Q-17."}, "verifier_independent_model": false}, "family": "scale-diverse-189-001", "id": "scale-diverse-189-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the meal is at Level 1 or Level 2 because setup is incomplete or a corrective cooking step is still required.", "true": "Yes — the meal is at Level 3 or Level 4 because all required cooking checks have passed and no corrective step remains."}, "instructions": "Use the latest temporal update and classify readiness on this ordered rubric: Level 1 = setup incomplete; Level 2 = cooking or another corrective step required; Level 3 = ready to serve; Level 4 = safely portioned and stored. Answer yes only if the meal is currently at Level 3 or Level 4. Is the meal currently ready to serve?", "type": "noul"}}, "state": "At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours. By the latest update, the rice and broccoli had each completed all cooking steps required for serving. The required portioning of the chicken with rice and broccoli was complete, as was the meal setup. Every Q-17 piece apart from the item represented by the new reading had reached 74°C. At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17. At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact violet seal marked Q-17. Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C. The latest issue review found no corrective task remaining other than continued chicken cooking prompted by the newly checked thick piece, if that piece fell within the batch governed by the rule."}, "method": "c2d", "provenance": {"source_id": "diverse-189", "source_is_synthetic": true, "source_sha256": "8db690328ed739930643b56a4c68d870d081e25d517614a57e4171b14e52bdb5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the recipe requirements, the one-low-piece kitchen rule, the current-meal readiness scope, and the latest-update timing. The two evidence spans are complete factual sentences. The counterfactual changes only the checked piece’s seal and therefore its membership in the finalized Q-17 batch; it does not create a duplicate or contradictory measurement. Neither context states a readiness label, answer code, proposition ID, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified completion or exclusion claims remain atomic rather than bundling a final classification. The focus atom is the factual membership relation between the thick piece and the meal’s chicken batch. Both assignments are realizable while changing only that membership fact: the 70°C observation can concern either a batch piece or a non-batch piece, with the other listed facts unchanged. The policy evidence correctly preserves the state-originating recipe threshold and batch-return rule; the readiness rubric and temporal instructions need not be repeated because they remain in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the 70°C piece belongs to the meal’s chicken batch. Because 70°C is below the required 74°C and the state’s kitchen rule sends the entire batch back for cooking when any piece is low, a corrective cooking step remains. The meal therefore cannot be at Level 3 or 4.", "rule_index": 0, "sound": true}, {"reason": "Refuting batch membership makes the 70°C observation inapplicable to this meal’s chicken batch. The remaining conditions establish that all batch pieces reached 74°C, the other required cooking steps and setup are complete, portioning is complete, and no other corrective step remains. Under the stated policy, the low-piece cooking correction is not triggered, so these conditions are sufficient for readiness.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The newly checked thick piece at 6:43 p.m. belongs to the chicken batch for this meal."}, {"id": "a2", "statement": "The newly checked thick piece had a measured temperature of 70°C at 6:43 p.m."}, {"id": "a3", "statement": "At the latest update, every chicken piece in the meal's batch other than the newly checked thick piece has reached 74°C."}, {"id": "a4", "statement": "At the latest update, the rice has completed all cooking steps required for serving."}, {"id": "a5", "statement": "At the latest update, the broccoli has completed all cooking steps required for serving."}, {"id": "a6", "statement": "At the latest update, the required portioning of the chicken with rice and broccoli is complete."}, {"id": "a7", "statement": "At the latest update, the meal's setup is complete."}, {"id": "a8", "statement": "At the latest update, no corrective step other than continued chicken cooking prompted by the newly checked thick piece remains."}], "base_state_json": "\"At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours. By the latest update, the rice and broccoli had each completed all cooking steps required for serving. The required portioning of the chicken with rice and broccoli was complete, as was the meal setup. Every Q-17 piece apart from the item represented by the new reading had reached 74°C. At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17. At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact violet seal marked Q-17. Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C. The latest issue review found no corrective task remaining other than continued chicken cooking prompted by the newly checked thick piece, if that piece fell within the batch governed by the rule.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17."}, {"path": [], "text": "At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact violet seal marked Q-17."}], "policy_evidence": [{"path": [], "text": "At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours."}, {"path": [], "text": "Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C."}], "rules": [{"justification": "The 70°C piece belongs to the meal's chicken batch, so one batch piece is below 74°C. The kitchen rule therefore requires the entire batch to return to cooking, leaving a corrective cooking step and placing the meal at Level 2.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "The 70°C observation does not apply to the meal's chicken batch; every piece that does belong to that batch has reached 74°C. Setup, the other cooking steps, and required portioning are complete, and no applicable corrective step remains, so the meal is at Level 3.", "target": "true", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17.", "negative_left": "At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17.", "negative_right": "At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact amber seal marked M-28 instead of a Q-17 seal.", "right": "At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact violet seal marked Q-17."}, "verifier_independent_model": false}, "family": "scale-diverse-189-001", "id": "scale-diverse-189-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the meal is at Level 1 or Level 2 because setup is incomplete or a corrective cooking step is still required.", "true": "Yes — the meal is at Level 3 or Level 4 because all required cooking checks have passed and no corrective step remains."}, "instructions": "Use the latest temporal update and classify readiness on this ordered rubric: Level 1 = setup incomplete; Level 2 = cooking or another corrective step required; Level 3 = ready to serve; Level 4 = safely portioned and stored. Answer yes only if the meal is currently at Level 3 or Level 4. Is the meal currently ready to serve?", "type": "noul"}}, "state": "At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours. By the latest update, the rice and broccoli had each completed all cooking steps required for serving. The required portioning of the chicken with rice and broccoli was complete, as was the meal setup. Every Q-17 piece apart from the item represented by the new reading had reached 74°C. At 6:43 p.m., the finalized tray inventory identified the chicken batch for this meal as exactly the pieces carrying an intact violet seal marked Q-17. At 6:43 p.m., the newly checked thick piece measured 70°C and carried an intact amber seal marked M-28 instead of a Q-17 seal. Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C. The latest issue review found no corrective task remaining other than continued chicken cooking prompted by the newly checked thick piece, if that piece fell within the batch governed by the rule."}, "method": "c2d", "provenance": {"source_id": "diverse-189", "source_is_synthetic": true, "source_sha256": "8db690328ed739930643b56a4c68d870d081e25d517614a57e4171b14e52bdb5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 74°C recipe requirement, portioning and refrigeration requirements, and the kitchen-wide corrective-cooking rule without adding exceptions or defaults. The question remains about the same meal’s current readiness at the operational handoff. The two evidence spans are complete factual sentences. In the counterfactual, the checked HN-317 piece is outside the meal batch identified exclusively as QV-842, so its 70°C measurement does not conflict with the statement that the meal-batch pieces passed; “every other piece” is slightly awkward but does not create an explicit contradictory measurement. Neither constructed context states a yes/no result, rubric level, answer code, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified completion or exclusion claims remain atomic rather than bundling a final classification. The focus atom is the factual membership relation between the thick piece and the meal’s chicken batch. Both assignments are realizable while changing only that membership fact: the 70°C observation can concern either a batch piece or a non-batch piece, with the other listed facts unchanged. The policy evidence correctly preserves the state-originating recipe threshold and batch-return rule; the readiness rubric and temporal instructions need not be repeated because they remain in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the 70°C piece belongs to the meal’s chicken batch. Because 70°C is below the required 74°C and the state’s kitchen rule sends the entire batch back for cooking when any piece is low, a corrective cooking step remains. The meal therefore cannot be at Level 3 or 4.", "rule_index": 0, "sound": true}, {"reason": "Refuting batch membership makes the 70°C observation inapplicable to this meal’s chicken batch. The remaining conditions establish that all batch pieces reached 74°C, the other required cooking steps and setup are complete, portioning is complete, and no other corrective step remains. Under the stated policy, the low-piece cooking correction is not triggered, so these conditions are sufficient for readiness.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The newly checked thick piece at 6:43 p.m. belongs to the chicken batch for this meal."}, {"id": "a2", "statement": "The newly checked thick piece had a measured temperature of 70°C at 6:43 p.m."}, {"id": "a3", "statement": "At the latest update, every chicken piece in the meal's batch other than the newly checked thick piece has reached 74°C."}, {"id": "a4", "statement": "At the latest update, the rice has completed all cooking steps required for serving."}, {"id": "a5", "statement": "At the latest update, the broccoli has completed all cooking steps required for serving."}, {"id": "a6", "statement": "At the latest update, the required portioning of the chicken with rice and broccoli is complete."}, {"id": "a7", "statement": "At the latest update, the meal's setup is complete."}, {"id": "a8", "statement": "At the latest update, no corrective step other than continued chicken cooking prompted by the newly checked thick piece remains."}], "base_state_json": "\"Operational handoff, 6:45 p.m. At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317. The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code HN-317. The calibrated probe log records that piece at 70°C at 6:43 p.m. The latest inspection confirms that every other piece in the meal’s chicken batch has reached 74°C. Rice and broccoli have each completed all cooking steps required for serving. Required portioning of the chicken with rice and broccoli is complete, and the meal setup is complete. No unresolved corrective action is listed apart from continued chicken cooking if the newly checked thick piece applies to this meal’s batch. At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours. Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317."}, {"path": [], "text": "The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code HN-317."}], "policy_evidence": [{"path": [], "text": "At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours."}, {"path": [], "text": "Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C."}], "rules": [{"justification": "The 70°C piece belongs to the meal's chicken batch, so one batch piece is below 74°C. The kitchen rule therefore requires the entire batch to return to cooking, leaving a corrective cooking step and placing the meal at Level 2.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "The 70°C observation does not apply to the meal's chicken batch; every piece that does belong to that batch has reached 74°C. Setup, the other cooking steps, and required portioning are complete, and no applicable corrective step remains, so the meal is at Level 3.", "target": "true", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317.", "negative_left": "At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317.", "negative_right": "The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code QV-842.", "right": "The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code HN-317."}, "verifier_independent_model": false}, "family": "scale-diverse-189-003", "id": "scale-diverse-189-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the meal is at Level 1 or Level 2 because setup is incomplete or a corrective cooking step is still required.", "true": "Yes — the meal is at Level 3 or Level 4 because all required cooking checks have passed and no corrective step remains."}, "instructions": "Use the latest temporal update and classify readiness on this ordered rubric: Level 1 = setup incomplete; Level 2 = cooking or another corrective step required; Level 3 = ready to serve; Level 4 = safely portioned and stored. Answer yes only if the meal is currently at Level 3 or Level 4. Is the meal currently ready to serve?", "type": "noul"}}, "state": "Operational handoff, 6:45 p.m. At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317. The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code HN-317. The calibrated probe log records that piece at 70°C at 6:43 p.m. The latest inspection confirms that every other piece in the meal’s chicken batch has reached 74°C. Rice and broccoli have each completed all cooking steps required for serving. Required portioning of the chicken with rice and broccoli is complete, and the meal setup is complete. No unresolved corrective action is listed apart from continued chicken cooking if the newly checked thick piece applies to this meal’s batch. At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours. Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C."}, "method": "c2d", "provenance": {"source_id": "diverse-189", "source_is_synthetic": true, "source_sha256": "8db690328ed739930643b56a4c68d870d081e25d517614a57e4171b14e52bdb5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 74°C recipe requirement, portioning and refrigeration requirements, and the kitchen-wide corrective-cooking rule without adding exceptions or defaults. The question remains about the same meal’s current readiness at the operational handoff. The two evidence spans are complete factual sentences. In the counterfactual, the checked HN-317 piece is outside the meal batch identified exclusively as QV-842, so its 70°C measurement does not conflict with the statement that the meal-batch pieces passed; “every other piece” is slightly awkward but does not create an explicit contradictory measurement. Neither constructed context states a yes/no result, rubric level, answer code, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified completion or exclusion claims remain atomic rather than bundling a final classification. The focus atom is the factual membership relation between the thick piece and the meal’s chicken batch. Both assignments are realizable while changing only that membership fact: the 70°C observation can concern either a batch piece or a non-batch piece, with the other listed facts unchanged. The policy evidence correctly preserves the state-originating recipe threshold and batch-return rule; the readiness rubric and temporal instructions need not be repeated because they remain in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the 70°C piece belongs to the meal’s chicken batch. Because 70°C is below the required 74°C and the state’s kitchen rule sends the entire batch back for cooking when any piece is low, a corrective cooking step remains. The meal therefore cannot be at Level 3 or 4.", "rule_index": 0, "sound": true}, {"reason": "Refuting batch membership makes the 70°C observation inapplicable to this meal’s chicken batch. The remaining conditions establish that all batch pieces reached 74°C, the other required cooking steps and setup are complete, portioning is complete, and no other corrective step remains. Under the stated policy, the low-piece cooking correction is not triggered, so these conditions are sufficient for readiness.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The newly checked thick piece at 6:43 p.m. belongs to the chicken batch for this meal."}, {"id": "a2", "statement": "The newly checked thick piece had a measured temperature of 70°C at 6:43 p.m."}, {"id": "a3", "statement": "At the latest update, every chicken piece in the meal's batch other than the newly checked thick piece has reached 74°C."}, {"id": "a4", "statement": "At the latest update, the rice has completed all cooking steps required for serving."}, {"id": "a5", "statement": "At the latest update, the broccoli has completed all cooking steps required for serving."}, {"id": "a6", "statement": "At the latest update, the required portioning of the chicken with rice and broccoli is complete."}, {"id": "a7", "statement": "At the latest update, the meal's setup is complete."}, {"id": "a8", "statement": "At the latest update, no corrective step other than continued chicken cooking prompted by the newly checked thick piece remains."}], "base_state_json": "\"Operational handoff, 6:45 p.m. At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317. The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code HN-317. The calibrated probe log records that piece at 70°C at 6:43 p.m. The latest inspection confirms that every other piece in the meal’s chicken batch has reached 74°C. Rice and broccoli have each completed all cooking steps required for serving. Required portioning of the chicken with rice and broccoli is complete, and the meal setup is complete. No unresolved corrective action is listed apart from continued chicken cooking if the newly checked thick piece applies to this meal’s batch. At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours. Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317."}, {"path": [], "text": "The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code HN-317."}], "policy_evidence": [{"path": [], "text": "At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours."}, {"path": [], "text": "Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C."}], "rules": [{"justification": "The 70°C piece belongs to the meal's chicken batch, so one batch piece is below 74°C. The kitchen rule therefore requires the entire batch to return to cooking, leaving a corrective cooking step and placing the meal at Level 2.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "The 70°C observation does not apply to the meal's chicken batch; every piece that does belong to that batch has reached 74°C. Setup, the other cooking steps, and required portioning are complete, and no applicable corrective step remains, so the meal is at Level 3.", "target": "true", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317.", "negative_left": "At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317.", "negative_right": "The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code QV-842.", "right": "The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code HN-317."}, "verifier_independent_model": false}, "family": "scale-diverse-189-003", "id": "scale-diverse-189-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the meal is at Level 1 or Level 2 because setup is incomplete or a corrective cooking step is still required.", "true": "Yes — the meal is at Level 3 or Level 4 because all required cooking checks have passed and no corrective step remains."}, "instructions": "Use the latest temporal update and classify readiness on this ordered rubric: Level 1 = setup incomplete; Level 2 = cooking or another corrective step required; Level 3 = ready to serve; Level 4 = safely portioned and stored. Answer yes only if the meal is currently at Level 3 or Level 4. Is the meal currently ready to serve?", "type": "noul"}}, "state": "Operational handoff, 6:45 p.m. At 6:43 p.m., the newly checked thick piece bore the single sealed batch code HN-317. The 6:41 p.m. operational handoff record identifies the chicken batch for this meal as exactly the pieces bearing sealed batch code QV-842. The calibrated probe log records that piece at 70°C at 6:43 p.m. The latest inspection confirms that every other piece in the meal’s chicken batch has reached 74°C. Rice and broccoli have each completed all cooking steps required for serving. Required portioning of the chicken with rice and broccoli is complete, and the meal setup is complete. No unresolved corrective action is listed apart from continued chicken cooking if the newly checked thick piece applies to this meal’s batch. At 6:10 p.m., the meal planner issued the recipe: roast all chicken pieces until every piece reaches 74°C, then portion with rice and broccoli; refrigerate extra portions within two hours. Under the kitchen rule, one low piece means the entire chicken batch returns to the cooking station until every piece passes 74°C."}, "method": "c2d", "provenance": {"source_id": "diverse-189", "source_is_synthetic": true, "source_sha256": "8db690328ed739930643b56a4c68d870d081e25d517614a57e4171b14e52bdb5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing rubric, kitchen rule, request scope, meal entity, date, and readiness question are preserved verbatim. The two focus-evidence spans are complete factual sentences. The counterfactual changes cooling onset from 19:42 to 20:19 while retaining the 18:07 cooking completion time; this is coherent and creates a 2-hour-12-minute interval without contradicting any duplicate time or count. Neither context embeds an answer, output instruction, code, rule table, proposition ID, or label rationale; use of “compliant cooling” is permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.\",\"true\":\"Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions.\"},\"instructions\":\"Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.\",\"type\":\"noul\"}},\"state\":{\"context\":\"A chronological kitchen log documents Mara’s preparation of paprika chicken and rice on 14 May 2026.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"The log records that Mara completed simmering the sauce before she began cooking the chicken.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"Mara used that probe on the chicken and obtained exactly two readings: 75°C and 74°C.\",\"Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock.\",\"Both of the meal's two designated storage portions began compliant cooling at 19:42 on 14 May 2026, according to the same synchronized clock.\",\"After cooking, Mara divided the entire meal into exactly four portions and verified that each contained the same amount of food. She designated exactly two portions for serving now and exactly two for storage. Both serving portions were plated and prepared for serving. Both storage portions were placed in labeled shallow containers to begin compliant cooling.\"],\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["state", "evidence", "4"], "text": "Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock."}, {"path": ["state", "evidence", "5"], "text": "Both of the meal's two designated storage portions began compliant cooling at 19:42 on 14 May 2026, according to the same synchronized clock."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock.", "negative_left": "Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock.", "negative_right": "Both of the meal's two designated storage portions began compliant cooling at 20:19 on 14 May 2026, according to the same synchronized clock.", "right": "Both of the meal's two designated storage portions began compliant cooling at 19:42 on 14 May 2026, according to the same synchronized clock."}, "verifier_independent_model": false}, "family": "scale-diverse-190-001", "id": "scale-diverse-190-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "A chronological kitchen log documents Mara’s preparation of paprika chicken and rice on 14 May 2026.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "The log records that Mara completed simmering the sauce before she began cooking the chicken.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara used that probe on the chicken and obtained exactly two readings: 75°C and 74°C.", "Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock.", "Both of the meal's two designated storage portions began compliant cooling at 19:42 on 14 May 2026, according to the same synchronized clock.", "After cooking, Mara divided the entire meal into exactly four portions and verified that each contained the same amount of food. She designated exactly two portions for serving now and exactly two for storage. Both serving portions were plated and prepared for serving. Both storage portions were placed in labeled shallow containers to begin compliant cooling."], "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing rubric, kitchen rule, request scope, meal entity, date, and readiness question are preserved verbatim. The two focus-evidence spans are complete factual sentences. The counterfactual changes cooling onset from 19:42 to 20:19 while retaining the 18:07 cooking completion time; this is coherent and creates a 2-hour-12-minute interval without contradicting any duplicate time or count. Neither context embeds an answer, output instruction, code, rule table, proposition ID, or label rationale; use of “compliant cooling” is permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.\",\"true\":\"Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions.\"},\"instructions\":\"Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.\",\"type\":\"noul\"}},\"state\":{\"context\":\"A chronological kitchen log documents Mara’s preparation of paprika chicken and rice on 14 May 2026.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"The log records that Mara completed simmering the sauce before she began cooking the chicken.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"Mara used that probe on the chicken and obtained exactly two readings: 75°C and 74°C.\",\"Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock.\",\"Both of the meal's two designated storage portions began compliant cooling at 19:42 on 14 May 2026, according to the same synchronized clock.\",\"After cooking, Mara divided the entire meal into exactly four portions and verified that each contained the same amount of food. She designated exactly two portions for serving now and exactly two for storage. Both serving portions were plated and prepared for serving. Both storage portions were placed in labeled shallow containers to begin compliant cooling.\"],\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["state", "evidence", "4"], "text": "Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock."}, {"path": ["state", "evidence", "5"], "text": "Both of the meal's two designated storage portions began compliant cooling at 19:42 on 14 May 2026, according to the same synchronized clock."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock.", "negative_left": "Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock.", "negative_right": "Both of the meal's two designated storage portions began compliant cooling at 20:19 on 14 May 2026, according to the same synchronized clock.", "right": "Both of the meal's two designated storage portions began compliant cooling at 19:42 on 14 May 2026, according to the same synchronized clock."}, "verifier_independent_model": false}, "family": "scale-diverse-190-001", "id": "scale-diverse-190-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "A chronological kitchen log documents Mara’s preparation of paprika chicken and rice on 14 May 2026.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "The log records that Mara completed simmering the sauce before she began cooking the chicken.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara used that probe on the chicken and obtained exactly two readings: 75°C and 74°C.", "Mara finished cooking the meal at 18:07 on 14 May 2026, as recorded by the kitchen's synchronized clock.", "Both of the meal's two designated storage portions began compliant cooling at 20:19 on 14 May 2026, according to the same synchronized clock.", "After cooking, Mara divided the entire meal into exactly four portions and verified that each contained the same amount of food. She designated exactly two portions for serving now and exactly two for storage. Both serving portions were plated and prepared for serving. Both storage portions were placed in labeled shallow containers to begin compliant cooling."], "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the questions verbatim and preserve the recipe requirements and probe-over-color kitchen rule from the original state. The request remains bound to Mara’s same meal, portions, readiness level, and cooling timeline. The two focus spans are complete factual sentences. Changing the cooling start from 19:49 to 20:29 is coherent with the unchanged 18:12 cooking-end time and creates no duplicate measurement or timing assertion; “compliant cooling” can describe the cooling method while the separate two-hour timeliness criterion remains governed by the question. Neither context includes an answer, code, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.\",\"true\":\"Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions.\"},\"instructions\":\"Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.\",\"type\":\"noul\"}},\"context\":\"Field note for Mara’s paprika chicken and rice, with prep by Sol, compliance checked by Inez, and cooling handled by Dev. The note was completed on the cooking day, and ‘today’ below refers to that same day.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Mara completed simmering the sauce before she began cooking the chicken.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"Mara used that probe on the chicken and obtained exactly two readings: 75°C and 74°C.\",\"After cooking, Mara divided the meal into exactly four portions containing equal amounts. Exactly two were designated and prepared for serving now; the other two were designated for storage.\",\"Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00.\",\"Both of that meal's two designated storage portions began compliant cooling at 19:49 on 14 August 2026, UTC−05:00.\"],\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "5"], "text": "Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00."}, {"path": ["evidence", "6"], "text": "Both of that meal's two designated storage portions began compliant cooling at 19:49 on 14 August 2026, UTC−05:00."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00.", "negative_left": "Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00.", "negative_right": "Both of that meal's two designated storage portions began compliant cooling at 20:29 on 14 August 2026, UTC−05:00.", "right": "Both of that meal's two designated storage portions began compliant cooling at 19:49 on 14 August 2026, UTC−05:00."}, "verifier_independent_model": false}, "family": "scale-diverse-190-004", "id": "scale-diverse-190-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Field note for Mara’s paprika chicken and rice, with prep by Sol, compliance checked by Inez, and cooling handled by Dev. The note was completed on the cooking day, and ‘today’ below refers to that same day.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Mara completed simmering the sauce before she began cooking the chicken.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara used that probe on the chicken and obtained exactly two readings: 75°C and 74°C.", "After cooking, Mara divided the meal into exactly four portions containing equal amounts. Exactly two were designated and prepared for serving now; the other two were designated for storage.", "Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00.", "Both of that meal's two designated storage portions began compliant cooling at 19:49 on 14 August 2026, UTC−05:00."], "questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the questions verbatim and preserve the recipe requirements and probe-over-color kitchen rule from the original state. The request remains bound to Mara’s same meal, portions, readiness level, and cooling timeline. The two focus spans are complete factual sentences. Changing the cooling start from 19:49 to 20:29 is coherent with the unchanged 18:12 cooking-end time and creates no duplicate measurement or timing assertion; “compliant cooling” can describe the cooling method while the separate two-hour timeliness criterion remains governed by the question. Neither context includes an answer, code, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"false\":\"No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.\",\"true\":\"Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions.\"},\"instructions\":\"Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.\",\"type\":\"noul\"}},\"context\":\"Field note for Mara’s paprika chicken and rice, with prep by Sol, compliance checked by Inez, and cooling handled by Dev. The note was completed on the cooking day, and ‘today’ below refers to that same day.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Mara completed simmering the sauce before she began cooking the chicken.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"Mara used that probe on the chicken and obtained exactly two readings: 75°C and 74°C.\",\"After cooking, Mara divided the meal into exactly four portions containing equal amounts. Exactly two were designated and prepared for serving now; the other two were designated for storage.\",\"Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00.\",\"Both of that meal's two designated storage portions began compliant cooling at 19:49 on 14 August 2026, UTC−05:00.\"],\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "5"], "text": "Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00."}, {"path": ["evidence", "6"], "text": "Both of that meal's two designated storage portions began compliant cooling at 19:49 on 14 August 2026, UTC−05:00."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00.", "negative_left": "Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00.", "negative_right": "Both of that meal's two designated storage portions began compliant cooling at 20:29 on 14 August 2026, UTC−05:00.", "right": "Both of that meal's two designated storage portions began compliant cooling at 19:49 on 14 August 2026, UTC−05:00."}, "verifier_independent_model": false}, "family": "scale-diverse-190-004", "id": "scale-diverse-190-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Field note for Mara’s paprika chicken and rice, with prep by Sol, compliance checked by Inez, and cooling handled by Dev. The note was completed on the cooking day, and ‘today’ below refers to that same day.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Mara completed simmering the sauce before she began cooking the chicken.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara used that probe on the chicken and obtained exactly two readings: 75°C and 74°C.", "After cooking, Mara divided the meal into exactly four portions containing equal amounts. Exactly two were designated and prepared for serving now; the other two were designated for storage.", "Mara ended cooking the meal at 18:12 on 14 August 2026, UTC−05:00.", "Both of that meal's two designated storage portions began compliant cooling at 20:29 on 14 August 2026, UTC−05:00."], "questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same chicken-temperature threshold, thickest-piece measurement requirement, prohibition on substitute indicators, further-cooking condition, and refrigeration requirement. The decision remains bound to the lemon chicken, thermometer event M, first serving/portioning event B, four dinners, and the specified date and times; only B’s observed timestamp changes. The two evidence spans are complete factual event-log sentences. In the counterfactual, B occurs before M, which represents noncompliant sequencing but does not contradict the single measurement, event count, or other unchanged assertions. Neither context contains an answer code, gold label, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Handoff coordinator\",\"text\":\"At 19:00, the home cook noted that the required internal-temperature reading was undocumented; this was before either logged event. Up to thermometer event M, no internal-temperature reading for the lemon chicken had been documented or announced. The execution route scheduled exactly one thermometer event for this chicken, designated M, with further cooking required if M showed below 74°C. Chicken-handling event B was defined as the first serving or portioning into the four dinners.\"},{\"speaker\":\"Probe operator\",\"text\":\"At M, the probe was placed in the thickest piece and recorded 74.6°C. The disposition record confirms that every unserved portion was refrigerated within the required two-hour window.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned.\"},{\"speaker\":\"Event logger\",\"text\":\"The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC.\"},{\"speaker\":\"Event logger\",\"text\":\"The kitchen event log records chicken-handling event B at 19:17:42 on 17 September 2026, UTC.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC."}, {"path": ["5", "text"], "text": "The kitchen event log records chicken-handling event B at 19:17:42 on 17 September 2026, UTC."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC.", "negative_left": "The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC.", "negative_right": "The kitchen event log records chicken-handling event B at 19:11:35 on 17 September 2026, UTC.", "right": "The kitchen event log records chicken-handling event B at 19:17:42 on 17 September 2026, UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-191-003", "id": "scale-diverse-191-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Handoff coordinator", "text": "At 19:00, the home cook noted that the required internal-temperature reading was undocumented; this was before either logged event. Up to thermometer event M, no internal-temperature reading for the lemon chicken had been documented or announced. The execution route scheduled exactly one thermometer event for this chicken, designated M, with further cooking required if M showed below 74°C. Chicken-handling event B was defined as the first serving or portioning into the four dinners."}, {"speaker": "Probe operator", "text": "At M, the probe was placed in the thickest piece and recorded 74.6°C. The disposition record confirms that every unserved portion was refrigerated within the required two-hour window."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}, {"speaker": "Event logger", "text": "The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC."}, {"speaker": "Event logger", "text": "The kitchen event log records chicken-handling event B at 19:17:42 on 17 September 2026, UTC."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same chicken-temperature threshold, thickest-piece measurement requirement, prohibition on substitute indicators, further-cooking condition, and refrigeration requirement. The decision remains bound to the lemon chicken, thermometer event M, first serving/portioning event B, four dinners, and the specified date and times; only B’s observed timestamp changes. The two evidence spans are complete factual event-log sentences. In the counterfactual, B occurs before M, which represents noncompliant sequencing but does not contradict the single measurement, event count, or other unchanged assertions. Neither context contains an answer code, gold label, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Handoff coordinator\",\"text\":\"At 19:00, the home cook noted that the required internal-temperature reading was undocumented; this was before either logged event. Up to thermometer event M, no internal-temperature reading for the lemon chicken had been documented or announced. The execution route scheduled exactly one thermometer event for this chicken, designated M, with further cooking required if M showed below 74°C. Chicken-handling event B was defined as the first serving or portioning into the four dinners.\"},{\"speaker\":\"Probe operator\",\"text\":\"At M, the probe was placed in the thickest piece and recorded 74.6°C. The disposition record confirms that every unserved portion was refrigerated within the required two-hour window.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned.\"},{\"speaker\":\"Event logger\",\"text\":\"The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC.\"},{\"speaker\":\"Event logger\",\"text\":\"The kitchen event log records chicken-handling event B at 19:17:42 on 17 September 2026, UTC.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC."}, {"path": ["5", "text"], "text": "The kitchen event log records chicken-handling event B at 19:17:42 on 17 September 2026, UTC."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC.", "negative_left": "The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC.", "negative_right": "The kitchen event log records chicken-handling event B at 19:11:35 on 17 September 2026, UTC.", "right": "The kitchen event log records chicken-handling event B at 19:17:42 on 17 September 2026, UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-191-003", "id": "scale-diverse-191-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Handoff coordinator", "text": "At 19:00, the home cook noted that the required internal-temperature reading was undocumented; this was before either logged event. Up to thermometer event M, no internal-temperature reading for the lemon chicken had been documented or announced. The execution route scheduled exactly one thermometer event for this chicken, designated M, with further cooking required if M showed below 74°C. Chicken-handling event B was defined as the first serving or portioning into the four dinners."}, {"speaker": "Probe operator", "text": "At M, the probe was placed in the thickest piece and recorded 74.6°C. The disposition record confirms that every unserved portion was refrigerated within the required two-hour window."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}, {"speaker": "Event logger", "text": "The kitchen event log records thermometer event M at 19:14:08 on 17 September 2026, UTC."}, {"speaker": "Event logger", "text": "The kitchen event log records chicken-handling event B at 19:11:35 on 17 September 2026, UTC."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing temperature and refrigeration rules, while the unchanged questions preserve the decision criteria. The two evidence spans are complete factual sentences, and the sole counterfactual timing change coherently moves thermometer event M from before event B to after it without creating duplicate or contradictory measurements. The chicken, decision path, and event-date bindings remain fixed, and neither context contains a score, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Field note\",\"text\":\"Before either event occurred, the home cook recognized that the required internal-temperature reading was undocumented. No internal-temperature reading for the lemon chicken had been documented or announced at that point.\"},{\"speaker\":\"Kitchen log\",\"text\":\"The kitchen log records thermometer event M at 19:12:08 on 17 September 2026.\"},{\"speaker\":\"Kitchen log\",\"text\":\"The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned.\"},{\"speaker\":\"Field note\",\"text\":\"The selected route scheduled exactly one thermometer event for the lemon chicken, designated M. M measured the thickest piece and recorded 76°C; the route required further cooking if that result had been below 74°C. Event B was the first serving or portioning of the chicken into the four dinners. Every portion not served was refrigerated within the required two-hour window.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "The kitchen log records thermometer event M at 19:12:08 on 17 September 2026."}, {"path": ["3", "text"], "text": "The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The kitchen log records thermometer event M at 19:12:08 on 17 September 2026.", "negative_left": "The kitchen log records thermometer event M at 19:24:08 on 17 September 2026.", "negative_right": "The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026.", "right": "The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-191-004", "id": "scale-diverse-191-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Field note", "text": "Before either event occurred, the home cook recognized that the required internal-temperature reading was undocumented. No internal-temperature reading for the lemon chicken had been documented or announced at that point."}, {"speaker": "Kitchen log", "text": "The kitchen log records thermometer event M at 19:12:08 on 17 September 2026."}, {"speaker": "Kitchen log", "text": "The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}, {"speaker": "Field note", "text": "The selected route scheduled exactly one thermometer event for the lemon chicken, designated M. M measured the thickest piece and recorded 76°C; the route required further cooking if that result had been below 74°C. Event B was the first serving or portioning of the chicken into the four dinners. Every portion not served was refrigerated within the required two-hour window."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing temperature and refrigeration rules, while the unchanged questions preserve the decision criteria. The two evidence spans are complete factual sentences, and the sole counterfactual timing change coherently moves thermometer event M from before event B to after it without creating duplicate or contradictory measurements. The chicken, decision path, and event-date bindings remain fixed, and neither context contains a score, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Field note\",\"text\":\"Before either event occurred, the home cook recognized that the required internal-temperature reading was undocumented. No internal-temperature reading for the lemon chicken had been documented or announced at that point.\"},{\"speaker\":\"Kitchen log\",\"text\":\"The kitchen log records thermometer event M at 19:12:08 on 17 September 2026.\"},{\"speaker\":\"Kitchen log\",\"text\":\"The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned.\"},{\"speaker\":\"Field note\",\"text\":\"The selected route scheduled exactly one thermometer event for the lemon chicken, designated M. M measured the thickest piece and recorded 76°C; the route required further cooking if that result had been below 74°C. Event B was the first serving or portioning of the chicken into the four dinners. Every portion not served was refrigerated within the required two-hour window.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "The kitchen log records thermometer event M at 19:12:08 on 17 September 2026."}, {"path": ["3", "text"], "text": "The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The kitchen log records thermometer event M at 19:12:08 on 17 September 2026.", "negative_left": "The kitchen log records thermometer event M at 19:24:08 on 17 September 2026.", "negative_right": "The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026.", "right": "The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-191-004", "id": "scale-diverse-191-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Field note", "text": "Before either event occurred, the home cook recognized that the required internal-temperature reading was undocumented. No internal-temperature reading for the lemon chicken had been documented or announced at that point."}, {"speaker": "Kitchen log", "text": "The kitchen log records thermometer event M at 19:24:08 on 17 September 2026."}, {"speaker": "Kitchen log", "text": "The kitchen log records chicken-handling event B at 19:18:41 on 17 September 2026."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}, {"speaker": "Field note", "text": "The selected route scheduled exactly one thermometer event for the lemon chicken, designated M. M measured the thickest piece and recorded 76°C; the route required further cooking if that result had been below 74°C. Event B was the first serving or portioning of the chicken into the four dinners. Every portion not served was refrigerated within the required two-hour window."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the Tuesday-dinner readiness scope, required components, preparation requirements, and explicit exceptions; the unchanged questions retain the full scoring policy. The meal, assessment time, and relevant component bindings remain intact. The two evidence spans are complete factual sentences. In the counterfactual, unique tag Q63 coherently makes S17 the chicken piece outside B42, so its 68°C measurement does not conflict with the assertion that every chicken piece in B42 other than S17 reached at least 74°C. Neither context includes a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the universally quantified batch and contamination propositions. The focus is the factual membership relation between S17 and B42. The base and counter assignments differ only on that membership and are jointly realizable: in the base S17 is the under-temperature batch piece, while in the counter it is outside B42 and every actual B42 piece is at least 74°C. Policy evidence preserves the needed state-originated component bindings, preparation thresholds, scope, and exceptions; decision-level priorities and classifications remain automatically preserved in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that B42 is the required chicken component and that one of its pieces is at 68°C at assessment, below the state-specified mandatory 74°C threshold. This is sufficient for level 0 regardless of the other components.", "rule_index": 0, "sound": true}, {"reason": "Refuting S17's membership in B42, together with every other B42 piece being at least 74°C, entails that all chicken in the required batch meets the temperature requirement. The contamination atom excludes unresolved in-scope contamination, while the unportioned required rice guarantees a remaining completion step. Unspecified resting or draining facts cannot improve or worsen the result beyond level 1: failures there would also be incompleteness, not a stated safety failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the Tuesday-dinner readiness assessment, batch B42 is the meal planner's required lemon-chicken component."}, {"id": "a2", "statement": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 measures 68°C."}, {"id": "a3", "statement": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is a piece of chicken in batch B42."}, {"id": "a4", "statement": "At the Tuesday-dinner readiness assessment, every piece of chicken in batch B42 other than specimen S17 has reached at least 74°C."}, {"id": "a5", "statement": "At the Tuesday-dinner readiness assessment, zero of the three required meal components has unresolved contamination."}, {"id": "a6", "statement": "Before the Tuesday-dinner readiness assessment, the required rice rested for six minutes."}, {"id": "a7", "statement": "The required rice remained covered throughout its rest before the Tuesday-dinner readiness assessment."}, {"id": "a8", "statement": "At the Tuesday-dinner readiness assessment, the required broccoli is drained."}, {"id": "a9", "statement": "At the Tuesday-dinner readiness assessment, the required rice has not been portioned for immediate serving."}], "base_state_json": "\"For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli. The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving. Readiness covers only the three required meal components. Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions.\\n\\nAt the Tuesday-dinner readiness assessment, batch B42 was the meal planner’s required lemon-chicken component. At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63. At the Tuesday-dinner readiness assessment, a chicken piece in batch B42 bore identifier tag Q63, and no other item bore that tag. Specimen S17 measured 68°C. Every item that was both a chicken piece in B42 and not S17 had reached at least 74°C. Inspectors found unresolved contamination in zero of the three required components. Before the assessment, the required rice rested for six minutes and remained covered throughout. The required broccoli was drained. The required rice had not been portioned for immediate serving.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63."}, {"path": [], "text": "At the Tuesday-dinner readiness assessment, a chicken piece in batch B42 bore identifier tag Q63, and no other item bore that tag."}], "policy_evidence": [{"path": [], "text": "For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli."}, {"path": [], "text": "The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving."}, {"path": [], "text": "Readiness covers only the three required meal components."}, {"path": [], "text": "Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions."}], "rules": [{"justification": "S17 is part of the required chicken batch and measures only 68°C, below the mandatory 74°C requirement, so a mandatory food-safety requirement fails.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "Because S17 is not in B42, the universal temperature fact covers every piece of the required chicken batch; no required component has unresolved contamination, so the in-scope safety requirements pass. The required rice remains unportioned, making the meal safe but incomplete.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63.", "negative_left": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63.", "negative_right": "At the Tuesday-dinner readiness assessment, a chicken piece outside batch B42 bore identifier tag Q63, and no other item bore that tag.", "right": "At the Tuesday-dinner readiness assessment, a chicken piece in batch B42 bore identifier tag Q63, and no other item bore that tag."}, "verifier_independent_model": false}, "family": "scale-diverse-192-006", "id": "scale-diverse-192-006-base", "input": {"questions": {"decision": {"criteria": ["Not ready—at least one in-scope mandatory food-safety requirement fails, or contamination still requires correction; the meal must not be served.", "Safe but incomplete—all in-scope safety requirements pass, but at least one required component still needs cooking, resting, draining, or portioning.", "Ready with a minor in-scope presentation or workflow issue—all required components are safe and complete, but a non-safety serving step remains; explicit exceptions do not count.", "Fully ready—all required components are safe, cooked, rested or drained as specified, and portioned for immediate serving, with no unresolved in-scope issue."], "instructions": "Rate the meal’s current readiness to serve using the ordered levels below. Apply requirements only to the stated readiness scope; do not lower the rating for explicit exceptions. A mandatory food-safety failure takes precedence over otherwise completed components.", "type": "score"}}, "state": "For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli. The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving. Readiness covers only the three required meal components. Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions.\n\nAt the Tuesday-dinner readiness assessment, batch B42 was the meal planner’s required lemon-chicken component. At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63. At the Tuesday-dinner readiness assessment, a chicken piece in batch B42 bore identifier tag Q63, and no other item bore that tag. Specimen S17 measured 68°C. Every item that was both a chicken piece in B42 and not S17 had reached at least 74°C. Inspectors found unresolved contamination in zero of the three required components. Before the assessment, the required rice rested for six minutes and remained covered throughout. The required broccoli was drained. The required rice had not been portioned for immediate serving."}, "method": "c2d", "provenance": {"source_id": "diverse-192", "source_is_synthetic": true, "source_sha256": "f16699a58ce6f897e6254823f6038b87cdbce46e187582c3dc5640f72ae7c868", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the Tuesday-dinner readiness scope, required components, preparation requirements, and explicit exceptions; the unchanged questions retain the full scoring policy. The meal, assessment time, and relevant component bindings remain intact. The two evidence spans are complete factual sentences. In the counterfactual, unique tag Q63 coherently makes S17 the chicken piece outside B42, so its 68°C measurement does not conflict with the assertion that every chicken piece in B42 other than S17 reached at least 74°C. Neither context includes a gold answer, output instruction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the universally quantified batch and contamination propositions. The focus is the factual membership relation between S17 and B42. The base and counter assignments differ only on that membership and are jointly realizable: in the base S17 is the under-temperature batch piece, while in the counter it is outside B42 and every actual B42 piece is at least 74°C. Policy evidence preserves the needed state-originated component bindings, preparation thresholds, scope, and exceptions; decision-level priorities and classifications remain automatically preserved in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that B42 is the required chicken component and that one of its pieces is at 68°C at assessment, below the state-specified mandatory 74°C threshold. This is sufficient for level 0 regardless of the other components.", "rule_index": 0, "sound": true}, {"reason": "Refuting S17's membership in B42, together with every other B42 piece being at least 74°C, entails that all chicken in the required batch meets the temperature requirement. The contamination atom excludes unresolved in-scope contamination, while the unportioned required rice guarantees a remaining completion step. Unspecified resting or draining facts cannot improve or worsen the result beyond level 1: failures there would also be incompleteness, not a stated safety failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the Tuesday-dinner readiness assessment, batch B42 is the meal planner's required lemon-chicken component."}, {"id": "a2", "statement": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 measures 68°C."}, {"id": "a3", "statement": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is a piece of chicken in batch B42."}, {"id": "a4", "statement": "At the Tuesday-dinner readiness assessment, every piece of chicken in batch B42 other than specimen S17 has reached at least 74°C."}, {"id": "a5", "statement": "At the Tuesday-dinner readiness assessment, zero of the three required meal components has unresolved contamination."}, {"id": "a6", "statement": "Before the Tuesday-dinner readiness assessment, the required rice rested for six minutes."}, {"id": "a7", "statement": "The required rice remained covered throughout its rest before the Tuesday-dinner readiness assessment."}, {"id": "a8", "statement": "At the Tuesday-dinner readiness assessment, the required broccoli is drained."}, {"id": "a9", "statement": "At the Tuesday-dinner readiness assessment, the required rice has not been portioned for immediate serving."}], "base_state_json": "\"For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli. The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving. Readiness covers only the three required meal components. Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions.\\n\\nAt the Tuesday-dinner readiness assessment, batch B42 was the meal planner’s required lemon-chicken component. At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63. At the Tuesday-dinner readiness assessment, a chicken piece in batch B42 bore identifier tag Q63, and no other item bore that tag. Specimen S17 measured 68°C. Every item that was both a chicken piece in B42 and not S17 had reached at least 74°C. Inspectors found unresolved contamination in zero of the three required components. Before the assessment, the required rice rested for six minutes and remained covered throughout. The required broccoli was drained. The required rice had not been portioned for immediate serving.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63."}, {"path": [], "text": "At the Tuesday-dinner readiness assessment, a chicken piece in batch B42 bore identifier tag Q63, and no other item bore that tag."}], "policy_evidence": [{"path": [], "text": "For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli."}, {"path": [], "text": "The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving."}, {"path": [], "text": "Readiness covers only the three required meal components."}, {"path": [], "text": "Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions."}], "rules": [{"justification": "S17 is part of the required chicken batch and measures only 68°C, below the mandatory 74°C requirement, so a mandatory food-safety requirement fails.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "Because S17 is not in B42, the universal temperature fact covers every piece of the required chicken batch; no required component has unresolved contamination, so the in-scope safety requirements pass. The required rice remains unportioned, making the meal safe but incomplete.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63.", "negative_left": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63.", "negative_right": "At the Tuesday-dinner readiness assessment, a chicken piece outside batch B42 bore identifier tag Q63, and no other item bore that tag.", "right": "At the Tuesday-dinner readiness assessment, a chicken piece in batch B42 bore identifier tag Q63, and no other item bore that tag."}, "verifier_independent_model": false}, "family": "scale-diverse-192-006", "id": "scale-diverse-192-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready—at least one in-scope mandatory food-safety requirement fails, or contamination still requires correction; the meal must not be served.", "Safe but incomplete—all in-scope safety requirements pass, but at least one required component still needs cooking, resting, draining, or portioning.", "Ready with a minor in-scope presentation or workflow issue—all required components are safe and complete, but a non-safety serving step remains; explicit exceptions do not count.", "Fully ready—all required components are safe, cooked, rested or drained as specified, and portioned for immediate serving, with no unresolved in-scope issue."], "instructions": "Rate the meal’s current readiness to serve using the ordered levels below. Apply requirements only to the stated readiness scope; do not lower the rating for explicit exceptions. A mandatory food-safety failure takes precedence over otherwise completed components.", "type": "score"}}, "state": "For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli. The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving. Readiness covers only the three required meal components. Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions.\n\nAt the Tuesday-dinner readiness assessment, batch B42 was the meal planner’s required lemon-chicken component. At the Tuesday-dinner readiness assessment, thermometer specimen S17 bore identifier tag Q63. At the Tuesday-dinner readiness assessment, a chicken piece outside batch B42 bore identifier tag Q63, and no other item bore that tag. Specimen S17 measured 68°C. Every item that was both a chicken piece in B42 and not S17 had reached at least 74°C. Inspectors found unresolved contamination in zero of the three required components. Before the assessment, the required rice rested for six minutes and remained covered throughout. The required broccoli was drained. The required rice had not been portioned for immediate serving."}, "method": "c2d", "provenance": {"source_id": "diverse-192", "source_is_synthetic": true, "source_sha256": "f16699a58ce6f897e6254823f6038b87cdbce46e187582c3dc5640f72ae7c868", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring policy, while both contexts retain the original meal requirements, readiness scope, and explicit exceptions. The Tuesday-dinner meal, required components, assessment time, and decision path remain fixed. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only tray T9’s chicken from belonging to batch B42 to not belonging to B42; because S17 occupies T9, this makes S17 outside B42 and does not conflict with the temperature statement governing B42 pieces. Neither context contains an answer label, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the universally quantified batch and contamination propositions. The focus is the factual membership relation between S17 and B42. The base and counter assignments differ only on that membership and are jointly realizable: in the base S17 is the under-temperature batch piece, while in the counter it is outside B42 and every actual B42 piece is at least 74°C. Policy evidence preserves the needed state-originated component bindings, preparation thresholds, scope, and exceptions; decision-level priorities and classifications remain automatically preserved in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that B42 is the required chicken component and that one of its pieces is at 68°C at assessment, below the state-specified mandatory 74°C threshold. This is sufficient for level 0 regardless of the other components.", "rule_index": 0, "sound": true}, {"reason": "Refuting S17's membership in B42, together with every other B42 piece being at least 74°C, entails that all chicken in the required batch meets the temperature requirement. The contamination atom excludes unresolved in-scope contamination, while the unportioned required rice guarantees a remaining completion step. Unspecified resting or draining facts cannot improve or worsen the result beyond level 1: failures there would also be incompleteness, not a stated safety failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the Tuesday-dinner readiness assessment, batch B42 is the meal planner's required lemon-chicken component."}, {"id": "a2", "statement": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 measures 68°C."}, {"id": "a3", "statement": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is a piece of chicken in batch B42."}, {"id": "a4", "statement": "At the Tuesday-dinner readiness assessment, every piece of chicken in batch B42 other than specimen S17 has reached at least 74°C."}, {"id": "a5", "statement": "At the Tuesday-dinner readiness assessment, zero of the three required meal components has unresolved contamination."}, {"id": "a6", "statement": "Before the Tuesday-dinner readiness assessment, the required rice rested for six minutes."}, {"id": "a7", "statement": "The required rice remained covered throughout its rest before the Tuesday-dinner readiness assessment."}, {"id": "a8", "statement": "At the Tuesday-dinner readiness assessment, the required broccoli is drained."}, {"id": "a9", "statement": "At the Tuesday-dinner readiness assessment, the required rice has not been portioned for immediate serving."}], "base_state_json": "\"For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli. The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving. Readiness covers only the three required meal components. Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions. The kitchen handoff log identifies batch B42 as that required lemon-chicken component. At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9. At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken in batch B42. At the assessment, S17’s calibrated reading is 68°C. For each piece of chicken belonging to B42, if it is not S17, the logged temperature is at least 74°C. The safety audit records unresolved contamination for none of the three required meal components. Before the assessment, the required rice rested six minutes and remained covered throughout that rest. The required broccoli is drained. The required rice has not been portioned for immediate serving.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9."}, {"path": [], "text": "At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken in batch B42."}], "policy_evidence": [{"path": [], "text": "For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli."}, {"path": [], "text": "The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving."}, {"path": [], "text": "Readiness covers only the three required meal components."}, {"path": [], "text": "Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions."}], "rules": [{"justification": "S17 is part of the required chicken batch and measures only 68°C, below the mandatory 74°C requirement, so a mandatory food-safety requirement fails.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "Because S17 is not in B42, the universal temperature fact covers every piece of the required chicken batch; no required component has unresolved contamination, so the in-scope safety requirements pass. The required rice remains unportioned, making the meal safe but incomplete.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9.", "negative_left": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9.", "negative_right": "At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken that is not in batch B42.", "right": "At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken in batch B42."}, "verifier_independent_model": false}, "family": "scale-diverse-192-007", "id": "scale-diverse-192-007-base", "input": {"questions": {"decision": {"criteria": ["Not ready—at least one in-scope mandatory food-safety requirement fails, or contamination still requires correction; the meal must not be served.", "Safe but incomplete—all in-scope safety requirements pass, but at least one required component still needs cooking, resting, draining, or portioning.", "Ready with a minor in-scope presentation or workflow issue—all required components are safe and complete, but a non-safety serving step remains; explicit exceptions do not count.", "Fully ready—all required components are safe, cooked, rested or drained as specified, and portioned for immediate serving, with no unresolved in-scope issue."], "instructions": "Rate the meal’s current readiness to serve using the ordered levels below. Apply requirements only to the stated readiness scope; do not lower the rating for explicit exceptions. A mandatory food-safety failure takes precedence over otherwise completed components.", "type": "score"}}, "state": "For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli. The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving. Readiness covers only the three required meal components. Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions. The kitchen handoff log identifies batch B42 as that required lemon-chicken component. At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9. At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken in batch B42. At the assessment, S17’s calibrated reading is 68°C. For each piece of chicken belonging to B42, if it is not S17, the logged temperature is at least 74°C. The safety audit records unresolved contamination for none of the three required meal components. Before the assessment, the required rice rested six minutes and remained covered throughout that rest. The required broccoli is drained. The required rice has not been portioned for immediate serving."}, "method": "c2d", "provenance": {"source_id": "diverse-192", "source_is_synthetic": true, "source_sha256": "f16699a58ce6f897e6254823f6038b87cdbce46e187582c3dc5640f72ae7c868", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring policy, while both contexts retain the original meal requirements, readiness scope, and explicit exceptions. The Tuesday-dinner meal, required components, assessment time, and decision path remain fixed. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only tray T9’s chicken from belonging to batch B42 to not belonging to B42; because S17 occupies T9, this makes S17 outside B42 and does not conflict with the temperature statement governing B42 pieces. Neither context contains an answer label, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the universally quantified batch and contamination propositions. The focus is the factual membership relation between S17 and B42. The base and counter assignments differ only on that membership and are jointly realizable: in the base S17 is the under-temperature batch piece, while in the counter it is outside B42 and every actual B42 piece is at least 74°C. Policy evidence preserves the needed state-originated component bindings, preparation thresholds, scope, and exceptions; decision-level priorities and classifications remain automatically preserved in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that B42 is the required chicken component and that one of its pieces is at 68°C at assessment, below the state-specified mandatory 74°C threshold. This is sufficient for level 0 regardless of the other components.", "rule_index": 0, "sound": true}, {"reason": "Refuting S17's membership in B42, together with every other B42 piece being at least 74°C, entails that all chicken in the required batch meets the temperature requirement. The contamination atom excludes unresolved in-scope contamination, while the unportioned required rice guarantees a remaining completion step. Unspecified resting or draining facts cannot improve or worsen the result beyond level 1: failures there would also be incompleteness, not a stated safety failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the Tuesday-dinner readiness assessment, batch B42 is the meal planner's required lemon-chicken component."}, {"id": "a2", "statement": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 measures 68°C."}, {"id": "a3", "statement": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is a piece of chicken in batch B42."}, {"id": "a4", "statement": "At the Tuesday-dinner readiness assessment, every piece of chicken in batch B42 other than specimen S17 has reached at least 74°C."}, {"id": "a5", "statement": "At the Tuesday-dinner readiness assessment, zero of the three required meal components has unresolved contamination."}, {"id": "a6", "statement": "Before the Tuesday-dinner readiness assessment, the required rice rested for six minutes."}, {"id": "a7", "statement": "The required rice remained covered throughout its rest before the Tuesday-dinner readiness assessment."}, {"id": "a8", "statement": "At the Tuesday-dinner readiness assessment, the required broccoli is drained."}, {"id": "a9", "statement": "At the Tuesday-dinner readiness assessment, the required rice has not been portioned for immediate serving."}], "base_state_json": "\"For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli. The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving. Readiness covers only the three required meal components. Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions. The kitchen handoff log identifies batch B42 as that required lemon-chicken component. At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9. At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken in batch B42. At the assessment, S17’s calibrated reading is 68°C. For each piece of chicken belonging to B42, if it is not S17, the logged temperature is at least 74°C. The safety audit records unresolved contamination for none of the three required meal components. Before the assessment, the required rice rested six minutes and remained covered throughout that rest. The required broccoli is drained. The required rice has not been portioned for immediate serving.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9."}, {"path": [], "text": "At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken in batch B42."}], "policy_evidence": [{"path": [], "text": "For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli."}, {"path": [], "text": "The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving."}, {"path": [], "text": "Readiness covers only the three required meal components."}, {"path": [], "text": "Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions."}], "rules": [{"justification": "S17 is part of the required chicken batch and measures only 68°C, below the mandatory 74°C requirement, so a mandatory food-safety requirement fails.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "Because S17 is not in B42, the universal temperature fact covers every piece of the required chicken batch; no required component has unresolved contamination, so the in-scope safety requirements pass. The required rice remains unportioned, making the meal safe but incomplete.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9.", "negative_left": "At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9.", "negative_right": "At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken that is not in batch B42.", "right": "At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken in batch B42."}, "verifier_independent_model": false}, "family": "scale-diverse-192-007", "id": "scale-diverse-192-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready—at least one in-scope mandatory food-safety requirement fails, or contamination still requires correction; the meal must not be served.", "Safe but incomplete—all in-scope safety requirements pass, but at least one required component still needs cooking, resting, draining, or portioning.", "Ready with a minor in-scope presentation or workflow issue—all required components are safe and complete, but a non-safety serving step remains; explicit exceptions do not count.", "Fully ready—all required components are safe, cooked, rested or drained as specified, and portioned for immediate serving, with no unresolved in-scope issue."], "instructions": "Rate the meal’s current readiness to serve using the ordered levels below. Apply requirements only to the stated readiness scope; do not lower the rating for explicit exceptions. A mandatory food-safety failure takes precedence over otherwise completed components.", "type": "score"}}, "state": "For Tuesday dinner, the meal planner specifies lemon chicken, rice, and broccoli. The recipe requires chicken to reach 74°C, rice to rest covered for five minutes, and broccoli to be drained before serving. Readiness covers only the three required meal components. Optional parsley, unused equipment, and storage of leftovers after dinner are explicit exceptions. The kitchen handoff log identifies batch B42 as that required lemon-chicken component. At the Tuesday-dinner readiness assessment, thermometer specimen S17 is the food item occupying handoff tray T9. At the Tuesday-dinner readiness assessment, every food item occupying handoff tray T9 is a piece of chicken that is not in batch B42. At the assessment, S17’s calibrated reading is 68°C. For each piece of chicken belonging to B42, if it is not S17, the logged temperature is at least 74°C. The safety audit records unresolved contamination for none of the three required meal components. Before the assessment, the required rice rested six minutes and remained covered throughout that rest. The required broccoli is drained. The required rice has not been portioned for immediate serving."}, "method": "c2d", "provenance": {"source_id": "diverse-192", "source_is_synthetic": true, "source_sha256": "f16699a58ce6f897e6254823f6038b87cdbce46e187582c3dc5640f72ae7c868", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same current-routing task, fictional Alderline three-shelf unit, decision time, and latest-applicable-evidence policy without adding exceptions or defaults. The two evidence spans are complete factual sentences. The counterfactual changes only AL29’s wobble measurement from 0 to 6 millimetres; this remains compatible with the nonnegative measurement domain, zero fastening play, and otherwise correct parts and configuration. Neither context states a routing answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"On 18 May 2042, a component and configuration inspection at 14:23 UTC was the latest applicable inspection evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25. It found every instruction-required component present, every present component undamaged and required by the instructions, and every installed component identical to the one specified. Every observed assembly-configuration feature matched its corresponding instruction. The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC. The calibrated AL29 system expressed wobble amplitude in millimetres as a nonnegative real-number measurement. That test recorded maximum fastening play of 0 millimetres. Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 0 millimetres.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC."}, {"path": [], "text": "Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC.", "negative_left": "The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC.", "negative_right": "Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 6 millimetres.", "right": "Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 0 millimetres."}, "verifier_independent_model": false}, "family": "scale-diverse-194-001", "id": "scale-diverse-194-001-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "On 18 May 2042, a component and configuration inspection at 14:23 UTC was the latest applicable inspection evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25. It found every instruction-required component present, every present component undamaged and required by the instructions, and every installed component identical to the one specified. Every observed assembly-configuration feature matched its corresponding instruction. The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC. The calibrated AL29 system expressed wobble amplitude in millimetres as a nonnegative real-number measurement. That test recorded maximum fastening play of 0 millimetres. Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 0 millimetres."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same current-routing task, fictional Alderline three-shelf unit, decision time, and latest-applicable-evidence policy without adding exceptions or defaults. The two evidence spans are complete factual sentences. The counterfactual changes only AL29’s wobble measurement from 0 to 6 millimetres; this remains compatible with the nonnegative measurement domain, zero fastening play, and otherwise correct parts and configuration. Neither context states a routing answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"On 18 May 2042, a component and configuration inspection at 14:23 UTC was the latest applicable inspection evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25. It found every instruction-required component present, every present component undamaged and required by the instructions, and every installed component identical to the one specified. Every observed assembly-configuration feature matched its corresponding instruction. The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC. The calibrated AL29 system expressed wobble amplitude in millimetres as a nonnegative real-number measurement. That test recorded maximum fastening play of 0 millimetres. Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 0 millimetres.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC."}, {"path": [], "text": "Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC.", "negative_left": "The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC.", "negative_right": "Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 6 millimetres.", "right": "Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 0 millimetres."}, "verifier_independent_model": false}, "family": "scale-diverse-194-001", "id": "scale-diverse-194-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "On 18 May 2042, a component and configuration inspection at 14:23 UTC was the latest applicable inspection evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25. It found every instruction-required component present, every present component undamaged and required by the instructions, and every installed component identical to the one specified. Every observed assembly-configuration feature matched its corresponding instruction. The fictional Alderline three-shelf unit's current-routing decision immediately after 14:25 UTC on 18 May 2042 had exactly two applicable load-and-push tests completed beforehand, test AL17 at 14:11 UTC and the later test AL29 at 14:24 UTC. The calibrated AL29 system expressed wobble amplitude in millimetres as a nonnegative real-number measurement. That test recorded maximum fastening play of 0 millimetres. Test AL29 recorded the fictional Alderline three-shelf unit's wobble amplitude as 6 millimetres."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "tightening"}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original latest-evidence routing policy without adding exceptions, priorities, or missing-evidence defaults. They preserve the decision question’s Alderline-unit entity, current-routing scope, and immediately-after-14:25 time binding. The focus evidence contains exactly two complete factual sentences, identifying the latest applicable run and its base wobble measurement. The counterfactual changes only that measurement from 0 to 6 millimetres; this does not contradict the separate zero fastening-play measurement or the unchanged parts and configuration findings. Neither context includes a route label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"Operational handoff — current-routing decision immediately after 14:25. For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25. Inspection run LP-731 recorded a wobble amplitude of 0 millimetres. The wobble field is an instrument measurement in millimetres, represented by a real number within a range that begins at zero. The same run recorded maximum fastening play of 0 millimetres. The latest signed inventory confirms that every component required by the instructions is present; every present component is undamaged and required, with no extras. Identity checks confirm that every installed component is the one specified by the instructions. The accompanying configuration inspection covers all observed assembly features, including orientation, sequence, and hole choice, and records each as matching the corresponding instruction. Earlier workshop observations are archived and are not part of this handoff.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25."}, {"path": [], "text": "Inspection run LP-731 recorded a wobble amplitude of 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25.", "negative_left": "For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25.", "negative_right": "Inspection run LP-731 recorded a wobble amplitude of 6 millimetres.", "right": "Inspection run LP-731 recorded a wobble amplitude of 0 millimetres."}, "verifier_independent_model": false}, "family": "scale-diverse-194-003", "id": "scale-diverse-194-003-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "Operational handoff — current-routing decision immediately after 14:25. For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25. Inspection run LP-731 recorded a wobble amplitude of 0 millimetres. The wobble field is an instrument measurement in millimetres, represented by a real number within a range that begins at zero. The same run recorded maximum fastening play of 0 millimetres. The latest signed inventory confirms that every component required by the instructions is present; every present component is undamaged and required, with no extras. Identity checks confirm that every installed component is the one specified by the instructions. The accompanying configuration inspection covers all observed assembly features, including orientation, sequence, and hole choice, and records each as matching the corresponding instruction. Earlier workshop observations are archived and are not part of this handoff."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original latest-evidence routing policy without adding exceptions, priorities, or missing-evidence defaults. They preserve the decision question’s Alderline-unit entity, current-routing scope, and immediately-after-14:25 time binding. The focus evidence contains exactly two complete factual sentences, identifying the latest applicable run and its base wobble measurement. The counterfactual changes only that measurement from 0 to 6 millimetres; this does not contradict the separate zero fastening-play measurement or the unchanged parts and configuration findings. Neither context includes a route label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"Operational handoff — current-routing decision immediately after 14:25. For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25. Inspection run LP-731 recorded a wobble amplitude of 0 millimetres. The wobble field is an instrument measurement in millimetres, represented by a real number within a range that begins at zero. The same run recorded maximum fastening play of 0 millimetres. The latest signed inventory confirms that every component required by the instructions is present; every present component is undamaged and required, with no extras. Identity checks confirm that every installed component is the one specified by the instructions. The accompanying configuration inspection covers all observed assembly features, including orientation, sequence, and hole choice, and records each as matching the corresponding instruction. Earlier workshop observations are archived and are not part of this handoff.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25."}, {"path": [], "text": "Inspection run LP-731 recorded a wobble amplitude of 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25.", "negative_left": "For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25.", "negative_right": "Inspection run LP-731 recorded a wobble amplitude of 6 millimetres.", "right": "Inspection run LP-731 recorded a wobble amplitude of 0 millimetres."}, "verifier_independent_model": false}, "family": "scale-diverse-194-003", "id": "scale-diverse-194-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "Operational handoff — current-routing decision immediately after 14:25. For the fictional Alderline three-shelf unit, inspection run LP-731 was the latest applicable load-and-push test at the current-routing decision immediately after 14:25. Inspection run LP-731 recorded a wobble amplitude of 6 millimetres. The wobble field is an instrument measurement in millimetres, represented by a real number within a range that begins at zero. The same run recorded maximum fastening play of 0 millimetres. The latest signed inventory confirms that every component required by the instructions is present; every present component is undamaged and required, with no extras. Identity checks confirm that every installed component is the one specified by the instructions. The accompanying configuration inspection covers all observed assembly features, including orientation, sequence, and hole choice, and records each as matching the corresponding instruction. Earlier workshop observations are archived and are not part of this handoff."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "tightening"}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all routing criteria and latest-evidence instructions, while neither context removes or alters any policy from the original state. Both contexts retain the same fictional unit, current-routing request, and immediately-after-14:25 time binding. The two evidence spans are complete factual sentences identifying the latest applicable test and its measured wobble amplitude. Changing that amplitude from 0 to 6 millimetres is a coherent single factual change and does not contradict the unchanged zero fastening-play, component-audit, or configuration-inspection findings. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"Field note — 14:25 routing review for the fictional Alderline three-shelf unit. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test. The test instrument reports wobble amplitude in millimetres as a nonnegative real-number measurement. Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 0 millimetres. LP-47 also recorded maximum fastening play of 0 millimetres.\\n\\nThe latest applicable component audit was checked during the review. It found every instruction-required component present, every present component undamaged, and every present component required by the instructions. It also verified that each installed component had the identity specified by the instructions. The latest applicable configuration inspection found that every observed feature—including shelf orientation, assembly sequence, selected holes, and rear-brace placement—matched the corresponding instruction. No evidence predating these named checks was used for the current-routing decision.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test."}, {"path": [], "text": "Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test.", "negative_left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test.", "negative_right": "Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 6 millimetres.", "right": "Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 0 millimetres."}, "verifier_independent_model": false}, "family": "scale-diverse-194-004", "id": "scale-diverse-194-004-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "Field note — 14:25 routing review for the fictional Alderline three-shelf unit. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test. The test instrument reports wobble amplitude in millimetres as a nonnegative real-number measurement. Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 0 millimetres. LP-47 also recorded maximum fastening play of 0 millimetres.\n\nThe latest applicable component audit was checked during the review. It found every instruction-required component present, every present component undamaged, and every present component required by the instructions. It also verified that each installed component had the identity specified by the instructions. The latest applicable configuration inspection found that every observed feature—including shelf orientation, assembly sequence, selected holes, and rear-brace placement—matched the corresponding instruction. No evidence predating these named checks was used for the current-routing decision."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all routing criteria and latest-evidence instructions, while neither context removes or alters any policy from the original state. Both contexts retain the same fictional unit, current-routing request, and immediately-after-14:25 time binding. The two evidence spans are complete factual sentences identifying the latest applicable test and its measured wobble amplitude. Changing that amplitude from 0 to 6 millimetres is a coherent single factual change and does not contradict the unchanged zero fastening-play, component-audit, or configuration-inspection findings. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"Field note — 14:25 routing review for the fictional Alderline three-shelf unit. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test. The test instrument reports wobble amplitude in millimetres as a nonnegative real-number measurement. Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 0 millimetres. LP-47 also recorded maximum fastening play of 0 millimetres.\\n\\nThe latest applicable component audit was checked during the review. It found every instruction-required component present, every present component undamaged, and every present component required by the instructions. It also verified that each installed component had the identity specified by the instructions. The latest applicable configuration inspection found that every observed feature—including shelf orientation, assembly sequence, selected holes, and rear-brace placement—matched the corresponding instruction. No evidence predating these named checks was used for the current-routing decision.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test."}, {"path": [], "text": "Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test.", "negative_left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test.", "negative_right": "Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 6 millimetres.", "right": "Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 0 millimetres."}, "verifier_independent_model": false}, "family": "scale-diverse-194-004", "id": "scale-diverse-194-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "Field note — 14:25 routing review for the fictional Alderline three-shelf unit. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, field test LP-47, completed at 14:23, was the latest applicable load-and-push test. The test instrument reports wobble amplitude in millimetres as a nonnegative real-number measurement. Field test LP-47 on the fictional Alderline three-shelf unit recorded a wobble amplitude of 6 millimetres. LP-47 also recorded maximum fastening play of 0 millimetres.\n\nThe latest applicable component audit was checked during the review. It found every instruction-required component present, every present component undamaged, and every present component required by the instructions. It also verified that each installed component had the identity specified by the instructions. The latest applicable configuration inspection found that every observed feature—including shelf orientation, assembly sequence, selected holes, and rear-brace placement—matched the corresponding instruction. No evidence predating these named checks was used for the current-routing decision."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "tightening"}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts retain the original cabinet inspection rule, image-conflict policy, yes/no request, cabinet entity, and final-inspection scope. The two focus spans are complete factual sentences; the counterfactual changes only the later photograph’s panel orientation, remains consistent with the marker identifying the smooth face, and introduces no duplicate contradiction or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"The completion packet for the fictional Alderline C-4 cabinet was transferred from technician Mara Venn to the final-inspection desk. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.\",\"evidence\":[\"At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet.\",\"At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing toward the fictional Alderline C-4 cabinet’s interior.\",\"At 16:14 UTC, the handoff checklist records the first required anti-tip strap as installed and secured to its designated anchor.\",\"At 16:15 UTC, the handoff checklist separately records the second required anti-tip strap as installed and secured to its designated anchor.\",\"At 16:18 UTC, final-use tester Ivo Chen reports that the cabinet did not rock during testing on the prepared inspection surface.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet."}, {"path": ["evidence", "1"], "text": "At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing toward the fictional Alderline C-4 cabinet’s interior."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet.", "negative_left": "At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet.", "negative_right": "At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing away from the fictional Alderline C-4 cabinet’s interior.", "right": "At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing toward the fictional Alderline C-4 cabinet’s interior."}, "verifier_independent_model": false}, "family": "scale-diverse-195-003", "id": "scale-diverse-195-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "The completion packet for the fictional Alderline C-4 cabinet was transferred from technician Mara Venn to the final-inspection desk. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.", "evidence": ["At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet.", "At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing toward the fictional Alderline C-4 cabinet’s interior.", "At 16:14 UTC, the handoff checklist records the first required anti-tip strap as installed and secured to its designated anchor.", "At 16:15 UTC, the handoff checklist separately records the second required anti-tip strap as installed and secured to its designated anchor.", "At 16:18 UTC, final-use tester Ivo Chen reports that the cabinet did not rock during testing on the prepared inspection surface."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both constructed contexts retain the original cabinet inspection rule, image-conflict policy, yes/no request, cabinet entity, and final-inspection scope. The two focus spans are complete factual sentences; the counterfactual changes only the later photograph’s panel orientation, remains consistent with the marker identifying the smooth face, and introduces no duplicate contradiction or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"The completion packet for the fictional Alderline C-4 cabinet was transferred from technician Mara Venn to the final-inspection desk. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.\",\"evidence\":[\"At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet.\",\"At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing toward the fictional Alderline C-4 cabinet’s interior.\",\"At 16:14 UTC, the handoff checklist records the first required anti-tip strap as installed and secured to its designated anchor.\",\"At 16:15 UTC, the handoff checklist separately records the second required anti-tip strap as installed and secured to its designated anchor.\",\"At 16:18 UTC, final-use tester Ivo Chen reports that the cabinet did not rock during testing on the prepared inspection surface.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet."}, {"path": ["evidence", "1"], "text": "At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing toward the fictional Alderline C-4 cabinet’s interior."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet.", "negative_left": "At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet.", "negative_right": "At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing away from the fictional Alderline C-4 cabinet’s interior.", "right": "At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing toward the fictional Alderline C-4 cabinet’s interior."}, "verifier_independent_model": false}, "family": "scale-diverse-195-003", "id": "scale-diverse-195-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "The completion packet for the fictional Alderline C-4 cabinet was transferred from technician Mara Venn to the final-inspection desk. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.", "evidence": ["At 16:08 UTC during final inspection, technician Mara Venn placed orange marker 47Q on the smooth face of the installed back panel in the fictional Alderline C-4 cabinet.", "At 16:12 UTC during the same final inspection, handoff photograph 8831 shows the installed back-panel side bearing orange marker 47Q facing away from the fictional Alderline C-4 cabinet’s interior.", "At 16:14 UTC, the handoff checklist records the first required anti-tip strap as installed and secured to its designated anchor.", "At 16:15 UTC, the handoff checklist separately records the second required anti-tip strap as installed and secured to its designated anchor.", "At 16:18 UTC, final-use tester Ivo Chen reports that the cabinet did not rock during testing on the prepared inspection surface."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original inspection rule and image-conflict policy, while the unchanged questions object preserves the instructions and decision criteria. The request remains bound to the fictional Alderline C-4 cabinet and final build inspection. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the photographed orientation from interior to exterior; together with the unchanged tactile identification, this implies the smooth face points exterior without creating contradictory duplicate measurements or assertions. Neither context embeds an answer, code, proposition identifier, label rationale, rule table, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"Final field note for the fictional Alderline C-4 cabinet. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.\",\"evidence\":[\"At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet.\",\"At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s interior.\",\"At 14:26 UTC, the hardware inspector recorded the first required anti-tip strap as installed and secured to its designated anchor.\",\"At 14:27 UTC, the hardware inspector separately recorded the second required anti-tip strap as installed and secured to its designated anchor.\",\"At 14:31 UTC, the final-use tester applied pressure at each upper corner and observed no rocking.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet."}, {"path": ["evidence", "1"], "text": "At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s interior."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet.", "negative_left": "At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet.", "negative_right": "At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s exterior.", "right": "At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s interior."}, "verifier_independent_model": false}, "family": "scale-diverse-195-004", "id": "scale-diverse-195-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "Final field note for the fictional Alderline C-4 cabinet. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.", "evidence": ["At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet.", "At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s interior.", "At 14:26 UTC, the hardware inspector recorded the first required anti-tip strap as installed and secured to its designated anchor.", "At 14:27 UTC, the hardware inspector separately recorded the second required anti-tip strap as installed and secured to its designated anchor.", "At 14:31 UTC, the final-use tester applied pressure at each upper corner and observed no rocking."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original inspection rule and image-conflict policy, while the unchanged questions object preserves the instructions and decision criteria. The request remains bound to the fictional Alderline C-4 cabinet and final build inspection. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the photographed orientation from interior to exterior; together with the unchanged tactile identification, this implies the smooth face points exterior without creating contradictory duplicate measurements or assertions. Neither context embeds an answer, code, proposition identifier, label rationale, rule table, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"Final field note for the fictional Alderline C-4 cabinet. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.\",\"evidence\":[\"At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet.\",\"At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s interior.\",\"At 14:26 UTC, the hardware inspector recorded the first required anti-tip strap as installed and secured to its designated anchor.\",\"At 14:27 UTC, the hardware inspector separately recorded the second required anti-tip strap as installed and secured to its designated anchor.\",\"At 14:31 UTC, the final-use tester applied pressure at each upper corner and observed no rocking.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet."}, {"path": ["evidence", "1"], "text": "At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s interior."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet.", "negative_left": "At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet.", "negative_right": "At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s exterior.", "right": "At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s interior."}, "verifier_independent_model": false}, "family": "scale-diverse-195-004", "id": "scale-diverse-195-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "Final field note for the fictional Alderline C-4 cabinet. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.", "evidence": ["At 14:22 UTC on 18 May 2026, the final-inspection tactile scan identified the face bearing green tag G17 as the smooth face of the back panel installed in the fictional Alderline C-4 cabinet.", "At 14:24 UTC on 18 May 2026, a final-inspection image showed the back panel face bearing green tag G17 directed toward the fictional Alderline C-4 cabinet’s exterior.", "At 14:26 UTC, the hardware inspector recorded the first required anti-tip strap as installed and secured to its designated anchor.", "At 14:27 UTC, the hardware inspector separately recorded the second required anti-tip strap as installed and secured to its designated anchor.", "At 14:31 UTC, the final-use tester applied pressure at each upper corner and observed no rocking."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts leave the original verification rubric unchanged and retain the Larkspur unit, step-6 back-panel checkpoint, final-placement decision, and relevant dates. The two focus-evidence spans are complete factual sentences. In the base context, custody through C-418 links the inspected unit to the submitted Larkspur unit; the counterfactual coherently breaks that link by placing the inspected unit in separate crate D-627 through 09:10, while C-418’s contents are submitted at 08:40. This does not contradict the remaining statements, including the absence of confirmations other than session S. Neither constructed context states an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Quality lead\",\"text\":\"For the physical Larkspur unit submitted for final placement on 2026-09-17, every required safety test had a passing result at the placement decision. Every concealed-orientation checkpoint except the step-6 back-panel checkpoint had qualifying evidence then.\"},{\"speaker\":\"Records clerk\",\"text\":\"At the placement decision, zero required progress images were available for that unit’s step-6 back-panel checkpoint.\"},{\"speaker\":\"Inspector\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, I directly observed the inspected unit’s back-panel grooved face oriented inward. That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"},{\"speaker\":\"Custody log\",\"text\":\"Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate C-418, and uninterrupted custody footage shows that nothing entered or left C-418 before its opening at 08:40 on 2026-09-17.\"},{\"speaker\":\"Intake log\",\"text\":\"When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit.\"},{\"speaker\":\"Decision clerk\",\"text\":\"Excluding session S, zero current direct visual confirmations of the submitted Larkspur unit’s step-6 back-panel orientation were available at the placement decision.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["3", "text"], "text": "Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate C-418, and uninterrupted custody footage shows that nothing entered or left C-418 before its opening at 08:40 on 2026-09-17."}, {"path": ["4", "text"], "text": "When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate C-418, and uninterrupted custody footage shows that nothing entered or left C-418 before its opening at 08:40 on 2026-09-17.", "negative_left": "Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate D-627, which was physically separate from crate C-418 and remained locked with that unit inside until 09:10 on 2026-09-17.", "negative_right": "When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit.", "right": "When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit."}, "verifier_independent_model": false}, "family": "scale-diverse-196-001", "id": "scale-diverse-196-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Quality lead", "text": "For the physical Larkspur unit submitted for final placement on 2026-09-17, every required safety test had a passing result at the placement decision. Every concealed-orientation checkpoint except the step-6 back-panel checkpoint had qualifying evidence then."}, {"speaker": "Records clerk", "text": "At the placement decision, zero required progress images were available for that unit’s step-6 back-panel checkpoint."}, {"speaker": "Inspector", "text": "During pre-placement visual-inspection session S on 2026-09-16, I directly observed the inspected unit’s back-panel grooved face oriented inward. That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Custody log", "text": "Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate C-418, and uninterrupted custody footage shows that nothing entered or left C-418 before its opening at 08:40 on 2026-09-17."}, {"speaker": "Intake log", "text": "When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit."}, {"speaker": "Decision clerk", "text": "Excluding session S, zero current direct visual confirmations of the submitted Larkspur unit’s step-6 back-panel orientation were available at the placement decision."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts leave the original verification rubric unchanged and retain the Larkspur unit, step-6 back-panel checkpoint, final-placement decision, and relevant dates. The two focus-evidence spans are complete factual sentences. In the base context, custody through C-418 links the inspected unit to the submitted Larkspur unit; the counterfactual coherently breaks that link by placing the inspected unit in separate crate D-627 through 09:10, while C-418’s contents are submitted at 08:40. This does not contradict the remaining statements, including the absence of confirmations other than session S. Neither constructed context states an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Quality lead\",\"text\":\"For the physical Larkspur unit submitted for final placement on 2026-09-17, every required safety test had a passing result at the placement decision. Every concealed-orientation checkpoint except the step-6 back-panel checkpoint had qualifying evidence then.\"},{\"speaker\":\"Records clerk\",\"text\":\"At the placement decision, zero required progress images were available for that unit’s step-6 back-panel checkpoint.\"},{\"speaker\":\"Inspector\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, I directly observed the inspected unit’s back-panel grooved face oriented inward. That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"},{\"speaker\":\"Custody log\",\"text\":\"Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate C-418, and uninterrupted custody footage shows that nothing entered or left C-418 before its opening at 08:40 on 2026-09-17.\"},{\"speaker\":\"Intake log\",\"text\":\"When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit.\"},{\"speaker\":\"Decision clerk\",\"text\":\"Excluding session S, zero current direct visual confirmations of the submitted Larkspur unit’s step-6 back-panel orientation were available at the placement decision.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["3", "text"], "text": "Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate C-418, and uninterrupted custody footage shows that nothing entered or left C-418 before its opening at 08:40 on 2026-09-17."}, {"path": ["4", "text"], "text": "When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate C-418, and uninterrupted custody footage shows that nothing entered or left C-418 before its opening at 08:40 on 2026-09-17.", "negative_left": "Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate D-627, which was physically separate from crate C-418 and remained locked with that unit inside until 09:10 on 2026-09-17.", "negative_right": "When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit.", "right": "When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit."}, "verifier_independent_model": false}, "family": "scale-diverse-196-001", "id": "scale-diverse-196-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Quality lead", "text": "For the physical Larkspur unit submitted for final placement on 2026-09-17, every required safety test had a passing result at the placement decision. Every concealed-orientation checkpoint except the step-6 back-panel checkpoint had qualifying evidence then."}, {"speaker": "Records clerk", "text": "At the placement decision, zero required progress images were available for that unit’s step-6 back-panel checkpoint."}, {"speaker": "Inspector", "text": "During pre-placement visual-inspection session S on 2026-09-16, I directly observed the inspected unit’s back-panel grooved face oriented inward. That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Custody log", "text": "Immediately after pre-placement visual-inspection session S on 2026-09-16, the inspector locked the physical unit inspected during that session alone inside previously empty crate D-627, which was physically separate from crate C-418 and remained locked with that unit inside until 09:10 on 2026-09-17."}, {"speaker": "Intake log", "text": "When crate C-418 was opened at 08:40 on 2026-09-17, its sole physical contents were immediately submitted for final placement as the Larkspur unit."}, {"speaker": "Decision clerk", "text": "Excluding session S, zero current direct visual confirmations of the submitted Larkspur unit’s step-6 back-panel orientation were available at the placement decision."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original verification policy and remain bound to the Larkspur unit’s 2026-09-17 final-placement decision and step-6 back-panel orientation. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the submitted unit’s immutable identifier from 731-QL to 884-NV; because session S inspected the uniquely registered 731-QL unit, the inspection can remain documented in the decision file without being a confirmation of the differently identified submitted unit. Neither context contains an answer, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Safety records reviewer\",\"text\":\"As of the placement decision on 2026-09-17, every required safety test for the submitted Larkspur unit had a passing result.\"},{\"speaker\":\"Checkpoint reviewer\",\"text\":\"Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence at that time.\"},{\"speaker\":\"Image-file reviewer\",\"text\":\"The step-6 back-panel checkpoint had zero required progress images available for the decision.\"},{\"speaker\":\"Identifier reconciliation record\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit.\"},{\"speaker\":\"Submission record\",\"text\":\"At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 731-QL.\"},{\"speaker\":\"Inspector\",\"text\":\"During session S, I directly observed the inspected unit’s back-panel grooved face oriented inward.\"},{\"speaker\":\"Verification-file custodian\",\"text\":\"That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"},{\"speaker\":\"Final evidence reviewer\",\"text\":\"Excluding session S, zero current direct visual confirmations of the submitted Larkspur unit’s step-6 back-panel orientation were available at the time of the decision.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["3", "text"], "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit."}, {"path": ["4", "text"], "text": "At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 731-QL."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit.", "negative_left": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit.", "negative_right": "At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 884-NV.", "right": "At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 731-QL."}, "verifier_independent_model": false}, "family": "scale-diverse-196-002", "id": "scale-diverse-196-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Safety records reviewer", "text": "As of the placement decision on 2026-09-17, every required safety test for the submitted Larkspur unit had a passing result."}, {"speaker": "Checkpoint reviewer", "text": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence at that time."}, {"speaker": "Image-file reviewer", "text": "The step-6 back-panel checkpoint had zero required progress images available for the decision."}, {"speaker": "Identifier reconciliation record", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit."}, {"speaker": "Submission record", "text": "At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 731-QL."}, {"speaker": "Inspector", "text": "During session S, I directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"speaker": "Verification-file custodian", "text": "That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Final evidence reviewer", "text": "Excluding session S, zero current direct visual confirmations of the submitted Larkspur unit’s step-6 back-panel orientation were available at the time of the decision."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original verification policy and remain bound to the Larkspur unit’s 2026-09-17 final-placement decision and step-6 back-panel orientation. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the submitted unit’s immutable identifier from 731-QL to 884-NV; because session S inspected the uniquely registered 731-QL unit, the inspection can remain documented in the decision file without being a confirmation of the differently identified submitted unit. Neither context contains an answer, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Safety records reviewer\",\"text\":\"As of the placement decision on 2026-09-17, every required safety test for the submitted Larkspur unit had a passing result.\"},{\"speaker\":\"Checkpoint reviewer\",\"text\":\"Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence at that time.\"},{\"speaker\":\"Image-file reviewer\",\"text\":\"The step-6 back-panel checkpoint had zero required progress images available for the decision.\"},{\"speaker\":\"Identifier reconciliation record\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit.\"},{\"speaker\":\"Submission record\",\"text\":\"At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 731-QL.\"},{\"speaker\":\"Inspector\",\"text\":\"During session S, I directly observed the inspected unit’s back-panel grooved face oriented inward.\"},{\"speaker\":\"Verification-file custodian\",\"text\":\"That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"},{\"speaker\":\"Final evidence reviewer\",\"text\":\"Excluding session S, zero current direct visual confirmations of the submitted Larkspur unit’s step-6 back-panel orientation were available at the time of the decision.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["3", "text"], "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit."}, {"path": ["4", "text"], "text": "At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 731-QL."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit.", "negative_left": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit.", "negative_right": "At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 884-NV.", "right": "At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 731-QL."}, "verifier_independent_model": false}, "family": "scale-diverse-196-002", "id": "scale-diverse-196-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Safety records reviewer", "text": "As of the placement decision on 2026-09-17, every required safety test for the submitted Larkspur unit had a passing result."}, {"speaker": "Checkpoint reviewer", "text": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence at that time."}, {"speaker": "Image-file reviewer", "text": "The step-6 back-panel checkpoint had zero required progress images available for the decision."}, {"speaker": "Identifier reconciliation record", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded 731-QL as the inspected physical unit's sole immutable manufacturer identifier, and the manufacturer's complete identifier register recorded 731-QL as assigned to exactly one physical unit."}, {"speaker": "Submission record", "text": "At submission for final placement on 2026-09-17, the Larkspur unit's sole immutable manufacturer identifier was 884-NV."}, {"speaker": "Inspector", "text": "During session S, I directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"speaker": "Verification-file custodian", "text": "That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Final evidence reviewer", "text": "Excluding session S, zero current direct visual confirmations of the submitted Larkspur unit’s step-6 back-panel orientation were available at the time of the decision."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts leave the original verification rubric and request scope unchanged and retain the Larkspur-unit, final-placement, and 2026-09-17 decision bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the submitted unit’s custody path from the crate-contained inspected unit to a different staging-cart unit; this does not conflict with the crate remaining sealed until 09:42, and the inspection file can still have been used even though session S concerned the other unit. Neither constructed context contains an answer code, explicit desired classification, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Custody coordinator\",\"text\":\"At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17.\"},{\"speaker\":\"Handoff staff\",\"text\":\"At 09:42 on 2026-09-17, handoff staff opened custody crate Q-47 and designated the sole physical unit removed from it as the Larkspur unit submitted for final placement on 2026-09-17.\"},{\"speaker\":\"Inspection file\",\"text\":\"During session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward. The observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"},{\"speaker\":\"Placement reviewer\",\"text\":\"As of the placement decision on 2026-09-17, every safety test required for the submitted Larkspur unit had a passing result. Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence. For step 6, zero required progress images were available; excluding session S, zero current direct visual confirmations of the submitted unit’s back-panel orientation were available.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["0", "text"], "text": "At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17."}, {"path": ["1", "text"], "text": "At 09:42 on 2026-09-17, handoff staff opened custody crate Q-47 and designated the sole physical unit removed from it as the Larkspur unit submitted for final placement on 2026-09-17."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17.", "negative_left": "At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17.", "negative_right": "At 09:42 on 2026-09-17, handoff staff designated a physical unit that had remained on staging cart C-84 and had never been inside custody crate Q-47 as the Larkspur unit submitted for final placement on 2026-09-17.", "right": "At 09:42 on 2026-09-17, handoff staff opened custody crate Q-47 and designated the sole physical unit removed from it as the Larkspur unit submitted for final placement on 2026-09-17."}, "verifier_independent_model": false}, "family": "scale-diverse-196-003", "id": "scale-diverse-196-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Custody coordinator", "text": "At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17."}, {"speaker": "Handoff staff", "text": "At 09:42 on 2026-09-17, handoff staff opened custody crate Q-47 and designated the sole physical unit removed from it as the Larkspur unit submitted for final placement on 2026-09-17."}, {"speaker": "Inspection file", "text": "During session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward. The observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Placement reviewer", "text": "As of the placement decision on 2026-09-17, every safety test required for the submitted Larkspur unit had a passing result. Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence. For step 6, zero required progress images were available; excluding session S, zero current direct visual confirmations of the submitted unit’s back-panel orientation were available."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts leave the original verification rubric and request scope unchanged and retain the Larkspur-unit, final-placement, and 2026-09-17 decision bindings. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the submitted unit’s custody path from the crate-contained inspected unit to a different staging-cart unit; this does not conflict with the crate remaining sealed until 09:42, and the inspection file can still have been used even though session S concerned the other unit. Neither constructed context contains an answer code, explicit desired classification, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Custody coordinator\",\"text\":\"At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17.\"},{\"speaker\":\"Handoff staff\",\"text\":\"At 09:42 on 2026-09-17, handoff staff opened custody crate Q-47 and designated the sole physical unit removed from it as the Larkspur unit submitted for final placement on 2026-09-17.\"},{\"speaker\":\"Inspection file\",\"text\":\"During session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward. The observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"},{\"speaker\":\"Placement reviewer\",\"text\":\"As of the placement decision on 2026-09-17, every safety test required for the submitted Larkspur unit had a passing result. Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence. For step 6, zero required progress images were available; excluding session S, zero current direct visual confirmations of the submitted unit’s back-panel orientation were available.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["0", "text"], "text": "At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17."}, {"path": ["1", "text"], "text": "At 09:42 on 2026-09-17, handoff staff opened custody crate Q-47 and designated the sole physical unit removed from it as the Larkspur unit submitted for final placement on 2026-09-17."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17.", "negative_left": "At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17.", "negative_right": "At 09:42 on 2026-09-17, handoff staff designated a physical unit that had remained on staging cart C-84 and had never been inside custody crate Q-47 as the Larkspur unit submitted for final placement on 2026-09-17.", "right": "At 09:42 on 2026-09-17, handoff staff opened custody crate Q-47 and designated the sole physical unit removed from it as the Larkspur unit submitted for final placement on 2026-09-17."}, "verifier_independent_model": false}, "family": "scale-diverse-196-003", "id": "scale-diverse-196-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Custody coordinator", "text": "At 17:18 on 2026-09-16, the physical unit inspected during pre-placement visual-inspection session S was sealed as the sole contents of custody crate Q-47, which remained continuously sealed and unopened until 09:42 on 2026-09-17."}, {"speaker": "Handoff staff", "text": "At 09:42 on 2026-09-17, handoff staff designated a physical unit that had remained on staging cart C-84 and had never been inside custody crate Q-47 as the Larkspur unit submitted for final placement on 2026-09-17."}, {"speaker": "Inspection file", "text": "During session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward. The observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Placement reviewer", "text": "As of the placement decision on 2026-09-17, every safety test required for the submitted Larkspur unit had a passing result. Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence. For step 6, zero required progress images were available; excluding session S, zero current direct visual confirmations of the submitted unit’s back-panel orientation were available."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric, and neither context removes or alters policy originating elsewhere. Both contexts retain the Larkspur unit, step-6 back-panel checkpoint, and 2026-09-17 decision bindings. The two focus spans are complete factual sentences. The counterfactual coherently changes the submitted Larkspur unit’s sole identity code to QL-7391 while session S remains tied to the distinct QL-4827 unit, without creating contradictory measurements or identity assertions. Neither context contains an answer, output instruction, code, rule table, proposition ID, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Custody records officer\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit.\"},{\"speaker\":\"Placement intake officer\",\"text\":\"At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-4827 as its sole permanent identity code.\"},{\"speaker\":\"Inspector\",\"text\":\"During session S, I directly observed the inspected unit’s back-panel grooved face oriented inward. That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"},{\"speaker\":\"Placement reviewer\",\"text\":\"As of the 2026-09-17 decision, every required safety test had a passing result, and every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence.\"},{\"speaker\":\"Evidence custodian\",\"text\":\"At decision time, zero required progress images were available for the submitted Larkspur unit’s step-6 back-panel checkpoint. Excluding session S, zero current direct visual confirmations of that unit’s step-6 back-panel orientation were available.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["0", "text"], "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit."}, {"path": ["1", "text"], "text": "At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-4827 as its sole permanent identity code."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit.", "negative_left": "During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit.", "negative_right": "At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-7391 as its sole permanent identity code.", "right": "At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-4827 as its sole permanent identity code."}, "verifier_independent_model": false}, "family": "scale-diverse-196-004", "id": "scale-diverse-196-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Custody records officer", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit."}, {"speaker": "Placement intake officer", "text": "At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-4827 as its sole permanent identity code."}, {"speaker": "Inspector", "text": "During session S, I directly observed the inspected unit’s back-panel grooved face oriented inward. That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Placement reviewer", "text": "As of the 2026-09-17 decision, every required safety test had a passing result, and every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence."}, {"speaker": "Evidence custodian", "text": "At decision time, zero required progress images were available for the submitted Larkspur unit’s step-6 back-panel checkpoint. Excluding session S, zero current direct visual confirmations of that unit’s step-6 back-panel orientation were available."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric, and neither context removes or alters policy originating elsewhere. Both contexts retain the Larkspur unit, step-6 back-panel checkpoint, and 2026-09-17 decision bindings. The two focus spans are complete factual sentences. The counterfactual coherently changes the submitted Larkspur unit’s sole identity code to QL-7391 while session S remains tied to the distinct QL-4827 unit, without creating contradictory measurements or identity assertions. Neither context contains an answer, output instruction, code, rule table, proposition ID, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Custody records officer\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit.\"},{\"speaker\":\"Placement intake officer\",\"text\":\"At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-4827 as its sole permanent identity code.\"},{\"speaker\":\"Inspector\",\"text\":\"During session S, I directly observed the inspected unit’s back-panel grooved face oriented inward. That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"},{\"speaker\":\"Placement reviewer\",\"text\":\"As of the 2026-09-17 decision, every required safety test had a passing result, and every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence.\"},{\"speaker\":\"Evidence custodian\",\"text\":\"At decision time, zero required progress images were available for the submitted Larkspur unit’s step-6 back-panel checkpoint. Excluding session S, zero current direct visual confirmations of that unit’s step-6 back-panel orientation were available.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["0", "text"], "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit."}, {"path": ["1", "text"], "text": "At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-4827 as its sole permanent identity code."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit.", "negative_left": "During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit.", "negative_right": "At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-7391 as its sole permanent identity code.", "right": "At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-4827 as its sole permanent identity code."}, "verifier_independent_model": false}, "family": "scale-diverse-196-004", "id": "scale-diverse-196-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Custody records officer", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspected physical unit bore identity code QL-4827 as its sole permanent identity code, and custody register CR-19 recorded QL-4827 as assigned to exactly one physical unit."}, {"speaker": "Placement intake officer", "text": "At final-placement submission on 2026-09-17, the physical Larkspur unit bore QL-7391 as its sole permanent identity code."}, {"speaker": "Inspector", "text": "During session S, I directly observed the inspected unit’s back-panel grooved face oriented inward. That observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Placement reviewer", "text": "As of the 2026-09-17 decision, every required safety test had a passing result, and every concealed-orientation checkpoint other than the step-6 back-panel checkpoint had qualifying evidence."}, {"speaker": "Evidence custodian", "text": "At decision time, zero required progress images were available for the submitted Larkspur unit’s step-6 back-panel checkpoint. Excluding session S, zero current direct visual confirmations of that unit’s step-6 back-panel orientation were available."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing 4 mm limit and scope exception, while the unchanged questions preserve the full scoring policy. The same Alderline unit, UT-01 test, measurement path, and date are used; the two evidence spans are complete factual sentences. Changing the maximum coordinate from 23.3 mm to 22.4 mm is coherent with an initial 18.6 mm coordinate, nondecreasing readings, and nonzero displacement, and neither context contains a score, output instruction, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Operational handoff: Alderline five-shelf unit serial AL-001 completed inspection. Inventory reconciliation found every required in-scope part present. Each installed in-scope part had its specified identifier, and every installed structural part had its specified orientation. The step record shows steps 1 through 12 were all performed; each performed step conformed to its corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error existed.\\n\\nDuring unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis. The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 23.3 mm. The UT-01 event log also marks top-corner displacement as nonzero.\\n\\nThe manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis."}, {"path": [], "text": "The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 23.3 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis.", "negative_left": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis.", "negative_right": "The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 22.4 mm.", "right": "The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 23.3 mm."}, "verifier_independent_model": false}, "family": "scale-diverse-197-003", "id": "scale-diverse-197-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Operational handoff: Alderline five-shelf unit serial AL-001 completed inspection. Inventory reconciliation found every required in-scope part present. Each installed in-scope part had its specified identifier, and every installed structural part had its specified orientation. The step record shows steps 1 through 12 were all performed; each performed step conformed to its corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error existed.\n\nDuring unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis. The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 23.3 mm. The UT-01 event log also marks top-corner displacement as nonzero.\n\nThe manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing 4 mm limit and scope exception, while the unchanged questions preserve the full scoring policy. The same Alderline unit, UT-01 test, measurement path, and date are used; the two evidence spans are complete factual sentences. Changing the maximum coordinate from 23.3 mm to 22.4 mm is coherent with an initial 18.6 mm coordinate, nondecreasing readings, and nonzero displacement, and neither context contains a score, output instruction, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Operational handoff: Alderline five-shelf unit serial AL-001 completed inspection. Inventory reconciliation found every required in-scope part present. Each installed in-scope part had its specified identifier, and every installed structural part had its specified orientation. The step record shows steps 1 through 12 were all performed; each performed step conformed to its corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error existed.\\n\\nDuring unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis. The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 23.3 mm. The UT-01 event log also marks top-corner displacement as nonzero.\\n\\nThe manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis."}, {"path": [], "text": "The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 23.3 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis.", "negative_left": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis.", "negative_right": "The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 22.4 mm.", "right": "The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 23.3 mm."}, "verifier_independent_model": false}, "family": "scale-diverse-197-003", "id": "scale-diverse-197-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Operational handoff: Alderline five-shelf unit serial AL-001 completed inspection. Inventory reconciliation found every required in-scope part present. Each installed in-scope part had its specified identifier, and every installed structural part had its specified orientation. The step record shows steps 1 through 12 were all performed; each performed step conformed to its corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error existed.\n\nDuring unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 on 14 September 2026, the calibrated position gauge recorded the tested top corner at 18.6 mm initially, recorded only nondecreasing coordinates thereafter, and detected no motion perpendicular to its measurement axis. The highest top-corner coordinate recorded by the calibrated position gauge during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 was 22.4 mm. The UT-01 event log also marks top-corner displacement as nonzero.\n\nThe manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring policy, while both contexts retain the original 4 mm limit and scope exception. The same Alderline unit, UT-01 measurement path, designated corner, timestamps, and coordinate frame are maintained. The two evidence spans are complete factual sentences. Changing only the later coordinate from (140, 246) to (137, 246) yields a coherent nonzero displacement of 4 mm rather than 5 mm, without conflicting with any unchanged assertion. Neither context contains a score, answer code, proposition identifier, rule table, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Operational handoff for Alderline five-shelf unit serial AL-001: inventory reconciliation found every required in-scope part present, and every installed in-scope part carried its specified identifier. All installed structural parts had the specified orientation. Steps 1–12 were each performed, and every performed step conformed to its corresponding instruction. Inspection found no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was found. UT-01 also recorded nonzero top-corner movement. The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242). The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (140, 246) in that fixed millimetre coordinate frame. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242)."}, {"path": [], "text": "The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (140, 246) in that fixed millimetre coordinate frame."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242).", "negative_left": "The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242).", "negative_right": "The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (137, 246) in that fixed millimetre coordinate frame.", "right": "The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (140, 246) in that fixed millimetre coordinate frame."}, "verifier_independent_model": false}, "family": "scale-diverse-197-007", "id": "scale-diverse-197-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Operational handoff for Alderline five-shelf unit serial AL-001: inventory reconciliation found every required in-scope part present, and every installed in-scope part carried its specified identifier. All installed structural parts had the specified orientation. Steps 1–12 were each performed, and every performed step conformed to its corresponding instruction. Inspection found no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was found. UT-01 also recorded nonzero top-corner movement. The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242). The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (140, 246) in that fixed millimetre coordinate frame. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring policy, while both contexts retain the original 4 mm limit and scope exception. The same Alderline unit, UT-01 measurement path, designated corner, timestamps, and coordinate frame are maintained. The two evidence spans are complete factual sentences. Changing only the later coordinate from (140, 246) to (137, 246) yields a coherent nonzero displacement of 4 mm rather than 5 mm, without conflicting with any unchanged assertion. Neither context contains a score, answer code, proposition identifier, rule table, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Operational handoff for Alderline five-shelf unit serial AL-001: inventory reconciliation found every required in-scope part present, and every installed in-scope part carried its specified identifier. All installed structural parts had the specified orientation. Steps 1–12 were each performed, and every performed step conformed to its corresponding instruction. Inspection found no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was found. UT-01 also recorded nonzero top-corner movement. The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242). The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (140, 246) in that fixed millimetre coordinate frame. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242)."}, {"path": [], "text": "The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (140, 246) in that fixed millimetre coordinate frame."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242).", "negative_left": "The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242).", "negative_right": "The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (137, 246) in that fixed millimetre coordinate frame.", "right": "The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (140, 246) in that fixed millimetre coordinate frame."}, "verifier_independent_model": false}, "family": "scale-diverse-197-007", "id": "scale-diverse-197-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Operational handoff for Alderline five-shelf unit serial AL-001: inventory reconciliation found every required in-scope part present, and every installed in-scope part carried its specified identifier. All installed structural parts had the specified orientation. Steps 1–12 were each performed, and every performed step conformed to its corresponding instruction. Inspection found no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was found. UT-01 also recorded nonzero top-corner movement. The UT-01 measurement record for Alderline five-shelf unit serial AL-001 specifies the measured top-corner displacement as the straight-line distance between the designated top corner's coordinates at 14:06:20 and 14:06:28 in the same fixed millimetre coordinate frame, with the 14:06:20 coordinates recorded as (137, 242). The designated top corner's coordinates at 14:06:28 during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 were recorded as (137, 246) in that fixed millimetre coordinate frame. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete scoring policy, while both contexts retain the applicable 4 mm stability limit, scope exclusions, and correction exception. Entity, test, measurement path, and date remain fixed; the two evidence spans are complete factual sentences. The counterfactual changes only the final gauge reading from 73.7 mm to 71.9 mm, yielding coherent absolute displacements of 5.3 mm and 3.5 mm respectively, with no duplicate contradictory measurement. Neither context includes a score, gold answer, answer code, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Inspection field note — Alderline five-shelf unit serial AL-001. Inventory reconciliation found every required in-scope part present, and every installed in-scope part bore its specified identifier. Every installed structural part had its specified orientation. Records confirm that steps 1 through 12 were all performed and that each performed step conformed to its corresponding instruction. Examination found no structural damage, no severe instability, and no assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was found. The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading. At the conclusion of the same test, the calibrated top-corner gauge recorded 73.7 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading."}, {"path": [], "text": "At the conclusion of the same test, the calibrated top-corner gauge recorded 73.7 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading.", "negative_left": "The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading.", "negative_right": "At the conclusion of the same test, the calibrated top-corner gauge recorded 71.9 mm.", "right": "At the conclusion of the same test, the calibrated top-corner gauge recorded 73.7 mm."}, "verifier_independent_model": false}, "family": "scale-diverse-197-008", "id": "scale-diverse-197-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Inspection field note — Alderline five-shelf unit serial AL-001. Inventory reconciliation found every required in-scope part present, and every installed in-scope part bore its specified identifier. Every installed structural part had its specified orientation. Records confirm that steps 1 through 12 were all performed and that each performed step conformed to its corresponding instruction. Examination found no structural damage, no severe instability, and no assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was found. The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading. At the conclusion of the same test, the calibrated top-corner gauge recorded 73.7 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete scoring policy, while both contexts retain the applicable 4 mm stability limit, scope exclusions, and correction exception. Entity, test, measurement path, and date remain fixed; the two evidence spans are complete factual sentences. The counterfactual changes only the final gauge reading from 73.7 mm to 71.9 mm, yielding coherent absolute displacements of 5.3 mm and 3.5 mm respectively, with no duplicate contradictory measurement. Neither context includes a score, gold answer, answer code, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Inspection field note — Alderline five-shelf unit serial AL-001. Inventory reconciliation found every required in-scope part present, and every installed in-scope part bore its specified identifier. Every installed structural part had its specified orientation. Records confirm that steps 1 through 12 were all performed and that each performed step conformed to its corresponding instruction. Examination found no structural damage, no severe instability, and no assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was found. The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading. At the conclusion of the same test, the calibrated top-corner gauge recorded 73.7 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading."}, {"path": [], "text": "At the conclusion of the same test, the calibrated top-corner gauge recorded 73.7 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading.", "negative_left": "The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading.", "negative_right": "At the conclusion of the same test, the calibrated top-corner gauge recorded 71.9 mm.", "right": "At the conclusion of the same test, the calibrated top-corner gauge recorded 73.7 mm."}, "verifier_independent_model": false}, "family": "scale-diverse-197-008", "id": "scale-diverse-197-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Inspection field note — Alderline five-shelf unit serial AL-001. Inventory reconciliation found every required in-scope part present, and every installed in-scope part bore its specified identifier. Every installed structural part had its specified orientation. Records confirm that steps 1 through 12 were all performed and that each performed step conformed to its corresponding instruction. Examination found no structural damage, no severe instability, and no assembly error creating an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was found. The signed record for unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, conducted on 2026-08-14, gives an initial top-corner gauge reading of 68.4 mm and calculates the measured top-corner displacement solely as the absolute difference between that reading and the final reading. At the conclusion of the same test, the calibrated top-corner gauge recorded 71.9 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 3 mm tolerance and remain compatible with the original routing, completion, and rating policy without adding exceptions or defaults. They preserve the decision scope and Priya’s 45 cm cushion-cover binding; the permitted observation update changes only the marked-line measurement from 423 mm to 426 mm at the same timestamp and measurement path. The two focus-evidence spans are complete factual sentences. No unchanged assertion conflicts with the counterfactual measurement, and neither context contains an answer choice, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"Evidence reconciliation for Priya’s 45 cm cushion cover: The project sheet allows no more than 3 mm seam deviation. At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline. At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 423 mm from the left edge of the same cut panel along its horizontal centerline. The consolidated inspection logged at that same timestamp verified every required dimension and confirmed that every required cut piece and every required construction seam was present. Every construction requirement other than required-seam presence and the right-edge seam-deviation limit passed. Trimming, pressing, and final inspection each passed.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline."}, {"path": [], "text": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 423 mm from the left edge of the same cut panel along its horizontal centerline."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline.", "negative_left": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline.", "negative_right": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 426 mm from the left edge of the same cut panel along its horizontal centerline.", "right": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 423 mm from the left edge of the same cut panel along its horizontal centerline."}, "verifier_independent_model": false}, "family": "scale-diverse-199-002", "id": "scale-diverse-199-002-base", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "Evidence reconciliation for Priya’s 45 cm cushion cover: The project sheet allows no more than 3 mm seam deviation. At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline. At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 423 mm from the left edge of the same cut panel along its horizontal centerline. The consolidated inspection logged at that same timestamp verified every required dimension and confirmed that every required cut piece and every required construction seam was present. Every construction requirement other than required-seam presence and the right-edge seam-deviation limit passed. Trimming, pressing, and final inspection each passed."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 3 mm tolerance and remain compatible with the original routing, completion, and rating policy without adding exceptions or defaults. They preserve the decision scope and Priya’s 45 cm cushion-cover binding; the permitted observation update changes only the marked-line measurement from 423 mm to 426 mm at the same timestamp and measurement path. The two focus-evidence spans are complete factual sentences. No unchanged assertion conflicts with the counterfactual measurement, and neither context contains an answer choice, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"Evidence reconciliation for Priya’s 45 cm cushion cover: The project sheet allows no more than 3 mm seam deviation. At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline. At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 423 mm from the left edge of the same cut panel along its horizontal centerline. The consolidated inspection logged at that same timestamp verified every required dimension and confirmed that every required cut piece and every required construction seam was present. Every construction requirement other than required-seam presence and the right-edge seam-deviation limit passed. Trimming, pressing, and final inspection each passed.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline."}, {"path": [], "text": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 423 mm from the left edge of the same cut panel along its horizontal centerline."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline.", "negative_left": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline.", "negative_right": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 426 mm from the left edge of the same cut panel along its horizontal centerline.", "right": "At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 423 mm from the left edge of the same cut panel along its horizontal centerline."}, "verifier_independent_model": false}, "family": "scale-diverse-199-002", "id": "scale-diverse-199-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "Evidence reconciliation for Priya’s 45 cm cushion cover: The project sheet allows no more than 3 mm seam deviation. At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 428 mm from the left edge of the cut panel along its horizontal centerline. At 2026-09-16 14:32 UTC, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the corresponding marked line at 426 mm from the left edge of the same cut panel along its horizontal centerline. The consolidated inspection logged at that same timestamp verified every required dimension and confirmed that every required cut piece and every required construction seam was present. Every construction requirement other than required-seam presence and the right-edge seam-deviation limit passed. Trimming, pressing, and final inspection each passed."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_clean_4"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 3 mm tolerance and latest-evidence rule without adding policy exceptions or defaults, while preserving the decision scope and Priya’s 45 cm cushion-cover binding. The two evidence spans are complete factual sentences. The counterfactual changes only the marked-line measurement from 42 mm to 45 mm; with the seam remaining at 47 mm, this creates a coherent 2 mm deviation and does not conflict with any unchanged measurement or assertion. Neither context includes an answer choice, rating, route instruction, code, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"Operational handoff for Priya’s 45 cm cushion cover: the 16:20 UTC inspection record dated 14 September 2026 is the latest timestamped evidence. The project sheet allows no more than 3 mm seam deviation. The checklist confirms that every required dimension has been verified and every required cut piece is present. It also confirms that every required construction seam is present. All construction requirements apart from required-seam presence and the right-edge seam-deviation check have passed. Trimming has passed, pressing has passed, and final inspection has passed. At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge. At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 42 mm inward from the cover’s right edge.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge."}, {"path": [], "text": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 42 mm inward from the cover’s right edge."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge.", "negative_left": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge.", "negative_right": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 45 mm inward from the cover’s right edge.", "right": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 42 mm inward from the cover’s right edge."}, "verifier_independent_model": false}, "family": "scale-diverse-199-003", "id": "scale-diverse-199-003-base", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "Operational handoff for Priya’s 45 cm cushion cover: the 16:20 UTC inspection record dated 14 September 2026 is the latest timestamped evidence. The project sheet allows no more than 3 mm seam deviation. The checklist confirms that every required dimension has been verified and every required cut piece is present. It also confirms that every required construction seam is present. All construction requirements apart from required-seam presence and the right-edge seam-deviation check have passed. Trimming has passed, pressing has passed, and final inspection has passed. At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge. At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 42 mm inward from the cover’s right edge."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 3 mm tolerance and latest-evidence rule without adding policy exceptions or defaults, while preserving the decision scope and Priya’s 45 cm cushion-cover binding. The two evidence spans are complete factual sentences. The counterfactual changes only the marked-line measurement from 42 mm to 45 mm; with the seam remaining at 47 mm, this creates a coherent 2 mm deviation and does not conflict with any unchanged measurement or assertion. Neither context includes an answer choice, rating, route instruction, code, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"Operational handoff for Priya’s 45 cm cushion cover: the 16:20 UTC inspection record dated 14 September 2026 is the latest timestamped evidence. The project sheet allows no more than 3 mm seam deviation. The checklist confirms that every required dimension has been verified and every required cut piece is present. It also confirms that every required construction seam is present. All construction requirements apart from required-seam presence and the right-edge seam-deviation check have passed. Trimming has passed, pressing has passed, and final inspection has passed. At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge. At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 42 mm inward from the cover’s right edge.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge."}, {"path": [], "text": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 42 mm inward from the cover’s right edge."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge.", "negative_left": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge.", "negative_right": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 45 mm inward from the cover’s right edge.", "right": "At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 42 mm inward from the cover’s right edge."}, "verifier_independent_model": false}, "family": "scale-diverse-199-003", "id": "scale-diverse-199-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "Operational handoff for Priya’s 45 cm cushion cover: the 16:20 UTC inspection record dated 14 September 2026 is the latest timestamped evidence. The project sheet allows no more than 3 mm seam deviation. The checklist confirms that every required dimension has been verified and every required cut piece is present. It also confirms that every required construction seam is present. All construction requirements apart from required-seam presence and the right-edge seam-deviation check have passed. Trimming has passed, pressing has passed, and final inspection has passed. At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover measured the existing right-edge seam at 47 mm inward from the cover’s right edge. At 16:20 UTC on 14 September 2026, the latest timestamped evidence for Priya’s 45 cm cushion cover recorded the right-edge seam’s marked line at 45 mm inward from the cover’s right edge."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_clean_4"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same project specifications, request, decision scope, and Mara/cover binding. The two focus spans are complete factual sentences. The counterfactual changes only the registry sentence for IMG-2, assigning its object to QA-908 and explicitly separating it from MC-714; this is coherent with the unchanged evidence because the pictured puckering then concerns another object, while the remaining inspection statements can consistently describe Mara’s cover. Neither context includes an outcome label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover.\",\"At 09:20, direct inspection measured the cover's finished width as 45.3 cm and its finished height as 44.6 cm.\",\"At 09:45, direct inspection found every corner square and every thread trimmed; the inspection record also confirmed that the cover had been pressed and counted zero cosmetic defects.\",\"At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier MC-714.\",\"IMG-2 shows a puckered zipper seam on the object depicted in that image.\",\"By 11:00, every inspected portion of Mara's submitted cover's zipper seam that was not part of the object shown in IMG-2 was even. Direct inspection counted zero construction defects on Mara's cover outside any zipper-seam portion shown in IMG-2.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "0"], "text": "From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover."}, {"path": ["evidence", "3"], "text": "At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier MC-714."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover.", "negative_left": "From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover.", "negative_right": "At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier QA-908, and the registry designated QA-908 and MC-714 to physically separate objects throughout that day.", "right": "At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier MC-714."}, "verifier_independent_model": false}, "family": "scale-diverse-200-001", "id": "scale-diverse-200-001-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover.", "At 09:20, direct inspection measured the cover's finished width as 45.3 cm and its finished height as 44.6 cm.", "At 09:45, direct inspection found every corner square and every thread trimmed; the inspection record also confirmed that the cover had been pressed and counted zero cosmetic defects.", "At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier MC-714.", "IMG-2 shows a puckered zipper seam on the object depicted in that image.", "By 11:00, every inspected portion of Mara's submitted cover's zipper seam that was not part of the object shown in IMG-2 was even. Direct inspection counted zero construction defects on Mara's cover outside any zipper-seam portion shown in IMG-2."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same project specifications, request, decision scope, and Mara/cover binding. The two focus spans are complete factual sentences. The counterfactual changes only the registry sentence for IMG-2, assigning its object to QA-908 and explicitly separating it from MC-714; this is coherent with the unchanged evidence because the pictured puckering then concerns another object, while the remaining inspection statements can consistently describe Mara’s cover. Neither context includes an outcome label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover.\",\"At 09:20, direct inspection measured the cover's finished width as 45.3 cm and its finished height as 44.6 cm.\",\"At 09:45, direct inspection found every corner square and every thread trimmed; the inspection record also confirmed that the cover had been pressed and counted zero cosmetic defects.\",\"At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier MC-714.\",\"IMG-2 shows a puckered zipper seam on the object depicted in that image.\",\"By 11:00, every inspected portion of Mara's submitted cover's zipper seam that was not part of the object shown in IMG-2 was even. Direct inspection counted zero construction defects on Mara's cover outside any zipper-seam portion shown in IMG-2.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "0"], "text": "From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover."}, {"path": ["evidence", "3"], "text": "At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier MC-714."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover.", "negative_left": "From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover.", "negative_right": "At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier QA-908, and the registry designated QA-908 and MC-714 to physically separate objects throughout that day.", "right": "At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier MC-714."}, "verifier_independent_model": false}, "family": "scale-diverse-200-001", "id": "scale-diverse-200-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["From 09:00 through 12:00 on 14 May 2026, Workshop Quality Registry identifier MC-714 designated exclusively one physical object, Mara's submitted zippered cushion cover.", "At 09:20, direct inspection measured the cover's finished width as 45.3 cm and its finished height as 44.6 cm.", "At 09:45, direct inspection found every corner square and every thread trimmed; the inspection record also confirmed that the cover had been pressed and counted zero cosmetic defects.", "At 10:25 on 14 May 2026, the Workshop Quality Registry entry for the object shown in inspection image IMG-2 recorded identifier QA-908, and the registry designated QA-908 and MC-714 to physically separate objects throughout that day.", "IMG-2 shows a puckered zipper seam on the object depicted in that image.", "By 11:00, every inspected portion of Mara's submitted cover's zipper seam that was not part of the object shown in IMG-2 was even. Direct inspection counted zero construction defects on Mara's cover outside any zipper-seam portion shown in IMG-2."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing criteria and instructions without adding exceptions, priorities, or missing-evidence defaults, and they retain Mara’s submitted cushion cover, the same inspection request, and the same specification scope. The two focus spans are complete factual registry sentences. Changing Mara’s cover identifier from OI-731 to OI-946 coherently makes it distinct from the IMG-2 object under the stated one-to-one registry constraints; consequently, the puckered seam shown in IMG-2 no longer conflicts with the assertions that Mara’s own cover has an even seam and no defects. Neither context embeds a gold outcome, answer code, proposition identifier, rule table, label rationale, or classifier-output instruction beyond the preserved task instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"complete_acceptable\":\"Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.\",\"complete_excellent\":\"Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.\",\"complete_good\":\"Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.\",\"route_finishing\":\"Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.\",\"route_seam_correction\":\"Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect.\"},\"instructions\":\"Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.\",\"type\":\"choice\"}},\"state\":{\"context\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Direct inspection measured Mara's submitted zippered cushion cover at 44.8 cm wide and 45.2 cm high.\",\"The inspector verified that every corner was square, every thread was trimmed, and the submitted cover had been pressed.\",\"Direct inspection found zero cosmetic defects on the submitted cover.\",\"Every inspected portion of Mara's submitted cover's zipper seam that was not part of the object shown in IMG-2 was even; no construction defects were found outside any zipper-seam portion shown in IMG-2.\",\"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\",\"In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object.\",\"In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-731.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["state", "evidence", "5"], "text": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object."}, {"path": ["state", "evidence", "6"], "text": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-731."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object.", "negative_left": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object.", "negative_right": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-946.", "right": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-731."}, "verifier_independent_model": false}, "family": "scale-diverse-200-002", "id": "scale-diverse-200-002-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Direct inspection measured Mara's submitted zippered cushion cover at 44.8 cm wide and 45.2 cm high.", "The inspector verified that every corner was square, every thread was trimmed, and the submitted cover had been pressed.", "Direct inspection found zero cosmetic defects on the submitted cover.", "Every inspected portion of Mara's submitted cover's zipper seam that was not part of the object shown in IMG-2 was even; no construction defects were found outside any zipper-seam portion shown in IMG-2.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object.", "In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-731."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing criteria and instructions without adding exceptions, priorities, or missing-evidence defaults, and they retain Mara’s submitted cushion cover, the same inspection request, and the same specification scope. The two focus spans are complete factual registry sentences. Changing Mara’s cover identifier from OI-731 to OI-946 coherently makes it distinct from the IMG-2 object under the stated one-to-one registry constraints; consequently, the puckered seam shown in IMG-2 no longer conflicts with the assertions that Mara’s own cover has an even seam and no defects. Neither context embeds a gold outcome, answer code, proposition identifier, rule table, label rationale, or classifier-output instruction beyond the preserved task instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"complete_acceptable\":\"Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.\",\"complete_excellent\":\"Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.\",\"complete_good\":\"Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.\",\"route_finishing\":\"Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.\",\"route_seam_correction\":\"Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect.\"},\"instructions\":\"Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.\",\"type\":\"choice\"}},\"state\":{\"context\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Direct inspection measured Mara's submitted zippered cushion cover at 44.8 cm wide and 45.2 cm high.\",\"The inspector verified that every corner was square, every thread was trimmed, and the submitted cover had been pressed.\",\"Direct inspection found zero cosmetic defects on the submitted cover.\",\"Every inspected portion of Mara's submitted cover's zipper seam that was not part of the object shown in IMG-2 was even; no construction defects were found outside any zipper-seam portion shown in IMG-2.\",\"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\",\"In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object.\",\"In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-731.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["state", "evidence", "5"], "text": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object."}, {"path": ["state", "evidence", "6"], "text": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-731."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object.", "negative_left": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object.", "negative_right": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-946.", "right": "In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-731."}, "verifier_independent_model": false}, "family": "scale-diverse-200-002", "id": "scale-diverse-200-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Direct inspection measured Mara's submitted zippered cushion cover at 44.8 cm wide and 45.2 cm high.", "The inspector verified that every corner was square, every thread was trimmed, and the submitted cover had been pressed.", "Direct inspection found zero cosmetic defects on the submitted cover.", "Every inspected portion of Mara's submitted cover's zipper seam that was not part of the object shown in IMG-2 was even; no construction defects were found outside any zipper-seam portion shown in IMG-2.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "In the sealed inspection registry at 14:20 UTC on 12 May 2026, the object shown in inspection image IMG-2 had recorded object identifier OI-731, and at that time every registered physical object had exactly one object identifier and every recorded object identifier designated exactly one registered physical object.", "In the sealed inspection registry at 14:20 UTC on 12 May 2026, Mara's submitted zippered cushion cover had recorded object identifier OI-946."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full decision policy, while both contexts retain the original project specifications and request without adding exceptions or defaults. The two focus spans are complete factual sentences; the pouch change coherently makes IMG-2 depict a different cover from Mara’s submission, without creating duplicate contradictory measurements or assertions, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"The handoff inspection recorded Mara’s submitted cover’s directly measured finished width as 44.8 cm and finished height as 45.3 cm. All four corners were square, every thread was trimmed, the cover had been pressed, and direct inspection found zero cosmetic defects.\",\"At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff.\",\"At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-62 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded.\",\"Inspection image IMG-2 shows a puckered zipper seam on the depicted object.\",\"Every inspected portion of Mara’s submitted cover’s zipper seam that was not part of the object shown in IMG-2 was even. Inspectors found zero construction defects on Mara’s cover outside any zipper-seam portion shown in IMG-2.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff."}, {"path": ["evidence", "2"], "text": "At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-62 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff.", "negative_left": "At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff.", "negative_right": "At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-89 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded.", "right": "At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-62 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded."}, "verifier_independent_model": false}, "family": "scale-diverse-200-003", "id": "scale-diverse-200-003-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["The handoff inspection recorded Mara’s submitted cover’s directly measured finished width as 44.8 cm and finished height as 45.3 cm. All four corners were square, every thread was trimmed, the cover had been pressed, and direct inspection found zero cosmetic defects.", "At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff.", "At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-62 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded.", "Inspection image IMG-2 shows a puckered zipper seam on the depicted object.", "Every inspected portion of Mara’s submitted cover’s zipper seam that was not part of the object shown in IMG-2 was even. Inspectors found zero construction defects on Mara’s cover outside any zipper-seam portion shown in IMG-2."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full decision policy, while both contexts retain the original project specifications and request without adding exceptions or defaults. The two focus spans are complete factual sentences; the pouch change coherently makes IMG-2 depict a different cover from Mara’s submission, without creating duplicate contradictory measurements or assertions, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"The handoff inspection recorded Mara’s submitted cover’s directly measured finished width as 44.8 cm and finished height as 45.3 cm. All four corners were square, every thread was trimmed, the cover had been pressed, and direct inspection found zero cosmetic defects.\",\"At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff.\",\"At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-62 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded.\",\"Inspection image IMG-2 shows a puckered zipper seam on the depicted object.\",\"Every inspected portion of Mara’s submitted cover’s zipper seam that was not part of the object shown in IMG-2 was even. Inspectors found zero construction defects on Mara’s cover outside any zipper-seam portion shown in IMG-2.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff."}, {"path": ["evidence", "2"], "text": "At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-62 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff.", "negative_left": "At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff.", "negative_right": "At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-89 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded.", "right": "At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-62 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded."}, "verifier_independent_model": false}, "family": "scale-diverse-200-003", "id": "scale-diverse-200-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["The handoff inspection recorded Mara’s submitted cover’s directly measured finished width as 44.8 cm and finished height as 45.3 cm. All four corners were square, every thread was trimmed, the cover had been pressed, and direct inspection found zero cosmetic defects.", "At 14:08 on 12 May 2026, the object shown in inspection image IMG-2 was recorded as identifier QX-731 and sealed alone in evidence pouch P-62, while a different physical cushion cover was sealed alone in evidence pouch P-89; both pouches remained sealed until the 14:30 handoff.", "At the 14:30 handoff on 12 May 2026, the sole object removed from still-sealed evidence pouch P-89 was Mara's submitted zippered cushion cover, for which identifier MC-204 was recorded.", "Inspection image IMG-2 shows a puckered zipper seam on the depicted object.", "Every inspected portion of Mara’s submitted cover’s zipper seam that was not part of the object shown in IMG-2 was even. Inspectors found zero construction defects on Mara’s cover outside any zipper-seam portion shown in IMG-2."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all stated quality policies and preserve the decision’s blue-cushion-cover scope and requirements. The two focus spans are complete factual sentences. The counterfactual changes only Q-17’s identity at the bound exposure time, coherently making Photograph I depict a distinct beige table runner rather than the cushion cover without creating a duplicate conflicting measurement or assertion within that context. Neither context contains a gold answer, answer code, proposition ID, rule table, label rationale, or classifier-output instruction; the references to “Good” and verification are preserved natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual or evidentiary relationship; the universally quantified atoms remain single relationships over explicit evidence sets. The focus atom concerns whether a photograph depicts the blue cushion cover, so it is factual rather than policy. The base and counter assignments are jointly realizable with only that photograph-identity fact changing: Photograph I can show enclosed edges in either scenario while depicting the blue cover only in the base. The policy evidence properly preserves the governing requirements and evidence rules originating in the original state; instructions already retained in the questions object need not be duplicated. Both rules are sufficient for their targets, and uncovered assignments may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions provide recorded 45 × 45 cm dimensions, photographic evidence of a centered zipper and clipped exterior threads, and an inside-seam photograph identified as depicting the blue cover and showing all its raw edges enclosed. The no-visible-cosmetic-flaw condition also excludes a competing lower finish assessment based on submitted evidence. This is sufficient for completion as Good under the stated evidence policy.", "rule_index": 0, "sound": true}, {"reason": "Because Photograph I is the only submitted inside-seam photograph and is explicitly refuted as depicting the blue cover, its enclosed-edge evidence cannot verify that cover. Thus the blue cover lacks required inside-seam evidence, which blocks completion as Good.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The recorded final width of the blue cushion cover is 45 cm."}, {"id": "a2", "statement": "The recorded final height of the blue cushion cover is 45 cm."}, {"id": "a3", "statement": "The submitted exterior photograph of the blue cushion cover shows that its zipper is centered."}, {"id": "a4", "statement": "The submitted exterior photograph of the blue cushion cover shows that every exterior thread end is clipped."}, {"id": "a5", "statement": "Photograph I is the only submitted inside-seam photograph."}, {"id": "a6", "statement": "Photograph I shows that every raw edge in the photographed cover is enclosed."}, {"id": "a7", "statement": "Photograph I depicts the blue cushion cover."}, {"id": "a8", "statement": "No cosmetic flaw is visible in the submitted evidence for the blue cushion cover."}], "base_state_json": "[{\"speaker\":\"Project sheet\",\"text\":\"The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges.\"},{\"speaker\":\"Quality policy\",\"text\":\"Our scale is Needs correction, Acceptable, then Good.\"},{\"speaker\":\"Quality policy\",\"text\":\"Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct.\"},{\"speaker\":\"Quality policy\",\"text\":\"Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified.\"},{\"speaker\":\"Quality policy\",\"text\":\"No assumptions are allowed.\"},{\"speaker\":\"Quality policy\",\"text\":\"Without final measurements and inside-seam evidence, it must remain in measurement and seam verification.\"},{\"speaker\":\"Measurement record\",\"text\":\"The recorded final dimensions of the blue cushion cover are 45 cm wide and 45 cm high.\"},{\"speaker\":\"Exterior review\",\"text\":\"The submitted exterior photograph of the blue cushion cover shows a centered zipper and every exterior thread end clipped.\"},{\"speaker\":\"Submission log\",\"text\":\"Photograph I is the only submitted inside-seam photograph.\"},{\"speaker\":\"Seam review\",\"text\":\"Photograph I shows that every raw edge in the photographed cover is enclosed.\"},{\"speaker\":\"Exposure record\",\"text\":\"During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present.\"},{\"speaker\":\"Inventory record\",\"text\":\"At 14:08:17 on 6 May 2025, inventory item Q-17 was the blue cushion cover.\"},{\"speaker\":\"Visual audit\",\"text\":\"No cosmetic flaw is visible in the submitted evidence for the blue cushion cover.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": ["10", "text"], "text": "During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present."}, {"path": ["11", "text"], "text": "At 14:08:17 on 6 May 2025, inventory item Q-17 was the blue cushion cover."}], "policy_evidence": [{"path": ["0", "text"], "text": "The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges."}, {"path": ["1", "text"], "text": "Our scale is Needs correction, Acceptable, then Good."}, {"path": ["1", "text"], "text": "Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct."}, {"path": ["1", "text"], "text": "Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified."}, {"path": ["3", "text"], "text": "No assumptions are allowed."}, {"path": ["3", "text"], "text": "Without final measurements and inside-seam evidence, it must remain in measurement and seam verification."}], "rules": [{"justification": "The recorded dimensions meet 45 × 45 cm, the exterior evidence verifies the centered zipper and clipped threads, and the sole submitted inside-seam photograph both depicts the blue cushion cover and verifies its enclosed raw edges. With no visible cosmetic flaw, every sheet requirement has evidence and no competing Acceptable outcome is indicated.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Photograph I is the only submitted inside-seam photograph, but it does not depict the blue cushion cover. Therefore the blue cushion cover lacks inside-seam evidence verifying its enclosed raw edges. Missing evidence for any sheet requirement blocks completion as Good, and assumptions are prohibited.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present.", "negative_left": "During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present.", "negative_right": "At 14:08:17 on 6 May 2025, inventory item Q-17 was a beige table runner distinct from the blue cushion cover.", "right": "At 14:08:17 on 6 May 2025, inventory item Q-17 was the blue cushion cover."}, "verifier_independent_model": false}, "family": "scale-diverse-201-001", "id": "scale-diverse-201-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it complete as Good; keep it in measurement and seam verification.", "true": "Yes — mark the cushion cover complete and rate its finish quality Good."}, "instructions": "Decide whether the cushion cover should be marked complete with a finish-quality rating of Good. Answer yes or no using the stated evidence policy.", "type": "noul"}}, "state": [{"speaker": "Project sheet", "text": "The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges."}, {"speaker": "Quality policy", "text": "Our scale is Needs correction, Acceptable, then Good."}, {"speaker": "Quality policy", "text": "Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct."}, {"speaker": "Quality policy", "text": "Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified."}, {"speaker": "Quality policy", "text": "No assumptions are allowed."}, {"speaker": "Quality policy", "text": "Without final measurements and inside-seam evidence, it must remain in measurement and seam verification."}, {"speaker": "Measurement record", "text": "The recorded final dimensions of the blue cushion cover are 45 cm wide and 45 cm high."}, {"speaker": "Exterior review", "text": "The submitted exterior photograph of the blue cushion cover shows a centered zipper and every exterior thread end clipped."}, {"speaker": "Submission log", "text": "Photograph I is the only submitted inside-seam photograph."}, {"speaker": "Seam review", "text": "Photograph I shows that every raw edge in the photographed cover is enclosed."}, {"speaker": "Exposure record", "text": "During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present."}, {"speaker": "Inventory record", "text": "At 14:08:17 on 6 May 2025, inventory item Q-17 was the blue cushion cover."}, {"speaker": "Visual audit", "text": "No cosmetic flaw is visible in the submitted evidence for the blue cushion cover."}]}, "method": "c2d", "provenance": {"source_id": "diverse-201", "source_is_synthetic": true, "source_sha256": "948210211c01bfb1c47fea1b9af1b2983c05fcb405b70f31e920771698ff4af9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all stated quality policies and preserve the decision’s blue-cushion-cover scope and requirements. The two focus spans are complete factual sentences. The counterfactual changes only Q-17’s identity at the bound exposure time, coherently making Photograph I depict a distinct beige table runner rather than the cushion cover without creating a duplicate conflicting measurement or assertion within that context. Neither context contains a gold answer, answer code, proposition ID, rule table, label rationale, or classifier-output instruction; the references to “Good” and verification are preserved natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual or evidentiary relationship; the universally quantified atoms remain single relationships over explicit evidence sets. The focus atom concerns whether a photograph depicts the blue cushion cover, so it is factual rather than policy. The base and counter assignments are jointly realizable with only that photograph-identity fact changing: Photograph I can show enclosed edges in either scenario while depicting the blue cover only in the base. The policy evidence properly preserves the governing requirements and evidence rules originating in the original state; instructions already retained in the questions object need not be duplicated. Both rules are sufficient for their targets, and uncovered assignments may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions provide recorded 45 × 45 cm dimensions, photographic evidence of a centered zipper and clipped exterior threads, and an inside-seam photograph identified as depicting the blue cover and showing all its raw edges enclosed. The no-visible-cosmetic-flaw condition also excludes a competing lower finish assessment based on submitted evidence. This is sufficient for completion as Good under the stated evidence policy.", "rule_index": 0, "sound": true}, {"reason": "Because Photograph I is the only submitted inside-seam photograph and is explicitly refuted as depicting the blue cover, its enclosed-edge evidence cannot verify that cover. Thus the blue cover lacks required inside-seam evidence, which blocks completion as Good.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The recorded final width of the blue cushion cover is 45 cm."}, {"id": "a2", "statement": "The recorded final height of the blue cushion cover is 45 cm."}, {"id": "a3", "statement": "The submitted exterior photograph of the blue cushion cover shows that its zipper is centered."}, {"id": "a4", "statement": "The submitted exterior photograph of the blue cushion cover shows that every exterior thread end is clipped."}, {"id": "a5", "statement": "Photograph I is the only submitted inside-seam photograph."}, {"id": "a6", "statement": "Photograph I shows that every raw edge in the photographed cover is enclosed."}, {"id": "a7", "statement": "Photograph I depicts the blue cushion cover."}, {"id": "a8", "statement": "No cosmetic flaw is visible in the submitted evidence for the blue cushion cover."}], "base_state_json": "[{\"speaker\":\"Project sheet\",\"text\":\"The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges.\"},{\"speaker\":\"Quality policy\",\"text\":\"Our scale is Needs correction, Acceptable, then Good.\"},{\"speaker\":\"Quality policy\",\"text\":\"Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct.\"},{\"speaker\":\"Quality policy\",\"text\":\"Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified.\"},{\"speaker\":\"Quality policy\",\"text\":\"No assumptions are allowed.\"},{\"speaker\":\"Quality policy\",\"text\":\"Without final measurements and inside-seam evidence, it must remain in measurement and seam verification.\"},{\"speaker\":\"Measurement record\",\"text\":\"The recorded final dimensions of the blue cushion cover are 45 cm wide and 45 cm high.\"},{\"speaker\":\"Exterior review\",\"text\":\"The submitted exterior photograph of the blue cushion cover shows a centered zipper and every exterior thread end clipped.\"},{\"speaker\":\"Submission log\",\"text\":\"Photograph I is the only submitted inside-seam photograph.\"},{\"speaker\":\"Seam review\",\"text\":\"Photograph I shows that every raw edge in the photographed cover is enclosed.\"},{\"speaker\":\"Exposure record\",\"text\":\"During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present.\"},{\"speaker\":\"Inventory record\",\"text\":\"At 14:08:17 on 6 May 2025, inventory item Q-17 was the blue cushion cover.\"},{\"speaker\":\"Visual audit\",\"text\":\"No cosmetic flaw is visible in the submitted evidence for the blue cushion cover.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": ["10", "text"], "text": "During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present."}, {"path": ["11", "text"], "text": "At 14:08:17 on 6 May 2025, inventory item Q-17 was the blue cushion cover."}], "policy_evidence": [{"path": ["0", "text"], "text": "The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges."}, {"path": ["1", "text"], "text": "Our scale is Needs correction, Acceptable, then Good."}, {"path": ["1", "text"], "text": "Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct."}, {"path": ["1", "text"], "text": "Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified."}, {"path": ["3", "text"], "text": "No assumptions are allowed."}, {"path": ["3", "text"], "text": "Without final measurements and inside-seam evidence, it must remain in measurement and seam verification."}], "rules": [{"justification": "The recorded dimensions meet 45 × 45 cm, the exterior evidence verifies the centered zipper and clipped threads, and the sole submitted inside-seam photograph both depicts the blue cushion cover and verifies its enclosed raw edges. With no visible cosmetic flaw, every sheet requirement has evidence and no competing Acceptable outcome is indicated.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Photograph I is the only submitted inside-seam photograph, but it does not depict the blue cushion cover. Therefore the blue cushion cover lacks inside-seam evidence verifying its enclosed raw edges. Missing evidence for any sheet requirement blocks completion as Good, and assumptions are prohibited.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present.", "negative_left": "During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present.", "negative_right": "At 14:08:17 on 6 May 2025, inventory item Q-17 was a beige table runner distinct from the blue cushion cover.", "right": "At 14:08:17 on 6 May 2025, inventory item Q-17 was the blue cushion cover."}, "verifier_independent_model": false}, "family": "scale-diverse-201-001", "id": "scale-diverse-201-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it complete as Good; keep it in measurement and seam verification.", "true": "Yes — mark the cushion cover complete and rate its finish quality Good."}, "instructions": "Decide whether the cushion cover should be marked complete with a finish-quality rating of Good. Answer yes or no using the stated evidence policy.", "type": "noul"}}, "state": [{"speaker": "Project sheet", "text": "The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges."}, {"speaker": "Quality policy", "text": "Our scale is Needs correction, Acceptable, then Good."}, {"speaker": "Quality policy", "text": "Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct."}, {"speaker": "Quality policy", "text": "Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified."}, {"speaker": "Quality policy", "text": "No assumptions are allowed."}, {"speaker": "Quality policy", "text": "Without final measurements and inside-seam evidence, it must remain in measurement and seam verification."}, {"speaker": "Measurement record", "text": "The recorded final dimensions of the blue cushion cover are 45 cm wide and 45 cm high."}, {"speaker": "Exterior review", "text": "The submitted exterior photograph of the blue cushion cover shows a centered zipper and every exterior thread end clipped."}, {"speaker": "Submission log", "text": "Photograph I is the only submitted inside-seam photograph."}, {"speaker": "Seam review", "text": "Photograph I shows that every raw edge in the photographed cover is enclosed."}, {"speaker": "Exposure record", "text": "During Photograph I's exposure at 14:08:17 on 6 May 2025, the camera's entire field of view contained only inventory item Q-17 against a blank backdrop, with no reflections or displayed images present."}, {"speaker": "Inventory record", "text": "At 14:08:17 on 6 May 2025, inventory item Q-17 was a beige table runner distinct from the blue cushion cover."}, {"speaker": "Visual audit", "text": "No cosmetic flaw is visible in the submitted evidence for the blue cushion cover."}]}, "method": "c2d", "provenance": {"source_id": "diverse-201", "source_is_synthetic": true, "source_sha256": "948210211c01bfb1c47fea1b9af1b2983c05fcb405b70f31e920771698ff4af9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rating scale, evidence requirements, no-assumptions rule, request scope, and binding to the blue cushion cover. The two focus spans are complete factual sentences. The counterfactual coherently changes Q-62 from depicting the blue cover to depicting only a red cover; the inside-seam statement remains about the photographed cover, so it does not contradict that change, and no other unchanged statement identifies Q-62 as blue. Neither context embeds an answer code, gold decision, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual or evidentiary relationship; the universally quantified atoms remain single relationships over explicit evidence sets. The focus atom concerns whether a photograph depicts the blue cushion cover, so it is factual rather than policy. The base and counter assignments are jointly realizable with only that photograph-identity fact changing: Photograph I can show enclosed edges in either scenario while depicting the blue cover only in the base. The policy evidence properly preserves the governing requirements and evidence rules originating in the original state; instructions already retained in the questions object need not be duplicated. Both rules are sufficient for their targets, and uncovered assignments may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions provide recorded 45 × 45 cm dimensions, photographic evidence of a centered zipper and clipped exterior threads, and an inside-seam photograph identified as depicting the blue cover and showing all its raw edges enclosed. The no-visible-cosmetic-flaw condition also excludes a competing lower finish assessment based on submitted evidence. This is sufficient for completion as Good under the stated evidence policy.", "rule_index": 0, "sound": true}, {"reason": "Because Photograph I is the only submitted inside-seam photograph and is explicitly refuted as depicting the blue cover, its enclosed-edge evidence cannot verify that cover. Thus the blue cover lacks required inside-seam evidence, which blocks completion as Good.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The recorded final width of the blue cushion cover is 45 cm."}, {"id": "a2", "statement": "The recorded final height of the blue cushion cover is 45 cm."}, {"id": "a3", "statement": "The submitted exterior photograph of the blue cushion cover shows that its zipper is centered."}, {"id": "a4", "statement": "The submitted exterior photograph of the blue cushion cover shows that every exterior thread end is clipped."}, {"id": "a5", "statement": "Photograph I is the only submitted inside-seam photograph."}, {"id": "a6", "statement": "Photograph I shows that every raw edge in the photographed cover is enclosed."}, {"id": "a7", "statement": "Photograph I depicts the blue cushion cover."}, {"id": "a8", "statement": "No cosmetic flaw is visible in the submitted evidence for the blue cushion cover."}], "base_state_json": "[{\"speaker\":\"Submission clerk\",\"text\":\"The final measurement record for the blue cushion cover gives a width of 45 cm and a height of 45 cm. The submitted exterior photograph shows a centered zipper and every exterior thread end clipped. No cosmetic flaw is visible anywhere in the submitted evidence for that cover.\"},{\"speaker\":\"Evidence auditor\",\"text\":\"Photograph I is the only submitted inside-seam photograph. It shows that every raw edge in the photographed cover is enclosed.\"},{\"speaker\":\"Archive custodian\",\"text\":\"In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62.\"},{\"speaker\":\"Image examiner\",\"text\":\"Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows the blue cushion cover.\"},{\"speaker\":\"Project sheet\",\"text\":\"The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"Our scale is Needs correction, Acceptable, then Good.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"No assumptions are allowed.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"Without final measurements and inside-seam evidence, it must remain in measurement and seam verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": ["2", "text"], "text": "In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62."}, {"path": ["3", "text"], "text": "Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows the blue cushion cover."}], "policy_evidence": [{"path": ["0", "text"], "text": "The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges."}, {"path": ["1", "text"], "text": "Our scale is Needs correction, Acceptable, then Good."}, {"path": ["1", "text"], "text": "Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct."}, {"path": ["1", "text"], "text": "Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified."}, {"path": ["3", "text"], "text": "No assumptions are allowed."}, {"path": ["3", "text"], "text": "Without final measurements and inside-seam evidence, it must remain in measurement and seam verification."}], "rules": [{"justification": "The recorded dimensions meet 45 × 45 cm, the exterior evidence verifies the centered zipper and clipped threads, and the sole submitted inside-seam photograph both depicts the blue cushion cover and verifies its enclosed raw edges. With no visible cosmetic flaw, every sheet requirement has evidence and no competing Acceptable outcome is indicated.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Photograph I is the only submitted inside-seam photograph, but it does not depict the blue cushion cover. Therefore the blue cushion cover lacks inside-seam evidence verifying its enclosed raw edges. Missing evidence for any sheet requirement blocks completion as Good, and assumptions are prohibited.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62.", "negative_left": "In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62.", "negative_right": "Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows only a red cushion cover and contains no depiction of the blue cushion cover.", "right": "Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows the blue cushion cover."}, "verifier_independent_model": false}, "family": "scale-diverse-201-002", "id": "scale-diverse-201-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it complete as Good; keep it in measurement and seam verification.", "true": "Yes — mark the cushion cover complete and rate its finish quality Good."}, "instructions": "Decide whether the cushion cover should be marked complete with a finish-quality rating of Good. Answer yes or no using the stated evidence policy.", "type": "noul"}}, "state": [{"speaker": "Submission clerk", "text": "The final measurement record for the blue cushion cover gives a width of 45 cm and a height of 45 cm. The submitted exterior photograph shows a centered zipper and every exterior thread end clipped. No cosmetic flaw is visible anywhere in the submitted evidence for that cover."}, {"speaker": "Evidence auditor", "text": "Photograph I is the only submitted inside-seam photograph. It shows that every raw edge in the photographed cover is enclosed."}, {"speaker": "Archive custodian", "text": "In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62."}, {"speaker": "Image examiner", "text": "Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows the blue cushion cover."}, {"speaker": "Project sheet", "text": "The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges."}, {"speaker": "Pattern and measurement checker", "text": "Our scale is Needs correction, Acceptable, then Good."}, {"speaker": "Pattern and measurement checker", "text": "Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct."}, {"speaker": "Pattern and measurement checker", "text": "Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified."}, {"speaker": "Pattern and measurement checker", "text": "No assumptions are allowed."}, {"speaker": "Pattern and measurement checker", "text": "Without final measurements and inside-seam evidence, it must remain in measurement and seam verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-201", "source_is_synthetic": true, "source_sha256": "948210211c01bfb1c47fea1b9af1b2983c05fcb405b70f31e920771698ff4af9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rating scale, evidence requirements, no-assumptions rule, request scope, and binding to the blue cushion cover. The two focus spans are complete factual sentences. The counterfactual coherently changes Q-62 from depicting the blue cover to depicting only a red cover; the inside-seam statement remains about the photographed cover, so it does not contradict that change, and no other unchanged statement identifies Q-62 as blue. Neither context embeds an answer code, gold decision, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual or evidentiary relationship; the universally quantified atoms remain single relationships over explicit evidence sets. The focus atom concerns whether a photograph depicts the blue cushion cover, so it is factual rather than policy. The base and counter assignments are jointly realizable with only that photograph-identity fact changing: Photograph I can show enclosed edges in either scenario while depicting the blue cover only in the base. The policy evidence properly preserves the governing requirements and evidence rules originating in the original state; instructions already retained in the questions object need not be duplicated. Both rules are sufficient for their targets, and uncovered assignments may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions provide recorded 45 × 45 cm dimensions, photographic evidence of a centered zipper and clipped exterior threads, and an inside-seam photograph identified as depicting the blue cover and showing all its raw edges enclosed. The no-visible-cosmetic-flaw condition also excludes a competing lower finish assessment based on submitted evidence. This is sufficient for completion as Good under the stated evidence policy.", "rule_index": 0, "sound": true}, {"reason": "Because Photograph I is the only submitted inside-seam photograph and is explicitly refuted as depicting the blue cover, its enclosed-edge evidence cannot verify that cover. Thus the blue cover lacks required inside-seam evidence, which blocks completion as Good.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The recorded final width of the blue cushion cover is 45 cm."}, {"id": "a2", "statement": "The recorded final height of the blue cushion cover is 45 cm."}, {"id": "a3", "statement": "The submitted exterior photograph of the blue cushion cover shows that its zipper is centered."}, {"id": "a4", "statement": "The submitted exterior photograph of the blue cushion cover shows that every exterior thread end is clipped."}, {"id": "a5", "statement": "Photograph I is the only submitted inside-seam photograph."}, {"id": "a6", "statement": "Photograph I shows that every raw edge in the photographed cover is enclosed."}, {"id": "a7", "statement": "Photograph I depicts the blue cushion cover."}, {"id": "a8", "statement": "No cosmetic flaw is visible in the submitted evidence for the blue cushion cover."}], "base_state_json": "[{\"speaker\":\"Submission clerk\",\"text\":\"The final measurement record for the blue cushion cover gives a width of 45 cm and a height of 45 cm. The submitted exterior photograph shows a centered zipper and every exterior thread end clipped. No cosmetic flaw is visible anywhere in the submitted evidence for that cover.\"},{\"speaker\":\"Evidence auditor\",\"text\":\"Photograph I is the only submitted inside-seam photograph. It shows that every raw edge in the photographed cover is enclosed.\"},{\"speaker\":\"Archive custodian\",\"text\":\"In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62.\"},{\"speaker\":\"Image examiner\",\"text\":\"Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows the blue cushion cover.\"},{\"speaker\":\"Project sheet\",\"text\":\"The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"Our scale is Needs correction, Acceptable, then Good.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"No assumptions are allowed.\"},{\"speaker\":\"Pattern and measurement checker\",\"text\":\"Without final measurements and inside-seam evidence, it must remain in measurement and seam verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": ["2", "text"], "text": "In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62."}, {"path": ["3", "text"], "text": "Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows the blue cushion cover."}], "policy_evidence": [{"path": ["0", "text"], "text": "The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges."}, {"path": ["1", "text"], "text": "Our scale is Needs correction, Acceptable, then Good."}, {"path": ["1", "text"], "text": "Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct."}, {"path": ["1", "text"], "text": "Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified."}, {"path": ["3", "text"], "text": "No assumptions are allowed."}, {"path": ["3", "text"], "text": "Without final measurements and inside-seam evidence, it must remain in measurement and seam verification."}], "rules": [{"justification": "The recorded dimensions meet 45 × 45 cm, the exterior evidence verifies the centered zipper and clipped threads, and the sole submitted inside-seam photograph both depicts the blue cushion cover and verifies its enclosed raw edges. With no visible cosmetic flaw, every sheet requirement has evidence and no competing Acceptable outcome is indicated.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Photograph I is the only submitted inside-seam photograph, but it does not depict the blue cushion cover. Therefore the blue cushion cover lacks inside-seam evidence verifying its enclosed raw edges. Missing evidence for any sheet requirement blocks completion as Good, and assumptions are prohibited.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62.", "negative_left": "In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62.", "negative_right": "Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows only a red cushion cover and contains no depiction of the blue cushion cover.", "right": "Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows the blue cushion cover."}, "verifier_independent_model": false}, "family": "scale-diverse-201-002", "id": "scale-diverse-201-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it complete as Good; keep it in measurement and seam verification.", "true": "Yes — mark the cushion cover complete and rate its finish quality Good."}, "instructions": "Decide whether the cushion cover should be marked complete with a finish-quality rating of Good. Answer yes or no using the stated evidence policy.", "type": "noul"}}, "state": [{"speaker": "Submission clerk", "text": "The final measurement record for the blue cushion cover gives a width of 45 cm and a height of 45 cm. The submitted exterior photograph shows a centered zipper and every exterior thread end clipped. No cosmetic flaw is visible anywhere in the submitted evidence for that cover."}, {"speaker": "Evidence auditor", "text": "Photograph I is the only submitted inside-seam photograph. It shows that every raw edge in the photographed cover is enclosed."}, {"speaker": "Archive custodian", "text": "In the submission archive sealed at 14:00 UTC on 12 March 2026, Photograph I is the image assigned the unique archive identifier Q-62."}, {"speaker": "Image examiner", "text": "Within that sealed submission archive, the image assigned the unique archive identifier Q-62 shows only a red cushion cover and contains no depiction of the blue cushion cover."}, {"speaker": "Project sheet", "text": "The project sheet requires 45 × 45 cm finished dimensions, a centered zipper, clipped threads, and enclosed raw edges."}, {"speaker": "Pattern and measurement checker", "text": "Our scale is Needs correction, Acceptable, then Good."}, {"speaker": "Pattern and measurement checker", "text": "Good requires evidence for every sheet requirement; missing evidence blocks completion even when the visible work looks correct."}, {"speaker": "Pattern and measurement checker", "text": "Acceptable permits minor cosmetic flaws, but still requires dimensions and seam construction to be verified."}, {"speaker": "Pattern and measurement checker", "text": "No assumptions are allowed."}, {"speaker": "Pattern and measurement checker", "text": "Without final measurements and inside-seam evidence, it must remain in measurement and seam verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-201", "source_is_synthetic": true, "source_sha256": "948210211c01bfb1c47fea1b9af1b2983c05fcb405b70f31e920771698ff4af9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original scope, measurement tolerance, approved exclusion, quality ordering, and Good-quality defect limit without adding a policy rule or default. The project, covers, participants, and completion-and-Good decision remain bound to the same entities and requirements. The two evidence spans are complete factual sentences. The counterfactual coherently changes only C29’s logged measurement from 40.4 cm to 40.8 cm, with no duplicate conflicting measurement in that context. Neither context states an answer, answer code, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm. At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.4 cm. At 09:15, Inez confirmed that each cover’s envelope closure worked, every loose thread on both covers was trimmed, and every visible seam on both was pressed. At 09:22, she observed exactly one visible topstitch wobble across C17 and C29 and measured it at 3 mm. Only the hidden inner raw edges remained unfinished. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm."}, {"path": [], "text": "At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.4 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm.", "negative_left": "At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm.", "negative_right": "At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.8 cm.", "right": "At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.4 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-202-001", "id": "scale-diverse-202-001-base", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm. At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.4 cm. At 09:15, Inez confirmed that each cover’s envelope closure worked, every loose thread on both covers was trimmed, and every visible seam on both was pressed. At 09:22, she observed exactly one visible topstitch wobble across C17 and C29 and measured it at 3 mm. Only the hidden inner raw edges remained unfinished. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original scope, measurement tolerance, approved exclusion, quality ordering, and Good-quality defect limit without adding a policy rule or default. The project, covers, participants, and completion-and-Good decision remain bound to the same entities and requirements. The two evidence spans are complete factual sentences. The counterfactual coherently changes only C29’s logged measurement from 40.4 cm to 40.8 cm, with no duplicate conflicting measurement in that context. Neither context states an answer, answer code, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm. At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.4 cm. At 09:15, Inez confirmed that each cover’s envelope closure worked, every loose thread on both covers was trimmed, and every visible seam on both was pressed. At 09:22, she observed exactly one visible topstitch wobble across C17 and C29 and measured it at 3 mm. Only the hidden inner raw edges remained unfinished. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm."}, {"path": [], "text": "At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.4 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm.", "negative_left": "At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm.", "negative_right": "At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.8 cm.", "right": "At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.4 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-202-001", "id": "scale-diverse-202-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "At 09:00 on 14 May 2026, Mara's stated project scope consisted of exactly two cushion covers, C17 and C29, and the calibrated inspection log recorded C17's measurement as 39.7 cm. At 09:08 on 14 May 2026, the same calibrated inspection log recorded C29's measurement as 40.8 cm. At 09:15, Inez confirmed that each cover’s envelope closure worked, every loose thread on both covers was trimmed, and every visible seam on both was pressed. At 09:22, she observed exactly one visible topstitch wobble across C17 and C29 and measured it at 3 mm. Only the hidden inner raw edges remained unfinished. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scope requirements, approved exclusion, quality ordering, and Good/Excellent defect limits without adding policy defaults or priorities. The same project, participants, two-cover scope, and relevant measurement/quality decision remain bound throughout. The focus evidence contains exactly two complete factual sentences describing the scoped items and recorded measurements. The counterfactual changes only MC-24’s measurement from 40.4 cm to 40.8 cm and introduces no conflicting duplicate measurement or other inconsistency. Neither context states the decision, supplies an answer code, or instructs the classifier what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope. The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.4 cm. Inez tested both listed items: each envelope closure worked, every loose thread was trimmed, and every visible seam was pressed. She found one visible topstitch wobble across the pair and measured it at 3 mm. The only unfinished areas were the hidden inner raw edges. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope."}, {"path": [], "text": "The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.4 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope.", "negative_left": "Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope.", "negative_right": "The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.8 cm.", "right": "The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.4 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-202-002", "id": "scale-diverse-202-002-base", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope. The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.4 cm. Inez tested both listed items: each envelope closure worked, every loose thread was trimmed, and every visible seam was pressed. She found one visible topstitch wobble across the pair and measured it at 3 mm. The only unfinished areas were the hidden inner raw edges. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scope requirements, approved exclusion, quality ordering, and Good/Excellent defect limits without adding policy defaults or priorities. The same project, participants, two-cover scope, and relevant measurement/quality decision remain bound throughout. The focus evidence contains exactly two complete factual sentences describing the scoped items and recorded measurements. The counterfactual changes only MC-24’s measurement from 40.4 cm to 40.8 cm and introduces no conflicting duplicate measurement or other inconsistency. Neither context states the decision, supplies an answer code, or instructs the classifier what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope. The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.4 cm. Inez tested both listed items: each envelope closure worked, every loose thread was trimmed, and every visible seam was pressed. She found one visible topstitch wobble across the pair and measured it at 3 mm. The only unfinished areas were the hidden inner raw edges. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope."}, {"path": [], "text": "The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.4 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope.", "negative_left": "Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope.", "negative_right": "The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.8 cm.", "right": "The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.4 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-202-002", "id": "scale-diverse-202-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "Mara’s scope ledger dated 12 September 2026 identifies items MC-17 and MC-24 as the only two cushion covers in her stated project scope. The final measurement record dated 14 September 2026 lists MC-17 at 39.7 cm and MC-24 at 40.8 cm. Inez tested both listed items: each envelope closure worked, every loose thread was trimmed, and every visible seam was pressed. She found one visible topstitch wobble across the pair and measured it at 3 mm. The only unfinished areas were the hidden inner raw edges. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scope requirements, approved exclusion, quality ordering, and Good/Excellent defect limits, while the unchanged questions preserve the decision criteria and instructions. The same cushion-cover project, responsible people, identified covers, inspection path, and handoff date are used in both contexts; only C24’s recorded measurement changes. The evidence consists of exactly two complete factual sentences. The 40.8 cm counterfactual measurement is coherent and does not conflict with another measurement in that context. Neither context contains a gold answer, output instruction, answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24. The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.4 cm. Inez tested the envelope closure on each identified cover, and both operated correctly. Her final inspection found all loose threads trimmed and every visible seam pressed on both covers. Across the identified covers, Inez observed one visible topstitch wobble, measured at 3 mm. The hidden inner raw edges remained unfinished. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24."}, {"path": [], "text": "The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.4 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24.", "negative_left": "At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24.", "negative_right": "The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.8 cm.", "right": "The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.4 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-202-003", "id": "scale-diverse-202-003-base", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24. The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.4 cm. Inez tested the envelope closure on each identified cover, and both operated correctly. Her final inspection found all loose threads trimmed and every visible seam pressed on both covers. Across the identified covers, Inez observed one visible topstitch wobble, measured at 3 mm. The hidden inner raw edges remained unfinished. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scope requirements, approved exclusion, quality ordering, and Good/Excellent defect limits, while the unchanged questions preserve the decision criteria and instructions. The same cushion-cover project, responsible people, identified covers, inspection path, and handoff date are used in both contexts; only C24’s recorded measurement changes. The evidence consists of exactly two complete factual sentences. The 40.8 cm counterfactual measurement is coherent and does not conflict with another measurement in that context. Neither context contains a gold answer, output instruction, answer code, proposition identifier, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24. The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.4 cm. Inez tested the envelope closure on each identified cover, and both operated correctly. Her final inspection found all loose threads trimmed and every visible seam pressed on both covers. Across the identified covers, Inez observed one visible topstitch wobble, measured at 3 mm. The hidden inner raw edges remained unfinished. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24."}, {"path": [], "text": "The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.4 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24.", "negative_left": "At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24.", "negative_right": "The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.8 cm.", "right": "The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.4 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-202-003", "id": "scale-diverse-202-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "At the 2026-08-14 operational handoff, Mara’s stated project scope consisted exactly of cushion covers C17 and C24. The final measurement register dated 2026-08-14 records C17 at 39.7 cm and C24 at 40.8 cm. Inez tested the envelope closure on each identified cover, and both operated correctly. Her final inspection found all loose threads trimmed and every visible seam pressed on both covers. Across the identified covers, Inez observed one visible topstitch wobble, measured at 3 mm. The hidden inner raw edges remained unfinished. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring policy, and neither context alters or adds governing rules. Both contexts remain bound to the same cushion cover and final decision assessment. The two evidence spans are complete factual sentences. The counterfactual changes only the visibility of part of the sole identified correction and does not conflict with its completion, size, or uniqueness. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Final assessor\",\"text\":\"At the decision assessment, every required operation on the cushion cover was finished. The zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed.\"},{\"speaker\":\"Measurement inspector\",\"text\":\"The finished horizontal and vertical dimensions were each 40.0 cm. Maximum visible stitch drift was 0.8 mm, and no puckering was visible.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"Before that assessment, exactly one correction had been made to the cover: a seam correction that was completed before inspection.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27.\"},{\"speaker\":\"Final assessor\",\"text\":\"At the final assessment used for question decision, every point of segment H-27 was fully concealed from every exposed viewing direction.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["3", "text"], "text": "The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27."}, {"path": ["4", "text"], "text": "At the final assessment used for question decision, every point of segment H-27 was fully concealed from every exposed viewing direction."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27.", "negative_left": "The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27.", "negative_right": "At the final assessment used for question decision, a continuous 4-millimeter portion of segment H-27 was directly visible from an exposed viewing direction.", "right": "At the final assessment used for question decision, every point of segment H-27 was fully concealed from every exposed viewing direction."}, "verifier_independent_model": false}, "family": "scale-diverse-204-003", "id": "scale-diverse-204-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Final assessor", "text": "At the decision assessment, every required operation on the cushion cover was finished. The zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed."}, {"speaker": "Measurement inspector", "text": "The finished horizontal and vertical dimensions were each 40.0 cm. Maximum visible stitch drift was 0.8 mm, and no puckering was visible."}, {"speaker": "Handoff recorder", "text": "Before that assessment, exactly one correction had been made to the cover: a seam correction that was completed before inspection."}, {"speaker": "Handoff recorder", "text": "The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27."}, {"speaker": "Final assessor", "text": "At the final assessment used for question decision, every point of segment H-27 was fully concealed from every exposed viewing direction."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring policy, and neither context alters or adds governing rules. Both contexts remain bound to the same cushion cover and final decision assessment. The two evidence spans are complete factual sentences. The counterfactual changes only the visibility of part of the sole identified correction and does not conflict with its completion, size, or uniqueness. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Final assessor\",\"text\":\"At the decision assessment, every required operation on the cushion cover was finished. The zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed.\"},{\"speaker\":\"Measurement inspector\",\"text\":\"The finished horizontal and vertical dimensions were each 40.0 cm. Maximum visible stitch drift was 0.8 mm, and no puckering was visible.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"Before that assessment, exactly one correction had been made to the cover: a seam correction that was completed before inspection.\"},{\"speaker\":\"Handoff recorder\",\"text\":\"The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27.\"},{\"speaker\":\"Final assessor\",\"text\":\"At the final assessment used for question decision, every point of segment H-27 was fully concealed from every exposed viewing direction.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["3", "text"], "text": "The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27."}, {"path": ["4", "text"], "text": "At the final assessment used for question decision, every point of segment H-27 was fully concealed from every exposed viewing direction."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27.", "negative_left": "The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27.", "negative_right": "At the final assessment used for question decision, a continuous 4-millimeter portion of segment H-27 was directly visible from an exposed viewing direction.", "right": "At the final assessment used for question decision, every point of segment H-27 was fully concealed from every exposed viewing direction."}, "verifier_independent_model": false}, "family": "scale-diverse-204-003", "id": "scale-diverse-204-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Final assessor", "text": "At the decision assessment, every required operation on the cushion cover was finished. The zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed."}, {"speaker": "Measurement inspector", "text": "The finished horizontal and vertical dimensions were each 40.0 cm. Maximum visible stitch drift was 0.8 mm, and no puckering was visible."}, {"speaker": "Handoff recorder", "text": "Before that assessment, exactly one correction had been made to the cover: a seam correction that was completed before inspection."}, {"speaker": "Handoff recorder", "text": "The operational handoff record states that the sole correction on the cushion cover being rated consists entirely of the 17-millimeter restitched segment identified as H-27."}, {"speaker": "Final assessor", "text": "At the final assessment used for question decision, a continuous 4-millimeter portion of segment H-27 was directly visible from an exposed viewing direction."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring policy, while both contexts remain bound to the same cushion cover and final assessment. The two evidence spans are complete factual sentences. The counterfactual coherently replaces complete concealment of S-48 with a 2-millimeter visible portion without conflicting with the enclosed-raw-edge finding or the single documented correction. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Work log\",\"text\":\"Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48.\"},{\"speaker\":\"Construction inspector\",\"text\":\"At the final assessment, every required operation was finished. The zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed. The completed cover measured 40.0 cm horizontally and 40.1 cm vertically.\"},{\"speaker\":\"Facing inspector\",\"text\":\"At the final assessment used for question decision, every point of seam segment S-48 was hidden from view beneath opaque, intact facing.\"},{\"speaker\":\"Quality inspector\",\"text\":\"The maximum visible stitch drift at that assessment was 0.8 mm, and no puckering was visible. The production record documents exactly one correction: the completed resewing of a side seam. No other correction was made before the assessment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["0", "text"], "text": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48."}, {"path": ["2", "text"], "text": "At the final assessment used for question decision, every point of seam segment S-48 was hidden from view beneath opaque, intact facing."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48.", "negative_left": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48.", "negative_right": "At the final assessment used for question decision, a 2-millimeter portion of seam segment S-48 was exposed and plainly visible.", "right": "At the final assessment used for question decision, every point of seam segment S-48 was hidden from view beneath opaque, intact facing."}, "verifier_independent_model": false}, "family": "scale-diverse-204-004", "id": "scale-diverse-204-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Work log", "text": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48."}, {"speaker": "Construction inspector", "text": "At the final assessment, every required operation was finished. The zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed. The completed cover measured 40.0 cm horizontally and 40.1 cm vertically."}, {"speaker": "Facing inspector", "text": "At the final assessment used for question decision, every point of seam segment S-48 was hidden from view beneath opaque, intact facing."}, {"speaker": "Quality inspector", "text": "The maximum visible stitch drift at that assessment was 0.8 mm, and no puckering was visible. The production record documents exactly one correction: the completed resewing of a side seam. No other correction was made before the assessment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring policy, while both contexts remain bound to the same cushion cover and final assessment. The two evidence spans are complete factual sentences. The counterfactual coherently replaces complete concealment of S-48 with a 2-millimeter visible portion without conflicting with the enclosed-raw-edge finding or the single documented correction. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Work log\",\"text\":\"Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48.\"},{\"speaker\":\"Construction inspector\",\"text\":\"At the final assessment, every required operation was finished. The zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed. The completed cover measured 40.0 cm horizontally and 40.1 cm vertically.\"},{\"speaker\":\"Facing inspector\",\"text\":\"At the final assessment used for question decision, every point of seam segment S-48 was hidden from view beneath opaque, intact facing.\"},{\"speaker\":\"Quality inspector\",\"text\":\"The maximum visible stitch drift at that assessment was 0.8 mm, and no puckering was visible. The production record documents exactly one correction: the completed resewing of a side seam. No other correction was made before the assessment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["0", "text"], "text": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48."}, {"path": ["2", "text"], "text": "At the final assessment used for question decision, every point of seam segment S-48 was hidden from view beneath opaque, intact facing."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48.", "negative_left": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48.", "negative_right": "At the final assessment used for question decision, a 2-millimeter portion of seam segment S-48 was exposed and plainly visible.", "right": "At the final assessment used for question decision, every point of seam segment S-48 was hidden from view beneath opaque, intact facing."}, "verifier_independent_model": false}, "family": "scale-diverse-204-004", "id": "scale-diverse-204-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Work log", "text": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was confined exactly to the 7-millimeter resewn seam segment S-48."}, {"speaker": "Construction inspector", "text": "At the final assessment, every required operation was finished. The zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed. The completed cover measured 40.0 cm horizontally and 40.1 cm vertically."}, {"speaker": "Facing inspector", "text": "At the final assessment used for question decision, a 2-millimeter portion of seam segment S-48 was exposed and plainly visible."}, {"speaker": "Quality inspector", "text": "The maximum visible stitch drift at that assessment was 0.8 mm, and no puckering was visible. The production record documents exactly one correction: the completed resewing of a side seam. No other correction was made before the assessment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring rubric, inclusivity rule, scope, and cushion-cover decision task; neither context alters those policies or the assessed item and timepoint. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the final-assessment visibility of the identified correction and remains consistent with the unchanged correction count, completion facts, measurements, stitch-drift measurement, and lack of puckering. Neither context contains an answer, score code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Work chronology recorder\",\"text\":\"Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17.\"},{\"speaker\":\"Workshop log reviewer\",\"text\":\"The work log classifies that sole correction as a side-seam correction and records it as completed before the final assessment. It also confirms that every required operation was finished by assessment time.\"},{\"speaker\":\"Final assessment inspector\",\"text\":\"During the final assessment, the zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed. Calibrated measurements gave a horizontal finished dimension of 40.0 cm and a vertical finished dimension of 40.1 cm. The maximum visible stitch drift was 0.8 mm, and no puckering was visible anywhere on the cover.\"},{\"speaker\":\"Final assessment observer\",\"text\":\"At the final assessment used for question decision, all five stitches designated C17 were completely concealed from every viewing angle beneath an opaque folded seam allowance.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["0", "text"], "text": "Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17."}, {"path": ["3", "text"], "text": "At the final assessment used for question decision, all five stitches designated C17 were completely concealed from every viewing angle beneath an opaque folded seam allowance."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17.", "negative_left": "Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17.", "negative_right": "At the final assessment used for question decision, one of the five stitches designated C17 was unobstructed and visually discernible on the cover's exterior.", "right": "At the final assessment used for question decision, all five stitches designated C17 were completely concealed from every viewing angle beneath an opaque folded seam allowance."}, "verifier_independent_model": false}, "family": "scale-diverse-204-005", "id": "scale-diverse-204-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Work chronology recorder", "text": "Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17."}, {"speaker": "Workshop log reviewer", "text": "The work log classifies that sole correction as a side-seam correction and records it as completed before the final assessment. It also confirms that every required operation was finished by assessment time."}, {"speaker": "Final assessment inspector", "text": "During the final assessment, the zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed. Calibrated measurements gave a horizontal finished dimension of 40.0 cm and a vertical finished dimension of 40.1 cm. The maximum visible stitch drift was 0.8 mm, and no puckering was visible anywhere on the cover."}, {"speaker": "Final assessment observer", "text": "At the final assessment used for question decision, all five stitches designated C17 were completely concealed from every viewing angle beneath an opaque folded seam allowance."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring rubric, inclusivity rule, scope, and cushion-cover decision task; neither context alters those policies or the assessed item and timepoint. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the final-assessment visibility of the identified correction and remains consistent with the unchanged correction count, completion facts, measurements, stitch-drift measurement, and lack of puckering. Neither context contains an answer, score code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Work chronology recorder\",\"text\":\"Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17.\"},{\"speaker\":\"Workshop log reviewer\",\"text\":\"The work log classifies that sole correction as a side-seam correction and records it as completed before the final assessment. It also confirms that every required operation was finished by assessment time.\"},{\"speaker\":\"Final assessment inspector\",\"text\":\"During the final assessment, the zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed. Calibrated measurements gave a horizontal finished dimension of 40.0 cm and a vertical finished dimension of 40.1 cm. The maximum visible stitch drift was 0.8 mm, and no puckering was visible anywhere on the cover.\"},{\"speaker\":\"Final assessment observer\",\"text\":\"At the final assessment used for question decision, all five stitches designated C17 were completely concealed from every viewing angle beneath an opaque folded seam allowance.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["0", "text"], "text": "Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17."}, {"path": ["3", "text"], "text": "At the final assessment used for question decision, all five stitches designated C17 were completely concealed from every viewing angle beneath an opaque folded seam allowance."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17.", "negative_left": "Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17.", "negative_right": "At the final assessment used for question decision, one of the five stitches designated C17 was unobstructed and visually discernible on the cover's exterior.", "right": "At the final assessment used for question decision, all five stitches designated C17 were completely concealed from every viewing angle beneath an opaque folded seam allowance."}, "verifier_independent_model": false}, "family": "scale-diverse-204-005", "id": "scale-diverse-204-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Work chronology recorder", "text": "Before the final assessment used for question decision, the cushion cover being rated received exactly one correction, consisting entirely of five replacement stitches designated C17."}, {"speaker": "Workshop log reviewer", "text": "The work log classifies that sole correction as a side-seam correction and records it as completed before the final assessment. It also confirms that every required operation was finished by assessment time."}, {"speaker": "Final assessment inspector", "text": "During the final assessment, the zipper opened and closed through its full travel, every seam was secured, and every raw edge was enclosed. Calibrated measurements gave a horizontal finished dimension of 40.0 cm and a vertical finished dimension of 40.1 cm. The maximum visible stitch drift was 0.8 mm, and no puckering was visible anywhere on the cover."}, {"speaker": "Final assessment observer", "text": "At the final assessment used for question decision, one of the five stitches designated C17 was unobstructed and visually discernible on the cover's exterior."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring policy, and neither generated context changes or removes any additional governing policy from the original state. Both contexts concern the same cushion cover and final assessment; the two evidence spans are complete factual sentences, and the counterfactual changes only the visibility of part of Q-17 without conflicting with the completed-correction record or other measurements. Neither context contains a score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Correction log reviewer\",\"text\":\"Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area.\"},{\"speaker\":\"Final inspector\",\"text\":\"At the final assessment used for question decision, direct inspection found that no point within seam section Q-17 or seam section Q-23 could be seen.\"},{\"speaker\":\"Final inspector\",\"text\":\"For the assessment used to rate this cushion cover, every required operation was finished. The zipper opened and closed through its full travel; every seam was secured, and every raw edge was enclosed.\"},{\"speaker\":\"Measurement checker\",\"text\":\"At that same final assessment, the horizontal finished dimension was 40.0 cm and the vertical finished dimension was 40.1 cm. The maximum visible stitch drift was 0.8 mm, and no puckering was visible.\"},{\"speaker\":\"Work-record auditor\",\"text\":\"The records show exactly one correction before the final assessment: the seam correction described in the correction log. Work records and inspection confirm that this correction was completed before the final assessment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["0", "text"], "text": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area."}, {"path": ["1", "text"], "text": "At the final assessment used for question decision, direct inspection found that no point within seam section Q-17 or seam section Q-23 could be seen."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area.", "negative_left": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area.", "negative_right": "At the final assessment used for question decision, direct inspection found that a 4-millimeter portion within seam section Q-17 could be seen, while no point within seam section Q-23 could be seen.", "right": "At the final assessment used for question decision, direct inspection found that no point within seam section Q-17 or seam section Q-23 could be seen."}, "verifier_independent_model": false}, "family": "scale-diverse-204-006", "id": "scale-diverse-204-006-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Correction log reviewer", "text": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area."}, {"speaker": "Final inspector", "text": "At the final assessment used for question decision, direct inspection found that no point within seam section Q-17 or seam section Q-23 could be seen."}, {"speaker": "Final inspector", "text": "For the assessment used to rate this cushion cover, every required operation was finished. The zipper opened and closed through its full travel; every seam was secured, and every raw edge was enclosed."}, {"speaker": "Measurement checker", "text": "At that same final assessment, the horizontal finished dimension was 40.0 cm and the vertical finished dimension was 40.1 cm. The maximum visible stitch drift was 0.8 mm, and no puckering was visible."}, {"speaker": "Work-record auditor", "text": "The records show exactly one correction before the final assessment: the seam correction described in the correction log. Work records and inspection confirm that this correction was completed before the final assessment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full scoring policy, and neither generated context changes or removes any additional governing policy from the original state. Both contexts concern the same cushion cover and final assessment; the two evidence spans are complete factual sentences, and the counterfactual changes only the visibility of part of Q-17 without conflicting with the completed-correction record or other measurements. Neither context contains a score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Correction log reviewer\",\"text\":\"Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area.\"},{\"speaker\":\"Final inspector\",\"text\":\"At the final assessment used for question decision, direct inspection found that no point within seam section Q-17 or seam section Q-23 could be seen.\"},{\"speaker\":\"Final inspector\",\"text\":\"For the assessment used to rate this cushion cover, every required operation was finished. The zipper opened and closed through its full travel; every seam was secured, and every raw edge was enclosed.\"},{\"speaker\":\"Measurement checker\",\"text\":\"At that same final assessment, the horizontal finished dimension was 40.0 cm and the vertical finished dimension was 40.1 cm. The maximum visible stitch drift was 0.8 mm, and no puckering was visible.\"},{\"speaker\":\"Work-record auditor\",\"text\":\"The records show exactly one correction before the final assessment: the seam correction described in the correction log. Work records and inspection confirm that this correction was completed before the final assessment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["0", "text"], "text": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area."}, {"path": ["1", "text"], "text": "At the final assessment used for question decision, direct inspection found that no point within seam section Q-17 or seam section Q-23 could be seen."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area.", "negative_left": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area.", "negative_right": "At the final assessment used for question decision, direct inspection found that a 4-millimeter portion within seam section Q-17 could be seen, while no point within seam section Q-23 could be seen.", "right": "At the final assessment used for question decision, direct inspection found that no point within seam section Q-17 or seam section Q-23 could be seen."}, "verifier_independent_model": false}, "family": "scale-diverse-204-006", "id": "scale-diverse-204-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Correction log reviewer", "text": "Before the final assessment used for question decision, the sole correction on the cushion cover being rated was documented as comprising exactly seam sections Q-17 and Q-23 and no other area."}, {"speaker": "Final inspector", "text": "At the final assessment used for question decision, direct inspection found that a 4-millimeter portion within seam section Q-17 could be seen, while no point within seam section Q-23 could be seen."}, {"speaker": "Final inspector", "text": "For the assessment used to rate this cushion cover, every required operation was finished. The zipper opened and closed through its full travel; every seam was secured, and every raw edge was enclosed."}, {"speaker": "Measurement checker", "text": "At that same final assessment, the horizontal finished dimension was 40.0 cm and the vertical finished dimension was 40.1 cm. The maximum visible stitch drift was 0.8 mm, and no puckering was visible."}, {"speaker": "Work-record auditor", "text": "The records show exactly one correction before the final assessment: the seam correction described in the correction log. Work records and inspection confirm that this correction was completed before the final assessment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing, impact, and intervention policies and preserve the museum livestream, sponsor clip, left-channel, 18-second, intelligibility, cue, local-recording, and platform path bindings. The focus evidence contains exactly two complete factual sentences. The counterfactual changes only the local recording measurement from alternating nonzero samples to zero-valued samples; this is coherent with the unchanged nonzero source-program measurement and absent platform output because those describe distinct signal locations. Neither context embeds an answer option, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "supported", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "full_context_fact_states": {"base": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "supported", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "counterfactual": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "refuted", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "remove_left": {"a_local_left_retained": "unknown"}, "remove_right": {"a_local_left_retained": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_local_left_retained": "unknown"}, "negative_pair": {"a_local_left_retained": "refuted"}, "negative_sentence": {"a_local_left_retained": "unknown"}, "positive_pair": {"a_local_left_retained": "supported"}, "right": {"a_local_left_retained": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a bundled policy conclusion. The focus atom is the factual condition of whether the local recording retained the left channel. Base and counter assignments differ only on that fact and are both realizable: one describes a platform-only loss and the other a locally present loss. The policy evidence correctly preserves the substantive routing, impact, exception, and intervention rules originating in the original state; question-level criteria need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a platform-output left-channel defect with a clean local recording, a stereo specification excluding intentional mono, and a timely cue excluding missed-cue ownership. It also establishes an 18-second duration with intelligible speech, satisfying Minor. Under the stated criteria, this is sufficient for the streaming-technician, Minor, intervention bundle.", "rule_index": 0, "sound": true}, {"reason": "Refuting local retention while supporting the same left-channel loss on the stream establishes that the non-exempt stereo audio defect is present locally. The state policy therefore assigns it to the audio engineer, which makes both listed substantive bundles incorrect and is sufficient for none_of_above. The duration and intelligibility also establish Minor, though the differing route alone already distinguishes the result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_stream_left_loss", "statement": "During the scheduled museum livestream, the streaming platform output did not carry the sponsor clip's left-channel program audio."}, {"id": "a_local_left_retained", "statement": "During the sponsor clip, the local recording retained the clip's left-channel program audio."}, {"id": "a_stereo_spec", "statement": "The rundown specifies stereo presentation for the sponsor clip."}, {"id": "a_duration", "statement": "The sponsor clip's left-channel loss on the streaming platform lasted 18 seconds."}, {"id": "a_intelligible", "statement": "Speech remained intelligible during the sponsor clip's left-channel loss."}, {"id": "a_cue_timing", "statement": "The stage manager fired the sponsor clip cue at its scheduled time."}], "base_state_json": "\"At the scheduled museum livestream, the stage manager fired the sponsor clip cue at its scheduled time. The streaming-platform output carried no left-channel program audio from the clip during the affected interval. The platform loss lasted exactly 18 seconds, while speech remained intelligible throughout. From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second. For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second. The approved rundown specifies stereo presentation for the sponsor clip. Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention.\"", "base_states": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "supported"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}], "counter_states": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "refuted"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}], "focus_atom": "a_local_left_retained", "focus_evidence": [{"path": [], "text": "From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second."}, {"path": [], "text": "For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second."}], "policy_evidence": [{"path": [], "text": "Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention."}], "rules": [{"justification": "The left-channel defect is present in the streaming-platform output but absent from the local recording, the stereo specification excludes intentional mono, and the timely cue excludes a missed-cue route. The 18-second duration and intelligible speech satisfy Minor.", "target": "streaming_technician_minor_intervention", "when": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "supported"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}]}, {"justification": "The stereo clip's left-channel program audio is absent both from the platform output and from the local recording, so the defect is present locally and belongs to the audio engineer rather than either listed route. Although the 18-second intelligible incident is Minor, its correct route differs from both substantive bundles.", "target": "none_of_above", "when": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "refuted"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}]}]}, "verified_pair": {"left": "From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second.", "negative_left": "From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second.", "negative_right": "For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 zero-valued PCM samples at 48,000 samples per second.", "right": "For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second."}, "verifier_independent_model": false}, "family": "scale-diverse-217-001", "id": "scale-diverse-217-001-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when the correct route, intervention decision, or impact rating differs from both listed substantive decision bundles.", "stage_manager_minor_no_intervention": "Route to the stage manager, rate Minor, and take no intervention because the issue was a cueing problem that has ended.", "streaming_technician_minor_intervention": "Route to the streaming technician, rate Minor, and intervene because the defect was confined to the streaming platform while the local recording remained clean."}, "instructions": "Select the single option that correctly routes the incident, determines whether intervention is required, and rates its broadcast impact under the stated scope and exceptions.", "type": "choice"}}, "state": "At the scheduled museum livestream, the stage manager fired the sponsor clip cue at its scheduled time. The streaming-platform output carried no left-channel program audio from the clip during the affected interval. The platform loss lasted exactly 18 seconds, while speech remained intelligible throughout. From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second. For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second. The approved rundown specifies stereo presentation for the sponsor clip. Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention."}, "method": "c2d", "provenance": {"source_id": "diverse-217", "source_is_synthetic": true, "source_sha256": "56c54e57171feb278bcede4915e9a42909b58c37f80d15ac73b42a2c27244a6b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "streaming_technician_minor_intervention"}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing, impact, and intervention policies and preserve the museum livestream, sponsor clip, left-channel, 18-second, intelligibility, cue, local-recording, and platform path bindings. The focus evidence contains exactly two complete factual sentences. The counterfactual changes only the local recording measurement from alternating nonzero samples to zero-valued samples; this is coherent with the unchanged nonzero source-program measurement and absent platform output because those describe distinct signal locations. Neither context embeds an answer option, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "refuted", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "full_context_fact_states": {"base": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "supported", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "counterfactual": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "refuted", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "remove_left": {"a_local_left_retained": "unknown"}, "remove_right": {"a_local_left_retained": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_local_left_retained": "unknown"}, "negative_pair": {"a_local_left_retained": "refuted"}, "negative_sentence": {"a_local_left_retained": "unknown"}, "positive_pair": {"a_local_left_retained": "supported"}, "right": {"a_local_left_retained": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a bundled policy conclusion. The focus atom is the factual condition of whether the local recording retained the left channel. Base and counter assignments differ only on that fact and are both realizable: one describes a platform-only loss and the other a locally present loss. The policy evidence correctly preserves the substantive routing, impact, exception, and intervention rules originating in the original state; question-level criteria need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a platform-output left-channel defect with a clean local recording, a stereo specification excluding intentional mono, and a timely cue excluding missed-cue ownership. It also establishes an 18-second duration with intelligible speech, satisfying Minor. Under the stated criteria, this is sufficient for the streaming-technician, Minor, intervention bundle.", "rule_index": 0, "sound": true}, {"reason": "Refuting local retention while supporting the same left-channel loss on the stream establishes that the non-exempt stereo audio defect is present locally. The state policy therefore assigns it to the audio engineer, which makes both listed substantive bundles incorrect and is sufficient for none_of_above. The duration and intelligibility also establish Minor, though the differing route alone already distinguishes the result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_stream_left_loss", "statement": "During the scheduled museum livestream, the streaming platform output did not carry the sponsor clip's left-channel program audio."}, {"id": "a_local_left_retained", "statement": "During the sponsor clip, the local recording retained the clip's left-channel program audio."}, {"id": "a_stereo_spec", "statement": "The rundown specifies stereo presentation for the sponsor clip."}, {"id": "a_duration", "statement": "The sponsor clip's left-channel loss on the streaming platform lasted 18 seconds."}, {"id": "a_intelligible", "statement": "Speech remained intelligible during the sponsor clip's left-channel loss."}, {"id": "a_cue_timing", "statement": "The stage manager fired the sponsor clip cue at its scheduled time."}], "base_state_json": "\"At the scheduled museum livestream, the stage manager fired the sponsor clip cue at its scheduled time. The streaming-platform output carried no left-channel program audio from the clip during the affected interval. The platform loss lasted exactly 18 seconds, while speech remained intelligible throughout. From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second. For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second. The approved rundown specifies stereo presentation for the sponsor clip. Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention.\"", "base_states": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "supported"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}], "counter_states": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "refuted"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}], "focus_atom": "a_local_left_retained", "focus_evidence": [{"path": [], "text": "From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second."}, {"path": [], "text": "For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second."}], "policy_evidence": [{"path": [], "text": "Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention."}], "rules": [{"justification": "The left-channel defect is present in the streaming-platform output but absent from the local recording, the stereo specification excludes intentional mono, and the timely cue excludes a missed-cue route. The 18-second duration and intelligible speech satisfy Minor.", "target": "streaming_technician_minor_intervention", "when": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "supported"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}]}, {"justification": "The stereo clip's left-channel program audio is absent both from the platform output and from the local recording, so the defect is present locally and belongs to the audio engineer rather than either listed route. Although the 18-second intelligible incident is Minor, its correct route differs from both substantive bundles.", "target": "none_of_above", "when": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "refuted"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}]}]}, "verified_pair": {"left": "From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second.", "negative_left": "From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second.", "negative_right": "For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 zero-valued PCM samples at 48,000 samples per second.", "right": "For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second."}, "verifier_independent_model": false}, "family": "scale-diverse-217-001", "id": "scale-diverse-217-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when the correct route, intervention decision, or impact rating differs from both listed substantive decision bundles.", "stage_manager_minor_no_intervention": "Route to the stage manager, rate Minor, and take no intervention because the issue was a cueing problem that has ended.", "streaming_technician_minor_intervention": "Route to the streaming technician, rate Minor, and intervene because the defect was confined to the streaming platform while the local recording remained clean."}, "instructions": "Select the single option that correctly routes the incident, determines whether intervention is required, and rates its broadcast impact under the stated scope and exceptions.", "type": "choice"}}, "state": "At the scheduled museum livestream, the stage manager fired the sponsor clip cue at its scheduled time. The streaming-platform output carried no left-channel program audio from the clip during the affected interval. The platform loss lasted exactly 18 seconds, while speech remained intelligible throughout. From 20:14:06 through 20:14:24 UTC on 12 September 2026, the sponsor clip's left-channel program audio consisted of 864,000 PCM samples alternating between +1200 and -1200 at 48,000 samples per second. For the interval from 20:14:06 through 20:14:24 UTC on 12 September 2026, the local recording's left-channel track stored 864,000 zero-valued PCM samples at 48,000 samples per second. The approved rundown specifies stereo presentation for the sponsor clip. Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention."}, "method": "c2d", "provenance": {"source_id": "diverse-217", "source_is_synthetic": true, "source_sha256": "56c54e57171feb278bcede4915e9a42909b58c37f80d15ac73b42a2c27244a6b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy, impact scale, request, scheduled-livestream setting, incident unit, and 19:42 reset timing. The two focus-evidence spans are complete factual sentences. The counterfactual consistently makes U-9 the uniquely V-5831-tagged device while excluding L-204, so the exclusive equipment-register statement can identify U-9 as A-310 without contradiction. Neither context states an answer option, code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"context\":\"Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.\",\"evidence\":[\"During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag.\",\"During the scheduled awards livestream, local confidence display L-204 carried asset tag V-5831.\",\"Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.\",\"The equipment register records U-9 as exactly one of local confidence display L-204 and public-program audio processor A-310.\",\"U-9’s malfunction interrupted its output for eight seconds during the livestream.\",\"A technician reset U-9 at 19:42, and it returned to normal operation immediately afterward.\",\"The incident log records no fault other than U-9’s malfunction and no editorial decision during the livestream.\",\"L-204’s sole function during the livestream was providing the local confidence display.\",\"A-310’s output supplied the public program sound during the livestream.\",\"The audience-facing broadcast remained usable throughout U-9’s output interruption.\"],\"request\":\"Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag."}, {"path": ["evidence", "1"], "text": "During the scheduled awards livestream, local confidence display L-204 carried asset tag V-5831."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag.", "negative_left": "During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag.", "negative_right": "During the scheduled awards livestream, local confidence display L-204 carried asset tag R-7462 and did not carry asset tag V-5831.", "right": "During the scheduled awards livestream, local confidence display L-204 carried asset tag V-5831."}, "verifier_independent_model": false}, "family": "scale-diverse-218-002", "id": "scale-diverse-218-002-base", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.", "evidence": ["During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag.", "During the scheduled awards livestream, local confidence display L-204 carried asset tag V-5831.", "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.", "The equipment register records U-9 as exactly one of local confidence display L-204 and public-program audio processor A-310.", "U-9’s malfunction interrupted its output for eight seconds during the livestream.", "A technician reset U-9 at 19:42, and it returned to normal operation immediately afterward.", "The incident log records no fault other than U-9’s malfunction and no editorial decision during the livestream.", "L-204’s sole function during the livestream was providing the local confidence display.", "A-310’s output supplied the public program sound during the livestream.", "The audience-facing broadcast remained usable throughout U-9’s output interruption."], "request": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "streaming_level_1"}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy, impact scale, request, scheduled-livestream setting, incident unit, and 19:42 reset timing. The two focus-evidence spans are complete factual sentences. The counterfactual consistently makes U-9 the uniquely V-5831-tagged device while excluding L-204, so the exclusive equipment-register statement can identify U-9 as A-310 without contradiction. Neither context states an answer option, code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"context\":\"Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.\",\"evidence\":[\"During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag.\",\"During the scheduled awards livestream, local confidence display L-204 carried asset tag V-5831.\",\"Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.\",\"The equipment register records U-9 as exactly one of local confidence display L-204 and public-program audio processor A-310.\",\"U-9’s malfunction interrupted its output for eight seconds during the livestream.\",\"A technician reset U-9 at 19:42, and it returned to normal operation immediately afterward.\",\"The incident log records no fault other than U-9’s malfunction and no editorial decision during the livestream.\",\"L-204’s sole function during the livestream was providing the local confidence display.\",\"A-310’s output supplied the public program sound during the livestream.\",\"The audience-facing broadcast remained usable throughout U-9’s output interruption.\"],\"request\":\"Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag."}, {"path": ["evidence", "1"], "text": "During the scheduled awards livestream, local confidence display L-204 carried asset tag V-5831."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag.", "negative_left": "During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag.", "negative_right": "During the scheduled awards livestream, local confidence display L-204 carried asset tag R-7462 and did not carry asset tag V-5831.", "right": "During the scheduled awards livestream, local confidence display L-204 carried asset tag V-5831."}, "verifier_independent_model": false}, "family": "scale-diverse-218-002", "id": "scale-diverse-218-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.", "evidence": ["During the scheduled awards livestream, asset tag V-5831 was assigned to exactly one device, and malfunctioning unit U-9 carried that tag.", "During the scheduled awards livestream, local confidence display L-204 carried asset tag R-7462 and did not carry asset tag V-5831.", "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.", "The equipment register records U-9 as exactly one of local confidence display L-204 and public-program audio processor A-310.", "U-9’s malfunction interrupted its output for eight seconds during the livestream.", "A technician reset U-9 at 19:42, and it returned to normal operation immediately afterward.", "The incident log records no fault other than U-9’s malfunction and no editorial decision during the livestream.", "L-204’s sole function during the livestream was providing the local confidence display.", "A-310’s output supplied the public program sound during the livestream.", "The audience-facing broadcast remained usable throughout U-9’s output interruption."], "request": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "audio_level_2"}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all choice criteria and instructions, while both generated states retain the original routing policy and impact scale. The request, scheduled livestream, incident entity, and 19:42 reset remain bound consistently. Both focus spans are complete factual sentences. The base code assignments are coherent if U-9 and L-204 are two designations for the same device; the counterfactual assigns them distinct unique codes, which coherently makes them distinct and remains compatible with the exclusive identity statement. Neither context includes an answer key, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"audio_level_2\":\"Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.\",\"producer_level_0\":\"Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.\",\"stage_level_3\":\"Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.\",\"streaming_level_1\":\"Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted.\"},\"instructions\":\"Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.\",\"type\":\"choice\"}},\"state\":{\"context\":\"Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.\",\"evidence\":[\"During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to both the malfunctioning unit designated U-9 and the local confidence display designated L-204.\",\"For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices.\",\"Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.\",\"Operations records state that U-9 was exactly one of L-204 and public-program audio processor A-310.\",\"During the livestream, U-9 malfunctioned and its output was interrupted for eight seconds.\",\"A technician reset U-9 at 19:42, and normal operation resumed immediately afterward.\",\"No other fault and no editorial decision occurred during the livestream.\",\"L-204's sole function was providing the local confidence display, while A-310's output supplied the public program sound.\",\"The audience-facing broadcast remained usable throughout U-9's output interruption.\"],\"request\":\"Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to both the malfunctioning unit designated U-9 and the local confidence display designated L-204."}, {"path": ["state", "evidence", "1"], "text": "For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to both the malfunctioning unit designated U-9 and the local confidence display designated L-204.", "negative_left": "During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to the malfunctioning unit designated U-9 and device code RC-884 to the local confidence display designated L-204.", "negative_right": "For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices.", "right": "For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices."}, "verifier_independent_model": false}, "family": "scale-diverse-218-006", "id": "scale-diverse-218-006-base", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.", "evidence": ["During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to both the malfunctioning unit designated U-9 and the local confidence display designated L-204.", "For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices.", "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.", "Operations records state that U-9 was exactly one of L-204 and public-program audio processor A-310.", "During the livestream, U-9 malfunctioned and its output was interrupted for eight seconds.", "A technician reset U-9 at 19:42, and normal operation resumed immediately afterward.", "No other fault and no editorial decision occurred during the livestream.", "L-204's sole function was providing the local confidence display, while A-310's output supplied the public program sound.", "The audience-facing broadcast remained usable throughout U-9's output interruption."], "request": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "streaming_level_1"}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all choice criteria and instructions, while both generated states retain the original routing policy and impact scale. The request, scheduled livestream, incident entity, and 19:42 reset remain bound consistently. Both focus spans are complete factual sentences. The base code assignments are coherent if U-9 and L-204 are two designations for the same device; the counterfactual assigns them distinct unique codes, which coherently makes them distinct and remains compatible with the exclusive identity statement. Neither context includes an answer key, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"audio_level_2\":\"Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.\",\"producer_level_0\":\"Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.\",\"stage_level_3\":\"Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.\",\"streaming_level_1\":\"Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted.\"},\"instructions\":\"Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.\",\"type\":\"choice\"}},\"state\":{\"context\":\"Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.\",\"evidence\":[\"During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to both the malfunctioning unit designated U-9 and the local confidence display designated L-204.\",\"For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices.\",\"Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.\",\"Operations records state that U-9 was exactly one of L-204 and public-program audio processor A-310.\",\"During the livestream, U-9 malfunctioned and its output was interrupted for eight seconds.\",\"A technician reset U-9 at 19:42, and normal operation resumed immediately afterward.\",\"No other fault and no editorial decision occurred during the livestream.\",\"L-204's sole function was providing the local confidence display, while A-310's output supplied the public program sound.\",\"The audience-facing broadcast remained usable throughout U-9's output interruption.\"],\"request\":\"Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to both the malfunctioning unit designated U-9 and the local confidence display designated L-204."}, {"path": ["state", "evidence", "1"], "text": "For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to both the malfunctioning unit designated U-9 and the local confidence display designated L-204.", "negative_left": "During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to the malfunctioning unit designated U-9 and device code RC-884 to the local confidence display designated L-204.", "negative_right": "For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices.", "right": "For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices."}, "verifier_independent_model": false}, "family": "scale-diverse-218-006", "id": "scale-diverse-218-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.", "evidence": ["During the scheduled awards livestream, the reconciliation register assigned device code RC-731 to the malfunctioning unit designated U-9 and device code RC-884 to the local confidence display designated L-204.", "For the scheduled awards livestream, the reconciliation register assigned exactly one device code to each device and never assigned the same device code to two different devices.", "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.", "Operations records state that U-9 was exactly one of L-204 and public-program audio processor A-310.", "During the livestream, U-9 malfunctioned and its output was interrupted for eight seconds.", "A technician reset U-9 at 19:42, and normal operation resumed immediately afterward.", "No other fault and no editorial decision occurred during the livestream.", "L-204's sole function was providing the local confidence display, while A-310's output supplied the public program sound.", "The audience-facing broadcast remained usable throughout U-9's output interruption."], "request": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "audio_level_2"}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both generated inputs retain the questions verbatim, preserve the routing and impact policies from the original state, and keep the same request, incident, entities, and 19:42 event binding. The two focus-evidence spans are complete factual sentences. Changing U-9’s identifying code from 5836 to 9174 coherently changes its identity from L-204 to A-310 under the unchanged exclusive-identification facts, without creating duplicate or contradictory measurements. Neither context states a selected option, gold answer, output instruction, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"audio_level_2\":\"Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.\",\"producer_level_0\":\"Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.\",\"stage_level_3\":\"Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.\",\"streaming_level_1\":\"Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted.\"},\"instructions\":\"Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.\",\"type\":\"choice\"}},\"state\":{\"context\":\"Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.\",\"evidence\":[\"During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174.\",\"During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 5836.\",\"The malfunction interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, and normal operation resumed immediately afterward. No fault other than U-9's malfunction occurred, and no editorial decision occurred. Throughout the livestream, L-204's sole function was providing the local confidence display, while A-310's output supplied public-program sound. The audience-facing broadcast remained usable throughout U-9's output interruption.\",\"Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.\"],\"request\":\"Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174."}, {"path": ["state", "evidence", "1"], "text": "During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 5836."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174.", "negative_left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174.", "negative_right": "During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 9174.", "right": "During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 5836."}, "verifier_independent_model": false}, "family": "scale-diverse-218-008", "id": "scale-diverse-218-008-base", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.", "evidence": ["During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174.", "During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 5836.", "The malfunction interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, and normal operation resumed immediately afterward. No fault other than U-9's malfunction occurred, and no editorial decision occurred. Throughout the livestream, L-204's sole function was providing the local confidence display, while A-310's output supplied public-program sound. The audience-facing broadcast remained usable throughout U-9's output interruption.", "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."], "request": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "streaming_level_1"}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both generated inputs retain the questions verbatim, preserve the routing and impact policies from the original state, and keep the same request, incident, entities, and 19:42 event binding. The two focus-evidence spans are complete factual sentences. Changing U-9’s identifying code from 5836 to 9174 coherently changes its identity from L-204 to A-310 under the unchanged exclusive-identification facts, without creating duplicate or contradictory measurements. Neither context states a selected option, gold answer, output instruction, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":{\"audio_level_2\":\"Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.\",\"producer_level_0\":\"Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.\",\"stage_level_3\":\"Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.\",\"streaming_level_1\":\"Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted.\"},\"instructions\":\"Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.\",\"type\":\"choice\"}},\"state\":{\"context\":\"Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.\",\"evidence\":[\"During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174.\",\"During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 5836.\",\"The malfunction interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, and normal operation resumed immediately afterward. No fault other than U-9's malfunction occurred, and no editorial decision occurred. Throughout the livestream, L-204's sole function was providing the local confidence display, while A-310's output supplied public-program sound. The audience-facing broadcast remained usable throughout U-9's output interruption.\",\"Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped.\"],\"request\":\"Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174."}, {"path": ["state", "evidence", "1"], "text": "During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 5836."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174.", "negative_left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174.", "negative_right": "During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 9174.", "right": "During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 5836."}, "verifier_independent_model": false}, "family": "scale-diverse-218-008", "id": "scale-diverse-218-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer.", "evidence": ["During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310; L-204's unique device-identifying code was 5836, whereas A-310's was 9174.", "During the scheduled awards livestream, malfunctioning unit U-9's unique device-identifying code was 9174.", "The malfunction interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, and normal operation resumed immediately afterward. No fault other than U-9's malfunction occurred, and no editorial decision occurred. Throughout the livestream, L-204's sole function was providing the local confidence display, while A-310's output supplied public-program sound. The audience-facing broadcast remained usable throughout U-9's output interruption.", "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."], "request": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "audio_level_2"}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original decision scope: classification of a verified public-output video-freeze incident under the unchanged duration/content-loss rubric. The focus evidence consists of two complete factual sentences, the counterfactual coherently changes only the logged end time without creating duplicate or contradictory measurements, and neither context embeds a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"During the public preshow countdown, the control room was displaying the sponsor slate while the production team prepared for the scheduled opening.\"},{\"speaker\":\"Encoder operator\",\"text\":\"At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost.\"},{\"speaker\":\"Streaming technician\",\"text\":\"At 21:03:16.891 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output.\"},{\"speaker\":\"Archive reviewer\",\"text\":\"The incident identifier is consistent across the encoder entries and the encoded-output review; no separate freeze was logged during that portion of the countdown.\"},{\"speaker\":\"Event producer\",\"text\":\"After the incident ended, the public stream was stable and the preshow continued. The team retained the encoder record for classification and routing under the production rubric.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost."}, {"path": ["2", "text"], "text": "At 21:03:16.891 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost.", "negative_left": "At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost.", "negative_right": "At 21:03:15.990 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output.", "right": "At 21:03:16.891 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output."}, "verifier_independent_model": false}, "family": "scale-diverse-219-001", "id": "scale-diverse-219-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "During the public preshow countdown, the control room was displaying the sponsor slate while the production team prepared for the scheduled opening."}, {"speaker": "Encoder operator", "text": "At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost."}, {"speaker": "Streaming technician", "text": "At 21:03:16.891 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output."}, {"speaker": "Archive reviewer", "text": "The incident identifier is consistent across the encoder entries and the encoded-output review; no separate freeze was logged during that portion of the countdown."}, {"speaker": "Event producer", "text": "After the incident ended, the public stream was stable and the preshow continued. The team retained the encoder record for classification and routing under the production rubric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original decision scope: classification of a verified public-output video-freeze incident under the unchanged duration/content-loss rubric. The focus evidence consists of two complete factual sentences, the counterfactual coherently changes only the logged end time without creating duplicate or contradictory measurements, and neither context embeds a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"During the public preshow countdown, the control room was displaying the sponsor slate while the production team prepared for the scheduled opening.\"},{\"speaker\":\"Encoder operator\",\"text\":\"At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost.\"},{\"speaker\":\"Streaming technician\",\"text\":\"At 21:03:16.891 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output.\"},{\"speaker\":\"Archive reviewer\",\"text\":\"The incident identifier is consistent across the encoder entries and the encoded-output review; no separate freeze was logged during that portion of the countdown.\"},{\"speaker\":\"Event producer\",\"text\":\"After the incident ended, the public stream was stable and the preshow continued. The team retained the encoder record for classification and routing under the production rubric.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost."}, {"path": ["2", "text"], "text": "At 21:03:16.891 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost.", "negative_left": "At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost.", "negative_right": "At 21:03:15.990 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output.", "right": "At 21:03:16.891 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output."}, "verifier_independent_model": false}, "family": "scale-diverse-219-001", "id": "scale-diverse-219-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "During the public preshow countdown, the control room was displaying the sponsor slate while the production team prepared for the scheduled opening."}, {"speaker": "Encoder operator", "text": "At 21:03:14.240 UTC on 14 September 2026, the encoder log records the start of video-freeze incident F during the public preshow countdown, with no program content lost."}, {"speaker": "Streaming technician", "text": "At 21:03:15.990 UTC on 14 September 2026, the encoder log records the end of video-freeze incident F and confirms that F is present in the public encoded output."}, {"speaker": "Archive reviewer", "text": "The incident identifier is consistent across the encoder entries and the encoded-output review; no separate freeze was logged during that portion of the countdown."}, {"speaker": "Event producer", "text": "After the incident ended, the public stream was stable and the preshow continued. The team retained the encoder record for classification and routing under the production rubric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric, thresholds, evidence priority, scope, and routing policy, while neither context alters or contradicts them. Both contexts concern the same public-output video-freeze incident and preserve its entity, evidentiary path, and relevant timing structure; the counterfactual changes only the factual end timestamp. The two focus-evidence spans are complete factual sentences. The revised end time is later than the unchanged start time and creates no duplicate or contradictory measurement. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or directive telling the classifier which output to select.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"Operational handoff: incident F occurred while the sponsor slate was displayed in the public preshow countdown. Engineering has completed the output and continuity checks.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.842 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log entry containing these timestamps identifies F as a video freeze during the public preshow countdown. Review of the encoded-output recording confirms that F is visible in the public encoded output.\"},{\"speaker\":\"Event producer\",\"text\":\"The program continuity review found that no program content was lost during incident F. Every scheduled preshow element and cue remained present in the recording.\"},{\"speaker\":\"Stage manager\",\"text\":\"The live feed is currently stable. Keep incident F attached to the preshow handoff record for disposition under the stated impact rubric.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC."}, {"path": ["2", "text"], "text": "The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.842 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC.", "negative_left": "The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC.", "negative_right": "The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.125 UTC.", "right": "The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.842 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-219-003", "id": "scale-diverse-219-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "Operational handoff: incident F occurred while the sponsor slate was displayed in the public preshow countdown. Engineering has completed the output and continuity checks."}, {"speaker": "Streaming technician", "text": "The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC."}, {"speaker": "Streaming technician", "text": "The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.842 UTC."}, {"speaker": "Streaming technician", "text": "The encoder log entry containing these timestamps identifies F as a video freeze during the public preshow countdown. Review of the encoded-output recording confirms that F is visible in the public encoded output."}, {"speaker": "Event producer", "text": "The program continuity review found that no program content was lost during incident F. Every scheduled preshow element and cue remained present in the recording."}, {"speaker": "Stage manager", "text": "The live feed is currently stable. Keep incident F attached to the preshow handoff record for disposition under the stated impact rubric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric, thresholds, evidence priority, scope, and routing policy, while neither context alters or contradicts them. Both contexts concern the same public-output video-freeze incident and preserve its entity, evidentiary path, and relevant timing structure; the counterfactual changes only the factual end timestamp. The two focus-evidence spans are complete factual sentences. The revised end time is later than the unchanged start time and creates no duplicate or contradictory measurement. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or directive telling the classifier which output to select.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"Operational handoff: incident F occurred while the sponsor slate was displayed in the public preshow countdown. Engineering has completed the output and continuity checks.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.842 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log entry containing these timestamps identifies F as a video freeze during the public preshow countdown. Review of the encoded-output recording confirms that F is visible in the public encoded output.\"},{\"speaker\":\"Event producer\",\"text\":\"The program continuity review found that no program content was lost during incident F. Every scheduled preshow element and cue remained present in the recording.\"},{\"speaker\":\"Stage manager\",\"text\":\"The live feed is currently stable. Keep incident F attached to the preshow handoff record for disposition under the stated impact rubric.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC."}, {"path": ["2", "text"], "text": "The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.842 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC.", "negative_left": "The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC.", "negative_right": "The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.125 UTC.", "right": "The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.842 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-219-003", "id": "scale-diverse-219-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "Operational handoff: incident F occurred while the sponsor slate was displayed in the public preshow countdown. Engineering has completed the output and continuity checks."}, {"speaker": "Streaming technician", "text": "The encoder log records the start of video-freeze incident F at 2026-08-14 19:42:11.375 UTC."}, {"speaker": "Streaming technician", "text": "The encoder log records the end of video-freeze incident F at 2026-08-14 19:42:13.125 UTC."}, {"speaker": "Streaming technician", "text": "The encoder log entry containing these timestamps identifies F as a video freeze during the public preshow countdown. Review of the encoded-output recording confirms that F is visible in the public encoded output."}, {"speaker": "Event producer", "text": "The program continuity review found that no program content was lost during incident F. Every scheduled preshow element and cue remained present in the recording."}, {"speaker": "Stage manager", "text": "The live feed is currently stable. Keep incident F attached to the preshow handoff record for disposition under the stated impact rubric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric, source priority, scope, and routing policy; neither context alters or adds to them. Both contexts concern one verified public-output video freeze during the preshow countdown, measured by the authoritative encoder log with no program-content loss. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the logged end time from 19:04:13.684 to 19:04:13.102, with no duplicate or conflicting measurement elsewhere. Neither context states a level, decision, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Shift lead\",\"text\":\"Field note: incident F was assigned during the preshow review. Staff were asked to preserve the authoritative encoder record and avoid relying on rough visual estimates.\"},{\"speaker\":\"Streaming technician\",\"text\":\"Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content.\"},{\"speaker\":\"Encoder operator\",\"text\":\"Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.684 UTC.\"},{\"speaker\":\"Quality coordinator\",\"text\":\"The export of EL-47 passed its integrity check. Both entries carry the same incident identifier and refer to one continuous freeze; the encoder clock was synchronized before the event, and no manual timestamp corrections were applied.\"},{\"speaker\":\"Event producer\",\"text\":\"The scheduled opening had not yet begun when F occurred. Operations retained the incident record for classification and technician routing after confirming that the stream had returned to normal.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content."}, {"path": ["2", "text"], "text": "Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.684 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content.", "negative_left": "Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content.", "negative_right": "Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.102 UTC.", "right": "Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.684 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-219-004", "id": "scale-diverse-219-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Shift lead", "text": "Field note: incident F was assigned during the preshow review. Staff were asked to preserve the authoritative encoder record and avoid relying on rough visual estimates."}, {"speaker": "Streaming technician", "text": "Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content."}, {"speaker": "Encoder operator", "text": "Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.684 UTC."}, {"speaker": "Quality coordinator", "text": "The export of EL-47 passed its integrity check. Both entries carry the same incident identifier and refer to one continuous freeze; the encoder clock was synchronized before the event, and no manual timestamp corrections were applied."}, {"speaker": "Event producer", "text": "The scheduled opening had not yet begun when F occurred. Operations retained the incident record for classification and technician routing after confirming that the stream had returned to normal."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric, source priority, scope, and routing policy; neither context alters or adds to them. Both contexts concern one verified public-output video freeze during the preshow countdown, measured by the authoritative encoder log with no program-content loss. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the logged end time from 19:04:13.684 to 19:04:13.102, with no duplicate or conflicting measurement elsewhere. Neither context states a level, decision, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Shift lead\",\"text\":\"Field note: incident F was assigned during the preshow review. Staff were asked to preserve the authoritative encoder record and avoid relying on rough visual estimates.\"},{\"speaker\":\"Streaming technician\",\"text\":\"Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content.\"},{\"speaker\":\"Encoder operator\",\"text\":\"Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.684 UTC.\"},{\"speaker\":\"Quality coordinator\",\"text\":\"The export of EL-47 passed its integrity check. Both entries carry the same incident identifier and refer to one continuous freeze; the encoder clock was synchronized before the event, and no manual timestamp corrections were applied.\"},{\"speaker\":\"Event producer\",\"text\":\"The scheduled opening had not yet begun when F occurred. Operations retained the incident record for classification and technician routing after confirming that the stream had returned to normal.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content."}, {"path": ["2", "text"], "text": "Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.684 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content.", "negative_left": "Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content.", "negative_right": "Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.102 UTC.", "right": "Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.684 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-219-004", "id": "scale-diverse-219-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Shift lead", "text": "Field note: incident F was assigned during the preshow review. Staff were asked to preserve the authoritative encoder record and avoid relying on rough visual estimates."}, {"speaker": "Streaming technician", "text": "Encoder log EL-47 records public-output video-freeze incident F during the public preshow countdown as starting at 2026-08-14 19:04:11.275 UTC and causing no loss of program content."}, {"speaker": "Encoder operator", "text": "Encoder log EL-47 records public-output video-freeze incident F as ending at 2026-08-14 19:04:13.102 UTC."}, {"speaker": "Quality coordinator", "text": "The export of EL-47 passed its integrity check. Both entries carry the same incident identifier and refer to one continuous freeze; the encoder clock was synchronized before the event, and no manual timestamp corrections were applied."}, {"speaker": "Event producer", "text": "The scheduled opening had not yet begun when F occurred. Operations retained the incident record for classification and technician routing after confirming that the stream had returned to normal."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric, threshold, source preference, output-path requirement, and routing policy; neither context alters those rules. Both contexts concern a video-freeze incident documented by the public encoder during the preshow, preserving the relevant entity, path, and scope bindings. The two evidence spans are complete factual sentences. The counterfactual changes only the recorded end time from 19:42:19.114Z to 19:42:18.114Z, remains later than the unchanged start time, and creates no duplicate or contradictory measurement. Neither context contains a gold label, answer code, rule table, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Broadcast coordinator\",\"text\":\"During the preshow, the operations team opened the archived encoder record associated with the public feed and labeled the reviewed event as incident F.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output.\"},{\"speaker\":\"Content logger\",\"text\":\"A frame-by-frame review found that the countdown slate resumed at the expected point; every scheduled item remained available, and no portion of the program was omitted.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records video-freeze incident F ending at 2026-08-12T19:42:19.114Z, with no program content lost during F.\"},{\"speaker\":\"Broadcast coordinator\",\"text\":\"After the event, the public feed continued through the remaining countdown sequence and into the scheduled opening without another freeze.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output."}, {"path": ["3", "text"], "text": "The encoder log records video-freeze incident F ending at 2026-08-12T19:42:19.114Z, with no program content lost during F."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output.", "negative_left": "The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output.", "negative_right": "The encoder log records video-freeze incident F ending at 2026-08-12T19:42:18.114Z, with no program content lost during F.", "right": "The encoder log records video-freeze incident F ending at 2026-08-12T19:42:19.114Z, with no program content lost during F."}, "verifier_independent_model": false}, "family": "scale-diverse-219-005", "id": "scale-diverse-219-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Broadcast coordinator", "text": "During the preshow, the operations team opened the archived encoder record associated with the public feed and labeled the reviewed event as incident F."}, {"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output."}, {"speaker": "Content logger", "text": "A frame-by-frame review found that the countdown slate resumed at the expected point; every scheduled item remained available, and no portion of the program was omitted."}, {"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F ending at 2026-08-12T19:42:19.114Z, with no program content lost during F."}, {"speaker": "Broadcast coordinator", "text": "After the event, the public feed continued through the remaining countdown sequence and into the scheduled opening without another freeze."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric, threshold, source preference, output-path requirement, and routing policy; neither context alters those rules. Both contexts concern a video-freeze incident documented by the public encoder during the preshow, preserving the relevant entity, path, and scope bindings. The two evidence spans are complete factual sentences. The counterfactual changes only the recorded end time from 19:42:19.114Z to 19:42:18.114Z, remains later than the unchanged start time, and creates no duplicate or contradictory measurement. Neither context contains a gold label, answer code, rule table, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Broadcast coordinator\",\"text\":\"During the preshow, the operations team opened the archived encoder record associated with the public feed and labeled the reviewed event as incident F.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output.\"},{\"speaker\":\"Content logger\",\"text\":\"A frame-by-frame review found that the countdown slate resumed at the expected point; every scheduled item remained available, and no portion of the program was omitted.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records video-freeze incident F ending at 2026-08-12T19:42:19.114Z, with no program content lost during F.\"},{\"speaker\":\"Broadcast coordinator\",\"text\":\"After the event, the public feed continued through the remaining countdown sequence and into the scheduled opening without another freeze.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output."}, {"path": ["3", "text"], "text": "The encoder log records video-freeze incident F ending at 2026-08-12T19:42:19.114Z, with no program content lost during F."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output.", "negative_left": "The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output.", "negative_right": "The encoder log records video-freeze incident F ending at 2026-08-12T19:42:18.114Z, with no program content lost during F.", "right": "The encoder log records video-freeze incident F ending at 2026-08-12T19:42:19.114Z, with no program content lost during F."}, "verifier_independent_model": false}, "family": "scale-diverse-219-005", "id": "scale-diverse-219-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Broadcast coordinator", "text": "During the preshow, the operations team opened the archived encoder record associated with the public feed and labeled the reviewed event as incident F."}, {"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F beginning at 2026-08-12T19:42:16.730Z during the public preshow countdown and confirms that F appears in the public encoded output."}, {"speaker": "Content logger", "text": "A frame-by-frame review found that the countdown slate resumed at the expected point; every scheduled item remained available, and no portion of the program was omitted."}, {"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F ending at 2026-08-12T19:42:18.114Z, with no program content lost during F."}, {"speaker": "Broadcast coordinator", "text": "After the event, the public feed continued through the remaining countdown sequence and into the scheduled opening without another freeze."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing, verification, intervention, and impact policies without adding exceptions, priorities, or missing-evidence defaults. The Harbor Arts incident and the requested verification/routing/impact/intervention decision remain bound to the same entity and fault path; only case observations change. The two evidence spans are complete factual sentences. The counterfactual coherently changes HM-47 from corroborating the reported timing fault to recording an unrelated battery warning, while retaining exactly two logs and creating no duplicate or contradictory measurement. Neither context states the combined decision, an answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is an allowed quantified proposition over the explicit set of submitted log evidence. The focus a2 is factual rather than policy-based. The base and counter assignments differ only on a2 and are jointly realizable: the same confidence-monitor cue disruption can occur either with or without equivalent submitted log evidence. The policy evidence accurately preserves the substantive rules originating in the original state, while the unchanged questions automatically preserve the decision criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for every required action: a2 supplies equivalent log evidence and therefore verification; a1 triggers streaming-technician routing; verified a2 plus cue disruption a3 requires intervention; and a4 plus a5 directly satisfies the L1 condition. No audience effect also excludes the stated L2 audience-degradation condition and is incompatible with broadcast loss.", "rule_index": 0, "sound": true}, {"reason": "Refuted a2 entails that no submitted log records a fault equivalent to the reported fault. Therefore the report cannot be treated as verified by equivalent log evidence, making at least one required action incorrect and satisfying the false criterion. This does not treat unknown as false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The Harbor Arts livestream incident is a confidence-monitor video timing fault."}, {"id": "a2", "statement": "Among all log evidence submitted for the Harbor Arts livestream incident, at least one log records a fault equivalent to the fault in the stage manager's report."}, {"id": "a3", "statement": "The Harbor Arts livestream incident disrupts the presenter's visual cues."}, {"id": "a4", "statement": "The Harbor Arts livestream incident causes internal disruption."}, {"id": "a5", "statement": "The Harbor Arts livestream incident has no audience effect."}], "base_state_json": "\"Evidence reconciliation for the Harbor Arts livestream incident produced the following findings. The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026. The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day. Operations review classifies the incident as a confidence-monitor video timing fault. The incident caused the presenter to miss visual cues and disrupted internal production work. Audience monitoring found no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026."}, {"path": [], "text": "The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day."}], "policy_evidence": [{"path": [], "text": "Policy routes confidence-monitor video timing faults to the streaming technician."}, {"path": [], "text": "A report is verified when logs restate the same fault in equivalent words."}, {"path": [], "text": "Verified cue disruptions require intervention."}, {"path": [], "text": "Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}], "rules": [{"justification": "Equivalent log evidence verifies the report; the confidence-monitor video timing fault routes to the streaming technician; the verified cue disruption requires intervention; and internal disruption with no audience effect is L1.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "With no submitted log recording a fault equivalent to the reported fault, treating the report as verified is incorrect, so at least one required action is incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026.", "negative_left": "The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026.", "negative_right": "The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records a camera battery warning from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day.", "right": "The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day."}, "verifier_independent_model": false}, "family": "scale-diverse-220-002", "id": "scale-diverse-220-002-base", "input": {"questions": {"decision": {"criteria": {"false": "At least one of these actions is incorrect: verification, streaming-technician routing, L1 impact, or intervention.", "true": "The report is verified by equivalent log evidence; route it to the streaming technician, rate it L1, and intervene."}, "instructions": "Decide whether the incident should be treated as verified, routed to the streaming technician, assigned L1 impact, and marked for intervention.", "type": "noul"}}, "state": "Evidence reconciliation for the Harbor Arts livestream incident produced the following findings. The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026. The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day. Operations review classifies the incident as a confidence-monitor video timing fault. The incident caused the presenter to miss visual cues and disrupted internal production work. Audience monitoring found no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}, "method": "c2d", "provenance": {"source_id": "diverse-220", "source_is_synthetic": true, "source_sha256": "ba35b2ca55f4bdeb927287f576149f873d20519ea2ac4ac62e5acbb8a52c3730", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing, verification, intervention, and impact policies without adding exceptions, priorities, or missing-evidence defaults. The Harbor Arts incident and the requested verification/routing/impact/intervention decision remain bound to the same entity and fault path; only case observations change. The two evidence spans are complete factual sentences. The counterfactual coherently changes HM-47 from corroborating the reported timing fault to recording an unrelated battery warning, while retaining exactly two logs and creating no duplicate or contradictory measurement. Neither context states the combined decision, an answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is an allowed quantified proposition over the explicit set of submitted log evidence. The focus a2 is factual rather than policy-based. The base and counter assignments differ only on a2 and are jointly realizable: the same confidence-monitor cue disruption can occur either with or without equivalent submitted log evidence. The policy evidence accurately preserves the substantive rules originating in the original state, while the unchanged questions automatically preserve the decision criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for every required action: a2 supplies equivalent log evidence and therefore verification; a1 triggers streaming-technician routing; verified a2 plus cue disruption a3 requires intervention; and a4 plus a5 directly satisfies the L1 condition. No audience effect also excludes the stated L2 audience-degradation condition and is incompatible with broadcast loss.", "rule_index": 0, "sound": true}, {"reason": "Refuted a2 entails that no submitted log records a fault equivalent to the reported fault. Therefore the report cannot be treated as verified by equivalent log evidence, making at least one required action incorrect and satisfying the false criterion. This does not treat unknown as false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The Harbor Arts livestream incident is a confidence-monitor video timing fault."}, {"id": "a2", "statement": "Among all log evidence submitted for the Harbor Arts livestream incident, at least one log records a fault equivalent to the fault in the stage manager's report."}, {"id": "a3", "statement": "The Harbor Arts livestream incident disrupts the presenter's visual cues."}, {"id": "a4", "statement": "The Harbor Arts livestream incident causes internal disruption."}, {"id": "a5", "statement": "The Harbor Arts livestream incident has no audience effect."}], "base_state_json": "\"Evidence reconciliation for the Harbor Arts livestream incident produced the following findings. The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026. The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day. Operations review classifies the incident as a confidence-monitor video timing fault. The incident caused the presenter to miss visual cues and disrupted internal production work. Audience monitoring found no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026."}, {"path": [], "text": "The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day."}], "policy_evidence": [{"path": [], "text": "Policy routes confidence-monitor video timing faults to the streaming technician."}, {"path": [], "text": "A report is verified when logs restate the same fault in equivalent words."}, {"path": [], "text": "Verified cue disruptions require intervention."}, {"path": [], "text": "Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}], "rules": [{"justification": "Equivalent log evidence verifies the report; the confidence-monitor video timing fault routes to the streaming technician; the verified cue disruption requires intervention; and internal disruption with no audience effect is L1.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "With no submitted log recording a fault equivalent to the reported fault, treating the report as verified is incorrect, so at least one required action is incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026.", "negative_left": "The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026.", "negative_right": "The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records a camera battery warning from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day.", "right": "The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day."}, "verifier_independent_model": false}, "family": "scale-diverse-220-002", "id": "scale-diverse-220-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "At least one of these actions is incorrect: verification, streaming-technician routing, L1 impact, or intervention.", "true": "The report is verified by equivalent log evidence; route it to the streaming technician, rate it L1, and intervene."}, "instructions": "Decide whether the incident should be treated as verified, routed to the streaming technician, assigned L1 impact, and marked for intervention.", "type": "noul"}}, "state": "Evidence reconciliation for the Harbor Arts livestream incident produced the following findings. The stage manager's report for the Harbor Arts livestream incident identifies the fault as the confidence-monitor video displaying each frame 640 milliseconds after that frame appeared in the program feed from 20:14:08 through 20:14:36 UTC on 12 August 2026. The complete evidence manifest for the Harbor Arts livestream incident lists exactly two submitted logs: log HM-47 records a camera battery warning from 20:14:08 through 20:14:36 UTC on 12 August 2026, and log HM-52 records an audio-console fader movement at 20:19:03 UTC that day. Operations review classifies the incident as a confidence-monitor video timing fault. The incident caused the presenter to miss visual cues and disrupted internal production work. Audience monitoring found no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}, "method": "c2d", "provenance": {"source_id": "diverse-220", "source_is_synthetic": true, "source_sha256": "ba35b2ca55f4bdeb927287f576149f873d20519ea2ac4ac62e5acbb8a52c3730", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing rules for routing, verification, intervention, and impact, while the unchanged questions preserve the requested decision criteria. The Harbor Arts incident and relevant fault/impact scope remain bound to the original question. The two evidence spans are complete factual sentences. The counterfactual coherently changes the sole log entry to an intercom-static fault, so it no longer corroborates the reported confidence-monitor timing fault without contradicting any unchanged fact. Neither context contains an answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is an allowed quantified proposition over the explicit set of submitted log evidence. The focus a2 is factual rather than policy-based. The base and counter assignments differ only on a2 and are jointly realizable: the same confidence-monitor cue disruption can occur either with or without equivalent submitted log evidence. The policy evidence accurately preserves the substantive rules originating in the original state, while the unchanged questions automatically preserve the decision criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for every required action: a2 supplies equivalent log evidence and therefore verification; a1 triggers streaming-technician routing; verified a2 plus cue disruption a3 requires intervention; and a4 plus a5 directly satisfies the L1 condition. No audience effect also excludes the stated L2 audience-degradation condition and is incompatible with broadcast loss.", "rule_index": 0, "sound": true}, {"reason": "Refuted a2 entails that no submitted log records a fault equivalent to the reported fault. Therefore the report cannot be treated as verified by equivalent log evidence, making at least one required action incorrect and satisfying the false criterion. This does not treat unknown as false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The Harbor Arts livestream incident is a confidence-monitor video timing fault."}, {"id": "a2", "statement": "Among all log evidence submitted for the Harbor Arts livestream incident, at least one log records a fault equivalent to the fault in the stage manager's report."}, {"id": "a3", "statement": "The Harbor Arts livestream incident disrupts the presenter's visual cues."}, {"id": "a4", "statement": "The Harbor Arts livestream incident causes internal disruption."}, {"id": "a5", "statement": "The Harbor Arts livestream incident has no audience effect."}], "base_state_json": "\"At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence. Entry 8Q in log HL-304 records that the confidence-monitor video trailed program audio by 2.7 seconds during cue 46 of the Harbor Arts livestream incident. Operations classified the incident as a confidence-monitor video timing fault that disrupted the presenter's visual cues, caused internal disruption, and had no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence."}, {"path": [], "text": "Entry 8Q in log HL-304 records that the confidence-monitor video trailed program audio by 2.7 seconds during cue 46 of the Harbor Arts livestream incident."}], "policy_evidence": [{"path": [], "text": "Policy routes confidence-monitor video timing faults to the streaming technician."}, {"path": [], "text": "A report is verified when logs restate the same fault in equivalent words."}, {"path": [], "text": "Verified cue disruptions require intervention."}, {"path": [], "text": "Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}], "rules": [{"justification": "Equivalent log evidence verifies the report; the confidence-monitor video timing fault routes to the streaming technician; the verified cue disruption requires intervention; and internal disruption with no audience effect is L1.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "With no submitted log recording a fault equivalent to the reported fault, treating the report as verified is incorrect, so at least one required action is incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence.", "negative_left": "At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence.", "negative_right": "Entry 8Q in log HL-304 records that the intercom microphone produced static during cue 46 of the Harbor Arts livestream incident.", "right": "Entry 8Q in log HL-304 records that the confidence-monitor video trailed program audio by 2.7 seconds during cue 46 of the Harbor Arts livestream incident."}, "verifier_independent_model": false}, "family": "scale-diverse-220-003", "id": "scale-diverse-220-003-base", "input": {"questions": {"decision": {"criteria": {"false": "At least one of these actions is incorrect: verification, streaming-technician routing, L1 impact, or intervention.", "true": "The report is verified by equivalent log evidence; route it to the streaming technician, rate it L1, and intervene."}, "instructions": "Decide whether the incident should be treated as verified, routed to the streaming technician, assigned L1 impact, and marked for intervention.", "type": "noul"}}, "state": "At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence. Entry 8Q in log HL-304 records that the confidence-monitor video trailed program audio by 2.7 seconds during cue 46 of the Harbor Arts livestream incident. Operations classified the incident as a confidence-monitor video timing fault that disrupted the presenter's visual cues, caused internal disruption, and had no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}, "method": "c2d", "provenance": {"source_id": "diverse-220", "source_is_synthetic": true, "source_sha256": "ba35b2ca55f4bdeb927287f576149f873d20519ea2ac4ac62e5acbb8a52c3730", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing rules for routing, verification, intervention, and impact, while the unchanged questions preserve the requested decision criteria. The Harbor Arts incident and relevant fault/impact scope remain bound to the original question. The two evidence spans are complete factual sentences. The counterfactual coherently changes the sole log entry to an intercom-static fault, so it no longer corroborates the reported confidence-monitor timing fault without contradicting any unchanged fact. Neither context contains an answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is an allowed quantified proposition over the explicit set of submitted log evidence. The focus a2 is factual rather than policy-based. The base and counter assignments differ only on a2 and are jointly realizable: the same confidence-monitor cue disruption can occur either with or without equivalent submitted log evidence. The policy evidence accurately preserves the substantive rules originating in the original state, while the unchanged questions automatically preserve the decision criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for every required action: a2 supplies equivalent log evidence and therefore verification; a1 triggers streaming-technician routing; verified a2 plus cue disruption a3 requires intervention; and a4 plus a5 directly satisfies the L1 condition. No audience effect also excludes the stated L2 audience-degradation condition and is incompatible with broadcast loss.", "rule_index": 0, "sound": true}, {"reason": "Refuted a2 entails that no submitted log records a fault equivalent to the reported fault. Therefore the report cannot be treated as verified by equivalent log evidence, making at least one required action incorrect and satisfying the false criterion. This does not treat unknown as false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The Harbor Arts livestream incident is a confidence-monitor video timing fault."}, {"id": "a2", "statement": "Among all log evidence submitted for the Harbor Arts livestream incident, at least one log records a fault equivalent to the fault in the stage manager's report."}, {"id": "a3", "statement": "The Harbor Arts livestream incident disrupts the presenter's visual cues."}, {"id": "a4", "statement": "The Harbor Arts livestream incident causes internal disruption."}, {"id": "a5", "statement": "The Harbor Arts livestream incident has no audience effect."}], "base_state_json": "\"At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence. Entry 8Q in log HL-304 records that the confidence-monitor video trailed program audio by 2.7 seconds during cue 46 of the Harbor Arts livestream incident. Operations classified the incident as a confidence-monitor video timing fault that disrupted the presenter's visual cues, caused internal disruption, and had no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence."}, {"path": [], "text": "Entry 8Q in log HL-304 records that the confidence-monitor video trailed program audio by 2.7 seconds during cue 46 of the Harbor Arts livestream incident."}], "policy_evidence": [{"path": [], "text": "Policy routes confidence-monitor video timing faults to the streaming technician."}, {"path": [], "text": "A report is verified when logs restate the same fault in equivalent words."}, {"path": [], "text": "Verified cue disruptions require intervention."}, {"path": [], "text": "Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}], "rules": [{"justification": "Equivalent log evidence verifies the report; the confidence-monitor video timing fault routes to the streaming technician; the verified cue disruption requires intervention; and internal disruption with no audience effect is L1.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "With no submitted log recording a fault equivalent to the reported fault, treating the report as verified is incorrect, so at least one required action is incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence.", "negative_left": "At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence.", "negative_right": "Entry 8Q in log HL-304 records that the intercom microphone produced static during cue 46 of the Harbor Arts livestream incident.", "right": "Entry 8Q in log HL-304 records that the confidence-monitor video trailed program audio by 2.7 seconds during cue 46 of the Harbor Arts livestream incident."}, "verifier_independent_model": false}, "family": "scale-diverse-220-003", "id": "scale-diverse-220-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "At least one of these actions is incorrect: verification, streaming-technician routing, L1 impact, or intervention.", "true": "The report is verified by equivalent log evidence; route it to the streaming technician, rate it L1, and intervene."}, "instructions": "Decide whether the incident should be treated as verified, routed to the streaming technician, assigned L1 impact, and marked for intervention.", "type": "noul"}}, "state": "At the 2026-08-14 22:10 UTC operational handoff for the Harbor Arts livestream incident, the stage manager's report described the confidence-monitor video as trailing program audio by 2.7 seconds during cue 46, and the submission manifest identified entry 8Q in log HL-304 as the sole fault entry in all submitted log evidence. Entry 8Q in log HL-304 records that the intercom microphone produced static during cue 46 of the Harbor Arts livestream incident. Operations classified the incident as a confidence-monitor video timing fault that disrupted the presenter's visual cues, caused internal disruption, and had no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}, "method": "c2d", "provenance": {"source_id": "diverse-220", "source_is_synthetic": true, "source_sha256": "ba35b2ca55f4bdeb927287f576149f873d20519ea2ac4ac62e5acbb8a52c3730", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing routing, verification, intervention, and impact policies from the original state, while the unchanged questions preserve the request and criteria. The Harbor Arts incident and decision scope remain bound consistently; the two evidence spans are complete factual sentences, and the counterfactual coherently changes HL-47 so that no submitted log restates the reported timing fault without introducing a contradictory measurement. Neither context contains a gold label, answer code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is an allowed quantified proposition over the explicit set of submitted log evidence. The focus a2 is factual rather than policy-based. The base and counter assignments differ only on a2 and are jointly realizable: the same confidence-monitor cue disruption can occur either with or without equivalent submitted log evidence. The policy evidence accurately preserves the substantive rules originating in the original state, while the unchanged questions automatically preserve the decision criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for every required action: a2 supplies equivalent log evidence and therefore verification; a1 triggers streaming-technician routing; verified a2 plus cue disruption a3 requires intervention; and a4 plus a5 directly satisfies the L1 condition. No audience effect also excludes the stated L2 audience-degradation condition and is incompatible with broadcast loss.", "rule_index": 0, "sound": true}, {"reason": "Refuted a2 entails that no submitted log records a fault equivalent to the reported fault. Therefore the report cannot be treated as verified by equivalent log evidence, making at least one required action incorrect and satisfying the false criterion. This does not treat unknown as false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The Harbor Arts livestream incident is a confidence-monitor video timing fault."}, {"id": "a2", "statement": "Among all log evidence submitted for the Harbor Arts livestream incident, at least one log records a fault equivalent to the fault in the stage manager's report."}, {"id": "a3", "statement": "The Harbor Arts livestream incident disrupts the presenter's visual cues."}, {"id": "a4", "statement": "The Harbor Arts livestream incident causes internal disruption."}, {"id": "a5", "statement": "The Harbor Arts livestream incident has no audience effect."}], "base_state_json": "\"Harbor Arts field note, 12 June 2026: Technicians classified the livestream incident as a confidence-monitor video timing fault. For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82. For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the confidence-monitor video lagged the program video by 640 milliseconds, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low. The presenter missed private visual cues, disrupting the production team's internal workflow. Audience monitoring confirmed uninterrupted, synchronized program audio and video with no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82."}, {"path": [], "text": "For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the confidence-monitor video lagged the program video by 640 milliseconds, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low."}], "policy_evidence": [{"path": [], "text": "Policy routes confidence-monitor video timing faults to the streaming technician."}, {"path": [], "text": "A report is verified when logs restate the same fault in equivalent words."}, {"path": [], "text": "Verified cue disruptions require intervention."}, {"path": [], "text": "Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}], "rules": [{"justification": "Equivalent log evidence verifies the report; the confidence-monitor video timing fault routes to the streaming technician; the verified cue disruption requires intervention; and internal disruption with no audience effect is L1.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "With no submitted log recording a fault equivalent to the reported fault, treating the report as verified is incorrect, so at least one required action is incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82.", "negative_left": "For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82.", "negative_right": "For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the lobby display lost power, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low.", "right": "For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the confidence-monitor video lagged the program video by 640 milliseconds, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low."}, "verifier_independent_model": false}, "family": "scale-diverse-220-004", "id": "scale-diverse-220-004-base", "input": {"questions": {"decision": {"criteria": {"false": "At least one of these actions is incorrect: verification, streaming-technician routing, L1 impact, or intervention.", "true": "The report is verified by equivalent log evidence; route it to the streaming technician, rate it L1, and intervene."}, "instructions": "Decide whether the incident should be treated as verified, routed to the streaming technician, assigned L1 impact, and marked for intervention.", "type": "noul"}}, "state": "Harbor Arts field note, 12 June 2026: Technicians classified the livestream incident as a confidence-monitor video timing fault. For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82. For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the confidence-monitor video lagged the program video by 640 milliseconds, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low. The presenter missed private visual cues, disrupting the production team's internal workflow. Audience monitoring confirmed uninterrupted, synchronized program audio and video with no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}, "method": "c2d", "provenance": {"source_id": "diverse-220", "source_is_synthetic": true, "source_sha256": "ba35b2ca55f4bdeb927287f576149f873d20519ea2ac4ac62e5acbb8a52c3730", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing routing, verification, intervention, and impact policies from the original state, while the unchanged questions preserve the request and criteria. The Harbor Arts incident and decision scope remain bound consistently; the two evidence spans are complete factual sentences, and the counterfactual coherently changes HL-47 so that no submitted log restates the reported timing fault without introducing a contradictory measurement. Neither context contains a gold label, answer code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is an allowed quantified proposition over the explicit set of submitted log evidence. The focus a2 is factual rather than policy-based. The base and counter assignments differ only on a2 and are jointly realizable: the same confidence-monitor cue disruption can occur either with or without equivalent submitted log evidence. The policy evidence accurately preserves the substantive rules originating in the original state, while the unchanged questions automatically preserve the decision criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for every required action: a2 supplies equivalent log evidence and therefore verification; a1 triggers streaming-technician routing; verified a2 plus cue disruption a3 requires intervention; and a4 plus a5 directly satisfies the L1 condition. No audience effect also excludes the stated L2 audience-degradation condition and is incompatible with broadcast loss.", "rule_index": 0, "sound": true}, {"reason": "Refuted a2 entails that no submitted log records a fault equivalent to the reported fault. Therefore the report cannot be treated as verified by equivalent log evidence, making at least one required action incorrect and satisfying the false criterion. This does not treat unknown as false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The Harbor Arts livestream incident is a confidence-monitor video timing fault."}, {"id": "a2", "statement": "Among all log evidence submitted for the Harbor Arts livestream incident, at least one log records a fault equivalent to the fault in the stage manager's report."}, {"id": "a3", "statement": "The Harbor Arts livestream incident disrupts the presenter's visual cues."}, {"id": "a4", "statement": "The Harbor Arts livestream incident causes internal disruption."}, {"id": "a5", "statement": "The Harbor Arts livestream incident has no audience effect."}], "base_state_json": "\"Harbor Arts field note, 12 June 2026: Technicians classified the livestream incident as a confidence-monitor video timing fault. For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82. For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the confidence-monitor video lagged the program video by 640 milliseconds, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low. The presenter missed private visual cues, disrupting the production team's internal workflow. Audience monitoring confirmed uninterrupted, synchronized program audio and video with no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82."}, {"path": [], "text": "For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the confidence-monitor video lagged the program video by 640 milliseconds, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low."}], "policy_evidence": [{"path": [], "text": "Policy routes confidence-monitor video timing faults to the streaming technician."}, {"path": [], "text": "A report is verified when logs restate the same fault in equivalent words."}, {"path": [], "text": "Verified cue disruptions require intervention."}, {"path": [], "text": "Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}], "rules": [{"justification": "Equivalent log evidence verifies the report; the confidence-monitor video timing fault routes to the streaming technician; the verified cue disruption requires intervention; and internal disruption with no audience effect is L1.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "With no submitted log recording a fault equivalent to the reported fault, treating the report as verified is incorrect, so at least one required action is incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82.", "negative_left": "For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82.", "negative_right": "For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the lobby display lost power, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low.", "right": "For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the confidence-monitor video lagged the program video by 640 milliseconds, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low."}, "verifier_independent_model": false}, "family": "scale-diverse-220-004", "id": "scale-diverse-220-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "At least one of these actions is incorrect: verification, streaming-technician routing, L1 impact, or intervention.", "true": "The report is verified by equivalent log evidence; route it to the streaming technician, rate it L1, and intervene."}, "instructions": "Decide whether the incident should be treated as verified, routed to the streaming technician, assigned L1 impact, and marked for intervention.", "type": "noul"}}, "state": "Harbor Arts field note, 12 June 2026: Technicians classified the livestream incident as a confidence-monitor video timing fault. For the Harbor Arts livestream incident on 12 June 2026, the stage manager's 20:18 UTC report identifies the fault as the confidence-monitor video lagging the program video by 640 milliseconds, and the complete log evidence submitted by 21:00 UTC consists only of logs HL-47 and HL-82. For the Harbor Arts livestream incident on 12 June 2026, the entire fault-record content of HL-47 states that the lobby display lost power, while the entire fault-record content of HL-82 states that the stage-left intercom battery was low. The presenter missed private visual cues, disrupting the production team's internal workflow. Audience monitoring confirmed uninterrupted, synchronized program audio and video with no audience effect. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}, "method": "c2d", "provenance": {"source_id": "diverse-220", "source_is_synthetic": true, "source_sha256": "ba35b2ca55f4bdeb927287f576149f873d20519ea2ac4ac62e5acbb8a52c3730", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing production policy and preserve the museum-panel livestream, program-audio/delivery path, endpoint, and relevant event-time scope. The two focus spans are complete factual sentences. The counterfactual changes only the monitored minimum from 42 to 34 decibels; this is coherent with the 38-decibel threshold, the specified 14:22:04–14:22:05 interval, localization to P-17, and the statement that the stream remained usable. Neither context contains an indexed answer, proposition identifier, rule table, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level.\",\"During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 42 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17.\",\"From 14:20 through 14:24, the encoder log recorded no stream-transport fault, while the recorded source program-audio meter remained at normal speech level.\",\"Exactly one viewer reported audio loss during the scheduled museum-panel livestream.\",\"Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05.\",\"No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the livestream.\",\"The scheduled museum-panel livestream remained usable throughout the event.\",\"No program-feed or stream-delivery impairment occurred during the livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level."}, {"path": ["evidence", "1"], "text": "During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 42 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level.", "negative_left": "Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level.", "negative_right": "During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 34 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17.", "right": "During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 42 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17."}, "verifier_independent_model": false}, "family": "scale-diverse-221-001", "id": "scale-diverse-221-001-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level.", "During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 42 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17.", "From 14:20 through 14:24, the encoder log recorded no stream-transport fault, while the recorded source program-audio meter remained at normal speech level.", "Exactly one viewer reported audio loss during the scheduled museum-panel livestream.", "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05.", "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the livestream.", "The scheduled museum-panel livestream remained usable throughout the event.", "No program-feed or stream-delivery impairment occurred during the livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."]}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing production policy and preserve the museum-panel livestream, program-audio/delivery path, endpoint, and relevant event-time scope. The two focus spans are complete factual sentences. The counterfactual changes only the monitored minimum from 42 to 34 decibels; this is coherent with the 38-decibel threshold, the specified 14:22:04–14:22:05 interval, localization to P-17, and the statement that the stream remained usable. Neither context contains an indexed answer, proposition identifier, rule table, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level.\",\"During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 42 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17.\",\"From 14:20 through 14:24, the encoder log recorded no stream-transport fault, while the recorded source program-audio meter remained at normal speech level.\",\"Exactly one viewer reported audio loss during the scheduled museum-panel livestream.\",\"Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05.\",\"No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the livestream.\",\"The scheduled museum-panel livestream remained usable throughout the event.\",\"No program-feed or stream-delivery impairment occurred during the livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level."}, {"path": ["evidence", "1"], "text": "During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 42 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level.", "negative_left": "Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level.", "negative_right": "During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 34 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17.", "right": "During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 42 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17."}, "verifier_independent_model": false}, "family": "scale-diverse-221-001", "id": "scale-diverse-221-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["Throughout the scheduled museum-panel livestream, audience endpoint P-17's configured audibility threshold was 38 decibels sound-pressure level.", "During the scheduled museum-panel livestream, the calibrated endpoint monitor recorded 34 decibels sound-pressure level as the minimum delivered-program audio level at audience endpoint P-17.", "From 14:20 through 14:24, the encoder log recorded no stream-transport fault, while the recorded source program-audio meter remained at normal speech level.", "Exactly one viewer reported audio loss during the scheduled museum-panel livestream.", "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05.", "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the livestream.", "The scheduled museum-panel livestream remained usable throughout the event.", "No program-feed or stream-delivery impairment occurred during the livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."]}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same governing production policy, decision scope, livestream, endpoint, and relevant timing. The two focus spans are complete factual sentences. Changing P-17’s threshold from 61 to 67 units coherently changes the 64-unit minimum from above threshold to below threshold; the stated 14:22:04–14:22:05 interval, continued usability, and absence of other impairments remain compatible. Neither context includes a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"An incident review reconciled endpoint measurements, transport records, source metering, and the sole viewer report. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels.\",\"Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 61 audibility units on the measurement ledger's scale.\",\"Every delivered-program audio measurement below P-17's configured audibility threshold, if any, was timestamped from 14:22:04 through 14:22:05.\",\"No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the livestream.\",\"The scheduled museum-panel livestream remained usable throughout the event.\",\"The incident review found no program-feed or stream-delivery impairment other than any delivered-program audio below P-17's configured threshold from 14:22:04 through 14:22:05.\",\"Exactly one viewer reported audio loss during the livestream.\",\"The encoder log recorded no stream-transport fault from 14:20 through 14:24.\",\"The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels."}, {"path": ["evidence", "1"], "text": "Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 61 audibility units on the measurement ledger's scale."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels.", "negative_left": "The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels.", "negative_right": "Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 67 audibility units on the measurement ledger's scale.", "right": "Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 61 audibility units on the measurement ledger's scale."}, "verifier_independent_model": false}, "family": "scale-diverse-221-002", "id": "scale-diverse-221-002-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "An incident review reconciled endpoint measurements, transport records, source metering, and the sole viewer report. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels.", "Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 61 audibility units on the measurement ledger's scale.", "Every delivered-program audio measurement below P-17's configured audibility threshold, if any, was timestamped from 14:22:04 through 14:22:05.", "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the livestream.", "The scheduled museum-panel livestream remained usable throughout the event.", "The incident review found no program-feed or stream-delivery impairment other than any delivered-program audio below P-17's configured threshold from 14:22:04 through 14:22:05.", "Exactly one viewer reported audio loss during the livestream.", "The encoder log recorded no stream-transport fault from 14:20 through 14:24.", "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24."]}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same governing production policy, decision scope, livestream, endpoint, and relevant timing. The two focus spans are complete factual sentences. Changing P-17’s threshold from 61 to 67 units coherently changes the 64-unit minimum from above threshold to below threshold; the stated 14:22:04–14:22:05 interval, continued usability, and absence of other impairments remain compatible. Neither context includes a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"An incident review reconciled endpoint measurements, transport records, source metering, and the sole viewer report. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels.\",\"Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 61 audibility units on the measurement ledger's scale.\",\"Every delivered-program audio measurement below P-17's configured audibility threshold, if any, was timestamped from 14:22:04 through 14:22:05.\",\"No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the livestream.\",\"The scheduled museum-panel livestream remained usable throughout the event.\",\"The incident review found no program-feed or stream-delivery impairment other than any delivered-program audio below P-17's configured threshold from 14:22:04 through 14:22:05.\",\"Exactly one viewer reported audio loss during the livestream.\",\"The encoder log recorded no stream-transport fault from 14:20 through 14:24.\",\"The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels."}, {"path": ["evidence", "1"], "text": "Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 61 audibility units on the measurement ledger's scale."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels.", "negative_left": "The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels.", "negative_right": "Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 67 audibility units on the measurement ledger's scale.", "right": "Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 61 audibility units on the measurement ledger's scale."}, "verifier_independent_model": false}, "family": "scale-diverse-221-002", "id": "scale-diverse-221-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "An incident review reconciled endpoint measurements, transport records, source metering, and the sole viewer report. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["The measurement ledger lists 64 audibility units as the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream, on a scale where larger values represent higher audio levels.", "Audience endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was 67 audibility units on the measurement ledger's scale.", "Every delivered-program audio measurement below P-17's configured audibility threshold, if any, was timestamped from 14:22:04 through 14:22:05.", "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the livestream.", "The scheduled museum-panel livestream remained usable throughout the event.", "The incident review found no program-feed or stream-delivery impairment other than any delivered-program audio below P-17's configured threshold from 14:22:04 through 14:22:05.", "Exactly one viewer reported audio loss during the livestream.", "The encoder log recorded no stream-transport fault from 14:20 through 14:24.", "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24."]}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the questions object verbatim and preserve the operational roles and governing verification policy from the original state. The request, scheduled museum-panel livestream, relevant endpoint, and event-time scope remain properly bound. The two focus-evidence spans are complete factual sentences. Changing P-17’s threshold from 44 to 50 units coherently changes the relationship to the unchanged 47-unit minimum: the base permits no below-threshold measurement, while the counterfactual supports a brief localized interval, without conflicting with the stream remaining usable, the absence of other impairment, or the normal source and transport records. Neither context embeds a selected score, answer code, label rationale, proposition ID, or output instruction beyond the unchanged question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.\",\"1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.\",\"2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change.\"],\"instructions\":\"Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.\",\"type\":\"score\"}},\"state\":{\"context\":\"Operational handoff for the scheduled museum-panel livestream: The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units.\",\"The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 44 units on the event's linear measurement scale.\",\"Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the livestream was timestamped from 14:22:04 through 14:22:05.\",\"No other audience endpoint had a delivered-program audio measurement below its configured audibility threshold during the livestream.\",\"The scheduled museum-panel livestream remained usable throughout the event.\",\"No program-feed or stream-delivery impairment occurred other than any delivered-program audio below P-17's configured audibility threshold from 14:22:04 through 14:22:05.\",\"Exactly one viewer reported audio loss during the livestream.\",\"The encoder log recorded no stream-transport fault from 14:20 through 14:24.\",\"The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24.\"],\"request\":\"Choose the level that best reflects verified broadcast impact, required routing, and whether immediate intervention is warranted.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units."}, {"path": ["state", "evidence", "1"], "text": "The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 44 units on the event's linear measurement scale."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units.", "negative_left": "The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units.", "negative_right": "The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 50 units on the event's linear measurement scale.", "right": "The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 44 units on the event's linear measurement scale."}, "verifier_independent_model": false}, "family": "scale-diverse-221-003", "id": "scale-diverse-221-003-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "Operational handoff for the scheduled museum-panel livestream: The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units.", "The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 44 units on the event's linear measurement scale.", "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the livestream was timestamped from 14:22:04 through 14:22:05.", "No other audience endpoint had a delivered-program audio measurement below its configured audibility threshold during the livestream.", "The scheduled museum-panel livestream remained usable throughout the event.", "No program-feed or stream-delivery impairment occurred other than any delivered-program audio below P-17's configured audibility threshold from 14:22:04 through 14:22:05.", "Exactly one viewer reported audio loss during the livestream.", "The encoder log recorded no stream-transport fault from 14:20 through 14:24.", "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24."], "request": "Choose the level that best reflects verified broadcast impact, required routing, and whether immediate intervention is warranted."}}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the questions object verbatim and preserve the operational roles and governing verification policy from the original state. The request, scheduled museum-panel livestream, relevant endpoint, and event-time scope remain properly bound. The two focus-evidence spans are complete factual sentences. Changing P-17’s threshold from 44 to 50 units coherently changes the relationship to the unchanged 47-unit minimum: the base permits no below-threshold measurement, while the counterfactual supports a brief localized interval, without conflicting with the stream remaining usable, the absence of other impairment, or the normal source and transport records. Neither context embeds a selected score, answer code, label rationale, proposition ID, or output instruction beyond the unchanged question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.\",\"1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.\",\"2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change.\"],\"instructions\":\"Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.\",\"type\":\"score\"}},\"state\":{\"context\":\"Operational handoff for the scheduled museum-panel livestream: The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units.\",\"The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 44 units on the event's linear measurement scale.\",\"Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the livestream was timestamped from 14:22:04 through 14:22:05.\",\"No other audience endpoint had a delivered-program audio measurement below its configured audibility threshold during the livestream.\",\"The scheduled museum-panel livestream remained usable throughout the event.\",\"No program-feed or stream-delivery impairment occurred other than any delivered-program audio below P-17's configured audibility threshold from 14:22:04 through 14:22:05.\",\"Exactly one viewer reported audio loss during the livestream.\",\"The encoder log recorded no stream-transport fault from 14:20 through 14:24.\",\"The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24.\"],\"request\":\"Choose the level that best reflects verified broadcast impact, required routing, and whether immediate intervention is warranted.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["state", "evidence", "0"], "text": "The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units."}, {"path": ["state", "evidence", "1"], "text": "The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 44 units on the event's linear measurement scale."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units.", "negative_left": "The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units.", "negative_right": "The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 50 units on the event's linear measurement scale.", "right": "The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 44 units on the event's linear measurement scale."}, "verifier_independent_model": false}, "family": "scale-diverse-221-003", "id": "scale-diverse-221-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "Operational handoff for the scheduled museum-panel livestream: The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["The final operational handoff record states that, on the event's linear measurement scale where larger values represent higher audio levels, the minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was 47 units.", "The configuration snapshot retained for the scheduled museum-panel livestream records audience endpoint P-17's configured audibility threshold as 50 units on the event's linear measurement scale.", "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the livestream was timestamped from 14:22:04 through 14:22:05.", "No other audience endpoint had a delivered-program audio measurement below its configured audibility threshold during the livestream.", "The scheduled museum-panel livestream remained usable throughout the event.", "No program-feed or stream-delivery impairment occurred other than any delivered-program audio below P-17's configured audibility threshold from 14:22:04 through 14:22:05.", "Exactly one viewer reported audio loss during the livestream.", "The encoder log recorded no stream-transport fault from 14:20 through 14:24.", "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24."], "request": "Choose the level that best reflects verified broadcast impact, required routing, and whether immediate intervention is warranted."}}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the unchanged questions and the original governing role, verification, routing, and intervention policy. The event, endpoint, measurement type, and relevant time bindings remain fixed; the counterfactual changes only P-17’s lowest measurement from −34 dBFS to −41 dBFS, coherently creating a brief localized threshold crossing without conflicting with stream usability or the stated impairment exception. The two focus spans are complete factual sentences, and neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output-directing instruction beyond the legitimate request and natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS.\",\"The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −34 dBFS.\",\"Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05.\",\"No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream.\",\"The scheduled museum-panel livestream remained usable throughout the event.\",\"No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05.\",\"Exactly one viewer reported audio loss during the scheduled museum-panel livestream.\",\"The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream.\",\"The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream.\"],\"request\":\"Choose the level that best reflects verified broadcast impact, required routing, and whether immediate intervention is warranted.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS."}, {"path": ["evidence", "1"], "text": "The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −34 dBFS."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS.", "negative_left": "Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS.", "negative_right": "The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −41 dBFS.", "right": "The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −34 dBFS."}, "verifier_independent_model": false}, "family": "scale-diverse-221-004", "id": "scale-diverse-221-004-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS.", "The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −34 dBFS.", "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05.", "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream.", "The scheduled museum-panel livestream remained usable throughout the event.", "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05.", "Exactly one viewer reported audio loss during the scheduled museum-panel livestream.", "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream.", "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."], "request": "Choose the level that best reflects verified broadcast impact, required routing, and whether immediate intervention is warranted."}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs preserve the unchanged questions and the original governing role, verification, routing, and intervention policy. The event, endpoint, measurement type, and relevant time bindings remain fixed; the counterfactual changes only P-17’s lowest measurement from −34 dBFS to −41 dBFS, coherently creating a brief localized threshold crossing without conflicting with stream usability or the stated impairment exception. The two focus spans are complete factual sentences, and neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output-directing instruction beyond the legitimate request and natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS.\",\"The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −34 dBFS.\",\"Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05.\",\"No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream.\",\"The scheduled museum-panel livestream remained usable throughout the event.\",\"No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05.\",\"Exactly one viewer reported audio loss during the scheduled museum-panel livestream.\",\"The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream.\",\"The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream.\"],\"request\":\"Choose the level that best reflects verified broadcast impact, required routing, and whether immediate intervention is warranted.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS."}, {"path": ["evidence", "1"], "text": "The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −34 dBFS."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS.", "negative_left": "Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS.", "negative_right": "The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −41 dBFS.", "right": "The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −34 dBFS."}, "verifier_independent_model": false}, "family": "scale-diverse-221-004", "id": "scale-diverse-221-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["Audience endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was −36 dBFS.", "The lowest delivered-program audio measurement recorded at audience endpoint P-17 during the scheduled museum-panel livestream was −41 dBFS.", "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05.", "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream.", "The scheduled museum-panel livestream remained usable throughout the event.", "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05.", "Exactly one viewer reported audio loss during the scheduled museum-panel livestream.", "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream.", "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."], "request": "Choose the level that best reflects verified broadcast impact, required routing, and whether immediate intervention is warranted."}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing routing rubric and retain the same decision-maker, requested asset, transfer, and routing question. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes Q-71 from being the same asset as HP-3 to being distinct from HP-3; together with the unchanged transfer-membership and scope-note facts, this creates no contradictory duplicate measurement or assertion. Neither constructed context contains an explicit answer, output instruction, rule table, proposition ID, answer code, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71.\",\"At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 and the asset assigned Harbor Echo transfer identifier HP-3 as the same asset.\",\"At 09:31, the registrar confirmed that the requested file belongs to the Harbor Echo transfer. The technical inspection at 09:43 classified HP-3 as a project file that remains editable and links source media; it is neither raw captured sound nor a published final export. At 09:55, the transfer ledger’s scope note recorded that every Harbor Echo asset other than HP-3 is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71."}, {"path": ["evidence", "1"], "text": "At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 and the asset assigned Harbor Echo transfer identifier HP-3 as the same asset."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71.", "negative_left": "At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71.", "negative_right": "At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 as distinct from the asset assigned Harbor Echo transfer identifier HP-3.", "right": "At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 and the asset assigned Harbor Echo transfer identifier HP-3 as the same asset."}, "verifier_independent_model": false}, "family": "scale-diverse-223-001", "id": "scale-diverse-223-001-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71.", "At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 and the asset assigned Harbor Echo transfer identifier HP-3 as the same asset.", "At 09:31, the registrar confirmed that the requested file belongs to the Harbor Echo transfer. The technical inspection at 09:43 classified HP-3 as a project file that remains editable and links source media; it is neither raw captured sound nor a published final export. At 09:55, the transfer ledger’s scope note recorded that every Harbor Echo asset other than HP-3 is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the governing routing rubric and retain the same decision-maker, requested asset, transfer, and routing question. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes Q-71 from being the same asset as HP-3 to being distinct from HP-3; together with the unchanged transfer-membership and scope-note facts, this creates no contradictory duplicate measurement or assertion. Neither constructed context contains an explicit answer, output instruction, rule table, proposition ID, answer code, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71.\",\"At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 and the asset assigned Harbor Echo transfer identifier HP-3 as the same asset.\",\"At 09:31, the registrar confirmed that the requested file belongs to the Harbor Echo transfer. The technical inspection at 09:43 classified HP-3 as a project file that remains editable and links source media; it is neither raw captured sound nor a published final export. At 09:55, the transfer ledger’s scope note recorded that every Harbor Echo asset other than HP-3 is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71."}, {"path": ["evidence", "1"], "text": "At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 and the asset assigned Harbor Echo transfer identifier HP-3 as the same asset."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71.", "negative_left": "At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71.", "negative_right": "At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 as distinct from the asset assigned Harbor Echo transfer identifier HP-3.", "right": "At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 and the asset assigned Harbor Echo transfer identifier HP-3 as the same asset."}, "verifier_independent_model": false}, "family": "scale-diverse-223-001", "id": "scale-diverse-223-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["At 09:14 on 18 August 2026, the Harbor Echo intake log recorded the asset requested as harbor_mix.prproj under temporary tag Q-71.", "At 09:26 on 18 August 2026, the Harbor Echo transfer ledger recorded the asset under temporary tag Q-71 as distinct from the asset assigned Harbor Echo transfer identifier HP-3.", "At 09:31, the registrar confirmed that the requested file belongs to the Harbor Echo transfer. The technical inspection at 09:43 classified HP-3 as a project file that remains editable and links source media; it is neither raw captured sound nor a published final export. At 09:55, the transfer ledger’s scope note recorded that every Harbor Echo asset other than HP-3 is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing rubric, request, requested asset, transfer, and manifest timestamp. The two focus spans are complete factual sentences. Changing harbor_mix.prproj’s unique inventory key from Q7-184 to M2-639 coherently separates it from HP-3 under the manifest’s one-to-one key assignment; the reconciliation evidence then consistently characterizes it as a non-HP-3 asset. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key Q7-184.\",\"In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184.\",\"Technical inspection records HP-3 as an Adobe project file whose editable timeline links source media; the examiner explicitly ruled out both raw captured sound and a published final export.\",\"The request register confirms that harbor_mix.prproj belongs to the Harbor Echo transfer.\",\"The reconciliation log states that every Harbor Echo asset other than HP-3 is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key Q7-184."}, {"path": ["evidence", "1"], "text": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key Q7-184.", "negative_left": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key M2-639.", "negative_right": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184.", "right": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184."}, "verifier_independent_model": false}, "family": "scale-diverse-223-002", "id": "scale-diverse-223-002-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key Q7-184.", "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184.", "Technical inspection records HP-3 as an Adobe project file whose editable timeline links source media; the examiner explicitly ruled out both raw captured sound and a published final export.", "The request register confirms that harbor_mix.prproj belongs to the Harbor Echo transfer.", "The reconciliation log states that every Harbor Echo asset other than HP-3 is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing rubric, request, requested asset, transfer, and manifest timestamp. The two focus spans are complete factual sentences. Changing harbor_mix.prproj’s unique inventory key from Q7-184 to M2-639 coherently separates it from HP-3 under the manifest’s one-to-one key assignment; the reconciliation evidence then consistently characterizes it as a non-HP-3 asset. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key Q7-184.\",\"In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184.\",\"Technical inspection records HP-3 as an Adobe project file whose editable timeline links source media; the examiner explicitly ruled out both raw captured sound and a published final export.\",\"The request register confirms that harbor_mix.prproj belongs to the Harbor Echo transfer.\",\"The reconciliation log states that every Harbor Echo asset other than HP-3 is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key Q7-184."}, {"path": ["evidence", "1"], "text": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key Q7-184.", "negative_left": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key M2-639.", "negative_right": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184.", "right": "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184."}, "verifier_independent_model": false}, "family": "scale-diverse-223-002", "id": "scale-diverse-223-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, the asset requested as harbor_mix.prproj carries inventory key M2-639.", "In the Harbor Echo manifest frozen at 2026-08-14T16:20:00Z, which assigns each asset exactly one inventory key and each inventory key to exactly one asset, the asset assigned transfer identifier HP-3 carries inventory key Q7-184.", "Technical inspection records HP-3 as an Adobe project file whose editable timeline links source media; the examiner explicitly ruled out both raw captured sound and a published final export.", "The request register confirms that harbor_mix.prproj belongs to the Harbor Echo transfer.", "The reconciliation log states that every Harbor Echo asset other than HP-3 is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing rubric, request, asset identity, and relevant transfer/time framing, while the unchanged questions preserve the full decision criteria and instructions. The two focus spans are complete factual sentences. The counterfactual coherently changes the compartment-to-HP-3 link: harbor_mix.prproj remains the sole asset in compartment 41, while HP-3 is stated not to be there, so the unchanged evidence can consistently treat harbor_mix.prproj as a non-HP-3 Harbor Echo asset. Neither context includes an answer code, explicit classifier-output instruction, proposition identifier, rule table, or gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41.\",\"At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records the sole asset in transfer compartment 41 as the asset assigned Harbor Echo transfer identifier HP-3.\",\"The transfer registry classifies HP-3 as an editable project file that links source media. It expressly identifies the asset as neither raw captured sound nor a published final export.\",\"The transfer register confirms that the asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer.\",\"A completion audit states that every Harbor Echo asset other than HP-3 is a published final export; none of those other assets is raw captured sound or editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41."}, {"path": ["evidence", "1"], "text": "At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records the sole asset in transfer compartment 41 as the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41.", "negative_left": "At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41.", "negative_right": "At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records that the asset assigned Harbor Echo transfer identifier HP-3 was not placed in transfer compartment 41.", "right": "At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records the sole asset in transfer compartment 41 as the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "scale-diverse-223-007", "id": "scale-diverse-223-007-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41.", "At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records the sole asset in transfer compartment 41 as the asset assigned Harbor Echo transfer identifier HP-3.", "The transfer registry classifies HP-3 as an editable project file that links source media. It expressly identifies the asset as neither raw captured sound nor a published final export.", "The transfer register confirms that the asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer.", "A completion audit states that every Harbor Echo asset other than HP-3 is a published final export; none of those other assets is raw captured sound or editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing rubric, request, asset identity, and relevant transfer/time framing, while the unchanged questions preserve the full decision criteria and instructions. The two focus spans are complete factual sentences. The counterfactual coherently changes the compartment-to-HP-3 link: harbor_mix.prproj remains the sole asset in compartment 41, while HP-3 is stated not to be there, so the unchanged evidence can consistently treat harbor_mix.prproj as a non-HP-3 Harbor Echo asset. Neither context includes an answer code, explicit classifier-output instruction, proposition identifier, rule table, or gold-label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41.\",\"At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records the sole asset in transfer compartment 41 as the asset assigned Harbor Echo transfer identifier HP-3.\",\"The transfer registry classifies HP-3 as an editable project file that links source media. It expressly identifies the asset as neither raw captured sound nor a published final export.\",\"The transfer register confirms that the asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer.\",\"A completion audit states that every Harbor Echo asset other than HP-3 is a published final export; none of those other assets is raw captured sound or editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41."}, {"path": ["evidence", "1"], "text": "At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records the sole asset in transfer compartment 41 as the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41.", "negative_left": "At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41.", "negative_right": "At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records that the asset assigned Harbor Echo transfer identifier HP-3 was not placed in transfer compartment 41.", "right": "At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records the sole asset in transfer compartment 41 as the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "scale-diverse-223-007", "id": "scale-diverse-223-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["At 14:20 UTC on 8 July 2026, the Harbor Echo handoff log records the asset requested as harbor_mix.prproj as the sole asset placed in transfer compartment 41.", "At 14:20 UTC on 8 July 2026, the Harbor Echo custody manifest records that the asset assigned Harbor Echo transfer identifier HP-3 was not placed in transfer compartment 41.", "The transfer registry classifies HP-3 as an editable project file that links source media. It expressly identifies the asset as neither raw captured sound nor a published final export.", "The transfer register confirms that the asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer.", "A completion audit states that every Harbor Echo asset other than HP-3 is a published final export; none of those other assets is raw captured sound or editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full metadata-completeness policy, request scope, and the “Harbor Bells” record binding. The two evidence spans are complete factual sentences. In the base context, six explicitly present fields plus an empty caption are consistent with seven total populated fields because rights must also be populated; in the counterfactual, they are consistent with six total populated fields because rights is then unpopulated. Neither context states the resulting level or embeds an answer code, rationale, proposition ID, or output-directing exception.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "full_context_fact_states": {"base": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "counterfactual": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "remove_left": {"rights_present": "unknown"}, "remove_right": {"rights_present": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"rights_present": "unknown"}, "negative_pair": {"rights_present": "refuted"}, "negative_sentence": {"rights_present": "unknown"}, "positive_pair": {"rights_present": "supported"}, "right": {"rights_present": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one metadata-presence relationship. The focus atom, rights presence, is factual rather than a policy classification. The base and counter assignments differ only in rights presence and are both realizable: the base yields all required fields plus one optional field, while the counter yields exactly one missing required field. The cited policy evidence is a valid citation from the original state and accurately preserves the metadata-field bindings and level rules, although the unchanged question already contains the governing criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that exactly one required field—rights—is missing while the other five required fields are present. Level 3 applies regardless of optional-field status, so the target is entailed.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all six required fields are present and exactly one optional field, subject, is present while caption is absent. This is sufficient for level 4 and excludes level 5.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "title_present", "statement": "The supplied record for “Harbor Bells” has title metadata present."}, {"id": "creator_present", "statement": "The supplied record for “Harbor Bells” has creator metadata present."}, {"id": "date_present", "statement": "The supplied record for “Harbor Bells” has date metadata present."}, {"id": "media_type_present", "statement": "The supplied record for “Harbor Bells” has media type metadata present."}, {"id": "duration_present", "statement": "The supplied record for “Harbor Bells” has duration metadata present."}, {"id": "rights_present", "statement": "The supplied record for “Harbor Bells” has rights metadata present."}, {"id": "subject_present", "statement": "The supplied record for “Harbor Bells” has subject metadata present."}, {"id": "caption_present", "statement": "The supplied record for “Harbor Bells” has caption metadata present."}], "base_state_json": "[{\"speaker\":\"Reconciliation auditor\",\"text\":\"At 09:40 UTC on 14 August 2026, exactly seven of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values.\"},{\"speaker\":\"Record examiner\",\"text\":\"In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing.\"},{\"speaker\":\"Media librarian\",\"text\":\"Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata.\"}]", "base_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "counter_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "focus_atom": "rights_present", "focus_evidence": [{"path": ["0", "text"], "text": "At 09:40 UTC on 14 August 2026, exactly seven of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values."}, {"path": ["1", "text"], "text": "In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty."}], "policy_evidence": [{"path": ["1", "text"], "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}], "rules": [{"justification": "Rights is the only missing required field; the policy assigns level_3 when exactly one of the six required fields is missing, regardless of optional fields.", "target": "level_3", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}, {"justification": "All six required fields are present and exactly one optional field, subject, is present; the policy assigns level_4.", "target": "level_4", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:40 UTC on 14 August 2026, exactly seven of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values.", "negative_left": "At 09:40 UTC on 14 August 2026, exactly six of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values.", "negative_right": "In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty.", "right": "In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty."}, "verifier_independent_model": false}, "family": "scale-diverse-224-002", "id": "scale-diverse-224-002-base", "input": {"questions": {"decision": {"criteria": {"level_1": "Four or more of the six required metadata fields are missing.", "level_2": "Two or three of the six required metadata fields are missing.", "level_3": "Exactly one of the six required metadata fields is missing, regardless of optional fields.", "level_4": "All six required metadata fields are present, with zero or exactly one optional field present.", "level_5": "All six required metadata fields and both optional fields are present."}, "instructions": "Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata.", "type": "choice"}}, "state": [{"speaker": "Reconciliation auditor", "text": "At 09:40 UTC on 14 August 2026, exactly seven of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values."}, {"speaker": "Record examiner", "text": "In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty."}, {"speaker": "Metadata specialist", "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}, {"speaker": "Media librarian", "text": "Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata."}]}, "method": "c2d", "provenance": {"source_id": "diverse-224", "source_is_synthetic": true, "source_sha256": "42be2105c51ead25ff90f2762bb2402ddf400ae28dfdeec25ee9b7154cc0ab7b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_4"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full metadata-completeness policy, request scope, and the “Harbor Bells” record binding. The two evidence spans are complete factual sentences. In the base context, six explicitly present fields plus an empty caption are consistent with seven total populated fields because rights must also be populated; in the counterfactual, they are consistent with six total populated fields because rights is then unpopulated. Neither context states the resulting level or embeds an answer code, rationale, proposition ID, or output-directing exception.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "full_context_fact_states": {"base": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "counterfactual": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "remove_left": {"rights_present": "unknown"}, "remove_right": {"rights_present": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"rights_present": "unknown"}, "negative_pair": {"rights_present": "refuted"}, "negative_sentence": {"rights_present": "unknown"}, "positive_pair": {"rights_present": "supported"}, "right": {"rights_present": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one metadata-presence relationship. The focus atom, rights presence, is factual rather than a policy classification. The base and counter assignments differ only in rights presence and are both realizable: the base yields all required fields plus one optional field, while the counter yields exactly one missing required field. The cited policy evidence is a valid citation from the original state and accurately preserves the metadata-field bindings and level rules, although the unchanged question already contains the governing criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that exactly one required field—rights—is missing while the other five required fields are present. Level 3 applies regardless of optional-field status, so the target is entailed.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all six required fields are present and exactly one optional field, subject, is present while caption is absent. This is sufficient for level 4 and excludes level 5.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "title_present", "statement": "The supplied record for “Harbor Bells” has title metadata present."}, {"id": "creator_present", "statement": "The supplied record for “Harbor Bells” has creator metadata present."}, {"id": "date_present", "statement": "The supplied record for “Harbor Bells” has date metadata present."}, {"id": "media_type_present", "statement": "The supplied record for “Harbor Bells” has media type metadata present."}, {"id": "duration_present", "statement": "The supplied record for “Harbor Bells” has duration metadata present."}, {"id": "rights_present", "statement": "The supplied record for “Harbor Bells” has rights metadata present."}, {"id": "subject_present", "statement": "The supplied record for “Harbor Bells” has subject metadata present."}, {"id": "caption_present", "statement": "The supplied record for “Harbor Bells” has caption metadata present."}], "base_state_json": "[{\"speaker\":\"Reconciliation auditor\",\"text\":\"At 09:40 UTC on 14 August 2026, exactly seven of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values.\"},{\"speaker\":\"Record examiner\",\"text\":\"In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing.\"},{\"speaker\":\"Media librarian\",\"text\":\"Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata.\"}]", "base_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "counter_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "focus_atom": "rights_present", "focus_evidence": [{"path": ["0", "text"], "text": "At 09:40 UTC on 14 August 2026, exactly seven of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values."}, {"path": ["1", "text"], "text": "In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty."}], "policy_evidence": [{"path": ["1", "text"], "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}], "rules": [{"justification": "Rights is the only missing required field; the policy assigns level_3 when exactly one of the six required fields is missing, regardless of optional fields.", "target": "level_3", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}, {"justification": "All six required fields are present and exactly one optional field, subject, is present; the policy assigns level_4.", "target": "level_4", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:40 UTC on 14 August 2026, exactly seven of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values.", "negative_left": "At 09:40 UTC on 14 August 2026, exactly six of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values.", "negative_right": "In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty.", "right": "In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty."}, "verifier_independent_model": false}, "family": "scale-diverse-224-002", "id": "scale-diverse-224-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1": "Four or more of the six required metadata fields are missing.", "level_2": "Two or three of the six required metadata fields are missing.", "level_3": "Exactly one of the six required metadata fields is missing, regardless of optional fields.", "level_4": "All six required metadata fields are present, with zero or exactly one optional field present.", "level_5": "All six required metadata fields and both optional fields are present."}, "instructions": "Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata.", "type": "choice"}}, "state": [{"speaker": "Reconciliation auditor", "text": "At 09:40 UTC on 14 August 2026, exactly six of the supplied record for “Harbor Bells”’s eight metadata fields—title, creator, date, media type, duration, rights, subject, and caption—contained values."}, {"speaker": "Record examiner", "text": "In that supplied record at 09:40 UTC on 14 August 2026, the title, creator, date, media type, duration, and subject fields contained values, while the caption field was empty."}, {"speaker": "Metadata specialist", "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}, {"speaker": "Media librarian", "text": "Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata."}]}, "method": "c2d", "provenance": {"source_id": "diverse-224", "source_is_synthetic": true, "source_sha256": "42be2105c51ead25ff90f2762bb2402ddf400ae28dfdeec25ee9b7154cc0ab7b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_3"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing completeness policy, while the unchanged questions preserve the decision criteria and scope. The same record, rights-storage path, and repository snapshot time are maintained; the two evidence spans are complete factual sentences. The counterfactual coherently changes RB-731 from containing the rights value to being empty without conflicting with another counterfactual assertion. Neither context states the applicable level or otherwise embeds a gold answer; the scoring language is the governing policy expressed in natural terms.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "full_context_fact_states": {"base": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "counterfactual": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "remove_left": {"rights_present": "unknown"}, "remove_right": {"rights_present": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"rights_present": "unknown"}, "negative_pair": {"rights_present": "refuted"}, "negative_sentence": {"rights_present": "unknown"}, "positive_pair": {"rights_present": "supported"}, "right": {"rights_present": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one metadata-presence relationship. The focus atom, rights presence, is factual rather than a policy classification. The base and counter assignments differ only in rights presence and are both realizable: the base yields all required fields plus one optional field, while the counter yields exactly one missing required field. The cited policy evidence is a valid citation from the original state and accurately preserves the metadata-field bindings and level rules, although the unchanged question already contains the governing criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that exactly one required field—rights—is missing while the other five required fields are present. Level 3 applies regardless of optional-field status, so the target is entailed.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all six required fields are present and exactly one optional field, subject, is present while caption is absent. This is sufficient for level 4 and excludes level 5.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "title_present", "statement": "The supplied record for “Harbor Bells” has title metadata present."}, {"id": "creator_present", "statement": "The supplied record for “Harbor Bells” has creator metadata present."}, {"id": "date_present", "statement": "The supplied record for “Harbor Bells” has date metadata present."}, {"id": "media_type_present", "statement": "The supplied record for “Harbor Bells” has media type metadata present."}, {"id": "duration_present", "statement": "The supplied record for “Harbor Bells” has duration metadata present."}, {"id": "rights_present", "statement": "The supplied record for “Harbor Bells” has rights metadata present."}, {"id": "subject_present", "statement": "The supplied record for “Harbor Bells” has subject metadata present."}, {"id": "caption_present", "statement": "The supplied record for “Harbor Bells” has caption metadata present."}], "base_state_json": "[{\"speaker\":\"Catalog technician\",\"text\":\"Operational handoff: inspection of the supplied record confirms title “Harbor Bells,” creator Nia Voss, date 2031-06-14, media type WAV, and duration 02:00. The subject field contains “coastal sound.” The caption field is absent.\"},{\"speaker\":\"Repository technician\",\"text\":\"In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value.\"},{\"speaker\":\"Repository technician\",\"text\":\"In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 contains the value “Copyright 2024 Marina Sound Archive; CC BY 4.0.”\"},{\"speaker\":\"Metadata specialist\",\"text\":\"Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing.\"}]", "base_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "counter_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "focus_atom": "rights_present", "focus_evidence": [{"path": ["1", "text"], "text": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value."}, {"path": ["2", "text"], "text": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 contains the value “Copyright 2024 Marina Sound Archive; CC BY 4.0.”"}], "policy_evidence": [{"path": ["1", "text"], "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}], "rules": [{"justification": "Rights is the only missing required field; the policy assigns level_3 when exactly one of the six required fields is missing, regardless of optional fields.", "target": "level_3", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}, {"justification": "All six required fields are present and exactly one optional field, subject, is present; the policy assigns level_4.", "target": "level_4", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}]}, "verified_pair": {"left": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value.", "negative_left": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value.", "negative_right": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is empty and contains no value.", "right": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 contains the value “Copyright 2024 Marina Sound Archive; CC BY 4.0.”"}, "verifier_independent_model": false}, "family": "scale-diverse-224-007", "id": "scale-diverse-224-007-base", "input": {"questions": {"decision": {"criteria": {"level_1": "Four or more of the six required metadata fields are missing.", "level_2": "Two or three of the six required metadata fields are missing.", "level_3": "Exactly one of the six required metadata fields is missing, regardless of optional fields.", "level_4": "All six required metadata fields are present, with zero or exactly one optional field present.", "level_5": "All six required metadata fields and both optional fields are present."}, "instructions": "Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata.", "type": "choice"}}, "state": [{"speaker": "Catalog technician", "text": "Operational handoff: inspection of the supplied record confirms title “Harbor Bells,” creator Nia Voss, date 2031-06-14, media type WAV, and duration 02:00. The subject field contains “coastal sound.” The caption field is absent."}, {"speaker": "Repository technician", "text": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value."}, {"speaker": "Repository technician", "text": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 contains the value “Copyright 2024 Marina Sound Archive; CC BY 4.0.”"}, {"speaker": "Metadata specialist", "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}]}, "method": "c2d", "provenance": {"source_id": "diverse-224", "source_is_synthetic": true, "source_sha256": "42be2105c51ead25ff90f2762bb2402ddf400ae28dfdeec25ee9b7154cc0ab7b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_4"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing completeness policy, while the unchanged questions preserve the decision criteria and scope. The same record, rights-storage path, and repository snapshot time are maintained; the two evidence spans are complete factual sentences. The counterfactual coherently changes RB-731 from containing the rights value to being empty without conflicting with another counterfactual assertion. Neither context states the applicable level or otherwise embeds a gold answer; the scoring language is the governing policy expressed in natural terms.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "full_context_fact_states": {"base": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "counterfactual": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "remove_left": {"rights_present": "unknown"}, "remove_right": {"rights_present": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"rights_present": "unknown"}, "negative_pair": {"rights_present": "refuted"}, "negative_sentence": {"rights_present": "unknown"}, "positive_pair": {"rights_present": "supported"}, "right": {"rights_present": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one metadata-presence relationship. The focus atom, rights presence, is factual rather than a policy classification. The base and counter assignments differ only in rights presence and are both realizable: the base yields all required fields plus one optional field, while the counter yields exactly one missing required field. The cited policy evidence is a valid citation from the original state and accurately preserves the metadata-field bindings and level rules, although the unchanged question already contains the governing criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that exactly one required field—rights—is missing while the other five required fields are present. Level 3 applies regardless of optional-field status, so the target is entailed.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all six required fields are present and exactly one optional field, subject, is present while caption is absent. This is sufficient for level 4 and excludes level 5.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "title_present", "statement": "The supplied record for “Harbor Bells” has title metadata present."}, {"id": "creator_present", "statement": "The supplied record for “Harbor Bells” has creator metadata present."}, {"id": "date_present", "statement": "The supplied record for “Harbor Bells” has date metadata present."}, {"id": "media_type_present", "statement": "The supplied record for “Harbor Bells” has media type metadata present."}, {"id": "duration_present", "statement": "The supplied record for “Harbor Bells” has duration metadata present."}, {"id": "rights_present", "statement": "The supplied record for “Harbor Bells” has rights metadata present."}, {"id": "subject_present", "statement": "The supplied record for “Harbor Bells” has subject metadata present."}, {"id": "caption_present", "statement": "The supplied record for “Harbor Bells” has caption metadata present."}], "base_state_json": "[{\"speaker\":\"Catalog technician\",\"text\":\"Operational handoff: inspection of the supplied record confirms title “Harbor Bells,” creator Nia Voss, date 2031-06-14, media type WAV, and duration 02:00. The subject field contains “coastal sound.” The caption field is absent.\"},{\"speaker\":\"Repository technician\",\"text\":\"In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value.\"},{\"speaker\":\"Repository technician\",\"text\":\"In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 contains the value “Copyright 2024 Marina Sound Archive; CC BY 4.0.”\"},{\"speaker\":\"Metadata specialist\",\"text\":\"Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing.\"}]", "base_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "counter_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "focus_atom": "rights_present", "focus_evidence": [{"path": ["1", "text"], "text": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value."}, {"path": ["2", "text"], "text": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 contains the value “Copyright 2024 Marina Sound Archive; CC BY 4.0.”"}], "policy_evidence": [{"path": ["1", "text"], "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}], "rules": [{"justification": "Rights is the only missing required field; the policy assigns level_3 when exactly one of the six required fields is missing, regardless of optional fields.", "target": "level_3", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}, {"justification": "All six required fields are present and exactly one optional field, subject, is present; the policy assigns level_4.", "target": "level_4", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}]}, "verified_pair": {"left": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value.", "negative_left": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value.", "negative_right": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is empty and contains no value.", "right": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 contains the value “Copyright 2024 Marina Sound Archive; CC BY 4.0.”"}, "verifier_independent_model": false}, "family": "scale-diverse-224-007", "id": "scale-diverse-224-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1": "Four or more of the six required metadata fields are missing.", "level_2": "Two or three of the six required metadata fields are missing.", "level_3": "Exactly one of the six required metadata fields is missing, regardless of optional fields.", "level_4": "All six required metadata fields are present, with zero or exactly one optional field present.", "level_5": "All six required metadata fields and both optional fields are present."}, "instructions": "Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata.", "type": "choice"}}, "state": [{"speaker": "Catalog technician", "text": "Operational handoff: inspection of the supplied record confirms title “Harbor Bells,” creator Nia Voss, date 2031-06-14, media type WAV, and duration 02:00. The subject field contains “coastal sound.” The caption field is absent."}, {"speaker": "Repository technician", "text": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is the designated location for the supplied record for “Harbor Bells” to store its rights value."}, {"speaker": "Repository technician", "text": "In the 2026-09-17T14:26:00Z repository snapshot, storage cell RB-731 is empty and contains no value."}, {"speaker": "Metadata specialist", "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}]}, "method": "c2d", "provenance": {"source_id": "diverse-224", "source_is_synthetic": true, "source_sha256": "42be2105c51ead25ff90f2762bb2402ddf400ae28dfdeec25ee9b7154cc0ab7b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_3"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Level 3 decision scope and the same playable 1:42 MP4, while changing only the factual event-code mapping. The two evidence spans are complete factual sentences. Mapping Q-74 to the East Basin Regatta does not conflict with any unchanged assertion because Q-74 is the description’s sole event reference and no other sentence identifies its event. Neither context states a classification answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "supported", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "full_context_fact_states": {"base": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "supported", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "counterfactual": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "refuted", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "remove_left": {"d_event": "unknown"}, "remove_right": {"d_event": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"d_event": "unknown"}, "negative_pair": {"d_event": "refuted"}, "negative_sentence": {"d_event": "unknown"}, "positive_pair": {"d_event": "supported"}, "right": {"d_event": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual metadata or description relationship, and the focus atom is the factual relationship of whether the description conveys the North Quay Festival event. The base and counter assignments are jointly realizable and differ only in that focus atom. Empty policy_evidence is correct because the governing completeness criteria and Level 2/Level 3 rules are already retained in original_input.questions; no substantive state-originating policy is needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction explicitly requires all four administrative fields and all five descriptive facts to be supported. Under the unchanged question, that is sufficient for Level 3 and the true target.", "rule_index": 0, "sound": true}, {"reason": "The conjunction requires all administrative fields and four descriptive facts to be supported while the event fact is refuted. Thus exactly one required descriptive fact is absent, making Level 3 inapplicable and the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_creator", "statement": "The record for the playable 1:42 MP4 contains the creator field."}, {"id": "a_date", "statement": "The record for the playable 1:42 MP4 contains the date field."}, {"id": "a_duration", "statement": "The record for the playable 1:42 MP4 contains the duration field."}, {"id": "a_rights", "statement": "The record for the playable 1:42 MP4 contains the publication rights field."}, {"id": "d_performer", "statement": "The description of the playable 1:42 MP4 conveys the performer."}, {"id": "d_action", "statement": "The description of the playable 1:42 MP4 conveys the action."}, {"id": "d_landmark", "statement": "The description of the playable 1:42 MP4 conveys the landmark."}, {"id": "d_weather", "statement": "The description of the playable 1:42 MP4 conveys the weather."}, {"id": "d_event", "statement": "The description of the playable 1:42 MP4 conveys the event as the North Quay Festival."}], "base_state_json": "\"At 14:02 UTC on 6 May 2026, a cataloger opened the record for the playable 1:42 MP4 and completed playback without an error. At 14:05, the administrative panel showed four populated fields: Creator listed Lina Orr, Date listed 2031-06-12, Duration listed 1:42, and Publication rights listed Archive-cleared for public exhibition. At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference. The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the North Quay Festival. At 14:10, the cataloger confirmed that the quoted sentence was the complete description, not a truncated preview.\"", "base_states": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "supported"}], "counter_states": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "refuted"}], "focus_atom": "d_event", "focus_evidence": [{"path": [], "text": "At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference."}, {"path": [], "text": "The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the North Quay Festival."}], "policy_evidence": [], "rules": [{"justification": "All four required administrative fields and all five required descriptive facts are present, so Level 3 applies.", "target": "true", "when": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "supported"}]}, {"justification": "All administrative fields are present, but exactly one required content fact—the event—is omitted, so Level 2 rather than Level 3 applies.", "target": "false", "when": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference.", "negative_left": "At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference.", "negative_right": "The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the East Basin Regatta.", "right": "The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the North Quay Festival."}, "verifier_independent_model": false}, "family": "scale-diverse-225-001", "id": "scale-diverse-225-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not assign Level 3 because at least one required descriptive fact is missing; assign the lower applicable level.", "true": "Yes — assign Level 3 because every required administrative and descriptive fact is present or accurately paraphrased."}, "instructions": "Decide whether the media librarian should assign metadata completeness Level 3. Level 3 requires all four administrative fields and a description conveying the performer, action, landmark, weather, and event. Accurate synonyms and paraphrases count. Level 2 applies when all administrative fields are present but exactly one required content fact is omitted. Answer yes or no.", "type": "noul"}}, "state": "At 14:02 UTC on 6 May 2026, a cataloger opened the record for the playable 1:42 MP4 and completed playback without an error. At 14:05, the administrative panel showed four populated fields: Creator listed Lina Orr, Date listed 2031-06-12, Duration listed 1:42, and Publication rights listed Archive-cleared for public exhibition. At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference. The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the North Quay Festival. At 14:10, the cataloger confirmed that the quoted sentence was the complete description, not a truncated preview."}, "method": "c2d", "provenance": {"source_id": "diverse-225", "source_is_synthetic": true, "source_sha256": "35e14e30f7cc8db92c465c8d2163577ba11656c8c37062d56493dbbfbfd99bf0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Level 3 decision scope and the same playable 1:42 MP4, while changing only the factual event-code mapping. The two evidence spans are complete factual sentences. Mapping Q-74 to the East Basin Regatta does not conflict with any unchanged assertion because Q-74 is the description’s sole event reference and no other sentence identifies its event. Neither context states a classification answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "refuted", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "full_context_fact_states": {"base": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "supported", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "counterfactual": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "refuted", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "remove_left": {"d_event": "unknown"}, "remove_right": {"d_event": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"d_event": "unknown"}, "negative_pair": {"d_event": "refuted"}, "negative_sentence": {"d_event": "unknown"}, "positive_pair": {"d_event": "supported"}, "right": {"d_event": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual metadata or description relationship, and the focus atom is the factual relationship of whether the description conveys the North Quay Festival event. The base and counter assignments are jointly realizable and differ only in that focus atom. Empty policy_evidence is correct because the governing completeness criteria and Level 2/Level 3 rules are already retained in original_input.questions; no substantive state-originating policy is needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction explicitly requires all four administrative fields and all five descriptive facts to be supported. Under the unchanged question, that is sufficient for Level 3 and the true target.", "rule_index": 0, "sound": true}, {"reason": "The conjunction requires all administrative fields and four descriptive facts to be supported while the event fact is refuted. Thus exactly one required descriptive fact is absent, making Level 3 inapplicable and the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_creator", "statement": "The record for the playable 1:42 MP4 contains the creator field."}, {"id": "a_date", "statement": "The record for the playable 1:42 MP4 contains the date field."}, {"id": "a_duration", "statement": "The record for the playable 1:42 MP4 contains the duration field."}, {"id": "a_rights", "statement": "The record for the playable 1:42 MP4 contains the publication rights field."}, {"id": "d_performer", "statement": "The description of the playable 1:42 MP4 conveys the performer."}, {"id": "d_action", "statement": "The description of the playable 1:42 MP4 conveys the action."}, {"id": "d_landmark", "statement": "The description of the playable 1:42 MP4 conveys the landmark."}, {"id": "d_weather", "statement": "The description of the playable 1:42 MP4 conveys the weather."}, {"id": "d_event", "statement": "The description of the playable 1:42 MP4 conveys the event as the North Quay Festival."}], "base_state_json": "\"At 14:02 UTC on 6 May 2026, a cataloger opened the record for the playable 1:42 MP4 and completed playback without an error. At 14:05, the administrative panel showed four populated fields: Creator listed Lina Orr, Date listed 2031-06-12, Duration listed 1:42, and Publication rights listed Archive-cleared for public exhibition. At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference. The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the North Quay Festival. At 14:10, the cataloger confirmed that the quoted sentence was the complete description, not a truncated preview.\"", "base_states": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "supported"}], "counter_states": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "refuted"}], "focus_atom": "d_event", "focus_evidence": [{"path": [], "text": "At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference."}, {"path": [], "text": "The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the North Quay Festival."}], "policy_evidence": [], "rules": [{"justification": "All four required administrative fields and all five required descriptive facts are present, so Level 3 applies.", "target": "true", "when": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "supported"}]}, {"justification": "All administrative fields are present, but exactly one required content fact—the event—is omitted, so Level 2 rather than Level 3 applies.", "target": "false", "when": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference.", "negative_left": "At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference.", "negative_right": "The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the East Basin Regatta.", "right": "The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the North Quay Festival."}, "verifier_independent_model": false}, "family": "scale-diverse-225-001", "id": "scale-diverse-225-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not assign Level 3 because at least one required descriptive fact is missing; assign the lower applicable level.", "true": "Yes — assign Level 3 because every required administrative and descriptive fact is present or accurately paraphrased."}, "instructions": "Decide whether the media librarian should assign metadata completeness Level 3. Level 3 requires all four administrative fields and a description conveying the performer, action, landmark, weather, and event. Accurate synonyms and paraphrases count. Level 2 applies when all administrative fields are present but exactly one required content fact is omitted. Answer yes or no.", "type": "noul"}}, "state": "At 14:02 UTC on 6 May 2026, a cataloger opened the record for the playable 1:42 MP4 and completed playback without an error. At 14:05, the administrative panel showed four populated fields: Creator listed Lina Orr, Date listed 2031-06-12, Duration listed 1:42, and Publication rights listed Archive-cleared for public exhibition. At 14:08:32 UTC on 6 May 2026, the full description attached to the playable 1:42 MP4 read “Lena Ortiz sings beside Beacon Arch in steady rain during Q-74,” with Q-74 as its sole event reference. The exhaustive event-code key governing that record at 14:08:32 UTC on 6 May 2026 uniquely mapped Q-74 to the East Basin Regatta. At 14:10, the cataloger confirmed that the quoted sentence was the complete description, not a truncated preview."}, "method": "c2d", "provenance": {"source_id": "diverse-225", "source_is_synthetic": true, "source_sha256": "35e14e30f7cc8db92c465c8d2163577ba11656c8c37062d56493dbbfbfd99bf0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts leave the Level 3 policy and the decision scope unchanged, while retaining the binding to the playable 1:42 MP4. The counterfactual changes only the factual authority-register meaning of EVT-604, and that change does not contradict the unchanged description or other catalog facts. The two evidence spans are complete factual sentences rather than instructions or policy definitions. Neither context includes a gold answer, output direction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "supported", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "full_context_fact_states": {"base": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "supported", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "counterfactual": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "refuted", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "remove_left": {"d_event": "unknown"}, "remove_right": {"d_event": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"d_event": "unknown"}, "negative_pair": {"d_event": "refuted"}, "negative_sentence": {"d_event": "unknown"}, "positive_pair": {"d_event": "supported"}, "right": {"d_event": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual metadata or description relationship, and the focus atom is the factual relationship of whether the description conveys the North Quay Festival event. The base and counter assignments are jointly realizable and differ only in that focus atom. Empty policy_evidence is correct because the governing completeness criteria and Level 2/Level 3 rules are already retained in original_input.questions; no substantive state-originating policy is needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction explicitly requires all four administrative fields and all five descriptive facts to be supported. Under the unchanged question, that is sufficient for Level 3 and the true target.", "rule_index": 0, "sound": true}, {"reason": "The conjunction requires all administrative fields and four descriptive facts to be supported while the event fact is refuted. Thus exactly one required descriptive fact is absent, making Level 3 inapplicable and the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_creator", "statement": "The record for the playable 1:42 MP4 contains the creator field."}, {"id": "a_date", "statement": "The record for the playable 1:42 MP4 contains the date field."}, {"id": "a_duration", "statement": "The record for the playable 1:42 MP4 contains the duration field."}, {"id": "a_rights", "statement": "The record for the playable 1:42 MP4 contains the publication rights field."}, {"id": "d_performer", "statement": "The description of the playable 1:42 MP4 conveys the performer."}, {"id": "d_action", "statement": "The description of the playable 1:42 MP4 conveys the action."}, {"id": "d_landmark", "statement": "The description of the playable 1:42 MP4 conveys the landmark."}, {"id": "d_weather", "statement": "The description of the playable 1:42 MP4 conveys the weather."}, {"id": "d_event", "statement": "The description of the playable 1:42 MP4 conveys the event as the North Quay Festival."}], "base_state_json": "\"During evidence reconciliation, the librarian confirmed through playback that the reviewed object is the playable MP4 under assessment and that the quoted description is the complete description attached to it. No other descriptive text or event-identifying element appears in the catalog record. The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element. For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the North Quay Festival. The authority-register entry applies specifically to the tag as used in this description.\"", "base_states": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "supported"}], "counter_states": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "refuted"}], "focus_atom": "d_event", "focus_evidence": [{"path": [], "text": "The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element."}, {"path": [], "text": "For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the North Quay Festival."}], "policy_evidence": [], "rules": [{"justification": "All four required administrative fields and all five required descriptive facts are present, so Level 3 applies.", "target": "true", "when": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "supported"}]}, {"justification": "All administrative fields are present, but exactly one required content fact—the event—is omitted, so Level 2 rather than Level 3 applies.", "target": "false", "when": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "refuted"}]}]}, "verified_pair": {"left": "The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element.", "negative_left": "The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element.", "negative_right": "For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the East Pier Regatta, an event distinct from the North Quay Festival.", "right": "For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the North Quay Festival."}, "verifier_independent_model": false}, "family": "scale-diverse-225-002", "id": "scale-diverse-225-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not assign Level 3 because at least one required descriptive fact is missing; assign the lower applicable level.", "true": "Yes — assign Level 3 because every required administrative and descriptive fact is present or accurately paraphrased."}, "instructions": "Decide whether the media librarian should assign metadata completeness Level 3. Level 3 requires all four administrative fields and a description conveying the performer, action, landmark, weather, and event. Accurate synonyms and paraphrases count. Level 2 applies when all administrative fields are present but exactly one required content fact is omitted. Answer yes or no.", "type": "noul"}}, "state": "During evidence reconciliation, the librarian confirmed through playback that the reviewed object is the playable MP4 under assessment and that the quoted description is the complete description attached to it. No other descriptive text or event-identifying element appears in the catalog record. The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element. For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the North Quay Festival. The authority-register entry applies specifically to the tag as used in this description."}, "method": "c2d", "provenance": {"source_id": "diverse-225", "source_is_synthetic": true, "source_sha256": "35e14e30f7cc8db92c465c8d2163577ba11656c8c37062d56493dbbfbfd99bf0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts leave the Level 3 policy and the decision scope unchanged, while retaining the binding to the playable 1:42 MP4. The counterfactual changes only the factual authority-register meaning of EVT-604, and that change does not contradict the unchanged description or other catalog facts. The two evidence spans are complete factual sentences rather than instructions or policy definitions. Neither context includes a gold answer, output direction, answer code, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "refuted", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "full_context_fact_states": {"base": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "supported", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "counterfactual": {"a_creator": "supported", "a_date": "supported", "a_duration": "supported", "a_rights": "supported", "d_action": "supported", "d_event": "refuted", "d_landmark": "supported", "d_performer": "supported", "d_weather": "supported"}, "remove_left": {"d_event": "unknown"}, "remove_right": {"d_event": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"d_event": "unknown"}, "negative_pair": {"d_event": "refuted"}, "negative_sentence": {"d_event": "unknown"}, "positive_pair": {"d_event": "supported"}, "right": {"d_event": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual metadata or description relationship, and the focus atom is the factual relationship of whether the description conveys the North Quay Festival event. The base and counter assignments are jointly realizable and differ only in that focus atom. Empty policy_evidence is correct because the governing completeness criteria and Level 2/Level 3 rules are already retained in original_input.questions; no substantive state-originating policy is needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction explicitly requires all four administrative fields and all five descriptive facts to be supported. Under the unchanged question, that is sufficient for Level 3 and the true target.", "rule_index": 0, "sound": true}, {"reason": "The conjunction requires all administrative fields and four descriptive facts to be supported while the event fact is refuted. Thus exactly one required descriptive fact is absent, making Level 3 inapplicable and the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_creator", "statement": "The record for the playable 1:42 MP4 contains the creator field."}, {"id": "a_date", "statement": "The record for the playable 1:42 MP4 contains the date field."}, {"id": "a_duration", "statement": "The record for the playable 1:42 MP4 contains the duration field."}, {"id": "a_rights", "statement": "The record for the playable 1:42 MP4 contains the publication rights field."}, {"id": "d_performer", "statement": "The description of the playable 1:42 MP4 conveys the performer."}, {"id": "d_action", "statement": "The description of the playable 1:42 MP4 conveys the action."}, {"id": "d_landmark", "statement": "The description of the playable 1:42 MP4 conveys the landmark."}, {"id": "d_weather", "statement": "The description of the playable 1:42 MP4 conveys the weather."}, {"id": "d_event", "statement": "The description of the playable 1:42 MP4 conveys the event as the North Quay Festival."}], "base_state_json": "\"During evidence reconciliation, the librarian confirmed through playback that the reviewed object is the playable MP4 under assessment and that the quoted description is the complete description attached to it. No other descriptive text or event-identifying element appears in the catalog record. The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element. For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the North Quay Festival. The authority-register entry applies specifically to the tag as used in this description.\"", "base_states": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "supported"}], "counter_states": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "refuted"}], "focus_atom": "d_event", "focus_evidence": [{"path": [], "text": "The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element."}, {"path": [], "text": "For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the North Quay Festival."}], "policy_evidence": [], "rules": [{"justification": "All four required administrative fields and all five required descriptive facts are present, so Level 3 applies.", "target": "true", "when": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "supported"}]}, {"justification": "All administrative fields are present, but exactly one required content fact—the event—is omitted, so Level 2 rather than Level 3 applies.", "target": "false", "when": [{"atom_id": "a_creator", "state": "supported"}, {"atom_id": "a_date", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_rights", "state": "supported"}, {"atom_id": "d_performer", "state": "supported"}, {"atom_id": "d_action", "state": "supported"}, {"atom_id": "d_landmark", "state": "supported"}, {"atom_id": "d_weather", "state": "supported"}, {"atom_id": "d_event", "state": "refuted"}]}]}, "verified_pair": {"left": "The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element.", "negative_left": "The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element.", "negative_right": "For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the East Pier Regatta, an event distinct from the North Quay Festival.", "right": "For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the North Quay Festival."}, "verifier_independent_model": false}, "family": "scale-diverse-225-002", "id": "scale-diverse-225-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not assign Level 3 because at least one required descriptive fact is missing; assign the lower applicable level.", "true": "Yes — assign Level 3 because every required administrative and descriptive fact is present or accurately paraphrased."}, "instructions": "Decide whether the media librarian should assign metadata completeness Level 3. Level 3 requires all four administrative fields and a description conveying the performer, action, landmark, weather, and event. Accurate synonyms and paraphrases count. Level 2 applies when all administrative fields are present but exactly one required content fact is omitted. Answer yes or no.", "type": "noul"}}, "state": "During evidence reconciliation, the librarian confirmed through playback that the reviewed object is the playable MP4 under assessment and that the quoted description is the complete description attached to it. No other descriptive text or event-identifying element appears in the catalog record. The record for the playable 1:42 MP4 lists creator Mira Sol, date 14 June 2025, duration 1:42, and publication rights held by Quayside Archive, and its description is exactly “Lena Ortiz plays violin beside the Beacon Clocktower in steady drizzle; event tag EVT-604,” with EVT-604 as its only event-identifying element. For this MP4’s cataloging scope, the controlled event tag EVT-604 uniquely denotes the East Pier Regatta, an event distinct from the North Quay Festival. The authority-register entry applies specifically to the tag as used in this description."}, "method": "c2d", "provenance": {"source_id": "diverse-225", "source_is_synthetic": true, "source_sha256": "35e14e30f7cc8db92c465c8d2163577ba11656c8c37062d56493dbbfbfd99bf0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing policy and continue to concern Package Aurora_Interview’s ingestion readiness. The two focus-evidence spans are complete factual sentences linking recording IA-204 to its governing rights declaration and reporting that declaration’s operative content. The counterfactual changes only the authorization assertion, leaving a coherent package with exactly one failed mandatory requirement and no conflicting rights measurement. Neither context includes a score, answer code, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Intake auditor\",\"text\":\"At 09:05 UTC on 14 April 2026, every file submitted in Package Aurora_Interview opened and was readable, and every supplied checksum was valid.\"},{\"speaker\":\"Manifest reviewer\",\"text\":\"At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording.\"},{\"speaker\":\"Document examiner\",\"text\":\"At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder authorizes archive ingestion of recording IA-204.”\"},{\"speaker\":\"Metadata auditor\",\"text\":\"At 09:30 UTC, the package was confirmed identifiable and its core metadata was complete.\"},{\"speaker\":\"Project technician\",\"text\":\"At 09:38 UTC, every project-asset link resolved to its corresponding submitted media asset, and the package was confirmed usable.\"},{\"speaker\":\"Discovery reviewer\",\"text\":\"At 09:45 UTC, the reviewer confirmed that Package Aurora_Interview includes neither subject keywords nor a transcript or equivalent description.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording."}, {"path": ["2", "text"], "text": "At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder authorizes archive ingestion of recording IA-204.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording.", "negative_left": "At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording.", "negative_right": "At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder does not authorize archive ingestion of recording IA-204.”", "right": "At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder authorizes archive ingestion of recording IA-204.”"}, "verifier_independent_model": false}, "family": "scale-diverse-227-001", "id": "scale-diverse-227-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Intake auditor", "text": "At 09:05 UTC on 14 April 2026, every file submitted in Package Aurora_Interview opened and was readable, and every supplied checksum was valid."}, {"speaker": "Manifest reviewer", "text": "At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording."}, {"speaker": "Document examiner", "text": "At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder authorizes archive ingestion of recording IA-204.”"}, {"speaker": "Metadata auditor", "text": "At 09:30 UTC, the package was confirmed identifiable and its core metadata was complete."}, {"speaker": "Project technician", "text": "At 09:38 UTC, every project-asset link resolved to its corresponding submitted media asset, and the package was confirmed usable."}, {"speaker": "Discovery reviewer", "text": "At 09:45 UTC, the reviewer confirmed that Package Aurora_Interview includes neither subject keywords nor a transcript or equivalent description."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing policy and continue to concern Package Aurora_Interview’s ingestion readiness. The two focus-evidence spans are complete factual sentences linking recording IA-204 to its governing rights declaration and reporting that declaration’s operative content. The counterfactual changes only the authorization assertion, leaving a coherent package with exactly one failed mandatory requirement and no conflicting rights measurement. Neither context includes a score, answer code, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Intake auditor\",\"text\":\"At 09:05 UTC on 14 April 2026, every file submitted in Package Aurora_Interview opened and was readable, and every supplied checksum was valid.\"},{\"speaker\":\"Manifest reviewer\",\"text\":\"At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording.\"},{\"speaker\":\"Document examiner\",\"text\":\"At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder authorizes archive ingestion of recording IA-204.”\"},{\"speaker\":\"Metadata auditor\",\"text\":\"At 09:30 UTC, the package was confirmed identifiable and its core metadata was complete.\"},{\"speaker\":\"Project technician\",\"text\":\"At 09:38 UTC, every project-asset link resolved to its corresponding submitted media asset, and the package was confirmed usable.\"},{\"speaker\":\"Discovery reviewer\",\"text\":\"At 09:45 UTC, the reviewer confirmed that Package Aurora_Interview includes neither subject keywords nor a transcript or equivalent description.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording."}, {"path": ["2", "text"], "text": "At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder authorizes archive ingestion of recording IA-204.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording.", "negative_left": "At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording.", "negative_right": "At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder does not authorize archive ingestion of recording IA-204.”", "right": "At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder authorizes archive ingestion of recording IA-204.”"}, "verifier_independent_model": false}, "family": "scale-diverse-227-001", "id": "scale-diverse-227-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Intake auditor", "text": "At 09:05 UTC on 14 April 2026, every file submitted in Package Aurora_Interview opened and was readable, and every supplied checksum was valid."}, {"speaker": "Manifest reviewer", "text": "At 09:15 UTC on 14 April 2026, Package Aurora_Interview's manifest identified its interview as recording IA-204 and designated document RD-47 as the governing rights declaration for that recording."}, {"speaker": "Document examiner", "text": "At 09:22 UTC on 14 April 2026, the operative clause in document RD-47 stated, “The rights holder does not authorize archive ingestion of recording IA-204.”"}, {"speaker": "Metadata auditor", "text": "At 09:30 UTC, the package was confirmed identifiable and its core metadata was complete."}, {"speaker": "Project technician", "text": "At 09:38 UTC, every project-asset link resolved to its corresponding submitted media asset, and the package was confirmed usable."}, {"speaker": "Discovery reviewer", "text": "At 09:45 UTC, the reviewer confirmed that Package Aurora_Interview includes neither subject keywords nor a transcript or equivalent description."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring and routing criteria, while both contexts retain the original governing policy about evidence authority, mandatory requirements, and one-failure review. The package identity and relevant ingestion assessment remain bound to Aurora_Interview, with the added handoff time consistently applied to the rights evidence. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the operative rights authorization from granted to withheld and does not conflict with any unchanged factual assertion. Neither context includes a score, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Ingestion coordinator\",\"text\":\"Package Aurora_Interview is identifiable and usable. Every submitted file opens and is readable, every supplied checksum is valid, and the manifest provides complete core metadata. Every project-asset link resolves to its corresponding submitted media asset.\"},{\"speaker\":\"Rights custodian\",\"text\":\"At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903.\"},{\"speaker\":\"Rights custodian\",\"text\":\"At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 authorized archive ingestion of asset AV-903 without restriction.\"},{\"speaker\":\"Discovery review\",\"text\":\"Package Aurora_Interview has no subject keywords and includes neither a transcript nor an equivalent description.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903."}, {"path": ["2", "text"], "text": "At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 authorized archive ingestion of asset AV-903 without restriction."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903.", "negative_left": "At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903.", "negative_right": "At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 withheld authorization for archive ingestion of asset AV-903.", "right": "At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 authorized archive ingestion of asset AV-903 without restriction."}, "verifier_independent_model": false}, "family": "scale-diverse-227-003", "id": "scale-diverse-227-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Ingestion coordinator", "text": "Package Aurora_Interview is identifiable and usable. Every submitted file opens and is readable, every supplied checksum is valid, and the manifest provides complete core metadata. Every project-asset link resolves to its corresponding submitted media asset."}, {"speaker": "Rights custodian", "text": "At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903."}, {"speaker": "Rights custodian", "text": "At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 authorized archive ingestion of asset AV-903 without restriction."}, {"speaker": "Discovery review", "text": "Package Aurora_Interview has no subject keywords and includes neither a transcript nor an equivalent description."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring and routing criteria, while both contexts retain the original governing policy about evidence authority, mandatory requirements, and one-failure review. The package identity and relevant ingestion assessment remain bound to Aurora_Interview, with the added handoff time consistently applied to the rights evidence. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the operative rights authorization from granted to withheld and does not conflict with any unchanged factual assertion. Neither context includes a score, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Ingestion coordinator\",\"text\":\"Package Aurora_Interview is identifiable and usable. Every submitted file opens and is readable, every supplied checksum is valid, and the manifest provides complete core metadata. Every project-asset link resolves to its corresponding submitted media asset.\"},{\"speaker\":\"Rights custodian\",\"text\":\"At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903.\"},{\"speaker\":\"Rights custodian\",\"text\":\"At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 authorized archive ingestion of asset AV-903 without restriction.\"},{\"speaker\":\"Discovery review\",\"text\":\"Package Aurora_Interview has no subject keywords and includes neither a transcript nor an equivalent description.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903."}, {"path": ["2", "text"], "text": "At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 authorized archive ingestion of asset AV-903 without restriction."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903.", "negative_left": "At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903.", "negative_right": "At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 withheld authorization for archive ingestion of asset AV-903.", "right": "At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 authorized archive ingestion of asset AV-903 without restriction."}, "verifier_independent_model": false}, "family": "scale-diverse-227-003", "id": "scale-diverse-227-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Ingestion coordinator", "text": "Package Aurora_Interview is identifiable and usable. Every submitted file opens and is readable, every supplied checksum is valid, and the manifest provides complete core metadata. Every project-asset link resolves to its corresponding submitted media asset."}, {"speaker": "Rights custodian", "text": "At the 2026-08-14T16:20:00Z ingestion handoff, file RD-47 was the rights declaration governing the interview in Package Aurora_Interview, and that interview was identified within the file as asset AV-903."}, {"speaker": "Rights custodian", "text": "At the 2026-08-14T16:20:00Z ingestion handoff, the signed operative clause in file RD-47 withheld authorization for archive ingestion of asset AV-903."}, {"speaker": "Discovery review", "text": "Package Aurora_Interview has no subject keywords and includes neither a transcript nor an equivalent description."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged question criteria and preserve the original governing policy on evidence authority, mandatory requirements, and one-blocker routing. The package entity and ingestion-readiness decision scope remain unchanged. The two focus spans are complete factual sentences about the package-specific rights declaration and its recorded field value, rather than classifier instructions or abstract policy definitions. Changing RC-73 from “AUTHORIZED” to “DENIED” is a coherent single factual intervention: under the unchanged declaration semantics it changes rights clearance without conflicting with readability, checksums, metadata, asset links, usability, or optional-feature observations. Neither context states a score, answer code, proposition ID, intended label, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Digital archivist\",\"text\":\"Inspection of Package Aurora_Interview found every submitted file readable and every supplied checksum valid. Its core metadata is complete, and every project-asset link resolves to the corresponding submitted media asset.\"},{\"speaker\":\"Rights examiner\",\"text\":\"The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”\"},{\"speaker\":\"Rights system log\",\"text\":\"In that declaration, field RC-73 has the value “AUTHORIZED” at 16:42 UTC on 2026-08-19.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"The manifest uniquely identifies Package Aurora_Interview, and inspection confirms that the package is usable. The package contains no subject keywords and includes neither a transcript nor an equivalent description.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”"}, {"path": ["2", "text"], "text": "In that declaration, field RC-73 has the value “AUTHORIZED” at 16:42 UTC on 2026-08-19."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”", "negative_left": "The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”", "negative_right": "In that declaration, field RC-73 has the value “DENIED” at 16:42 UTC on 2026-08-19.", "right": "In that declaration, field RC-73 has the value “AUTHORIZED” at 16:42 UTC on 2026-08-19."}, "verifier_independent_model": false}, "family": "scale-diverse-227-004", "id": "scale-diverse-227-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Digital archivist", "text": "Inspection of Package Aurora_Interview found every submitted file readable and every supplied checksum valid. Its core metadata is complete, and every project-asset link resolves to the corresponding submitted media asset."}, {"speaker": "Rights examiner", "text": "The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”"}, {"speaker": "Rights system log", "text": "In that declaration, field RC-73 has the value “AUTHORIZED” at 16:42 UTC on 2026-08-19."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}, {"speaker": "Metadata specialist", "text": "The manifest uniquely identifies Package Aurora_Interview, and inspection confirms that the package is usable. The package contains no subject keywords and includes neither a transcript nor an equivalent description."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged question criteria and preserve the original governing policy on evidence authority, mandatory requirements, and one-blocker routing. The package entity and ingestion-readiness decision scope remain unchanged. The two focus spans are complete factual sentences about the package-specific rights declaration and its recorded field value, rather than classifier instructions or abstract policy definitions. Changing RC-73 from “AUTHORIZED” to “DENIED” is a coherent single factual intervention: under the unchanged declaration semantics it changes rights clearance without conflicting with readability, checksums, metadata, asset links, usability, or optional-feature observations. Neither context states a score, answer code, proposition ID, intended label, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Digital archivist\",\"text\":\"Inspection of Package Aurora_Interview found every submitted file readable and every supplied checksum valid. Its core metadata is complete, and every project-asset link resolves to the corresponding submitted media asset.\"},{\"speaker\":\"Rights examiner\",\"text\":\"The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”\"},{\"speaker\":\"Rights system log\",\"text\":\"In that declaration, field RC-73 has the value “AUTHORIZED” at 16:42 UTC on 2026-08-19.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"The manifest uniquely identifies Package Aurora_Interview, and inspection confirms that the package is usable. The package contains no subject keywords and includes neither a transcript nor an equivalent description.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”"}, {"path": ["2", "text"], "text": "In that declaration, field RC-73 has the value “AUTHORIZED” at 16:42 UTC on 2026-08-19."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”", "negative_left": "The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”", "negative_right": "In that declaration, field RC-73 has the value “DENIED” at 16:42 UTC on 2026-08-19.", "right": "In that declaration, field RC-73 has the value “AUTHORIZED” at 16:42 UTC on 2026-08-19."}, "verifier_independent_model": false}, "family": "scale-diverse-227-004", "id": "scale-diverse-227-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Digital archivist", "text": "Inspection of Package Aurora_Interview found every submitted file readable and every supplied checksum valid. Its core metadata is complete, and every project-asset link resolves to the corresponding submitted media asset."}, {"speaker": "Rights examiner", "text": "The sole rights declaration governing the interview in Package Aurora_Interview states that it grants affirmative clearance for archive ingestion if and only if field RC-73 has the value “AUTHORIZED.”"}, {"speaker": "Rights system log", "text": "In that declaration, field RC-73 has the value “DENIED” at 16:42 UTC on 2026-08-19."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}, {"speaker": "Metadata specialist", "text": "The manifest uniquely identifies Package Aurora_Interview, and inspection confirms that the package is usable. The package contains no subject keywords and includes neither a transcript nor an equivalent description."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing ingestion criteria and the donor’s conditional access instruction without adding exceptions, priorities, or missing-evidence defaults. The package, preservation video, person at 00:04:12, and event-date relationship remain fixed; only that person’s birth date changes. The two focus-evidence spans are complete factual sentences. The counterfactual birth date is coherent with the unchanged 2025-08-17 event date and creates no duplicate or contradictory measurement. Neither context contains a readiness code, gold answer, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"During accession review, staff reliably identified the MP4 preservation video, WAV audio, JPG poster, and PRPROJ editing file as the Harbor Day package’s required preservation files. Each required file was playable and uncorrupted, with an attached checksum matching its manifest checksum. The authoritative metadata record supplied title, creator, event-date, and rights-holder values, and recorded the format of every package file. The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17. The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2009-10-30. Frame inspection then found the person at 00:04:12 identifiable and unblurred; every other person appearing in the preservation video was blurred. The authoritative donor record contained the applicable access instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Rights documentation permits archival preservation. Finally, a separate public-access derivative passed usability testing.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17."}, {"path": [], "text": "The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2009-10-30."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17.", "negative_left": "The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17.", "negative_right": "The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2006-04-09.", "right": "The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2009-10-30."}, "verifier_independent_model": false}, "family": "scale-diverse-228-001", "id": "scale-diverse-228-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "During accession review, staff reliably identified the MP4 preservation video, WAV audio, JPG poster, and PRPROJ editing file as the Harbor Day package’s required preservation files. Each required file was playable and uncorrupted, with an attached checksum matching its manifest checksum. The authoritative metadata record supplied title, creator, event-date, and rights-holder values, and recorded the format of every package file. The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17. The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2009-10-30. Frame inspection then found the person at 00:04:12 identifiable and unblurred; every other person appearing in the preservation video was blurred. The authoritative donor record contained the applicable access instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Rights documentation permits archival preservation. Finally, a separate public-access derivative passed usability testing."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing ingestion criteria and the donor’s conditional access instruction without adding exceptions, priorities, or missing-evidence defaults. The package, preservation video, person at 00:04:12, and event-date relationship remain fixed; only that person’s birth date changes. The two focus-evidence spans are complete factual sentences. The counterfactual birth date is coherent with the unchanged 2025-08-17 event date and creates no duplicate or contradictory measurement. Neither context contains a readiness code, gold answer, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"During accession review, staff reliably identified the MP4 preservation video, WAV audio, JPG poster, and PRPROJ editing file as the Harbor Day package’s required preservation files. Each required file was playable and uncorrupted, with an attached checksum matching its manifest checksum. The authoritative metadata record supplied title, creator, event-date, and rights-holder values, and recorded the format of every package file. The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17. The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2009-10-30. Frame inspection then found the person at 00:04:12 identifiable and unblurred; every other person appearing in the preservation video was blurred. The authoritative donor record contained the applicable access instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Rights documentation permits archival preservation. Finally, a separate public-access derivative passed usability testing.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17."}, {"path": [], "text": "The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2009-10-30."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17.", "negative_left": "The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17.", "negative_right": "The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2006-04-09.", "right": "The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2009-10-30."}, "verifier_independent_model": false}, "family": "scale-diverse-228-001", "id": "scale-diverse-228-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "During accession review, staff reliably identified the MP4 preservation video, WAV audio, JPG poster, and PRPROJ editing file as the Harbor Day package’s required preservation files. Each required file was playable and uncorrupted, with an attached checksum matching its manifest checksum. The authoritative metadata record supplied title, creator, event-date, and rights-holder values, and recorded the format of every package file. The Harbor Day event documented by the Harbor Day preservation video occurred on 2025-08-17. The verified birth record for the person shown at 00:04:12 in the Harbor Day preservation video lists that person's date of birth as 2006-04-09. Frame inspection then found the person at 00:04:12 identifiable and unblurred; every other person appearing in the preservation video was blurred. The authoritative donor record contained the applicable access instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Rights documentation permits archival preservation. Finally, a separate public-access derivative passed usability testing."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor’s conditional-access rule and the original ingestion-readiness scope without adding exceptions, priorities, or missing-evidence defaults. They preserve the Harbor Day package, the identified person at 00:04:12, and the 14 June 2026 event-date binding; only the birth year changes from 2009 to 2007. The two evidence spans are complete factual sentences. The changed birth year creates no duplicate or contradictory measurement and remains coherent with the unchanged event date and visibility evidence. Neither context supplies a readiness score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"Repository reconciliation confirmed that every required preservation file is reliably identified as part of the Harbor Day package. Each required file is playable and uncorrupted, has an attached checksum, and matches its manifest checksum. The authoritative metadata record supplies the title, creator, rights holder, and a recorded format for every file. The governing rights permit archival preservation. The authoritative donor record supplies the applicable access instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” A usable public-access derivative is available.\\n\\nAt 00:04:12, reviewers can identify the depicted person, whose face is unblurred; every other person appearing in the preservation video is blurred. The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2009. The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2009."}, {"path": [], "text": "The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2009.", "negative_left": "The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2007.", "negative_right": "The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026.", "right": "The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-228-002", "id": "scale-diverse-228-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "Repository reconciliation confirmed that every required preservation file is reliably identified as part of the Harbor Day package. Each required file is playable and uncorrupted, has an attached checksum, and matches its manifest checksum. The authoritative metadata record supplies the title, creator, rights holder, and a recorded format for every file. The governing rights permit archival preservation. The authoritative donor record supplies the applicable access instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” A usable public-access derivative is available.\n\nAt 00:04:12, reviewers can identify the depicted person, whose face is unblurred; every other person appearing in the preservation video is blurred. The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2009. The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor’s conditional-access rule and the original ingestion-readiness scope without adding exceptions, priorities, or missing-evidence defaults. They preserve the Harbor Day package, the identified person at 00:04:12, and the 14 June 2026 event-date binding; only the birth year changes from 2009 to 2007. The two evidence spans are complete factual sentences. The changed birth year creates no duplicate or contradictory measurement and remains coherent with the unchanged event date and visibility evidence. Neither context supplies a readiness score, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"Repository reconciliation confirmed that every required preservation file is reliably identified as part of the Harbor Day package. Each required file is playable and uncorrupted, has an attached checksum, and matches its manifest checksum. The authoritative metadata record supplies the title, creator, rights holder, and a recorded format for every file. The governing rights permit archival preservation. The authoritative donor record supplies the applicable access instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” A usable public-access derivative is available.\\n\\nAt 00:04:12, reviewers can identify the depicted person, whose face is unblurred; every other person appearing in the preservation video is blurred. The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2009. The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2009."}, {"path": [], "text": "The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2009.", "negative_left": "The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2007.", "negative_right": "The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026.", "right": "The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-228-002", "id": "scale-diverse-228-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "Repository reconciliation confirmed that every required preservation file is reliably identified as part of the Harbor Day package. Each required file is playable and uncorrupted, has an attached checksum, and matches its manifest checksum. The authoritative metadata record supplies the title, creator, rights holder, and a recorded format for every file. The governing rights permit archival preservation. The authoritative donor record supplies the applicable access instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” A usable public-access derivative is available.\n\nAt 00:04:12, reviewers can identify the depicted person, whose face is unblurred; every other person appearing in the preservation video is blurred. The verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 23 November 2007. The authoritative metadata record dates the Harbor Day event depicted in the preservation video to 14 June 2026."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor’s governing conditional-access policy, while the unchanged questions preserve the ingestion criteria. The Harbor Day package, video, person, timestamp, and event-date bindings remain fixed; only Lina Voss’s birth date changes, coherently changing her age at the event without creating duplicate or contradictory facts. The two evidence spans are complete factual sentences, and neither context contains a gold answer, score, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"Operational handoff: The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025. The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2010. Visual inspection confirms that the identified person at 00:04:12 is unblurred, while every other person appearing in the video is blurred. Custody and manifest records reliably identify every required preservation file as belonging to the Harbor Day package. Every required file is playable and uncorrupted, has an attached checksum, and has a checksum matching its manifest checksum. The authoritative metadata record supplies title, creator, event-date, and rights-holder values and records every package file’s format. Rights documentation permits archival preservation, and the authoritative donor record supplies the applicable access instruction. A usable public-access derivative passed review. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025."}, {"path": [], "text": "The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2010."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025.", "negative_left": "The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025.", "negative_right": "The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2006.", "right": "The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2010."}, "verifier_independent_model": false}, "family": "scale-diverse-228-003", "id": "scale-diverse-228-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "Operational handoff: The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025. The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2010. Visual inspection confirms that the identified person at 00:04:12 is unblurred, while every other person appearing in the video is blurred. Custody and manifest records reliably identify every required preservation file as belonging to the Harbor Day package. Every required file is playable and uncorrupted, has an attached checksum, and has a checksum matching its manifest checksum. The authoritative metadata record supplies title, creator, event-date, and rights-holder values and records every package file’s format. Rights documentation permits archival preservation, and the authoritative donor record supplies the applicable access instruction. A usable public-access derivative passed review. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor’s governing conditional-access policy, while the unchanged questions preserve the ingestion criteria. The Harbor Day package, video, person, timestamp, and event-date bindings remain fixed; only Lina Voss’s birth date changes, coherently changing her age at the event without creating duplicate or contradictory facts. The two evidence spans are complete factual sentences, and neither context contains a gold answer, score, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"Operational handoff: The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025. The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2010. Visual inspection confirms that the identified person at 00:04:12 is unblurred, while every other person appearing in the video is blurred. Custody and manifest records reliably identify every required preservation file as belonging to the Harbor Day package. Every required file is playable and uncorrupted, has an attached checksum, and has a checksum matching its manifest checksum. The authoritative metadata record supplies title, creator, event-date, and rights-holder values and records every package file’s format. Rights documentation permits archival preservation, and the authoritative donor record supplies the applicable access instruction. A usable public-access derivative passed review. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025."}, {"path": [], "text": "The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2010."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025.", "negative_left": "The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025.", "negative_right": "The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2006.", "right": "The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2010."}, "verifier_independent_model": false}, "family": "scale-diverse-228-003", "id": "scale-diverse-228-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "Operational handoff: The Harbor Day event documented in the Harbor Day preservation video occurred on 23 August 2025. The completed identity handoff record identifies the person shown at 00:04:12 in the Harbor Day preservation video as Lina Voss, whose verified date of birth is 7 November 2006. Visual inspection confirms that the identified person at 00:04:12 is unblurred, while every other person appearing in the video is blurred. Custody and manifest records reliably identify every required preservation file as belonging to the Harbor Day package. Every required file is playable and uncorrupted, has an attached checksum, and has a checksum matching its manifest checksum. The authoritative metadata record supplies title, creator, event-date, and rights-holder values and records every package file’s format. Rights documentation permits archival preservation, and the authoritative donor record supplies the applicable access instruction. A usable public-access derivative passed review. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Harbor Day package, the relevant person and 00:04:12 timestamp, the ingestion-readiness request scope, and the donor’s governing conditional-access instruction; the unchanged questions preserve all scoring criteria. The two evidence spans are complete factual sentences. Changing only the event date from 2025-02-14 to 2027-10-01 is coherent with the unchanged 2008-09-23 birth date and creates no duplicate conflicting assertion within the counterfactual. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"Field note: A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23. The authoritative metadata record gives 2025-02-14 as the Harbor Day event date. At that timestamp, the person’s face and badge are clearly identifiable and unblurred; every other person appearing anywhere in the video is blurred. All required preservation files are reliably assigned to the Harbor Day package, open and play successfully, and pass corruption checks. Each has an attached checksum matching the manifest. The authoritative metadata record contains populated title, creator, event-date, and rights-holder fields and records the format of every package file. The authoritative donor record supplies the applicable instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Documented rights permit archival preservation. A tested public-access derivative is present and usable.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23."}, {"path": [], "text": "The authoritative metadata record gives 2025-02-14 as the Harbor Day event date."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23.", "negative_left": "A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23.", "negative_right": "The authoritative metadata record gives 2027-10-01 as the Harbor Day event date.", "right": "The authoritative metadata record gives 2025-02-14 as the Harbor Day event date."}, "verifier_independent_model": false}, "family": "scale-diverse-228-004", "id": "scale-diverse-228-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "Field note: A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23. The authoritative metadata record gives 2025-02-14 as the Harbor Day event date. At that timestamp, the person’s face and badge are clearly identifiable and unblurred; every other person appearing anywhere in the video is blurred. All required preservation files are reliably assigned to the Harbor Day package, open and play successfully, and pass corruption checks. Each has an attached checksum matching the manifest. The authoritative metadata record contains populated title, creator, event-date, and rights-holder fields and records the format of every package file. The authoritative donor record supplies the applicable instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Documented rights permit archival preservation. A tested public-access derivative is present and usable."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Harbor Day package, the relevant person and 00:04:12 timestamp, the ingestion-readiness request scope, and the donor’s governing conditional-access instruction; the unchanged questions preserve all scoring criteria. The two evidence spans are complete factual sentences. Changing only the event date from 2025-02-14 to 2027-10-01 is coherent with the unchanged 2008-09-23 birth date and creates no duplicate conflicting assertion within the counterfactual. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"Field note: A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23. The authoritative metadata record gives 2025-02-14 as the Harbor Day event date. At that timestamp, the person’s face and badge are clearly identifiable and unblurred; every other person appearing anywhere in the video is blurred. All required preservation files are reliably assigned to the Harbor Day package, open and play successfully, and pass corruption checks. Each has an attached checksum matching the manifest. The authoritative metadata record contains populated title, creator, event-date, and rights-holder fields and records the format of every package file. The authoritative donor record supplies the applicable instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Documented rights permit archival preservation. A tested public-access derivative is present and usable.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23."}, {"path": [], "text": "The authoritative metadata record gives 2025-02-14 as the Harbor Day event date."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23.", "negative_left": "A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23.", "negative_right": "The authoritative metadata record gives 2027-10-01 as the Harbor Day event date.", "right": "The authoritative metadata record gives 2025-02-14 as the Harbor Day event date."}, "verifier_independent_model": false}, "family": "scale-diverse-228-004", "id": "scale-diverse-228-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "Field note: A verified birth record identifies the person shown at 00:04:12 in the Harbor Day preservation video as born on 2008-09-23. The authoritative metadata record gives 2027-10-01 as the Harbor Day event date. At that timestamp, the person’s face and badge are clearly identifiable and unblurred; every other person appearing anywhere in the video is blurred. All required preservation files are reliably assigned to the Harbor Day package, open and play successfully, and pass corruption checks. Each has an attached checksum matching the manifest. The authoritative metadata record contains populated title, creator, event-date, and rights-holder fields and records the format of every package file. The authoritative donor record supplies the applicable instruction: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Documented rights permit archival preservation. A tested public-access derivative is present and usable."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the brief requirement, role-ownership rules, animatic-advancement rule, and complete urgency rubric without adding exceptions or defaults. They preserve the question’s focus on the remaining issue in the proposed revision of “Paper Moon,” while the counterfactual coherently changes which candidate record matches PM-731: R-V matches in the base, whereas the exact-one-match assertion entails that R-D matches in the counterfactual. The two focus-evidence spans are complete factual sentences. The review findings remain consistent with either a visual-continuity or non-brief-conflicting dialogue issue, and neither context contains an option code, explicit answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731. At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-731. The register contained exactly one remaining issue, and its identifier was identical to exactly one of the identifiers in candidate records R-V and R-D. R-V contained only the issue-category value visual continuity and only the responsible-role value Storyboard artist. R-D contained only the issue-category value dialogue and only the responsible-role value Writer. Final review found the remaining issue significant, assigned urgency 3, and confirmed that it did not conflict with the signed brief or block animatic advancement. The sequence was marked ready for animatic. the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731."}, {"path": [], "text": "At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-731."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731.", "negative_left": "At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731.", "negative_right": "At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-846.", "right": "At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-731."}, "verifier_independent_model": false}, "family": "scale-diverse-230-001", "id": "scale-diverse-230-001-base", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731. At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-731. The register contained exactly one remaining issue, and its identifier was identical to exactly one of the identifiers in candidate records R-V and R-D. R-V contained only the issue-category value visual continuity and only the responsible-role value Storyboard artist. R-D contained only the issue-category value dialogue and only the responsible-role value Writer. Final review found the remaining issue significant, assigned urgency 3, and confirmed that it did not conflict with the signed brief or block animatic advancement. The sequence was marked ready for animatic. the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "storyboard_artist_ready_3"}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the brief requirement, role-ownership rules, animatic-advancement rule, and complete urgency rubric without adding exceptions or defaults. They preserve the question’s focus on the remaining issue in the proposed revision of “Paper Moon,” while the counterfactual coherently changes which candidate record matches PM-731: R-V matches in the base, whereas the exact-one-match assertion entails that R-D matches in the counterfactual. The two focus-evidence spans are complete factual sentences. The review findings remain consistent with either a visual-continuity or non-brief-conflicting dialogue issue, and neither context contains an option code, explicit answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731. At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-731. The register contained exactly one remaining issue, and its identifier was identical to exactly one of the identifiers in candidate records R-V and R-D. R-V contained only the issue-category value visual continuity and only the responsible-role value Storyboard artist. R-D contained only the issue-category value dialogue and only the responsible-role value Writer. Final review found the remaining issue significant, assigned urgency 3, and confirmed that it did not conflict with the signed brief or block animatic advancement. The sequence was marked ready for animatic. the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731."}, {"path": [], "text": "At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-731."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731.", "negative_left": "At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731.", "negative_right": "At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-846.", "right": "At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-731."}, "verifier_independent_model": false}, "family": "scale-diverse-230-001", "id": "scale-diverse-230-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "At 09:14 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed the issue identifier PM-731. At 09:16 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed the issue identifier PM-846. The register contained exactly one remaining issue, and its identifier was identical to exactly one of the identifiers in candidate records R-V and R-D. R-V contained only the issue-category value visual continuity and only the responsible-role value Storyboard artist. R-D contained only the issue-category value dialogue and only the responsible-role value Writer. Final review found the remaining issue significant, assigned urgency 3, and confirmed that it did not conflict with the signed brief or block animatic advancement. The sequence was marked ready for animatic. the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "writer_ready_3"}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the creative constraint, role-ownership rules, animatic-readiness rule, urgency rubric, request scope, and the proposed-revision entity and handoff time. The two evidence spans are complete factual sentences. The counterfactual changes only R-V’s identifier; together with the unchanged exact-one-match assertion, this coherently makes R-D the matching record without creating a duplicate or contradictory measurement. Neither context contains an option key, gold label, proposition ID, output instruction, or impermissible rule-table encoding.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q. At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-47Q. The register identifier was identical to exactly one of the identifiers in candidate records R-V and R-D. The proposed revision had exactly one remaining issue. The sole issue-category value in R-V was visual continuity; its sole responsible-role value was Storyboard artist. The sole issue-category value in R-D was dialogue; its sole responsible-role value was Writer. The remaining issue was significant, did not conflict with the signed brief, and did not block animatic advancement. The sequence was ready for animatic, and the issue’s urgency value was 3. Creative constraint: the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q."}, {"path": [], "text": "At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-47Q."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q.", "negative_left": "At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q.", "negative_right": "At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-83L.", "right": "At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-47Q."}, "verifier_independent_model": false}, "family": "scale-diverse-230-003", "id": "scale-diverse-230-003-base", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q. At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-47Q. The register identifier was identical to exactly one of the identifiers in candidate records R-V and R-D. The proposed revision had exactly one remaining issue. The sole issue-category value in R-V was visual continuity; its sole responsible-role value was Storyboard artist. The sole issue-category value in R-D was dialogue; its sole responsible-role value was Writer. The remaining issue was significant, did not conflict with the signed brief, and did not block animatic advancement. The sequence was ready for animatic, and the issue’s urgency value was 3. Creative constraint: the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "storyboard_artist_ready_3"}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the creative constraint, role-ownership rules, animatic-readiness rule, urgency rubric, request scope, and the proposed-revision entity and handoff time. The two evidence spans are complete factual sentences. The counterfactual changes only R-V’s identifier; together with the unchanged exact-one-match assertion, this coherently makes R-D the matching record without creating a duplicate or contradictory measurement. Neither context contains an option key, gold label, proposition ID, output instruction, or impermissible rule-table encoding.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q. At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-47Q. The register identifier was identical to exactly one of the identifiers in candidate records R-V and R-D. The proposed revision had exactly one remaining issue. The sole issue-category value in R-V was visual continuity; its sole responsible-role value was Storyboard artist. The sole issue-category value in R-D was dialogue; its sole responsible-role value was Writer. The remaining issue was significant, did not conflict with the signed brief, and did not block animatic advancement. The sequence was ready for animatic, and the issue’s urgency value was 3. Creative constraint: the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q."}, {"path": [], "text": "At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-47Q."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q.", "negative_left": "At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q.", "negative_right": "At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-83L.", "right": "At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-47Q."}, "verifier_independent_model": false}, "family": "scale-diverse-230-003", "id": "scale-diverse-230-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "At the 2026-09-17 14:20 UTC operational handoff, the remaining-issue register for the proposed revision of “Paper Moon” showed issue identifier PM-47Q. At the 2026-09-17 14:20 UTC operational handoff, candidate record R-V showed issue identifier PM-83L. The register identifier was identical to exactly one of the identifiers in candidate records R-V and R-D. The proposed revision had exactly one remaining issue. The sole issue-category value in R-V was visual continuity; its sole responsible-role value was Storyboard artist. The sole issue-category value in R-D was dialogue; its sole responsible-role value was Writer. The remaining issue was significant, did not conflict with the signed brief, and did not block animatic advancement. The sequence was ready for animatic, and the issue’s urgency value was 3. Creative constraint: the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "writer_ready_3"}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing brief, ownership, advancement, and urgency policies, while the unchanged questions preserve all choice criteria. The same revision, register field, timestamp, and decision scope remain bound. The two evidence spans are complete factual sentences. In the counterfactual, changing R-V’s identifier to PM-2849 is coherent with the unchanged statement that exactly one of R-V and R-D matches PM-7316, which makes R-D the matching candidate without creating a duplicate contradiction. Neither context contains an answer code, explicit gold label, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"Field note, 14 May 2026: The proposed revision of “Paper Moon” has exactly one issue left in its remaining-issue register. In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316. In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-7316. A validation check records that the register identifier is identical to exactly one of the issue identifiers in candidate records R-V and R-D. R-V has one issue-category value, visual continuity, and one responsible-role value, Storyboard artist. R-D has one issue-category value, dialogue, and one responsible-role value, Writer. Review marks the remaining issue significant, confirms that it neither conflicts with the signed brief nor blocks animatic advancement, and records the sequence as ready for animatic. Its urgency value is 3.\\n\\nPolicy excerpts:\\nthe brief requires Mina’s realization to remain wordless.\\nWriter owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues.\\nA sequence advances to animatic only with no brief-blocking conflict.\\nRevision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316."}, {"path": [], "text": "In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-7316."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316.", "negative_left": "In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316.", "negative_right": "In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-2849.", "right": "In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-7316."}, "verifier_independent_model": false}, "family": "scale-diverse-230-004", "id": "scale-diverse-230-004-base", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "Field note, 14 May 2026: The proposed revision of “Paper Moon” has exactly one issue left in its remaining-issue register. In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316. In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-7316. A validation check records that the register identifier is identical to exactly one of the issue identifiers in candidate records R-V and R-D. R-V has one issue-category value, visual continuity, and one responsible-role value, Storyboard artist. R-D has one issue-category value, dialogue, and one responsible-role value, Writer. Review marks the remaining issue significant, confirms that it neither conflicts with the signed brief nor blocks animatic advancement, and records the sequence as ready for animatic. Its urgency value is 3.\n\nPolicy excerpts:\nthe brief requires Mina’s realization to remain wordless.\nWriter owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues.\nA sequence advances to animatic only with no brief-blocking conflict.\nRevision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "storyboard_artist_ready_3"}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing brief, ownership, advancement, and urgency policies, while the unchanged questions preserve all choice criteria. The same revision, register field, timestamp, and decision scope remain bound. The two evidence spans are complete factual sentences. In the counterfactual, changing R-V’s identifier to PM-2849 is coherent with the unchanged statement that exactly one of R-V and R-D matches PM-7316, which makes R-D the matching candidate without creating a duplicate contradiction. Neither context contains an answer code, explicit gold label, rule table, proposition identifier, output instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"Field note, 14 May 2026: The proposed revision of “Paper Moon” has exactly one issue left in its remaining-issue register. In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316. In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-7316. A validation check records that the register identifier is identical to exactly one of the issue identifiers in candidate records R-V and R-D. R-V has one issue-category value, visual continuity, and one responsible-role value, Storyboard artist. R-D has one issue-category value, dialogue, and one responsible-role value, Writer. Review marks the remaining issue significant, confirms that it neither conflicts with the signed brief nor blocks animatic advancement, and records the sequence as ready for animatic. Its urgency value is 3.\\n\\nPolicy excerpts:\\nthe brief requires Mina’s realization to remain wordless.\\nWriter owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues.\\nA sequence advances to animatic only with no brief-blocking conflict.\\nRevision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316."}, {"path": [], "text": "In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-7316."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316.", "negative_left": "In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316.", "negative_right": "In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-2849.", "right": "In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-7316."}, "verifier_independent_model": false}, "family": "scale-diverse-230-004", "id": "scale-diverse-230-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "Field note, 14 May 2026: The proposed revision of “Paper Moon” has exactly one issue left in its remaining-issue register. In the 14 May 2026 09:20 UTC audit export, the issue-identifier field of the remaining-issue register for the proposed revision of “Paper Moon” contains PM-7316. In the same 14 May 2026 09:20 UTC audit export, the issue-identifier field of candidate record R-V contains PM-2849. A validation check records that the register identifier is identical to exactly one of the issue identifiers in candidate records R-V and R-D. R-V has one issue-category value, visual continuity, and one responsible-role value, Storyboard artist. R-D has one issue-category value, dialogue, and one responsible-role value, Writer. Review marks the remaining issue significant, confirms that it neither conflicts with the signed brief nor blocks animatic advancement, and records the sequence as ready for animatic. Its urgency value is 3.\n\nPolicy excerpts:\nthe brief requires Mina’s realization to remain wordless.\nWriter owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues.\nA sequence advances to animatic only with no brief-blocking conflict.\nRevision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "writer_ready_3"}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing brief, ownership, animatic-readiness, and urgency policies from the original state, while the unchanged questions preserve all choice criteria. The work, remaining-issue decision path, and relevant timestamps remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes the selected remaining issue from PM-742 to PM-319; both candidate records already exist, so this does not create contradictory duplicate measurements or assertions. Neither context includes an answer code, explicit option selection, proposition ID, classifier instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"At 09:10 UTC on 17 September 2026, the revision log showed exactly one issue remaining for the proposed revision of “Paper Moon.” At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier. At 09:15 UTC, R-V listed visual continuity as its sole issue category and Storyboard artist as its sole responsible role; R-D listed PM-319 as its sole issue identifier, dialogue as its sole issue category, and Writer as its sole responsible role. At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier. Reviewers subsequently recorded that the sole remaining issue was significant, did not conflict with the signed brief, and did not block animatic advancement. The sequence was marked ready for animatic, with urgency 3. Production policy states that the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier."}, {"path": [], "text": "At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier.", "negative_left": "At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier.", "negative_right": "At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-319 as its sole issue identifier.", "right": "At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier."}, "verifier_independent_model": false}, "family": "scale-diverse-230-005", "id": "scale-diverse-230-005-base", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "At 09:10 UTC on 17 September 2026, the revision log showed exactly one issue remaining for the proposed revision of “Paper Moon.” At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier. At 09:15 UTC, R-V listed visual continuity as its sole issue category and Storyboard artist as its sole responsible role; R-D listed PM-319 as its sole issue identifier, dialogue as its sole issue category, and Writer as its sole responsible role. At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier. Reviewers subsequently recorded that the sole remaining issue was significant, did not conflict with the signed brief, and did not block animatic advancement. The sequence was marked ready for animatic, with urgency 3. Production policy states that the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "storyboard_artist_ready_3"}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing brief, ownership, animatic-readiness, and urgency policies from the original state, while the unchanged questions preserve all choice criteria. The work, remaining-issue decision path, and relevant timestamps remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes the selected remaining issue from PM-742 to PM-319; both candidate records already exist, so this does not create contradictory duplicate measurements or assertions. Neither context includes an answer code, explicit option selection, proposition ID, classifier instruction, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"At 09:10 UTC on 17 September 2026, the revision log showed exactly one issue remaining for the proposed revision of “Paper Moon.” At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier. At 09:15 UTC, R-V listed visual continuity as its sole issue category and Storyboard artist as its sole responsible role; R-D listed PM-319 as its sole issue identifier, dialogue as its sole issue category, and Writer as its sole responsible role. At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier. Reviewers subsequently recorded that the sole remaining issue was significant, did not conflict with the signed brief, and did not block animatic advancement. The sequence was marked ready for animatic, with urgency 3. Production policy states that the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier."}, {"path": [], "text": "At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier.", "negative_left": "At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier.", "negative_right": "At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-319 as its sole issue identifier.", "right": "At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier."}, "verifier_independent_model": false}, "family": "scale-diverse-230-005", "id": "scale-diverse-230-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "At 09:10 UTC on 17 September 2026, the revision log showed exactly one issue remaining for the proposed revision of “Paper Moon.” At 09:14 UTC on 17 September 2026, candidate record R-V for the proposed revision of “Paper Moon” displayed PM-742 as its sole issue identifier. At 09:15 UTC, R-V listed visual continuity as its sole issue category and Storyboard artist as its sole responsible role; R-D listed PM-319 as its sole issue identifier, dialogue as its sole issue category, and Writer as its sole responsible role. At 09:18 UTC on 17 September 2026, the remaining-issue register for the proposed revision of “Paper Moon” displayed PM-319 as its sole issue identifier. Reviewers subsequently recorded that the sole remaining issue was significant, did not conflict with the signed brief, and did not block animatic advancement. The sequence was marked ready for animatic, with urgency 3. Production policy states that the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "writer_ready_3"}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy from the original state, while the unchanged questions preserve all decision criteria and instructions. The work, revision, remaining-issue scope, timestamp, and decision path remain fixed. The focus evidence contains exactly two complete factual sentences. The counterfactual changes only R-V’s identifier from Q-731 to M-408; together with the unchanged exact-one-match validation, this coherently makes R-D the matching candidate without creating a duplicate or contradiction. Neither context includes an option code, explicit classifier instruction, rule table, proposition ID, or stated gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"Operational handoff: At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier. At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed Q-731 as the sole issue identifier in candidate record R-V. The handoff confirms that exactly one issue remains in the proposed revision. An identifier validation confirms that the register identifier is identical to exactly one of the identifiers in candidate records R-V and R-D. Candidate record R-V has visual continuity as its sole issue-category value and Storyboard artist as its sole responsible-role value. Candidate record R-D has dialogue as its sole issue-category value and Writer as its sole responsible-role value. Reviewers marked the remaining issue significant, nonblocking for animatic advancement, and not in conflict with the signed brief. The sequence is ready for animatic, and the remaining issue has urgency 3.\\n\\nGoverning policy: the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier."}, {"path": [], "text": "At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed Q-731 as the sole issue identifier in candidate record R-V."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier.", "negative_left": "At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier.", "negative_right": "At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed M-408 as the sole issue identifier in candidate record R-V.", "right": "At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed Q-731 as the sole issue identifier in candidate record R-V."}, "verifier_independent_model": false}, "family": "scale-diverse-230-007", "id": "scale-diverse-230-007-base", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "Operational handoff: At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier. At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed Q-731 as the sole issue identifier in candidate record R-V. The handoff confirms that exactly one issue remains in the proposed revision. An identifier validation confirms that the register identifier is identical to exactly one of the identifiers in candidate records R-V and R-D. Candidate record R-V has visual continuity as its sole issue-category value and Storyboard artist as its sole responsible-role value. Candidate record R-D has dialogue as its sole issue-category value and Writer as its sole responsible-role value. Reviewers marked the remaining issue significant, nonblocking for animatic advancement, and not in conflict with the signed brief. The sequence is ready for animatic, and the remaining issue has urgency 3.\n\nGoverning policy: the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "storyboard_artist_ready_3"}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy from the original state, while the unchanged questions preserve all decision criteria and instructions. The work, revision, remaining-issue scope, timestamp, and decision path remain fixed. The focus evidence contains exactly two complete factual sentences. The counterfactual changes only R-V’s identifier from Q-731 to M-408; together with the unchanged exact-one-match validation, this coherently makes R-D the matching candidate without creating a duplicate or contradiction. Neither context includes an option code, explicit classifier instruction, rule table, proposition ID, or stated gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"Operational handoff: At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier. At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed Q-731 as the sole issue identifier in candidate record R-V. The handoff confirms that exactly one issue remains in the proposed revision. An identifier validation confirms that the register identifier is identical to exactly one of the identifiers in candidate records R-V and R-D. Candidate record R-V has visual continuity as its sole issue-category value and Storyboard artist as its sole responsible-role value. Candidate record R-D has dialogue as its sole issue-category value and Writer as its sole responsible-role value. Reviewers marked the remaining issue significant, nonblocking for animatic advancement, and not in conflict with the signed brief. The sequence is ready for animatic, and the remaining issue has urgency 3.\\n\\nGoverning policy: the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier."}, {"path": [], "text": "At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed Q-731 as the sole issue identifier in candidate record R-V."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier.", "negative_left": "At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier.", "negative_right": "At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed M-408 as the sole issue identifier in candidate record R-V.", "right": "At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed Q-731 as the sole issue identifier in candidate record R-V."}, "verifier_independent_model": false}, "family": "scale-diverse-230-007", "id": "scale-diverse-230-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "Operational handoff: At 14:20 UTC on 17 September 2026, the operational-handoff snapshot of the remaining-issue register for the proposed revision of “Paper Moon” listed Q-731 as its sole issue identifier. At 14:20 UTC on 17 September 2026, the same operational-handoff snapshot listed M-408 as the sole issue identifier in candidate record R-V. The handoff confirms that exactly one issue remains in the proposed revision. An identifier validation confirms that the register identifier is identical to exactly one of the identifiers in candidate records R-V and R-D. Candidate record R-V has visual continuity as its sole issue-category value and Storyboard artist as its sole responsible-role value. Candidate record R-D has dialogue as its sole issue-category value and Writer as its sole responsible-role value. Reviewers marked the remaining issue significant, nonblocking for animatic advancement, and not in conflict with the signed brief. The sequence is ready for animatic, and the remaining issue has urgency 3.\n\nGoverning policy: the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "writer_ready_3"}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing brief, role-ownership rules, animatic-advancement rule, and urgency rubric; the unchanged questions preserve all choice criteria and instructions. The same work, remaining-issue path, and dated snapshot bindings are maintained. The two evidence spans are complete factual sentences. In the counterfactual, changing R-V’s identifier to PM-844 remains coherent because the cross-check can be satisfied by R-D having PM-731, whose identifier is otherwise unstated; the unchanged dialogue, readiness, and urgency facts do not contradict that mapping. Neither context contains a gold answer, option code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"Field note: The change log records exactly one remaining issue in the proposed revision of “Paper Moon.” The register cross-check confirms that its identifier is identical to exactly one of the identifiers in candidate records R-V and R-D. In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731. In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-731. Candidate record R-V has visual continuity as its sole issue category and Storyboard artist as its sole responsible role. Candidate record R-D has dialogue as its sole issue category and Writer as its sole responsible role. Review finds the remaining issue significant but not in conflict with the signed brief and not blocking animatic advancement. The sequence is ready for animatic, and the issue’s urgency value is 3.\\n\\nthe brief requires Mina’s realization to remain wordless.\\nWriter owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues.\\nA sequence advances to animatic only with no brief-blocking conflict.\\nRevision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731."}, {"path": [], "text": "In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-731."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731.", "negative_left": "In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731.", "negative_right": "In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-844.", "right": "In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-731."}, "verifier_independent_model": false}, "family": "scale-diverse-230-008", "id": "scale-diverse-230-008-base", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "Field note: The change log records exactly one remaining issue in the proposed revision of “Paper Moon.” The register cross-check confirms that its identifier is identical to exactly one of the identifiers in candidate records R-V and R-D. In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731. In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-731. Candidate record R-V has visual continuity as its sole issue category and Storyboard artist as its sole responsible role. Candidate record R-D has dialogue as its sole issue category and Writer as its sole responsible role. Review finds the remaining issue significant but not in conflict with the signed brief and not blocking animatic advancement. The sequence is ready for animatic, and the issue’s urgency value is 3.\n\nthe brief requires Mina’s realization to remain wordless.\nWriter owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues.\nA sequence advances to animatic only with no brief-blocking conflict.\nRevision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "storyboard_artist_ready_3"}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing brief, role-ownership rules, animatic-advancement rule, and urgency rubric; the unchanged questions preserve all choice criteria and instructions. The same work, remaining-issue path, and dated snapshot bindings are maintained. The two evidence spans are complete factual sentences. In the counterfactual, changing R-V’s identifier to PM-844 remains coherent because the cross-check can be satisfied by R-D having PM-731, whose identifier is otherwise unstated; the unchanged dialogue, readiness, and urgency facts do not contradict that mapping. Neither context contains a gold answer, option code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"Field note: The change log records exactly one remaining issue in the proposed revision of “Paper Moon.” The register cross-check confirms that its identifier is identical to exactly one of the identifiers in candidate records R-V and R-D. In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731. In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-731. Candidate record R-V has visual continuity as its sole issue category and Storyboard artist as its sole responsible role. Candidate record R-D has dialogue as its sole issue category and Writer as its sole responsible role. Review finds the remaining issue significant but not in conflict with the signed brief and not blocking animatic advancement. The sequence is ready for animatic, and the issue’s urgency value is 3.\\n\\nthe brief requires Mina’s realization to remain wordless.\\nWriter owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues.\\nA sequence advances to animatic only with no brief-blocking conflict.\\nRevision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731."}, {"path": [], "text": "In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-731."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731.", "negative_left": "In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731.", "negative_right": "In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-844.", "right": "In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-731."}, "verifier_independent_model": false}, "family": "scale-diverse-230-008", "id": "scale-diverse-230-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "Field note: The change log records exactly one remaining issue in the proposed revision of “Paper Moon.” The register cross-check confirms that its identifier is identical to exactly one of the identifiers in candidate records R-V and R-D. In the finalized register snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is the exact character string PM-731. In the finalized candidate-record snapshot dated 17 September 2026 at 14:20 UTC, the issue identifier in candidate record R-V is the exact character string PM-844. Candidate record R-V has visual continuity as its sole issue category and Storyboard artist as its sole responsible role. Candidate record R-D has dialogue as its sole issue category and Writer as its sole responsible role. Review finds the remaining issue significant but not in conflict with the signed brief and not blocking animatic advancement. The sequence is ready for animatic, and the issue’s urgency value is 3.\n\nthe brief requires Mina’s realization to remain wordless.\nWriter owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues.\nA sequence advances to animatic only with no brief-blocking conflict.\nRevision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "writer_ready_3"}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original brief, writer specification, request, review scale, and unchanged questions object, so the governing readiness policy remains intact. The sequence-18 entity and decision scope are unchanged. The two focus spans are complete factual sentences; the counterfactual changes only the register value from 4 to 2, which coherently changes the designated continuity rating without creating a duplicate or contradictory measurement within either context. Neither context embeds an answer, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified holding condition remains one relationship over an explicit interval. The focus is the factual continuity rating, not the readiness policy. The base and counter assignments are realizable with only that rating crossing the threshold. Policy evidence preserves the state-origin brief and review scale needed to interpret the unchanged question; no question-origin policy needed to be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction covers all state-origin story requirements: Mira retains the compass in her left hand until a depicted handoff to Tavi, the gull startles her and causes the map rather than the compass to fall, and both clarity and continuity are at least 3. These conditions are sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a continuity rating of at least 3 entails a rating below 3. Under the preserved four-point scale, that means 1 or 2, which independently requires the false outcome. The RR-18 assignment binds the rating to sequence 18.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before the depicted transfer of the brass compass from Mira to Tavi in sequence 18, Mira holds the brass compass in her left hand at every depicted moment."}, {"id": "a2", "statement": "Sequence 18 depicts Mira giving the brass compass to Tavi."}, {"id": "a3", "statement": "Sequence 18 depicts a gull startling Mira."}, {"id": "a4", "statement": "In sequence 18, the gull startling Mira causes Mira’s map to fall."}, {"id": "a5", "statement": "Sequence 18 does not depict the brass compass falling."}, {"id": "a6", "statement": "Readiness review RR-18 is assigned to sequence 18."}, {"id": "a7", "statement": "Readiness review RR-18 gives sequence 18 a clarity rating of at least 3."}, {"id": "a8", "statement": "Readiness review RR-18 gives sequence 18 a continuity rating of at least 3."}], "base_state_json": "{\"context\":\"For the animated short “Northlight,” the Director’s brief says sequence 18 must show Mira holding the brass compass in her left hand until she gives it to Tavi. The Writer specifies that a gull startles Mira, causing her map—not the compass—to fall. The Animation producer is deciding whether the sequence can enter animatic.\",\"evidence\":[\"Across every depicted moment before the transfer in sequence 18, Mira holds the brass compass in her left hand.\",\"The revised panels then depict Mira giving the brass compass to Tavi. They also show a gull startling Mira and thereby causing her map to fall; no panel depicts the brass compass falling.\",\"Readiness review RR-18 records a clarity rating of 3 for sequence 18.\",\"At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18.\",\"From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 4.\",\"Review scale: 1=unclear/broken, 2=major revision, 3=clear with minor polish, 4=fully clear and continuous.\"],\"request\":\"Is sequence 18 ready to enter animatic under the supplied readiness rule?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "3"], "text": "At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18."}, {"path": ["evidence", "4"], "text": "From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 4."}], "policy_evidence": [{"path": ["context"], "text": "For the animated short “Northlight,” the Director’s brief says sequence 18 must show Mira holding the brass compass in her left hand until she gives it to Tavi. The Writer specifies that a gull startles Mira, causing her map—not the compass—to fall. The Animation producer is deciding whether the sequence can enter animatic."}, {"path": ["evidence", "2"], "text": "Review scale: 1=unclear/broken, 2=major revision, 3=clear with minor polish, 4=fully clear and continuous."}, {"path": ["request"], "text": "Is sequence 18 ready to enter animatic under the supplied readiness rule?"}], "rules": [{"justification": "Sequence 18 satisfies every story requirement from the Director’s brief and Writer’s specification, and its assigned readiness review rates both clarity and continuity at least 3.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "RR-18 is assigned to sequence 18, and explicit refutation of a continuity rating of at least 3 places continuity below 3; on the supplied four-point scale that is a rating of 1 or 2, which requires revision before animatic.", "target": "false", "when": [{"atom_id": "a6", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18.", "negative_left": "At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18.", "negative_right": "From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 2.", "right": "From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 4."}, "verifier_independent_model": false}, "family": "scale-diverse-231-004", "id": "scale-diverse-231-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required action is absent or either clarity or continuity rates 1 or 2, so revision is required before animatic.", "true": "Yes — the sequence satisfies every required story action, clarity and continuity are each rated 3 or 4, and it may enter animatic."}, "instructions": "Verify the revision only against the brief and scene evidence. A sequence is ready only if both clarity and continuity rate at least 3 and no required story action is missing. Answer yes or no; route any blocking continuity note to the Storyboard artist.", "type": "noul"}}, "state": {"context": "For the animated short “Northlight,” the Director’s brief says sequence 18 must show Mira holding the brass compass in her left hand until she gives it to Tavi. The Writer specifies that a gull startles Mira, causing her map—not the compass—to fall. The Animation producer is deciding whether the sequence can enter animatic.", "evidence": ["Across every depicted moment before the transfer in sequence 18, Mira holds the brass compass in her left hand.", "The revised panels then depict Mira giving the brass compass to Tavi. They also show a gull startling Mira and thereby causing her map to fall; no panel depicts the brass compass falling.", "Readiness review RR-18 records a clarity rating of 3 for sequence 18.", "At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18.", "From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 4.", "Review scale: 1=unclear/broken, 2=major revision, 3=clear with minor polish, 4=fully clear and continuous."], "request": "Is sequence 18 ready to enter animatic under the supplied readiness rule?"}}, "method": "c2d", "provenance": {"source_id": "diverse-231", "source_is_synthetic": true, "source_sha256": "b78a6c06bbf1f9e0b63724e3e9e4a807d0d30d0c7fa0fc3a043ae5820bb6f67e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original brief, writer specification, request, review scale, and unchanged questions object, so the governing readiness policy remains intact. The sequence-18 entity and decision scope are unchanged. The two focus spans are complete factual sentences; the counterfactual changes only the register value from 4 to 2, which coherently changes the designated continuity rating without creating a duplicate or contradictory measurement within either context. Neither context embeds an answer, answer code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified holding condition remains one relationship over an explicit interval. The focus is the factual continuity rating, not the readiness policy. The base and counter assignments are realizable with only that rating crossing the threshold. Policy evidence preserves the state-origin brief and review scale needed to interpret the unchanged question; no question-origin policy needed to be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction covers all state-origin story requirements: Mira retains the compass in her left hand until a depicted handoff to Tavi, the gull startles her and causes the map rather than the compass to fall, and both clarity and continuity are at least 3. These conditions are sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a continuity rating of at least 3 entails a rating below 3. Under the preserved four-point scale, that means 1 or 2, which independently requires the false outcome. The RR-18 assignment binds the rating to sequence 18.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before the depicted transfer of the brass compass from Mira to Tavi in sequence 18, Mira holds the brass compass in her left hand at every depicted moment."}, {"id": "a2", "statement": "Sequence 18 depicts Mira giving the brass compass to Tavi."}, {"id": "a3", "statement": "Sequence 18 depicts a gull startling Mira."}, {"id": "a4", "statement": "In sequence 18, the gull startling Mira causes Mira’s map to fall."}, {"id": "a5", "statement": "Sequence 18 does not depict the brass compass falling."}, {"id": "a6", "statement": "Readiness review RR-18 is assigned to sequence 18."}, {"id": "a7", "statement": "Readiness review RR-18 gives sequence 18 a clarity rating of at least 3."}, {"id": "a8", "statement": "Readiness review RR-18 gives sequence 18 a continuity rating of at least 3."}], "base_state_json": "{\"context\":\"For the animated short “Northlight,” the Director’s brief says sequence 18 must show Mira holding the brass compass in her left hand until she gives it to Tavi. The Writer specifies that a gull startles Mira, causing her map—not the compass—to fall. The Animation producer is deciding whether the sequence can enter animatic.\",\"evidence\":[\"Across every depicted moment before the transfer in sequence 18, Mira holds the brass compass in her left hand.\",\"The revised panels then depict Mira giving the brass compass to Tavi. They also show a gull startling Mira and thereby causing her map to fall; no panel depicts the brass compass falling.\",\"Readiness review RR-18 records a clarity rating of 3 for sequence 18.\",\"At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18.\",\"From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 4.\",\"Review scale: 1=unclear/broken, 2=major revision, 3=clear with minor polish, 4=fully clear and continuous.\"],\"request\":\"Is sequence 18 ready to enter animatic under the supplied readiness rule?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "3"], "text": "At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18."}, {"path": ["evidence", "4"], "text": "From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 4."}], "policy_evidence": [{"path": ["context"], "text": "For the animated short “Northlight,” the Director’s brief says sequence 18 must show Mira holding the brass compass in her left hand until she gives it to Tavi. The Writer specifies that a gull startles Mira, causing her map—not the compass—to fall. The Animation producer is deciding whether the sequence can enter animatic."}, {"path": ["evidence", "2"], "text": "Review scale: 1=unclear/broken, 2=major revision, 3=clear with minor polish, 4=fully clear and continuous."}, {"path": ["request"], "text": "Is sequence 18 ready to enter animatic under the supplied readiness rule?"}], "rules": [{"justification": "Sequence 18 satisfies every story requirement from the Director’s brief and Writer’s specification, and its assigned readiness review rates both clarity and continuity at least 3.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "RR-18 is assigned to sequence 18, and explicit refutation of a continuity rating of at least 3 places continuity below 3; on the supplied four-point scale that is a rating of 1 or 2, which requires revision before animatic.", "target": "false", "when": [{"atom_id": "a6", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18.", "negative_left": "At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18.", "negative_right": "From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 2.", "right": "From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 4."}, "verifier_independent_model": false}, "family": "scale-diverse-231-004", "id": "scale-diverse-231-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required action is absent or either clarity or continuity rates 1 or 2, so revision is required before animatic.", "true": "Yes — the sequence satisfies every required story action, clarity and continuity are each rated 3 or 4, and it may enter animatic."}, "instructions": "Verify the revision only against the brief and scene evidence. A sequence is ready only if both clarity and continuity rate at least 3 and no required story action is missing. Answer yes or no; route any blocking continuity note to the Storyboard artist.", "type": "noul"}}, "state": {"context": "For the animated short “Northlight,” the Director’s brief says sequence 18 must show Mira holding the brass compass in her left hand until she gives it to Tavi. The Writer specifies that a gull startles Mira, causing her map—not the compass—to fall. The Animation producer is deciding whether the sequence can enter animatic.", "evidence": ["Across every depicted moment before the transfer in sequence 18, Mira holds the brass compass in her left hand.", "The revised panels then depict Mira giving the brass compass to Tavi. They also show a gull startling Mira and thereby causing her map to fall; no panel depicts the brass compass falling.", "Readiness review RR-18 records a clarity rating of 3 for sequence 18.", "At 14:10 UTC on 12 June 2026, readiness review RR-18 designated the integer in the locked review register's uniquely labeled field C-47 as its continuity rating for sequence 18.", "From 14:00 through 14:15 UTC on 12 June 2026, the locked review register's uniquely labeled field C-47 contained the integer 2.", "Review scale: 1=unclear/broken, 2=major revision, 3=clear with minor polish, 4=fully clear and continuous."], "request": "Is sequence 18 ready to enter animatic under the supplied readiness rule?"}}, "method": "c2d", "provenance": {"source_id": "diverse-231", "source_is_synthetic": true, "source_sha256": "b78a6c06bbf1f9e0b63724e3e9e4a807d0d30d0c7fa0fc3a043ae5820bb6f67e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the brief and do not alter or invent the urgency, routing, readiness, or issue-classification policies from the original question. They remain bound to the proposed revision, its sole remaining annotation note, the Storyboard artist, and animatic advancement. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Mara’s first-seeing time; seeing the compass before the hatch opens violates the brief but is not logically inconsistent with the glance audit, because a glance targeting the scraping sound could still result in incidental visual perception of the compass. Neither context contains an explicit decision, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Event-log editor\",\"text\":\"In the proposed revision's event log, the attic hatch first opens at 14:07:12.400.\"},{\"speaker\":\"Event-log editor\",\"text\":\"In the proposed revision's event log, Mara first sees the brass compass at 14:07:19.650.\"},{\"speaker\":\"Continuity reviewer\",\"text\":\"The glance audit confirms that every glance Mara makes before the hatch first opens targets only the scraping sound.\"},{\"speaker\":\"Revision coordinator\",\"text\":\"The proposed revision has exactly one minor issue. Its sole remaining note identifies a missing storyboard annotation, and the note is marked non-blocking.\"},{\"speaker\":\"Production manager\",\"text\":\"No hold applies to the proposed revision, and the routing check found that no remaining note requires any role other than the Storyboard artist.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.400."}, {"path": ["2", "text"], "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:19.650."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.400.", "negative_left": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.400.", "negative_right": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:08.275.", "right": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:19.650."}, "verifier_independent_model": false}, "family": "scale-diverse-232-001", "id": "scale-diverse-232-001-base", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Event-log editor", "text": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.400."}, {"speaker": "Event-log editor", "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:19.650."}, {"speaker": "Continuity reviewer", "text": "The glance audit confirms that every glance Mara makes before the hatch first opens targets only the scraping sound."}, {"speaker": "Revision coordinator", "text": "The proposed revision has exactly one minor issue. Its sole remaining note identifies a missing storyboard annotation, and the note is marked non-blocking."}, {"speaker": "Production manager", "text": "No hold applies to the proposed revision, and the routing check found that no remaining note requires any role other than the Storyboard artist."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the brief and do not alter or invent the urgency, routing, readiness, or issue-classification policies from the original question. They remain bound to the proposed revision, its sole remaining annotation note, the Storyboard artist, and animatic advancement. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Mara’s first-seeing time; seeing the compass before the hatch opens violates the brief but is not logically inconsistent with the glance audit, because a glance targeting the scraping sound could still result in incidental visual perception of the compass. Neither context contains an explicit decision, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Event-log editor\",\"text\":\"In the proposed revision's event log, the attic hatch first opens at 14:07:12.400.\"},{\"speaker\":\"Event-log editor\",\"text\":\"In the proposed revision's event log, Mara first sees the brass compass at 14:07:19.650.\"},{\"speaker\":\"Continuity reviewer\",\"text\":\"The glance audit confirms that every glance Mara makes before the hatch first opens targets only the scraping sound.\"},{\"speaker\":\"Revision coordinator\",\"text\":\"The proposed revision has exactly one minor issue. Its sole remaining note identifies a missing storyboard annotation, and the note is marked non-blocking.\"},{\"speaker\":\"Production manager\",\"text\":\"No hold applies to the proposed revision, and the routing check found that no remaining note requires any role other than the Storyboard artist.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.400."}, {"path": ["2", "text"], "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:19.650."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.400.", "negative_left": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.400.", "negative_right": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:08.275.", "right": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:19.650."}, "verifier_independent_model": false}, "family": "scale-diverse-232-001", "id": "scale-diverse-232-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Event-log editor", "text": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.400."}, {"speaker": "Event-log editor", "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:08.275."}, {"speaker": "Continuity reviewer", "text": "The glance audit confirms that every glance Mara makes before the hatch first opens targets only the scraping sound."}, {"speaker": "Revision coordinator", "text": "The proposed revision has exactly one minor issue. Its sole remaining note identifies a missing storyboard annotation, and the note is marked non-blocking."}, {"speaker": "Production manager", "text": "No hold applies to the proposed revision, and the routing check found that no remaining note requires any role other than the Storyboard artist."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same brief and do not alter or invent any governing decision, urgency, routing, or readiness policy. The proposed revision, remaining annotation note, Storyboard-artist routing, and animatic decision remain bound to the original question. The two evidence spans are complete factual event-log sentences. The counterfactual changes only Mara’s first-seeing timestamp to before the hatch opening; this creates a brief violation rather than an internal factual contradiction, because seeing the compass need not contradict the separate assertion that her glance target was the scraping sound. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Event logger\",\"text\":\"The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC.\"},{\"speaker\":\"Event logger\",\"text\":\"The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.950 UTC.\"},{\"speaker\":\"Continuity reviewer\",\"text\":\"The glance audit confirms that every glance by Mara before the attic hatch first opens targets only the scraping sound.\"},{\"speaker\":\"Review coordinator\",\"text\":\"The review ledger contains exactly one minor issue. The sole remaining note concerns a missing storyboard annotation and is non-blocking.\"},{\"speaker\":\"Routing coordinator\",\"text\":\"No hold applies to the proposed revision. The routing check found that no remaining revision note requires routing to a role other than the Storyboard artist.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC."}, {"path": ["2", "text"], "text": "The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.950 UTC."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC.", "negative_left": "The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC.", "negative_right": "The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.100 UTC.", "right": "The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.950 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-232-003", "id": "scale-diverse-232-003-base", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Event logger", "text": "The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC."}, {"speaker": "Event logger", "text": "The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.950 UTC."}, {"speaker": "Continuity reviewer", "text": "The glance audit confirms that every glance by Mara before the attic hatch first opens targets only the scraping sound."}, {"speaker": "Review coordinator", "text": "The review ledger contains exactly one minor issue. The sole remaining note concerns a missing storyboard annotation and is non-blocking."}, {"speaker": "Routing coordinator", "text": "No hold applies to the proposed revision. The routing check found that no remaining revision note requires routing to a role other than the Storyboard artist."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same brief and do not alter or invent any governing decision, urgency, routing, or readiness policy. The proposed revision, remaining annotation note, Storyboard-artist routing, and animatic decision remain bound to the original question. The two evidence spans are complete factual event-log sentences. The counterfactual changes only Mara’s first-seeing timestamp to before the hatch opening; this creates a brief violation rather than an internal factual contradiction, because seeing the compass need not contradict the separate assertion that her glance target was the scraping sound. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Event logger\",\"text\":\"The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC.\"},{\"speaker\":\"Event logger\",\"text\":\"The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.950 UTC.\"},{\"speaker\":\"Continuity reviewer\",\"text\":\"The glance audit confirms that every glance by Mara before the attic hatch first opens targets only the scraping sound.\"},{\"speaker\":\"Review coordinator\",\"text\":\"The review ledger contains exactly one minor issue. The sole remaining note concerns a missing storyboard annotation and is non-blocking.\"},{\"speaker\":\"Routing coordinator\",\"text\":\"No hold applies to the proposed revision. The routing check found that no remaining revision note requires routing to a role other than the Storyboard artist.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC."}, {"path": ["2", "text"], "text": "The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.950 UTC."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC.", "negative_left": "The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC.", "negative_right": "The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.100 UTC.", "right": "The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.950 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-232-003", "id": "scale-diverse-232-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Event logger", "text": "The event log for the proposed revision records that the attic hatch first opens at 2026-09-17 14:07:12.400 UTC."}, {"speaker": "Event logger", "text": "The event log for the proposed revision records that Mara first sees the brass compass at 2026-09-17 14:07:12.100 UTC."}, {"speaker": "Continuity reviewer", "text": "The glance audit confirms that every glance by Mara before the attic hatch first opens targets only the scraping sound."}, {"speaker": "Review coordinator", "text": "The review ledger contains exactly one minor issue. The sole remaining note concerns a missing storyboard annotation and is non-blocking."}, {"speaker": "Routing coordinator", "text": "No hold applies to the proposed revision. The routing check found that no remaining revision note requires routing to a role other than the Storyboard artist."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Director’s governing brief and the producer’s advancement condition, while the unchanged questions preserve all scoring and routing policy. The same proposed sequence, characters, objects, event path, and timestamped date remain bound to the question. The two evidence spans are complete factual sentences. The counterfactual replaces only Mara’s first-sight timestamp; it creates a brief violation rather than a contradictory duplicate measurement, and the statement about pre-opening glances targeting the sound does not logically preclude sight occurring by another means. Neither context contains an answer code, output instruction, rule table, proposition ID, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Continuity logger\",\"text\":\"The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026.\"},{\"speaker\":\"Continuity logger\",\"text\":\"The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:15.900 UTC on 18 February 2026.\"},{\"speaker\":\"Review note\",\"text\":\"Every glance by Mara before the hatch first opens targets only the scraping sound. The revision has exactly one minor issue: its sole remaining note concerns a missing storyboard annotation. The note is non-blocking, and no remaining note needs routing to any role other than the Storyboard artist.\"},{\"speaker\":\"Production coordinator\",\"text\":\"No hold applies to the proposed revision.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026."}, {"path": ["2", "text"], "text": "The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:15.900 UTC on 18 February 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026.", "negative_left": "The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026.", "negative_right": "The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:10.100 UTC on 18 February 2026.", "right": "The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:15.900 UTC on 18 February 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-232-004", "id": "scale-diverse-232-004-base", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Continuity logger", "text": "The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026."}, {"speaker": "Continuity logger", "text": "The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:15.900 UTC on 18 February 2026."}, {"speaker": "Review note", "text": "Every glance by Mara before the hatch first opens targets only the scraping sound. The revision has exactly one minor issue: its sole remaining note concerns a missing storyboard annotation. The note is non-blocking, and no remaining note needs routing to any role other than the Storyboard artist."}, {"speaker": "Production coordinator", "text": "No hold applies to the proposed revision."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Director’s governing brief and the producer’s advancement condition, while the unchanged questions preserve all scoring and routing policy. The same proposed sequence, characters, objects, event path, and timestamped date remain bound to the question. The two evidence spans are complete factual sentences. The counterfactual replaces only Mara’s first-sight timestamp; it creates a brief violation rather than a contradictory duplicate measurement, and the statement about pre-opening glances targeting the sound does not logically preclude sight occurring by another means. Neither context contains an answer code, output instruction, rule table, proposition ID, or explicit gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Continuity logger\",\"text\":\"The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026.\"},{\"speaker\":\"Continuity logger\",\"text\":\"The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:15.900 UTC on 18 February 2026.\"},{\"speaker\":\"Review note\",\"text\":\"Every glance by Mara before the hatch first opens targets only the scraping sound. The revision has exactly one minor issue: its sole remaining note concerns a missing storyboard annotation. The note is non-blocking, and no remaining note needs routing to any role other than the Storyboard artist.\"},{\"speaker\":\"Production coordinator\",\"text\":\"No hold applies to the proposed revision.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026."}, {"path": ["2", "text"], "text": "The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:15.900 UTC on 18 February 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026.", "negative_left": "The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026.", "negative_right": "The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:10.100 UTC on 18 February 2026.", "right": "The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:15.900 UTC on 18 February 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-232-004", "id": "scale-diverse-232-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Continuity logger", "text": "The event log for the proposed revision records the attic hatch's first opening at 14:08:12.400 UTC on 18 February 2026."}, {"speaker": "Continuity logger", "text": "The event log for the proposed revision records Mara's first sight of the brass compass at 14:08:10.100 UTC on 18 February 2026."}, {"speaker": "Review note", "text": "Every glance by Mara before the hatch first opens targets only the scraping sound. The revision has exactly one minor issue: its sole remaining note concerns a missing storyboard annotation. The note is non-blocking, and no remaining note needs routing to any role other than the Storyboard artist."}, {"speaker": "Production coordinator", "text": "No hold applies to the proposed revision."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all decision rules, while both contexts retain the director’s governing brief and the relevant production condition. The same proposed sequence, remaining note, routing role, and event-time frame remain bound across contexts. The two evidence spans are complete factual event-log sentences. The counterfactual changes only Mara’s first-seeing time; seeing the compass before the hatch opens violates the brief but does not contradict the separate assertion that her pre-opening glances target the scraping sound, because visual perception need not be the glance’s target. Neither context contains a gold label, answer code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Continuity logger\",\"text\":\"In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026.\"},{\"speaker\":\"Continuity logger\",\"text\":\"In the proposed revision's event log, Mara first sees the brass compass at 14:07:15.930 UTC on 18 September 2026.\"},{\"speaker\":\"Review coordinator\",\"text\":\"Every glance by Mara before the attic hatch first opens targets only the scraping sound. The review found exactly one minor issue: the remaining note concerns a missing storyboard annotation. That note is non-blocking, and no remaining revision note requires routing to any role other than the Storyboard artist. The proposed revision is not subject to a hold.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026."}, {"path": ["2", "text"], "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:15.930 UTC on 18 September 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026.", "negative_left": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026.", "negative_right": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:09.610 UTC on 18 September 2026.", "right": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:15.930 UTC on 18 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-232-005", "id": "scale-diverse-232-005-base", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Continuity logger", "text": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026."}, {"speaker": "Continuity logger", "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:15.930 UTC on 18 September 2026."}, {"speaker": "Review coordinator", "text": "Every glance by Mara before the attic hatch first opens targets only the scraping sound. The review found exactly one minor issue: the remaining note concerns a missing storyboard annotation. That note is non-blocking, and no remaining revision note requires routing to any role other than the Storyboard artist. The proposed revision is not subject to a hold."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all decision rules, while both contexts retain the director’s governing brief and the relevant production condition. The same proposed sequence, remaining note, routing role, and event-time frame remain bound across contexts. The two evidence spans are complete factual event-log sentences. The counterfactual changes only Mara’s first-seeing time; seeing the compass before the hatch opens violates the brief but does not contradict the separate assertion that her pre-opening glances target the scraping sound, because visual perception need not be the glance’s target. Neither context contains a gold label, answer code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Continuity logger\",\"text\":\"In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026.\"},{\"speaker\":\"Continuity logger\",\"text\":\"In the proposed revision's event log, Mara first sees the brass compass at 14:07:15.930 UTC on 18 September 2026.\"},{\"speaker\":\"Review coordinator\",\"text\":\"Every glance by Mara before the attic hatch first opens targets only the scraping sound. The review found exactly one minor issue: the remaining note concerns a missing storyboard annotation. That note is non-blocking, and no remaining revision note requires routing to any role other than the Storyboard artist. The proposed revision is not subject to a hold.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026."}, {"path": ["2", "text"], "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:15.930 UTC on 18 September 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026.", "negative_left": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026.", "negative_right": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:09.610 UTC on 18 September 2026.", "right": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:15.930 UTC on 18 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-232-005", "id": "scale-diverse-232-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Continuity logger", "text": "In the proposed revision's event log, the attic hatch first opens at 14:07:12.480 UTC on 18 September 2026."}, {"speaker": "Continuity logger", "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:07:09.610 UTC on 18 September 2026."}, {"speaker": "Review coordinator", "text": "Every glance by Mara before the attic hatch first opens targets only the scraping sound. The review found exactly one minor issue: the remaining note concerns a missing storyboard annotation. That note is non-blocking, and no remaining revision note requires routing to any role other than the Storyboard artist. The proposed revision is not subject to a hold."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Director’s conditional policy and the original readiness request’s sequence, entity, and production-stage bindings. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only the whistle source’s measured position from 74 degrees to 14 degrees; this coherently changes whether Mara’s 14-to-74-degree eye-line shift points toward the source without creating duplicate or contradictory measurements. The statement that there is no separate unmet requirement remains compatible with the identified eye-line issue. Neither context includes a score, answer code, rule table, proposition identifier, output instruction, or explicit gold-answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; a9 is an allowed quantified absence over an explicit requirement set rather than a final classification. The focus concerns what the boards depict, not the scoring policy. The base and counter assignments can differ only in whether the eye-line shift is depicted while retaining a previously specified contingency correction. Policy evidence preserves the substantive state-originating brief and Director condition; the omitted artist/timing statements are case observations that may be replaced by synthetic observations, while all scale criteria and instructions remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The whistle precedes the shot depicting the door opening, the required eye-line shift is depicted, and a9's refutation excludes every other unmet brief, continuity, staging, or notation requirement. These conditions are sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The whistle-order condition is satisfied, while the eye-line requirement is specifically refuted. The conjunction identifies that as the only unmet requirement, supplies a localized correction and responsible role, and excludes timing or structural changes. This is sufficient for level 3 rather than a story-level block.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The proposed revised boards depict Mara’s eye-line shifting toward the sound source identified as Pip’s three-note whistle."}, {"id": "a2", "statement": "The proposed revision establishes Pip’s three-note whistle before the designated cellar-door shot."}, {"id": "a3", "statement": "The designated cellar-door shot depicts Mara opening the cellar door."}, {"id": "a4", "statement": "The specified correction for a failed eye-line depiction is the addition of Mara’s eye-line arrow toward the registered whistle source."}, {"id": "a5", "statement": "The Storyboard artist is the creative role responsible for the specified eye-line correction."}, {"id": "a6", "statement": "The specified eye-line correction is confined to the cellar-door shot."}, {"id": "a7", "statement": "The specified eye-line correction changes the sequence timing."}, {"id": "a8", "statement": "The specified eye-line correction requires a change to the story structure."}, {"id": "a9", "statement": "The proposed revised sequence has an unmet brief, continuity, staging, or notation requirement other than the required eye-line depiction."}], "base_state_json": "\"For the animated short “Moonlit Parcel,” the brief requires viewers to hear Pip’s three-note whistle before Mara opens the cellar door, and the boards must show Mara’s eye-line shifting toward the sound. “If the whistle is established before the door shot, the sequence may advance once any missing performance notation is added; otherwise, return it to the Writer for structural revision.” The revision log confirms that the proposed sequence establishes the whistle before the designated cellar-door shot, which depicts Mara opening the cellar door. In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale. On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 74 degrees. The final audit identifies no separate unmet brief, continuity, staging, or notation requirement. If the eye-line depiction fails review, the specified correction is adding Mara’s eye-line arrow toward the registered whistle source. The Storyboard artist owns that correction; it is confined to the cellar-door shot and changes neither sequence timing nor story structure.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale."}, {"path": [], "text": "On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 74 degrees."}], "policy_evidence": [{"path": [], "text": "For the animated short “Moonlit Parcel,” the brief requires viewers to hear Pip’s three-note whistle before Mara opens the cellar door, and the boards must show Mara’s eye-line shifting toward the sound."}, {"path": [], "text": "“If the whistle is established before the door shot, the sequence may advance once any missing performance notation is added; otherwise, return it to the Writer for structural revision.”"}], "rules": [{"justification": "The whistle precedes the shot in which Mara opens the door, the boards depict the required eye-line shift, and no other brief, continuity, staging, or notation requirement is unmet; the sequence can advance immediately.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The whistle story condition is satisfied, and the only unmet requirement is the eye-line depiction; its correction is clearly specified, localized to one shot, assigned to the Storyboard artist, and requires neither timing nor story-structure changes.", "target": "3", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale.", "negative_left": "In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale.", "negative_right": "On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 14 degrees.", "right": "On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 74 degrees."}, "verifier_independent_model": false}, "family": "scale-diverse-233-001", "id": "scale-diverse-233-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not viable: the sequence contradicts the core brief and needs replacement or extensive reconception before further review.", "1 — Major revision required: multiple story or staging problems remain and the sequence must return to the Writer or Director for substantial changes and full re-review.", "2 — Blocked pending creative decision: a substantive continuity or intent issue remains that cannot be fixed through straightforward board execution and requires Writer or Director resolution before advancing.", "3 — Conditionally ready: the story condition is satisfied, but one localized, clearly specified board or performance fix must be completed by the relevant creative role before handoff; no story rethink is needed.", "4 — Fully ready: the sequence meets all brief, continuity, staging, and notation requirements and can advance immediately without revision."], "instructions": "Rate the sequence’s readiness for the next production stage using the ordered scale. Apply the Director’s conditional intent, verify the proposed revision against the brief and scene evidence, and identify the appropriate routing implied by the selected level.", "type": "score"}}, "state": "For the animated short “Moonlit Parcel,” the brief requires viewers to hear Pip’s three-note whistle before Mara opens the cellar door, and the boards must show Mara’s eye-line shifting toward the sound. “If the whistle is established before the door shot, the sequence may advance once any missing performance notation is added; otherwise, return it to the Writer for structural revision.” The revision log confirms that the proposed sequence establishes the whistle before the designated cellar-door shot, which depicts Mara opening the cellar door. In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale. On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 74 degrees. The final audit identifies no separate unmet brief, continuity, staging, or notation requirement. If the eye-line depiction fails review, the specified correction is adding Mara’s eye-line arrow toward the registered whistle source. The Storyboard artist owns that correction; it is confined to the cellar-door shot and changes neither sequence timing nor story structure."}, "method": "c2d", "provenance": {"source_id": "diverse-233", "source_is_synthetic": true, "source_sha256": "f13f86e2792e19a89df1340faef87cc18ec3e23d6b7ba37bd58fc334f8de415f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the Director’s conditional policy and the original readiness request’s sequence, entity, and production-stage bindings. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only the whistle source’s measured position from 74 degrees to 14 degrees; this coherently changes whether Mara’s 14-to-74-degree eye-line shift points toward the source without creating duplicate or contradictory measurements. The statement that there is no separate unmet requirement remains compatible with the identified eye-line issue. Neither context includes a score, answer code, rule table, proposition identifier, output instruction, or explicit gold-answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; a9 is an allowed quantified absence over an explicit requirement set rather than a final classification. The focus concerns what the boards depict, not the scoring policy. The base and counter assignments can differ only in whether the eye-line shift is depicted while retaining a previously specified contingency correction. Policy evidence preserves the substantive state-originating brief and Director condition; the omitted artist/timing statements are case observations that may be replaced by synthetic observations, while all scale criteria and instructions remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The whistle precedes the shot depicting the door opening, the required eye-line shift is depicted, and a9's refutation excludes every other unmet brief, continuity, staging, or notation requirement. These conditions are sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The whistle-order condition is satisfied, while the eye-line requirement is specifically refuted. The conjunction identifies that as the only unmet requirement, supplies a localized correction and responsible role, and excludes timing or structural changes. This is sufficient for level 3 rather than a story-level block.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The proposed revised boards depict Mara’s eye-line shifting toward the sound source identified as Pip’s three-note whistle."}, {"id": "a2", "statement": "The proposed revision establishes Pip’s three-note whistle before the designated cellar-door shot."}, {"id": "a3", "statement": "The designated cellar-door shot depicts Mara opening the cellar door."}, {"id": "a4", "statement": "The specified correction for a failed eye-line depiction is the addition of Mara’s eye-line arrow toward the registered whistle source."}, {"id": "a5", "statement": "The Storyboard artist is the creative role responsible for the specified eye-line correction."}, {"id": "a6", "statement": "The specified eye-line correction is confined to the cellar-door shot."}, {"id": "a7", "statement": "The specified eye-line correction changes the sequence timing."}, {"id": "a8", "statement": "The specified eye-line correction requires a change to the story structure."}, {"id": "a9", "statement": "The proposed revised sequence has an unmet brief, continuity, staging, or notation requirement other than the required eye-line depiction."}], "base_state_json": "\"For the animated short “Moonlit Parcel,” the brief requires viewers to hear Pip’s three-note whistle before Mara opens the cellar door, and the boards must show Mara’s eye-line shifting toward the sound. “If the whistle is established before the door shot, the sequence may advance once any missing performance notation is added; otherwise, return it to the Writer for structural revision.” The revision log confirms that the proposed sequence establishes the whistle before the designated cellar-door shot, which depicts Mara opening the cellar door. In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale. On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 74 degrees. The final audit identifies no separate unmet brief, continuity, staging, or notation requirement. If the eye-line depiction fails review, the specified correction is adding Mara’s eye-line arrow toward the registered whistle source. The Storyboard artist owns that correction; it is confined to the cellar-door shot and changes neither sequence timing nor story structure.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale."}, {"path": [], "text": "On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 74 degrees."}], "policy_evidence": [{"path": [], "text": "For the animated short “Moonlit Parcel,” the brief requires viewers to hear Pip’s three-note whistle before Mara opens the cellar door, and the boards must show Mara’s eye-line shifting toward the sound."}, {"path": [], "text": "“If the whistle is established before the door shot, the sequence may advance once any missing performance notation is added; otherwise, return it to the Writer for structural revision.”"}], "rules": [{"justification": "The whistle precedes the shot in which Mara opens the door, the boards depict the required eye-line shift, and no other brief, continuity, staging, or notation requirement is unmet; the sequence can advance immediately.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The whistle story condition is satisfied, and the only unmet requirement is the eye-line depiction; its correction is clearly specified, localized to one shot, assigned to the Storyboard artist, and requires neither timing nor story-structure changes.", "target": "3", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale.", "negative_left": "In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale.", "negative_right": "On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 14 degrees.", "right": "On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 74 degrees."}, "verifier_independent_model": false}, "family": "scale-diverse-233-001", "id": "scale-diverse-233-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not viable: the sequence contradicts the core brief and needs replacement or extensive reconception before further review.", "1 — Major revision required: multiple story or staging problems remain and the sequence must return to the Writer or Director for substantial changes and full re-review.", "2 — Blocked pending creative decision: a substantive continuity or intent issue remains that cannot be fixed through straightforward board execution and requires Writer or Director resolution before advancing.", "3 — Conditionally ready: the story condition is satisfied, but one localized, clearly specified board or performance fix must be completed by the relevant creative role before handoff; no story rethink is needed.", "4 — Fully ready: the sequence meets all brief, continuity, staging, and notation requirements and can advance immediately without revision."], "instructions": "Rate the sequence’s readiness for the next production stage using the ordered scale. Apply the Director’s conditional intent, verify the proposed revision against the brief and scene evidence, and identify the appropriate routing implied by the selected level.", "type": "score"}}, "state": "For the animated short “Moonlit Parcel,” the brief requires viewers to hear Pip’s three-note whistle before Mara opens the cellar door, and the boards must show Mara’s eye-line shifting toward the sound. “If the whistle is established before the door shot, the sequence may advance once any missing performance notation is added; otherwise, return it to the Writer for structural revision.” The revision log confirms that the proposed sequence establishes the whistle before the designated cellar-door shot, which depicts Mara opening the cellar door. In the proposed revised boards for “Moonlit Parcel,” consecutive panels 271 and 272, in that chronological order, depict Mara’s eye-line at 14 degrees and then 74 degrees on the boards’ fixed shot-coordinate scale. On the fixed shot-coordinate scale used in panels 271 and 272 of the proposed revised boards for “Moonlit Parcel,” the depicted source of Pip’s three-note whistle is at 14 degrees. The final audit identifies no separate unmet brief, continuity, staging, or notation requirement. If the eye-line depiction fails review, the specified correction is adding Mara’s eye-line arrow toward the registered whistle source. The Storyboard artist owns that correction; it is confined to the cellar-door shot and changes neither sequence timing nor story structure."}, "method": "c2d", "provenance": {"source_id": "diverse-233", "source_is_synthetic": true, "source_sha256": "f13f86e2792e19a89df1340faef87cc18ec3e23d6b7ba37bd58fc334f8de415f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same readiness criteria, request scope, sequence, current-version binding, and handoff requirements. The two focus spans are complete factual sentences. Changing CN-682 from attached to absent coherently changes panel 17’s camera-note status without creating a duplicate or conflicting measurement; the role-confirmation report can remain factually reported despite the independently documented defect. Neither context states a readiness score, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "supported", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "full_context_fact_states": {"base": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "supported", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "counterfactual": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "refuted", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "remove_left": {"a_panel17_camera_note": "unknown"}, "remove_right": {"a_panel17_camera_note": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_panel17_camera_note": "unknown"}, "negative_pair": {"a_panel17_camera_note": "refuted"}, "negative_sentence": {"a_panel17_camera_note": "unknown"}, "positive_pair": {"a_panel17_camera_note": "supported"}, "right": {"a_panel17_camera_note": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universal camera-note and confirmation atoms remain single relations over their stated sets. The focus—whether panel 17 has a camera note—is factual rather than policy-based. The base and counter assignments differ only on that focus and are realizable: confirmations can report no unresolved revisions even if an objective panel-note defect exists, and the correction can remain non-redesigning in either assignment. Policy evidence preserves the substantive state-originating brief requirements. Supersession, scoring criteria, priorities, and task instructions are already retained in the unchanged questions object and correctly need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the latest assessable version satisfies every state-originating brief requirement: pursuit direction, rescue choice, exactly 24 numbered panels, and camera notes on all panels. They also establish that all relevant role confirmations report no unresolved revisions. This is sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all core story, continuity, and panel-count requirements are met and that panel 17 is the sole camera-note defect. They further establish that correcting it requires no sequence redesign, making it a single localized handoff defect sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_current", "statement": "Version 3 is the latest applicable version of the rooftop escape sequence."}, {"id": "a_storyboard", "statement": "The current storyboard for Version 3 is available for review."}, {"id": "a_brief", "statement": "The essential brief for the rooftop escape sequence is available for review."}, {"id": "a_direction", "statement": "Version 3 maintains readable left-to-right pursuit."}, {"id": "a_rescue", "statement": "Version 3 visibly depicts Lina choosing to rescue the trapped bird before escaping."}, {"id": "a_panel_count", "statement": "Version 3 contains exactly 24 numbered panels."}, {"id": "a_other_camera_notes", "statement": "Every numbered panel in Version 3 other than panel 17 has an attached camera note."}, {"id": "a_panel17_camera_note", "statement": "Numbered panel 17 in Version 3 has an attached camera note."}, {"id": "a_local_correction", "statement": "Attaching a camera note to panel 17 requires no redesign of the sequence."}, {"id": "a_understandable", "statement": "The sequence in Version 3 is understandable."}, {"id": "a_confirmations", "statement": "Every relevant role confirmation for Version 3 reports no unresolved revisions."}], "base_state_json": "{\"context\":\"The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff.\",\"evidence\":[\"Operations has made both the current Version 3 storyboard and the essential sequence brief available for review; Version 3 is the latest applicable version.\",\"Review confirms that Version 3 is understandable, maintains readable left-to-right pursuit, and visibly shows Lina choosing to rescue the trapped bird before escaping.\",\"The board contains exactly 24 numbered panels. Every numbered panel other than panel 17 has an attached camera note.\",\"In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682.\",\"Record CN-682 in that register has the status “camera-note file attached.”\",\"Every relevant role confirmation for Version 3 reports no unresolved revisions. Attaching a camera note to panel 17, if required, is a localized correction and requires no redesign of the sequence.\"],\"request\":\"Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required.\"}", "base_states": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "supported"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}], "counter_states": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "refuted"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}], "focus_atom": "a_panel17_camera_note", "focus_evidence": [{"path": ["evidence", "3"], "text": "In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682."}, {"path": ["evidence", "4"], "text": "Record CN-682 in that register has the status “camera-note file attached.”"}], "policy_evidence": [{"path": ["context"], "text": "The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff."}, {"path": ["request"], "text": "Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required."}], "rules": [{"justification": "The latest applicable version is assessable and meets every stated story, continuity, panel-count, and camera-note requirement, while all relevant confirmations report no unresolved revisions.", "target": "4", "when": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}]}, {"justification": "All core story, continuity, and panel-count requirements are met, but panel 17 is the sole panel without a camera note; attaching that note is a localized correction that requires no sequence redesign.", "target": "3", "when": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "refuted"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}]}]}, "verified_pair": {"left": "In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682.", "negative_left": "In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682.", "negative_right": "Record CN-682 in that register has the status “camera-note file absent.”", "right": "Record CN-682 in that register has the status “camera-note file attached.”"}, "verifier_independent_model": false}, "family": "scale-diverse-234-003", "id": "scale-diverse-234-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable: the current storyboard or essential brief is absent, so readiness cannot be checked and the package must not advance.", "1 — Major rework required: the current version contradicts a core story or continuity requirement, or is too incomplete for meaningful handoff; route substantive revisions to the responsible creative roles.", "2 — Significant revision required: the sequence is understandable but still has at least one unresolved brief requirement or multiple handoff defects; keep it out of layout and route specific notes.", "3 — Conditionally ready: all core story and continuity requirements are met, but one minor, localized handoff defect remains that can be corrected without redesigning the sequence; advance only after that correction is verified.", "4 — Fully ready: the latest version meets every stated story, continuity, panel-count, and camera-note requirement, and relevant role confirmations show no unresolved revisions; approve for layout with no further note routing."], "instructions": "Select one readiness level. Later evidence supersedes earlier evidence when it explicitly replaces a prior version. Evaluate the current sequence against every stated brief and handoff requirement; older resolved notes must not lower the score.", "type": "score"}}, "state": {"context": "The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff.", "evidence": ["Operations has made both the current Version 3 storyboard and the essential sequence brief available for review; Version 3 is the latest applicable version.", "Review confirms that Version 3 is understandable, maintains readable left-to-right pursuit, and visibly shows Lina choosing to rescue the trapped bird before escaping.", "The board contains exactly 24 numbered panels. Every numbered panel other than panel 17 has an attached camera note.", "In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682.", "Record CN-682 in that register has the status “camera-note file attached.”", "Every relevant role confirmation for Version 3 reports no unresolved revisions. Attaching a camera note to panel 17, if required, is a localized correction and requires no redesign of the sequence."], "request": "Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required."}}, "method": "c2d", "provenance": {"source_id": "diverse-234", "source_is_synthetic": true, "source_sha256": "e4655a6f708b74648e4a2cae39e134d37b2fbee776787cd553fa4dad14a35a6e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same readiness criteria, request scope, sequence, current-version binding, and handoff requirements. The two focus spans are complete factual sentences. Changing CN-682 from attached to absent coherently changes panel 17’s camera-note status without creating a duplicate or conflicting measurement; the role-confirmation report can remain factually reported despite the independently documented defect. Neither context states a readiness score, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "refuted", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "full_context_fact_states": {"base": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "supported", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "counterfactual": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "refuted", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "remove_left": {"a_panel17_camera_note": "unknown"}, "remove_right": {"a_panel17_camera_note": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_panel17_camera_note": "unknown"}, "negative_pair": {"a_panel17_camera_note": "refuted"}, "negative_sentence": {"a_panel17_camera_note": "unknown"}, "positive_pair": {"a_panel17_camera_note": "supported"}, "right": {"a_panel17_camera_note": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universal camera-note and confirmation atoms remain single relations over their stated sets. The focus—whether panel 17 has a camera note—is factual rather than policy-based. The base and counter assignments differ only on that focus and are realizable: confirmations can report no unresolved revisions even if an objective panel-note defect exists, and the correction can remain non-redesigning in either assignment. Policy evidence preserves the substantive state-originating brief requirements. Supersession, scoring criteria, priorities, and task instructions are already retained in the unchanged questions object and correctly need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the latest assessable version satisfies every state-originating brief requirement: pursuit direction, rescue choice, exactly 24 numbered panels, and camera notes on all panels. They also establish that all relevant role confirmations report no unresolved revisions. This is sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all core story, continuity, and panel-count requirements are met and that panel 17 is the sole camera-note defect. They further establish that correcting it requires no sequence redesign, making it a single localized handoff defect sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_current", "statement": "Version 3 is the latest applicable version of the rooftop escape sequence."}, {"id": "a_storyboard", "statement": "The current storyboard for Version 3 is available for review."}, {"id": "a_brief", "statement": "The essential brief for the rooftop escape sequence is available for review."}, {"id": "a_direction", "statement": "Version 3 maintains readable left-to-right pursuit."}, {"id": "a_rescue", "statement": "Version 3 visibly depicts Lina choosing to rescue the trapped bird before escaping."}, {"id": "a_panel_count", "statement": "Version 3 contains exactly 24 numbered panels."}, {"id": "a_other_camera_notes", "statement": "Every numbered panel in Version 3 other than panel 17 has an attached camera note."}, {"id": "a_panel17_camera_note", "statement": "Numbered panel 17 in Version 3 has an attached camera note."}, {"id": "a_local_correction", "statement": "Attaching a camera note to panel 17 requires no redesign of the sequence."}, {"id": "a_understandable", "statement": "The sequence in Version 3 is understandable."}, {"id": "a_confirmations", "statement": "Every relevant role confirmation for Version 3 reports no unresolved revisions."}], "base_state_json": "{\"context\":\"The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff.\",\"evidence\":[\"Operations has made both the current Version 3 storyboard and the essential sequence brief available for review; Version 3 is the latest applicable version.\",\"Review confirms that Version 3 is understandable, maintains readable left-to-right pursuit, and visibly shows Lina choosing to rescue the trapped bird before escaping.\",\"The board contains exactly 24 numbered panels. Every numbered panel other than panel 17 has an attached camera note.\",\"In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682.\",\"Record CN-682 in that register has the status “camera-note file attached.”\",\"Every relevant role confirmation for Version 3 reports no unresolved revisions. Attaching a camera note to panel 17, if required, is a localized correction and requires no redesign of the sequence.\"],\"request\":\"Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required.\"}", "base_states": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "supported"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}], "counter_states": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "refuted"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}], "focus_atom": "a_panel17_camera_note", "focus_evidence": [{"path": ["evidence", "3"], "text": "In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682."}, {"path": ["evidence", "4"], "text": "Record CN-682 in that register has the status “camera-note file attached.”"}], "policy_evidence": [{"path": ["context"], "text": "The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff."}, {"path": ["request"], "text": "Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required."}], "rules": [{"justification": "The latest applicable version is assessable and meets every stated story, continuity, panel-count, and camera-note requirement, while all relevant confirmations report no unresolved revisions.", "target": "4", "when": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}]}, {"justification": "All core story, continuity, and panel-count requirements are met, but panel 17 is the sole panel without a camera note; attaching that note is a localized correction that requires no sequence redesign.", "target": "3", "when": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "refuted"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}]}]}, "verified_pair": {"left": "In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682.", "negative_left": "In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682.", "negative_right": "Record CN-682 in that register has the status “camera-note file absent.”", "right": "Record CN-682 in that register has the status “camera-note file attached.”"}, "verifier_independent_model": false}, "family": "scale-diverse-234-003", "id": "scale-diverse-234-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable: the current storyboard or essential brief is absent, so readiness cannot be checked and the package must not advance.", "1 — Major rework required: the current version contradicts a core story or continuity requirement, or is too incomplete for meaningful handoff; route substantive revisions to the responsible creative roles.", "2 — Significant revision required: the sequence is understandable but still has at least one unresolved brief requirement or multiple handoff defects; keep it out of layout and route specific notes.", "3 — Conditionally ready: all core story and continuity requirements are met, but one minor, localized handoff defect remains that can be corrected without redesigning the sequence; advance only after that correction is verified.", "4 — Fully ready: the latest version meets every stated story, continuity, panel-count, and camera-note requirement, and relevant role confirmations show no unresolved revisions; approve for layout with no further note routing."], "instructions": "Select one readiness level. Later evidence supersedes earlier evidence when it explicitly replaces a prior version. Evaluate the current sequence against every stated brief and handoff requirement; older resolved notes must not lower the score.", "type": "score"}}, "state": {"context": "The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff.", "evidence": ["Operations has made both the current Version 3 storyboard and the essential sequence brief available for review; Version 3 is the latest applicable version.", "Review confirms that Version 3 is understandable, maintains readable left-to-right pursuit, and visibly shows Lina choosing to rescue the trapped bird before escaping.", "The board contains exactly 24 numbered panels. Every numbered panel other than panel 17 has an attached camera note.", "In the Version 3 handoff register frozen at 14:20 UTC on 8 September 2026, numbered panel 17 is uniquely mapped to record CN-682.", "Record CN-682 in that register has the status “camera-note file absent.”", "Every relevant role confirmation for Version 3 reports no unresolved revisions. Attaching a camera note to panel 17, if required, is a localized correction and requires no redesign of the sequence."], "request": "Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required."}}, "method": "c2d", "provenance": {"source_id": "diverse-234", "source_is_synthetic": true, "source_sha256": "e4655a6f708b74648e4a2cae39e134d37b2fbee776787cd553fa4dad14a35a6e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing protocol and the unchanged question preserves the full routing criteria and instructions. The request remains bound to the same trial, treatment-control validity decision, routing task, and relevant trial period. The two focus spans are complete factual sentences. The counterfactual coherently changes the documentation for one identified affected event without contradicting the roster, outage report, or master-ledger assertion, since a ledger entry is not necessarily a qualifying independent dated manual record. Neither context embeds an answer, output code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal quantification over an explicit set does not make the atoms non-atomic. A6 is a factual manual-record coverage relation rather than a policy classification. The base and counter assignments can coexist with all non-focus atoms unchanged: in the counter case, an affected event can lack a manual entry while any existing corresponding entries remain dated and independent, and A3 can still hold through a different kind of watering-record entry. The state-derived policy evidence preserves the substantive protocol requirements and sensor-gap exception. Rules originating in the retained questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports all four stated core requirements and establishes the sensor gap plus complete, dated, independent manual coverage of every affected watering event. It therefore suffices for valid_documented_exception and excludes the enumerated protocol failures.", "rule_index": 0, "sound": true}, {"reason": "With the sensor gap established, refutation of A6 entails that at least one affected watering event lacks a corresponding manual-log entry. This is explicit failure of the required sensor exception, not merely unknown documentation, and suffices for invalid_protocol_failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assignment of trial benches 1–4 to drought or control before treatment began was randomized."}, {"id": "A2", "statement": "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days."}, {"id": "A3", "statement": "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records."}, {"id": "A4", "statement": "Each of the trial’s 48 plants has a final dry-mass measurement."}, {"id": "A5", "statement": "The watering-sensor stream for drought bench 3 is missing on trial days 4–6."}, {"id": "A6", "statement": "Every drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 has a corresponding manual-log entry."}, {"id": "A7", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated."}, {"id": "A8", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."}], "base_state_json": "{\"context\":\"The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event.\",\"evidence\":[\"Before treatment, a signed allocation record shows that a random draw assigned trial benches 1–4 between drought and control.\",\"Daily issue sheets confirm that every trial pot received 300 mL on each of the seven pre-trial days.\",\"The reconciled master ledger contains an entry for every drought-pot watering event during the treatment period.\",\"The closeout file contains a final dry-mass measurement for each of the trial’s 48 plants.\",\"An outage report confirms that the watering-sensor stream for drought bench 3 is missing on trial days 4–6.\",\"The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6.\",\"The trial manual log contains entries ML-209, ML-214, and ML-221 corresponding respectively to watering events DB3-417, DB3-428, and DB3-436.\",\"Each manual-log entry identified in the evidence is dated and was created contemporaneously on paper, without input from or reconstruction using the sensor stream.\"],\"request\":\"Classify the trial’s treatment-control validity under the stated protocol and route it accordingly.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": ["evidence", "5"], "text": "The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6."}, {"path": ["evidence", "6"], "text": "The trial manual log contains entries ML-209, ML-214, and ML-221 corresponding respectively to watering events DB3-417, DB3-428, and DB3-436."}], "policy_evidence": [{"path": ["context"], "text": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements."}, {"path": ["context"], "text": "One missing sensor stream is allowed only when dated manual records independently document every affected watering event."}, {"path": ["request"], "text": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}], "rules": [{"justification": "All core requirements are explicitly supported, and every event affected by the established sensor gap has a dated manual-log entry independent of the missing sensor stream.", "target": "valid_documented_exception", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}, {"justification": "The established sensor gap has at least one affected watering event without a corresponding manual-log entry, so the claimed sensor exception fails and the protocol has an uncovered affected event.", "target": "invalid_protocol_failure", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6.", "negative_left": "The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6.", "negative_right": "The trial manual log contains entries ML-209 and ML-214 corresponding respectively to watering events DB3-417 and DB3-428, but contains no entry corresponding to watering event DB3-436.", "right": "The trial manual log contains entries ML-209, ML-214, and ML-221 corresponding respectively to watering events DB3-417, DB3-428, and DB3-436."}, "verifier_independent_model": false}, "family": "scale-diverse-241-003", "id": "scale-diverse-241-003-base", "input": {"questions": {"decision": {"criteria": {"insufficient_evidence_hold": "Place on evidence hold: the record neither proves a core violation nor supplies enough explicit documentation to confirm all core requirements or a claimed sensor exception.", "invalid_protocol_failure": "Route for protocol failure: explicit evidence shows nonrandom assignment, unequal pre-trial watering, an uncovered watering event, missing final growth measurements, or another violated core requirement.", "valid_documented_exception": "Accept as treatment-control valid and route to analysis: every core requirement is explicitly supported, and any sensor gap is fully covered by qualifying dated manual records."}, "instructions": "Select exactly one routing option. Apply the explicit protocol rule: all core requirements must be evidenced; a sensor gap is acceptable only if independently documented manual records cover every affected watering event.", "type": "choice"}}, "state": {"context": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event.", "evidence": ["Before treatment, a signed allocation record shows that a random draw assigned trial benches 1–4 between drought and control.", "Daily issue sheets confirm that every trial pot received 300 mL on each of the seven pre-trial days.", "The reconciled master ledger contains an entry for every drought-pot watering event during the treatment period.", "The closeout file contains a final dry-mass measurement for each of the trial’s 48 plants.", "An outage report confirms that the watering-sensor stream for drought bench 3 is missing on trial days 4–6.", "The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6.", "The trial manual log contains entries ML-209, ML-214, and ML-221 corresponding respectively to watering events DB3-417, DB3-428, and DB3-436.", "Each manual-log entry identified in the evidence is dated and was created contemporaneously on paper, without input from or reconstruction using the sensor stream."], "request": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}}, "method": "c2d", "provenance": {"source_id": "diverse-241", "source_is_synthetic": true, "source_sha256": "1a2bbfd79de14bae070972939aa5cb346cad567f60d5260d99a6e72327b34f05", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_documented_exception"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing protocol and the unchanged question preserves the full routing criteria and instructions. The request remains bound to the same trial, treatment-control validity decision, routing task, and relevant trial period. The two focus spans are complete factual sentences. The counterfactual coherently changes the documentation for one identified affected event without contradicting the roster, outage report, or master-ledger assertion, since a ledger entry is not necessarily a qualifying independent dated manual record. Neither context embeds an answer, output code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal quantification over an explicit set does not make the atoms non-atomic. A6 is a factual manual-record coverage relation rather than a policy classification. The base and counter assignments can coexist with all non-focus atoms unchanged: in the counter case, an affected event can lack a manual entry while any existing corresponding entries remain dated and independent, and A3 can still hold through a different kind of watering-record entry. The state-derived policy evidence preserves the substantive protocol requirements and sensor-gap exception. Rules originating in the retained questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports all four stated core requirements and establishes the sensor gap plus complete, dated, independent manual coverage of every affected watering event. It therefore suffices for valid_documented_exception and excludes the enumerated protocol failures.", "rule_index": 0, "sound": true}, {"reason": "With the sensor gap established, refutation of A6 entails that at least one affected watering event lacks a corresponding manual-log entry. This is explicit failure of the required sensor exception, not merely unknown documentation, and suffices for invalid_protocol_failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assignment of trial benches 1–4 to drought or control before treatment began was randomized."}, {"id": "A2", "statement": "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days."}, {"id": "A3", "statement": "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records."}, {"id": "A4", "statement": "Each of the trial’s 48 plants has a final dry-mass measurement."}, {"id": "A5", "statement": "The watering-sensor stream for drought bench 3 is missing on trial days 4–6."}, {"id": "A6", "statement": "Every drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 has a corresponding manual-log entry."}, {"id": "A7", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated."}, {"id": "A8", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."}], "base_state_json": "{\"context\":\"The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event.\",\"evidence\":[\"Before treatment, a signed allocation record shows that a random draw assigned trial benches 1–4 between drought and control.\",\"Daily issue sheets confirm that every trial pot received 300 mL on each of the seven pre-trial days.\",\"The reconciled master ledger contains an entry for every drought-pot watering event during the treatment period.\",\"The closeout file contains a final dry-mass measurement for each of the trial’s 48 plants.\",\"An outage report confirms that the watering-sensor stream for drought bench 3 is missing on trial days 4–6.\",\"The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6.\",\"The trial manual log contains entries ML-209, ML-214, and ML-221 corresponding respectively to watering events DB3-417, DB3-428, and DB3-436.\",\"Each manual-log entry identified in the evidence is dated and was created contemporaneously on paper, without input from or reconstruction using the sensor stream.\"],\"request\":\"Classify the trial’s treatment-control validity under the stated protocol and route it accordingly.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": ["evidence", "5"], "text": "The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6."}, {"path": ["evidence", "6"], "text": "The trial manual log contains entries ML-209, ML-214, and ML-221 corresponding respectively to watering events DB3-417, DB3-428, and DB3-436."}], "policy_evidence": [{"path": ["context"], "text": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements."}, {"path": ["context"], "text": "One missing sensor stream is allowed only when dated manual records independently document every affected watering event."}, {"path": ["request"], "text": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}], "rules": [{"justification": "All core requirements are explicitly supported, and every event affected by the established sensor gap has a dated manual-log entry independent of the missing sensor stream.", "target": "valid_documented_exception", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}, {"justification": "The established sensor gap has at least one affected watering event without a corresponding manual-log entry, so the claimed sensor exception fails and the protocol has an uncovered affected event.", "target": "invalid_protocol_failure", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6.", "negative_left": "The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6.", "negative_right": "The trial manual log contains entries ML-209 and ML-214 corresponding respectively to watering events DB3-417 and DB3-428, but contains no entry corresponding to watering event DB3-436.", "right": "The trial manual log contains entries ML-209, ML-214, and ML-221 corresponding respectively to watering events DB3-417, DB3-428, and DB3-436."}, "verifier_independent_model": false}, "family": "scale-diverse-241-003", "id": "scale-diverse-241-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"insufficient_evidence_hold": "Place on evidence hold: the record neither proves a core violation nor supplies enough explicit documentation to confirm all core requirements or a claimed sensor exception.", "invalid_protocol_failure": "Route for protocol failure: explicit evidence shows nonrandom assignment, unequal pre-trial watering, an uncovered watering event, missing final growth measurements, or another violated core requirement.", "valid_documented_exception": "Accept as treatment-control valid and route to analysis: every core requirement is explicitly supported, and any sensor gap is fully covered by qualifying dated manual records."}, "instructions": "Select exactly one routing option. Apply the explicit protocol rule: all core requirements must be evidenced; a sensor gap is acceptable only if independently documented manual records cover every affected watering event.", "type": "choice"}}, "state": {"context": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event.", "evidence": ["Before treatment, a signed allocation record shows that a random draw assigned trial benches 1–4 between drought and control.", "Daily issue sheets confirm that every trial pot received 300 mL on each of the seven pre-trial days.", "The reconciled master ledger contains an entry for every drought-pot watering event during the treatment period.", "The closeout file contains a final dry-mass measurement for each of the trial’s 48 plants.", "An outage report confirms that the watering-sensor stream for drought bench 3 is missing on trial days 4–6.", "The completed operational handoff roster lists watering events DB3-417, DB3-428, and DB3-436 as the only drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6.", "The trial manual log contains entries ML-209 and ML-214 corresponding respectively to watering events DB3-417 and DB3-428, but contains no entry corresponding to watering event DB3-436.", "Each manual-log entry identified in the evidence is dated and was created contemporaneously on paper, without input from or reconstruction using the sensor stream."], "request": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}}, "method": "c2d", "provenance": {"source_id": "diverse-241", "source_is_synthetic": true, "source_sha256": "1a2bbfd79de14bae070972939aa5cb346cad567f60d5260d99a6e72327b34f05", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_protocol_failure"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the trial’s measurement requirements without altering the decision rubric or introducing exceptions, priorities, or missing-evidence defaults. The same 20-tray, 10-day bean trial and treatment-control validity question remain in scope. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the watering record from a documented tray-07 crossover to schedules matching the documented assignments, while leaving the 187-of-200 reading count and complete endpoints unchanged. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"trial protocol\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"},{\"speaker\":\"records auditor\",\"text\":\"Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present.\"},{\"speaker\":\"watering-log reviewer\",\"text\":\"During the 10-day bean-tray trial, tray 07 received the control-group watering schedule, trays 01 through 06 and 08 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["1", "text"], "text": "Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present."}, {"path": ["2", "text"], "text": "During the 10-day bean-tray trial, tray 07 received the control-group watering schedule, trays 01 through 06 and 08 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present.", "negative_left": "Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present.", "negative_right": "During the 10-day bean-tray trial, trays 01 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule.", "right": "During the 10-day bean-tray trial, tray 07 received the control-group watering schedule, trays 01 through 06 and 08 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule."}, "verifier_independent_model": false}, "family": "scale-diverse-242-001", "id": "scale-diverse-242-001-base", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "trial protocol", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}, {"speaker": "records auditor", "text": "Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present."}, {"speaker": "watering-log reviewer", "text": "During the 10-day bean-tray trial, tray 07 received the control-group watering schedule, trays 01 through 06 and 08 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_documented_crossover"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the trial’s measurement requirements without altering the decision rubric or introducing exceptions, priorities, or missing-evidence defaults. The same 20-tray, 10-day bean trial and treatment-control validity question remain in scope. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the watering record from a documented tray-07 crossover to schedules matching the documented assignments, while leaving the 187-of-200 reading count and complete endpoints unchanged. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"trial protocol\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"},{\"speaker\":\"records auditor\",\"text\":\"Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present.\"},{\"speaker\":\"watering-log reviewer\",\"text\":\"During the 10-day bean-tray trial, tray 07 received the control-group watering schedule, trays 01 through 06 and 08 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["1", "text"], "text": "Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present."}, {"path": ["2", "text"], "text": "During the 10-day bean-tray trial, tray 07 received the control-group watering schedule, trays 01 through 06 and 08 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present.", "negative_left": "Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present.", "negative_right": "During the 10-day bean-tray trial, trays 01 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule.", "right": "During the 10-day bean-tray trial, tray 07 received the control-group watering schedule, trays 01 through 06 and 08 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule."}, "verifier_independent_model": false}, "family": "scale-diverse-242-001", "id": "scale-diverse-242-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "trial protocol", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}, {"speaker": "records auditor", "text": "Before watering began in the 10-day bean-tray trial, trays 01 through 09 were documented as randomly assigned to the drought-treatment group and trays 10 through 20 were documented as randomly assigned to the control group; 187 of the 200 required daily soil-moisture readings and the required final height and final biomass measurements for every tray were present."}, {"speaker": "watering-log reviewer", "text": "During the 10-day bean-tray trial, trays 01 through 09 received the drought-treatment-group watering schedule, and trays 10 through 20 received the control-group watering schedule."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_at_threshold"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same trial protocol, assignment structure, 20-tray scope, 10-day period, measurement requirements, and decision rubric without adding exceptions or defaults. The two evidence spans are complete factual sentences: one documents assignments and schedules, and the other documents delivered watering. The counterfactual coherently changes B07’s recorded watering from the control schedule to its assigned drought schedule, with no conflicting duplicate measurement or assertion elsewhere. Neither context contains a decision label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"study protocol\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"},{\"speaker\":\"randomization record\",\"text\":\"Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily.\"},{\"speaker\":\"watering record\",\"text\":\"The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B07, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B09, B11, B13, B15, B17, and B19.\"},{\"speaker\":\"data audit\",\"text\":\"The archived dataset contains 180 of the 200 required daily soil-moisture readings. Final height and final biomass measurements are present for every tray from B01 through B20.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["1", "text"], "text": "Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily."}, {"path": ["2", "text"], "text": "The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B07, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B09, B11, B13, B15, B17, and B19."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily.", "negative_left": "Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily.", "negative_right": "The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19.", "right": "The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B07, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B09, B11, B13, B15, B17, and B19."}, "verifier_independent_model": false}, "family": "scale-diverse-242-002", "id": "scale-diverse-242-002-base", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "study protocol", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}, {"speaker": "randomization record", "text": "Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily."}, {"speaker": "watering record", "text": "The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B07, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B09, B11, B13, B15, B17, and B19."}, {"speaker": "data audit", "text": "The archived dataset contains 180 of the 200 required daily soil-moisture readings. Final height and final biomass measurements are present for every tray from B01 through B20."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_documented_crossover"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same trial protocol, assignment structure, 20-tray scope, 10-day period, measurement requirements, and decision rubric without adding exceptions or defaults. The two evidence spans are complete factual sentences: one documents assignments and schedules, and the other documents delivered watering. The counterfactual coherently changes B07’s recorded watering from the control schedule to its assigned drought schedule, with no conflicting duplicate measurement or assertion elsewhere. Neither context contains a decision label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"study protocol\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"},{\"speaker\":\"randomization record\",\"text\":\"Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily.\"},{\"speaker\":\"watering record\",\"text\":\"The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B07, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B09, B11, B13, B15, B17, and B19.\"},{\"speaker\":\"data audit\",\"text\":\"The archived dataset contains 180 of the 200 required daily soil-moisture readings. Final height and final biomass measurements are present for every tray from B01 through B20.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["1", "text"], "text": "Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily."}, {"path": ["2", "text"], "text": "The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B07, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B09, B11, B13, B15, B17, and B19."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily.", "negative_left": "Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily.", "negative_right": "The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19.", "right": "The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B07, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B09, B11, B13, B15, B17, and B19."}, "verifier_independent_model": false}, "family": "scale-diverse-242-002", "id": "scale-diverse-242-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "study protocol", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}, {"speaker": "randomization record", "text": "Before watering began in the 10-day bean-tray trial conducted from 1 June through 10 June 2026, the documented randomized assignments placed trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19 in the drought-treatment group and trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 in the control group, with the drought-treatment schedule set at 20 milliliters at 09:00 daily and the control schedule set at 60 milliliters at 09:00 daily."}, {"speaker": "watering record", "text": "The complete watering log for the 10-day bean-tray trial records 60 milliliters at 09:00 on each trial day for trays B02, B04, B06, B08, B10, B12, B14, B16, B18, and B20 and 20 milliliters at 09:00 on each trial day for trays B01, B03, B05, B07, B09, B11, B13, B15, B17, and B19."}, {"speaker": "data audit", "text": "The archived dataset contains 180 of the 200 required daily soil-moisture readings. Final height and final biomass measurements are present for every tray from B01 through B20."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_at_threshold"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete decision rubric, while both contexts retain the trial’s measurement requirements and the same bean-tray entity, treatment path, and 10-day time scope. The two evidence spans are complete factual sentences; the counterfactual coherently changes only the watering-schedule account to remove the documented tray 11/12 crossover, without conflicting with the unchanged assignment, completeness count, or endpoint assertions. Neither context contains an answer label, code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"records coordinator\",\"text\":\"Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group.\"},{\"speaker\":\"irrigation log reviewer\",\"text\":\"Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 12, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 11, 14, 16, 18, and 20 received the control-group watering schedule.\"},{\"speaker\":\"plant ecologist\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"},{\"speaker\":\"data custodian\",\"text\":\"The locked archive contains 184 of the 200 required daily soil-moisture readings. The endpoint register contains both a final height measurement and a final biomass measurement for each of the 20 trays, with every entry linked to its tray identifier.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group."}, {"path": ["1", "text"], "text": "Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 12, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 11, 14, 16, 18, and 20 received the control-group watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group.", "negative_left": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group.", "negative_right": "Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 received the control-group watering schedule.", "right": "Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 12, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 11, 14, 16, 18, and 20 received the control-group watering schedule."}, "verifier_independent_model": false}, "family": "scale-diverse-242-003", "id": "scale-diverse-242-003-base", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "records coordinator", "text": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group."}, {"speaker": "irrigation log reviewer", "text": "Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 12, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 11, 14, 16, 18, and 20 received the control-group watering schedule."}, {"speaker": "plant ecologist", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}, {"speaker": "data custodian", "text": "The locked archive contains 184 of the 200 required daily soil-moisture readings. The endpoint register contains both a final height measurement and a final biomass measurement for each of the 20 trays, with every entry linked to its tray identifier."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_documented_crossover"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete decision rubric, while both contexts retain the trial’s measurement requirements and the same bean-tray entity, treatment path, and 10-day time scope. The two evidence spans are complete factual sentences; the counterfactual coherently changes only the watering-schedule account to remove the documented tray 11/12 crossover, without conflicting with the unchanged assignment, completeness count, or endpoint assertions. Neither context contains an answer label, code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"records coordinator\",\"text\":\"Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group.\"},{\"speaker\":\"irrigation log reviewer\",\"text\":\"Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 12, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 11, 14, 16, 18, and 20 received the control-group watering schedule.\"},{\"speaker\":\"plant ecologist\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"},{\"speaker\":\"data custodian\",\"text\":\"The locked archive contains 184 of the 200 required daily soil-moisture readings. The endpoint register contains both a final height measurement and a final biomass measurement for each of the 20 trays, with every entry linked to its tray identifier.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group."}, {"path": ["1", "text"], "text": "Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 12, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 11, 14, 16, 18, and 20 received the control-group watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group.", "negative_left": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group.", "negative_right": "Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 received the control-group watering schedule.", "right": "Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 12, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 11, 14, 16, 18, and 20 received the control-group watering schedule."}, "verifier_independent_model": false}, "family": "scale-diverse-242-003", "id": "scale-diverse-242-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "records coordinator", "text": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 in the drought-treatment group and trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 in the control group."}, {"speaker": "irrigation log reviewer", "text": "Throughout the 10-day bean-tray trial, trays 01, 03, 05, 07, 09, 11, 13, 15, 17, and 19 received the drought-group watering schedule, while trays 02, 04, 06, 08, 10, 12, 14, 16, 18, and 20 received the control-group watering schedule."}, {"speaker": "plant ecologist", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}, {"speaker": "data custodian", "text": "The locked archive contains 184 of the 200 required daily soil-moisture readings. The endpoint register contains both a final height measurement and a final biomass measurement for each of the 20 trays, with every entry linked to its tray identifier."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_at_threshold"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the trial requirements and rely on the unchanged questions object for the governing decision rubric. They preserve the same bean-tray trial, treatment assignment, 10-day period, measurements, and decision scope. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the watering record from a documented deviation involving B17 to full adherence, without conflicting with the unchanged assignment, sensor count, or endpoint record. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"field coordinator\",\"text\":\"Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily.\"},{\"speaker\":\"records technician\",\"text\":\"Watering logs for all 10 days record that trays B01–B10 and B17 received the drought-group schedule, while trays B11–B16 and B18–B20 received the control-group schedule.\"},{\"speaker\":\"protocol custodian\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"},{\"speaker\":\"data reviewer\",\"text\":\"The finalized sensor export contains 181 of the 200 required daily soil-moisture readings. Each retained reading has a recognized tray identifier and a trial-day entry.\"},{\"speaker\":\"endpoint reviewer\",\"text\":\"The endpoint register has both a final height measurement and a final biomass measurement for each of the 20 trays; no required endpoint field is blank.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily."}, {"path": ["1", "text"], "text": "Watering logs for all 10 days record that trays B01–B10 and B17 received the drought-group schedule, while trays B11–B16 and B18–B20 received the control-group schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily.", "negative_left": "Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily.", "negative_right": "Watering logs for all 10 days record that trays B01–B10 received the drought-group schedule, while trays B11–B20 received the control-group schedule.", "right": "Watering logs for all 10 days record that trays B01–B10 and B17 received the drought-group schedule, while trays B11–B16 and B18–B20 received the control-group schedule."}, "verifier_independent_model": false}, "family": "scale-diverse-242-004", "id": "scale-diverse-242-004-base", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "field coordinator", "text": "Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily."}, {"speaker": "records technician", "text": "Watering logs for all 10 days record that trays B01–B10 and B17 received the drought-group schedule, while trays B11–B16 and B18–B20 received the control-group schedule."}, {"speaker": "protocol custodian", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}, {"speaker": "data reviewer", "text": "The finalized sensor export contains 181 of the 200 required daily soil-moisture readings. Each retained reading has a recognized tray identifier and a trial-day entry."}, {"speaker": "endpoint reviewer", "text": "The endpoint register has both a final height measurement and a final biomass measurement for each of the 20 trays; no required endpoint field is blank."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_documented_crossover"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the trial requirements and rely on the unchanged questions object for the governing decision rubric. They preserve the same bean-tray trial, treatment assignment, 10-day period, measurements, and decision scope. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes the watering record from a documented deviation involving B17 to full adherence, without conflicting with the unchanged assignment, sensor count, or endpoint record. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"field coordinator\",\"text\":\"Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily.\"},{\"speaker\":\"records technician\",\"text\":\"Watering logs for all 10 days record that trays B01–B10 and B17 received the drought-group schedule, while trays B11–B16 and B18–B20 received the control-group schedule.\"},{\"speaker\":\"protocol custodian\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"},{\"speaker\":\"data reviewer\",\"text\":\"The finalized sensor export contains 181 of the 200 required daily soil-moisture readings. Each retained reading has a recognized tray identifier and a trial-day entry.\"},{\"speaker\":\"endpoint reviewer\",\"text\":\"The endpoint register has both a final height measurement and a final biomass measurement for each of the 20 trays; no required endpoint field is blank.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily."}, {"path": ["1", "text"], "text": "Watering logs for all 10 days record that trays B01–B10 and B17 received the drought-group schedule, while trays B11–B16 and B18–B20 received the control-group schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily.", "negative_left": "Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily.", "negative_right": "Watering logs for all 10 days record that trays B01–B10 received the drought-group schedule, while trays B11–B20 received the control-group schedule.", "right": "Watering logs for all 10 days record that trays B01–B10 and B17 received the drought-group schedule, while trays B11–B16 and B18–B20 received the control-group schedule."}, "verifier_independent_model": false}, "family": "scale-diverse-242-004", "id": "scale-diverse-242-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "field coordinator", "text": "Before watering began in the 10-day bean-tray trial, the documented randomization assigned trays B01–B10 to the drought group and trays B11–B20 to the control group; the drought-group schedule was 18 milliliters on days 1 and 6 only, and the control-group schedule was 18 milliliters daily."}, {"speaker": "records technician", "text": "Watering logs for all 10 days record that trays B01–B10 received the drought-group schedule, while trays B11–B20 received the control-group schedule."}, {"speaker": "protocol custodian", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}, {"speaker": "data reviewer", "text": "The finalized sensor export contains 181 of the 200 required daily soil-moisture readings. Each retained reading has a recognized tray identifier and a trial-day entry."}, {"speaker": "endpoint reviewer", "text": "The endpoint register has both a final height measurement and a final biomass measurement for each of the 20 trays; no required endpoint field is blank."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_at_threshold"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same conditional assignment policy and the same tray, calibration, timing, and protocol-intent question bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes the calibration outcome: B-1 through B-3 passed and exactly one of the four governed sensors failed, necessarily B-4, without contradicting unchanged facts. Neither context contains an answer code, proposition ID, rule table, classifier instruction, or explicit gold-answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4. Mira then documented the assignment policy: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The calibration was completed at 08:00, before either tray received water. The individual entries for B-1, B-2, and B-3 each show a passing result. Mira's completed 08:00 pre-watering calibration log records that every governed soil-moisture sensor passed. Afterward, the greenhouse team prepared the trays for treatment. The reviewer is checking whether tray B was protocol-intended for drought treatment under Mira's documented conditional assignment, rather than evaluating later watering practices or growth measurements.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4."}, {"path": [], "text": "Mira's completed 08:00 pre-watering calibration log records that every governed soil-moisture sensor passed."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4.", "negative_left": "Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4.", "negative_right": "Mira's completed 08:00 pre-watering calibration log records that B-1, B-2, and B-3 passed and exactly one governed soil-moisture sensor failed.", "right": "Mira's completed 08:00 pre-watering calibration log records that every governed soil-moisture sensor passed."}, "verifier_independent_model": false}, "family": "scale-diverse-243-001", "id": "scale-diverse-243-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4. Mira then documented the assignment policy: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The calibration was completed at 08:00, before either tray received water. The individual entries for B-1, B-2, and B-3 each show a passing result. Mira's completed 08:00 pre-watering calibration log records that every governed soil-moisture sensor passed. Afterward, the greenhouse team prepared the trays for treatment. The reviewer is checking whether tray B was protocol-intended for drought treatment under Mira's documented conditional assignment, rather than evaluating later watering practices or growth measurements."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same conditional assignment policy and the same tray, calibration, timing, and protocol-intent question bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes the calibration outcome: B-1 through B-3 passed and exactly one of the four governed sensors failed, necessarily B-4, without contradicting unchanged facts. Neither context contains an answer code, proposition ID, rule table, classifier instruction, or explicit gold-answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4. Mira then documented the assignment policy: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The calibration was completed at 08:00, before either tray received water. The individual entries for B-1, B-2, and B-3 each show a passing result. Mira's completed 08:00 pre-watering calibration log records that every governed soil-moisture sensor passed. Afterward, the greenhouse team prepared the trays for treatment. The reviewer is checking whether tray B was protocol-intended for drought treatment under Mira's documented conditional assignment, rather than evaluating later watering practices or growth measurements.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4."}, {"path": [], "text": "Mira's completed 08:00 pre-watering calibration log records that every governed soil-moisture sensor passed."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4.", "negative_left": "Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4.", "negative_right": "Mira's completed 08:00 pre-watering calibration log records that B-1, B-2, and B-3 passed and exactly one governed soil-moisture sensor failed.", "right": "Mira's completed 08:00 pre-watering calibration log records that every governed soil-moisture sensor passed."}, "verifier_independent_model": false}, "family": "scale-diverse-243-001", "id": "scale-diverse-243-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before watering began, Mira's calibration roster identified the complete set of soil-moisture sensors governed by her 08:00 calibration as exactly B-1, B-2, B-3, and B-4. Mira then documented the assignment policy: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The calibration was completed at 08:00, before either tray received water. The individual entries for B-1, B-2, and B-3 each show a passing result. Mira's completed 08:00 pre-watering calibration log records that B-1, B-2, and B-3 passed and exactly one governed soil-moisture sensor failed. Afterward, the greenhouse team prepared the trays for treatment. The reviewer is checking whether tray B was protocol-intended for drought treatment under Mira's documented conditional assignment, rather than evaluating later watering practices or growth measurements."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve Mira’s assignment rule, the pre-watering assessment scope, and the tray B/08:00 calibration bindings. The two evidence spans are complete factual sentences. The base coherently establishes four passes, while the counterfactual changes that count to three, consistent with B-1, B-2, and B-3 passing and without creating a contradictory measurement. Neither context includes an explicit answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"During a pretrial evidence reconciliation, the reviewer relied on Mira’s signed calibration register and assignment memo, both completed before watering began. Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result. Exactly four members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration. The assignment memo states: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” Subsequent watering actions and complete growth measurements were recorded separately as implementation history; the reviewer is assessing the documented assignment intent from the pre-watering calibration evidence and memo.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result."}, {"path": [], "text": "Exactly four members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result.", "negative_left": "Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result.", "negative_right": "Exactly three members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration.", "right": "Exactly four members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration."}, "verifier_independent_model": false}, "family": "scale-diverse-243-002", "id": "scale-diverse-243-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "During a pretrial evidence reconciliation, the reviewer relied on Mira’s signed calibration register and assignment memo, both completed before watering began. Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result. Exactly four members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration. The assignment memo states: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” Subsequent watering actions and complete growth measurements were recorded separately as implementation history; the reviewer is assessing the documented assignment intent from the pre-watering calibration evidence and memo."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve Mira’s assignment rule, the pre-watering assessment scope, and the tray B/08:00 calibration bindings. The two evidence spans are complete factual sentences. The base coherently establishes four passes, while the counterfactual changes that count to three, consistent with B-1, B-2, and B-3 passing and without creating a contradictory measurement. Neither context includes an explicit answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"During a pretrial evidence reconciliation, the reviewer relied on Mira’s signed calibration register and assignment memo, both completed before watering began. Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result. Exactly four members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration. The assignment memo states: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” Subsequent watering actions and complete growth measurements were recorded separately as implementation history; the reviewer is assessing the documented assignment intent from the pre-watering calibration evidence and memo.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result."}, {"path": [], "text": "Exactly four members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result.", "negative_left": "Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result.", "negative_right": "Exactly three members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration.", "right": "Exactly four members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration."}, "verifier_independent_model": false}, "family": "scale-diverse-243-002", "id": "scale-diverse-243-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "During a pretrial evidence reconciliation, the reviewer relied on Mira’s signed calibration register and assignment memo, both completed before watering began. Mira’s complete governed set for the 08:00 pre-watering calibration consisted exactly of sensors B-1, B-2, B-3, and B-4, with B-1, B-2, and B-3 each receiving a passing result. Exactly three members of Mira’s complete governed sensor set received passing results in the 08:00 pre-watering calibration. The assignment memo states: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” Subsequent watering actions and complete growth measurements were recorded separately as implementation history; the reviewer is assessing the documented assignment intent from the pre-watering calibration evidence and memo."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same protocol, update rule, request, trial entity, event identities, and dates. The two focus spans are complete factual sentences. The counterfactual changes only D-31’s delivered volume from 235 mL to 234 mL, without creating a duplicate or conflicting measurement; all unchanged evidence remains coherent. Neither constructed context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"An unsigned entry placed random assignment at 08:40. A later signed correction records baseline height at 09:00 and random assignment at 09:20 on 11 May; the controller audit records the assignment command at 09:20:14.\",\"The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL.\",\"Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 235 mL for event D-31, and 264 mL for event C-38.\",\"Schedule reconciliation identifies 14 May as the trial’s only scheduled moisture-sensor reading gap. Every scheduled reading absent because of that gap has a gravimetric backup, and the completed growth table contains every protocol-required final-height measurement.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "1"], "text": "The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL."}, {"path": ["evidence", "2"], "text": "Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 235 mL for event D-31, and 264 mL for event C-38."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL.", "negative_left": "The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL.", "negative_right": "Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 234 mL for event D-31, and 264 mL for event C-38.", "right": "Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 235 mL for event D-31, and 264 mL for event C-38."}, "verifier_independent_model": false}, "family": "scale-diverse-244-001", "id": "scale-diverse-244-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "evidence": ["An unsigned entry placed random assignment at 08:40. A later signed correction records baseline height at 09:00 and random assignment at 09:20 on 11 May; the controller audit records the assignment command at 09:20:14.", "The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL.", "Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 235 mL for event D-31, and 264 mL for event C-38.", "Schedule reconciliation identifies 14 May as the trial’s only scheduled moisture-sensor reading gap. Every scheduled reading absent because of that gap has a gravimetric backup, and the completed growth table contains every protocol-required final-height measurement."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same protocol, update rule, request, trial entity, event identities, and dates. The two focus spans are complete factual sentences. The counterfactual changes only D-31’s delivered volume from 235 mL to 234 mL, without creating a duplicate or conflicting measurement; all unchanged evidence remains coherent. Neither constructed context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"An unsigned entry placed random assignment at 08:40. A later signed correction records baseline height at 09:00 and random assignment at 09:20 on 11 May; the controller audit records the assignment command at 09:20:14.\",\"The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL.\",\"Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 235 mL for event D-31, and 264 mL for event C-38.\",\"Schedule reconciliation identifies 14 May as the trial’s only scheduled moisture-sensor reading gap. Every scheduled reading absent because of that gap has a gravimetric backup, and the completed growth table contains every protocol-required final-height measurement.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "1"], "text": "The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL."}, {"path": ["evidence", "2"], "text": "Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 235 mL for event D-31, and 264 mL for event C-38."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL.", "negative_left": "The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL.", "negative_right": "Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 234 mL for event D-31, and 264 mL for event C-38.", "right": "Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 235 mL for event D-31, and 264 mL for event C-38."}, "verifier_independent_model": false}, "family": "scale-diverse-244-001", "id": "scale-diverse-244-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "evidence": ["An unsigned entry placed random assignment at 08:40. A later signed correction records baseline height at 09:00 and random assignment at 09:20 on 11 May; the controller audit records the assignment command at 09:20:14.", "The complete set of drought-arm and control-arm watering events in the fictional tomato drought trial comprised drought-arm event D-17 on 12 May 2026 at 08:00 with a protocol-prescribed volume of 240 mL, control-arm event C-24 on 12 May 2026 at 08:15 with a protocol-prescribed volume of 260 mL, drought-arm event D-31 on 15 May 2026 at 08:00 with a protocol-prescribed volume of 245 mL, and control-arm event C-38 on 15 May 2026 at 08:15 with a protocol-prescribed volume of 255 mL.", "Calibrated flow-meter records show delivered volumes of 247 mL for event D-17, 252 mL for event C-24, 234 mL for event D-31, and 264 mL for event C-38.", "Schedule reconciliation identifies 14 May as the trial’s only scheduled moisture-sensor reading gap. Every scheduled reading absent because of that gap has a gravimetric backup, and the completed growth table contains every protocol-required final-height measurement."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the baseline-before-assignment, watering-tolerance, sensor-gap backup, and signed-correction policies, as well as the same trial, readiness question, and temporal scope. The two focus-evidence spans are complete factual sentences. The counterfactual changes only D-18’s delivered volume from 307 mL to 329 mL; this is coherent with the unchanged prescribed volume and does not create a duplicate or contradictory measurement. Neither context embeds an answer, output instruction, code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"A reconciliation note was prepared for the fictional tomato drought trial. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively.\",\"The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 307 mL for D-18, 269 mL for C-41, and 330 mL for C-42.\",\"An unsigned entry placed baseline height at 09:00 and assignment at 08:40 on 11 May. A later signed correction places assignment at 09:20 that day, supported by the controller audit; no correction changes the baseline time.\",\"The schedule and exception log identify exactly one moisture-sensor reading gap, on 14 May. Gravimetric backups cover every scheduled reading missing during that gap.\",\"The finalized growth table records every protocol-required final-height measurement.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively."}, {"path": ["evidence", "1"], "text": "The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 307 mL for D-18, 269 mL for C-41, and 330 mL for C-42."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively.", "negative_left": "The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively.", "negative_right": "The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 329 mL for D-18, 269 mL for C-41, and 330 mL for C-42.", "right": "The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 307 mL for D-18, 269 mL for C-41, and 330 mL for C-42."}, "verifier_independent_model": false}, "family": "scale-diverse-244-002", "id": "scale-diverse-244-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "A reconciliation note was prepared for the fictional tomato drought trial. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "evidence": ["The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively.", "The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 307 mL for D-18, 269 mL for C-41, and 330 mL for C-42.", "An unsigned entry placed baseline height at 09:00 and assignment at 08:40 on 11 May. A later signed correction places assignment at 09:20 that day, supported by the controller audit; no correction changes the baseline time.", "The schedule and exception log identify exactly one moisture-sensor reading gap, on 14 May. Gravimetric backups cover every scheduled reading missing during that gap.", "The finalized growth table records every protocol-required final-height measurement."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the baseline-before-assignment, watering-tolerance, sensor-gap backup, and signed-correction policies, as well as the same trial, readiness question, and temporal scope. The two focus-evidence spans are complete factual sentences. The counterfactual changes only D-18’s delivered volume from 307 mL to 329 mL; this is coherent with the unchanged prescribed volume and does not create a duplicate or contradictory measurement. Neither context embeds an answer, output instruction, code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"A reconciliation note was prepared for the fictional tomato drought trial. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively.\",\"The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 307 mL for D-18, 269 mL for C-41, and 330 mL for C-42.\",\"An unsigned entry placed baseline height at 09:00 and assignment at 08:40 on 11 May. A later signed correction places assignment at 09:20 that day, supported by the controller audit; no correction changes the baseline time.\",\"The schedule and exception log identify exactly one moisture-sensor reading gap, on 14 May. Gravimetric backups cover every scheduled reading missing during that gap.\",\"The finalized growth table records every protocol-required final-height measurement.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively."}, {"path": ["evidence", "1"], "text": "The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 307 mL for D-18, 269 mL for C-41, and 330 mL for C-42."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively.", "negative_left": "The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively.", "negative_right": "The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 329 mL for D-18, 269 mL for C-41, and 330 mL for C-42.", "right": "The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 307 mL for D-18, 269 mL for C-41, and 330 mL for C-42."}, "verifier_independent_model": false}, "family": "scale-diverse-244-002", "id": "scale-diverse-244-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "A reconciliation note was prepared for the fictional tomato drought trial. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "evidence": ["The fictional tomato drought trial's complete event register contains exactly four watering events: drought-arm events D-17 and D-18, with protocol-prescribed volumes of 240 mL and 315 mL respectively, and control-arm events C-41 and C-42, with protocol-prescribed volumes of 260 mL and 335 mL respectively.", "The instrument-audited delivery record gives delivered volumes of 247 mL for D-17, 329 mL for D-18, 269 mL for C-41, and 330 mL for C-42.", "An unsigned entry placed baseline height at 09:00 and assignment at 08:40 on 11 May. A later signed correction places assignment at 09:20 that day, supported by the controller audit; no correction changes the baseline time.", "The schedule and exception log identify exactly one moisture-sensor reading gap, on 14 May. Gravimetric backups cover every scheduled reading missing during that gap.", "The finalized growth table records every protocol-required final-height measurement."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the original governing protocol and update rule, while the unchanged questions object preserves the decision criteria and instructions. The request remains bound to the same fictional tomato drought trial, readiness decision, and relevant event and timing path. The two focus-evidence spans are complete factual sentences: one establishes the complete event set and prescribed volumes, and the other reports audited delivered volumes. The counterfactual changes only D17's delivered volume from 436 mL to 447 mL and does not create a duplicate or conflicting measurement elsewhere. Neither context includes a gold answer, output code, proposition identifier, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"Operational handoff notes show the signed baseline-height sheet was completed on 11 May at 09:00. An earlier unsigned entry placed random assignment at 08:40, but a later signed correction records assignment at 09:20, corroborated by the controller audit at 09:20:14.\",\"The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively.\",\"The audited delivery records list delivered volumes of 436 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18.\",\"Schedule reconciliation identified exactly one moisture-sensor reading gap during the trial, dated 14 May. The coverage ledger contains a gravimetric backup for every scheduled reading absent during that gap.\",\"The final growth table contains every final-height measurement required by the protocol. The data-quality reviewer verified the controller audit, backup weights, and growth table on 15 May.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "1"], "text": "The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively."}, {"path": ["evidence", "2"], "text": "The audited delivery records list delivered volumes of 436 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively.", "negative_left": "The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively.", "negative_right": "The audited delivery records list delivered volumes of 447 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18.", "right": "The audited delivery records list delivered volumes of 436 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18."}, "verifier_independent_model": false}, "family": "scale-diverse-244-003", "id": "scale-diverse-244-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "evidence": ["Operational handoff notes show the signed baseline-height sheet was completed on 11 May at 09:00. An earlier unsigned entry placed random assignment at 08:40, but a later signed correction records assignment at 09:20, corroborated by the controller audit at 09:20:14.", "The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively.", "The audited delivery records list delivered volumes of 436 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18.", "Schedule reconciliation identified exactly one moisture-sensor reading gap during the trial, dated 14 May. The coverage ledger contains a gravimetric backup for every scheduled reading absent during that gap.", "The final growth table contains every final-height measurement required by the protocol. The data-quality reviewer verified the controller audit, backup weights, and growth table on 15 May."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the original governing protocol and update rule, while the unchanged questions object preserves the decision criteria and instructions. The request remains bound to the same fictional tomato drought trial, readiness decision, and relevant event and timing path. The two focus-evidence spans are complete factual sentences: one establishes the complete event set and prescribed volumes, and the other reports audited delivered volumes. The counterfactual changes only D17's delivered volume from 436 mL to 447 mL and does not create a duplicate or conflicting measurement elsewhere. Neither context includes a gold answer, output code, proposition identifier, rule table, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"Operational handoff notes show the signed baseline-height sheet was completed on 11 May at 09:00. An earlier unsigned entry placed random assignment at 08:40, but a later signed correction records assignment at 09:20, corroborated by the controller audit at 09:20:14.\",\"The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively.\",\"The audited delivery records list delivered volumes of 436 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18.\",\"Schedule reconciliation identified exactly one moisture-sensor reading gap during the trial, dated 14 May. The coverage ledger contains a gravimetric backup for every scheduled reading absent during that gap.\",\"The final growth table contains every final-height measurement required by the protocol. The data-quality reviewer verified the controller audit, backup weights, and growth table on 15 May.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "1"], "text": "The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively."}, {"path": ["evidence", "2"], "text": "The audited delivery records list delivered volumes of 436 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively.", "negative_left": "The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively.", "negative_right": "The audited delivery records list delivered volumes of 447 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18.", "right": "The audited delivery records list delivered volumes of 436 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18."}, "verifier_independent_model": false}, "family": "scale-diverse-244-003", "id": "scale-diverse-244-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "evidence": ["Operational handoff notes show the signed baseline-height sheet was completed on 11 May at 09:00. An earlier unsigned entry placed random assignment at 08:40, but a later signed correction records assignment at 09:20, corroborated by the controller audit at 09:20:14.", "The complete set of watering events in the fictional tomato drought trial consists of drought-arm events D17 and D18, with protocol-prescribed volumes of 430 mL and 275 mL respectively, and control-arm events C17 and C18, with protocol-prescribed volumes of 510 mL and 365 mL respectively.", "The audited delivery records list delivered volumes of 447 mL for D17, 268 mL for D18, 502 mL for C17, and 374 mL for C18.", "Schedule reconciliation identified exactly one moisture-sensor reading gap during the trial, dated 14 May. The coverage ledger contains a gravimetric backup for every scheduled reading absent during that gap.", "The final growth table contains every final-height measurement required by the protocol. The data-quality reviewer verified the controller audit, backup weights, and growth table on 15 May."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original DR-18 treatment-control validity question and its governing scoring policy without adding exceptions, priorities, or missing-evidence defaults. The changed pot identifiers and assignment observations are permissible case-observation changes; the counterfactual introduces a single unresolved disagreement for P-42 between the two designated primary records, without contradicting the statement that no protocol deviation itself causes assignment ambiguity. The two focus-evidence spans are complete factual sentences. Neither context contains a gold score, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"records administrator\",\"text\":\"For DR-18, the signed bench map and the timestamped pre-treatment RFID assignment export are each designated a primary assignment record. No other DR-18 record is designated a primary assignment record.\"},{\"speaker\":\"submission clerk\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted.\"},{\"speaker\":\"analysis custodian\",\"text\":\"At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68.\"},{\"speaker\":\"assignment auditor\",\"text\":\"The DR-18 signed bench map dated 2026-03-02 and the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z each list P-17 as treatment, P-42 as control, and P-68 as treatment.\"},{\"speaker\":\"protocol monitor\",\"text\":\"A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18. No documented protocol deviation makes the assignment identity of any pot included in that analysis ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68."}, {"path": ["3", "text"], "text": "The DR-18 signed bench map dated 2026-03-02 and the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z each list P-17 as treatment, P-42 as control, and P-68 as treatment."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68.", "negative_left": "At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68.", "negative_right": "The DR-18 signed bench map dated 2026-03-02 lists P-17 as treatment, P-42 as control, and P-68 as treatment, while the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z lists P-17 as treatment, P-42 as treatment, and P-68 as treatment.", "right": "The DR-18 signed bench map dated 2026-03-02 and the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z each list P-17 as treatment, P-42 as control, and P-68 as treatment."}, "verifier_independent_model": false}, "family": "scale-diverse-245-001", "id": "scale-diverse-245-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "records administrator", "text": "For DR-18, the signed bench map and the timestamped pre-treatment RFID assignment export are each designated a primary assignment record. No other DR-18 record is designated a primary assignment record."}, {"speaker": "submission clerk", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"speaker": "analysis custodian", "text": "At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68."}, {"speaker": "assignment auditor", "text": "The DR-18 signed bench map dated 2026-03-02 and the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z each list P-17 as treatment, P-42 as control, and P-68 as treatment."}, {"speaker": "protocol monitor", "text": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18. No documented protocol deviation makes the assignment identity of any pot included in that analysis ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original DR-18 treatment-control validity question and its governing scoring policy without adding exceptions, priorities, or missing-evidence defaults. The changed pot identifiers and assignment observations are permissible case-observation changes; the counterfactual introduces a single unresolved disagreement for P-42 between the two designated primary records, without contradicting the statement that no protocol deviation itself causes assignment ambiguity. The two focus-evidence spans are complete factual sentences. Neither context contains a gold score, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"records administrator\",\"text\":\"For DR-18, the signed bench map and the timestamped pre-treatment RFID assignment export are each designated a primary assignment record. No other DR-18 record is designated a primary assignment record.\"},{\"speaker\":\"submission clerk\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted.\"},{\"speaker\":\"analysis custodian\",\"text\":\"At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68.\"},{\"speaker\":\"assignment auditor\",\"text\":\"The DR-18 signed bench map dated 2026-03-02 and the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z each list P-17 as treatment, P-42 as control, and P-68 as treatment.\"},{\"speaker\":\"protocol monitor\",\"text\":\"A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18. No documented protocol deviation makes the assignment identity of any pot included in that analysis ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68."}, {"path": ["3", "text"], "text": "The DR-18 signed bench map dated 2026-03-02 and the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z each list P-17 as treatment, P-42 as control, and P-68 as treatment."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68.", "negative_left": "At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68.", "negative_right": "The DR-18 signed bench map dated 2026-03-02 lists P-17 as treatment, P-42 as control, and P-68 as treatment, while the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z lists P-17 as treatment, P-42 as treatment, and P-68 as treatment.", "right": "The DR-18 signed bench map dated 2026-03-02 and the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z each list P-17 as treatment, P-42 as control, and P-68 as treatment."}, "verifier_independent_model": false}, "family": "scale-diverse-245-001", "id": "scale-diverse-245-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "records administrator", "text": "For DR-18, the signed bench map and the timestamped pre-treatment RFID assignment export are each designated a primary assignment record. No other DR-18 record is designated a primary assignment record."}, {"speaker": "submission clerk", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"speaker": "analysis custodian", "text": "At the 2026-04-19 analysis freeze for DR-18, the final growth analysis included exactly pots P-17, P-42, and P-68."}, {"speaker": "assignment auditor", "text": "The DR-18 signed bench map dated 2026-03-02 lists P-17 as treatment, P-42 as control, and P-68 as treatment, while the pre-treatment RFID assignment export timestamped 2026-03-03T07:14:00Z lists P-17 as treatment, P-42 as treatment, and P-68 as treatment."}, {"speaker": "protocol monitor", "text": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18. No documented protocol deviation makes the assignment identity of any pot included in that analysis ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric and request scope without adding exceptions, priorities, or missing-evidence defaults, and they concern the same 21-day, 24-seedling trial and treatment-control-validity evaluation. The two evidence spans are complete factual sentences. The counterfactual changes only TA-17's presence in the evidence package; this is coherent with treatment identities being recorded elsewhere and with every other critical assignment record being available. Neither context states a score, answer code, proposition ID, classifier instruction, or explicit label rationale; terms such as “critical” and “noncritical supporting detail” are permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"At intake, the trial register classified document TA-17 as a critical assignment record. At 16:40 UTC on 8 May 2026, the complete evidence package for the trial contained document TA-17. Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The reviewer then confirmed that treatment identities and outcome measurements were recorded for all 24 seedlings. Every other critical assignment record was available, as were every critical watering record, every critical sensor record, and every critical outcome record for the trial. During the final inventory, the original bench map could not be located and remained unavailable. The documentation index classifies that map as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail used for that evaluation was available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 16:40 UTC on 8 May 2026, the complete evidence package for the trial contained document TA-17."}, {"path": [], "text": "Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:40 UTC on 8 May 2026, the complete evidence package for the trial contained document TA-17.", "negative_left": "At 16:40 UTC on 8 May 2026, the complete evidence package for the trial did not contain document TA-17.", "negative_right": "Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings.", "right": "Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings."}, "verifier_independent_model": false}, "family": "scale-diverse-246-001", "id": "scale-diverse-246-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "At intake, the trial register classified document TA-17 as a critical assignment record. At 16:40 UTC on 8 May 2026, the complete evidence package for the trial contained document TA-17. Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The reviewer then confirmed that treatment identities and outcome measurements were recorded for all 24 seedlings. Every other critical assignment record was available, as were every critical watering record, every critical sensor record, and every critical outcome record for the trial. During the final inventory, the original bench map could not be located and remained unavailable. The documentation index classifies that map as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail used for that evaluation was available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric and request scope without adding exceptions, priorities, or missing-evidence defaults, and they concern the same 21-day, 24-seedling trial and treatment-control-validity evaluation. The two evidence spans are complete factual sentences. The counterfactual changes only TA-17's presence in the evidence package; this is coherent with treatment identities being recorded elsewhere and with every other critical assignment record being available. Neither context states a score, answer code, proposition ID, classifier instruction, or explicit label rationale; terms such as “critical” and “noncritical supporting detail” are permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"At intake, the trial register classified document TA-17 as a critical assignment record. At 16:40 UTC on 8 May 2026, the complete evidence package for the trial contained document TA-17. Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The reviewer then confirmed that treatment identities and outcome measurements were recorded for all 24 seedlings. Every other critical assignment record was available, as were every critical watering record, every critical sensor record, and every critical outcome record for the trial. During the final inventory, the original bench map could not be located and remained unavailable. The documentation index classifies that map as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail used for that evaluation was available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 16:40 UTC on 8 May 2026, the complete evidence package for the trial contained document TA-17."}, {"path": [], "text": "Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:40 UTC on 8 May 2026, the complete evidence package for the trial contained document TA-17.", "negative_left": "At 16:40 UTC on 8 May 2026, the complete evidence package for the trial did not contain document TA-17.", "negative_right": "Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings.", "right": "Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings."}, "verifier_independent_model": false}, "family": "scale-diverse-246-001", "id": "scale-diverse-246-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "At intake, the trial register classified document TA-17 as a critical assignment record. At 16:40 UTC on 8 May 2026, the complete evidence package for the trial did not contain document TA-17. Document TA-17 is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The reviewer then confirmed that treatment identities and outcome measurements were recorded for all 24 seedlings. Every other critical assignment record was available, as were every critical watering record, every critical sensor record, and every critical outcome record for the trial. During the final inventory, the original bench map could not be located and remained unavailable. The documentation index classifies that map as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail used for that evaluation was available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric and the trial, evidence-completeness, treatment-control-validity, and 21-day/24-seedling bindings while changing only case observations. The two evidence spans are complete factual sentences. Removing TS-742 from the explicitly complete package inventory coherently makes the uniquely identified treatment-assignment sheet unavailable; this does not conflict with the statement that every other critical assignment record is available. Neither context contains a score, answer code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"Evidence reconciliation for greenhouse trial records: At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers TS-742, WR-318, SR-509, and OR-266. For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742. Trial documentation classifies that sheet as a critical assignment record. The treatment identity and outcome measurements for each of the 24 seedlings are recorded. Every other critical assignment record is available. All critical watering records, critical sensor records, and critical outcome records are also available. The original bench map is unavailable. Documentation classifies that map as a noncritical supporting detail for evaluating treatment-control validity, and every other supporting detail for that evaluation is available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers TS-742, WR-318, SR-509, and OR-266."}, {"path": [], "text": "For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers TS-742, WR-318, SR-509, and OR-266.", "negative_left": "At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers WR-318, SR-509, and OR-266.", "negative_right": "For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742.", "right": "For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742."}, "verifier_independent_model": false}, "family": "scale-diverse-246-002", "id": "scale-diverse-246-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "Evidence reconciliation for greenhouse trial records: At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers TS-742, WR-318, SR-509, and OR-266. For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742. Trial documentation classifies that sheet as a critical assignment record. The treatment identity and outcome measurements for each of the 24 seedlings are recorded. Every other critical assignment record is available. All critical watering records, critical sensor records, and critical outcome records are also available. The original bench map is unavailable. Documentation classifies that map as a noncritical supporting detail for evaluating treatment-control validity, and every other supporting detail for that evaluation is available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric and the trial, evidence-completeness, treatment-control-validity, and 21-day/24-seedling bindings while changing only case observations. The two evidence spans are complete factual sentences. Removing TS-742 from the explicitly complete package inventory coherently makes the uniquely identified treatment-assignment sheet unavailable; this does not conflict with the statement that every other critical assignment record is available. Neither context contains a score, answer code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"Evidence reconciliation for greenhouse trial records: At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers TS-742, WR-318, SR-509, and OR-266. For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742. Trial documentation classifies that sheet as a critical assignment record. The treatment identity and outcome measurements for each of the 24 seedlings are recorded. Every other critical assignment record is available. All critical watering records, critical sensor records, and critical outcome records are also available. The original bench map is unavailable. Documentation classifies that map as a noncritical supporting detail for evaluating treatment-control validity, and every other supporting detail for that evaluation is available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers TS-742, WR-318, SR-509, and OR-266."}, {"path": [], "text": "For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers TS-742, WR-318, SR-509, and OR-266.", "negative_left": "At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers WR-318, SR-509, and OR-266.", "negative_right": "For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742.", "right": "For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742."}, "verifier_independent_model": false}, "family": "scale-diverse-246-002", "id": "scale-diverse-246-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "Evidence reconciliation for greenhouse trial records: At 16:00 UTC on 12 May 2026, the complete inventory of evidence package EP-47 listed exactly the records with accession identifiers WR-318, SR-509, and OR-266. For the 21-day trial of 24 tomato seedlings, the treatment-assignment sheet was uniquely designated by accession identifier TS-742. Trial documentation classifies that sheet as a critical assignment record. The treatment identity and outcome measurements for each of the 24 seedlings are recorded. Every other critical assignment record is available. All critical watering records, critical sensor records, and critical outcome records are also available. The original bench map is unavailable. Documentation classifies that map as a noncritical supporting detail for evaluating treatment-control validity, and every other supporting detail for that evaluation is available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring policy and request, while both contexts retain the same 21-day, 24-seedling trial and treatment-control-validity evaluation. The two evidence spans are complete factual sentences. The counterfactual coherently changes R-417 from the treatment-assignment sheet to a height worksheet, making the assignment sheet unavailable without conflicting with the package count, archive availability statements, or other unchanged facts. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592. In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings, and R-592 identifies its irrigation log. The handoff inventory classifies the treatment-assignment sheet as a critical assignment record. Apart from that sheet, every critical assignment record is available through the linked trial archive. The archive records each seedling’s treatment identity and outcome measurements. Every critical watering, sensor, and outcome record is also available. The original bench map is unavailable; the custodian classifies it as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail for that evaluation is available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592."}, {"path": [], "text": "In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings, and R-592 identifies its irrigation log."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592.", "negative_left": "At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592.", "negative_right": "In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the seedling-height worksheet, R-592 identifies the irrigation log, and neither record is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings.", "right": "In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings, and R-592 identifies its irrigation log."}, "verifier_independent_model": false}, "family": "scale-diverse-246-003", "id": "scale-diverse-246-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592. In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings, and R-592 identifies its irrigation log. The handoff inventory classifies the treatment-assignment sheet as a critical assignment record. Apart from that sheet, every critical assignment record is available through the linked trial archive. The archive records each seedling’s treatment identity and outcome measurements. Every critical watering, sensor, and outcome record is also available. The original bench map is unavailable; the custodian classifies it as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail for that evaluation is available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring policy and request, while both contexts retain the same 21-day, 24-seedling trial and treatment-control-validity evaluation. The two evidence spans are complete factual sentences. The counterfactual coherently changes R-417 from the treatment-assignment sheet to a height worksheet, making the assignment sheet unavailable without conflicting with the package count, archive availability statements, or other unchanged facts. Neither context contains a score, gold answer, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592. In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings, and R-592 identifies its irrigation log. The handoff inventory classifies the treatment-assignment sheet as a critical assignment record. Apart from that sheet, every critical assignment record is available through the linked trial archive. The archive records each seedling’s treatment identity and outcome measurements. Every critical watering, sensor, and outcome record is also available. The original bench map is unavailable; the custodian classifies it as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail for that evaluation is available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592."}, {"path": [], "text": "In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings, and R-592 identifies its irrigation log."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592.", "negative_left": "At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592.", "negative_right": "In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the seedling-height worksheet, R-592 identifies the irrigation log, and neither record is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings.", "right": "In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings, and R-592 identifies its irrigation log."}, "verifier_independent_model": false}, "family": "scale-diverse-246-003", "id": "scale-diverse-246-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "At the 2026-09-16 14:35 UTC operational handoff, the sealed evidence package for the 21-day trial of 24 tomato seedlings contained exactly the records bearing identifiers R-417 and R-592. In the trial archive index finalized at 2026-09-16 14:20 UTC, R-417 identifies the seedling-height worksheet, R-592 identifies the irrigation log, and neither record is the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The handoff inventory classifies the treatment-assignment sheet as a critical assignment record. Apart from that sheet, every critical assignment record is available through the linked trial archive. The archive records each seedling’s treatment identity and outcome measurements. Every critical watering, sensor, and outcome record is also available. The original bench map is unavailable; the custodian classifies it as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail for that evaluation is available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the scoring policy, instructions, and criteria, while both contexts remain bound to the same 21-day, 24-seedling treatment-assignment evaluation. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently removes TS-17—the registered treatment-assignment sheet—from the inventoried package without conflicting with the statement that every other critical assignment record is available. Neither context includes a score, answer code, rule table, proposition identifier, output instruction, or explicit gold answer; its critical/noncritical terminology is permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"Field note, 6 August 2026. The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-24 without gaps. The sheet is designated a critical assignment record for this trial. Treatment identities for all 24 seedlings are recorded, as are their outcome measurements. Every critical assignment record other than the treatment-assignment sheet is available. Every critical watering record, every critical sensor record, and every critical outcome record is available. The original bench map is unavailable. For evaluating treatment-control validity, the map is classified as a noncritical supporting detail. Every supporting detail for that evaluation other than the original bench map is available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings."}, {"path": [], "text": "The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-24 without gaps."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings.", "negative_left": "The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings.", "negative_right": "The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-16 and TS-18 through TS-24, with no other exhibits present.", "right": "The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-24 without gaps."}, "verifier_independent_model": false}, "family": "scale-diverse-246-004", "id": "scale-diverse-246-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "Field note, 6 August 2026. The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-24 without gaps. The sheet is designated a critical assignment record for this trial. Treatment identities for all 24 seedlings are recorded, as are their outcome measurements. Every critical assignment record other than the treatment-assignment sheet is available. Every critical watering record, every critical sensor record, and every critical outcome record is available. The original bench map is unavailable. For evaluating treatment-control validity, the map is classified as a noncritical supporting detail. Every supporting detail for that evaluation other than the original bench map is available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the scoring policy, instructions, and criteria, while both contexts remain bound to the same 21-day, 24-seedling treatment-assignment evaluation. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently removes TS-17—the registered treatment-assignment sheet—from the inventoried package without conflicting with the statement that every other critical assignment record is available. Neither context includes a score, answer code, rule table, proposition identifier, output instruction, or explicit gold answer; its critical/noncritical terminology is permissible natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"Field note, 6 August 2026. The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-24 without gaps. The sheet is designated a critical assignment record for this trial. Treatment identities for all 24 seedlings are recorded, as are their outcome measurements. Every critical assignment record other than the treatment-assignment sheet is available. Every critical watering record, every critical sensor record, and every critical outcome record is available. The original bench map is unavailable. For evaluating treatment-control validity, the map is classified as a noncritical supporting detail. Every supporting detail for that evaluation other than the original bench map is available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings."}, {"path": [], "text": "The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-24 without gaps."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings.", "negative_left": "The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings.", "negative_right": "The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-16 and TS-18 through TS-24, with no other exhibits present.", "right": "The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-24 without gaps."}, "verifier_independent_model": false}, "family": "scale-diverse-246-004", "id": "scale-diverse-246-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "Field note, 6 August 2026. The records register assigns exhibit code TS-17 to the treatment-assignment sheet for the 21-day trial of 24 tomato seedlings. The complete evidence package inventoried at 14:35 on 6 August 2026 contained exhibits TS-01 through TS-16 and TS-18 through TS-24, with no other exhibits present. The sheet is designated a critical assignment record for this trial. Treatment identities for all 24 seedlings are recorded, as are their outcome measurements. Every critical assignment record other than the treatment-assignment sheet is available. Every critical watering record, every critical sensor record, and every critical outcome record is available. The original bench map is unavailable. For evaluating treatment-control validity, the map is classified as a noncritical supporting detail. Every supporting detail for that evaluation other than the original bench map is available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original completeness, routing, precedence, and confidence policy without adding an exception or changed priority. They preserve the decision scope and Survey D247/site and voucher bindings; changing the effort and S88 observation is permitted. The two evidence spans are complete factual sentences. The counterfactual makes S88 agree with C19 and does not contradict the unchanged count of two distinct operative entries or the auditor's statements. Neither context states a choice, answer code, rule table, proposition ID, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universal claims such as every lot having a count remain atomic. The focus atom concerns a factual post-precedence inconsistency, not a policy conclusion. The base and counter assignments are realizable with only a10 changing: an explicit unresolved status can coexist either with an additional post-precedence conflicting status or with no such conflict, while other fields remain consistent. Empty policy_evidence is correct because the governing completeness, routing, precedence, and confidence rules are already retained in the questions object; the state contributes case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail completeness: effort, exact site-label match, numeric counts for every lot, and all four habitat observations are supported. The post-precedence inconsistency in voucher V4's species-resolution status is a remaining contradiction affecting whether Identification routing applies, so Data Review takes priority. Because that needed routing fact conflicts, Low confidence follows.", "rule_index": 0, "sound": true}, {"reason": "The conditions entail completeness. Refutation of a10 entails no post-precedence inconsistency in V4's species-resolution status, while a11 excludes contradictions elsewhere. The explicitly unresolved voucher therefore requires Identification. All routing facts are explicit, and the consistency conditions exclude conflicts, so High confidence follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Survey D247's recorded effort is at least 30 minutes."}, {"id": "a2", "statement": "Survey D247's field-sheet site label exactly matches Survey D247's registered site label."}, {"id": "a3", "statement": "Every specimen lot collected in Survey D247 has a numeric count."}, {"id": "a4", "statement": "Survey D247 has a recorded flow observation."}, {"id": "a5", "statement": "Survey D247 has a recorded substrate observation."}, {"id": "a6", "statement": "Survey D247 has a recorded canopy-shade observation."}, {"id": "a7", "statement": "Survey D247 has a recorded wetted-width observation."}, {"id": "a8", "statement": "Voucher V4 in Survey D247 is explicitly unresolved at species level."}, {"id": "a9", "statement": "Every fact needed to route Survey D247 is explicitly recorded rather than inferred."}, {"id": "a10", "statement": "After identification-correction precedence is applied, Survey D247's complete record assigns inconsistent species-resolution statuses to voucher V4."}, {"id": "a11", "statement": "After identification-correction precedence is applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent."}], "base_state_json": "[{\"speaker\":\"field operations clerk\",\"text\":\"The final field sheet for Survey D247 records 34 minutes of effort at site ST-08, and the registered site label is also exactly ST-08. Each collected specimen lot has a numeric count.\"},{\"speaker\":\"habitat records clerk\",\"text\":\"Survey D247 directly records observations for flow, substrate, canopy shade, and wetted width. The reconciliation audit confirms that every fact required for routing is explicitly recorded, with none inferred.\"},{\"speaker\":\"records archivist\",\"text\":\"At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88.\"},{\"speaker\":\"status-table custodian\",\"text\":\"Survey D247's status-table entry S88 recorded voucher V4 as resolved at species level.\"},{\"speaker\":\"reconciliation auditor\",\"text\":\"The post-precedence comparison found every field in Survey D247 other than voucher V4's species-resolution status internally consistent. No omitted amendments, duplicate operative entries, or conflicts affect the effort, site, lot counts, habitat observations, or any other record field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["2", "text"], "text": "At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88."}, {"path": ["3", "text"], "text": "Survey D247's status-table entry S88 recorded voucher V4 as resolved at species level."}], "policy_evidence": [], "rules": [{"justification": "All completeness requirements are satisfied. After correction precedence, voucher V4 has conflicting species-resolution statuses, so a contradiction remains in a fact needed to decide whether the survey routes to Identification; the required outcome is therefore Data Review with Low confidence.", "target": "data_review_yes_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All completeness requirements are satisfied. Voucher V4 is explicitly unresolved, its post-precedence species-resolution status is not inconsistent, every other record field is consistent, and every routing fact is explicit; the required outcome is therefore Identification with High confidence.", "target": "identification_yes_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88.", "negative_left": "At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88.", "negative_right": "Survey D247's status-table entry S88 recorded voucher V4 as unresolved at species level.", "right": "Survey D247's status-table entry S88 recorded voucher V4 as resolved at species level."}, "verifier_independent_model": false}, "family": "scale-diverse-247-002", "id": "scale-diverse-247-002-base", "input": {"questions": {"decision": {"criteria": {"data_review_yes_low": "Completeness Yes; Data-review queue; Low confidence. Choose only if a record contradiction remains after correction precedence is applied and a needed routing fact conflicts.", "identification_yes_high": "Completeness Yes; Identification queue; High confidence. Choose only if the complete record has no remaining contradiction, at least one taxon is explicitly unresolved, and all routing facts are direct and consistent.", "sampling_no_medium": "Completeness No; Sampling queue; Medium confidence. Choose only if at least one required completeness element is absent and that conclusion requires one inference.", "validated_yes_medium": "Completeness Yes; Validated; Medium confidence. Choose only if no contradiction or unresolved taxon remains and one required routing fact must be inferred."}, "instructions": "Select the combined completeness judgment, queue, and confidence rating. Completeness is Yes only when effort is at least 30 minutes, the site label exactly matches the register, every specimen lot has a numeric count, and all four habitat observations are present. Route incomplete surveys to Sampling; otherwise route remaining record contradictions to Data Review; otherwise route any unresolved taxon to Identification; otherwise mark Validated. A later explicit statement that an identification is not confirmed controls over its quoted tentative note and is not itself a contradiction. Confidence is ordered High > Medium > Low: High requires all routing facts to be explicit and consistent; Medium requires one inference; Low means a needed fact conflicts.", "type": "choice"}}, "state": [{"speaker": "field operations clerk", "text": "The final field sheet for Survey D247 records 34 minutes of effort at site ST-08, and the registered site label is also exactly ST-08. Each collected specimen lot has a numeric count."}, {"speaker": "habitat records clerk", "text": "Survey D247 directly records observations for flow, substrate, canopy shade, and wetted width. The reconciliation audit confirms that every fact required for routing is explicitly recorded, with none inferred."}, {"speaker": "records archivist", "text": "At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88."}, {"speaker": "status-table custodian", "text": "Survey D247's status-table entry S88 recorded voucher V4 as resolved at species level."}, {"speaker": "reconciliation auditor", "text": "The post-precedence comparison found every field in Survey D247 other than voucher V4's species-resolution status internally consistent. No omitted amendments, duplicate operative entries, or conflicts affect the effort, site, lot counts, habitat observations, or any other record field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-247", "source_is_synthetic": true, "source_sha256": "9d6d41d1934faa1fee1cda0be4e6d613633e53a4f44a8bad39c426cd84cb0f2d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "data_review_yes_low"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original completeness, routing, precedence, and confidence policy without adding an exception or changed priority. They preserve the decision scope and Survey D247/site and voucher bindings; changing the effort and S88 observation is permitted. The two evidence spans are complete factual sentences. The counterfactual makes S88 agree with C19 and does not contradict the unchanged count of two distinct operative entries or the auditor's statements. Neither context states a choice, answer code, rule table, proposition ID, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universal claims such as every lot having a count remain atomic. The focus atom concerns a factual post-precedence inconsistency, not a policy conclusion. The base and counter assignments are realizable with only a10 changing: an explicit unresolved status can coexist either with an additional post-precedence conflicting status or with no such conflict, while other fields remain consistent. Empty policy_evidence is correct because the governing completeness, routing, precedence, and confidence rules are already retained in the questions object; the state contributes case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail completeness: effort, exact site-label match, numeric counts for every lot, and all four habitat observations are supported. The post-precedence inconsistency in voucher V4's species-resolution status is a remaining contradiction affecting whether Identification routing applies, so Data Review takes priority. Because that needed routing fact conflicts, Low confidence follows.", "rule_index": 0, "sound": true}, {"reason": "The conditions entail completeness. Refutation of a10 entails no post-precedence inconsistency in V4's species-resolution status, while a11 excludes contradictions elsewhere. The explicitly unresolved voucher therefore requires Identification. All routing facts are explicit, and the consistency conditions exclude conflicts, so High confidence follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Survey D247's recorded effort is at least 30 minutes."}, {"id": "a2", "statement": "Survey D247's field-sheet site label exactly matches Survey D247's registered site label."}, {"id": "a3", "statement": "Every specimen lot collected in Survey D247 has a numeric count."}, {"id": "a4", "statement": "Survey D247 has a recorded flow observation."}, {"id": "a5", "statement": "Survey D247 has a recorded substrate observation."}, {"id": "a6", "statement": "Survey D247 has a recorded canopy-shade observation."}, {"id": "a7", "statement": "Survey D247 has a recorded wetted-width observation."}, {"id": "a8", "statement": "Voucher V4 in Survey D247 is explicitly unresolved at species level."}, {"id": "a9", "statement": "Every fact needed to route Survey D247 is explicitly recorded rather than inferred."}, {"id": "a10", "statement": "After identification-correction precedence is applied, Survey D247's complete record assigns inconsistent species-resolution statuses to voucher V4."}, {"id": "a11", "statement": "After identification-correction precedence is applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent."}], "base_state_json": "[{\"speaker\":\"field operations clerk\",\"text\":\"The final field sheet for Survey D247 records 34 minutes of effort at site ST-08, and the registered site label is also exactly ST-08. Each collected specimen lot has a numeric count.\"},{\"speaker\":\"habitat records clerk\",\"text\":\"Survey D247 directly records observations for flow, substrate, canopy shade, and wetted width. The reconciliation audit confirms that every fact required for routing is explicitly recorded, with none inferred.\"},{\"speaker\":\"records archivist\",\"text\":\"At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88.\"},{\"speaker\":\"status-table custodian\",\"text\":\"Survey D247's status-table entry S88 recorded voucher V4 as resolved at species level.\"},{\"speaker\":\"reconciliation auditor\",\"text\":\"The post-precedence comparison found every field in Survey D247 other than voucher V4's species-resolution status internally consistent. No omitted amendments, duplicate operative entries, or conflicts affect the effort, site, lot counts, habitat observations, or any other record field.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["2", "text"], "text": "At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88."}, {"path": ["3", "text"], "text": "Survey D247's status-table entry S88 recorded voucher V4 as resolved at species level."}], "policy_evidence": [], "rules": [{"justification": "All completeness requirements are satisfied. After correction precedence, voucher V4 has conflicting species-resolution statuses, so a contradiction remains in a fact needed to decide whether the survey routes to Identification; the required outcome is therefore Data Review with Low confidence.", "target": "data_review_yes_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All completeness requirements are satisfied. Voucher V4 is explicitly unresolved, its post-precedence species-resolution status is not inconsistent, every other record field is consistent, and every routing fact is explicit; the required outcome is therefore Identification with High confidence.", "target": "identification_yes_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88.", "negative_left": "At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88.", "negative_right": "Survey D247's status-table entry S88 recorded voucher V4 as unresolved at species level.", "right": "Survey D247's status-table entry S88 recorded voucher V4 as resolved at species level."}, "verifier_independent_model": false}, "family": "scale-diverse-247-002", "id": "scale-diverse-247-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_yes_low": "Completeness Yes; Data-review queue; Low confidence. Choose only if a record contradiction remains after correction precedence is applied and a needed routing fact conflicts.", "identification_yes_high": "Completeness Yes; Identification queue; High confidence. Choose only if the complete record has no remaining contradiction, at least one taxon is explicitly unresolved, and all routing facts are direct and consistent.", "sampling_no_medium": "Completeness No; Sampling queue; Medium confidence. Choose only if at least one required completeness element is absent and that conclusion requires one inference.", "validated_yes_medium": "Completeness Yes; Validated; Medium confidence. Choose only if no contradiction or unresolved taxon remains and one required routing fact must be inferred."}, "instructions": "Select the combined completeness judgment, queue, and confidence rating. Completeness is Yes only when effort is at least 30 minutes, the site label exactly matches the register, every specimen lot has a numeric count, and all four habitat observations are present. Route incomplete surveys to Sampling; otherwise route remaining record contradictions to Data Review; otherwise route any unresolved taxon to Identification; otherwise mark Validated. A later explicit statement that an identification is not confirmed controls over its quoted tentative note and is not itself a contradiction. Confidence is ordered High > Medium > Low: High requires all routing facts to be explicit and consistent; Medium requires one inference; Low means a needed fact conflicts.", "type": "choice"}}, "state": [{"speaker": "field operations clerk", "text": "The final field sheet for Survey D247 records 34 minutes of effort at site ST-08, and the registered site label is also exactly ST-08. Each collected specimen lot has a numeric count."}, {"speaker": "habitat records clerk", "text": "Survey D247 directly records observations for flow, substrate, canopy shade, and wetted width. The reconciliation audit confirms that every fact required for routing is explicitly recorded, with none inferred."}, {"speaker": "records archivist", "text": "At 16:40 UTC on 12 August 2026, Survey D247's frozen complete record contained exactly two operative species-resolution entries for voucher V4: correction C19, which was marked as taking precedence over every earlier identification note and recorded V4 as unresolved at species level, and status-table entry S88."}, {"speaker": "status-table custodian", "text": "Survey D247's status-table entry S88 recorded voucher V4 as unresolved at species level."}, {"speaker": "reconciliation auditor", "text": "The post-precedence comparison found every field in Survey D247 other than voucher V4's species-resolution status internally consistent. No omitted amendments, duplicate operative entries, or conflicts affect the effort, site, lot counts, habitat observations, or any other record field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-247", "source_is_synthetic": true, "source_sha256": "9d6d41d1934faa1fee1cda0be4e6d613633e53a4f44a8bad39c426cd84cb0f2d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_yes_high"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing decision criteria and instructions, while neither context alters policy or introduces a new rule. Survey D247, voucher V4, site ST-08, dates, times, and relevant record paths remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes E17 from confirmed to unresolved, making it consistent with E23 without conflicting with the unchanged audit facts. Neither context contains an answer choice, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universal claims such as every lot having a count remain atomic. The focus atom concerns a factual post-precedence inconsistency, not a policy conclusion. The base and counter assignments are realizable with only a10 changing: an explicit unresolved status can coexist either with an additional post-precedence conflicting status or with no such conflict, while other fields remain consistent. Empty policy_evidence is correct because the governing completeness, routing, precedence, and confidence rules are already retained in the questions object; the state contributes case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail completeness: effort, exact site-label match, numeric counts for every lot, and all four habitat observations are supported. The post-precedence inconsistency in voucher V4's species-resolution status is a remaining contradiction affecting whether Identification routing applies, so Data Review takes priority. Because that needed routing fact conflicts, Low confidence follows.", "rule_index": 0, "sound": true}, {"reason": "The conditions entail completeness. Refutation of a10 entails no post-precedence inconsistency in V4's species-resolution status, while a11 excludes contradictions elsewhere. The explicitly unresolved voucher therefore requires Identification. All routing facts are explicit, and the consistency conditions exclude conflicts, so High confidence follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Survey D247's recorded effort is at least 30 minutes."}, {"id": "a2", "statement": "Survey D247's field-sheet site label exactly matches Survey D247's registered site label."}, {"id": "a3", "statement": "Every specimen lot collected in Survey D247 has a numeric count."}, {"id": "a4", "statement": "Survey D247 has a recorded flow observation."}, {"id": "a5", "statement": "Survey D247 has a recorded substrate observation."}, {"id": "a6", "statement": "Survey D247 has a recorded canopy-shade observation."}, {"id": "a7", "statement": "Survey D247 has a recorded wetted-width observation."}, {"id": "a8", "statement": "Voucher V4 in Survey D247 is explicitly unresolved at species level."}, {"id": "a9", "statement": "Every fact needed to route Survey D247 is explicitly recorded rather than inferred."}, {"id": "a10", "statement": "After identification-correction precedence is applied, Survey D247's complete record assigns inconsistent species-resolution statuses to voucher V4."}, {"id": "a11", "statement": "After identification-correction precedence is applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent."}], "base_state_json": "[{\"speaker\":\"field recorder\",\"text\":\"Survey D247 ended at 08:40 UTC on 18 April 2026 after exactly 30 minutes of recorded effort. The field-sheet site label and registered site label were both exactly ST-08.\"},{\"speaker\":\"records officer\",\"text\":\"Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as confirmed, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4.\"},{\"speaker\":\"identification recorder\",\"text\":\"Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved.\"},{\"speaker\":\"audit clerk\",\"text\":\"The survey contained three specimen lots, and each had a numeric count: 12, 7, and 6. Flow, substrate, canopy shade, and wetted width observations were recorded. The 11:30 UTC audit was conducted after identification-correction precedence had been applied. It documented every fact needed for routing directly rather than by inference and found every record field other than V4's species-resolution status internally consistent.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["1", "text"], "text": "Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as confirmed, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4."}, {"path": ["2", "text"], "text": "Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved."}], "policy_evidence": [], "rules": [{"justification": "All completeness requirements are satisfied. After correction precedence, voucher V4 has conflicting species-resolution statuses, so a contradiction remains in a fact needed to decide whether the survey routes to Identification; the required outcome is therefore Data Review with Low confidence.", "target": "data_review_yes_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All completeness requirements are satisfied. Voucher V4 is explicitly unresolved, its post-precedence species-resolution status is not inconsistent, every other record field is consistent, and every routing fact is explicit; the required outcome is therefore Identification with High confidence.", "target": "identification_yes_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as confirmed, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4.", "negative_left": "Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as unresolved, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4.", "negative_right": "Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved.", "right": "Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved."}, "verifier_independent_model": false}, "family": "scale-diverse-247-005", "id": "scale-diverse-247-005-base", "input": {"questions": {"decision": {"criteria": {"data_review_yes_low": "Completeness Yes; Data-review queue; Low confidence. Choose only if a record contradiction remains after correction precedence is applied and a needed routing fact conflicts.", "identification_yes_high": "Completeness Yes; Identification queue; High confidence. Choose only if the complete record has no remaining contradiction, at least one taxon is explicitly unresolved, and all routing facts are direct and consistent.", "sampling_no_medium": "Completeness No; Sampling queue; Medium confidence. Choose only if at least one required completeness element is absent and that conclusion requires one inference.", "validated_yes_medium": "Completeness Yes; Validated; Medium confidence. Choose only if no contradiction or unresolved taxon remains and one required routing fact must be inferred."}, "instructions": "Select the combined completeness judgment, queue, and confidence rating. Completeness is Yes only when effort is at least 30 minutes, the site label exactly matches the register, every specimen lot has a numeric count, and all four habitat observations are present. Route incomplete surveys to Sampling; otherwise route remaining record contradictions to Data Review; otherwise route any unresolved taxon to Identification; otherwise mark Validated. A later explicit statement that an identification is not confirmed controls over its quoted tentative note and is not itself a contradiction. Confidence is ordered High > Medium > Low: High requires all routing facts to be explicit and consistent; Medium requires one inference; Low means a needed fact conflicts.", "type": "choice"}}, "state": [{"speaker": "field recorder", "text": "Survey D247 ended at 08:40 UTC on 18 April 2026 after exactly 30 minutes of recorded effort. The field-sheet site label and registered site label were both exactly ST-08."}, {"speaker": "records officer", "text": "Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as confirmed, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4."}, {"speaker": "identification recorder", "text": "Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved."}, {"speaker": "audit clerk", "text": "The survey contained three specimen lots, and each had a numeric count: 12, 7, and 6. Flow, substrate, canopy shade, and wetted width observations were recorded. The 11:30 UTC audit was conducted after identification-correction precedence had been applied. It documented every fact needed for routing directly rather than by inference and found every record field other than V4's species-resolution status internally consistent."}]}, "method": "c2d", "provenance": {"source_id": "diverse-247", "source_is_synthetic": true, "source_sha256": "9d6d41d1934faa1fee1cda0be4e6d613633e53a4f44a8bad39c426cd84cb0f2d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "data_review_yes_low"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing decision criteria and instructions, while neither context alters policy or introduces a new rule. Survey D247, voucher V4, site ST-08, dates, times, and relevant record paths remain bound consistently. The two evidence spans are complete factual sentences. The counterfactual coherently changes E17 from confirmed to unresolved, making it consistent with E23 without conflicting with the unchanged audit facts. Neither context contains an answer choice, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universal claims such as every lot having a count remain atomic. The focus atom concerns a factual post-precedence inconsistency, not a policy conclusion. The base and counter assignments are realizable with only a10 changing: an explicit unresolved status can coexist either with an additional post-precedence conflicting status or with no such conflict, while other fields remain consistent. Empty policy_evidence is correct because the governing completeness, routing, precedence, and confidence rules are already retained in the questions object; the state contributes case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail completeness: effort, exact site-label match, numeric counts for every lot, and all four habitat observations are supported. The post-precedence inconsistency in voucher V4's species-resolution status is a remaining contradiction affecting whether Identification routing applies, so Data Review takes priority. Because that needed routing fact conflicts, Low confidence follows.", "rule_index": 0, "sound": true}, {"reason": "The conditions entail completeness. Refutation of a10 entails no post-precedence inconsistency in V4's species-resolution status, while a11 excludes contradictions elsewhere. The explicitly unresolved voucher therefore requires Identification. All routing facts are explicit, and the consistency conditions exclude conflicts, so High confidence follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Survey D247's recorded effort is at least 30 minutes."}, {"id": "a2", "statement": "Survey D247's field-sheet site label exactly matches Survey D247's registered site label."}, {"id": "a3", "statement": "Every specimen lot collected in Survey D247 has a numeric count."}, {"id": "a4", "statement": "Survey D247 has a recorded flow observation."}, {"id": "a5", "statement": "Survey D247 has a recorded substrate observation."}, {"id": "a6", "statement": "Survey D247 has a recorded canopy-shade observation."}, {"id": "a7", "statement": "Survey D247 has a recorded wetted-width observation."}, {"id": "a8", "statement": "Voucher V4 in Survey D247 is explicitly unresolved at species level."}, {"id": "a9", "statement": "Every fact needed to route Survey D247 is explicitly recorded rather than inferred."}, {"id": "a10", "statement": "After identification-correction precedence is applied, Survey D247's complete record assigns inconsistent species-resolution statuses to voucher V4."}, {"id": "a11", "statement": "After identification-correction precedence is applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent."}], "base_state_json": "[{\"speaker\":\"field recorder\",\"text\":\"Survey D247 ended at 08:40 UTC on 18 April 2026 after exactly 30 minutes of recorded effort. The field-sheet site label and registered site label were both exactly ST-08.\"},{\"speaker\":\"records officer\",\"text\":\"Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as confirmed, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4.\"},{\"speaker\":\"identification recorder\",\"text\":\"Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved.\"},{\"speaker\":\"audit clerk\",\"text\":\"The survey contained three specimen lots, and each had a numeric count: 12, 7, and 6. Flow, substrate, canopy shade, and wetted width observations were recorded. The 11:30 UTC audit was conducted after identification-correction precedence had been applied. It documented every fact needed for routing directly rather than by inference and found every record field other than V4's species-resolution status internally consistent.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["1", "text"], "text": "Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as confirmed, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4."}, {"path": ["2", "text"], "text": "Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved."}], "policy_evidence": [], "rules": [{"justification": "All completeness requirements are satisfied. After correction precedence, voucher V4 has conflicting species-resolution statuses, so a contradiction remains in a fact needed to decide whether the survey routes to Identification; the required outcome is therefore Data Review with Low confidence.", "target": "data_review_yes_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All completeness requirements are satisfied. Voucher V4 is explicitly unresolved, its post-precedence species-resolution status is not inconsistent, every other record field is consistent, and every routing fact is explicit; the required outcome is therefore Identification with High confidence.", "target": "identification_yes_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as confirmed, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4.", "negative_left": "Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as unresolved, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4.", "negative_right": "Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved.", "right": "Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved."}, "verifier_independent_model": false}, "family": "scale-diverse-247-005", "id": "scale-diverse-247-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_yes_low": "Completeness Yes; Data-review queue; Low confidence. Choose only if a record contradiction remains after correction precedence is applied and a needed routing fact conflicts.", "identification_yes_high": "Completeness Yes; Identification queue; High confidence. Choose only if the complete record has no remaining contradiction, at least one taxon is explicitly unresolved, and all routing facts are direct and consistent.", "sampling_no_medium": "Completeness No; Sampling queue; Medium confidence. Choose only if at least one required completeness element is absent and that conclusion requires one inference.", "validated_yes_medium": "Completeness Yes; Validated; Medium confidence. Choose only if no contradiction or unresolved taxon remains and one required routing fact must be inferred."}, "instructions": "Select the combined completeness judgment, queue, and confidence rating. Completeness is Yes only when effort is at least 30 minutes, the site label exactly matches the register, every specimen lot has a numeric count, and all four habitat observations are present. Route incomplete surveys to Sampling; otherwise route remaining record contradictions to Data Review; otherwise route any unresolved taxon to Identification; otherwise mark Validated. A later explicit statement that an identification is not confirmed controls over its quoted tentative note and is not itself a contradiction. Confidence is ordered High > Medium > Low: High requires all routing facts to be explicit and consistent; Medium requires one inference; Low means a needed fact conflicts.", "type": "choice"}}, "state": [{"speaker": "field recorder", "text": "Survey D247 ended at 08:40 UTC on 18 April 2026 after exactly 30 minutes of recorded effort. The field-sheet site label and registered site label were both exactly ST-08."}, {"speaker": "records officer", "text": "Entry E17, posted to Survey D247 at 09:12 UTC on 18 April 2026, marked voucher V4's species-level identification as unresolved, and the 11:30 UTC post-precedence audit found that E17 and E23 were the complete record's only operative species-resolution entries for V4."}, {"speaker": "identification recorder", "text": "Entry E23, posted to Survey D247 at 10:46 UTC on 18 April 2026, explicitly marked voucher V4's species-level identification as unresolved."}, {"speaker": "audit clerk", "text": "The survey contained three specimen lots, and each had a numeric count: 12, 7, and 6. Flow, substrate, canopy shade, and wetted width observations were recorded. The 11:30 UTC audit was conducted after identification-correction precedence had been applied. It documented every fact needed for routing directly rather than by inference and found every record field other than V4's species-resolution status internally consistent."}]}, "method": "c2d", "provenance": {"source_id": "diverse-247", "source_is_synthetic": true, "source_sha256": "9d6d41d1934faa1fee1cda0be4e6d613633e53a4f44a8bad39c426cd84cb0f2d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_yes_high"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara’s routing policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the Pine Fork survey, PF-1/PF-2, submission/transfer-log path, and routing-time framing. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the transfer-log status of QV-91 while leaving the six-vial submission and all other records unchanged, creating no duplicate or contradictory measurement. Neither context states an answer option, code, proposition ID, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Evidence reconciliation was completed at routing time. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91. At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91. Each site received the required 15-minute kick sample and 5-minute hand search. The kick-sample and hand-search effort records were complete for each of PF-1 and PF-2. Specimen-count entries were complete on both the field sheets and custody record, with each record showing a total of 47. Flow, substrate, canopy, and temperature observations were complete for both sites. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91."}, {"path": [], "text": "At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91.", "negative_right": "At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, and QV-68 but contained no entry for QV-91.", "right": "At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91."}, "verifier_independent_model": false}, "family": "scale-diverse-248-002", "id": "scale-diverse-248-002-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Evidence reconciliation was completed at routing time. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91. At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91. Each site received the required 15-minute kick sample and 5-minute hand search. The kick-sample and hand-search effort records were complete for each of PF-1 and PF-2. Specimen-count entries were complete on both the field sheets and custody record, with each record showing a total of 47. Flow, substrate, canopy, and temperature observations were complete for both sites. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara’s routing policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the Pine Fork survey, PF-1/PF-2, submission/transfer-log path, and routing-time framing. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes only the transfer-log status of QV-91 while leaving the six-vial submission and all other records unchanged, creating no duplicate or contradictory measurement. Neither context states an answer option, code, proposition ID, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Evidence reconciliation was completed at routing time. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91. At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91. Each site received the required 15-minute kick sample and 5-minute hand search. The kick-sample and hand-search effort records were complete for each of PF-1 and PF-2. Specimen-count entries were complete on both the field sheets and custody record, with each record showing a total of 47. Flow, substrate, canopy, and temperature observations were complete for both sites. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91."}, {"path": [], "text": "At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91.", "negative_right": "At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, and QV-68 but contained no entry for QV-91.", "right": "At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91."}, "verifier_independent_model": false}, "family": "scale-diverse-248-002", "id": "scale-diverse-248-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Evidence reconciliation was completed at routing time. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six vial codes QV-14, QV-27, QV-39, QV-52, QV-68, and QV-91. At routing time, that submission’s transfer log contained entries for QV-14, QV-27, QV-39, QV-52, and QV-68 but contained no entry for QV-91. Each site received the required 15-minute kick sample and 5-minute hand search. The kick-sample and hand-search effort records were complete for each of PF-1 and PF-2. Specimen-count entries were complete on both the field sheets and custody record, with each record showing a total of 47. Flow, substrate, canopy, and temperature observations were complete for both sites. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara’s governing routing policy, the same Pine Fork survey, sites PF-1/PF-2, routing-time scope, and relevant record types. The two evidence spans are complete factual sentences. The counterfactual coherently changes only QF-79’s transfer-log status without conflicting with the six-vial inventory or any unchanged count, effort, or habitat assertion. Neither context embeds a gold option, answer code, proposition identifier, rationale, or output instruction; the queue terminology appears only as part of the preserved natural-language policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79. At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79. The routing checklist marks the required 15-minute kick-sample effort record and the required 5-minute hand-search effort record complete for each site. The field-sheet specimen-count record and custody specimen-count record are complete; each reports 47 specimens. For both PF-1 and PF-2, flow, substrate, canopy, and temperature observations are complete. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79."}, {"path": [], "text": "At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79.", "negative_right": "At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, and QF-66 but contains no entry for code QF-79.", "right": "At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79."}, "verifier_independent_model": false}, "family": "scale-diverse-248-003", "id": "scale-diverse-248-003-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79. At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79. The routing checklist marks the required 15-minute kick-sample effort record and the required 5-minute hand-search effort record complete for each site. The field-sheet specimen-count record and custody specimen-count record are complete; each reports 47 specimens. For both PF-1 and PF-2, flow, substrate, canopy, and temperature observations are complete. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara’s governing routing policy, the same Pine Fork survey, sites PF-1/PF-2, routing-time scope, and relevant record types. The two evidence spans are complete factual sentences. The counterfactual coherently changes only QF-79’s transfer-log status without conflicting with the six-vial inventory or any unchanged count, effort, or habitat assertion. Neither context embeds a gold option, answer code, proposition identifier, rationale, or output instruction; the queue terminology appears only as part of the preserved natural-language policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79. At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79. The routing checklist marks the required 15-minute kick-sample effort record and the required 5-minute hand-search effort record complete for each site. The field-sheet specimen-count record and custody specimen-count record are complete; each reports 47 specimens. For both PF-1 and PF-2, flow, substrate, canopy, and temperature observations are complete. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79."}, {"path": [], "text": "At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79.", "negative_right": "At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, and QF-66 but contains no entry for code QF-79.", "right": "At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79."}, "verifier_independent_model": false}, "family": "scale-diverse-248-003", "id": "scale-diverse-248-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: QF-17, QF-24, QF-38, QF-51, QF-66, and QF-79. At routing time, Mara’s Pine Fork survey submission’s transfer log contains an entry for each of codes QF-17, QF-24, QF-38, QF-51, and QF-66 but contains no entry for code QF-79. The routing checklist marks the required 15-minute kick-sample effort record and the required 5-minute hand-search effort record complete for each site. The field-sheet specimen-count record and custody specimen-count record are complete; each reports 47 specimens. For both PF-1 and PF-2, flow, substrate, canopy, and temperature observations are complete. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara’s original governing routing policy verbatim, while the unchanged questions preserve all choice criteria and instructions. The Pine Fork, PF-1/PF-2, submission, and routing-time bindings remain fixed; only the transfer-log observation changes. The two evidence spans are complete factual sentences. The counterfactual coherently changes the log from containing all six vial codes to containing five with PFV-83 absent, without conflicting with the unchanged effort, count, or habitat facts. Neither context includes an answer choice, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Routing field note for Mara’s Pine Fork submission covering PF-1 and PF-2: At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83. At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83. Each site received the required 15-minute kick sample and 5-minute hand search. The routing-time effort record marks both required methods complete at PF-1 and PF-2. The field-sheet and custody specimen-count records are complete at routing time, and both show a total of 47 specimens. Flow, substrate, canopy, and temperature observations are each complete for both sites at routing time. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83."}, {"path": [], "text": "At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83.", "negative_right": "At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, and PFV-68 but does not list PFV-83.", "right": "At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83."}, "verifier_independent_model": false}, "family": "scale-diverse-248-004", "id": "scale-diverse-248-004-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Routing field note for Mara’s Pine Fork submission covering PF-1 and PF-2: At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83. At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83. Each site received the required 15-minute kick sample and 5-minute hand search. The routing-time effort record marks both required methods complete at PF-1 and PF-2. The field-sheet and custody specimen-count records are complete at routing time, and both show a total of 47 specimens. Flow, substrate, canopy, and temperature observations are each complete for both sites at routing time. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara’s original governing routing policy verbatim, while the unchanged questions preserve all choice criteria and instructions. The Pine Fork, PF-1/PF-2, submission, and routing-time bindings remain fixed; only the transfer-log observation changes. The two evidence spans are complete factual sentences. The counterfactual coherently changes the log from containing all six vial codes to containing five with PFV-83 absent, without conflicting with the unchanged effort, count, or habitat facts. Neither context includes an answer choice, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Routing field note for Mara’s Pine Fork submission covering PF-1 and PF-2: At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83. At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83. Each site received the required 15-minute kick sample and 5-minute hand search. The routing-time effort record marks both required methods complete at PF-1 and PF-2. The field-sheet and custody specimen-count records are complete at routing time, and both show a total of 47 specimens. Flow, substrate, canopy, and temperature observations are each complete for both sites at routing time. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83."}, {"path": [], "text": "At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83.", "negative_right": "At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, and PFV-68 but does not list PFV-83.", "right": "At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83."}, "verifier_independent_model": false}, "family": "scale-diverse-248-004", "id": "scale-diverse-248-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Routing field note for Mara’s Pine Fork submission covering PF-1 and PF-2: At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: PFV-14, PFV-27, PFV-39, PFV-52, PFV-68, and PFV-83. At routing time, that submission’s transfer log lists PFV-14, PFV-27, PFV-39, PFV-52, and PFV-68 but does not list PFV-83. Each site received the required 15-minute kick sample and 5-minute hand search. The routing-time effort record marks both required methods complete at PF-1 and PF-2. The field-sheet and custody specimen-count records are complete at routing time, and both show a total of 47 specimens. Flow, substrate, canopy, and temperature observations are each complete for both sites at routing time. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara’s governing routing policy, while the unchanged questions preserve all choice criteria and instructions. The Pine Fork survey, Mara, sites PF-1/PF-2, routing-time scope, and relevant record paths remain bound to the original question. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the transfer-log contents from six listed vial codes to five, without contradicting the unchanged assertion that the submission contains six vials. Neither context includes an answer code, proposition ID, rule table, label rationale, or output instruction; Mara’s quoted routing statement is preserved governing policy rather than answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C. At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C. The routing packet marked the required 15-minute kick-sample effort record and required 5-minute hand-search effort record complete for each site. Mara’s submission had complete field-sheet and custody specimen-count records for both sites, and their specimen totals were equal. At routing time, flow, substrate, canopy, and temperature observations were complete for each site. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C."}, {"path": [], "text": "At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C.", "negative_right": "At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, and PF-66A.", "right": "At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C."}, "verifier_independent_model": false}, "family": "scale-diverse-248-005", "id": "scale-diverse-248-005-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C. At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C. The routing packet marked the required 15-minute kick-sample effort record and required 5-minute hand-search effort record complete for each site. Mara’s submission had complete field-sheet and custody specimen-count records for both sites, and their specimen totals were equal. At routing time, flow, substrate, canopy, and temperature observations were complete for each site. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara’s governing routing policy, while the unchanged questions preserve all choice criteria and instructions. The Pine Fork survey, Mara, sites PF-1/PF-2, routing-time scope, and relevant record paths remain bound to the original question. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the transfer-log contents from six listed vial codes to five, without contradicting the unchanged assertion that the submission contains six vials. Neither context includes an answer code, proposition ID, rule table, label rationale, or output instruction; Mara’s quoted routing statement is preserved governing policy rather than answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C. At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C. The routing packet marked the required 15-minute kick-sample effort record and required 5-minute hand-search effort record complete for each site. Mara’s submission had complete field-sheet and custody specimen-count records for both sites, and their specimen totals were equal. At routing time, flow, substrate, canopy, and temperature observations were complete for each site. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C."}, {"path": [], "text": "At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C.", "negative_right": "At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, and PF-66A.", "right": "At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C."}, "verifier_independent_model": false}, "family": "scale-diverse-248-005", "id": "scale-diverse-248-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contained exactly the six distinct vial codes PF-17A, PF-24C, PF-38B, PF-51D, PF-66A, and PF-79C. At routing time, that submission’s transfer log contained entries for exactly the vial codes PF-17A, PF-24C, PF-38B, PF-51D, and PF-66A. The routing packet marked the required 15-minute kick-sample effort record and required 5-minute hand-search effort record complete for each site. Mara’s submission had complete field-sheet and custody specimen-count records for both sites, and their specimen totals were equal. At routing time, flow, substrate, canopy, and temperature observations were complete for each site. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original completion and conflict-resolution policy without adding exceptions, priorities, or missing-evidence defaults, and they remain bound to the same SR-04 survey and routing decision. The focus evidence consists of exactly two complete factual sentences. Changing the physical-label inventory from 47 to 46 is coherent: it does not contradict the separately stated 48-line raw log, 47-specimen field tally, or the exactly-one raw-log-versus-field-tally discrepancy; it merely leaves a different specimen-count inconsistency. Neither context states the required output, confidence selection, answer code, rule table, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "supported", "site_label_assignment": "supported"}, "full_context_fact_states": {"base": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "supported", "site_label_assignment": "supported"}, "counterfactual": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "refuted", "site_label_assignment": "supported"}, "remove_left": {"physical_deduplicated_count_agreement": "unknown"}, "remove_right": {"physical_deduplicated_count_agreement": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"physical_deduplicated_count_agreement": "unknown"}, "negative_pair": {"physical_deduplicated_count_agreement": "refuted"}, "negative_sentence": {"physical_deduplicated_count_agreement": "unknown"}, "positive_pair": {"physical_deduplicated_count_agreement": "supported"}, "right": {"physical_deduplicated_count_agreement": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally scoped statements over the three replicates remain atomic. The focus atom is a factual count-agreement relation rather than a policy conclusion. The base and counter assignments are jointly realizable: the same exactly-one field-versus-log conflict and duplicate entry can coexist either with an agreeing physical-sequence total or with a conflicting one. Empty policy_evidence is correct because the governing completeness, conflict-resolution, and confidence rules are all contained in the questions object; the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction directly documents sampling effort, unambiguous replicate labels, a specimen tally, and habitat observations. It also establishes an exactly-one count conflict, a duplicated log entry, and agreement between the physical sequence and the log after deduplication. Together with the absence of another required-field dispute, these conditions are sufficient for Yes with High confidence under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "The required fields remain documented, but the physical sequence is explicitly established not to agree with the deduplicated log count. Therefore the exactly-one count conflict is not resolved by the required combination of physical sequence and duplicated record, which is sufficient for routing to data review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "effort_replicate_count", "statement": "The supplied records document that the survey at site SR-04 completed three kick-net replicates."}, {"id": "effort_replicate_duration", "statement": "The supplied records document a duration of 2 minutes for every completed kick-net replicate in the survey at site SR-04."}, {"id": "site_label_assignment", "statement": "The supplied records assign labels SR-04/R1, SR-04/R2, and SR-04/R3 one-to-one to the three completed replicates at site SR-04."}, {"id": "field_specimen_count", "statement": "The supplied records document a field specimen tally of 47 for the survey at site SR-04."}, {"id": "habitat_observations", "statement": "The supplied records document habitat observations for the survey at site SR-04."}, {"id": "exactly_one_count_conflict", "statement": "For the survey at site SR-04, the specimen-log total exceeds the field specimen tally by exactly one."}, {"id": "duplicated_log_record", "statement": "The two specimen-log lines bearing identifier SR-04/R2-017 are duplicate entries of one record."}, {"id": "physical_deduplicated_count_agreement", "statement": "For the survey at site SR-04, the continuous physical-label-sequence total equals the specimen-log count after one of the two duplicate SR-04/R2-017 lines is excluded."}, {"id": "no_competing_required_field_dispute", "statement": "In the supplied records for the survey at site SR-04, no required field other than specimen-count consistency is disputed."}], "base_state_json": "[{\"speaker\":\"field survey lead\",\"text\":\"On 13 May 2026, site SR-04 completed three kick-net replicates, each lasting 2 minutes. The three replicates were assigned one-to-one to SR-04/R1, SR-04/R2, and SR-04/R3. The field tally was 47 specimens. Habitat observations recorded gravel riffle, shaded bank, 14°C water, and moderate flow.\"},{\"speaker\":\"records auditor\",\"text\":\"On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record.\"},{\"speaker\":\"collection custodian\",\"text\":\"On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 47 labels.\"},{\"speaker\":\"quality reviewer\",\"text\":\"On 16 May 2026, review confirmed that the raw-log versus field-tally discrepancy was exactly one. Specimen-count consistency was the only disputed required field; all other required fields were uncontested.\"}]", "base_states": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "supported"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}], "counter_states": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "refuted"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}], "focus_atom": "physical_deduplicated_count_agreement", "focus_evidence": [{"path": ["1", "text"], "text": "On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record."}, {"path": ["2", "text"], "text": "On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 47 labels."}], "policy_evidence": [], "rules": [{"justification": "All required survey fields are directly documented, and the exactly-one specimen-count conflict is resolved jointly by the duplicated record and the agreeing physical sequence.", "target": "true", "when": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "supported"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}]}, {"justification": "Although the required fields and duplicated record are documented, the physical sequence contradicts the deduplicated log count, so the exactly-one specimen-count conflict remains unresolved.", "target": "false", "when": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "refuted"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}]}]}, "verified_pair": {"left": "On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record.", "negative_left": "On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record.", "negative_right": "On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 46 labels.", "right": "On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 47 labels."}, "verifier_independent_model": false}, "family": "scale-diverse-250-001", "id": "scale-diverse-250-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route to data review because a required field is missing or the specimen-count conflict remains unresolved.", "true": "Yes — route to the taxonomic identification queue as complete, with High confidence."}, "instructions": "Decide whether this survey should be routed to the taxonomic identification queue as complete. A survey is complete when sampling effort, unambiguous site labels, specimen count, and habitat observations are supplied. A count conflict of exactly one is acceptable only when the supplied physical sequence and duplicated record resolve it; otherwise route to data review. Confidence is ordered Low < Medium < High. Select Yes with High confidence only when every criterion and the conflict resolution are directly documented in the supplied records.", "type": "noul"}}, "state": [{"speaker": "field survey lead", "text": "On 13 May 2026, site SR-04 completed three kick-net replicates, each lasting 2 minutes. The three replicates were assigned one-to-one to SR-04/R1, SR-04/R2, and SR-04/R3. The field tally was 47 specimens. Habitat observations recorded gravel riffle, shaded bank, 14°C water, and moderate flow."}, {"speaker": "records auditor", "text": "On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record."}, {"speaker": "collection custodian", "text": "On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 47 labels."}, {"speaker": "quality reviewer", "text": "On 16 May 2026, review confirmed that the raw-log versus field-tally discrepancy was exactly one. Specimen-count consistency was the only disputed required field; all other required fields were uncontested."}]}, "method": "c2d", "provenance": {"source_id": "diverse-250", "source_is_synthetic": true, "source_sha256": "d6f27723b730632c53f25cae707f9ca8cc621630985ba69f06620839f4c0c116", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original completion and conflict-resolution policy without adding exceptions, priorities, or missing-evidence defaults, and they remain bound to the same SR-04 survey and routing decision. The focus evidence consists of exactly two complete factual sentences. Changing the physical-label inventory from 47 to 46 is coherent: it does not contradict the separately stated 48-line raw log, 47-specimen field tally, or the exactly-one raw-log-versus-field-tally discrepancy; it merely leaves a different specimen-count inconsistency. Neither context states the required output, confidence selection, answer code, rule table, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "refuted", "site_label_assignment": "supported"}, "full_context_fact_states": {"base": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "supported", "site_label_assignment": "supported"}, "counterfactual": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "refuted", "site_label_assignment": "supported"}, "remove_left": {"physical_deduplicated_count_agreement": "unknown"}, "remove_right": {"physical_deduplicated_count_agreement": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"physical_deduplicated_count_agreement": "unknown"}, "negative_pair": {"physical_deduplicated_count_agreement": "refuted"}, "negative_sentence": {"physical_deduplicated_count_agreement": "unknown"}, "positive_pair": {"physical_deduplicated_count_agreement": "supported"}, "right": {"physical_deduplicated_count_agreement": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally scoped statements over the three replicates remain atomic. The focus atom is a factual count-agreement relation rather than a policy conclusion. The base and counter assignments are jointly realizable: the same exactly-one field-versus-log conflict and duplicate entry can coexist either with an agreeing physical-sequence total or with a conflicting one. Empty policy_evidence is correct because the governing completeness, conflict-resolution, and confidence rules are all contained in the questions object; the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction directly documents sampling effort, unambiguous replicate labels, a specimen tally, and habitat observations. It also establishes an exactly-one count conflict, a duplicated log entry, and agreement between the physical sequence and the log after deduplication. Together with the absence of another required-field dispute, these conditions are sufficient for Yes with High confidence under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "The required fields remain documented, but the physical sequence is explicitly established not to agree with the deduplicated log count. Therefore the exactly-one count conflict is not resolved by the required combination of physical sequence and duplicated record, which is sufficient for routing to data review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "effort_replicate_count", "statement": "The supplied records document that the survey at site SR-04 completed three kick-net replicates."}, {"id": "effort_replicate_duration", "statement": "The supplied records document a duration of 2 minutes for every completed kick-net replicate in the survey at site SR-04."}, {"id": "site_label_assignment", "statement": "The supplied records assign labels SR-04/R1, SR-04/R2, and SR-04/R3 one-to-one to the three completed replicates at site SR-04."}, {"id": "field_specimen_count", "statement": "The supplied records document a field specimen tally of 47 for the survey at site SR-04."}, {"id": "habitat_observations", "statement": "The supplied records document habitat observations for the survey at site SR-04."}, {"id": "exactly_one_count_conflict", "statement": "For the survey at site SR-04, the specimen-log total exceeds the field specimen tally by exactly one."}, {"id": "duplicated_log_record", "statement": "The two specimen-log lines bearing identifier SR-04/R2-017 are duplicate entries of one record."}, {"id": "physical_deduplicated_count_agreement", "statement": "For the survey at site SR-04, the continuous physical-label-sequence total equals the specimen-log count after one of the two duplicate SR-04/R2-017 lines is excluded."}, {"id": "no_competing_required_field_dispute", "statement": "In the supplied records for the survey at site SR-04, no required field other than specimen-count consistency is disputed."}], "base_state_json": "[{\"speaker\":\"field survey lead\",\"text\":\"On 13 May 2026, site SR-04 completed three kick-net replicates, each lasting 2 minutes. The three replicates were assigned one-to-one to SR-04/R1, SR-04/R2, and SR-04/R3. The field tally was 47 specimens. Habitat observations recorded gravel riffle, shaded bank, 14°C water, and moderate flow.\"},{\"speaker\":\"records auditor\",\"text\":\"On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record.\"},{\"speaker\":\"collection custodian\",\"text\":\"On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 47 labels.\"},{\"speaker\":\"quality reviewer\",\"text\":\"On 16 May 2026, review confirmed that the raw-log versus field-tally discrepancy was exactly one. Specimen-count consistency was the only disputed required field; all other required fields were uncontested.\"}]", "base_states": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "supported"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}], "counter_states": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "refuted"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}], "focus_atom": "physical_deduplicated_count_agreement", "focus_evidence": [{"path": ["1", "text"], "text": "On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record."}, {"path": ["2", "text"], "text": "On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 47 labels."}], "policy_evidence": [], "rules": [{"justification": "All required survey fields are directly documented, and the exactly-one specimen-count conflict is resolved jointly by the duplicated record and the agreeing physical sequence.", "target": "true", "when": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "supported"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}]}, {"justification": "Although the required fields and duplicated record are documented, the physical sequence contradicts the deduplicated log count, so the exactly-one specimen-count conflict remains unresolved.", "target": "false", "when": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "refuted"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}]}]}, "verified_pair": {"left": "On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record.", "negative_left": "On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record.", "negative_right": "On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 46 labels.", "right": "On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 47 labels."}, "verifier_independent_model": false}, "family": "scale-diverse-250-001", "id": "scale-diverse-250-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route to data review because a required field is missing or the specimen-count conflict remains unresolved.", "true": "Yes — route to the taxonomic identification queue as complete, with High confidence."}, "instructions": "Decide whether this survey should be routed to the taxonomic identification queue as complete. A survey is complete when sampling effort, unambiguous site labels, specimen count, and habitat observations are supplied. A count conflict of exactly one is acceptable only when the supplied physical sequence and duplicated record resolve it; otherwise route to data review. Confidence is ordered Low < Medium < High. Select Yes with High confidence only when every criterion and the conflict resolution are directly documented in the supplied records.", "type": "noul"}}, "state": [{"speaker": "field survey lead", "text": "On 13 May 2026, site SR-04 completed three kick-net replicates, each lasting 2 minutes. The three replicates were assigned one-to-one to SR-04/R1, SR-04/R2, and SR-04/R3. The field tally was 47 specimens. Habitat observations recorded gravel riffle, shaded bank, 14°C water, and moderate flow."}, {"speaker": "records auditor", "text": "On 14 May 2026, the specimen log for the survey at site SR-04 contained 48 lines, including exactly two lines bearing identifier SR-04/R2-017 that were duplicate entries of one record."}, {"speaker": "collection custodian", "text": "On 15 May 2026, an uninterrupted inventory of the physical labels from the survey at site SR-04 counted 46 labels."}, {"speaker": "quality reviewer", "text": "On 16 May 2026, review confirmed that the raw-log versus field-tally discrepancy was exactly one. Specimen-count consistency was the only disputed required field; all other required fields were uncontested."}]}, "method": "c2d", "provenance": {"source_id": "diverse-250", "source_is_synthetic": true, "source_sha256": "d6f27723b730632c53f25cae707f9ca8cc621630985ba69f06620839f4c0c116", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original completion and count-conflict policy without adding exceptions or defaults. The counterfactual changes only the physical-label count from 47 to 46 while preserving the SR-04 survey, audit time, request scope, and relevant evidence path. The two focus-evidence spans are complete factual sentences. The revised count creates an unresolved cross-record discrepancy rather than a contradictory duplicate measurement, so it remains coherent with the unchanged statements. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "supported", "site_label_assignment": "supported"}, "full_context_fact_states": {"base": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "supported", "site_label_assignment": "supported"}, "counterfactual": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "refuted", "site_label_assignment": "supported"}, "remove_left": {"physical_deduplicated_count_agreement": "unknown"}, "remove_right": {"physical_deduplicated_count_agreement": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"physical_deduplicated_count_agreement": "unknown"}, "negative_pair": {"physical_deduplicated_count_agreement": "refuted"}, "negative_sentence": {"physical_deduplicated_count_agreement": "unknown"}, "positive_pair": {"physical_deduplicated_count_agreement": "supported"}, "right": {"physical_deduplicated_count_agreement": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally scoped statements over the three replicates remain atomic. The focus atom is a factual count-agreement relation rather than a policy conclusion. The base and counter assignments are jointly realizable: the same exactly-one field-versus-log conflict and duplicate entry can coexist either with an agreeing physical-sequence total or with a conflicting one. Empty policy_evidence is correct because the governing completeness, conflict-resolution, and confidence rules are all contained in the questions object; the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction directly documents sampling effort, unambiguous replicate labels, a specimen tally, and habitat observations. It also establishes an exactly-one count conflict, a duplicated log entry, and agreement between the physical sequence and the log after deduplication. Together with the absence of another required-field dispute, these conditions are sufficient for Yes with High confidence under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "The required fields remain documented, but the physical sequence is explicitly established not to agree with the deduplicated log count. Therefore the exactly-one count conflict is not resolved by the required combination of physical sequence and duplicated record, which is sufficient for routing to data review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "effort_replicate_count", "statement": "The supplied records document that the survey at site SR-04 completed three kick-net replicates."}, {"id": "effort_replicate_duration", "statement": "The supplied records document a duration of 2 minutes for every completed kick-net replicate in the survey at site SR-04."}, {"id": "site_label_assignment", "statement": "The supplied records assign labels SR-04/R1, SR-04/R2, and SR-04/R3 one-to-one to the three completed replicates at site SR-04."}, {"id": "field_specimen_count", "statement": "The supplied records document a field specimen tally of 47 for the survey at site SR-04."}, {"id": "habitat_observations", "statement": "The supplied records document habitat observations for the survey at site SR-04."}, {"id": "exactly_one_count_conflict", "statement": "For the survey at site SR-04, the specimen-log total exceeds the field specimen tally by exactly one."}, {"id": "duplicated_log_record", "statement": "The two specimen-log lines bearing identifier SR-04/R2-017 are duplicate entries of one record."}, {"id": "physical_deduplicated_count_agreement", "statement": "For the survey at site SR-04, the continuous physical-label-sequence total equals the specimen-log count after one of the two duplicate SR-04/R2-017 lines is excluded."}, {"id": "no_competing_required_field_dispute", "statement": "In the supplied records for the survey at site SR-04, no required field other than specimen-count consistency is disputed."}], "base_state_json": "[{\"speaker\":\"field documentation reviewer\",\"text\":\"The SR-04 completion sheet records three completed kick-net replicates, each lasting 2 minutes. It assigns SR-04/R1, SR-04/R2, and SR-04/R3 one-to-one to those replicates, gives a field specimen tally of 47, and includes habitat observations covering gravel substrate, moderate flow, bank shade, and 14°C water.\"},{\"speaker\":\"physical collection auditor\",\"text\":\"For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 47 labels in the continuous physical-label sequence.\"},{\"speaker\":\"specimen records auditor\",\"text\":\"For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines.\"},{\"speaker\":\"reconciliation reviewer\",\"text\":\"The reconciliation worksheet notes that the uncorrected specimen-log total exceeds the field specimen tally by exactly one. Its required-field review identifies specimen-count consistency as the only disputed field in the supplied SR-04 records.\"}]", "base_states": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "supported"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}], "counter_states": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "refuted"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}], "focus_atom": "physical_deduplicated_count_agreement", "focus_evidence": [{"path": ["1", "text"], "text": "For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 47 labels in the continuous physical-label sequence."}, {"path": ["2", "text"], "text": "For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines."}], "policy_evidence": [], "rules": [{"justification": "All required survey fields are directly documented, and the exactly-one specimen-count conflict is resolved jointly by the duplicated record and the agreeing physical sequence.", "target": "true", "when": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "supported"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}]}, {"justification": "Although the required fields and duplicated record are documented, the physical sequence contradicts the deduplicated log count, so the exactly-one specimen-count conflict remains unresolved.", "target": "false", "when": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "refuted"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}]}]}, "verified_pair": {"left": "For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 47 labels in the continuous physical-label sequence.", "negative_left": "For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 46 labels in the continuous physical-label sequence.", "negative_right": "For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines.", "right": "For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines."}, "verifier_independent_model": false}, "family": "scale-diverse-250-002", "id": "scale-diverse-250-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route to data review because a required field is missing or the specimen-count conflict remains unresolved.", "true": "Yes — route to the taxonomic identification queue as complete, with High confidence."}, "instructions": "Decide whether this survey should be routed to the taxonomic identification queue as complete. A survey is complete when sampling effort, unambiguous site labels, specimen count, and habitat observations are supplied. A count conflict of exactly one is acceptable only when the supplied physical sequence and duplicated record resolve it; otherwise route to data review. Confidence is ordered Low < Medium < High. Select Yes with High confidence only when every criterion and the conflict resolution are directly documented in the supplied records.", "type": "noul"}}, "state": [{"speaker": "field documentation reviewer", "text": "The SR-04 completion sheet records three completed kick-net replicates, each lasting 2 minutes. It assigns SR-04/R1, SR-04/R2, and SR-04/R3 one-to-one to those replicates, gives a field specimen tally of 47, and includes habitat observations covering gravel substrate, moderate flow, bank shade, and 14°C water."}, {"speaker": "physical collection auditor", "text": "For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 47 labels in the continuous physical-label sequence."}, {"speaker": "specimen records auditor", "text": "For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines."}, {"speaker": "reconciliation reviewer", "text": "The reconciliation worksheet notes that the uncorrected specimen-log total exceeds the field specimen tally by exactly one. Its required-field review identifies specimen-count consistency as the only disputed field in the supplied SR-04 records."}]}, "method": "c2d", "provenance": {"source_id": "diverse-250", "source_is_synthetic": true, "source_sha256": "d6f27723b730632c53f25cae707f9ca8cc621630985ba69f06620839f4c0c116", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original completion and count-conflict policy without adding exceptions or defaults. The counterfactual changes only the physical-label count from 47 to 46 while preserving the SR-04 survey, audit time, request scope, and relevant evidence path. The two focus-evidence spans are complete factual sentences. The revised count creates an unresolved cross-record discrepancy rather than a contradictory duplicate measurement, so it remains coherent with the unchanged statements. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "refuted", "site_label_assignment": "supported"}, "full_context_fact_states": {"base": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "supported", "site_label_assignment": "supported"}, "counterfactual": {"duplicated_log_record": "supported", "effort_replicate_count": "supported", "effort_replicate_duration": "supported", "exactly_one_count_conflict": "supported", "field_specimen_count": "supported", "habitat_observations": "supported", "no_competing_required_field_dispute": "supported", "physical_deduplicated_count_agreement": "refuted", "site_label_assignment": "supported"}, "remove_left": {"physical_deduplicated_count_agreement": "unknown"}, "remove_right": {"physical_deduplicated_count_agreement": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"physical_deduplicated_count_agreement": "unknown"}, "negative_pair": {"physical_deduplicated_count_agreement": "refuted"}, "negative_sentence": {"physical_deduplicated_count_agreement": "unknown"}, "positive_pair": {"physical_deduplicated_count_agreement": "supported"}, "right": {"physical_deduplicated_count_agreement": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally scoped statements over the three replicates remain atomic. The focus atom is a factual count-agreement relation rather than a policy conclusion. The base and counter assignments are jointly realizable: the same exactly-one field-versus-log conflict and duplicate entry can coexist either with an agreeing physical-sequence total or with a conflicting one. Empty policy_evidence is correct because the governing completeness, conflict-resolution, and confidence rules are all contained in the questions object; the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction directly documents sampling effort, unambiguous replicate labels, a specimen tally, and habitat observations. It also establishes an exactly-one count conflict, a duplicated log entry, and agreement between the physical sequence and the log after deduplication. Together with the absence of another required-field dispute, these conditions are sufficient for Yes with High confidence under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "The required fields remain documented, but the physical sequence is explicitly established not to agree with the deduplicated log count. Therefore the exactly-one count conflict is not resolved by the required combination of physical sequence and duplicated record, which is sufficient for routing to data review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "effort_replicate_count", "statement": "The supplied records document that the survey at site SR-04 completed three kick-net replicates."}, {"id": "effort_replicate_duration", "statement": "The supplied records document a duration of 2 minutes for every completed kick-net replicate in the survey at site SR-04."}, {"id": "site_label_assignment", "statement": "The supplied records assign labels SR-04/R1, SR-04/R2, and SR-04/R3 one-to-one to the three completed replicates at site SR-04."}, {"id": "field_specimen_count", "statement": "The supplied records document a field specimen tally of 47 for the survey at site SR-04."}, {"id": "habitat_observations", "statement": "The supplied records document habitat observations for the survey at site SR-04."}, {"id": "exactly_one_count_conflict", "statement": "For the survey at site SR-04, the specimen-log total exceeds the field specimen tally by exactly one."}, {"id": "duplicated_log_record", "statement": "The two specimen-log lines bearing identifier SR-04/R2-017 are duplicate entries of one record."}, {"id": "physical_deduplicated_count_agreement", "statement": "For the survey at site SR-04, the continuous physical-label-sequence total equals the specimen-log count after one of the two duplicate SR-04/R2-017 lines is excluded."}, {"id": "no_competing_required_field_dispute", "statement": "In the supplied records for the survey at site SR-04, no required field other than specimen-count consistency is disputed."}], "base_state_json": "[{\"speaker\":\"field documentation reviewer\",\"text\":\"The SR-04 completion sheet records three completed kick-net replicates, each lasting 2 minutes. It assigns SR-04/R1, SR-04/R2, and SR-04/R3 one-to-one to those replicates, gives a field specimen tally of 47, and includes habitat observations covering gravel substrate, moderate flow, bank shade, and 14°C water.\"},{\"speaker\":\"physical collection auditor\",\"text\":\"For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 47 labels in the continuous physical-label sequence.\"},{\"speaker\":\"specimen records auditor\",\"text\":\"For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines.\"},{\"speaker\":\"reconciliation reviewer\",\"text\":\"The reconciliation worksheet notes that the uncorrected specimen-log total exceeds the field specimen tally by exactly one. Its required-field review identifies specimen-count consistency as the only disputed field in the supplied SR-04 records.\"}]", "base_states": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "supported"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}], "counter_states": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "refuted"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}], "focus_atom": "physical_deduplicated_count_agreement", "focus_evidence": [{"path": ["1", "text"], "text": "For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 47 labels in the continuous physical-label sequence."}, {"path": ["2", "text"], "text": "For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines."}], "policy_evidence": [], "rules": [{"justification": "All required survey fields are directly documented, and the exactly-one specimen-count conflict is resolved jointly by the duplicated record and the agreeing physical sequence.", "target": "true", "when": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "supported"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}]}, {"justification": "Although the required fields and duplicated record are documented, the physical sequence contradicts the deduplicated log count, so the exactly-one specimen-count conflict remains unresolved.", "target": "false", "when": [{"atom_id": "effort_replicate_count", "state": "supported"}, {"atom_id": "effort_replicate_duration", "state": "supported"}, {"atom_id": "site_label_assignment", "state": "supported"}, {"atom_id": "field_specimen_count", "state": "supported"}, {"atom_id": "habitat_observations", "state": "supported"}, {"atom_id": "exactly_one_count_conflict", "state": "supported"}, {"atom_id": "duplicated_log_record", "state": "supported"}, {"atom_id": "physical_deduplicated_count_agreement", "state": "refuted"}, {"atom_id": "no_competing_required_field_dispute", "state": "supported"}]}]}, "verified_pair": {"left": "For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 47 labels in the continuous physical-label sequence.", "negative_left": "For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 46 labels in the continuous physical-label sequence.", "negative_right": "For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines.", "right": "For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines."}, "verifier_independent_model": false}, "family": "scale-diverse-250-002", "id": "scale-diverse-250-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route to data review because a required field is missing or the specimen-count conflict remains unresolved.", "true": "Yes — route to the taxonomic identification queue as complete, with High confidence."}, "instructions": "Decide whether this survey should be routed to the taxonomic identification queue as complete. A survey is complete when sampling effort, unambiguous site labels, specimen count, and habitat observations are supplied. A count conflict of exactly one is acceptable only when the supplied physical sequence and duplicated record resolve it; otherwise route to data review. Confidence is ordered Low < Medium < High. Select Yes with High confidence only when every criterion and the conflict resolution are directly documented in the supplied records.", "type": "noul"}}, "state": [{"speaker": "field documentation reviewer", "text": "The SR-04 completion sheet records three completed kick-net replicates, each lasting 2 minutes. It assigns SR-04/R1, SR-04/R2, and SR-04/R3 one-to-one to those replicates, gives a field specimen tally of 47, and includes habitat observations covering gravel substrate, moderate flow, bank shade, and 14°C water."}, {"speaker": "physical collection auditor", "text": "For the survey at site SR-04, the audit completed at 16:20 UTC on 14 June 2025 counted exactly 46 labels in the continuous physical-label sequence."}, {"speaker": "specimen records auditor", "text": "For the survey at site SR-04, the specimen log finalized at 16:20 UTC on 14 June 2025 contains 48 lines, exactly two of which bear identifier SR-04/R2-017 and are duplicate entries of one record, so excluding one of those two lines leaves 47 lines."}, {"speaker": "reconciliation reviewer", "text": "The reconciliation worksheet notes that the uncorrected specimen-log total exceeds the field specimen tally by exactly one. Its required-field review identifies specimen-count consistency as the only disputed field in the supplied SR-04 records."}]}, "method": "c2d", "provenance": {"source_id": "diverse-250", "source_is_synthetic": true, "source_sha256": "d6f27723b730632c53f25cae707f9ca8cc621630985ba69f06620839f4c0c116", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions and preserve the original workflow without adding exceptions or defaults. The contexts remain bound to Mara’s single stream-survey completeness decision, and the changed site, counts, timestamp, and field contents are permissible observation changes. The focus evidence consists of exactly two complete factual sentences. Changing E17 from a nonblank numeric entry to blank is a coherent single-sentence counterfactual and does not conflict with the unchanged statement that E17 is the sole sampling-effort field or with the absence of attachments and amendments. Neither context states an assigned score, route, gold answer, proposition identifier, or output instruction; the general routing language is preserved policy rather than answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_counts": "supported", "a_effort": "supported", "a_habitat": "refuted", "a_site": "supported"}, "full_context_fact_states": {"base": {"a_counts": "supported", "a_effort": "supported", "a_habitat": "refuted", "a_site": "supported"}, "counterfactual": {"a_counts": "supported", "a_effort": "refuted", "a_habitat": "refuted", "a_site": "supported"}, "remove_left": {"a_effort": "unknown"}, "remove_right": {"a_effort": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_effort": "unknown"}, "negative_pair": {"a_effort": "refuted"}, "negative_sentence": {"a_effort": "unknown"}, "positive_pair": {"a_effort": "supported"}, "right": {"a_effort": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship about the survey record, and the focus atom concerns the factual presence of sampling-effort information rather than a policy conclusion. The base and counter assignments are realizable with only that fact changing: the base yields level 2 and the counter yields level 1. The policy evidence accurately cites a workflow rule from the original state; all remaining governing criteria are already preserved verbatim in the questions object, while state-specific observations and role assignments need not be carried into synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A usable site label and numerical specimen counts are present, sampling effort is supplied, and habitat observations are absent. Exactly one of the two additional categories is therefore supplied, which is sufficient for level 2 under the rubric.", "rule_index": 0, "sound": true}, {"reason": "A usable site label and numerical specimen counts are present, while both sampling effort and habitat observations are absent. This is sufficient for level 1 under the rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_site", "statement": "The one stream survey record submitted by Mara for this completeness decision provides a usable stream-site label."}, {"id": "a_counts", "statement": "The one stream survey record submitted by Mara for this completeness decision provides numerical specimen counts."}, {"id": "a_effort", "statement": "The one stream survey record submitted by Mara for this completeness decision supplies sampling-effort information."}, {"id": "a_habitat", "statement": "The one stream survey record submitted by Mara for this completeness decision supplies habitat observations."}], "base_state_json": "\"Field note: At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a nonblank E17 entry of “36.” In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes. The site registry confirms that Brook-47 is a current, unambiguous label for one stream location. Review found no continuation pages, attachments, marginal notes, or later amendments associated with Mara’s submission. Ecology data manager Chen will assign completeness and routing. Taxonomic identifier Ivo remains available for organism review, and watershed coordinator Sal receives completed records. Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification.\"", "base_states": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "supported"}, {"atom_id": "a_habitat", "state": "refuted"}], "counter_states": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "refuted"}, {"atom_id": "a_habitat", "state": "refuted"}], "focus_atom": "a_effort", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a nonblank E17 entry of “36.”"}, {"path": [], "text": "In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes."}], "policy_evidence": [{"path": [], "text": "Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification."}], "rules": [{"justification": "A usable site label and numerical specimen counts are present, and sampling effort is supplied while habitat observations are missing. Thus exactly one of the two additional categories is supplied, which is sufficient for score 2 and excludes scores 0, 1, 3, and 4.", "target": "2", "when": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "supported"}, {"atom_id": "a_habitat", "state": "refuted"}]}, {"justification": "A usable site label and numerical specimen counts are present, while both sampling-effort information and habitat observations are missing. This is sufficient for score 1 and excludes score 0 and all scores requiring either additional category.", "target": "1", "when": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "refuted"}, {"atom_id": "a_habitat", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a nonblank E17 entry of “36.”", "negative_left": "At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a blank E17 entry.", "negative_right": "In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes.", "right": "In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes."}, "verifier_independent_model": false}, "family": "scale-diverse-251-004", "id": "scale-diverse-251-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: the record lacks either a usable stream-site label or any specimen counts; route to the sampling queue.", "1 — Minimally complete: a usable site label and specimen counts are present, but both sampling-effort information and habitat observations are missing; route to the sampling queue.", "2 — Partially complete: a usable site label and specimen counts are present, and exactly one of sampling effort or habitat observations is supplied; route to the data-review queue.", "3 — Nearly complete: site label, specimen counts, sampling effort, and habitat observations are all present, but at least one supplied category is vague or non-quantitative, such as “sampled briefly” or “normal habitat”; route to the data-review queue.", "4 — Fully complete: the record provides a usable site label, numerical specimen counts, quantified sampling effort, and specific habitat observations; route to the identification queue."], "instructions": "Assign the survey’s completeness level using the ordered rubric, then apply the stated routing rule. Treat blank fields as missing; do not infer sampling effort or habitat from the specimen counts.", "type": "score"}}, "state": "Field note: At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a nonblank E17 entry of “36.” In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes. The site registry confirms that Brook-47 is a current, unambiguous label for one stream location. Review found no continuation pages, attachments, marginal notes, or later amendments associated with Mara’s submission. Ecology data manager Chen will assign completeness and routing. Taxonomic identifier Ivo remains available for organism review, and watershed coordinator Sal receives completed records. Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification."}, "method": "c2d", "provenance": {"source_id": "diverse-251", "source_is_synthetic": true, "source_sha256": "b5f8de89182cbde93b1bee659ab602b8887a792686d45bd079fa28b2eb909afd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions and preserve the original workflow without adding exceptions or defaults. The contexts remain bound to Mara’s single stream-survey completeness decision, and the changed site, counts, timestamp, and field contents are permissible observation changes. The focus evidence consists of exactly two complete factual sentences. Changing E17 from a nonblank numeric entry to blank is a coherent single-sentence counterfactual and does not conflict with the unchanged statement that E17 is the sole sampling-effort field or with the absence of attachments and amendments. Neither context states an assigned score, route, gold answer, proposition identifier, or output instruction; the general routing language is preserved policy rather than answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_counts": "supported", "a_effort": "refuted", "a_habitat": "refuted", "a_site": "supported"}, "full_context_fact_states": {"base": {"a_counts": "supported", "a_effort": "supported", "a_habitat": "refuted", "a_site": "supported"}, "counterfactual": {"a_counts": "supported", "a_effort": "refuted", "a_habitat": "refuted", "a_site": "supported"}, "remove_left": {"a_effort": "unknown"}, "remove_right": {"a_effort": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_effort": "unknown"}, "negative_pair": {"a_effort": "refuted"}, "negative_sentence": {"a_effort": "unknown"}, "positive_pair": {"a_effort": "supported"}, "right": {"a_effort": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship about the survey record, and the focus atom concerns the factual presence of sampling-effort information rather than a policy conclusion. The base and counter assignments are realizable with only that fact changing: the base yields level 2 and the counter yields level 1. The policy evidence accurately cites a workflow rule from the original state; all remaining governing criteria are already preserved verbatim in the questions object, while state-specific observations and role assignments need not be carried into synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A usable site label and numerical specimen counts are present, sampling effort is supplied, and habitat observations are absent. Exactly one of the two additional categories is therefore supplied, which is sufficient for level 2 under the rubric.", "rule_index": 0, "sound": true}, {"reason": "A usable site label and numerical specimen counts are present, while both sampling effort and habitat observations are absent. This is sufficient for level 1 under the rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_site", "statement": "The one stream survey record submitted by Mara for this completeness decision provides a usable stream-site label."}, {"id": "a_counts", "statement": "The one stream survey record submitted by Mara for this completeness decision provides numerical specimen counts."}, {"id": "a_effort", "statement": "The one stream survey record submitted by Mara for this completeness decision supplies sampling-effort information."}, {"id": "a_habitat", "statement": "The one stream survey record submitted by Mara for this completeness decision supplies habitat observations."}], "base_state_json": "\"Field note: At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a nonblank E17 entry of “36.” In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes. The site registry confirms that Brook-47 is a current, unambiguous label for one stream location. Review found no continuation pages, attachments, marginal notes, or later amendments associated with Mara’s submission. Ecology data manager Chen will assign completeness and routing. Taxonomic identifier Ivo remains available for organism review, and watershed coordinator Sal receives completed records. Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification.\"", "base_states": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "supported"}, {"atom_id": "a_habitat", "state": "refuted"}], "counter_states": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "refuted"}, {"atom_id": "a_habitat", "state": "refuted"}], "focus_atom": "a_effort", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a nonblank E17 entry of “36.”"}, {"path": [], "text": "In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes."}], "policy_evidence": [{"path": [], "text": "Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification."}], "rules": [{"justification": "A usable site label and numerical specimen counts are present, and sampling effort is supplied while habitat observations are missing. Thus exactly one of the two additional categories is supplied, which is sufficient for score 2 and excludes scores 0, 1, 3, and 4.", "target": "2", "when": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "supported"}, {"atom_id": "a_habitat", "state": "refuted"}]}, {"justification": "A usable site label and numerical specimen counts are present, while both sampling-effort information and habitat observations are missing. This is sufficient for score 1 and excludes score 0 and all scores requiring either additional category.", "target": "1", "when": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "refuted"}, {"atom_id": "a_habitat", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a nonblank E17 entry of “36.”", "negative_left": "At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a blank E17 entry.", "negative_right": "In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes.", "right": "In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes."}, "verifier_independent_model": false}, "family": "scale-diverse-251-004", "id": "scale-diverse-251-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: the record lacks either a usable stream-site label or any specimen counts; route to the sampling queue.", "1 — Minimally complete: a usable site label and specimen counts are present, but both sampling-effort information and habitat observations are missing; route to the sampling queue.", "2 — Partially complete: a usable site label and specimen counts are present, and exactly one of sampling effort or habitat observations is supplied; route to the data-review queue.", "3 — Nearly complete: site label, specimen counts, sampling effort, and habitat observations are all present, but at least one supplied category is vague or non-quantitative, such as “sampled briefly” or “normal habitat”; route to the data-review queue.", "4 — Fully complete: the record provides a usable site label, numerical specimen counts, quantified sampling effort, and specific habitat observations; route to the identification queue."], "instructions": "Assign the survey’s completeness level using the ordered rubric, then apply the stated routing rule. Treat blank fields as missing; do not infer sampling effort or habitat from the specimen counts.", "type": "score"}}, "state": "Field note: At 14:20 UTC on 8 May 2026, the one stream survey record submitted by Mara for this completeness decision listed site label Brook-47, specimen counts of 13 minnows and 6 darters, no habitat observations, and a blank E17 entry. In the one stream survey record submitted by Mara for this completeness decision, E17 is the sole field designated to report sampling effort in minutes. The site registry confirms that Brook-47 is a current, unambiguous label for one stream location. Review found no continuation pages, attachments, marginal notes, or later amendments associated with Mara’s submission. Ecology data manager Chen will assign completeness and routing. Taxonomic identifier Ivo remains available for organism review, and watershed coordinator Sal receives completed records. Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification."}, "method": "c2d", "provenance": {"source_id": "diverse-251", "source_is_synthetic": true, "source_sha256": "b5f8de89182cbde93b1bee659ab602b8887a792686d45bd079fa28b2eb909afd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts are evaluated with the unchanged questions object, which preserves the complete scoring rubric and instructions, and both retain the original state's workflow routing policy without alteration. The request remains bound to Mara's single stream-survey completeness decision; changing the site, counts, and sampling-effort observation is permitted case variation. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently replaces documented quantified sampling effort with an explicitly blank effort field and absence of effort information elsewhere, while remaining consistent with the unchanged missing habitat observations. Neither context states a resulting score, route, gold label, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_counts": "supported", "a_effort": "supported", "a_habitat": "refuted", "a_site": "supported"}, "full_context_fact_states": {"base": {"a_counts": "supported", "a_effort": "supported", "a_habitat": "refuted", "a_site": "supported"}, "counterfactual": {"a_counts": "supported", "a_effort": "refuted", "a_habitat": "refuted", "a_site": "supported"}, "remove_left": {"a_effort": "unknown"}, "remove_right": {"a_effort": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_effort": "unknown"}, "negative_pair": {"a_effort": "refuted"}, "negative_sentence": {"a_effort": "unknown"}, "positive_pair": {"a_effort": "supported"}, "right": {"a_effort": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship about the survey record, and the focus atom concerns the factual presence of sampling-effort information rather than a policy conclusion. The base and counter assignments are realizable with only that fact changing: the base yields level 2 and the counter yields level 1. The policy evidence accurately cites a workflow rule from the original state; all remaining governing criteria are already preserved verbatim in the questions object, while state-specific observations and role assignments need not be carried into synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A usable site label and numerical specimen counts are present, sampling effort is supplied, and habitat observations are absent. Exactly one of the two additional categories is therefore supplied, which is sufficient for level 2 under the rubric.", "rule_index": 0, "sound": true}, {"reason": "A usable site label and numerical specimen counts are present, while both sampling effort and habitat observations are absent. This is sufficient for level 1 under the rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_site", "statement": "The one stream survey record submitted by Mara for this completeness decision provides a usable stream-site label."}, {"id": "a_counts", "statement": "The one stream survey record submitted by Mara for this completeness decision provides numerical specimen counts."}, {"id": "a_effort", "statement": "The one stream survey record submitted by Mara for this completeness decision supplies sampling-effort information."}, {"id": "a_habitat", "statement": "The one stream survey record submitted by Mara for this completeness decision supplies habitat observations."}], "base_state_json": "\"At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank. The station registry confirmed that the label uniquely identified an active stream site. No continuation sheets or attachments accompanied QN-583, and the records clerk found no habitat observations elsewhere in the record. At 09:06 UTC on 14 May 2026, an audit of file QN-583 found that its sampling entry documented two 18-minute net passes covering 90 meters. Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification.\"", "base_states": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "supported"}, {"atom_id": "a_habitat", "state": "refuted"}], "counter_states": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "refuted"}, {"atom_id": "a_habitat", "state": "refuted"}], "focus_atom": "a_effort", "focus_evidence": [{"path": [], "text": "At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank."}, {"path": [], "text": "At 09:06 UTC on 14 May 2026, an audit of file QN-583 found that its sampling entry documented two 18-minute net passes covering 90 meters."}], "policy_evidence": [{"path": [], "text": "Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification."}], "rules": [{"justification": "A usable site label and numerical specimen counts are present, and sampling effort is supplied while habitat observations are missing. Thus exactly one of the two additional categories is supplied, which is sufficient for score 2 and excludes scores 0, 1, 3, and 4.", "target": "2", "when": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "supported"}, {"atom_id": "a_habitat", "state": "refuted"}]}, {"justification": "A usable site label and numerical specimen counts are present, while both sampling-effort information and habitat observations are missing. This is sufficient for score 1 and excludes score 0 and all scores requiring either additional category.", "target": "1", "when": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "refuted"}, {"atom_id": "a_habitat", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank.", "negative_left": "At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank.", "negative_right": "At 09:06 UTC on 14 May 2026, an audit of file QN-583 found its sampling-effort field blank and found no sampling duration, distance, method, or other effort information elsewhere in the file.", "right": "At 09:06 UTC on 14 May 2026, an audit of file QN-583 found that its sampling entry documented two 18-minute net passes covering 90 meters."}, "verifier_independent_model": false}, "family": "scale-diverse-251-005", "id": "scale-diverse-251-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: the record lacks either a usable stream-site label or any specimen counts; route to the sampling queue.", "1 — Minimally complete: a usable site label and specimen counts are present, but both sampling-effort information and habitat observations are missing; route to the sampling queue.", "2 — Partially complete: a usable site label and specimen counts are present, and exactly one of sampling effort or habitat observations is supplied; route to the data-review queue.", "3 — Nearly complete: site label, specimen counts, sampling effort, and habitat observations are all present, but at least one supplied category is vague or non-quantitative, such as “sampled briefly” or “normal habitat”; route to the data-review queue.", "4 — Fully complete: the record provides a usable site label, numerical specimen counts, quantified sampling effort, and specific habitat observations; route to the identification queue."], "instructions": "Assign the survey’s completeness level using the ordered rubric, then apply the stated routing rule. Treat blank fields as missing; do not infer sampling effort or habitat from the specimen counts.", "type": "score"}}, "state": "At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank. The station registry confirmed that the label uniquely identified an active stream site. No continuation sheets or attachments accompanied QN-583, and the records clerk found no habitat observations elsewhere in the record. At 09:06 UTC on 14 May 2026, an audit of file QN-583 found that its sampling entry documented two 18-minute net passes covering 90 meters. Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification."}, "method": "c2d", "provenance": {"source_id": "diverse-251", "source_is_synthetic": true, "source_sha256": "b5f8de89182cbde93b1bee659ab602b8887a792686d45bd079fa28b2eb909afd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts are evaluated with the unchanged questions object, which preserves the complete scoring rubric and instructions, and both retain the original state's workflow routing policy without alteration. The request remains bound to Mara's single stream-survey completeness decision; changing the site, counts, and sampling-effort observation is permitted case variation. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently replaces documented quantified sampling effort with an explicitly blank effort field and absence of effort information elsewhere, while remaining consistent with the unchanged missing habitat observations. Neither context states a resulting score, route, gold label, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_counts": "supported", "a_effort": "refuted", "a_habitat": "refuted", "a_site": "supported"}, "full_context_fact_states": {"base": {"a_counts": "supported", "a_effort": "supported", "a_habitat": "refuted", "a_site": "supported"}, "counterfactual": {"a_counts": "supported", "a_effort": "refuted", "a_habitat": "refuted", "a_site": "supported"}, "remove_left": {"a_effort": "unknown"}, "remove_right": {"a_effort": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_effort": "unknown"}, "negative_pair": {"a_effort": "refuted"}, "negative_sentence": {"a_effort": "unknown"}, "positive_pair": {"a_effort": "supported"}, "right": {"a_effort": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship about the survey record, and the focus atom concerns the factual presence of sampling-effort information rather than a policy conclusion. The base and counter assignments are realizable with only that fact changing: the base yields level 2 and the counter yields level 1. The policy evidence accurately cites a workflow rule from the original state; all remaining governing criteria are already preserved verbatim in the questions object, while state-specific observations and role assignments need not be carried into synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A usable site label and numerical specimen counts are present, sampling effort is supplied, and habitat observations are absent. Exactly one of the two additional categories is therefore supplied, which is sufficient for level 2 under the rubric.", "rule_index": 0, "sound": true}, {"reason": "A usable site label and numerical specimen counts are present, while both sampling effort and habitat observations are absent. This is sufficient for level 1 under the rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_site", "statement": "The one stream survey record submitted by Mara for this completeness decision provides a usable stream-site label."}, {"id": "a_counts", "statement": "The one stream survey record submitted by Mara for this completeness decision provides numerical specimen counts."}, {"id": "a_effort", "statement": "The one stream survey record submitted by Mara for this completeness decision supplies sampling-effort information."}, {"id": "a_habitat", "statement": "The one stream survey record submitted by Mara for this completeness decision supplies habitat observations."}], "base_state_json": "\"At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank. The station registry confirmed that the label uniquely identified an active stream site. No continuation sheets or attachments accompanied QN-583, and the records clerk found no habitat observations elsewhere in the record. At 09:06 UTC on 14 May 2026, an audit of file QN-583 found that its sampling entry documented two 18-minute net passes covering 90 meters. Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification.\"", "base_states": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "supported"}, {"atom_id": "a_habitat", "state": "refuted"}], "counter_states": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "refuted"}, {"atom_id": "a_habitat", "state": "refuted"}], "focus_atom": "a_effort", "focus_evidence": [{"path": [], "text": "At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank."}, {"path": [], "text": "At 09:06 UTC on 14 May 2026, an audit of file QN-583 found that its sampling entry documented two 18-minute net passes covering 90 meters."}], "policy_evidence": [{"path": [], "text": "Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification."}], "rules": [{"justification": "A usable site label and numerical specimen counts are present, and sampling effort is supplied while habitat observations are missing. Thus exactly one of the two additional categories is supplied, which is sufficient for score 2 and excludes scores 0, 1, 3, and 4.", "target": "2", "when": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "supported"}, {"atom_id": "a_habitat", "state": "refuted"}]}, {"justification": "A usable site label and numerical specimen counts are present, while both sampling-effort information and habitat observations are missing. This is sufficient for score 1 and excludes score 0 and all scores requiring either additional category.", "target": "1", "when": [{"atom_id": "a_site", "state": "supported"}, {"atom_id": "a_counts", "state": "supported"}, {"atom_id": "a_effort", "state": "refuted"}, {"atom_id": "a_habitat", "state": "refuted"}]}]}, "verified_pair": {"left": "At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank.", "negative_left": "At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank.", "negative_right": "At 09:06 UTC on 14 May 2026, an audit of file QN-583 found its sampling-effort field blank and found no sampling duration, distance, method, or other effort information elsewhere in the file.", "right": "At 09:06 UTC on 14 May 2026, an audit of file QN-583 found that its sampling entry documented two 18-minute net passes covering 90 meters."}, "verifier_independent_model": false}, "family": "scale-diverse-251-005", "id": "scale-diverse-251-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: the record lacks either a usable stream-site label or any specimen counts; route to the sampling queue.", "1 — Minimally complete: a usable site label and specimen counts are present, but both sampling-effort information and habitat observations are missing; route to the sampling queue.", "2 — Partially complete: a usable site label and specimen counts are present, and exactly one of sampling effort or habitat observations is supplied; route to the data-review queue.", "3 — Nearly complete: site label, specimen counts, sampling effort, and habitat observations are all present, but at least one supplied category is vague or non-quantitative, such as “sampled briefly” or “normal habitat”; route to the data-review queue.", "4 — Fully complete: the record provides a usable site label, numerical specimen counts, quantified sampling effort, and specific habitat observations; route to the identification queue."], "instructions": "Assign the survey’s completeness level using the ordered rubric, then apply the stated routing rule. Treat blank fields as missing; do not infer sampling effort or habitat from the specimen counts.", "type": "score"}}, "state": "At 08:42 UTC on 14 May 2026, Mara submitted exactly one stream survey record for the completeness decision, file QN-583, which labeled the site “Willow Bend 7,” reported 26 dace and 11 sculpin, and left its habitat-observations section blank. The station registry confirmed that the label uniquely identified an active stream site. No continuation sheets or attachments accompanied QN-583, and the records clerk found no habitat observations elsewhere in the record. At 09:06 UTC on 14 May 2026, an audit of file QN-583 found its sampling-effort field blank and found no sampling duration, distance, method, or other effort information elsewhere in the file. Under the workflow, ratings 0–1 go to the sampling queue, 2–3 to data review, and 4 to identification."}, "method": "c2d", "provenance": {"source_id": "diverse-251", "source_is_synthetic": true, "source_sha256": "b5f8de89182cbde93b1bee659ab602b8887a792686d45bd079fa28b2eb909afd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same request, three-site Pine Run scope, wet/dry classifications, and governing completeness rubric without adding exceptions, priorities, or missing-evidence defaults. The two evidence spans are complete factual sentences. The counterfactual changes only PR-2’s finalized count-sheet value from 37 to 34; this creates a coherent disagreement between two distinct records rather than contradictory duplicate assertions, and the closing-audit statement expressly accommodates that possible discrepancy. Neither context contains a score, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"On 18 June 2026, the route sheet scoped exactly Pine Run sites PR-1, PR-2, and PR-3. Field staff classified PR-1 and PR-2 as wet and PR-3 as dry.\",\"chronology\":[\"At PR-1, staff logged a 20-minute kick-net effort, a mappable site label, and habitat observations. Its jar and finalized count sheet each recorded 24 specimens.\",\"At PR-2, staff logged a 20-minute kick-net effort, a mappable site label, and habitat observations.\",\"At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens.\",\"At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 37 specimens.\",\"For PR-3, the submitted dry-site photo was dated 18 June 2026 and documented the absence of flowing water; habitat observations accompanied it.\",\"The closing audit found no unresolved required-field deficiency across the three sites except any possible deficiency represented by the PR-2 specimen-count comparison. It likewise found no unresolved record discrepancy except any possible discrepancy between those PR-2 records. No nonrequired formatting, wording, or organizational irregularity remained in the packet.\"],\"request\":\"Rate the packet’s completeness confidence under the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["chronology", "2"], "text": "At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens."}, {"path": ["chronology", "3"], "text": "At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 37 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens.", "negative_left": "At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens.", "negative_right": "At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 34 specimens.", "right": "At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 37 specimens."}, "verifier_independent_model": false}, "family": "scale-diverse-252-001", "id": "scale-diverse-252-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"chronology": ["At PR-1, staff logged a 20-minute kick-net effort, a mappable site label, and habitat observations. Its jar and finalized count sheet each recorded 24 specimens.", "At PR-2, staff logged a 20-minute kick-net effort, a mappable site label, and habitat observations.", "At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens.", "At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 37 specimens.", "For PR-3, the submitted dry-site photo was dated 18 June 2026 and documented the absence of flowing water; habitat observations accompanied it.", "The closing audit found no unresolved required-field deficiency across the three sites except any possible deficiency represented by the PR-2 specimen-count comparison. It likewise found no unresolved record discrepancy except any possible discrepancy between those PR-2 records. No nonrequired formatting, wording, or organizational irregularity remained in the packet."], "context": "On 18 June 2026, the route sheet scoped exactly Pine Run sites PR-1, PR-2, and PR-3. Field staff classified PR-1 and PR-2 as wet and PR-3 as dry.", "request": "Rate the packet’s completeness confidence under the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same request, three-site Pine Run scope, wet/dry classifications, and governing completeness rubric without adding exceptions, priorities, or missing-evidence defaults. The two evidence spans are complete factual sentences. The counterfactual changes only PR-2’s finalized count-sheet value from 37 to 34; this creates a coherent disagreement between two distinct records rather than contradictory duplicate assertions, and the closing-audit statement expressly accommodates that possible discrepancy. Neither context contains a score, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"On 18 June 2026, the route sheet scoped exactly Pine Run sites PR-1, PR-2, and PR-3. Field staff classified PR-1 and PR-2 as wet and PR-3 as dry.\",\"chronology\":[\"At PR-1, staff logged a 20-minute kick-net effort, a mappable site label, and habitat observations. Its jar and finalized count sheet each recorded 24 specimens.\",\"At PR-2, staff logged a 20-minute kick-net effort, a mappable site label, and habitat observations.\",\"At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens.\",\"At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 37 specimens.\",\"For PR-3, the submitted dry-site photo was dated 18 June 2026 and documented the absence of flowing water; habitat observations accompanied it.\",\"The closing audit found no unresolved required-field deficiency across the three sites except any possible deficiency represented by the PR-2 specimen-count comparison. It likewise found no unresolved record discrepancy except any possible discrepancy between those PR-2 records. No nonrequired formatting, wording, or organizational irregularity remained in the packet.\"],\"request\":\"Rate the packet’s completeness confidence under the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["chronology", "2"], "text": "At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens."}, {"path": ["chronology", "3"], "text": "At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 37 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens.", "negative_left": "At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens.", "negative_right": "At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 34 specimens.", "right": "At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 37 specimens."}, "verifier_independent_model": false}, "family": "scale-diverse-252-001", "id": "scale-diverse-252-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"chronology": ["At PR-1, staff logged a 20-minute kick-net effort, a mappable site label, and habitat observations. Its jar and finalized count sheet each recorded 24 specimens.", "At PR-2, staff logged a 20-minute kick-net effort, a mappable site label, and habitat observations.", "At 14:20 UTC on 18 June 2026, the completed inventory of PR-2's jar recorded 37 specimens.", "At 14:35 UTC on 18 June 2026, PR-2's finalized count sheet recorded 34 specimens.", "For PR-3, the submitted dry-site photo was dated 18 June 2026 and documented the absence of flowing water; habitat observations accompanied it.", "The closing audit found no unresolved required-field deficiency across the three sites except any possible deficiency represented by the PR-2 specimen-count comparison. It likewise found no unresolved record discrepancy except any possible discrepancy between those PR-2 records. No nonrequired formatting, wording, or organizational irregularity remained in the packet."], "context": "On 18 June 2026, the route sheet scoped exactly Pine Run sites PR-1, PR-2, and PR-3. Field staff classified PR-1 and PR-2 as wet and PR-3 as dry.", "request": "Rate the packet’s completeness confidence under the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric, three-site scope, request, and PR-2 specimen-count comparison without adding exceptions, priorities, or missing-evidence defaults. The two focus spans are complete factual sentences; changing only the finalized PR-2 count-sheet value from 47 to 46 creates a coherent jar-versus-sheet discrepancy rather than contradictory duplicate measurements, and neither context embeds an answer, code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager reconciled the field packet against the route sheet before routing. The route sheet’s complete scoped-site list is PR-1, PR-2, and PR-3.\",\"evidence\":[\"PR-1 is wet. Its record shows a 20-minute kick-net effort, a mappable site label, 31 specimens on both the jar and count sheet, and habitat observations.\",\"PR-2 is wet. Its record includes a 20-minute kick-net effort, a mappable site label, and habitat observations.\",\"At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar.\",\"The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 47 specimens for PR-2.\",\"PR-3 is dry. Its submitted dry-site photo is dated and documents the absence of flowing water at PR-3; habitat observations are also present.\",\"Across the three scoped sites, no unresolved required-field deficiency exists except any possible deficiency represented by the PR-2 specimen-count comparison. No unresolved record discrepancy exists except any possible discrepancy between that jar count and count-sheet count.\",\"The reconciliation found no remaining nonrequired formatting, wording, or organizational irregularity in the packet.\"],\"request\":\"Rate the packet’s completeness confidence under the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar."}, {"path": ["evidence", "3"], "text": "The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 47 specimens for PR-2."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar.", "negative_left": "At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar.", "negative_right": "The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 46 specimens for PR-2.", "right": "The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 47 specimens for PR-2."}, "verifier_independent_model": false}, "family": "scale-diverse-252-002", "id": "scale-diverse-252-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reconciled the field packet against the route sheet before routing. The route sheet’s complete scoped-site list is PR-1, PR-2, and PR-3.", "evidence": ["PR-1 is wet. Its record shows a 20-minute kick-net effort, a mappable site label, 31 specimens on both the jar and count sheet, and habitat observations.", "PR-2 is wet. Its record includes a 20-minute kick-net effort, a mappable site label, and habitat observations.", "At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar.", "The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 47 specimens for PR-2.", "PR-3 is dry. Its submitted dry-site photo is dated and documents the absence of flowing water at PR-3; habitat observations are also present.", "Across the three scoped sites, no unresolved required-field deficiency exists except any possible deficiency represented by the PR-2 specimen-count comparison. No unresolved record discrepancy exists except any possible discrepancy between that jar count and count-sheet count.", "The reconciliation found no remaining nonrequired formatting, wording, or organizational irregularity in the packet."], "request": "Rate the packet’s completeness confidence under the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric, three-site scope, request, and PR-2 specimen-count comparison without adding exceptions, priorities, or missing-evidence defaults. The two focus spans are complete factual sentences; changing only the finalized PR-2 count-sheet value from 47 to 46 creates a coherent jar-versus-sheet discrepancy rather than contradictory duplicate measurements, and neither context embeds an answer, code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager reconciled the field packet against the route sheet before routing. The route sheet’s complete scoped-site list is PR-1, PR-2, and PR-3.\",\"evidence\":[\"PR-1 is wet. Its record shows a 20-minute kick-net effort, a mappable site label, 31 specimens on both the jar and count sheet, and habitat observations.\",\"PR-2 is wet. Its record includes a 20-minute kick-net effort, a mappable site label, and habitat observations.\",\"At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar.\",\"The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 47 specimens for PR-2.\",\"PR-3 is dry. Its submitted dry-site photo is dated and documents the absence of flowing water at PR-3; habitat observations are also present.\",\"Across the three scoped sites, no unresolved required-field deficiency exists except any possible deficiency represented by the PR-2 specimen-count comparison. No unresolved record discrepancy exists except any possible discrepancy between that jar count and count-sheet count.\",\"The reconciliation found no remaining nonrequired formatting, wording, or organizational irregularity in the packet.\"],\"request\":\"Rate the packet’s completeness confidence under the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar."}, {"path": ["evidence", "3"], "text": "The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 47 specimens for PR-2."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar.", "negative_left": "At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar.", "negative_right": "The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 46 specimens for PR-2.", "right": "The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 47 specimens for PR-2."}, "verifier_independent_model": false}, "family": "scale-diverse-252-002", "id": "scale-diverse-252-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reconciled the field packet against the route sheet before routing. The route sheet’s complete scoped-site list is PR-1, PR-2, and PR-3.", "evidence": ["PR-1 is wet. Its record shows a 20-minute kick-net effort, a mappable site label, 31 specimens on both the jar and count sheet, and habitat observations.", "PR-2 is wet. Its record includes a 20-minute kick-net effort, a mappable site label, and habitat observations.", "At 14:10 UTC on 8 September 2026, a complete physical count found exactly 47 specimens in PR-2's jar.", "The finalized PR-2 count sheet, signed at 14:25 UTC on 8 September 2026, records exactly 46 specimens for PR-2.", "PR-3 is dry. Its submitted dry-site photo is dated and documents the absence of flowing water at PR-3; habitat observations are also present.", "Across the three scoped sites, no unresolved required-field deficiency exists except any possible deficiency represented by the PR-2 specimen-count comparison. No unresolved record discrepancy exists except any possible discrepancy between that jar count and count-sheet count.", "The reconciliation found no remaining nonrequired formatting, wording, or organizational irregularity in the packet."], "request": "Rate the packet’s completeness confidence under the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scoring criteria, wet-site requirements, dry-site exception, discrepancy treatment, three-site scope, score request, and PR-1/PR-2/PR-3 bindings. The two focus spans are complete factual sentences. The counterfactual changes only PR-2’s count-sheet total from 37 to 41, creating one coherent jar-versus-sheet disagreement without conflicting with the hedged closing-audit statement or any unchanged evidence. Neither context states a score, answer code, classifier instruction, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.\",\"1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.\",\"2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.\",\"3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.\",\"4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain.\"],\"instructions\":\"Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.\",\"type\":\"score\"}},\"state\":{\"context\":\"The survey packet is ready for an operational handoff review. The route sheet scopes exactly PR-1, PR-2, and PR-3; all observations below were checked against the submitted field records.\",\"evidence\":[\"PR-1 is wet. It has recorded kick-net effort, a site label that resolves to the route map, habitat observations, and identical specimen totals on its jar and count sheet.\",\"At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens.\",\"At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 37 specimens.\",\"PR-2 is wet and has recorded kick-net effort, a mappable site label, and habitat observations.\",\"PR-3 is dry. Its submitted dry-site photo is dated and documents no flowing water at the site; habitat observations are also present.\",\"The closing audit found no unresolved required-field deficiency across the three sites except any possible deficiency represented by the PR-2 specimen-count comparison, and no unresolved record discrepancy except any possible discrepancy between those two PR-2 records.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"],\"request\":\"Rate the packet’s completeness confidence under the supplied rubric.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens."}, {"path": ["state", "evidence", "2"], "text": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 37 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens.", "negative_left": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens.", "negative_right": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 41 specimens.", "right": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 37 specimens."}, "verifier_independent_model": false}, "family": "scale-diverse-252-003", "id": "scale-diverse-252-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "The survey packet is ready for an operational handoff review. The route sheet scopes exactly PR-1, PR-2, and PR-3; all observations below were checked against the submitted field records.", "evidence": ["PR-1 is wet. It has recorded kick-net effort, a site label that resolves to the route map, habitat observations, and identical specimen totals on its jar and count sheet.", "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens.", "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 37 specimens.", "PR-2 is wet and has recorded kick-net effort, a mappable site label, and habitat observations.", "PR-3 is dry. Its submitted dry-site photo is dated and documents no flowing water at the site; habitat observations are also present.", "The closing audit found no unresolved required-field deficiency across the three sites except any possible deficiency represented by the PR-2 specimen-count comparison, and no unresolved record discrepancy except any possible discrepancy between those two PR-2 records.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."], "request": "Rate the packet’s completeness confidence under the supplied rubric."}}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original scoring criteria, wet-site requirements, dry-site exception, discrepancy treatment, three-site scope, score request, and PR-1/PR-2/PR-3 bindings. The two focus spans are complete factual sentences. The counterfactual changes only PR-2’s count-sheet total from 37 to 41, creating one coherent jar-versus-sheet disagreement without conflicting with the hedged closing-audit statement or any unchanged evidence. Neither context states a score, answer code, classifier instruction, proposition identifier, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"questions\":{\"decision\":{\"criteria\":[\"0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.\",\"1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.\",\"2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.\",\"3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.\",\"4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain.\"],\"instructions\":\"Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.\",\"type\":\"score\"}},\"state\":{\"context\":\"The survey packet is ready for an operational handoff review. The route sheet scopes exactly PR-1, PR-2, and PR-3; all observations below were checked against the submitted field records.\",\"evidence\":[\"PR-1 is wet. It has recorded kick-net effort, a site label that resolves to the route map, habitat observations, and identical specimen totals on its jar and count sheet.\",\"At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens.\",\"At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 37 specimens.\",\"PR-2 is wet and has recorded kick-net effort, a mappable site label, and habitat observations.\",\"PR-3 is dry. Its submitted dry-site photo is dated and documents no flowing water at the site; habitat observations are also present.\",\"The closing audit found no unresolved required-field deficiency across the three sites except any possible deficiency represented by the PR-2 specimen-count comparison, and no unresolved record discrepancy except any possible discrepancy between those two PR-2 records.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"],\"request\":\"Rate the packet’s completeness confidence under the supplied rubric.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["state", "evidence", "1"], "text": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens."}, {"path": ["state", "evidence", "2"], "text": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 37 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens.", "negative_left": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens.", "negative_right": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 41 specimens.", "right": "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 37 specimens."}, "verifier_independent_model": false}, "family": "scale-diverse-252-003", "id": "scale-diverse-252-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "The survey packet is ready for an operational handoff review. The route sheet scopes exactly PR-1, PR-2, and PR-3; all observations below were checked against the submitted field records.", "evidence": ["PR-1 is wet. It has recorded kick-net effort, a site label that resolves to the route map, habitat observations, and identical specimen totals on its jar and count sheet.", "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole specimen jar contained exactly 37 specimens.", "At the 2026-08-14 16:20 UTC operational handoff, PR-2's sole count sheet listed exactly 41 specimens.", "PR-2 is wet and has recorded kick-net effort, a mappable site label, and habitat observations.", "PR-3 is dry. Its submitted dry-site photo is dated and documents no flowing water at the site; habitat observations are also present.", "The closing audit found no unresolved required-field deficiency across the three sites except any possible deficiency represented by the PR-2 specimen-count comparison, and no unresolved record discrepancy except any possible discrepancy between those two PR-2 records.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."], "request": "Rate the packet’s completeness confidence under the supplied rubric."}}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, request scope, and relevant segmentation-rerun bindings, while the unchanged questions preserve the full criteria and instructions. The two focus-evidence spans are complete factual sentences. The counterfactual changes only SR-47's count from 2,675 to 2,600, which remains coherent with CR-46's count of 2,500 and creates no duplicate or conflicting measurement. Neither context contains an answer, answer code, proposition identifier, label rationale, or output instruction beyond the preserved request.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500.\",\"The validated output for segmentation rerun SR-47 records a cell count of 2,675.\",\"The current calibration release is marked pass. The latest same-session bead test measured a channel offset of 0.4 px, with no later bead test logged.\",\"The most recent relevant slide review found a fixed slide defect in 2 of 30 reviewed fields.\",\"The current technical QC is marked pass, and the most recent relevant replicate coefficient of variation is 8%.\",\"Operations confirmed that these records are validated and current for the handoff.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500."}, {"path": ["evidence", "1"], "text": "The validated output for segmentation rerun SR-47 records a cell count of 2,675."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500.", "negative_left": "The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500.", "negative_right": "The validated output for segmentation rerun SR-47 records a cell count of 2,600.", "right": "The validated output for segmentation rerun SR-47 records a cell count of 2,675."}, "verifier_independent_model": false}, "family": "scale-diverse-259-003", "id": "scale-diverse-259-003-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500.", "The validated output for segmentation rerun SR-47 records a cell count of 2,675.", "The current calibration release is marked pass. The latest same-session bead test measured a channel offset of 0.4 px, with no later bead test logged.", "The most recent relevant slide review found a fixed slide defect in 2 of 30 reviewed fields.", "The current technical QC is marked pass, and the most recent relevant replicate coefficient of variation is 8%.", "Operations confirmed that these records are validated and current for the handoff."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, request scope, and relevant segmentation-rerun bindings, while the unchanged questions preserve the full criteria and instructions. The two focus-evidence spans are complete factual sentences. The counterfactual changes only SR-47's count from 2,675 to 2,600, which remains coherent with CR-46's count of 2,500 and creates no duplicate or conflicting measurement. Neither context contains an answer, answer code, proposition identifier, label rationale, or output instruction beyond the preserved request.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500.\",\"The validated output for segmentation rerun SR-47 records a cell count of 2,675.\",\"The current calibration release is marked pass. The latest same-session bead test measured a channel offset of 0.4 px, with no later bead test logged.\",\"The most recent relevant slide review found a fixed slide defect in 2 of 30 reviewed fields.\",\"The current technical QC is marked pass, and the most recent relevant replicate coefficient of variation is 8%.\",\"Operations confirmed that these records are validated and current for the handoff.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500."}, {"path": ["evidence", "1"], "text": "The validated output for segmentation rerun SR-47 records a cell count of 2,675."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500.", "negative_left": "The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500.", "negative_right": "The validated output for segmentation rerun SR-47 records a cell count of 2,600.", "right": "The validated output for segmentation rerun SR-47 records a cell count of 2,675."}, "verifier_independent_model": false}, "family": "scale-diverse-259-003", "id": "scale-diverse-259-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["The validated handoff log identifies segmentation rerun SR-47, completed at 2026-09-17T14:20:00Z, as the most recent relevant segmentation rerun, identifies run CR-46 as its comparison run on the same raw image dataset, and records CR-46's cell count as 2,500.", "The validated output for segmentation rerun SR-47 records a cell count of 2,600.", "The current calibration release is marked pass. The latest same-session bead test measured a channel offset of 0.4 px, with no later bead test logged.", "The most recent relevant slide review found a fixed slide defect in 2 of 30 reviewed fields.", "The current technical QC is marked pass, and the most recent relevant replicate coefficient of variation is 8%.", "Operations confirmed that these records are validated and current for the handoff."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, request scope, latest-evidence requirement, and segmentation comparison on unchanged raw images. The two focus spans are complete factual sentences; changing R-41 from 2,532 to 2,518 cells creates no duplicate or contradictory measurement, remains coherent with the unchanged 2,400-cell comparison run, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41.\",\"Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,532 cells in unchanged raw image dataset D-8.\",\"The latest same-session bead test measured a channel offset of 0.6 px.\",\"The most recent relevant slide review found a fixed defect in 3 of 40 reviewed fields.\",\"Calibration record CAL-9 is current and marked pass.\",\"The most recent relevant replicate coefficient of variation was 11%, and the current technical QC is marked pass.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41."}, {"path": ["evidence", "1"], "text": "Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,532 cells in unchanged raw image dataset D-8."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41.", "negative_left": "As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41.", "negative_right": "Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,518 cells in unchanged raw image dataset D-8.", "right": "Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,532 cells in unchanged raw image dataset D-8."}, "verifier_independent_model": false}, "family": "scale-diverse-259-004", "id": "scale-diverse-259-004-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41.", "Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,532 cells in unchanged raw image dataset D-8.", "The latest same-session bead test measured a channel offset of 0.6 px.", "The most recent relevant slide review found a fixed defect in 3 of 40 reviewed fields.", "Calibration record CAL-9 is current and marked pass.", "The most recent relevant replicate coefficient of variation was 11%, and the current technical QC is marked pass."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, request scope, latest-evidence requirement, and segmentation comparison on unchanged raw images. The two focus spans are complete factual sentences; changing R-41 from 2,532 to 2,518 cells creates no duplicate or contradictory measurement, remains coherent with the unchanged 2,400-cell comparison run, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41.\",\"Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,532 cells in unchanged raw image dataset D-8.\",\"The latest same-session bead test measured a channel offset of 0.6 px.\",\"The most recent relevant slide review found a fixed defect in 3 of 40 reviewed fields.\",\"Calibration record CAL-9 is current and marked pass.\",\"The most recent relevant replicate coefficient of variation was 11%, and the current technical QC is marked pass.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41."}, {"path": ["evidence", "1"], "text": "Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,532 cells in unchanged raw image dataset D-8."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41.", "negative_left": "As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41.", "negative_right": "Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,518 cells in unchanged raw image dataset D-8.", "right": "Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,532 cells in unchanged raw image dataset D-8."}, "verifier_independent_model": false}, "family": "scale-diverse-259-004", "id": "scale-diverse-259-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["As of 2026-09-17 14:00 UTC, segmentation run C-41, which counted 2,400 cells in raw image dataset D-8, is the comparison run for rerun R-41.", "Rerun R-41, completed at 2026-09-17 13:30 UTC and the most recent relevant segmentation rerun as of 2026-09-17 14:00 UTC, counted 2,518 cells in unchanged raw image dataset D-8.", "The latest same-session bead test measured a channel offset of 0.6 px.", "The most recent relevant slide review found a fixed defect in 3 of 40 reviewed fields.", "Calibration record CAL-9 is current and marked pass.", "The most recent relevant replicate coefficient of variation was 11%, and the current technical QC is marked pass."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the routing rubric, threshold semantics, precedence rules, request scope, and case-workflow question. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the recalculated drift from 2.7% to 1.6% for the same log and routing time, without conflicting duplicate evidence or unchanged facts. Neither context embeds a gold answer, output code, rule table, proposition identifier, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"case coordinator\",\"text\":\"The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow.\"},{\"speaker\":\"image analyst\",\"text\":\"At routing time, direct measurements from the supplied raw images showed 2.6% image-confirmed bubble coverage and 0.7% saturated pixels for the case workflow.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 2.7%.\"},{\"speaker\":\"image analyst\",\"text\":\"At routing time, replicate segmentation for the case workflow showed 6.1% disagreement.\"},{\"speaker\":\"assay scientist\",\"text\":\"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["1", "text"], "text": "The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow."}, {"path": ["3", "text"], "text": "At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 2.7%."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow.", "negative_left": "The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow.", "negative_right": "At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 1.6%.", "right": "At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 2.7%."}, "verifier_independent_model": false}, "family": "scale-diverse-260-001", "id": "scale-diverse-260-001-base", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "case coordinator", "text": "The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow."}, {"speaker": "image analyst", "text": "At routing time, direct measurements from the supplied raw images showed 2.6% image-confirmed bubble coverage and 0.7% saturated pixels for the case workflow."}, {"speaker": "calibration reviewer", "text": "At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 2.7%."}, {"speaker": "image analyst", "text": "At routing time, replicate segmentation for the case workflow showed 6.1% disagreement."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_acquisition"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the routing rubric, threshold semantics, precedence rules, request scope, and case-workflow question. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the recalculated drift from 2.7% to 1.6% for the same log and routing time, without conflicting duplicate evidence or unchanged facts. Neither context embeds a gold answer, output code, rule table, proposition identifier, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"case coordinator\",\"text\":\"The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow.\"},{\"speaker\":\"image analyst\",\"text\":\"At routing time, direct measurements from the supplied raw images showed 2.6% image-confirmed bubble coverage and 0.7% saturated pixels for the case workflow.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 2.7%.\"},{\"speaker\":\"image analyst\",\"text\":\"At routing time, replicate segmentation for the case workflow showed 6.1% disagreement.\"},{\"speaker\":\"assay scientist\",\"text\":\"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["1", "text"], "text": "The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow."}, {"path": ["3", "text"], "text": "At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 2.7%."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow.", "negative_left": "The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow.", "negative_right": "At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 1.6%.", "right": "At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 2.7%."}, "verifier_independent_model": false}, "family": "scale-diverse-260-001", "id": "scale-diverse-260-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "case coordinator", "text": "The routing time for the case workflow was 14:32:10 UTC on 8 April 2026, when calibration log CL-47Q was associated with that workflow."}, {"speaker": "image analyst", "text": "At routing time, direct measurements from the supplied raw images showed 2.6% image-confirmed bubble coverage and 0.7% saturated pixels for the case workflow."}, {"speaker": "calibration reviewer", "text": "At 14:32:10 UTC on 8 April 2026, recalculation from calibration log CL-47Q yielded a calibration drift of 1.6%."}, {"speaker": "image analyst", "text": "At routing time, replicate segmentation for the case workflow showed 6.1% disagreement."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the full routing rubric, threshold semantics, ordering, and evidence-precedence rules without adding exceptions or defaults. They retain the same case-routing question and workflow entity, with the added routing timestamp consistently bound across the relevant observations. The two focus-evidence spans are complete factual sentences describing the discrepancy and its recalculation formula. Changing the discrepancy from 1.4 to 0.8 micrometres is coherent with the unchanged 50.0-micrometre reference and creates no duplicate or contradictory measurement. Neither context embeds a selected route, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"image analyst\",\"text\":\"For the case workflow at routing time, direct inspection of the supplied raw images measured exactly 3.0% image-confirmed bubble coverage and exactly 1.0% saturated pixels. Replicate segmentation showed 5.2% disagreement.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 1.4 micrometres.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres.\"},{\"speaker\":\"assay scientist\",\"text\":\"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 1.4 micrometres."}, {"path": ["3", "text"], "text": "At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 1.4 micrometres.", "negative_left": "At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 0.8 micrometres.", "negative_right": "At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres.", "right": "At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres."}, "verifier_independent_model": false}, "family": "scale-diverse-260-002", "id": "scale-diverse-260-002-base", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "For the case workflow at routing time, direct inspection of the supplied raw images measured exactly 3.0% image-confirmed bubble coverage and exactly 1.0% saturated pixels. Replicate segmentation showed 5.2% disagreement."}, {"speaker": "calibration reviewer", "text": "At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 1.4 micrometres."}, {"speaker": "calibration reviewer", "text": "At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_acquisition"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts reproduce the full routing rubric, threshold semantics, ordering, and evidence-precedence rules without adding exceptions or defaults. They retain the same case-routing question and workflow entity, with the added routing timestamp consistently bound across the relevant observations. The two focus-evidence spans are complete factual sentences describing the discrepancy and its recalculation formula. Changing the discrepancy from 1.4 to 0.8 micrometres is coherent with the unchanged 50.0-micrometre reference and creates no duplicate or contradictory measurement. Neither context embeds a selected route, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"image analyst\",\"text\":\"For the case workflow at routing time, direct inspection of the supplied raw images measured exactly 3.0% image-confirmed bubble coverage and exactly 1.0% saturated pixels. Replicate segmentation showed 5.2% disagreement.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 1.4 micrometres.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres.\"},{\"speaker\":\"assay scientist\",\"text\":\"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 1.4 micrometres."}, {"path": ["3", "text"], "text": "At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 1.4 micrometres.", "negative_left": "At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 0.8 micrometres.", "negative_right": "At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres.", "right": "At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres."}, "verifier_independent_model": false}, "family": "scale-diverse-260-002", "id": "scale-diverse-260-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "For the case workflow at routing time, direct inspection of the supplied raw images measured exactly 3.0% image-confirmed bubble coverage and exactly 1.0% saturated pixels. Replicate segmentation showed 5.2% disagreement."}, {"speaker": "calibration reviewer", "text": "At the case workflow's routing time of 2026-09-17 14:20 UTC, the recalculation from its calibration log used an absolute calibration discrepancy of 0.8 micrometres."}, {"speaker": "calibration reviewer", "text": "At the case workflow's routing time of 2026-09-17 14:20 UTC, calibration drift was recalculated as 100 times the absolute calibration discrepancy divided by the logged reference length of 50.0 micrometres."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing routing rubric, precedence rules, thresholds, request scope, case workflow, and routing-time binding, while the unchanged questions object preserves the original decision criteria. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the comparison response from 776.0 to 792.0 units, creating no duplicate or contradictory measurement, and neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"image analyst\",\"text\":\"At routing time, direct measurements from the supplied raw images for the case workflow show 2.7% image-confirmed bubble coverage and 0.9% saturated pixels. Replicate segmentation shows 5.2% disagreement.\"},{\"speaker\":\"calibration custodian\",\"text\":\"At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units.\"},{\"speaker\":\"calibration custodian\",\"text\":\"The routing-time recalculation entry in calibration log M-47 records a comparison response of 776.0 units for the case workflow.\"},{\"speaker\":\"assay scientist\",\"text\":\"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units."}, {"path": ["3", "text"], "text": "The routing-time recalculation entry in calibration log M-47 records a comparison response of 776.0 units for the case workflow."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units.", "negative_left": "At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units.", "negative_right": "The routing-time recalculation entry in calibration log M-47 records a comparison response of 792.0 units for the case workflow.", "right": "The routing-time recalculation entry in calibration log M-47 records a comparison response of 776.0 units for the case workflow."}, "verifier_independent_model": false}, "family": "scale-diverse-260-003", "id": "scale-diverse-260-003-base", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "At routing time, direct measurements from the supplied raw images for the case workflow show 2.7% image-confirmed bubble coverage and 0.9% saturated pixels. Replicate segmentation shows 5.2% disagreement."}, {"speaker": "calibration custodian", "text": "At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units."}, {"speaker": "calibration custodian", "text": "The routing-time recalculation entry in calibration log M-47 records a comparison response of 776.0 units for the case workflow."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_acquisition"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing routing rubric, precedence rules, thresholds, request scope, case workflow, and routing-time binding, while the unchanged questions object preserves the original decision criteria. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the comparison response from 776.0 to 792.0 units, creating no duplicate or contradictory measurement, and neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"image analyst\",\"text\":\"At routing time, direct measurements from the supplied raw images for the case workflow show 2.7% image-confirmed bubble coverage and 0.9% saturated pixels. Replicate segmentation shows 5.2% disagreement.\"},{\"speaker\":\"calibration custodian\",\"text\":\"At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units.\"},{\"speaker\":\"calibration custodian\",\"text\":\"The routing-time recalculation entry in calibration log M-47 records a comparison response of 776.0 units for the case workflow.\"},{\"speaker\":\"assay scientist\",\"text\":\"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units."}, {"path": ["3", "text"], "text": "The routing-time recalculation entry in calibration log M-47 records a comparison response of 776.0 units for the case workflow."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units.", "negative_left": "At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units.", "negative_right": "The routing-time recalculation entry in calibration log M-47 records a comparison response of 792.0 units for the case workflow.", "right": "The routing-time recalculation entry in calibration log M-47 records a comparison response of 776.0 units for the case workflow."}, "verifier_independent_model": false}, "family": "scale-diverse-260-003", "id": "scale-diverse-260-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "At routing time, direct measurements from the supplied raw images for the case workflow show 2.7% image-confirmed bubble coverage and 0.9% saturated pixels. Replicate segmentation shows 5.2% disagreement."}, {"speaker": "calibration custodian", "text": "At routing time, calibration log M-47 associated with the case workflow identifies 800.0 units as the baseline response and records that its calibration-drift field equals 100 times the absolute difference between the comparison response and 800.0 units, divided by 800.0 units."}, {"speaker": "calibration custodian", "text": "The routing-time recalculation entry in calibration log M-47 records a comparison response of 792.0 units for the case workflow."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing rubric, thresholds, precedence rules, scope, and case-workflow routing-time binding; the unchanged questions preserve the choice criteria and instructions. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently replaces the sole recalculated drift measurement from 2.6% to 1.8% without introducing a duplicate or contradiction, while all other observations remain unchanged. Neither context includes a gold answer, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"field reviewer\",\"text\":\"At routing time, direct measurements from the supplied raw images for the case workflow showed 2.7% image-confirmed bubble coverage and 0.6% saturated pixels. Replicate segmentation for the case workflow showed 6.4% disagreement at that time.\"},{\"speaker\":\"records custodian\",\"text\":\"At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 2.6%.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84."}, {"path": ["3", "text"], "text": "A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 2.6%."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84.", "negative_left": "At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84.", "negative_right": "A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 1.8%.", "right": "A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 2.6%."}, "verifier_independent_model": false}, "family": "scale-diverse-260-004", "id": "scale-diverse-260-004-base", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "field reviewer", "text": "At routing time, direct measurements from the supplied raw images for the case workflow showed 2.7% image-confirmed bubble coverage and 0.6% saturated pixels. Replicate segmentation for the case workflow showed 6.4% disagreement at that time."}, {"speaker": "records custodian", "text": "At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84."}, {"speaker": "calibration reviewer", "text": "A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 2.6%."}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_acquisition"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing rubric, thresholds, precedence rules, scope, and case-workflow routing-time binding; the unchanged questions preserve the choice criteria and instructions. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently replaces the sole recalculated drift measurement from 2.6% to 1.8% without introducing a duplicate or contradiction, while all other observations remain unchanged. Neither context includes a gold answer, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"field reviewer\",\"text\":\"At routing time, direct measurements from the supplied raw images for the case workflow showed 2.7% image-confirmed bubble coverage and 0.6% saturated pixels. Replicate segmentation for the case workflow showed 6.4% disagreement at that time.\"},{\"speaker\":\"records custodian\",\"text\":\"At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 2.6%.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84."}, {"path": ["3", "text"], "text": "A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 2.6%."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84.", "negative_left": "At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84.", "negative_right": "A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 1.8%.", "right": "A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 2.6%."}, "verifier_independent_model": false}, "family": "scale-diverse-260-004", "id": "scale-diverse-260-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "field reviewer", "text": "At routing time, direct measurements from the supplied raw images for the case workflow showed 2.7% image-confirmed bubble coverage and 0.6% saturated pixels. Replicate segmentation for the case workflow showed 6.4% disagreement at that time."}, {"speaker": "records custodian", "text": "At its routing time of 2026-04-11 14:20 UTC, the case workflow was associated with calibration log CL-84."}, {"speaker": "calibration reviewer", "text": "A recalculation performed at 2026-04-11 14:20 UTC solely from entries in calibration log CL-84 yielded a calibration drift of 1.8%."}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same QC requirements and handling of unverified criteria while changing only case observations. They preserve the acceptance question concerning Lena’s microscopy workflow and Priya’s decision. The two evidence spans are complete factual sentences. The counterfactual’s subtraction of 0.004 µm per pixel from the 0.522 preliminary output is coherent and does not contradict another recorded measurement. Neither context includes a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a1 is an allowed universal claim over the explicit three-replicate set. The focus a4 is a factual pixel-size relationship rather than a policy proposition. Base and counter assignments are realizable with only a4 changing: the same calibration image can have an in-tolerance measurement in the base case and an out-of-tolerance measurement in the countercase. The policy evidence preserves the substantive rubric, threshold, and unverified-criterion rule originating in the original state. The synthetic identifier C-17 does not introduce a conflicting policy assumption.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that segmentation and artifact requirements pass and that the calibration image exists, while refutation of a4 establishes that its recorded pixel size is outside the allowed 2% tolerance. Thus a required criterion fails, which is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all three rubric requirements: every replicate meets the 0.90 segmentation threshold, artifact burden is low or moderate, and a submitted calibration image has a recorded pixel size within 2% of 0.50 µm. This is sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At Priya's acceptance decision, every one of Lena's three submitted fluorescent tissue-slide replicates has a segmentation Dice score of at least 0.90."}, {"id": "a2", "statement": "At Priya's acceptance decision, the artifact-burden rating for Lena's submitted three-replicate microscopy workflow belongs to the set {low, moderate}."}, {"id": "a3", "statement": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17."}, {"id": "a4", "statement": "At Priya's acceptance decision, the pixel-size measurement recorded for the submitted calibration image C-17 differs by no more than 2% from the certified 0.50 µm value."}], "base_state_json": "\"On 14 August 2026, Omar completed analysis of Lena’s three submitted fluorescent tissue-slide replicates. Their segmentation Dice scores were 0.93, 0.92, and 0.94, each satisfying the required 0.90 minimum. He rated the submitted workflow’s artifact burden low. At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel. The submission manifest also listed calibration image C-17. At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.016 µm per pixel. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel."}, {"path": [], "text": "At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.016 µm per pixel."}], "policy_evidence": [{"path": [], "text": "the required 0.90 minimum"}, {"path": [], "text": "The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value."}, {"path": [], "text": "Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}], "rules": [{"justification": "All other documented criteria pass, but the submitted calibration image's recorded pixel size is explicitly outside the permitted 2% tolerance around the rubric's certified 0.50 µm value, so a required criterion fails.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}, {"justification": "Every submitted replicate meets the required 0.90 segmentation threshold, artifact burden has an allowed rating, and the submitted calibration image records a pixel size within 2% of the expressly required certified 0.50 µm value.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel.", "negative_left": "At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel.", "negative_right": "At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.004 µm per pixel.", "right": "At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.016 µm per pixel."}, "verifier_independent_model": false}, "family": "scale-diverse-261-001", "id": "scale-diverse-261-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required criterion fails or cannot be verified from the supplied information.", "true": "Yes — all required acceptance criteria are documented and pass."}, "instructions": "Is this microscopy workflow ready for acceptance under the supplied QC rubric? Answer yes or no.", "type": "noul"}}, "state": "On 14 August 2026, Omar completed analysis of Lena’s three submitted fluorescent tissue-slide replicates. Their segmentation Dice scores were 0.93, 0.92, and 0.94, each satisfying the required 0.90 minimum. He rated the submitted workflow’s artifact burden low. At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel. The submission manifest also listed calibration image C-17. At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.016 µm per pixel. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}, "method": "c2d", "provenance": {"source_id": "diverse-261", "source_is_synthetic": true, "source_sha256": "5342cad74fb77dfe1b1e4d029535fa8b1ca8f1d3c3996871dd4943e57ed0d56b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same QC requirements and handling of unverified criteria while changing only case observations. They preserve the acceptance question concerning Lena’s microscopy workflow and Priya’s decision. The two evidence spans are complete factual sentences. The counterfactual’s subtraction of 0.004 µm per pixel from the 0.522 preliminary output is coherent and does not contradict another recorded measurement. Neither context includes a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a1 is an allowed universal claim over the explicit three-replicate set. The focus a4 is a factual pixel-size relationship rather than a policy proposition. Base and counter assignments are realizable with only a4 changing: the same calibration image can have an in-tolerance measurement in the base case and an out-of-tolerance measurement in the countercase. The policy evidence preserves the substantive rubric, threshold, and unverified-criterion rule originating in the original state. The synthetic identifier C-17 does not introduce a conflicting policy assumption.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that segmentation and artifact requirements pass and that the calibration image exists, while refutation of a4 establishes that its recorded pixel size is outside the allowed 2% tolerance. Thus a required criterion fails, which is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all three rubric requirements: every replicate meets the 0.90 segmentation threshold, artifact burden is low or moderate, and a submitted calibration image has a recorded pixel size within 2% of 0.50 µm. This is sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At Priya's acceptance decision, every one of Lena's three submitted fluorescent tissue-slide replicates has a segmentation Dice score of at least 0.90."}, {"id": "a2", "statement": "At Priya's acceptance decision, the artifact-burden rating for Lena's submitted three-replicate microscopy workflow belongs to the set {low, moderate}."}, {"id": "a3", "statement": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17."}, {"id": "a4", "statement": "At Priya's acceptance decision, the pixel-size measurement recorded for the submitted calibration image C-17 differs by no more than 2% from the certified 0.50 µm value."}], "base_state_json": "\"On 14 August 2026, Omar completed analysis of Lena’s three submitted fluorescent tissue-slide replicates. Their segmentation Dice scores were 0.93, 0.92, and 0.94, each satisfying the required 0.90 minimum. He rated the submitted workflow’s artifact burden low. At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel. The submission manifest also listed calibration image C-17. At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.016 µm per pixel. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel."}, {"path": [], "text": "At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.016 µm per pixel."}], "policy_evidence": [{"path": [], "text": "the required 0.90 minimum"}, {"path": [], "text": "The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value."}, {"path": [], "text": "Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}], "rules": [{"justification": "All other documented criteria pass, but the submitted calibration image's recorded pixel size is explicitly outside the permitted 2% tolerance around the rubric's certified 0.50 µm value, so a required criterion fails.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}, {"justification": "Every submitted replicate meets the required 0.90 segmentation threshold, artifact burden has an allowed rating, and the submitted calibration image records a pixel size within 2% of the expressly required certified 0.50 µm value.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel.", "negative_left": "At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel.", "negative_right": "At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.004 µm per pixel.", "right": "At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.016 µm per pixel."}, "verifier_independent_model": false}, "family": "scale-diverse-261-001", "id": "scale-diverse-261-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required criterion fails or cannot be verified from the supplied information.", "true": "Yes — all required acceptance criteria are documented and pass."}, "instructions": "Is this microscopy workflow ready for acceptance under the supplied QC rubric? Answer yes or no.", "type": "noul"}}, "state": "On 14 August 2026, Omar completed analysis of Lena’s three submitted fluorescent tissue-slide replicates. Their segmentation Dice scores were 0.93, 0.92, and 0.94, each satisfying the required 0.90 minimum. He rated the submitted workflow’s artifact burden low. At 09:10 on 14 August 2026, the preliminary instrument output for Lena's submitted calibration image C-17 was 0.522 µm per pixel, and C-17's certificate listed 0.500 µm per pixel. The submission manifest also listed calibration image C-17. At Priya's acceptance decision at 09:25 on 14 August 2026, the submission ledger specified C-17's recorded pixel-size measurement as the preliminary instrument output for C-17 minus 0.004 µm per pixel. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}, "method": "c2d", "provenance": {"source_id": "diverse-261", "source_is_synthetic": true, "source_sha256": "5342cad74fb77dfe1b1e4d029535fa8b1ca8f1d3c3996871dd4943e57ed0d56b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original QC requirements and the rule for unverified criteria, while the unchanged questions object preserves the decision instructions and scoring criteria. The workflow, people, calibration asset, record, decision time, and certified-value bindings remain fixed; only the recorded pixel-size measurement changes from 0.507 µm to 0.516 µm. The two evidence spans are complete factual sentences. The changed measurement does not conflict with another measurement in the counterfactual context, and neither context includes an answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a1 is an allowed universal claim over the explicit three-replicate set. The focus a4 is a factual pixel-size relationship rather than a policy proposition. Base and counter assignments are realizable with only a4 changing: the same calibration image can have an in-tolerance measurement in the base case and an out-of-tolerance measurement in the countercase. The policy evidence preserves the substantive rubric, threshold, and unverified-criterion rule originating in the original state. The synthetic identifier C-17 does not introduce a conflicting policy assumption.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that segmentation and artifact requirements pass and that the calibration image exists, while refutation of a4 establishes that its recorded pixel size is outside the allowed 2% tolerance. Thus a required criterion fails, which is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all three rubric requirements: every replicate meets the 0.90 segmentation threshold, artifact burden is low or moderate, and a submitted calibration image has a recorded pixel size within 2% of 0.50 µm. This is sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At Priya's acceptance decision, every one of Lena's three submitted fluorescent tissue-slide replicates has a segmentation Dice score of at least 0.90."}, {"id": "a2", "statement": "At Priya's acceptance decision, the artifact-burden rating for Lena's submitted three-replicate microscopy workflow belongs to the set {low, moderate}."}, {"id": "a3", "statement": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17."}, {"id": "a4", "statement": "At Priya's acceptance decision, the pixel-size measurement recorded for the submitted calibration image C-17 differs by no more than 2% from the certified 0.50 µm value."}], "base_state_json": "\"Operational handoff note for Priya’s acceptance review: Omar’s locked analysis report lists Lena’s three submitted fluorescent tissue-slide replicates at Dice scores 0.93, 0.92, and 0.94; each meets the required 0.90 minimum. The report rates artifact burden for that submitted workflow low. At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88. At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.507 µm for asset M-42 against its certified value of 0.50 µm. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88."}, {"path": [], "text": "At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.507 µm for asset M-42 against its certified value of 0.50 µm."}], "policy_evidence": [{"path": [], "text": "the required 0.90 minimum"}, {"path": [], "text": "The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value."}, {"path": [], "text": "Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}], "rules": [{"justification": "All other documented criteria pass, but the submitted calibration image's recorded pixel size is explicitly outside the permitted 2% tolerance around the rubric's certified 0.50 µm value, so a required criterion fails.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}, {"justification": "Every submitted replicate meets the required 0.90 segmentation threshold, artifact burden has an allowed rating, and the submitted calibration image records a pixel size within 2% of the expressly required certified 0.50 µm value.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88.", "negative_left": "At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88.", "negative_right": "At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.516 µm for asset M-42 against its certified value of 0.50 µm.", "right": "At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.507 µm for asset M-42 against its certified value of 0.50 µm."}, "verifier_independent_model": false}, "family": "scale-diverse-261-003", "id": "scale-diverse-261-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required criterion fails or cannot be verified from the supplied information.", "true": "Yes — all required acceptance criteria are documented and pass."}, "instructions": "Is this microscopy workflow ready for acceptance under the supplied QC rubric? Answer yes or no.", "type": "noul"}}, "state": "Operational handoff note for Priya’s acceptance review: Omar’s locked analysis report lists Lena’s three submitted fluorescent tissue-slide replicates at Dice scores 0.93, 0.92, and 0.94; each meets the required 0.90 minimum. The report rates artifact burden for that submitted workflow low. At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88. At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.507 µm for asset M-42 against its certified value of 0.50 µm. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}, "method": "c2d", "provenance": {"source_id": "diverse-261", "source_is_synthetic": true, "source_sha256": "5342cad74fb77dfe1b1e4d029535fa8b1ca8f1d3c3996871dd4943e57ed0d56b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original QC requirements and the rule for unverified criteria, while the unchanged questions object preserves the decision instructions and scoring criteria. The workflow, people, calibration asset, record, decision time, and certified-value bindings remain fixed; only the recorded pixel-size measurement changes from 0.507 µm to 0.516 µm. The two evidence spans are complete factual sentences. The changed measurement does not conflict with another measurement in the counterfactual context, and neither context includes an answer, answer code, proposition identifier, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a1 is an allowed universal claim over the explicit three-replicate set. The focus a4 is a factual pixel-size relationship rather than a policy proposition. Base and counter assignments are realizable with only a4 changing: the same calibration image can have an in-tolerance measurement in the base case and an out-of-tolerance measurement in the countercase. The policy evidence preserves the substantive rubric, threshold, and unverified-criterion rule originating in the original state. The synthetic identifier C-17 does not introduce a conflicting policy assumption.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that segmentation and artifact requirements pass and that the calibration image exists, while refutation of a4 establishes that its recorded pixel size is outside the allowed 2% tolerance. Thus a required criterion fails, which is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all three rubric requirements: every replicate meets the 0.90 segmentation threshold, artifact burden is low or moderate, and a submitted calibration image has a recorded pixel size within 2% of 0.50 µm. This is sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At Priya's acceptance decision, every one of Lena's three submitted fluorescent tissue-slide replicates has a segmentation Dice score of at least 0.90."}, {"id": "a2", "statement": "At Priya's acceptance decision, the artifact-burden rating for Lena's submitted three-replicate microscopy workflow belongs to the set {low, moderate}."}, {"id": "a3", "statement": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17."}, {"id": "a4", "statement": "At Priya's acceptance decision, the pixel-size measurement recorded for the submitted calibration image C-17 differs by no more than 2% from the certified 0.50 µm value."}], "base_state_json": "\"Operational handoff note for Priya’s acceptance review: Omar’s locked analysis report lists Lena’s three submitted fluorescent tissue-slide replicates at Dice scores 0.93, 0.92, and 0.94; each meets the required 0.90 minimum. The report rates artifact burden for that submitted workflow low. At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88. At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.507 µm for asset M-42 against its certified value of 0.50 µm. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88."}, {"path": [], "text": "At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.507 µm for asset M-42 against its certified value of 0.50 µm."}], "policy_evidence": [{"path": [], "text": "the required 0.90 minimum"}, {"path": [], "text": "The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value."}, {"path": [], "text": "Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}], "rules": [{"justification": "All other documented criteria pass, but the submitted calibration image's recorded pixel size is explicitly outside the permitted 2% tolerance around the rubric's certified 0.50 µm value, so a required criterion fails.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}, {"justification": "Every submitted replicate meets the required 0.90 segmentation threshold, artifact burden has an allowed rating, and the submitted calibration image records a pixel size within 2% of the expressly required certified 0.50 µm value.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88.", "negative_left": "At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88.", "negative_right": "At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.516 µm for asset M-42 against its certified value of 0.50 µm.", "right": "At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.507 µm for asset M-42 against its certified value of 0.50 µm."}, "verifier_independent_model": false}, "family": "scale-diverse-261-003", "id": "scale-diverse-261-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required criterion fails or cannot be verified from the supplied information.", "true": "Yes — all required acceptance criteria are documented and pass."}, "instructions": "Is this microscopy workflow ready for acceptance under the supplied QC rubric? Answer yes or no.", "type": "noul"}}, "state": "Operational handoff note for Priya’s acceptance review: Omar’s locked analysis report lists Lena’s three submitted fluorescent tissue-slide replicates at Dice scores 0.93, 0.92, and 0.94; each meets the required 0.90 minimum. The report rates artifact burden for that submitted workflow low. At Priya's acceptance decision, Lena's submitted calibration image C-17 is logged as asset M-42 in sealed operational handoff record H-88. At Priya's acceptance decision, operational handoff record H-88 records a pixel-size measurement of 0.516 µm for asset M-42 against its certified value of 0.50 µm. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}, "method": "c2d", "provenance": {"source_id": "diverse-261", "source_is_synthetic": true, "source_sha256": "5342cad74fb77dfe1b1e4d029535fa8b1ca8f1d3c3996871dd4943e57ed0d56b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy, exception, request, and case/workflow scope. The two focus-evidence spans are complete factual sentences. In the base context, IMG-C is detector run 47 and therefore belongs to the exact raw-image manifest; in the counterfactual, IMG-C is explicitly distinct from every manifested file, so its 2.7% calibration deviation does not contradict the statement that the manifested raw images pass. No context contains an answer label, output instruction, rule table, proposition ID, or explicit gold-answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A11": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including permissible universally quantified relationships over the explicit workflow set. A1 is a factual scope/membership proposition rather than a policy conclusion. The base and counter assignments are jointly realizable: IMG-C may retain the stated calibration deviation while changing only whether it is an in-scope raw image. The policy evidence preserves the routing rules, exception, and capture thresholds originating in the original state; extra state citations do not create a completeness defect.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "IMG-C is an in-scope raw image with a calibration deviation above the allowed 2%, establishing an acquisition failure. The apparent-blur exception cannot apply because not all raw images meet capture criteria. The conjunction also excludes the stated preparation and analysis alternatives, so routing away from the image analyst is entailed.", "rule_index": 0, "sound": true}, {"reason": "Because IMG-C is not a raw image in the workflow, every in-scope raw image is covered by A3–A5 and meets all stated capture thresholds. Apparent blur appearing only in segmentation overlays therefore satisfies the explicit exception and entails routing to analysis. The stated preparation grounds are excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Image file IMG-C is a raw image in the flagged three-replicate fluorescence workflow."}, {"id": "A2", "statement": "Image file IMG-C has a 0.50-µm calibration-check deviation greater than 2%."}, {"id": "A3", "statement": "Every raw image other than IMG-C in the flagged three-replicate fluorescence workflow has focus of at least 0.80."}, {"id": "A4", "statement": "Every raw image other than IMG-C in the flagged three-replicate fluorescence workflow has exposure saturation of at most 1%."}, {"id": "A5", "statement": "Every raw image other than IMG-C in the flagged three-replicate fluorescence workflow has a 0.50-µm calibration-check deviation of at most 2%."}, {"id": "A6", "statement": "The flagged cell-edge halo in the three-replicate fluorescence workflow is apparent blur."}, {"id": "A7", "statement": "The flagged cell-edge halo appears only in segmentation overlays of the three-replicate fluorescence workflow."}, {"id": "A8", "statement": "No slide in the flagged three-replicate fluorescence workflow has a fold."}, {"id": "A9", "statement": "No slide in the flagged three-replicate fluorescence workflow has stain debris."}, {"id": "A10", "statement": "No segmentation-mask defect exists in the flagged three-replicate fluorescence workflow."}, {"id": "A11", "statement": "No scoring defect exists in the flagged three-replicate fluorescence workflow."}], "base_state_json": "{\"context\":\"Routing policy assigns slide folds or stain debris to preparation, raw focus/exposure or calibration failures to acquisition, and mask or scoring defects to analysis. Exception: apparent blur is routed to analysis when raw images meet capture criteria and the defect appears only in segmentation overlays.\",\"evidence\":[\"At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47.\",\"Image file IMG-C and the file produced by detector run 47 are the same file.\",\"At 09:18 UTC, IMG-C’s 0.50-µm calibration-check deviation was measured at 2.7%.\",\"For every raw image other than IMG-C, the checks recorded focus ≥0.80, exposure saturation ≤1%, and the 0.50-µm calibration check within ±2%.\",\"At 09:25 UTC, the technician classified the flagged cell-edge halo as apparent blur and observed it only in segmentation overlays.\",\"Slide inspection found no fold and no stain debris.\",\"Subsequent review found no segmentation-mask defect and no scoring defect.\",\"The report also cites the assay scientist’s acceptance rule of all three within 10 points.\"],\"request\":\"Should the case be routed to the image analyst?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47."}, {"path": ["evidence", "1"], "text": "Image file IMG-C and the file produced by detector run 47 are the same file."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy assigns slide folds or stain debris to preparation, raw focus/exposure or calibration failures to acquisition, and mask or scoring defects to analysis."}, {"path": ["context"], "text": "Exception: apparent blur is routed to analysis when raw images meet capture criteria and the defect appears only in segmentation overlays."}, {"path": ["evidence", "1"], "text": "focus ≥0.80, exposure saturation ≤1%, and the 0.50-µm calibration check within ±2%."}, {"path": ["evidence", "2"], "text": "the assay scientist’s acceptance rule of all three within 10 points."}, {"path": ["request"], "text": "Should the case be routed to the image analyst?"}], "rules": [{"justification": "IMG-C is an in-scope raw image and exceeds the permitted calibration deviation, so the workflow has an acquisition failure and cannot satisfy the apparent-blur exception's requirement that its raw images meet capture criteria. The remaining conditions exclude slide-fold, stain-debris, mask-defect, and scoring-defect grounds for another route.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}]}, {"justification": "IMG-C is not an in-scope raw image, so its calibration deviation does not affect this workflow. Every in-scope raw image consequently meets each stated capture threshold, and the apparent blur appears only in segmentation overlays. The explicit exception therefore routes the workflow to analysis, while the other conditions exclude competing preparation defects.", "target": "true", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47.", "negative_left": "At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47.", "negative_right": "Image file IMG-C is distinct from IMG-A, IMG-B, and the file produced by detector run 47.", "right": "Image file IMG-C and the file produced by detector run 47 are the same file."}, "verifier_independent_model": false}, "family": "scale-diverse-262-001", "id": "scale-diverse-262-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route it to a preparation or acquisition specialist instead.", "true": "Yes — route the workflow to the image analyst for segmentation and scoring review."}, "instructions": "Answer yes or no using the routing scope and its stated exception.", "type": "noul"}}, "state": {"context": "Routing policy assigns slide folds or stain debris to preparation, raw focus/exposure or calibration failures to acquisition, and mask or scoring defects to analysis. Exception: apparent blur is routed to analysis when raw images meet capture criteria and the defect appears only in segmentation overlays.", "evidence": ["At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47.", "Image file IMG-C and the file produced by detector run 47 are the same file.", "At 09:18 UTC, IMG-C’s 0.50-µm calibration-check deviation was measured at 2.7%.", "For every raw image other than IMG-C, the checks recorded focus ≥0.80, exposure saturation ≤1%, and the 0.50-µm calibration check within ±2%.", "At 09:25 UTC, the technician classified the flagged cell-edge halo as apparent blur and observed it only in segmentation overlays.", "Slide inspection found no fold and no stain debris.", "Subsequent review found no segmentation-mask defect and no scoring defect.", "The report also cites the assay scientist’s acceptance rule of all three within 10 points."], "request": "Should the case be routed to the image analyst?"}}, "method": "c2d", "provenance": {"source_id": "diverse-262", "source_is_synthetic": true, "source_sha256": "8fca1c2a153a4585b7f7f1ff393c32ed67ca088ef30f96bfd30ab6bd8879aa94", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy, exception, request, and case/workflow scope. The two focus-evidence spans are complete factual sentences. In the base context, IMG-C is detector run 47 and therefore belongs to the exact raw-image manifest; in the counterfactual, IMG-C is explicitly distinct from every manifested file, so its 2.7% calibration deviation does not contradict the statement that the manifested raw images pass. No context contains an answer label, output instruction, rule table, proposition ID, or explicit gold-answer rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A11": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A11": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including permissible universally quantified relationships over the explicit workflow set. A1 is a factual scope/membership proposition rather than a policy conclusion. The base and counter assignments are jointly realizable: IMG-C may retain the stated calibration deviation while changing only whether it is an in-scope raw image. The policy evidence preserves the routing rules, exception, and capture thresholds originating in the original state; extra state citations do not create a completeness defect.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "IMG-C is an in-scope raw image with a calibration deviation above the allowed 2%, establishing an acquisition failure. The apparent-blur exception cannot apply because not all raw images meet capture criteria. The conjunction also excludes the stated preparation and analysis alternatives, so routing away from the image analyst is entailed.", "rule_index": 0, "sound": true}, {"reason": "Because IMG-C is not a raw image in the workflow, every in-scope raw image is covered by A3–A5 and meets all stated capture thresholds. Apparent blur appearing only in segmentation overlays therefore satisfies the explicit exception and entails routing to analysis. The stated preparation grounds are excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Image file IMG-C is a raw image in the flagged three-replicate fluorescence workflow."}, {"id": "A2", "statement": "Image file IMG-C has a 0.50-µm calibration-check deviation greater than 2%."}, {"id": "A3", "statement": "Every raw image other than IMG-C in the flagged three-replicate fluorescence workflow has focus of at least 0.80."}, {"id": "A4", "statement": "Every raw image other than IMG-C in the flagged three-replicate fluorescence workflow has exposure saturation of at most 1%."}, {"id": "A5", "statement": "Every raw image other than IMG-C in the flagged three-replicate fluorescence workflow has a 0.50-µm calibration-check deviation of at most 2%."}, {"id": "A6", "statement": "The flagged cell-edge halo in the three-replicate fluorescence workflow is apparent blur."}, {"id": "A7", "statement": "The flagged cell-edge halo appears only in segmentation overlays of the three-replicate fluorescence workflow."}, {"id": "A8", "statement": "No slide in the flagged three-replicate fluorescence workflow has a fold."}, {"id": "A9", "statement": "No slide in the flagged three-replicate fluorescence workflow has stain debris."}, {"id": "A10", "statement": "No segmentation-mask defect exists in the flagged three-replicate fluorescence workflow."}, {"id": "A11", "statement": "No scoring defect exists in the flagged three-replicate fluorescence workflow."}], "base_state_json": "{\"context\":\"Routing policy assigns slide folds or stain debris to preparation, raw focus/exposure or calibration failures to acquisition, and mask or scoring defects to analysis. Exception: apparent blur is routed to analysis when raw images meet capture criteria and the defect appears only in segmentation overlays.\",\"evidence\":[\"At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47.\",\"Image file IMG-C and the file produced by detector run 47 are the same file.\",\"At 09:18 UTC, IMG-C’s 0.50-µm calibration-check deviation was measured at 2.7%.\",\"For every raw image other than IMG-C, the checks recorded focus ≥0.80, exposure saturation ≤1%, and the 0.50-µm calibration check within ±2%.\",\"At 09:25 UTC, the technician classified the flagged cell-edge halo as apparent blur and observed it only in segmentation overlays.\",\"Slide inspection found no fold and no stain debris.\",\"Subsequent review found no segmentation-mask defect and no scoring defect.\",\"The report also cites the assay scientist’s acceptance rule of all three within 10 points.\"],\"request\":\"Should the case be routed to the image analyst?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47."}, {"path": ["evidence", "1"], "text": "Image file IMG-C and the file produced by detector run 47 are the same file."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy assigns slide folds or stain debris to preparation, raw focus/exposure or calibration failures to acquisition, and mask or scoring defects to analysis."}, {"path": ["context"], "text": "Exception: apparent blur is routed to analysis when raw images meet capture criteria and the defect appears only in segmentation overlays."}, {"path": ["evidence", "1"], "text": "focus ≥0.80, exposure saturation ≤1%, and the 0.50-µm calibration check within ±2%."}, {"path": ["evidence", "2"], "text": "the assay scientist’s acceptance rule of all three within 10 points."}, {"path": ["request"], "text": "Should the case be routed to the image analyst?"}], "rules": [{"justification": "IMG-C is an in-scope raw image and exceeds the permitted calibration deviation, so the workflow has an acquisition failure and cannot satisfy the apparent-blur exception's requirement that its raw images meet capture criteria. The remaining conditions exclude slide-fold, stain-debris, mask-defect, and scoring-defect grounds for another route.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}]}, {"justification": "IMG-C is not an in-scope raw image, so its calibration deviation does not affect this workflow. Every in-scope raw image consequently meets each stated capture threshold, and the apparent blur appears only in segmentation overlays. The explicit exception therefore routes the workflow to analysis, while the other conditions exclude competing preparation defects.", "target": "true", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47.", "negative_left": "At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47.", "negative_right": "Image file IMG-C is distinct from IMG-A, IMG-B, and the file produced by detector run 47.", "right": "Image file IMG-C and the file produced by detector run 47 are the same file."}, "verifier_independent_model": false}, "family": "scale-diverse-262-001", "id": "scale-diverse-262-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route it to a preparation or acquisition specialist instead.", "true": "Yes — route the workflow to the image analyst for segmentation and scoring review."}, "instructions": "Answer yes or no using the routing scope and its stated exception.", "type": "noul"}}, "state": {"context": "Routing policy assigns slide folds or stain debris to preparation, raw focus/exposure or calibration failures to acquisition, and mask or scoring defects to analysis. Exception: apparent blur is routed to analysis when raw images meet capture criteria and the defect appears only in segmentation overlays.", "evidence": ["At 09:12 UTC on 14 August 2026, the complete raw-image manifest for the flagged three-replicate fluorescence workflow contained exactly IMG-A, IMG-B, and the file produced by detector run 47.", "Image file IMG-C is distinct from IMG-A, IMG-B, and the file produced by detector run 47.", "At 09:18 UTC, IMG-C’s 0.50-µm calibration-check deviation was measured at 2.7%.", "For every raw image other than IMG-C, the checks recorded focus ≥0.80, exposure saturation ≤1%, and the 0.50-µm calibration check within ±2%.", "At 09:25 UTC, the technician classified the flagged cell-edge halo as apparent blur and observed it only in segmentation overlays.", "Slide inspection found no fold and no stain debris.", "Subsequent review found no segmentation-mask defect and no scoring defect.", "The report also cites the assay scientist’s acceptance rule of all three within 10 points."], "request": "Should the case be routed to the image analyst?"}}, "method": "c2d", "provenance": {"source_id": "diverse-262", "source_is_synthetic": true, "source_sha256": "8fca1c2a153a4585b7f7f1ff393c32ed67ca088ef30f96bfd30ab6bd8879aa94", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same workflow, scored replicates, log-entry relationship, and governing reliability policy without adding exceptions or defaults. The two evidence spans are complete factual sentences; changing entry 47 from 0.95 to 0.85 coherently changes replicate C’s focus score without conflicting with the statement that every other required check passed, and neither context exposes a score label or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"records auditor\",\"text\":\"At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score.\"},{\"speaker\":\"log reviewer\",\"text\":\"Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.95.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score."}, {"path": ["1", "text"], "text": "Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.95."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score.", "negative_left": "At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score.", "negative_right": "Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.85.", "right": "Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.95."}, "verifier_independent_model": false}, "family": "scale-diverse-263-001", "id": "scale-diverse-263-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "records auditor", "text": "At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score."}, {"speaker": "log reviewer", "text": "Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.95."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same workflow, scored replicates, log-entry relationship, and governing reliability policy without adding exceptions or defaults. The two evidence spans are complete factual sentences; changing entry 47 from 0.95 to 0.85 coherently changes replicate C’s focus score without conflicting with the statement that every other required check passed, and neither context exposes a score label or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"records auditor\",\"text\":\"At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score.\"},{\"speaker\":\"log reviewer\",\"text\":\"Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.95.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score."}, {"path": ["1", "text"], "text": "Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.95."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score.", "negative_left": "At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score.", "negative_right": "Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.85.", "right": "Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.95."}, "verifier_independent_model": false}, "family": "scale-diverse-263-001", "id": "scale-diverse-263-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "records auditor", "text": "At 2026-08-19T11:03:40Z, the signed chronological log for this workflow run reported focus scores of 0.96 and 0.93 for scored fluorescence replicates A and B respectively, confirmed that every other required capture check and every required equipment-calibration and tissue-ROI segmentation check had passed for scored fluorescence replicates A, B, and C, recorded zero counted artifacts inside each replicate's tissue ROI and cell-count CVs of 5.7%, 6.4%, and 7.1% respectively, and identified entry 47, created during capture of scored fluorescence replicate C at 2026-08-19T10:42:18Z, as that replicate's recorded focus score."}, {"speaker": "log reviewer", "text": "Entry 47 in the signed chronological log for this workflow run contains the numeric value 0.85."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Neither constructed context alters or supplements the governing reliability criteria, thresholds, scope, or interpretation instructions. The same workflow run, replicate C, ledger entry FS-47, and focus-score evidentiary path are preserved; the two evidence spans are complete factual sentences. The counterfactual changes only FS-47's value from 0.93 to 0.87, while the separate focus assertion is limited to replicates A and B, so no duplicate measurement or unchanged assertion contradicts it. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"capture-record reviewer\",\"text\":\"In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry.\"},{\"speaker\":\"ledger reviewer\",\"text\":\"In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.93.\"},{\"speaker\":\"quality specialist\",\"text\":\"The capture records separately confirm that scored fluorescence replicates A and B each recorded a focus score of at least 0.90. Every required capture check other than the focus-score check passed for each of replicates A, B, and C.\"},{\"speaker\":\"equipment specialist\",\"text\":\"Every required calibration check passed for the equipment used to capture the three scored fluorescence replicates.\"},{\"speaker\":\"image-analysis reviewer\",\"text\":\"For every tissue ROI in replicates A, B, and C, all required segmentation checks passed. Artifact reconciliation counted zero artifacts inside each replicate's tissue ROI.\"},{\"speaker\":\"assay reviewer\",\"text\":\"The recorded cell-count CVs were 7.6% for replicate A, 7.9% for replicate B, and 8.0% for replicate C.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry."}, {"path": ["1", "text"], "text": "In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.93."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry.", "negative_left": "In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry.", "negative_right": "In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.87.", "right": "In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.93."}, "verifier_independent_model": false}, "family": "scale-diverse-263-002", "id": "scale-diverse-263-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "capture-record reviewer", "text": "In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry."}, {"speaker": "ledger reviewer", "text": "In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.93."}, {"speaker": "quality specialist", "text": "The capture records separately confirm that scored fluorescence replicates A and B each recorded a focus score of at least 0.90. Every required capture check other than the focus-score check passed for each of replicates A, B, and C."}, {"speaker": "equipment specialist", "text": "Every required calibration check passed for the equipment used to capture the three scored fluorescence replicates."}, {"speaker": "image-analysis reviewer", "text": "For every tissue ROI in replicates A, B, and C, all required segmentation checks passed. Artifact reconciliation counted zero artifacts inside each replicate's tissue ROI."}, {"speaker": "assay reviewer", "text": "The recorded cell-count CVs were 7.6% for replicate A, 7.9% for replicate B, and 8.0% for replicate C."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Neither constructed context alters or supplements the governing reliability criteria, thresholds, scope, or interpretation instructions. The same workflow run, replicate C, ledger entry FS-47, and focus-score evidentiary path are preserved; the two evidence spans are complete factual sentences. The counterfactual changes only FS-47's value from 0.93 to 0.87, while the separate focus assertion is limited to replicates A and B, so no duplicate measurement or unchanged assertion contradicts it. Neither context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"capture-record reviewer\",\"text\":\"In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry.\"},{\"speaker\":\"ledger reviewer\",\"text\":\"In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.93.\"},{\"speaker\":\"quality specialist\",\"text\":\"The capture records separately confirm that scored fluorescence replicates A and B each recorded a focus score of at least 0.90. Every required capture check other than the focus-score check passed for each of replicates A, B, and C.\"},{\"speaker\":\"equipment specialist\",\"text\":\"Every required calibration check passed for the equipment used to capture the three scored fluorescence replicates.\"},{\"speaker\":\"image-analysis reviewer\",\"text\":\"For every tissue ROI in replicates A, B, and C, all required segmentation checks passed. Artifact reconciliation counted zero artifacts inside each replicate's tissue ROI.\"},{\"speaker\":\"assay reviewer\",\"text\":\"The recorded cell-count CVs were 7.6% for replicate A, 7.9% for replicate B, and 8.0% for replicate C.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry."}, {"path": ["1", "text"], "text": "In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.93."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry.", "negative_left": "In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry.", "negative_right": "In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.87.", "right": "In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.93."}, "verifier_independent_model": false}, "family": "scale-diverse-263-002", "id": "scale-diverse-263-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "capture-record reviewer", "text": "In the workflow run completed at 2026-08-14T16:20:00Z, the capture record for scored fluorescence replicate C designates ledger entry FS-47 as that replicate's recorded focus-score entry."}, {"speaker": "ledger reviewer", "text": "In the capture ledger for the workflow run completed at 2026-08-14T16:20:00Z, entry FS-47 records the numeric value 0.87."}, {"speaker": "quality specialist", "text": "The capture records separately confirm that scored fluorescence replicates A and B each recorded a focus score of at least 0.90. Every required capture check other than the focus-score check passed for each of replicates A, B, and C."}, {"speaker": "equipment specialist", "text": "Every required calibration check passed for the equipment used to capture the three scored fluorescence replicates."}, {"speaker": "image-analysis reviewer", "text": "For every tissue ROI in replicates A, B, and C, all required segmentation checks passed. Artifact reconciliation counted zero artifacts inside each replicate's tissue ROI."}, {"speaker": "assay reviewer", "text": "The recorded cell-count CVs were 7.6% for replicate A, 7.9% for replicate B, and 8.0% for replicate C."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original reliability criteria, scope, workflow, replicates, timestamp, and frame identity without adding exceptions or missing-evidence defaults. The evidence consists of exactly two complete factual sentences, and the counterfactual coherently changes only replicate C’s focus score from 0.94 to 0.84 without creating duplicate contradictory measurements; neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"microscopy technician\",\"text\":\"At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731.\"},{\"speaker\":\"capture ledger\",\"text\":\"The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.94.\"},{\"speaker\":\"quality-control handoff\",\"text\":\"Scored fluorescence replicates A and B each recorded a focus score of at least 0.90. For A, B, and C, every required capture check other than the focus-score check passed. Every required calibration check passed for the equipment used to capture all three replicates.\"},{\"speaker\":\"image-analysis handoff\",\"text\":\"Every required segmentation check passed for each tissue ROI in A, B, and C. Each replicate had zero counted artifacts inside its tissue ROI. Cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively. A bright streak observed outside C’s tissue ROI was not included in the artifact count.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731."}, {"path": ["1", "text"], "text": "The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.94."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731.", "negative_left": "At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731.", "negative_right": "The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.84.", "right": "The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.94."}, "verifier_independent_model": false}, "family": "scale-diverse-263-003", "id": "scale-diverse-263-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "microscopy technician", "text": "At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731."}, {"speaker": "capture ledger", "text": "The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.94."}, {"speaker": "quality-control handoff", "text": "Scored fluorescence replicates A and B each recorded a focus score of at least 0.90. For A, B, and C, every required capture check other than the focus-score check passed. Every required calibration check passed for the equipment used to capture all three replicates."}, {"speaker": "image-analysis handoff", "text": "Every required segmentation check passed for each tissue ROI in A, B, and C. Each replicate had zero counted artifacts inside its tissue ROI. Cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively. A bright streak observed outside C’s tissue ROI was not included in the artifact count."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original reliability criteria, scope, workflow, replicates, timestamp, and frame identity without adding exceptions or missing-evidence defaults. The evidence consists of exactly two complete factual sentences, and the counterfactual coherently changes only replicate C’s focus score from 0.94 to 0.84 without creating duplicate contradictory measurements; neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"microscopy technician\",\"text\":\"At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731.\"},{\"speaker\":\"capture ledger\",\"text\":\"The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.94.\"},{\"speaker\":\"quality-control handoff\",\"text\":\"Scored fluorescence replicates A and B each recorded a focus score of at least 0.90. For A, B, and C, every required capture check other than the focus-score check passed. Every required calibration check passed for the equipment used to capture all three replicates.\"},{\"speaker\":\"image-analysis handoff\",\"text\":\"Every required segmentation check passed for each tissue ROI in A, B, and C. Each replicate had zero counted artifacts inside its tissue ROI. Cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively. A bright streak observed outside C’s tissue ROI was not included in the artifact count.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731."}, {"path": ["1", "text"], "text": "The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.94."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731.", "negative_left": "At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731.", "negative_right": "The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.84.", "right": "The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.94."}, "verifier_independent_model": false}, "family": "scale-diverse-263-003", "id": "scale-diverse-263-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "microscopy technician", "text": "At 2026-08-14T09:17:26Z in this workflow run, scored fluorescence replicate C was captured exclusively as image frame QF-731."}, {"speaker": "capture ledger", "text": "The capture ledger for this workflow run records image frame QF-731 at 2026-08-14T09:17:26Z with a focus score of 0.84."}, {"speaker": "quality-control handoff", "text": "Scored fluorescence replicates A and B each recorded a focus score of at least 0.90. For A, B, and C, every required capture check other than the focus-score check passed. Every required calibration check passed for the equipment used to capture all three replicates."}, {"speaker": "image-analysis handoff", "text": "Every required segmentation check passed for each tissue ROI in A, B, and C. Each replicate had zero counted artifacts inside its tissue ROI. Cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively. A bright streak observed outside C’s tissue ROI was not included in the artifact count."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions, while both contexts retain the same workflow, replicate C, capture-record FC-47, and timestamp bindings. The two evidence spans are complete factual sentences. The counterfactual changes only FC-47's focus score from 0.94 to 0.86; the statement that all checks apart from focus passed does not contradict that change. Neither context contains an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"capture technician\",\"text\":\"During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47.\"},{\"speaker\":\"records reviewer\",\"text\":\"The focus-score field in capture record FC-47 contains the value 0.94 for the capture at 2026-08-14T10:32:18Z.\"},{\"speaker\":\"capture technician\",\"text\":\"The capture records give focus scores of 0.93 for scored fluorescence replicate A and 0.91 for scored fluorescence replicate B. Apart from the focus-score check, every required capture check passed for each of A, B, and C.\"},{\"speaker\":\"equipment specialist\",\"text\":\"Every required calibration check passed for the equipment used to capture all three scored fluorescence replicates.\"},{\"speaker\":\"image analyst\",\"text\":\"Every required segmentation check passed for each tissue ROI in A, B, and C. Each replicate had zero counted artifacts inside its tissue ROI; the bright streak associated with C lay outside its ROI and was excluded from the count.\"},{\"speaker\":\"assay scientist\",\"text\":\"Replicates A, B, and C remained scored and had cell-count CVs of 7.6%, 7.9%, and 8.0%, respectively.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47."}, {"path": ["1", "text"], "text": "The focus-score field in capture record FC-47 contains the value 0.94 for the capture at 2026-08-14T10:32:18Z."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47.", "negative_left": "During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47.", "negative_right": "The focus-score field in capture record FC-47 contains the value 0.86 for the capture at 2026-08-14T10:32:18Z.", "right": "The focus-score field in capture record FC-47 contains the value 0.94 for the capture at 2026-08-14T10:32:18Z."}, "verifier_independent_model": false}, "family": "scale-diverse-263-004", "id": "scale-diverse-263-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "capture technician", "text": "During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47."}, {"speaker": "records reviewer", "text": "The focus-score field in capture record FC-47 contains the value 0.94 for the capture at 2026-08-14T10:32:18Z."}, {"speaker": "capture technician", "text": "The capture records give focus scores of 0.93 for scored fluorescence replicate A and 0.91 for scored fluorescence replicate B. Apart from the focus-score check, every required capture check passed for each of A, B, and C."}, {"speaker": "equipment specialist", "text": "Every required calibration check passed for the equipment used to capture all three scored fluorescence replicates."}, {"speaker": "image analyst", "text": "Every required segmentation check passed for each tissue ROI in A, B, and C. Each replicate had zero counted artifacts inside its tissue ROI; the bright streak associated with C lay outside its ROI and was excluded from the count."}, {"speaker": "assay scientist", "text": "Replicates A, B, and C remained scored and had cell-count CVs of 7.6%, 7.9%, and 8.0%, respectively."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions, while both contexts retain the same workflow, replicate C, capture-record FC-47, and timestamp bindings. The two evidence spans are complete factual sentences. The counterfactual changes only FC-47's focus score from 0.94 to 0.86; the statement that all checks apart from focus passed does not contradict that change. Neither context contains an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"capture technician\",\"text\":\"During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47.\"},{\"speaker\":\"records reviewer\",\"text\":\"The focus-score field in capture record FC-47 contains the value 0.94 for the capture at 2026-08-14T10:32:18Z.\"},{\"speaker\":\"capture technician\",\"text\":\"The capture records give focus scores of 0.93 for scored fluorescence replicate A and 0.91 for scored fluorescence replicate B. Apart from the focus-score check, every required capture check passed for each of A, B, and C.\"},{\"speaker\":\"equipment specialist\",\"text\":\"Every required calibration check passed for the equipment used to capture all three scored fluorescence replicates.\"},{\"speaker\":\"image analyst\",\"text\":\"Every required segmentation check passed for each tissue ROI in A, B, and C. Each replicate had zero counted artifacts inside its tissue ROI; the bright streak associated with C lay outside its ROI and was excluded from the count.\"},{\"speaker\":\"assay scientist\",\"text\":\"Replicates A, B, and C remained scored and had cell-count CVs of 7.6%, 7.9%, and 8.0%, respectively.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47."}, {"path": ["1", "text"], "text": "The focus-score field in capture record FC-47 contains the value 0.94 for the capture at 2026-08-14T10:32:18Z."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47.", "negative_left": "During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47.", "negative_right": "The focus-score field in capture record FC-47 contains the value 0.86 for the capture at 2026-08-14T10:32:18Z.", "right": "The focus-score field in capture record FC-47 contains the value 0.94 for the capture at 2026-08-14T10:32:18Z."}, "verifier_independent_model": false}, "family": "scale-diverse-263-004", "id": "scale-diverse-263-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "capture technician", "text": "During this workflow run, scored fluorescence replicate C was captured at 2026-08-14T10:32:18Z under the unique capture-record identifier FC-47."}, {"speaker": "records reviewer", "text": "The focus-score field in capture record FC-47 contains the value 0.86 for the capture at 2026-08-14T10:32:18Z."}, {"speaker": "capture technician", "text": "The capture records give focus scores of 0.93 for scored fluorescence replicate A and 0.91 for scored fluorescence replicate B. Apart from the focus-score check, every required capture check passed for each of A, B, and C."}, {"speaker": "equipment specialist", "text": "Every required calibration check passed for the equipment used to capture all three scored fluorescence replicates."}, {"speaker": "image analyst", "text": "Every required segmentation check passed for each tissue ROI in A, B, and C. Each replicate had zero counted artifacts inside its tissue ROI; the bright streak associated with C lay outside its ROI and was excluded from the count."}, {"speaker": "assay scientist", "text": "Replicates A, B, and C remained scored and had cell-count CVs of 7.6%, 7.9%, and 8.0%, respectively."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration gate, the 10% threshold, and its precedence over other quality indicators while keeping the decision scoped to the microscopy workflow’s quantitative size and density measurements. The two evidence spans are complete factual sentences about CR-47; the counterfactual changes only the absolute error from 0.24 to 0.18 micrometres per pixel, yielding coherent errors of 12% and 9% respectively against the unchanged 2.00 reference, with no contradictory duplicates. Neither context states a reliability level, answer code, proposition ID, classifier instruction, or case-specific label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.24 micrometres per pixel. Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100. Preparation inspection found no major preparation defect that materially biased the workflow’s measurements. Capture QA likewise found no major capture defect that materially biased the measurements. Segmentation QA found no major segmentation defect that materially biased the measurements. Review of the session images classified artifacts as negligible. The workflow’s segmentation met its validation criteria, and replicate measurements were consistent across the reviewed fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.24 micrometres per pixel."}, {"path": [], "text": "Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.24 micrometres per pixel.", "negative_left": "At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.18 micrometres per pixel.", "negative_right": "Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100.", "right": "Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100."}, "verifier_independent_model": false}, "family": "scale-diverse-264-003", "id": "scale-diverse-264-003-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.24 micrometres per pixel. Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100. Preparation inspection found no major preparation defect that materially biased the workflow’s measurements. Capture QA likewise found no major capture defect that materially biased the measurements. Segmentation QA found no major segmentation defect that materially biased the measurements. Review of the session images classified artifacts as negligible. The workflow’s segmentation met its validation criteria, and replicate measurements were consistent across the reviewed fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration gate, the 10% threshold, and its precedence over other quality indicators while keeping the decision scoped to the microscopy workflow’s quantitative size and density measurements. The two evidence spans are complete factual sentences about CR-47; the counterfactual changes only the absolute error from 0.24 to 0.18 micrometres per pixel, yielding coherent errors of 12% and 9% respectively against the unchanged 2.00 reference, with no contradictory duplicates. Neither context states a reliability level, answer code, proposition ID, classifier instruction, or case-specific label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.24 micrometres per pixel. Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100. Preparation inspection found no major preparation defect that materially biased the workflow’s measurements. Capture QA likewise found no major capture defect that materially biased the measurements. Segmentation QA found no major segmentation defect that materially biased the measurements. Review of the session images classified artifacts as negligible. The workflow’s segmentation met its validation criteria, and replicate measurements were consistent across the reviewed fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.24 micrometres per pixel."}, {"path": [], "text": "Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.24 micrometres per pixel.", "negative_left": "At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.18 micrometres per pixel.", "negative_right": "Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100.", "right": "Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100."}, "verifier_independent_model": false}, "family": "scale-diverse-264-003", "id": "scale-diverse-264-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "At the 2026-08-14 operational handoff, mandatory calibration record CR-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded an absolute scale error of 0.18 micrometres per pixel. Calibration record CR-47 specifies a reference scale of 2.00 micrometres per pixel and calculates percentage scale error as the absolute scale error divided by that reference scale, multiplied by 100. Preparation inspection found no major preparation defect that materially biased the workflow’s measurements. Capture QA likewise found no major capture defect that materially biased the measurements. Segmentation QA found no major segmentation defect that materially biased the measurements. Review of the session images classified artifacts as negligible. The workflow’s segmentation met its validation criteria, and replicate measurements were consistent across the reviewed fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete scoring rubric and instructions, while both contexts retain the original mandatory-calibration policy without alteration. The workflow scope and calibration-record/time bindings are consistent; the two evidence spans are complete factual sentences; and the counterfactual coherently changes only the apparent spacing from 11.20 to 10.80 micrometers, producing a noncontradictory change from 12% to 8% under the stated formula. Neither context includes a level, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"Field note — microscopy workflow capture review\\n\\nAt 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers. For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 11.20 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing.\\n\\nReview found no major preparation defect, major capture defect, or major segmentation defect that materially biased the workflow’s measurements. Image artifacts were negligible. Segmentation met the validation criteria, and replicate measurements were consistent.\\n\\nPolicy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "At 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers."}, {"path": [], "text": "For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 11.20 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers.", "negative_left": "At 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers.", "negative_right": "For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 10.80 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing.", "right": "For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 11.20 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing."}, "verifier_independent_model": false}, "family": "scale-diverse-264-004", "id": "scale-diverse-264-004-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "Field note — microscopy workflow capture review\n\nAt 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers. For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 11.20 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing.\n\nReview found no major preparation defect, major capture defect, or major segmentation defect that materially biased the workflow’s measurements. Image artifacts were negligible. Segmentation met the validation criteria, and replicate measurements were consistent.\n\nPolicy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete scoring rubric and instructions, while both contexts retain the original mandatory-calibration policy without alteration. The workflow scope and calibration-record/time bindings are consistent; the two evidence spans are complete factual sentences; and the counterfactual coherently changes only the apparent spacing from 11.20 to 10.80 micrometers, producing a noncontradictory change from 12% to 8% under the stated formula. Neither context includes a level, answer code, proposition identifier, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"Field note — microscopy workflow capture review\\n\\nAt 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers. For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 11.20 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing.\\n\\nReview found no major preparation defect, major capture defect, or major segmentation defect that materially biased the workflow’s measurements. Image artifacts were negligible. Segmentation met the validation criteria, and replicate measurements were consistent.\\n\\nPolicy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "At 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers."}, {"path": [], "text": "For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 11.20 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers.", "negative_left": "At 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers.", "negative_right": "For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 10.80 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing.", "right": "For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 11.20 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing."}, "verifier_independent_model": false}, "family": "scale-diverse-264-004", "id": "scale-diverse-264-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "Field note — microscopy workflow capture review\n\nAt 2026-08-14T09:20:00Z, calibration record MC-47, the mandatory record used for quantitative size and density measurements in the microscopy workflow’s capture session, listed a certified feature spacing of 10.00 micrometers. For calibration record MC-47 at 2026-08-14T09:20:00Z, the capture log recorded an apparent feature spacing of 10.80 micrometers and calculated scale error as the absolute difference between apparent and certified spacing divided by certified spacing.\n\nReview found no major preparation defect, major capture defect, or major segmentation defect that materially biased the workflow’s measurements. Image artifacts were negligible. Segmentation met the validation criteria, and replicate measurements were consistent.\n\nPolicy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration-gate policy from the original state, while the unchanged questions preserve the full scoring rubric and instructions. The workflow, quantitative size/density measurement scope, calibration record, capture-session relationship, and relevant times remain aligned across the contexts. The two evidence spans are complete factual sentences reporting the certified spacing, calculation method recorded for the session, and measured spacing. The counterfactual changes only the measured spacing from 56.35 to 54.25 micrometres, without creating duplicate or contradictory measurements; all other facts remain coherent. Neither context states a reliability level, answer code, proposition identifier, case-specific label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%. At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 56.35 micrometres. At 09:15 UTC, review found no major preparation defect, capture defect, or segmentation defect materially biasing the microscopy workflow’s measurements. Image inspection found artifacts negligible. Segmentation met the documented validation criteria, and replicate measurements were consistent across the reviewed fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%."}, {"path": [], "text": "At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 56.35 micrometres."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%.", "negative_left": "At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%.", "negative_right": "At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 54.25 micrometres.", "right": "At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 56.35 micrometres."}, "verifier_independent_model": false}, "family": "scale-diverse-264-005", "id": "scale-diverse-264-005-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%. At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 56.35 micrometres. At 09:15 UTC, review found no major preparation defect, capture defect, or segmentation defect materially biasing the microscopy workflow’s measurements. Image inspection found artifacts negligible. Segmentation met the documented validation criteria, and replicate measurements were consistent across the reviewed fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration-gate policy from the original state, while the unchanged questions preserve the full scoring rubric and instructions. The workflow, quantitative size/density measurement scope, calibration record, capture-session relationship, and relevant times remain aligned across the contexts. The two evidence spans are complete factual sentences reporting the certified spacing, calculation method recorded for the session, and measured spacing. The counterfactual changes only the measured spacing from 56.35 to 54.25 micrometres, without creating duplicate or contradictory measurements; all other facts remain coherent. Neither context states a reliability level, answer code, proposition identifier, case-specific label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%. At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 56.35 micrometres. At 09:15 UTC, review found no major preparation defect, capture defect, or segmentation defect materially biasing the microscopy workflow’s measurements. Image inspection found artifacts negligible. Segmentation met the documented validation criteria, and replicate measurements were consistent across the reviewed fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%."}, {"path": [], "text": "At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 56.35 micrometres."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%.", "negative_left": "At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%.", "negative_right": "At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 54.25 micrometres.", "right": "At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 56.35 micrometres."}, "verifier_independent_model": false}, "family": "scale-diverse-264-005", "id": "scale-diverse-264-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "At 08:40 UTC on 14 May 2026, mandatory calibration record MC-47 for quantitative size and density measurements in the microscopy workflow’s capture session recorded a certified graticule spacing of 50.00 micrometres and specified its scale error as the absolute difference between measured and certified spacing divided by certified spacing, multiplied by 100%. At 08:52 UTC on 14 May 2026, calibration record MC-47 recorded the graticule’s measured spacing as 54.25 micrometres. At 09:15 UTC, review found no major preparation defect, capture defect, or segmentation defect materially biasing the microscopy workflow’s measurements. Image inspection found artifacts negligible. Segmentation met the documented validation criteria, and replicate measurements were consistent across the reviewed fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the complete protocol and exception without adding priorities, exceptions, or missing-evidence defaults, and they retain the same experiment, compliance question, and 14-day protocol scope. The two focus spans are complete factual measurement sentences. The sole counterfactual change moves the excursion end from 15:37 to 15:49, producing a coherent 97-minute excursion instead of the base context’s 85-minute excursion, with no duplicate or contradictory timing assertion. Neither context embeds an answer, code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"The replication reviewer reconciled the setup sheet, sampling register, and chamber audit for the experiment under review.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"The study record classifies the experiment as a moisture microcosm and shows that its incubation ran for 14 days.\",\"Setup records show that every jar began between 58% and 62% water-holding capacity, inclusive.\",\"The sampling register contains CO₂ measurements recorded on days 3, 7, and 14.\",\"The chamber audit identifies exactly one consecutive excursion outside 19°C to 21°C. It affected every treatment and paired control; at all other times, incubation remained between 19°C and 21°C, inclusive.\",\"The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026.\",\"The same incubation recorder logged the end of that temperature excursion at 15:37 UTC on 8 May 2026.\"],\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "6"], "text": "The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026."}, {"path": ["evidence", "7"], "text": "The same incubation recorder logged the end of that temperature excursion at 15:37 UTC on 8 May 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026.", "negative_left": "The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026.", "negative_right": "The same incubation recorder logged the end of that temperature excursion at 15:49 UTC on 8 May 2026.", "right": "The same incubation recorder logged the end of that temperature excursion at 15:37 UTC on 8 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-267-002", "id": "scale-diverse-267-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "The replication reviewer reconciled the setup sheet, sampling register, and chamber audit for the experiment under review.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "The study record classifies the experiment as a moisture microcosm and shows that its incubation ran for 14 days.", "Setup records show that every jar began between 58% and 62% water-holding capacity, inclusive.", "The sampling register contains CO₂ measurements recorded on days 3, 7, and 14.", "The chamber audit identifies exactly one consecutive excursion outside 19°C to 21°C. It affected every treatment and paired control; at all other times, incubation remained between 19°C and 21°C, inclusive.", "The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026.", "The same incubation recorder logged the end of that temperature excursion at 15:37 UTC on 8 May 2026."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the complete protocol and exception without adding priorities, exceptions, or missing-evidence defaults, and they retain the same experiment, compliance question, and 14-day protocol scope. The two focus spans are complete factual measurement sentences. The sole counterfactual change moves the excursion end from 15:37 to 15:49, producing a coherent 97-minute excursion instead of the base context’s 85-minute excursion, with no duplicate or contradictory timing assertion. Neither context embeds an answer, code, proposition ID, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"The replication reviewer reconciled the setup sheet, sampling register, and chamber audit for the experiment under review.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"The study record classifies the experiment as a moisture microcosm and shows that its incubation ran for 14 days.\",\"Setup records show that every jar began between 58% and 62% water-holding capacity, inclusive.\",\"The sampling register contains CO₂ measurements recorded on days 3, 7, and 14.\",\"The chamber audit identifies exactly one consecutive excursion outside 19°C to 21°C. It affected every treatment and paired control; at all other times, incubation remained between 19°C and 21°C, inclusive.\",\"The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026.\",\"The same incubation recorder logged the end of that temperature excursion at 15:37 UTC on 8 May 2026.\"],\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "6"], "text": "The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026."}, {"path": ["evidence", "7"], "text": "The same incubation recorder logged the end of that temperature excursion at 15:37 UTC on 8 May 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026.", "negative_left": "The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026.", "negative_right": "The same incubation recorder logged the end of that temperature excursion at 15:49 UTC on 8 May 2026.", "right": "The same incubation recorder logged the end of that temperature excursion at 15:37 UTC on 8 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-267-002", "id": "scale-diverse-267-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "The replication reviewer reconciled the setup sheet, sampling register, and chamber audit for the experiment under review.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "The study record classifies the experiment as a moisture microcosm and shows that its incubation ran for 14 days.", "Setup records show that every jar began between 58% and 62% water-holding capacity, inclusive.", "The sampling register contains CO₂ measurements recorded on days 3, 7, and 14.", "The chamber audit identifies exactly one consecutive excursion outside 19°C to 21°C. It affected every treatment and paired control; at all other times, incubation remained between 19°C and 21°C, inclusive.", "The incubation recorder logged the start of the single temperature excursion in the experiment under review at 14:12 UTC on 8 May 2026.", "The same incubation recorder logged the end of that temperature excursion at 15:49 UTC on 8 May 2026."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol requirements and temperature-excursion exception without alteration. The request remains bound to the same experiment, compliance decision, and 14-day protocol scope. The two focus-evidence spans are complete factual sentences recording the excursion's start and end. Changing the end time from 11:37 UTC to 12:04 UTC creates a coherent 106-minute excursion and does not conflict with the unchanged assertion that there was exactly one shared consecutive excursion. Neither context embeds an answer code, gold answer, rule table, proposition identifier, classifier instruction, or non-policy label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"An operations reviewer is completing the handoff check for a soil-science replication using the run sheet, startup inventory, chamber trace, and measurement register.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"The run sheet identifies the experiment under review as a moisture microcosm completed over an incubation duration of 14 days.\",\"The startup inventory confirms that every jar began between 58% and 62% water-holding capacity, inclusive.\",\"For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026.\",\"For the experiment under review, the operational handoff log records the same temperature excursion as ending at 11:37 UTC on 12 May 2026.\",\"The chamber trace documents exactly one consecutive temperature excursion outside 19°C to 21°C during incubation, shared by every treatment and paired control.\",\"At every recorded time outside that excursion, the incubation temperature remained between 19°C and 21°C, inclusive.\",\"The measurement register contains CO₂ readings recorded separately on day 3, day 7, and day 14.\",\"The reviewer has confirmed that the trace and registers belong to this experiment and cover its full incubation period.\"],\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "4"], "text": "For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026."}, {"path": ["evidence", "5"], "text": "For the experiment under review, the operational handoff log records the same temperature excursion as ending at 11:37 UTC on 12 May 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026.", "negative_left": "For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026.", "negative_right": "For the experiment under review, the operational handoff log records the same temperature excursion as ending at 12:04 UTC on 12 May 2026.", "right": "For the experiment under review, the operational handoff log records the same temperature excursion as ending at 11:37 UTC on 12 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-267-003", "id": "scale-diverse-267-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "An operations reviewer is completing the handoff check for a soil-science replication using the run sheet, startup inventory, chamber trace, and measurement register.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "The run sheet identifies the experiment under review as a moisture microcosm completed over an incubation duration of 14 days.", "The startup inventory confirms that every jar began between 58% and 62% water-holding capacity, inclusive.", "For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026.", "For the experiment under review, the operational handoff log records the same temperature excursion as ending at 11:37 UTC on 12 May 2026.", "The chamber trace documents exactly one consecutive temperature excursion outside 19°C to 21°C during incubation, shared by every treatment and paired control.", "At every recorded time outside that excursion, the incubation temperature remained between 19°C and 21°C, inclusive.", "The measurement register contains CO₂ readings recorded separately on day 3, day 7, and day 14.", "The reviewer has confirmed that the trace and registers belong to this experiment and cover its full incubation period."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol requirements and temperature-excursion exception without alteration. The request remains bound to the same experiment, compliance decision, and 14-day protocol scope. The two focus-evidence spans are complete factual sentences recording the excursion's start and end. Changing the end time from 11:37 UTC to 12:04 UTC creates a coherent 106-minute excursion and does not conflict with the unchanged assertion that there was exactly one shared consecutive excursion. Neither context embeds an answer code, gold answer, rule table, proposition identifier, classifier instruction, or non-policy label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"An operations reviewer is completing the handoff check for a soil-science replication using the run sheet, startup inventory, chamber trace, and measurement register.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"The run sheet identifies the experiment under review as a moisture microcosm completed over an incubation duration of 14 days.\",\"The startup inventory confirms that every jar began between 58% and 62% water-holding capacity, inclusive.\",\"For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026.\",\"For the experiment under review, the operational handoff log records the same temperature excursion as ending at 11:37 UTC on 12 May 2026.\",\"The chamber trace documents exactly one consecutive temperature excursion outside 19°C to 21°C during incubation, shared by every treatment and paired control.\",\"At every recorded time outside that excursion, the incubation temperature remained between 19°C and 21°C, inclusive.\",\"The measurement register contains CO₂ readings recorded separately on day 3, day 7, and day 14.\",\"The reviewer has confirmed that the trace and registers belong to this experiment and cover its full incubation period.\"],\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "4"], "text": "For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026."}, {"path": ["evidence", "5"], "text": "For the experiment under review, the operational handoff log records the same temperature excursion as ending at 11:37 UTC on 12 May 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026.", "negative_left": "For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026.", "negative_right": "For the experiment under review, the operational handoff log records the same temperature excursion as ending at 12:04 UTC on 12 May 2026.", "right": "For the experiment under review, the operational handoff log records the same temperature excursion as ending at 11:37 UTC on 12 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-267-003", "id": "scale-diverse-267-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "An operations reviewer is completing the handoff check for a soil-science replication using the run sheet, startup inventory, chamber trace, and measurement register.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "The run sheet identifies the experiment under review as a moisture microcosm completed over an incubation duration of 14 days.", "The startup inventory confirms that every jar began between 58% and 62% water-holding capacity, inclusive.", "For the experiment under review, the operational handoff log records the single temperature excursion as starting at 10:18 UTC on 12 May 2026.", "For the experiment under review, the operational handoff log records the same temperature excursion as ending at 12:04 UTC on 12 May 2026.", "The chamber trace documents exactly one consecutive temperature excursion outside 19°C to 21°C during incubation, shared by every treatment and paired control.", "At every recorded time outside that excursion, the incubation temperature remained between 19°C and 21°C, inclusive.", "The measurement register contains CO₂ readings recorded separately on day 3, day 7, and day 14.", "The reviewer has confirmed that the trace and registers belong to this experiment and cover its full incubation period."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing rubric, excluded-jar scope, and Cedar-centered decomposition-trial decision. The two evidence spans are complete factual sentences. The counterfactual changes only Cedar’s correction time from 07:10 to 05:50 on 15 April 2026, yielding a coherent continuous episode of 71 hours 15 minutes rather than 72 hours 35 minutes, with no duplicate or conflicting measurement. Neither context contains an answer code, explicit case answer, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"trial coordinator\",\"text\":\"The decomposition trial contained exactly two jars, Cedar and Elm. Cedar was the sealed test jar, while Elm was excluded from replication judgments.\"},{\"speaker\":\"record keeper\",\"text\":\"Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026.\"},{\"speaker\":\"record keeper\",\"text\":\"Cedar's correction in the decomposition trial was completed at 07:10 UTC on 15 April 2026.\"},{\"speaker\":\"quality officer\",\"text\":\"The incubation log documents exactly one continuous episode for Cedar outside 55–65% water-holding capacity. It began with that first low reading and ended when Cedar was corrected. Every sealing check for Cedar passed, as did every instrument check relevant to Cedar's moisture readings.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026."}, {"path": ["2", "text"], "text": "Cedar's correction in the decomposition trial was completed at 07:10 UTC on 15 April 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026.", "negative_left": "Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026.", "negative_right": "Cedar's correction in the decomposition trial was completed at 05:50 UTC on 15 April 2026.", "right": "Cedar's correction in the decomposition trial was completed at 07:10 UTC on 15 April 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-268-001", "id": "scale-diverse-268-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "trial coordinator", "text": "The decomposition trial contained exactly two jars, Cedar and Elm. Cedar was the sealed test jar, while Elm was excluded from replication judgments."}, {"speaker": "record keeper", "text": "Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026."}, {"speaker": "record keeper", "text": "Cedar's correction in the decomposition trial was completed at 07:10 UTC on 15 April 2026."}, {"speaker": "quality officer", "text": "The incubation log documents exactly one continuous episode for Cedar outside 55–65% water-holding capacity. It began with that first low reading and ended when Cedar was corrected. Every sealing check for Cedar passed, as did every instrument check relevant to Cedar's moisture readings."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original routing rubric, excluded-jar scope, and Cedar-centered decomposition-trial decision. The two evidence spans are complete factual sentences. The counterfactual changes only Cedar’s correction time from 07:10 to 05:50 on 15 April 2026, yielding a coherent continuous episode of 71 hours 15 minutes rather than 72 hours 35 minutes, with no duplicate or conflicting measurement. Neither context contains an answer code, explicit case answer, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"trial coordinator\",\"text\":\"The decomposition trial contained exactly two jars, Cedar and Elm. Cedar was the sealed test jar, while Elm was excluded from replication judgments.\"},{\"speaker\":\"record keeper\",\"text\":\"Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026.\"},{\"speaker\":\"record keeper\",\"text\":\"Cedar's correction in the decomposition trial was completed at 07:10 UTC on 15 April 2026.\"},{\"speaker\":\"quality officer\",\"text\":\"The incubation log documents exactly one continuous episode for Cedar outside 55–65% water-holding capacity. It began with that first low reading and ended when Cedar was corrected. Every sealing check for Cedar passed, as did every instrument check relevant to Cedar's moisture readings.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026."}, {"path": ["2", "text"], "text": "Cedar's correction in the decomposition trial was completed at 07:10 UTC on 15 April 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026.", "negative_left": "Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026.", "negative_right": "Cedar's correction in the decomposition trial was completed at 05:50 UTC on 15 April 2026.", "right": "Cedar's correction in the decomposition trial was completed at 07:10 UTC on 15 April 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-268-001", "id": "scale-diverse-268-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "trial coordinator", "text": "The decomposition trial contained exactly two jars, Cedar and Elm. Cedar was the sealed test jar, while Elm was excluded from replication judgments."}, {"speaker": "record keeper", "text": "Cedar's first low moisture reading on day 2 of the decomposition trial was recorded at 06:35 UTC on 12 April 2026."}, {"speaker": "record keeper", "text": "Cedar's correction in the decomposition trial was completed at 05:50 UTC on 15 April 2026."}, {"speaker": "quality officer", "text": "The incubation log documents exactly one continuous episode for Cedar outside 55–65% water-holding capacity. It began with that first low reading and ended when Cedar was corrected. Every sealing check for Cedar passed, as did every instrument check relevant to Cedar's moisture readings."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing rubric without adding exceptions or defaults and preserve the decision about Cedar’s decomposition-trial incubation episode. The two focus-evidence spans are complete factual timestamp sentences. The sole counterfactual change moves Cedar’s correction from 08:35 to 05:50 UTC on 17 May 2026, which remains consistent with the single continuous episode and creates no duplicate or contradictory measurement. Neither context contains a gold answer, output code, proposition identifier, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"trial registrar\",\"text\":\"The decomposition-trial inventory contains exactly two jars: Cedar and Elm. Cedar is recorded as a sealed test jar, while Elm is excluded from replication judgments.\"},{\"speaker\":\"moisture log reviewer\",\"text\":\"Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. That episode began with its first low moisture reading on day 2 and ended when Cedar was corrected.\"},{\"speaker\":\"timestamp record\",\"text\":\"In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.\"},{\"speaker\":\"correction record\",\"text\":\"In the decomposition trial, Cedar's correction was completed at 08:35 UTC on 17 May 2026.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Every sealing check for Cedar during the trial passed. Every instrument check relevant to Cedar's moisture readings also passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026."}, {"path": ["3", "text"], "text": "In the decomposition trial, Cedar's correction was completed at 08:35 UTC on 17 May 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.", "negative_left": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.", "negative_right": "In the decomposition trial, Cedar's correction was completed at 05:50 UTC on 17 May 2026.", "right": "In the decomposition trial, Cedar's correction was completed at 08:35 UTC on 17 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-268-002", "id": "scale-diverse-268-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "trial registrar", "text": "The decomposition-trial inventory contains exactly two jars: Cedar and Elm. Cedar is recorded as a sealed test jar, while Elm is excluded from replication judgments."}, {"speaker": "moisture log reviewer", "text": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. That episode began with its first low moisture reading on day 2 and ended when Cedar was corrected."}, {"speaker": "timestamp record", "text": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026."}, {"speaker": "correction record", "text": "In the decomposition trial, Cedar's correction was completed at 08:35 UTC on 17 May 2026."}, {"speaker": "quality reviewer", "text": "Every sealing check for Cedar during the trial passed. Every instrument check relevant to Cedar's moisture readings also passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing rubric without adding exceptions or defaults and preserve the decision about Cedar’s decomposition-trial incubation episode. The two focus-evidence spans are complete factual timestamp sentences. The sole counterfactual change moves Cedar’s correction from 08:35 to 05:50 UTC on 17 May 2026, which remains consistent with the single continuous episode and creates no duplicate or contradictory measurement. Neither context contains a gold answer, output code, proposition identifier, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"trial registrar\",\"text\":\"The decomposition-trial inventory contains exactly two jars: Cedar and Elm. Cedar is recorded as a sealed test jar, while Elm is excluded from replication judgments.\"},{\"speaker\":\"moisture log reviewer\",\"text\":\"Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. That episode began with its first low moisture reading on day 2 and ended when Cedar was corrected.\"},{\"speaker\":\"timestamp record\",\"text\":\"In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.\"},{\"speaker\":\"correction record\",\"text\":\"In the decomposition trial, Cedar's correction was completed at 08:35 UTC on 17 May 2026.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Every sealing check for Cedar during the trial passed. Every instrument check relevant to Cedar's moisture readings also passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026."}, {"path": ["3", "text"], "text": "In the decomposition trial, Cedar's correction was completed at 08:35 UTC on 17 May 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.", "negative_left": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.", "negative_right": "In the decomposition trial, Cedar's correction was completed at 05:50 UTC on 17 May 2026.", "right": "In the decomposition trial, Cedar's correction was completed at 08:35 UTC on 17 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-268-002", "id": "scale-diverse-268-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "trial registrar", "text": "The decomposition-trial inventory contains exactly two jars: Cedar and Elm. Cedar is recorded as a sealed test jar, while Elm is excluded from replication judgments."}, {"speaker": "moisture log reviewer", "text": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. That episode began with its first low moisture reading on day 2 and ended when Cedar was corrected."}, {"speaker": "timestamp record", "text": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026."}, {"speaker": "correction record", "text": "In the decomposition trial, Cedar's correction was completed at 05:50 UTC on 17 May 2026."}, {"speaker": "quality reviewer", "text": "Every sealing check for Cedar during the trial passed. Every instrument check relevant to Cedar's moisture readings also passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original incubation-routing threshold, exact-72-hours qualification, excluded-jar rule, and precedence condition without adding exceptions or defaults. The question remains bound to whether Cedar should be routed as an incubation issue based on its continuous excursion and correction time. The two focus-evidence spans are complete factual sentences. Changing Cedar’s correction from 07:05 to 05:50 on 14 April yields a coherent duration change from 72 hours 45 minutes to 71 hours 30 minutes, with no conflicting duplicate timestamp. Neither context states a gold decision, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"handoff coordinator\",\"text\":\"Trial inventory closure lists exactly two jars: Cedar and Elm. Cedar is designated as the sealed test jar, while Elm is excluded from replication judgments.\"},{\"speaker\":\"incubation technician\",\"text\":\"The excursion register shows Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. It began with Cedar's first low moisture reading on day 2 and ended when Cedar was corrected.\"},{\"speaker\":\"operations recorder\",\"text\":\"The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026.\"},{\"speaker\":\"operations recorder\",\"text\":\"The operational handoff log records Cedar's correction at 07:05 UTC on 14 April 2026.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Every sealing check for Cedar during the decomposition trial passed. Every instrument check relevant to Cedar's moisture readings during the trial also passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026."}, {"path": ["3", "text"], "text": "The operational handoff log records Cedar's correction at 07:05 UTC on 14 April 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026.", "negative_left": "The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026.", "negative_right": "The operational handoff log records Cedar's correction at 05:50 UTC on 14 April 2026.", "right": "The operational handoff log records Cedar's correction at 07:05 UTC on 14 April 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-268-003", "id": "scale-diverse-268-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "handoff coordinator", "text": "Trial inventory closure lists exactly two jars: Cedar and Elm. Cedar is designated as the sealed test jar, while Elm is excluded from replication judgments."}, {"speaker": "incubation technician", "text": "The excursion register shows Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. It began with Cedar's first low moisture reading on day 2 and ended when Cedar was corrected."}, {"speaker": "operations recorder", "text": "The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026."}, {"speaker": "operations recorder", "text": "The operational handoff log records Cedar's correction at 07:05 UTC on 14 April 2026."}, {"speaker": "quality reviewer", "text": "Every sealing check for Cedar during the decomposition trial passed. Every instrument check relevant to Cedar's moisture readings during the trial also passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original incubation-routing threshold, exact-72-hours qualification, excluded-jar rule, and precedence condition without adding exceptions or defaults. The question remains bound to whether Cedar should be routed as an incubation issue based on its continuous excursion and correction time. The two focus-evidence spans are complete factual sentences. Changing Cedar’s correction from 07:05 to 05:50 on 14 April yields a coherent duration change from 72 hours 45 minutes to 71 hours 30 minutes, with no conflicting duplicate timestamp. Neither context states a gold decision, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"handoff coordinator\",\"text\":\"Trial inventory closure lists exactly two jars: Cedar and Elm. Cedar is designated as the sealed test jar, while Elm is excluded from replication judgments.\"},{\"speaker\":\"incubation technician\",\"text\":\"The excursion register shows Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. It began with Cedar's first low moisture reading on day 2 and ended when Cedar was corrected.\"},{\"speaker\":\"operations recorder\",\"text\":\"The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026.\"},{\"speaker\":\"operations recorder\",\"text\":\"The operational handoff log records Cedar's correction at 07:05 UTC on 14 April 2026.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Every sealing check for Cedar during the decomposition trial passed. Every instrument check relevant to Cedar's moisture readings during the trial also passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026."}, {"path": ["3", "text"], "text": "The operational handoff log records Cedar's correction at 07:05 UTC on 14 April 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026.", "negative_left": "The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026.", "negative_right": "The operational handoff log records Cedar's correction at 05:50 UTC on 14 April 2026.", "right": "The operational handoff log records Cedar's correction at 07:05 UTC on 14 April 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-268-003", "id": "scale-diverse-268-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "handoff coordinator", "text": "Trial inventory closure lists exactly two jars: Cedar and Elm. Cedar is designated as the sealed test jar, while Elm is excluded from replication judgments."}, {"speaker": "incubation technician", "text": "The excursion register shows Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. It began with Cedar's first low moisture reading on day 2 and ended when Cedar was corrected."}, {"speaker": "operations recorder", "text": "The operational handoff log records Cedar's first low moisture reading on day 2 at 06:20 UTC on 11 April 2026."}, {"speaker": "operations recorder", "text": "The operational handoff log records Cedar's correction at 05:50 UTC on 14 April 2026."}, {"speaker": "quality reviewer", "text": "Every sealing check for Cedar during the decomposition trial passed. Every instrument check relevant to Cedar's moisture readings during the trial also passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the unchanged question and reproduce the governing routing policy without adding exceptions, priorities, or missing-evidence defaults. The decomposition-trial, Cedar, incubation-duration, and routing-decision bindings remain intact; the two evidence spans are complete factual sentences; and the sole counterfactual change moves Cedar’s correction time from 07:05 to 05:50 while remaining consistent with the unchanged start time and continuous-episode description. Neither context embeds an answer, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"field registrar\",\"text\":\"The decomposition trial’s complete jar roster contained Cedar and Elm only. Cedar was the sealed test jar, while Elm was excluded from replication judgments.\"},{\"speaker\":\"incubation logger\",\"text\":\"In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.\"},{\"speaker\":\"incubation logger\",\"text\":\"Cedar's moisture correction was completed at 07:05 UTC on 17 May 2026.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. That episode began with its first low moisture reading on day 2 and ended when Cedar was corrected. Every sealing check for Cedar passed, as did every instrument check relevant to Cedar’s moisture readings.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026."}, {"path": ["2", "text"], "text": "Cedar's moisture correction was completed at 07:05 UTC on 17 May 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.", "negative_left": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.", "negative_right": "Cedar's moisture correction was completed at 05:50 UTC on 17 May 2026.", "right": "Cedar's moisture correction was completed at 07:05 UTC on 17 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-268-004", "id": "scale-diverse-268-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "field registrar", "text": "The decomposition trial’s complete jar roster contained Cedar and Elm only. Cedar was the sealed test jar, while Elm was excluded from replication judgments."}, {"speaker": "incubation logger", "text": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026."}, {"speaker": "incubation logger", "text": "Cedar's moisture correction was completed at 07:05 UTC on 17 May 2026."}, {"speaker": "quality reviewer", "text": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. That episode began with its first low moisture reading on day 2 and ended when Cedar was corrected. Every sealing check for Cedar passed, as did every instrument check relevant to Cedar’s moisture readings."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the unchanged question and reproduce the governing routing policy without adding exceptions, priorities, or missing-evidence defaults. The decomposition-trial, Cedar, incubation-duration, and routing-decision bindings remain intact; the two evidence spans are complete factual sentences; and the sole counterfactual change moves Cedar’s correction time from 07:05 to 05:50 while remaining consistent with the unchanged start time and continuous-episode description. Neither context embeds an answer, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"field registrar\",\"text\":\"The decomposition trial’s complete jar roster contained Cedar and Elm only. Cedar was the sealed test jar, while Elm was excluded from replication judgments.\"},{\"speaker\":\"incubation logger\",\"text\":\"In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.\"},{\"speaker\":\"incubation logger\",\"text\":\"Cedar's moisture correction was completed at 07:05 UTC on 17 May 2026.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. That episode began with its first low moisture reading on day 2 and ended when Cedar was corrected. Every sealing check for Cedar passed, as did every instrument check relevant to Cedar’s moisture readings.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026."}, {"path": ["2", "text"], "text": "Cedar's moisture correction was completed at 07:05 UTC on 17 May 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.", "negative_left": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026.", "negative_right": "Cedar's moisture correction was completed at 05:50 UTC on 17 May 2026.", "right": "Cedar's moisture correction was completed at 07:05 UTC on 17 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-268-004", "id": "scale-diverse-268-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "field registrar", "text": "The decomposition trial’s complete jar roster contained Cedar and Elm only. Cedar was the sealed test jar, while Elm was excluded from replication judgments."}, {"speaker": "incubation logger", "text": "In the decomposition trial, Cedar's first low moisture reading on day 2 was logged at 06:20 UTC on 14 May 2026."}, {"speaker": "incubation logger", "text": "Cedar's moisture correction was completed at 05:50 UTC on 17 May 2026."}, {"speaker": "quality reviewer", "text": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. That episode began with its first low moisture reading on day 2 and ended when Cedar was corrected. Every sealing check for Cedar passed, as did every instrument check relevant to Cedar’s moisture readings."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing hierarchy, substitution-approval scope, shipment/PO/ASN entities, and immediate-routing request. The two focus spans are complete factual sentences; the counterfactual changes only the scanner-recorded identifier from VX-4092 to VX-4087, which coherently removes the identity discrepancy without conflicting with the unchanged count, condition, or supplier-fault evidence, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"A reconciliation review at dock 3 compares shipment R-184 with PO-771 and ASN-771A before receipt posting.\",\"evidence\":[\"For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item.\",\"At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4092 as the physical item identifier on shipment R-184.\",\"PO-771 and ASN-771A specify the same physical item identifier and each lists 480 units. The verified physical count is also 480 units. Dock inspection recorded neither damage nor a safety concern. PO-771 was released on 6 May, and Inventory Control has not verified supplier fault for this shipment.\",\"Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\",\"Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\"],\"request\":\"Select the single immediate routing classification for this shipment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item."}, {"path": ["evidence", "1"], "text": "At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4092 as the physical item identifier on shipment R-184."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item.", "negative_left": "For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item.", "negative_right": "At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4087 as the physical item identifier on shipment R-184.", "right": "At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4092 as the physical item identifier on shipment R-184."}, "verifier_independent_model": false}, "family": "scale-diverse-277-002", "id": "scale-diverse-277-002-base", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "A reconciliation review at dock 3 compares shipment R-184 with PO-771 and ASN-771A before receipt posting.", "evidence": ["For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item.", "At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4092 as the physical item identifier on shipment R-184.", "PO-771 and ASN-771A specify the same physical item identifier and each lists 480 units. The verified physical count is also 480 units. Dock inspection recorded neither damage nor a safety concern. PO-771 was released on 6 May, and Inventory Control has not verified supplier fault for this shipment.", "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.", "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."], "request": "Select the single immediate routing classification for this shipment."}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "INVENTORY_IDENTITY_HOLD"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing hierarchy, substitution-approval scope, shipment/PO/ASN entities, and immediate-routing request. The two focus spans are complete factual sentences; the counterfactual changes only the scanner-recorded identifier from VX-4092 to VX-4087, which coherently removes the identity discrepancy without conflicting with the unchanged count, condition, or supplier-fault evidence, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"A reconciliation review at dock 3 compares shipment R-184 with PO-771 and ASN-771A before receipt posting.\",\"evidence\":[\"For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item.\",\"At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4092 as the physical item identifier on shipment R-184.\",\"PO-771 and ASN-771A specify the same physical item identifier and each lists 480 units. The verified physical count is also 480 units. Dock inspection recorded neither damage nor a safety concern. PO-771 was released on 6 May, and Inventory Control has not verified supplier fault for this shipment.\",\"Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\",\"Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\"],\"request\":\"Select the single immediate routing classification for this shipment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item."}, {"path": ["evidence", "1"], "text": "At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4092 as the physical item identifier on shipment R-184."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item.", "negative_left": "For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item.", "negative_right": "At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4087 as the physical item identifier on shipment R-184.", "right": "At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4092 as the physical item identifier on shipment R-184."}, "verifier_independent_model": false}, "family": "scale-diverse-277-002", "id": "scale-diverse-277-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "A reconciliation review at dock 3 compares shipment R-184 with PO-771 and ASN-771A before receipt posting.", "evidence": ["For shipment R-184, PO-771 specifies item identifier VX-4087 for the V40 item.", "At 14:32 UTC on 8 April 2026, the dock 3 scanner recorded VX-4087 as the physical item identifier on shipment R-184.", "PO-771 and ASN-771A specify the same physical item identifier and each lists 480 units. The verified physical count is also 480 units. Dock inspection recorded neither damage nor a safety concern. PO-771 was released on 6 May, and Inventory Control has not verified supplier fault for this shipment.", "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.", "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."], "request": "Select the single immediate routing classification for this shipment."}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "RECEIVING_CLEAN"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original approval scope and routing policy without adding exceptions, priorities, or missing-evidence defaults. The shipment, PO, ASN, dock, request, and immediate-routing decision remain bound consistently; the two focus spans are complete factual sentences, and the counterfactual changes only the observed physical identifier so that it coherently matches the unchanged PO and ASN identifier. Neither context states a selected classification, answer code, proposition ID, or classifier instruction beyond the preserved request and natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"Receiving staff prepared an operational handoff for shipment R-184 at dock 3 after checking the delivered goods against PO-771 and ASN-771A.\",\"evidence\":[\"During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-4827.\",\"PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184.\",\"ASN-771A carries the same item identifier as PO-771. PO-771 and ASN-771A each specify 480 units, and the dock count found 480 units. Inspection found no damage and no safety concern. PO-771 was released on 6 May. Inventory Control has not verified supplier fault for R-184.\",\"Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\",\"Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\"],\"request\":\"Select the single immediate routing classification for this shipment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-4827."}, {"path": ["evidence", "1"], "text": "PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-4827.", "negative_left": "During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-9364.", "negative_right": "PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184.", "right": "PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184."}, "verifier_independent_model": false}, "family": "scale-diverse-277-003", "id": "scale-diverse-277-003-base", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "Receiving staff prepared an operational handoff for shipment R-184 at dock 3 after checking the delivered goods against PO-771 and ASN-771A.", "evidence": ["During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-4827.", "PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184.", "ASN-771A carries the same item identifier as PO-771. PO-771 and ASN-771A each specify 480 units, and the dock count found 480 units. Inspection found no damage and no safety concern. PO-771 was released on 6 May. Inventory Control has not verified supplier fault for R-184.", "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.", "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."], "request": "Select the single immediate routing classification for this shipment."}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "INVENTORY_IDENTITY_HOLD"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original approval scope and routing policy without adding exceptions, priorities, or missing-evidence defaults. The shipment, PO, ASN, dock, request, and immediate-routing decision remain bound consistently; the two focus spans are complete factual sentences, and the counterfactual changes only the observed physical identifier so that it coherently matches the unchanged PO and ASN identifier. Neither context states a selected classification, answer code, proposition ID, or classifier instruction beyond the preserved request and natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"Receiving staff prepared an operational handoff for shipment R-184 at dock 3 after checking the delivered goods against PO-771 and ASN-771A.\",\"evidence\":[\"During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-4827.\",\"PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184.\",\"ASN-771A carries the same item identifier as PO-771. PO-771 and ASN-771A each specify 480 units, and the dock count found 480 units. Inspection found no damage and no safety concern. PO-771 was released on 6 May. Inventory Control has not verified supplier fault for R-184.\",\"Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\",\"Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\"],\"request\":\"Select the single immediate routing classification for this shipment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-4827."}, {"path": ["evidence", "1"], "text": "PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-4827.", "negative_left": "During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-9364.", "negative_right": "PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184.", "right": "PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184."}, "verifier_independent_model": false}, "family": "scale-diverse-277-003", "id": "scale-diverse-277-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "Receiving staff prepared an operational handoff for shipment R-184 at dock 3 after checking the delivered goods against PO-771 and ASN-771A.", "evidence": ["During the operational handoff at dock 3 on 8 May 2026 at 14:32 UTC, the physical item identifier observed for shipment R-184 was QX-9364.", "PO-771 specifies QX-9364 as the V40 item identifier for shipment R-184.", "ASN-771A carries the same item identifier as PO-771. PO-771 and ASN-771A each specify 480 units, and the dock count found 480 units. Inspection found no damage and no safety concern. PO-771 was released on 6 May. Inventory Control has not verified supplier fault for R-184.", "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.", "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."], "request": "Select the single immediate routing classification for this shipment."}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "RECEIVING_CLEAN"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions object and preserve the original routing and substitution policies without adding exceptions or defaults. The shipment, PO, ASN, dock, request scope, and relevant dates remain bound consistently. The two focus spans are complete factual sentences. The counterfactual changes only the physical scan identifier from V40B to V40, which is coherent with the unchanged quantity, document, damage, and supplier-fault evidence and creates no duplicate contradictory measurement. Neither context embeds a gold label, answer code, rationale, proposition ID, rule table, or output instruction beyond the legitimate request and natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"Dock 3 maintained a chronological receiving record for shipment R-184 against PO-771 and ASN-771A.\",\"evidence\":[\"PO-771 was released on 6 May 2026, before 10 May.\",\"At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40.\",\"On 8 May, document reconciliation confirmed that ASN-771A and PO-771 used the same physical item identifier for R-184.\",\"Both PO-771 and ASN-771A specified 480 units for the shipment.\",\"At receiving, the physical count was 480 units, matching the quantity in PO-771.\",\"At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40B.\",\"The completed dock inspection recorded no observed damage and no safety concern.\",\"Inventory Control confirmed that supplier fault had not been verified.\",\"Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\",\"Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\"],\"request\":\"Select the single immediate routing classification for this shipment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40."}, {"path": ["evidence", "5"], "text": "At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40B."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40.", "negative_left": "At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40.", "negative_right": "At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40.", "right": "At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40B."}, "verifier_independent_model": false}, "family": "scale-diverse-277-005", "id": "scale-diverse-277-005-base", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "Dock 3 maintained a chronological receiving record for shipment R-184 against PO-771 and ASN-771A.", "evidence": ["PO-771 was released on 6 May 2026, before 10 May.", "At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40.", "On 8 May, document reconciliation confirmed that ASN-771A and PO-771 used the same physical item identifier for R-184.", "Both PO-771 and ASN-771A specified 480 units for the shipment.", "At receiving, the physical count was 480 units, matching the quantity in PO-771.", "At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40B.", "The completed dock inspection recorded no observed damage and no safety concern.", "Inventory Control confirmed that supplier fault had not been verified.", "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.", "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."], "request": "Select the single immediate routing classification for this shipment."}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "INVENTORY_IDENTITY_HOLD"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions object and preserve the original routing and substitution policies without adding exceptions or defaults. The shipment, PO, ASN, dock, request scope, and relevant dates remain bound consistently. The two focus spans are complete factual sentences. The counterfactual changes only the physical scan identifier from V40B to V40, which is coherent with the unchanged quantity, document, damage, and supplier-fault evidence and creates no duplicate contradictory measurement. Neither context embeds a gold label, answer code, rationale, proposition ID, rule table, or output instruction beyond the legitimate request and natural policy language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"Dock 3 maintained a chronological receiving record for shipment R-184 against PO-771 and ASN-771A.\",\"evidence\":[\"PO-771 was released on 6 May 2026, before 10 May.\",\"At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40.\",\"On 8 May, document reconciliation confirmed that ASN-771A and PO-771 used the same physical item identifier for R-184.\",\"Both PO-771 and ASN-771A specified 480 units for the shipment.\",\"At receiving, the physical count was 480 units, matching the quantity in PO-771.\",\"At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40B.\",\"The completed dock inspection recorded no observed damage and no safety concern.\",\"Inventory Control confirmed that supplier fault had not been verified.\",\"Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\",\"Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\"],\"request\":\"Select the single immediate routing classification for this shipment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40."}, {"path": ["evidence", "5"], "text": "At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40B."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40.", "negative_left": "At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40.", "negative_right": "At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40.", "right": "At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40B."}, "verifier_independent_model": false}, "family": "scale-diverse-277-005", "id": "scale-diverse-277-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "Dock 3 maintained a chronological receiving record for shipment R-184 against PO-771 and ASN-771A.", "evidence": ["PO-771 was released on 6 May 2026, before 10 May.", "At 09:20 UTC on 7 May 2026, the item line for shipment R-184 in PO-771 specified the physical item identifier V40.", "On 8 May, document reconciliation confirmed that ASN-771A and PO-771 used the same physical item identifier for R-184.", "Both PO-771 and ASN-771A specified 480 units for the shipment.", "At receiving, the physical count was 480 units, matching the quantity in PO-771.", "At 14:35 UTC on 12 May 2026, a dock 3 scan of every unit in shipment R-184 recorded the physical item identifier V40.", "The completed dock inspection recorded no observed damage and no safety concern.", "Inventory Control confirmed that supplier fault had not been verified.", "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.", "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."], "request": "Select the single immediate routing classification for this shipment."}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "RECEIVING_CLEAN"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing priority and substitution thresholds without adding exceptions or defaults, and they continue to concern PO-731 and the same routing-and-severity decision scope. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only the affected substitution quantity from 37 to 44 units; this does not conflict with the unchanged order quantity, absence of damage, lack of approval, or audit findings. Neither context embeds an answer choice, code, proposition ID, classifier instruction, or answer-specific rationale; the routing and threshold language is permissible governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving record\",\"text\":\"At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731.\"},{\"speaker\":\"Procurement reviewer\",\"text\":\"At 10:25 UTC, procurement confirmed that no substitution request associated with this delivery had been approved. The receiving audit was then completed against the purchase order and approval file.\"},{\"speaker\":\"Receiving audit\",\"text\":\"At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 37 of the units received under PO-731.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731."}, {"path": ["2", "text"], "text": "At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 37 of the units received under PO-731."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731.", "negative_left": "At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731.", "negative_right": "At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 44 of the units received under PO-731.", "right": "At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 37 of the units received under PO-731."}, "verifier_independent_model": false}, "family": "scale-diverse-278-001", "id": "scale-diverse-278-001-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving record", "text": "At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731."}, {"speaker": "Procurement reviewer", "text": "At 10:25 UTC, procurement confirmed that no substitution request associated with this delivery had been approved. The receiving audit was then completed against the purchase order and approval file."}, {"speaker": "Receiving audit", "text": "At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 37 of the units received under PO-731."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing priority and substitution thresholds without adding exceptions or defaults, and they continue to concern PO-731 and the same routing-and-severity decision scope. The focus evidence consists of exactly two complete factual sentences. The counterfactual changes only the affected substitution quantity from 37 to 44 units; this does not conflict with the unchanged order quantity, absence of damage, lack of approval, or audit findings. Neither context embeds an answer choice, code, proposition ID, classifier instruction, or answer-specific rationale; the routing and threshold language is permissible governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving record\",\"text\":\"At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731.\"},{\"speaker\":\"Procurement reviewer\",\"text\":\"At 10:25 UTC, procurement confirmed that no substitution request associated with this delivery had been approved. The receiving audit was then completed against the purchase order and approval file.\"},{\"speaker\":\"Receiving audit\",\"text\":\"At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 37 of the units received under PO-731.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731."}, {"path": ["2", "text"], "text": "At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 37 of the units received under PO-731."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731.", "negative_left": "At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731.", "negative_right": "At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 44 of the units received under PO-731.", "right": "At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 37 of the units received under PO-731."}, "verifier_independent_model": false}, "family": "scale-diverse-278-001", "id": "scale-diverse-278-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving record", "text": "At 09:10 UTC on 14 August 2026, the receiving record for PO-731 documented 370 units ordered and no physical damage among the units received under PO-731."}, {"speaker": "Procurement reviewer", "text": "At 10:25 UTC, procurement confirmed that no substitution request associated with this delivery had been approved. The receiving audit was then completed against the purchase order and approval file."}, {"speaker": "Receiving audit", "text": "At 11:40 UTC on 14 August 2026, the completed receiving audit identified an unauthorized substitution affecting 44 of the units received under PO-731."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing priority and substitution threshold policy without adding exceptions, priorities, or missing-evidence defaults. They remain bound to PO-731 and the same receipt-routing decision scope; the counterfactual changes only the substitution count from 61 to 73 while leaving the 680-unit order and all other observations coherent. The two focus-evidence spans are complete factual sentences rather than instructions or policy definitions. Neither context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Records custodian\",\"text\":\"The receiving file is closed and authenticated. Its signed reconciliation and final count concern the same completed receipt under PO-731, cover all received lots, and have not been superseded or amended.\"},{\"speaker\":\"Receiving auditor\",\"text\":\"The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units.\"},{\"speaker\":\"Procurement compliance officer\",\"text\":\"For this receipt, the term unauthorized substitution denotes a received unit supplied in place of the ordered item without written purchasing approval. The final count uses that designation consistently with the signed reconciliation.\"},{\"speaker\":\"Count supervisor\",\"text\":\"The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 61 received units as unauthorized substitutions.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Cross-checking found no duplicate count lines, omitted receipt lots, unresolved item classifications, or later inspection findings. The two signed records therefore form the complete evidence set for routing this receipt.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units."}, {"path": ["3", "text"], "text": "The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 61 received units as unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units.", "negative_left": "The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units.", "negative_right": "The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 73 received units as unauthorized substitutions.", "right": "The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 61 received units as unauthorized substitutions."}, "verifier_independent_model": false}, "family": "scale-diverse-278-002", "id": "scale-diverse-278-002-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Records custodian", "text": "The receiving file is closed and authenticated. Its signed reconciliation and final count concern the same completed receipt under PO-731, cover all received lots, and have not been superseded or amended."}, {"speaker": "Receiving auditor", "text": "The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units."}, {"speaker": "Procurement compliance officer", "text": "For this receipt, the term unauthorized substitution denotes a received unit supplied in place of the ordered item without written purchasing approval. The final count uses that designation consistently with the signed reconciliation."}, {"speaker": "Count supervisor", "text": "The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 61 received units as unauthorized substitutions."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Inventory control analyst", "text": "Cross-checking found no duplicate count lines, omitted receipt lots, unresolved item classifications, or later inspection findings. The two signed records therefore form the complete evidence set for routing this receipt."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing priority and substitution threshold policy without adding exceptions, priorities, or missing-evidence defaults. They remain bound to PO-731 and the same receipt-routing decision scope; the counterfactual changes only the substitution count from 61 to 73 while leaving the 680-unit order and all other observations coherent. The two focus-evidence spans are complete factual sentences rather than instructions or policy definitions. Neither context contains a gold answer, answer code, proposition identifier, label rationale, rule table, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Records custodian\",\"text\":\"The receiving file is closed and authenticated. Its signed reconciliation and final count concern the same completed receipt under PO-731, cover all received lots, and have not been superseded or amended.\"},{\"speaker\":\"Receiving auditor\",\"text\":\"The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units.\"},{\"speaker\":\"Procurement compliance officer\",\"text\":\"For this receipt, the term unauthorized substitution denotes a received unit supplied in place of the ordered item without written purchasing approval. The final count uses that designation consistently with the signed reconciliation.\"},{\"speaker\":\"Count supervisor\",\"text\":\"The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 61 received units as unauthorized substitutions.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Cross-checking found no duplicate count lines, omitted receipt lots, unresolved item classifications, or later inspection findings. The two signed records therefore form the complete evidence set for routing this receipt.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units."}, {"path": ["3", "text"], "text": "The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 61 received units as unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units.", "negative_left": "The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units.", "negative_right": "The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 73 received units as unauthorized substitutions.", "right": "The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 61 received units as unauthorized substitutions."}, "verifier_independent_model": false}, "family": "scale-diverse-278-002", "id": "scale-diverse-278-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Records custodian", "text": "The receiving file is closed and authenticated. Its signed reconciliation and final count concern the same completed receipt under PO-731, cover all received lots, and have not been superseded or amended."}, {"speaker": "Receiving auditor", "text": "The signed reconciliation for PO-731, completed on 14 August 2026, records 680 units ordered, no physical damage among the units received, and the presence of unauthorized substitutions among those units."}, {"speaker": "Procurement compliance officer", "text": "For this receipt, the term unauthorized substitution denotes a received unit supplied in place of the ordered item without written purchasing approval. The final count uses that designation consistently with the signed reconciliation."}, {"speaker": "Count supervisor", "text": "The final receiving count for PO-731, completed on 15 August 2026, identifies exactly 73 received units as unauthorized substitutions."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Inventory control analyst", "text": "Cross-checking found no duplicate count lines, omitted receipt lots, unresolved item classifications, or later inspection findings. The two signed records therefore form the complete evidence set for routing this receipt."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all choice criteria and instructions, while both contexts retain the original routing priority and substitution threshold without alteration. Both remain bound to PO-731 and the same decision path; the two evidence spans are complete factual sentences, and changing substitutions from 37 to 52 against 430 ordered units creates a coherent threshold-crossing counterfactual without conflicting measurements or embedded answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving lead\",\"text\":\"The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received.\"},{\"speaker\":\"Inspection lead\",\"text\":\"The inspection completed at 14:35 UTC on 12 August 2026 verified that 37 units received under PO-731 were unauthorized substitutions.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received."}, {"path": ["1", "text"], "text": "The inspection completed at 14:35 UTC on 12 August 2026 verified that 37 units received under PO-731 were unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received.", "negative_left": "The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received.", "negative_right": "The inspection completed at 14:35 UTC on 12 August 2026 verified that 52 units received under PO-731 were unauthorized substitutions.", "right": "The inspection completed at 14:35 UTC on 12 August 2026 verified that 37 units received under PO-731 were unauthorized substitutions."}, "verifier_independent_model": false}, "family": "scale-diverse-278-003", "id": "scale-diverse-278-003-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving lead", "text": "The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received."}, {"speaker": "Inspection lead", "text": "The inspection completed at 14:35 UTC on 12 August 2026 verified that 37 units received under PO-731 were unauthorized substitutions."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Operations coordinator", "text": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all choice criteria and instructions, while both contexts retain the original routing priority and substitution threshold without alteration. Both remain bound to PO-731 and the same decision path; the two evidence spans are complete factual sentences, and changing substitutions from 37 to 52 against 430 ordered units creates a coherent threshold-crossing counterfactual without conflicting measurements or embedded answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving lead\",\"text\":\"The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received.\"},{\"speaker\":\"Inspection lead\",\"text\":\"The inspection completed at 14:35 UTC on 12 August 2026 verified that 37 units received under PO-731 were unauthorized substitutions.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received."}, {"path": ["1", "text"], "text": "The inspection completed at 14:35 UTC on 12 August 2026 verified that 37 units received under PO-731 were unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received.", "negative_left": "The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received.", "negative_right": "The inspection completed at 14:35 UTC on 12 August 2026 verified that 52 units received under PO-731 were unauthorized substitutions.", "right": "The inspection completed at 14:35 UTC on 12 August 2026 verified that 37 units received under PO-731 were unauthorized substitutions."}, "verifier_independent_model": false}, "family": "scale-diverse-278-003", "id": "scale-diverse-278-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving lead", "text": "The receiving handoff completed at 14:20 UTC on 12 August 2026 for PO-731 records 430 units ordered and zero physically damaged units among those received."}, {"speaker": "Inspection lead", "text": "The inspection completed at 14:35 UTC on 12 August 2026 verified that 52 units received under PO-731 were unauthorized substitutions."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Operations coordinator", "text": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all question-level instructions and criteria, while both contexts retain the original routing priority and substitution thresholds verbatim. The same PO-731 shipment and decision scope are maintained; the two evidence spans are complete factual sentences, and the counterfactual coherently changes only the affected substitution count from 19 to 31 against an unchanged total of 260 without creating duplicate or contradictory measurements. Neither context supplies a selected choice, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving record\",\"text\":\"PO-731 ordered 260 units in total.\"},{\"speaker\":\"Discrepancy record\",\"text\":\"The unauthorized substitution under PO-731 affected 19 units.\"},{\"speaker\":\"Receiving inspector\",\"text\":\"Inspection covered every delivered unit and its packaging. The signed report records no crushing, cracking, moisture, deformation, or other physical damage. All received items were intact when the inspection ended.\"},{\"speaker\":\"Procurement specialist\",\"text\":\"The identity check confirms that Q-9 seals were delivered where PO-731 required P-9 seals. The approval register contains no authorized substitute request or approval for this order.\"},{\"speaker\":\"Records auditor\",\"text\":\"The purchase order, receiving log, discrepancy record, inspection report, and approval register all refer to the same shipment. Their identifiers and signatures were checked against the dock file.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"The receiving reconciliation is complete. No shortage, overage, documentation ambiguity, or additional discrepancy was recorded for the shipment.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 260 units in total."}, {"path": ["1", "text"], "text": "The unauthorized substitution under PO-731 affected 19 units."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 260 units in total.", "negative_left": "PO-731 ordered 260 units in total.", "negative_right": "The unauthorized substitution under PO-731 affected 31 units.", "right": "The unauthorized substitution under PO-731 affected 19 units."}, "verifier_independent_model": false}, "family": "scale-diverse-278-004", "id": "scale-diverse-278-004-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving record", "text": "PO-731 ordered 260 units in total."}, {"speaker": "Discrepancy record", "text": "The unauthorized substitution under PO-731 affected 19 units."}, {"speaker": "Receiving inspector", "text": "Inspection covered every delivered unit and its packaging. The signed report records no crushing, cracking, moisture, deformation, or other physical damage. All received items were intact when the inspection ended."}, {"speaker": "Procurement specialist", "text": "The identity check confirms that Q-9 seals were delivered where PO-731 required P-9 seals. The approval register contains no authorized substitute request or approval for this order."}, {"speaker": "Records auditor", "text": "The purchase order, receiving log, discrepancy record, inspection report, and approval register all refer to the same shipment. Their identifiers and signatures were checked against the dock file."}, {"speaker": "Supplier claims coordinator", "text": "The receiving reconciliation is complete. No shortage, overage, documentation ambiguity, or additional discrepancy was recorded for the shipment."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all question-level instructions and criteria, while both contexts retain the original routing priority and substitution thresholds verbatim. The same PO-731 shipment and decision scope are maintained; the two evidence spans are complete factual sentences, and the counterfactual coherently changes only the affected substitution count from 19 to 31 against an unchanged total of 260 without creating duplicate or contradictory measurements. Neither context supplies a selected choice, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving record\",\"text\":\"PO-731 ordered 260 units in total.\"},{\"speaker\":\"Discrepancy record\",\"text\":\"The unauthorized substitution under PO-731 affected 19 units.\"},{\"speaker\":\"Receiving inspector\",\"text\":\"Inspection covered every delivered unit and its packaging. The signed report records no crushing, cracking, moisture, deformation, or other physical damage. All received items were intact when the inspection ended.\"},{\"speaker\":\"Procurement specialist\",\"text\":\"The identity check confirms that Q-9 seals were delivered where PO-731 required P-9 seals. The approval register contains no authorized substitute request or approval for this order.\"},{\"speaker\":\"Records auditor\",\"text\":\"The purchase order, receiving log, discrepancy record, inspection report, and approval register all refer to the same shipment. Their identifiers and signatures were checked against the dock file.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"The receiving reconciliation is complete. No shortage, overage, documentation ambiguity, or additional discrepancy was recorded for the shipment.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 260 units in total."}, {"path": ["1", "text"], "text": "The unauthorized substitution under PO-731 affected 19 units."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 260 units in total.", "negative_left": "PO-731 ordered 260 units in total.", "negative_right": "The unauthorized substitution under PO-731 affected 31 units.", "right": "The unauthorized substitution under PO-731 affected 19 units."}, "verifier_independent_model": false}, "family": "scale-diverse-278-004", "id": "scale-diverse-278-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving record", "text": "PO-731 ordered 260 units in total."}, {"speaker": "Discrepancy record", "text": "The unauthorized substitution under PO-731 affected 31 units."}, {"speaker": "Receiving inspector", "text": "Inspection covered every delivered unit and its packaging. The signed report records no crushing, cracking, moisture, deformation, or other physical damage. All received items were intact when the inspection ended."}, {"speaker": "Procurement specialist", "text": "The identity check confirms that Q-9 seals were delivered where PO-731 required P-9 seals. The approval register contains no authorized substitute request or approval for this order."}, {"speaker": "Records auditor", "text": "The purchase order, receiving log, discrepancy record, inspection report, and approval register all refer to the same shipment. Their identifiers and signatures were checked against the dock file."}, {"speaker": "Supplier claims coordinator", "text": "The receiving reconciliation is complete. No shortage, overage, documentation ambiguity, or additional discrepancy was recorded for the shipment."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy and identical request concerning the shipment associated with PO 7714 and ASN 7714-A. The two focus-evidence spans are complete factual sentences. The counterfactual changes the PO SKU to NQ-4826, making it match the received SKU without creating a duplicate contradictory measurement; the remaining reference to alternate goods is still coherent because the ASN specifies a different SKU. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\",\"evidence\":[\"At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-7319 for the goods in that order.\",\"At 11:40 UTC on 7 April 2026, the issued ASN 7714-A specified SKU NQ-6504 for the shipment associated with PO 7714.\",\"At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826.\",\"At 14:50 UTC, the receiving inspection recorded that the alternate goods were undamaged.\",\"At 15:20 UTC, the records custodian completed a search and confirmed that no written authorization exists for substituting the received goods.\"],\"request\":\"Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826."}, {"path": ["evidence", "0"], "text": "At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-7319 for the goods in that order."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826.", "negative_left": "At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826.", "negative_right": "At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-4826 for the goods in that order.", "right": "At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-7319 for the goods in that order."}, "verifier_independent_model": false}, "family": "scale-diverse-280-001", "id": "scale-diverse-280-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-7319 for the goods in that order.", "At 11:40 UTC on 7 April 2026, the issued ASN 7714-A specified SKU NQ-6504 for the shipment associated with PO 7714.", "At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826.", "At 14:50 UTC, the receiving inspection recorded that the alternate goods were undamaged.", "At 15:20 UTC, the records custodian completed a search and confirmed that no written authorization exists for substituting the received goods."], "request": "Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?"}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy and identical request concerning the shipment associated with PO 7714 and ASN 7714-A. The two focus-evidence spans are complete factual sentences. The counterfactual changes the PO SKU to NQ-4826, making it match the received SKU without creating a duplicate contradictory measurement; the remaining reference to alternate goods is still coherent because the ASN specifies a different SKU. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\",\"evidence\":[\"At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-7319 for the goods in that order.\",\"At 11:40 UTC on 7 April 2026, the issued ASN 7714-A specified SKU NQ-6504 for the shipment associated with PO 7714.\",\"At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826.\",\"At 14:50 UTC, the receiving inspection recorded that the alternate goods were undamaged.\",\"At 15:20 UTC, the records custodian completed a search and confirmed that no written authorization exists for substituting the received goods.\"],\"request\":\"Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826."}, {"path": ["evidence", "0"], "text": "At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-7319 for the goods in that order."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826.", "negative_left": "At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826.", "negative_right": "At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-4826 for the goods in that order.", "right": "At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-7319 for the goods in that order."}, "verifier_independent_model": false}, "family": "scale-diverse-280-001", "id": "scale-diverse-280-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["At 09:15 UTC on 6 April 2026, the issued version of PO 7714 specified SKU NQ-4826 for the goods in that order.", "At 11:40 UTC on 7 April 2026, the issued ASN 7714-A specified SKU NQ-6504 for the shipment associated with PO 7714.", "At 14:32 UTC on 8 April 2026, North Quay Warehouse's receiving record identified the goods in the shipment associated with PO 7714 and ASN 7714-A as SKU NQ-4826.", "At 14:50 UTC, the receiving inspection recorded that the alternate goods were undamaged.", "At 15:20 UTC, the records custodian completed a search and confirmed that no written authorization exists for substituting the received goods."], "request": "Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?"}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same routing policy, request, warehouse, shipment identifiers, and routing target. Each evidence span is a complete factual sentence. The counterfactual changes only the received SKU from RX-842 to GT-506; this is coherent with the unchanged purchase-order record and creates no duplicate contradictory measurement or assertion. Neither context includes a gold answer, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\",\"evidence\":[\"For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU RX-842, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization.\",\"The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714.\"],\"request\":\"Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU RX-842, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization."}, {"path": ["evidence", "1"], "text": "The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU RX-842, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization.", "negative_left": "For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU GT-506, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization.", "negative_right": "The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714.", "right": "The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714."}, "verifier_independent_model": false}, "family": "scale-diverse-280-002", "id": "scale-diverse-280-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU RX-842, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization.", "The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714."], "request": "Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?"}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the same routing policy, request, warehouse, shipment identifiers, and routing target. Each evidence span is a complete factual sentence. The counterfactual changes only the received SKU from RX-842 to GT-506; this is coherent with the unchanged purchase-order record and creates no duplicate contradictory measurement or assertion. Neither context includes a gold answer, answer code, proposition ID, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\",\"evidence\":[\"For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU RX-842, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization.\",\"The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714.\"],\"request\":\"Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU RX-842, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization."}, {"path": ["evidence", "1"], "text": "The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU RX-842, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization.", "negative_left": "For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU GT-506, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization.", "negative_right": "The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714.", "right": "The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714."}, "verifier_independent_model": false}, "family": "scale-diverse-280-002", "id": "scale-diverse-280-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["For the shipment associated with PO 7714 and ASN 7714-A, the receiving inspection completed at North Quay Warehouse at 10:25 UTC on 14 August 2026 recorded received SKU GT-506, ASN-listed SKU VK-319, undamaged goods, and no written substitution authorization.", "The signed purchase-order record in effect for the shipment associated with PO 7714 and ASN 7714-A at 10:25 UTC on 14 August 2026 specifies SKU GT-506 for PO 7714."], "request": "Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?"}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, request, shipment identifiers (PO 7714 and ASN 7714-A), and relevant timing, while the unchanged questions object preserves the decision instructions and criteria. Each context contains exactly two complete factual evidence sentences. The counterfactual changes only the received SKU from VT-3917 to NX-5842; this coherently makes it match the purchase order while leaving the distinct ASN SKU and other facts unchanged, with no contradictory duplicate measurement or assertion. Neither context includes a gold answer, answer code, proposition ID, rule table, label rationale, or output instruction; the routing language is permissible governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\",\"evidence\":[\"At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A.\",\"At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU VT-3917, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed.\"],\"request\":\"Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A."}, {"path": ["evidence", "1"], "text": "At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU VT-3917, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A.", "negative_left": "At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A.", "negative_right": "At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU NX-5842, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed.", "right": "At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU VT-3917, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed."}, "verifier_independent_model": false}, "family": "scale-diverse-280-004", "id": "scale-diverse-280-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A.", "At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU VT-3917, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed."], "request": "Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?"}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing policy, request, shipment identifiers (PO 7714 and ASN 7714-A), and relevant timing, while the unchanged questions object preserves the decision instructions and criteria. Each context contains exactly two complete factual evidence sentences. The counterfactual changes only the received SKU from VT-3917 to NX-5842; this coherently makes it match the purchase order while leaving the distinct ASN SKU and other facts unchanged, with no contradictory duplicate measurement or assertion. Neither context includes a gold answer, answer code, proposition ID, rule table, label rationale, or output instruction; the routing language is permissible governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\",\"evidence\":[\"At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A.\",\"At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU VT-3917, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed.\"],\"request\":\"Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A."}, {"path": ["evidence", "1"], "text": "At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU VT-3917, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A.", "negative_left": "At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A.", "negative_right": "At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU NX-5842, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed.", "right": "At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU VT-3917, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed."}, "verifier_independent_model": false}, "family": "scale-diverse-280-004", "id": "scale-diverse-280-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["At 09:40 UTC on 12 August 2026, the purchase-order record for PO 7714 specified SKU NX-5842 for the shipment identified by ASN 7714-A.", "At 09:55 UTC on 12 August 2026, the completed receiving record for the shipment associated with PO 7714 and ASN 7714-A listed the received, undamaged goods as SKU NX-5842, listed ASN 7714-A as specifying SKU CR-2206, and confirmed that no written substitution authorization existed."], "request": "Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?"}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the documented-buyer-approval policy without adding exceptions or defaults, and they preserve the PO-7714 shipment decision scope and the relevant evidence paths. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the ordered revision from Q-41 to Q-38, making the PO, ASN, and carton revision assertions mutually consistent; the unchanged absence of substitution approval remains coherent because no substitution is asserted there. Neither context contains a discrepancy-level answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express factual relationships rather than outcome classifications; broad absence atom a9 is still a single factual claim about whether any other follow-up discrepancy exists. The focus a1 is a factual revision-comparison relation. Base and counter assignments can coexist with the same non-focus facts: a differing but uniformly labeled revision can pass inspection, while the counter can use the ordered revision without changing the other facts. The sole state-origin policy rule needed for interpretation—documented buyer approval for a substituted revision—is preserved in policy_evidence; rules already contained in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish evidence of a wrong-revision substitution without documented buyer approval, usable quantity, and no immediate interruption. They also explicitly refute every Level 3 trigger, so Level 2 is the highest supported level.", "rule_index": 0, "sound": true}, {"reason": "Together, the conditions establish agreement among the PO, ASN, labels, physical count, and inspection, while a9 excludes any other discrepancy requiring follow-up. The revision difference is refuted and the Level 3 triggers are refuted, making Level 0 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For shipment PO-7714, the single revision uniformly displayed on all carton labels differs from the revision ordered by purchase order PO-7714."}, {"id": "a2", "statement": "All carton labels for shipment PO-7714 display one uniform revision."}, {"id": "a3", "statement": "ASN-7714 and purchase order PO-7714 specify the same quantity."}, {"id": "a4", "statement": "ASN-7714 and purchase order PO-7714 specify the same item identity."}, {"id": "a5", "statement": "ASN-7714 and purchase order PO-7714 specify the same revision."}, {"id": "a6", "statement": "The physical count for shipment PO-7714 equals the quantity ordered by purchase order PO-7714."}, {"id": "a7", "statement": "The item identity on every carton label for shipment PO-7714 matches the item identity ordered by purchase order PO-7714."}, {"id": "a8", "statement": "The inspection result for shipment PO-7714 conforms to the inspection requirements for the ordered goods."}, {"id": "a9", "statement": "Shipment PO-7714 has zero discrepancies requiring follow-up apart from any difference between its uniform carton-label revision and its ordered purchase-order revision."}, {"id": "a10", "statement": "The quantity delivered in shipment PO-7714 is usable."}, {"id": "a11", "statement": "Shipment PO-7714 is expected to cause an immediate operational interruption."}, {"id": "a12", "statement": "The receiving packet for shipment PO-7714 contains documented buyer approval for the carton-label revision as a substitution for the ordered revision."}, {"id": "a13", "statement": "Shipment PO-7714 has a safety-critical mismatch."}, {"id": "a14", "statement": "Shipment PO-7714 has widespread damage."}, {"id": "a15", "statement": "Shipment PO-7714 has a material shortage."}, {"id": "a16", "statement": "Shipment PO-7714 contains wrong goods likely to stop operations."}], "base_state_json": "[{\"speaker\":\"Inspection log\",\"text\":\"At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision.\"},{\"speaker\":\"Purchasing record\",\"text\":\"When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-41 for the ordered goods.\"},{\"speaker\":\"Receiving review\",\"text\":\"ASN-7714 matched PO-7714 in quantity, item identity, and revision. The physical count equaled the ordered quantity, and the item identity on every carton label matched the order. The delivered quantity was usable. Inspection results conformed to the requirements for the ordered goods, with no safety-critical mismatch or widespread damage.\"},{\"speaker\":\"Operations review\",\"text\":\"There was no material shortage, no wrong goods likely to stop operations, and no immediate operational interruption expected. No discrepancy required follow-up apart from any variance found in the carton-revision comparison. The receiving packet contained no documented buyer approval for a revision substitution.\"},{\"speaker\":\"Policy excerpt\",\"text\":\"our policy requires documented buyer approval before any substituted revision can enter available inventory.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision."}, {"path": ["1", "text"], "text": "When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-41 for the ordered goods."}], "policy_evidence": [{"path": ["3", "text"], "text": "our policy requires documented buyer approval before any substituted revision can enter available inventory."}], "rules": [{"justification": "A carton-label revision different from the ordered revision, without the required documented buyer approval, is an undocumented wrong-revision substitution. Because the quantity is usable, no immediate operational interruption exists, and every Level 3 trigger is explicitly absent, Level 2 is the highest supported level.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}]}, {"justification": "The uniform carton-label revision does not differ from the ordered revision; the PO, ASN, labels, count, and inspection agree; no other discrepancy requires follow-up; and all higher-severity triggers are absent. Level 0 therefore applies.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision.", "negative_left": "At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision.", "negative_right": "When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-38 for the ordered goods.", "right": "When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-41 for the ordered goods."}, "verifier_independent_model": false}, "family": "scale-diverse-281-001", "id": "scale-diverse-281-001-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Receive normally: PO, ASN, labels, count, and inspection all agree, with no discrepancy requiring follow-up.", "Level 1 — Routine receiving correction: a minor clerical or carton-level variance is fully resolved by supplied documents, with item identity and authorization confirmed; receive stock and log the correction.", "Level 2 — Quarantine and route to Inventory Control: quantity is usable and no immediate operational interruption exists, but evidence shows a wrong revision, wrong item, or undocumented substitution lacking required buyer approval; block inventory pending verification.", "Level 3 — Escalate to Dock Supervisor and Supplier Claims: evidence shows safety-critical mismatch, widespread damage, material shortage, or wrong goods likely to stop operations; reject or hold the load and open a supplier claim immediately."], "instructions": "Assign the shipment to exactly one ordered discrepancy level using only the explicit evidence. Apply the highest level whose conditions are supported. This is a routing-and-severity decision; an undocumented revision substitution is not cleared merely by passing inspection or by a supplier's equivalence claim.", "type": "score"}}, "state": [{"speaker": "Inspection log", "text": "At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision."}, {"speaker": "Purchasing record", "text": "When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-41 for the ordered goods."}, {"speaker": "Receiving review", "text": "ASN-7714 matched PO-7714 in quantity, item identity, and revision. The physical count equaled the ordered quantity, and the item identity on every carton label matched the order. The delivered quantity was usable. Inspection results conformed to the requirements for the ordered goods, with no safety-critical mismatch or widespread damage."}, {"speaker": "Operations review", "text": "There was no material shortage, no wrong goods likely to stop operations, and no immediate operational interruption expected. No discrepancy required follow-up apart from any variance found in the carton-revision comparison. The receiving packet contained no documented buyer approval for a revision substitution."}, {"speaker": "Policy excerpt", "text": "our policy requires documented buyer approval before any substituted revision can enter available inventory."}]}, "method": "c2d", "provenance": {"source_id": "diverse-281", "source_is_synthetic": true, "source_sha256": "9852c08cb6247bcdb48316d4bd130d10d68e029336afd24cacd7f50bb05c452c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the documented-buyer-approval policy without adding exceptions or defaults, and they preserve the PO-7714 shipment decision scope and the relevant evidence paths. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the ordered revision from Q-41 to Q-38, making the PO, ASN, and carton revision assertions mutually consistent; the unchanged absence of substitution approval remains coherent because no substitution is asserted there. Neither context contains a discrepancy-level answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express factual relationships rather than outcome classifications; broad absence atom a9 is still a single factual claim about whether any other follow-up discrepancy exists. The focus a1 is a factual revision-comparison relation. Base and counter assignments can coexist with the same non-focus facts: a differing but uniformly labeled revision can pass inspection, while the counter can use the ordered revision without changing the other facts. The sole state-origin policy rule needed for interpretation—documented buyer approval for a substituted revision—is preserved in policy_evidence; rules already contained in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish evidence of a wrong-revision substitution without documented buyer approval, usable quantity, and no immediate interruption. They also explicitly refute every Level 3 trigger, so Level 2 is the highest supported level.", "rule_index": 0, "sound": true}, {"reason": "Together, the conditions establish agreement among the PO, ASN, labels, physical count, and inspection, while a9 excludes any other discrepancy requiring follow-up. The revision difference is refuted and the Level 3 triggers are refuted, making Level 0 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For shipment PO-7714, the single revision uniformly displayed on all carton labels differs from the revision ordered by purchase order PO-7714."}, {"id": "a2", "statement": "All carton labels for shipment PO-7714 display one uniform revision."}, {"id": "a3", "statement": "ASN-7714 and purchase order PO-7714 specify the same quantity."}, {"id": "a4", "statement": "ASN-7714 and purchase order PO-7714 specify the same item identity."}, {"id": "a5", "statement": "ASN-7714 and purchase order PO-7714 specify the same revision."}, {"id": "a6", "statement": "The physical count for shipment PO-7714 equals the quantity ordered by purchase order PO-7714."}, {"id": "a7", "statement": "The item identity on every carton label for shipment PO-7714 matches the item identity ordered by purchase order PO-7714."}, {"id": "a8", "statement": "The inspection result for shipment PO-7714 conforms to the inspection requirements for the ordered goods."}, {"id": "a9", "statement": "Shipment PO-7714 has zero discrepancies requiring follow-up apart from any difference between its uniform carton-label revision and its ordered purchase-order revision."}, {"id": "a10", "statement": "The quantity delivered in shipment PO-7714 is usable."}, {"id": "a11", "statement": "Shipment PO-7714 is expected to cause an immediate operational interruption."}, {"id": "a12", "statement": "The receiving packet for shipment PO-7714 contains documented buyer approval for the carton-label revision as a substitution for the ordered revision."}, {"id": "a13", "statement": "Shipment PO-7714 has a safety-critical mismatch."}, {"id": "a14", "statement": "Shipment PO-7714 has widespread damage."}, {"id": "a15", "statement": "Shipment PO-7714 has a material shortage."}, {"id": "a16", "statement": "Shipment PO-7714 contains wrong goods likely to stop operations."}], "base_state_json": "[{\"speaker\":\"Inspection log\",\"text\":\"At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision.\"},{\"speaker\":\"Purchasing record\",\"text\":\"When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-41 for the ordered goods.\"},{\"speaker\":\"Receiving review\",\"text\":\"ASN-7714 matched PO-7714 in quantity, item identity, and revision. The physical count equaled the ordered quantity, and the item identity on every carton label matched the order. The delivered quantity was usable. Inspection results conformed to the requirements for the ordered goods, with no safety-critical mismatch or widespread damage.\"},{\"speaker\":\"Operations review\",\"text\":\"There was no material shortage, no wrong goods likely to stop operations, and no immediate operational interruption expected. No discrepancy required follow-up apart from any variance found in the carton-revision comparison. The receiving packet contained no documented buyer approval for a revision substitution.\"},{\"speaker\":\"Policy excerpt\",\"text\":\"our policy requires documented buyer approval before any substituted revision can enter available inventory.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision."}, {"path": ["1", "text"], "text": "When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-41 for the ordered goods."}], "policy_evidence": [{"path": ["3", "text"], "text": "our policy requires documented buyer approval before any substituted revision can enter available inventory."}], "rules": [{"justification": "A carton-label revision different from the ordered revision, without the required documented buyer approval, is an undocumented wrong-revision substitution. Because the quantity is usable, no immediate operational interruption exists, and every Level 3 trigger is explicitly absent, Level 2 is the highest supported level.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}]}, {"justification": "The uniform carton-label revision does not differ from the ordered revision; the PO, ASN, labels, count, and inspection agree; no other discrepancy requires follow-up; and all higher-severity triggers are absent. Level 0 therefore applies.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision.", "negative_left": "At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision.", "negative_right": "When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-38 for the ordered goods.", "right": "When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-41 for the ordered goods."}, "verifier_independent_model": false}, "family": "scale-diverse-281-001", "id": "scale-diverse-281-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Receive normally: PO, ASN, labels, count, and inspection all agree, with no discrepancy requiring follow-up.", "Level 1 — Routine receiving correction: a minor clerical or carton-level variance is fully resolved by supplied documents, with item identity and authorization confirmed; receive stock and log the correction.", "Level 2 — Quarantine and route to Inventory Control: quantity is usable and no immediate operational interruption exists, but evidence shows a wrong revision, wrong item, or undocumented substitution lacking required buyer approval; block inventory pending verification.", "Level 3 — Escalate to Dock Supervisor and Supplier Claims: evidence shows safety-critical mismatch, widespread damage, material shortage, or wrong goods likely to stop operations; reject or hold the load and open a supplier claim immediately."], "instructions": "Assign the shipment to exactly one ordered discrepancy level using only the explicit evidence. Apply the highest level whose conditions are supported. This is a routing-and-severity decision; an undocumented revision substitution is not cleared merely by passing inspection or by a supplier's equivalence claim.", "type": "score"}}, "state": [{"speaker": "Inspection log", "text": "At 09:14 UTC on 16 September 2026, inspectors recorded that every carton label in shipment PO-7714 displayed revision Q-38 and no other revision."}, {"speaker": "Purchasing record", "text": "When purchase order PO-7714 was issued at 15:40 UTC on 8 September 2026, it listed revision Q-38 for the ordered goods."}, {"speaker": "Receiving review", "text": "ASN-7714 matched PO-7714 in quantity, item identity, and revision. The physical count equaled the ordered quantity, and the item identity on every carton label matched the order. The delivered quantity was usable. Inspection results conformed to the requirements for the ordered goods, with no safety-critical mismatch or widespread damage."}, {"speaker": "Operations review", "text": "There was no material shortage, no wrong goods likely to stop operations, and no immediate operational interruption expected. No discrepancy required follow-up apart from any variance found in the carton-revision comparison. The receiving packet contained no documented buyer approval for a revision substitution."}, {"speaker": "Policy excerpt", "text": "our policy requires documented buyer approval before any substituted revision can enter available inventory."}]}, "method": "c2d", "provenance": {"source_id": "diverse-281", "source_is_synthetic": true, "source_sha256": "9852c08cb6247bcdb48316d4bd130d10d68e029336afd24cacd7f50bb05c452c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the buyer-approval policy and the same shipment-level decision scope and PO-7714 binding; only case observations change. The two focus-evidence spans are complete factual sentences. Changing the ordered revision from R6 to R7 makes the counterfactual consistent with the uniformly R7 labels and does not conflict with the ASN, count, inspection, exception, or operations statements. Neither context contains a gold answer, level code, rule table, proposition identifier, output instruction, or impermissible label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express factual relationships rather than outcome classifications; broad absence atom a9 is still a single factual claim about whether any other follow-up discrepancy exists. The focus a1 is a factual revision-comparison relation. Base and counter assignments can coexist with the same non-focus facts: a differing but uniformly labeled revision can pass inspection, while the counter can use the ordered revision without changing the other facts. The sole state-origin policy rule needed for interpretation—documented buyer approval for a substituted revision—is preserved in policy_evidence; rules already contained in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish evidence of a wrong-revision substitution without documented buyer approval, usable quantity, and no immediate interruption. They also explicitly refute every Level 3 trigger, so Level 2 is the highest supported level.", "rule_index": 0, "sound": true}, {"reason": "Together, the conditions establish agreement among the PO, ASN, labels, physical count, and inspection, while a9 excludes any other discrepancy requiring follow-up. The revision difference is refuted and the Level 3 triggers are refuted, making Level 0 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For shipment PO-7714, the single revision uniformly displayed on all carton labels differs from the revision ordered by purchase order PO-7714."}, {"id": "a2", "statement": "All carton labels for shipment PO-7714 display one uniform revision."}, {"id": "a3", "statement": "ASN-7714 and purchase order PO-7714 specify the same quantity."}, {"id": "a4", "statement": "ASN-7714 and purchase order PO-7714 specify the same item identity."}, {"id": "a5", "statement": "ASN-7714 and purchase order PO-7714 specify the same revision."}, {"id": "a6", "statement": "The physical count for shipment PO-7714 equals the quantity ordered by purchase order PO-7714."}, {"id": "a7", "statement": "The item identity on every carton label for shipment PO-7714 matches the item identity ordered by purchase order PO-7714."}, {"id": "a8", "statement": "The inspection result for shipment PO-7714 conforms to the inspection requirements for the ordered goods."}, {"id": "a9", "statement": "Shipment PO-7714 has zero discrepancies requiring follow-up apart from any difference between its uniform carton-label revision and its ordered purchase-order revision."}, {"id": "a10", "statement": "The quantity delivered in shipment PO-7714 is usable."}, {"id": "a11", "statement": "Shipment PO-7714 is expected to cause an immediate operational interruption."}, {"id": "a12", "statement": "The receiving packet for shipment PO-7714 contains documented buyer approval for the carton-label revision as a substitution for the ordered revision."}, {"id": "a13", "statement": "Shipment PO-7714 has a safety-critical mismatch."}, {"id": "a14", "statement": "Shipment PO-7714 has widespread damage."}, {"id": "a15", "statement": "Shipment PO-7714 has a material shortage."}, {"id": "a16", "statement": "Shipment PO-7714 contains wrong goods likely to stop operations."}], "base_state_json": "[{\"speaker\":\"Receiving handoff\",\"text\":\"At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels.\"},{\"speaker\":\"Order record\",\"text\":\"The signed line item in purchase order PO-7714 ordered revision R6 for shipment PO-7714.\"},{\"speaker\":\"Document check\",\"text\":\"ASN-7714 matches purchase order PO-7714 in quantity, item identity, and revision. Every carton label carries the ordered item identity.\"},{\"speaker\":\"Count and inspection\",\"text\":\"The physical count equals the purchase-order quantity, and the delivered quantity is usable. Inspection conforms to the requirements for the ordered goods.\"},{\"speaker\":\"Exception review\",\"text\":\"Apart from any difference between the uniform carton-label revision and the ordered revision, the shipment has zero discrepancies requiring follow-up. The receiving packet contains no documented buyer approval for a substitution.\"},{\"speaker\":\"Operations review\",\"text\":\"No immediate operational interruption is expected. Review found no safety-critical mismatch, widespread damage, material shortage, or wrong goods likely to stop operations.\"},{\"speaker\":\"Policy note\",\"text\":\"our policy requires documented buyer approval before any substituted revision can enter available inventory.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels."}, {"path": ["1", "text"], "text": "The signed line item in purchase order PO-7714 ordered revision R6 for shipment PO-7714."}], "policy_evidence": [{"path": ["3", "text"], "text": "our policy requires documented buyer approval before any substituted revision can enter available inventory."}], "rules": [{"justification": "A carton-label revision different from the ordered revision, without the required documented buyer approval, is an undocumented wrong-revision substitution. Because the quantity is usable, no immediate operational interruption exists, and every Level 3 trigger is explicitly absent, Level 2 is the highest supported level.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}]}, {"justification": "The uniform carton-label revision does not differ from the ordered revision; the PO, ASN, labels, count, and inspection agree; no other discrepancy requires follow-up; and all higher-severity triggers are absent. Level 0 therefore applies.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels.", "negative_left": "At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels.", "negative_right": "The signed line item in purchase order PO-7714 ordered revision R7 for shipment PO-7714.", "right": "The signed line item in purchase order PO-7714 ordered revision R6 for shipment PO-7714."}, "verifier_independent_model": false}, "family": "scale-diverse-281-003", "id": "scale-diverse-281-003-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Receive normally: PO, ASN, labels, count, and inspection all agree, with no discrepancy requiring follow-up.", "Level 1 — Routine receiving correction: a minor clerical or carton-level variance is fully resolved by supplied documents, with item identity and authorization confirmed; receive stock and log the correction.", "Level 2 — Quarantine and route to Inventory Control: quantity is usable and no immediate operational interruption exists, but evidence shows a wrong revision, wrong item, or undocumented substitution lacking required buyer approval; block inventory pending verification.", "Level 3 — Escalate to Dock Supervisor and Supplier Claims: evidence shows safety-critical mismatch, widespread damage, material shortage, or wrong goods likely to stop operations; reject or hold the load and open a supplier claim immediately."], "instructions": "Assign the shipment to exactly one ordered discrepancy level using only the explicit evidence. Apply the highest level whose conditions are supported. This is a routing-and-severity decision; an undocumented revision substitution is not cleared merely by passing inspection or by a supplier's equivalence claim.", "type": "score"}}, "state": [{"speaker": "Receiving handoff", "text": "At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels."}, {"speaker": "Order record", "text": "The signed line item in purchase order PO-7714 ordered revision R6 for shipment PO-7714."}, {"speaker": "Document check", "text": "ASN-7714 matches purchase order PO-7714 in quantity, item identity, and revision. Every carton label carries the ordered item identity."}, {"speaker": "Count and inspection", "text": "The physical count equals the purchase-order quantity, and the delivered quantity is usable. Inspection conforms to the requirements for the ordered goods."}, {"speaker": "Exception review", "text": "Apart from any difference between the uniform carton-label revision and the ordered revision, the shipment has zero discrepancies requiring follow-up. The receiving packet contains no documented buyer approval for a substitution."}, {"speaker": "Operations review", "text": "No immediate operational interruption is expected. Review found no safety-critical mismatch, widespread damage, material shortage, or wrong goods likely to stop operations."}, {"speaker": "Policy note", "text": "our policy requires documented buyer approval before any substituted revision can enter available inventory."}]}, "method": "c2d", "provenance": {"source_id": "diverse-281", "source_is_synthetic": true, "source_sha256": "9852c08cb6247bcdb48316d4bd130d10d68e029336afd24cacd7f50bb05c452c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the buyer-approval policy and the same shipment-level decision scope and PO-7714 binding; only case observations change. The two focus-evidence spans are complete factual sentences. Changing the ordered revision from R6 to R7 makes the counterfactual consistent with the uniformly R7 labels and does not conflict with the ASN, count, inspection, exception, or operations statements. Neither context contains a gold answer, level code, rule table, proposition identifier, output instruction, or impermissible label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "refuted", "a12": "refuted", "a13": "refuted", "a14": "refuted", "a15": "refuted", "a16": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express factual relationships rather than outcome classifications; broad absence atom a9 is still a single factual claim about whether any other follow-up discrepancy exists. The focus a1 is a factual revision-comparison relation. Base and counter assignments can coexist with the same non-focus facts: a differing but uniformly labeled revision can pass inspection, while the counter can use the ordered revision without changing the other facts. The sole state-origin policy rule needed for interpretation—documented buyer approval for a substituted revision—is preserved in policy_evidence; rules already contained in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish evidence of a wrong-revision substitution without documented buyer approval, usable quantity, and no immediate interruption. They also explicitly refute every Level 3 trigger, so Level 2 is the highest supported level.", "rule_index": 0, "sound": true}, {"reason": "Together, the conditions establish agreement among the PO, ASN, labels, physical count, and inspection, while a9 excludes any other discrepancy requiring follow-up. The revision difference is refuted and the Level 3 triggers are refuted, making Level 0 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For shipment PO-7714, the single revision uniformly displayed on all carton labels differs from the revision ordered by purchase order PO-7714."}, {"id": "a2", "statement": "All carton labels for shipment PO-7714 display one uniform revision."}, {"id": "a3", "statement": "ASN-7714 and purchase order PO-7714 specify the same quantity."}, {"id": "a4", "statement": "ASN-7714 and purchase order PO-7714 specify the same item identity."}, {"id": "a5", "statement": "ASN-7714 and purchase order PO-7714 specify the same revision."}, {"id": "a6", "statement": "The physical count for shipment PO-7714 equals the quantity ordered by purchase order PO-7714."}, {"id": "a7", "statement": "The item identity on every carton label for shipment PO-7714 matches the item identity ordered by purchase order PO-7714."}, {"id": "a8", "statement": "The inspection result for shipment PO-7714 conforms to the inspection requirements for the ordered goods."}, {"id": "a9", "statement": "Shipment PO-7714 has zero discrepancies requiring follow-up apart from any difference between its uniform carton-label revision and its ordered purchase-order revision."}, {"id": "a10", "statement": "The quantity delivered in shipment PO-7714 is usable."}, {"id": "a11", "statement": "Shipment PO-7714 is expected to cause an immediate operational interruption."}, {"id": "a12", "statement": "The receiving packet for shipment PO-7714 contains documented buyer approval for the carton-label revision as a substitution for the ordered revision."}, {"id": "a13", "statement": "Shipment PO-7714 has a safety-critical mismatch."}, {"id": "a14", "statement": "Shipment PO-7714 has widespread damage."}, {"id": "a15", "statement": "Shipment PO-7714 has a material shortage."}, {"id": "a16", "statement": "Shipment PO-7714 contains wrong goods likely to stop operations."}], "base_state_json": "[{\"speaker\":\"Receiving handoff\",\"text\":\"At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels.\"},{\"speaker\":\"Order record\",\"text\":\"The signed line item in purchase order PO-7714 ordered revision R6 for shipment PO-7714.\"},{\"speaker\":\"Document check\",\"text\":\"ASN-7714 matches purchase order PO-7714 in quantity, item identity, and revision. Every carton label carries the ordered item identity.\"},{\"speaker\":\"Count and inspection\",\"text\":\"The physical count equals the purchase-order quantity, and the delivered quantity is usable. Inspection conforms to the requirements for the ordered goods.\"},{\"speaker\":\"Exception review\",\"text\":\"Apart from any difference between the uniform carton-label revision and the ordered revision, the shipment has zero discrepancies requiring follow-up. The receiving packet contains no documented buyer approval for a substitution.\"},{\"speaker\":\"Operations review\",\"text\":\"No immediate operational interruption is expected. Review found no safety-critical mismatch, widespread damage, material shortage, or wrong goods likely to stop operations.\"},{\"speaker\":\"Policy note\",\"text\":\"our policy requires documented buyer approval before any substituted revision can enter available inventory.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels."}, {"path": ["1", "text"], "text": "The signed line item in purchase order PO-7714 ordered revision R6 for shipment PO-7714."}], "policy_evidence": [{"path": ["3", "text"], "text": "our policy requires documented buyer approval before any substituted revision can enter available inventory."}], "rules": [{"justification": "A carton-label revision different from the ordered revision, without the required documented buyer approval, is an undocumented wrong-revision substitution. Because the quantity is usable, no immediate operational interruption exists, and every Level 3 trigger is explicitly absent, Level 2 is the highest supported level.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}]}, {"justification": "The uniform carton-label revision does not differ from the ordered revision; the PO, ASN, labels, count, and inspection agree; no other discrepancy requires follow-up; and all higher-severity triggers are absent. Level 0 therefore applies.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels.", "negative_left": "At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels.", "negative_right": "The signed line item in purchase order PO-7714 ordered revision R7 for shipment PO-7714.", "right": "The signed line item in purchase order PO-7714 ordered revision R6 for shipment PO-7714."}, "verifier_independent_model": false}, "family": "scale-diverse-281-003", "id": "scale-diverse-281-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Receive normally: PO, ASN, labels, count, and inspection all agree, with no discrepancy requiring follow-up.", "Level 1 — Routine receiving correction: a minor clerical or carton-level variance is fully resolved by supplied documents, with item identity and authorization confirmed; receive stock and log the correction.", "Level 2 — Quarantine and route to Inventory Control: quantity is usable and no immediate operational interruption exists, but evidence shows a wrong revision, wrong item, or undocumented substitution lacking required buyer approval; block inventory pending verification.", "Level 3 — Escalate to Dock Supervisor and Supplier Claims: evidence shows safety-critical mismatch, widespread damage, material shortage, or wrong goods likely to stop operations; reject or hold the load and open a supplier claim immediately."], "instructions": "Assign the shipment to exactly one ordered discrepancy level using only the explicit evidence. Apply the highest level whose conditions are supported. This is a routing-and-severity decision; an undocumented revision substitution is not cleared merely by passing inspection or by a supplier's equivalence claim.", "type": "score"}}, "state": [{"speaker": "Receiving handoff", "text": "At 14:20 UTC on 17 September 2026, receiving handoff RH-308 recorded that every carton label in shipment PO-7714 displayed revision R7 and that no other revision appeared on those labels."}, {"speaker": "Order record", "text": "The signed line item in purchase order PO-7714 ordered revision R7 for shipment PO-7714."}, {"speaker": "Document check", "text": "ASN-7714 matches purchase order PO-7714 in quantity, item identity, and revision. Every carton label carries the ordered item identity."}, {"speaker": "Count and inspection", "text": "The physical count equals the purchase-order quantity, and the delivered quantity is usable. Inspection conforms to the requirements for the ordered goods."}, {"speaker": "Exception review", "text": "Apart from any difference between the uniform carton-label revision and the ordered revision, the shipment has zero discrepancies requiring follow-up. The receiving packet contains no documented buyer approval for a substitution."}, {"speaker": "Operations review", "text": "No immediate operational interruption is expected. Review found no safety-critical mismatch, widespread damage, material shortage, or wrong goods likely to stop operations."}, {"speaker": "Policy note", "text": "our policy requires documented buyer approval before any substituted revision can enter available inventory."}]}, "method": "c2d", "provenance": {"source_id": "diverse-281", "source_is_synthetic": true, "source_sha256": "9852c08cb6247bcdb48316d4bd130d10d68e029336afd24cacd7f50bb05c452c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same receiving-routing policy and request scope through the unchanged original question, while preserving the bound supplier, location, PO, shipment event, and decision path. The changed SKU and quantity observations are permissible case-observation changes. The two evidence spans are complete factual sentences. The counterfactual changes only the exhaustive unit-scan result and is consistent with the unchanged count, condition, PO, and ASN facts. Neither context includes a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"The receiving clerk completed the count and condition review before beginning the unit-level scan, treating the records and physical shipment as one receiving event. At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU. At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded 63 units as SKU AX-317 and one unit as SKU BQ-908. The clerk then closed the inspection record and referred the receiving event for routing under the applicable criteria.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU."}, {"path": [], "text": "At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded 63 units as SKU AX-317 and one unit as SKU BQ-908."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU.", "negative_left": "At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU.", "negative_right": "At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded all 64 units as SKU AX-317.", "right": "At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded 63 units as SKU AX-317 and one unit as SKU BQ-908."}, "verifier_independent_model": false}, "family": "scale-diverse-282-001", "id": "scale-diverse-282-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "The receiving clerk completed the count and condition review before beginning the unit-level scan, treating the records and physical shipment as one receiving event. At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU. At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded 63 units as SKU AX-317 and one unit as SKU BQ-908. The clerk then closed the inspection record and referred the receiving event for routing under the applicable criteria."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same receiving-routing policy and request scope through the unchanged original question, while preserving the bound supplier, location, PO, shipment event, and decision path. The changed SKU and quantity observations are permissible case-observation changes. The two evidence spans are complete factual sentences. The counterfactual changes only the exhaustive unit-scan result and is consistent with the unchanged count, condition, PO, and ASN facts. Neither context includes a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"The receiving clerk completed the count and condition review before beginning the unit-level scan, treating the records and physical shipment as one receiving event. At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU. At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded 63 units as SKU AX-317 and one unit as SKU BQ-908. The clerk then closed the inspection record and referred the receiving event for routing under the applicable criteria.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU."}, {"path": [], "text": "At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded 63 units as SKU AX-317 and one unit as SKU BQ-908."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU.", "negative_left": "At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU.", "negative_right": "At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded all 64 units as SKU AX-317.", "right": "At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded 63 units as SKU AX-317 and one unit as SKU BQ-908."}, "verifier_independent_model": false}, "family": "scale-diverse-282-001", "id": "scale-diverse-282-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "The receiving clerk completed the count and condition review before beginning the unit-level scan, treating the records and physical shipment as one receiving event. At Door 6 on 14 September 2026, PO-1842 required 64 units of SKU AX-317 in intact condition, Northstar’s ASN-7731 listed 64 units of SKU AX-317 in intact condition and authorized no substitute, and the receiving record documented 64 intact physical units, no unresolved discrepancy unrelated to physical-unit SKU, and no physical receiving risk unrelated to physical-unit SKU. At 08:42 on 14 September 2026, an exhaustive scan of all 64 physical units in Northstar’s shipment at Door 6 recorded all 64 units as SKU AX-317. The clerk then closed the inspection record and referred the receiving event for routing under the applicable criteria."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original receiving policy without adding exceptions, priorities, or missing-evidence defaults, while preserving the shipment, supplier, door, PO, decision scope, and shared timestamp. The two evidence spans are complete factual sentences. The counterfactual changes only the PO-required SKU from NX-419 to NX-407, which remains coherent with the physical census, ASN match, quantity, condition, and discrepancy assertions. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407. The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-419 for every physical unit at 09:17 on 14 September 2026. The ASN’s listed SKU matches the signed PO line. The ASN explicitly states that no substitution is authorized. The quantity required by the PO and the quantity listed by the ASN each equal the complete physical census total. Both records specify factory-new, intact goods, and inspection found every unit and package in that condition. Reconciliation of the count sheet, ASN, PO, seal record, and condition report found no unresolved discrepancy independent of the physical-SKU comparison. Inspection also found no damage, unsafe unloading condition, recount need, containment need, or other independent physical receiving risk.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407."}, {"path": [], "text": "The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-419 for every physical unit at 09:17 on 14 September 2026."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407.", "negative_left": "At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407.", "negative_right": "The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-407 for every physical unit at 09:17 on 14 September 2026.", "right": "The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-419 for every physical unit at 09:17 on 14 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-282-002", "id": "scale-diverse-282-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407. The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-419 for every physical unit at 09:17 on 14 September 2026. The ASN’s listed SKU matches the signed PO line. The ASN explicitly states that no substitution is authorized. The quantity required by the PO and the quantity listed by the ASN each equal the complete physical census total. Both records specify factory-new, intact goods, and inspection found every unit and package in that condition. Reconciliation of the count sheet, ASN, PO, seal record, and condition report found no unresolved discrepancy independent of the physical-SKU comparison. Inspection also found no damage, unsafe unloading condition, recount need, containment need, or other independent physical receiving risk."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original receiving policy without adding exceptions, priorities, or missing-evidence defaults, while preserving the shipment, supplier, door, PO, decision scope, and shared timestamp. The two evidence spans are complete factual sentences. The counterfactual changes only the PO-required SKU from NX-419 to NX-407, which remains coherent with the physical census, ASN match, quantity, condition, and discrepancy assertions. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407. The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-419 for every physical unit at 09:17 on 14 September 2026. The ASN’s listed SKU matches the signed PO line. The ASN explicitly states that no substitution is authorized. The quantity required by the PO and the quantity listed by the ASN each equal the complete physical census total. Both records specify factory-new, intact goods, and inspection found every unit and package in that condition. Reconciliation of the count sheet, ASN, PO, seal record, and condition report found no unresolved discrepancy independent of the physical-SKU comparison. Inspection also found no damage, unsafe unloading condition, recount need, containment need, or other independent physical receiving risk.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407."}, {"path": [], "text": "The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-419 for every physical unit at 09:17 on 14 September 2026."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407.", "negative_left": "At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407.", "negative_right": "The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-407 for every physical unit at 09:17 on 14 September 2026.", "right": "The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-419 for every physical unit at 09:17 on 14 September 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-282-002", "id": "scale-diverse-282-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "At 09:17 on 14 September 2026, a complete physical census established that Northstar’s shipment at Door 6 consisted of exactly 31 units, each bearing SKU NX-407. The signed line item of PO-1842 applicable to Northstar’s shipment at Door 6 required SKU NX-407 for every physical unit at 09:17 on 14 September 2026. The ASN’s listed SKU matches the signed PO line. The ASN explicitly states that no substitution is authorized. The quantity required by the PO and the quantity listed by the ASN each equal the complete physical census total. Both records specify factory-new, intact goods, and inspection found every unit and package in that condition. Reconciliation of the count sheet, ASN, PO, seal record, and condition report found no unresolved discrepancy independent of the physical-SKU comparison. Inspection also found no damage, unsafe unloading condition, recount need, containment need, or other independent physical receiving risk."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all routing criteria and instructions, while both contexts retain the shipment, Northstar, Door 6, and PO-1842 bindings; the altered SKU and quantity are permissible case-observation changes. The two evidence spans are complete factual sentences, the counterfactual coherently changes the inspected SKU without creating contradictory measurements or assertions, and neither context contains an answer code, gold label, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"This signed operational handoff is the complete receiving record submitted for routing. At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-731, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842. For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-731, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842."}, {"path": [], "text": "For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-731, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842.", "negative_left": "At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-732, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842.", "negative_right": "For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731.", "right": "For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731."}, "verifier_independent_model": false}, "family": "scale-diverse-282-003", "id": "scale-diverse-282-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "This signed operational handoff is the complete receiving record submitted for routing. At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-731, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842. For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all routing criteria and instructions, while both contexts retain the shipment, Northstar, Door 6, and PO-1842 bindings; the altered SKU and quantity are permissible case-observation changes. The two evidence spans are complete factual sentences, the counterfactual coherently changes the inspected SKU without creating contradictory measurements or assertions, and neither context contains an answer code, gold label, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"This signed operational handoff is the complete receiving record submitted for routing. At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-731, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842. For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-731, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842."}, {"path": [], "text": "For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-731, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842.", "negative_left": "At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-732, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842.", "negative_right": "For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731.", "right": "For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731."}, "verifier_independent_model": false}, "family": "scale-diverse-282-003", "id": "scale-diverse-282-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "This signed operational handoff is the complete receiving record submitted for routing. At 14:20 UTC on 8 September 2026, the completed Door 6 handoff inspection recorded Northstar’s shipment as exactly 28 physical units, each marked SKU NS-732, with every unit and carton sealed, dry, and undamaged; it also recorded no unloading hazard, containment need, or unresolved receiving discrepancy independent of SKU conformity to PO-1842. For Northstar’s shipment inspected at Door 6 at 14:20 UTC on 8 September 2026, PO-1842 requires exactly 28 units of SKU NS-732 in sealed, dry, and undamaged condition, while the ASN lists the same SKU, quantity, and condition and contains no explicit authorization to substitute SKU NS-731."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing routing policy and criteria, while neither generated context removes or alters any additional governing policy from the original state. The generic decision request remains bound to the shipment evidence, and the changed shipment identifiers, quantities, and observations are permissible case-observation changes. The two evidence spans are complete factual sentences. The counterfactual coherently replaces the mixed-SKU count with 24 units of the specified SKU without conflicting with the unchanged quantity, condition, or safety assertions. Neither context includes a gold answer, score code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"Door 6 field note. At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment. Both records required factory-sealed units with dry, intact packaging and no visible damage. At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 23 sealed physical units bearing SKU NX-741 and one sealed physical unit bearing SKU NX-742. The same inspection found every unit and carton in the condition required by both records; unloading was completed normally without a safety issue, recount need, or containment concern.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment."}, {"path": [], "text": "At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 23 sealed physical units bearing SKU NX-741 and one sealed physical unit bearing SKU NX-742."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment.", "negative_left": "At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment.", "negative_right": "At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 24 sealed physical units bearing SKU NX-741.", "right": "At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 23 sealed physical units bearing SKU NX-741 and one sealed physical unit bearing SKU NX-742."}, "verifier_independent_model": false}, "family": "scale-diverse-282-004", "id": "scale-diverse-282-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "Door 6 field note. At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment. Both records required factory-sealed units with dry, intact packaging and no visible damage. At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 23 sealed physical units bearing SKU NX-741 and one sealed physical unit bearing SKU NX-742. The same inspection found every unit and carton in the condition required by both records; unloading was completed normally without a safety issue, recount need, or containment concern."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing routing policy and criteria, while neither generated context removes or alters any additional governing policy from the original state. The generic decision request remains bound to the shipment evidence, and the changed shipment identifiers, quantities, and observations are permissible case-observation changes. The two evidence spans are complete factual sentences. The counterfactual coherently replaces the mixed-SKU count with 24 units of the specified SKU without conflicting with the unchanged quantity, condition, or safety assertions. Neither context includes a gold answer, score code, rule table, proposition identifier, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"Door 6 field note. At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment. Both records required factory-sealed units with dry, intact packaging and no visible damage. At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 23 sealed physical units bearing SKU NX-741 and one sealed physical unit bearing SKU NX-742. The same inspection found every unit and carton in the condition required by both records; unloading was completed normally without a safety issue, recount need, or containment concern.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment."}, {"path": [], "text": "At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 23 sealed physical units bearing SKU NX-741 and one sealed physical unit bearing SKU NX-742."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment.", "negative_left": "At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment.", "negative_right": "At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 24 sealed physical units bearing SKU NX-741.", "right": "At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 23 sealed physical units bearing SKU NX-741 and one sealed physical unit bearing SKU NX-742."}, "verifier_independent_model": false}, "family": "scale-diverse-282-004", "id": "scale-diverse-282-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "Door 6 field note. At 09:10 on 17 September 2026, PO-1842 and Northstar ASN NS-6307 each specified 24 sealed units of SKU NX-741, the ASN authorized no substitute SKU, and the Door 6 receiving log recorded no unresolved non-SKU discrepancy or independent physical receiving risk for the shipment. Both records required factory-sealed units with dry, intact packaging and no visible damage. At 09:12 on 17 September 2026, the Door 6 inspection counted Northstar’s shipment as 24 sealed physical units bearing SKU NX-741. The same inspection found every unit and carton in the condition required by both records; unloading was completed normally without a safety issue, recount need, or containment concern."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release and routing policy without adding exceptions, priorities, or missing-evidence defaults. They preserve the decision scope and the Line 4, Mint-500, and 14:00 bindings while permissibly changing the quality-check observations. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Q-846 from 112 to 119, which is coherent with its stated inclusive range and does not create a duplicate or contradictory measurement. Neither context contains an answer label, code, rule table, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "full_context_fact_states": {"base": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "counterfactual": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "remove_left": {"quality_checks_passed": "unknown"}, "remove_right": {"quality_checks_passed": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"quality_checks_passed": "unknown"}, "negative_pair": {"quality_checks_passed": "refuted"}, "negative_sentence": {"quality_checks_passed": "unknown"}, "positive_pair": {"quality_checks_passed": "supported"}, "right": {"quality_checks_passed": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship, including permissible universally quantified relationships over required materials, records, or checks. The focus is a factual quality-check status. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state adds case observations but no additional interpretive policy that must be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports every readiness condition expressly required for release: current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault.", "rule_index": 0, "sound": true}, {"reason": "A refuted required-quality-check condition prevents release. With materials and every other readiness condition supported, the blocker is quality rather than material for the scheduled Mint-500 run, so materials support is excluded and none_of_above is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_ready", "statement": "Every material required for Line 4’s currently scheduled Mint-500 run at 14:00 is staged in the required quantity."}, {"id": "tooling_approved", "statement": "The tooling installed on Line 4 for the currently scheduled Mint-500 run at 14:00 matches the approved Mint-500 tooling specification."}, {"id": "cleaning_records_complete", "statement": "Every cleaning record required for Line 4’s changeover to the currently scheduled Mint-500 run at 14:00 is complete."}, {"id": "staffing_adequate", "statement": "The number of trained operators assigned to Line 4 for the currently scheduled Mint-500 run at 14:00 meets or exceeds that run’s staffing requirement."}, {"id": "quality_checks_passed", "statement": "Every quality-check result required for Line 4 after cleaning and before the currently scheduled Mint-500 run at 14:00 satisfies its corresponding signed pass criterion."}, {"id": "no_open_maintenance_fault", "statement": "No maintenance fault affecting Line 4 is open at the readiness decision for the currently scheduled Mint-500 run at 14:00."}], "base_state_json": "\"The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00. The staging reconciliation confirms that every material required for the run is staged in the required quantity. Inspection confirms that the tooling installed on Line 4 matches the approved Mint-500 tooling specification. Document control confirms that every cleaning record required for the changeover is complete. The assignment roster lists four trained operators, meeting the run’s staffing requirement of four. At the readiness decision, the maintenance log shows no open fault affecting Line 4; yesterday’s conveyor sensor repair is closed after successful testing. Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 112 against its signed inclusive pass range of 108 to 116.\"", "base_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "counter_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "focus_atom": "quality_checks_passed", "focus_evidence": [{"path": [], "text": "The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00."}, {"path": [], "text": "Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 112 against its signed inclusive pass range of 108 to 116."}], "policy_evidence": [], "rules": [{"justification": "All expressly required readiness conditions are satisfied, including the required post-clean quality checks.", "target": "release_to_production", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}, {"justification": "A failed required quality check prevents release, while supported current-job material readiness means the sole blocker is not missing or deficient material and therefore does not qualify for materials support.", "target": "none_of_above", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}]}, "verified_pair": {"left": "The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00.", "negative_left": "The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00.", "negative_right": "Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 119 against its signed inclusive pass range of 108 to 116.", "right": "Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 112 against its signed inclusive pass range of 108 to 116."}, "verifier_independent_model": false}, "family": "scale-diverse-289-003", "id": "scale-diverse-289-003-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Select when the line cannot be released and the evidence does not meet the materials-support rubric, including when an unresolved quality or maintenance requirement is the blocker.", "release_to_production": "Release the line to production only if every required readiness condition is satisfied, including a passed post-clean quality check.", "route_to_materials_support": "Do not release; route to materials support only if the sole blocker is material required for the currently scheduled Mint-500 run."}, "instructions": "Choose the single applicable routing/readiness outcome. Release requires current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault. Route to materials support only when the sole release blocker is missing or deficient material for the scheduled job.", "type": "choice"}}, "state": "The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00. The staging reconciliation confirms that every material required for the run is staged in the required quantity. Inspection confirms that the tooling installed on Line 4 matches the approved Mint-500 tooling specification. Document control confirms that every cleaning record required for the changeover is complete. The assignment roster lists four trained operators, meeting the run’s staffing requirement of four. At the readiness decision, the maintenance log shows no open fault affecting Line 4; yesterday’s conveyor sensor repair is closed after successful testing. Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 112 against its signed inclusive pass range of 108 to 116."}, "method": "c2d", "provenance": {"source_id": "diverse-289", "source_is_synthetic": true, "source_sha256": "884345667ad55a5ef86e09cc2ba68f6b13c8ccb20beb9fbed251a080f3d46c34", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "release_to_production"}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release and routing policy without adding exceptions, priorities, or missing-evidence defaults. They preserve the decision scope and the Line 4, Mint-500, and 14:00 bindings while permissibly changing the quality-check observations. The two focus-evidence spans are complete factual sentences. The counterfactual changes only Q-846 from 112 to 119, which is coherent with its stated inclusive range and does not create a duplicate or contradictory measurement. Neither context contains an answer label, code, rule table, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "full_context_fact_states": {"base": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "counterfactual": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "remove_left": {"quality_checks_passed": "unknown"}, "remove_right": {"quality_checks_passed": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"quality_checks_passed": "unknown"}, "negative_pair": {"quality_checks_passed": "refuted"}, "negative_sentence": {"quality_checks_passed": "unknown"}, "positive_pair": {"quality_checks_passed": "supported"}, "right": {"quality_checks_passed": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship, including permissible universally quantified relationships over required materials, records, or checks. The focus is a factual quality-check status. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state adds case observations but no additional interpretive policy that must be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports every readiness condition expressly required for release: current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault.", "rule_index": 0, "sound": true}, {"reason": "A refuted required-quality-check condition prevents release. With materials and every other readiness condition supported, the blocker is quality rather than material for the scheduled Mint-500 run, so materials support is excluded and none_of_above is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_ready", "statement": "Every material required for Line 4’s currently scheduled Mint-500 run at 14:00 is staged in the required quantity."}, {"id": "tooling_approved", "statement": "The tooling installed on Line 4 for the currently scheduled Mint-500 run at 14:00 matches the approved Mint-500 tooling specification."}, {"id": "cleaning_records_complete", "statement": "Every cleaning record required for Line 4’s changeover to the currently scheduled Mint-500 run at 14:00 is complete."}, {"id": "staffing_adequate", "statement": "The number of trained operators assigned to Line 4 for the currently scheduled Mint-500 run at 14:00 meets or exceeds that run’s staffing requirement."}, {"id": "quality_checks_passed", "statement": "Every quality-check result required for Line 4 after cleaning and before the currently scheduled Mint-500 run at 14:00 satisfies its corresponding signed pass criterion."}, {"id": "no_open_maintenance_fault", "statement": "No maintenance fault affecting Line 4 is open at the readiness decision for the currently scheduled Mint-500 run at 14:00."}], "base_state_json": "\"The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00. The staging reconciliation confirms that every material required for the run is staged in the required quantity. Inspection confirms that the tooling installed on Line 4 matches the approved Mint-500 tooling specification. Document control confirms that every cleaning record required for the changeover is complete. The assignment roster lists four trained operators, meeting the run’s staffing requirement of four. At the readiness decision, the maintenance log shows no open fault affecting Line 4; yesterday’s conveyor sensor repair is closed after successful testing. Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 112 against its signed inclusive pass range of 108 to 116.\"", "base_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "counter_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "focus_atom": "quality_checks_passed", "focus_evidence": [{"path": [], "text": "The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00."}, {"path": [], "text": "Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 112 against its signed inclusive pass range of 108 to 116."}], "policy_evidence": [], "rules": [{"justification": "All expressly required readiness conditions are satisfied, including the required post-clean quality checks.", "target": "release_to_production", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}, {"justification": "A failed required quality check prevents release, while supported current-job material readiness means the sole blocker is not missing or deficient material and therefore does not qualify for materials support.", "target": "none_of_above", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}]}, "verified_pair": {"left": "The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00.", "negative_left": "The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00.", "negative_right": "Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 119 against its signed inclusive pass range of 108 to 116.", "right": "Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 112 against its signed inclusive pass range of 108 to 116."}, "verifier_independent_model": false}, "family": "scale-diverse-289-003", "id": "scale-diverse-289-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Select when the line cannot be released and the evidence does not meet the materials-support rubric, including when an unresolved quality or maintenance requirement is the blocker.", "release_to_production": "Release the line to production only if every required readiness condition is satisfied, including a passed post-clean quality check.", "route_to_materials_support": "Do not release; route to materials support only if the sole blocker is material required for the currently scheduled Mint-500 run."}, "instructions": "Choose the single applicable routing/readiness outcome. Release requires current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault. Route to materials support only when the sole release blocker is missing or deficient material for the scheduled job.", "type": "choice"}}, "state": "The 13:45 operational handoff register records cleaning of Line 4 as completed at 13:20 and identifies Q-731 and Q-846 as the complete set of quality checks required between that cleaning and the currently scheduled Mint-500 run at 14:00. The staging reconciliation confirms that every material required for the run is staged in the required quantity. Inspection confirms that the tooling installed on Line 4 matches the approved Mint-500 tooling specification. Document control confirms that every cleaning record required for the changeover is complete. The assignment roster lists four trained operators, meeting the run’s staffing requirement of four. At the readiness decision, the maintenance log shows no open fault affecting Line 4; yesterday’s conveyor sensor repair is closed after successful testing. Line 4's Q-731 result at 13:32 was 7.4 against its signed inclusive pass range of 7.0 to 8.0, and its Q-846 result at 13:39 was 119 against its signed inclusive pass range of 108 to 116."}, "method": "c2d", "provenance": {"source_id": "diverse-289", "source_is_synthetic": true, "source_sha256": "884345667ad55a5ef86e09cc2ba68f6b13c8ccb20beb9fbed251a080f3d46c34", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original readiness rubric and remain bound to the Citrus-500 to Berry-500 changeover scheduled for 14:00. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes the specified tooling from TK-731 to the distinct TK-884 while leaving TK-731 installed, creating a tooling mismatch rather than contradictory duplicate measurements. Neither context states a Green/Amber/Red classification, release decision, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 13:05, the staging ledger recorded every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 as staged, including Berry concentrate lot B771 and 18,000 clean bottles for the planned 17,500-bottle run.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 13:12, all six operators required by the staffing plan were present, and cleaning record CL-204 carried the required signature.\"},{\"speaker\":\"Changeover coordinator\",\"text\":\"At 13:20, the signed changeover sheet specified tooling kit TK-731 for the Citrus-500 to Berry-500 changeover scheduled at 14:00.\"},{\"speaker\":\"Maintenance planner\",\"text\":\"At 13:32, the maintenance register showed WO-619, the sole work order required before release, as closed with its required torque reading and technician signature recorded.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 13:41, the quality file confirmed that every prerequisite required before release had passed, including allergen swab Q-882 and label-code verification.\"},{\"speaker\":\"Line technician\",\"text\":\"At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["5", "text"], "text": "At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}, {"path": ["2", "text"], "text": "At 13:20, the signed changeover sheet specified tooling kit TK-731 for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00.", "negative_left": "At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00.", "negative_right": "At 13:20, the signed changeover sheet specified tooling kit TK-884, a physical kit distinct from TK-731, for the Citrus-500 to Berry-500 changeover scheduled at 14:00.", "right": "At 13:20, the signed changeover sheet specified tooling kit TK-731 for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}, "verifier_independent_model": false}, "family": "scale-diverse-291-001", "id": "scale-diverse-291-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "At 13:05, the staging ledger recorded every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 as staged, including Berry concentrate lot B771 and 18,000 clean bottles for the planned 17,500-bottle run."}, {"speaker": "Line supervisor", "text": "At 13:12, all six operators required by the staffing plan were present, and cleaning record CL-204 carried the required signature."}, {"speaker": "Changeover coordinator", "text": "At 13:20, the signed changeover sheet specified tooling kit TK-731 for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}, {"speaker": "Maintenance planner", "text": "At 13:32, the maintenance register showed WO-619, the sole work order required before release, as closed with its required torque reading and technician signature recorded."}, {"speaker": "Quality inspector", "text": "At 13:41, the quality file confirmed that every prerequisite required before release had passed, including allergen swab Q-882 and label-code verification."}, {"speaker": "Line technician", "text": "At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original readiness rubric and remain bound to the Citrus-500 to Berry-500 changeover scheduled for 14:00. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes the specified tooling from TK-731 to the distinct TK-884 while leaving TK-731 installed, creating a tooling mismatch rather than contradictory duplicate measurements. Neither context states a Green/Amber/Red classification, release decision, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 13:05, the staging ledger recorded every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 as staged, including Berry concentrate lot B771 and 18,000 clean bottles for the planned 17,500-bottle run.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 13:12, all six operators required by the staffing plan were present, and cleaning record CL-204 carried the required signature.\"},{\"speaker\":\"Changeover coordinator\",\"text\":\"At 13:20, the signed changeover sheet specified tooling kit TK-731 for the Citrus-500 to Berry-500 changeover scheduled at 14:00.\"},{\"speaker\":\"Maintenance planner\",\"text\":\"At 13:32, the maintenance register showed WO-619, the sole work order required before release, as closed with its required torque reading and technician signature recorded.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 13:41, the quality file confirmed that every prerequisite required before release had passed, including allergen swab Q-882 and label-code verification.\"},{\"speaker\":\"Line technician\",\"text\":\"At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["5", "text"], "text": "At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}, {"path": ["2", "text"], "text": "At 13:20, the signed changeover sheet specified tooling kit TK-731 for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00.", "negative_left": "At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00.", "negative_right": "At 13:20, the signed changeover sheet specified tooling kit TK-884, a physical kit distinct from TK-731, for the Citrus-500 to Berry-500 changeover scheduled at 14:00.", "right": "At 13:20, the signed changeover sheet specified tooling kit TK-731 for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}, "verifier_independent_model": false}, "family": "scale-diverse-291-001", "id": "scale-diverse-291-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "At 13:05, the staging ledger recorded every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 as staged, including Berry concentrate lot B771 and 18,000 clean bottles for the planned 17,500-bottle run."}, {"speaker": "Line supervisor", "text": "At 13:12, all six operators required by the staffing plan were present, and cleaning record CL-204 carried the required signature."}, {"speaker": "Changeover coordinator", "text": "At 13:20, the signed changeover sheet specified tooling kit TK-884, a physical kit distinct from TK-731, for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}, {"speaker": "Maintenance planner", "text": "At 13:32, the maintenance register showed WO-619, the sole work order required before release, as closed with its required torque reading and technician signature recorded."}, {"speaker": "Quality inspector", "text": "At 13:41, the quality file confirmed that every prerequisite required before release had passed, including allergen swab Q-882 and label-code verification."}, {"speaker": "Line technician", "text": "At 13:47, tooling kit TK-731 was installed for the Citrus-500 to Berry-500 changeover scheduled at 14:00."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original readiness rubric and remain bound to the Citrus-500 to Berry-500 changeover at 14:00. The two evidence spans are complete factual sentences. The counterfactual coherently changes the specified tooling to TK-8842 while leaving the installed tooling as TK-7316; these are distinct facts about specified versus installed equipment, not contradictory duplicate measurements. Neither context states a classification, release decision, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"For the Citrus-500 to Berry-500 changeover at 14:00, the complete schedule lists Berry concentrate lot B771 and 18,000 clean bottles as every scheduled material, and both are staged. The planned Berry-500 run requires 17,500 bottles.\"},{\"speaker\":\"Installation recorder\",\"text\":\"The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316.\"},{\"speaker\":\"Setup planner\",\"text\":\"The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-7316.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The changeover requires six operators, and six required operators are present. Cleaning record CL-204 bears the required signature.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release of this changeover has passed, including the allergen swab and label-code check.\"},{\"speaker\":\"Maintenance coordinator\",\"text\":\"Every maintenance work order required before release of this changeover is closed. Every such work order has its required torque reading and required technician signature recorded.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316."}, {"path": ["2", "text"], "text": "The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-7316."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316.", "negative_left": "The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316.", "negative_right": "The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-8842, which is distinct from plant asset TK-7316.", "right": "The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-7316."}, "verifier_independent_model": false}, "family": "scale-diverse-291-002", "id": "scale-diverse-291-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "For the Citrus-500 to Berry-500 changeover at 14:00, the complete schedule lists Berry concentrate lot B771 and 18,000 clean bottles as every scheduled material, and both are staged. The planned Berry-500 run requires 17,500 bottles."}, {"speaker": "Installation recorder", "text": "The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316."}, {"speaker": "Setup planner", "text": "The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-7316."}, {"speaker": "Line supervisor", "text": "The changeover requires six operators, and six required operators are present. Cleaning record CL-204 bears the required signature."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release of this changeover has passed, including the allergen swab and label-code check."}, {"speaker": "Maintenance coordinator", "text": "Every maintenance work order required before release of this changeover is closed. Every such work order has its required torque reading and required technician signature recorded."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original readiness rubric and remain bound to the Citrus-500 to Berry-500 changeover at 14:00. The two evidence spans are complete factual sentences. The counterfactual coherently changes the specified tooling to TK-8842 while leaving the installed tooling as TK-7316; these are distinct facts about specified versus installed equipment, not contradictory duplicate measurements. Neither context states a classification, release decision, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"For the Citrus-500 to Berry-500 changeover at 14:00, the complete schedule lists Berry concentrate lot B771 and 18,000 clean bottles as every scheduled material, and both are staged. The planned Berry-500 run requires 17,500 bottles.\"},{\"speaker\":\"Installation recorder\",\"text\":\"The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316.\"},{\"speaker\":\"Setup planner\",\"text\":\"The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-7316.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The changeover requires six operators, and six required operators are present. Cleaning record CL-204 bears the required signature.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release of this changeover has passed, including the allergen swab and label-code check.\"},{\"speaker\":\"Maintenance coordinator\",\"text\":\"Every maintenance work order required before release of this changeover is closed. Every such work order has its required torque reading and required technician signature recorded.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316."}, {"path": ["2", "text"], "text": "The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-7316."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316.", "negative_left": "The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316.", "negative_right": "The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-8842, which is distinct from plant asset TK-7316.", "right": "The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-7316."}, "verifier_independent_model": false}, "family": "scale-diverse-291-002", "id": "scale-diverse-291-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "For the Citrus-500 to Berry-500 changeover at 14:00, the complete schedule lists Berry concentrate lot B771 and 18,000 clean bottles as every scheduled material, and both are staged. The planned Berry-500 run requires 17,500 bottles."}, {"speaker": "Installation recorder", "text": "The 13:47 installation scan identifies the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 as the unique plant asset numbered TK-7316."}, {"speaker": "Setup planner", "text": "The signed setup sheet for the Citrus-500 to Berry-500 changeover at 14:00 identifies the specified tooling kit as the unique plant asset numbered TK-8842, which is distinct from plant asset TK-7316."}, {"speaker": "Line supervisor", "text": "The changeover requires six operators, and six required operators are present. Cleaning record CL-204 bears the required signature."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release of this changeover has passed, including the allergen swab and label-code check."}, {"speaker": "Maintenance coordinator", "text": "Every maintenance work order required before release of this changeover is closed. Every such work order has its required torque reading and required technician signature recorded."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full release rubric, and neither context alters or invents governing policy. Both contexts retain the Citrus-500 to Berry-500 changeover, 14:00 timing, and current-release decision scope. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes the approved tooling specification from TK-271 to TK-846 while leaving TK-271 recorded as installed; these are distinct specification and installation assertions that create a mismatch rather than a contradictory duplicate measurement. Neither context states a classification, release answer, answer code, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Handoff recorder\",\"text\":\"At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00.\"},{\"speaker\":\"Changeover planner\",\"text\":\"The approved changeover sheet CS-914 specifies tooling kit asset TK-271 for the Citrus-500 to Berry-500 changeover at 14:00.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"The schedule identifies Berry concentrate lot B771 and clean bottles as the only materials needed for the 14:00 changeover, and both are staged. The bottle count is 18,000 against 17,500 required for the planned Berry-500 run.\"},{\"speaker\":\"Shift coordinator\",\"text\":\"The changeover requires six operators, and all six required operators are present. Cleaning record CL-204 bears its required signature.\"},{\"speaker\":\"Quality coordinator\",\"text\":\"The completed quality register confirms that every prerequisite required before release has passed, including the allergen swab, label-code check, and first-piece inspection.\"},{\"speaker\":\"Maintenance coordinator\",\"text\":\"WO-619 is the only maintenance work order required before release of this changeover. It is closed, with the required torque reading and technician signature recorded in the maintenance system.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00."}, {"path": ["1", "text"], "text": "The approved changeover sheet CS-914 specifies tooling kit asset TK-271 for the Citrus-500 to Berry-500 changeover at 14:00."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00.", "negative_left": "At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00.", "negative_right": "The approved changeover sheet CS-914 specifies tooling kit asset TK-846 for the Citrus-500 to Berry-500 changeover at 14:00.", "right": "The approved changeover sheet CS-914 specifies tooling kit asset TK-271 for the Citrus-500 to Berry-500 changeover at 14:00."}, "verifier_independent_model": false}, "family": "scale-diverse-291-003", "id": "scale-diverse-291-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Handoff recorder", "text": "At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00."}, {"speaker": "Changeover planner", "text": "The approved changeover sheet CS-914 specifies tooling kit asset TK-271 for the Citrus-500 to Berry-500 changeover at 14:00."}, {"speaker": "Materials coordinator", "text": "The schedule identifies Berry concentrate lot B771 and clean bottles as the only materials needed for the 14:00 changeover, and both are staged. The bottle count is 18,000 against 17,500 required for the planned Berry-500 run."}, {"speaker": "Shift coordinator", "text": "The changeover requires six operators, and all six required operators are present. Cleaning record CL-204 bears its required signature."}, {"speaker": "Quality coordinator", "text": "The completed quality register confirms that every prerequisite required before release has passed, including the allergen swab, label-code check, and first-piece inspection."}, {"speaker": "Maintenance coordinator", "text": "WO-619 is the only maintenance work order required before release of this changeover. It is closed, with the required torque reading and technician signature recorded in the maintenance system."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full release rubric, and neither context alters or invents governing policy. Both contexts retain the Citrus-500 to Berry-500 changeover, 14:00 timing, and current-release decision scope. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes the approved tooling specification from TK-271 to TK-846 while leaving TK-271 recorded as installed; these are distinct specification and installation assertions that create a mismatch rather than a contradictory duplicate measurement. Neither context states a classification, release answer, answer code, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Handoff recorder\",\"text\":\"At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00.\"},{\"speaker\":\"Changeover planner\",\"text\":\"The approved changeover sheet CS-914 specifies tooling kit asset TK-271 for the Citrus-500 to Berry-500 changeover at 14:00.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"The schedule identifies Berry concentrate lot B771 and clean bottles as the only materials needed for the 14:00 changeover, and both are staged. The bottle count is 18,000 against 17,500 required for the planned Berry-500 run.\"},{\"speaker\":\"Shift coordinator\",\"text\":\"The changeover requires six operators, and all six required operators are present. Cleaning record CL-204 bears its required signature.\"},{\"speaker\":\"Quality coordinator\",\"text\":\"The completed quality register confirms that every prerequisite required before release has passed, including the allergen swab, label-code check, and first-piece inspection.\"},{\"speaker\":\"Maintenance coordinator\",\"text\":\"WO-619 is the only maintenance work order required before release of this changeover. It is closed, with the required torque reading and technician signature recorded in the maintenance system.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00."}, {"path": ["1", "text"], "text": "The approved changeover sheet CS-914 specifies tooling kit asset TK-271 for the Citrus-500 to Berry-500 changeover at 14:00."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00.", "negative_left": "At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00.", "negative_right": "The approved changeover sheet CS-914 specifies tooling kit asset TK-846 for the Citrus-500 to Berry-500 changeover at 14:00.", "right": "The approved changeover sheet CS-914 specifies tooling kit asset TK-271 for the Citrus-500 to Berry-500 changeover at 14:00."}, "verifier_independent_model": false}, "family": "scale-diverse-291-003", "id": "scale-diverse-291-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Handoff recorder", "text": "At 13:47, handoff log HL-638 recorded tooling kit asset TK-271 as installed for the Citrus-500 to Berry-500 changeover at 14:00."}, {"speaker": "Changeover planner", "text": "The approved changeover sheet CS-914 specifies tooling kit asset TK-846 for the Citrus-500 to Berry-500 changeover at 14:00."}, {"speaker": "Materials coordinator", "text": "The schedule identifies Berry concentrate lot B771 and clean bottles as the only materials needed for the 14:00 changeover, and both are staged. The bottle count is 18,000 against 17,500 required for the planned Berry-500 run."}, {"speaker": "Shift coordinator", "text": "The changeover requires six operators, and all six required operators are present. Cleaning record CL-204 bears its required signature."}, {"speaker": "Quality coordinator", "text": "The completed quality register confirms that every prerequisite required before release has passed, including the allergen swab, label-code check, and first-piece inspection."}, {"speaker": "Maintenance coordinator", "text": "WO-619 is the only maintenance work order required before release of this changeover. It is closed, with the required torque reading and technician signature recorded in the maintenance system."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question’s governing release rubric and preserve the Citrus-500-to-Berry-500 changeover at 14:00. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the tooling serial specified by the signed sheet to TK-946 while the installed kit remains TK-731; this creates a tooling mismatch rather than a contradictory duplicate assertion, because the sheet-specified kit need not be the installed kit. Neither context includes a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Inventory clerk\",\"text\":\"At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial.\"},{\"speaker\":\"Changeover coordinator\",\"text\":\"The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-731 for the Citrus-500 to Berry-500 changeover at 14:00.\"},{\"speaker\":\"Materials controller\",\"text\":\"Staging ledger SL-88 confirms every material scheduled for the 14:00 Citrus-500 to Berry-500 changeover is staged. It records 18,000 clean bottles against 17,500 required for the planned Berry-500 run.\"},{\"speaker\":\"Line supervisor\",\"text\":\"Attendance records show six required operators present, meeting the required staffing level. Cleaning record CL-204 has the required signature.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release of the 14:00 changeover has passed.\"},{\"speaker\":\"Maintenance planner\",\"text\":\"Every maintenance work order required before release of the 14:00 changeover is closed. Each has its required torque reading and technician signature recorded.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial."}, {"path": ["1", "text"], "text": "The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-731 for the Citrus-500 to Berry-500 changeover at 14:00."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial.", "negative_left": "At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial.", "negative_right": "The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-946 for the Citrus-500 to Berry-500 changeover at 14:00.", "right": "The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-731 for the Citrus-500 to Berry-500 changeover at 14:00."}, "verifier_independent_model": false}, "family": "scale-diverse-291-004", "id": "scale-diverse-291-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Inventory clerk", "text": "At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial."}, {"speaker": "Changeover coordinator", "text": "The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-731 for the Citrus-500 to Berry-500 changeover at 14:00."}, {"speaker": "Materials controller", "text": "Staging ledger SL-88 confirms every material scheduled for the 14:00 Citrus-500 to Berry-500 changeover is staged. It records 18,000 clean bottles against 17,500 required for the planned Berry-500 run."}, {"speaker": "Line supervisor", "text": "Attendance records show six required operators present, meeting the required staffing level. Cleaning record CL-204 has the required signature."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release of the 14:00 changeover has passed."}, {"speaker": "Maintenance planner", "text": "Every maintenance work order required before release of the 14:00 changeover is closed. Each has its required torque reading and technician signature recorded."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question’s governing release rubric and preserve the Citrus-500-to-Berry-500 changeover at 14:00. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes only the tooling serial specified by the signed sheet to TK-946 while the installed kit remains TK-731; this creates a tooling mismatch rather than a contradictory duplicate assertion, because the sheet-specified kit need not be the installed kit. Neither context includes a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Inventory clerk\",\"text\":\"At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial.\"},{\"speaker\":\"Changeover coordinator\",\"text\":\"The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-731 for the Citrus-500 to Berry-500 changeover at 14:00.\"},{\"speaker\":\"Materials controller\",\"text\":\"Staging ledger SL-88 confirms every material scheduled for the 14:00 Citrus-500 to Berry-500 changeover is staged. It records 18,000 clean bottles against 17,500 required for the planned Berry-500 run.\"},{\"speaker\":\"Line supervisor\",\"text\":\"Attendance records show six required operators present, meeting the required staffing level. Cleaning record CL-204 has the required signature.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release of the 14:00 changeover has passed.\"},{\"speaker\":\"Maintenance planner\",\"text\":\"Every maintenance work order required before release of the 14:00 changeover is closed. Each has its required torque reading and technician signature recorded.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial."}, {"path": ["1", "text"], "text": "The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-731 for the Citrus-500 to Berry-500 changeover at 14:00."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial.", "negative_left": "At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial.", "negative_right": "The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-946 for the Citrus-500 to Berry-500 changeover at 14:00.", "right": "The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-731 for the Citrus-500 to Berry-500 changeover at 14:00."}, "verifier_independent_model": false}, "family": "scale-diverse-291-004", "id": "scale-diverse-291-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Inventory clerk", "text": "At 13:47, inventory check IC-639 recorded that the tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 has manufacturer serial TK-731, that each tooling kit has exactly one manufacturer serial, and that no two tooling kits share a manufacturer serial."}, {"speaker": "Changeover coordinator", "text": "The signed changeover sheet CS-482 specifies a tooling kit with manufacturer serial TK-946 for the Citrus-500 to Berry-500 changeover at 14:00."}, {"speaker": "Materials controller", "text": "Staging ledger SL-88 confirms every material scheduled for the 14:00 Citrus-500 to Berry-500 changeover is staged. It records 18,000 clean bottles against 17,500 required for the planned Berry-500 run."}, {"speaker": "Line supervisor", "text": "Attendance records show six required operators present, meeting the required staffing level. Cleaning record CL-204 has the required signature."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release of the 14:00 changeover has passed."}, {"speaker": "Maintenance planner", "text": "Every maintenance work order required before release of the 14:00 changeover is closed. Each has its required torque reading and technician signature recorded."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release policy and preserve the Line 4, scheduled 14:00 blue-to-white cap change, production-release decision, and 13:55 decision timing. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes the finalized roster from nine to seven trained assigned operators while retaining the unchanged requirement for eight positions, without creating duplicate or contradictory measurements. Neither context includes an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions. At 13:20, the signed record showed that all required cleaning of Line 4 had been completed. By 13:24, every material required for the change, including the white resin and labels, had been staged. At 13:30, the specified white-cap mold was installed, and every required tooling check had passed. At 13:35, the maintenance closeout showed that no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed nine assigned operators, each recorded as trained for that change. At 13:50, the quality inspector approved the first piece for the scheduled white-cap production. These records were assembled for the release decision at 13:55. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions."}, {"path": [], "text": "At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed nine assigned operators, each recorded as trained for that change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions.", "negative_left": "At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions.", "negative_right": "At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed seven assigned operators, each recorded as trained for that change.", "right": "At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed nine assigned operators, each recorded as trained for that change."}, "verifier_independent_model": false}, "family": "scale-diverse-292-001", "id": "scale-diverse-292-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions. At 13:20, the signed record showed that all required cleaning of Line 4 had been completed. By 13:24, every material required for the change, including the white resin and labels, had been staged. At 13:30, the specified white-cap mold was installed, and every required tooling check had passed. At 13:35, the maintenance closeout showed that no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed nine assigned operators, each recorded as trained for that change. At 13:50, the quality inspector approved the first piece for the scheduled white-cap production. These records were assembled for the release decision at 13:55. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release policy and preserve the Line 4, scheduled 14:00 blue-to-white cap change, production-release decision, and 13:55 decision timing. The focus evidence consists of exactly two complete factual sentences. The counterfactual coherently changes the finalized roster from nine to seven trained assigned operators while retaining the unchanged requirement for eight positions, without creating duplicate or contradictory measurements. Neither context includes an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions. At 13:20, the signed record showed that all required cleaning of Line 4 had been completed. By 13:24, every material required for the change, including the white resin and labels, had been staged. At 13:30, the specified white-cap mold was installed, and every required tooling check had passed. At 13:35, the maintenance closeout showed that no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed nine assigned operators, each recorded as trained for that change. At 13:50, the quality inspector approved the first piece for the scheduled white-cap production. These records were assembled for the release decision at 13:55. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions."}, {"path": [], "text": "At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed nine assigned operators, each recorded as trained for that change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions.", "negative_left": "At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions.", "negative_right": "At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed seven assigned operators, each recorded as trained for that change.", "right": "At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed nine assigned operators, each recorded as trained for that change."}, "verifier_independent_model": false}, "family": "scale-diverse-292-001", "id": "scale-diverse-292-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At 13:18, the signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed eight required positions. At 13:20, the signed record showed that all required cleaning of Line 4 had been completed. By 13:24, every material required for the change, including the white resin and labels, had been staged. At 13:30, the specified white-cap mold was installed, and every required tooling check had passed. At 13:35, the maintenance closeout showed that no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. At 13:42, the finalized assignment roster for Line 4's scheduled 14:00 change from blue valve caps to white valve caps listed seven assigned operators, each recorded as trained for that change. At 13:50, the quality inspector approved the first piece for the scheduled white-cap production. These records were assembled for the release decision at 13:55. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the release policy and retain the same Line 4, production-release, product-change, and scheduled-time bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the staffing requirement from six to eight positions while retaining seven assigned operators, without creating duplicate conflicting measurements. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Before the release decision at Northstar Components, records for Line 4’s scheduled 14:00 change from blue valve caps to white valve caps were reconciled. Warehouse record WR-88 confirms every required material was staged for the change. Cleaning record CR-41 shows the required cleaning was completed at 13:20. The mold specified for the white valve caps was installed, and the tooling sheet records that every required tooling check passed. Maintenance log ML-17 confirms that no required maintenance work remained open and that no defect affecting Line 4 remained unresolved. Quality inspector Mara Chen approved the first white-cap piece before release. At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change. Signed staffing specification SS-19 required six operator positions for Line 4's scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change."}, {"path": [], "text": "Signed staffing specification SS-19 required six operator positions for Line 4's scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_left": "At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_right": "Signed staffing specification SS-19 required eight operator positions for Line 4's scheduled 14:00 change.", "right": "Signed staffing specification SS-19 required six operator positions for Line 4's scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "scale-diverse-292-002", "id": "scale-diverse-292-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Before the release decision at Northstar Components, records for Line 4’s scheduled 14:00 change from blue valve caps to white valve caps were reconciled. Warehouse record WR-88 confirms every required material was staged for the change. Cleaning record CR-41 shows the required cleaning was completed at 13:20. The mold specified for the white valve caps was installed, and the tooling sheet records that every required tooling check passed. Maintenance log ML-17 confirms that no required maintenance work remained open and that no defect affecting Line 4 remained unresolved. Quality inspector Mara Chen approved the first white-cap piece before release. At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change. Signed staffing specification SS-19 required six operator positions for Line 4's scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the release policy and retain the same Line 4, production-release, product-change, and scheduled-time bindings. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the staffing requirement from six to eight positions while retaining seven assigned operators, without creating duplicate conflicting measurements. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Before the release decision at Northstar Components, records for Line 4’s scheduled 14:00 change from blue valve caps to white valve caps were reconciled. Warehouse record WR-88 confirms every required material was staged for the change. Cleaning record CR-41 shows the required cleaning was completed at 13:20. The mold specified for the white valve caps was installed, and the tooling sheet records that every required tooling check passed. Maintenance log ML-17 confirms that no required maintenance work remained open and that no defect affecting Line 4 remained unresolved. Quality inspector Mara Chen approved the first white-cap piece before release. At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change. Signed staffing specification SS-19 required six operator positions for Line 4's scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change."}, {"path": [], "text": "Signed staffing specification SS-19 required six operator positions for Line 4's scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_left": "At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_right": "Signed staffing specification SS-19 required eight operator positions for Line 4's scheduled 14:00 change.", "right": "Signed staffing specification SS-19 required six operator positions for Line 4's scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "scale-diverse-292-002", "id": "scale-diverse-292-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Before the release decision at Northstar Components, records for Line 4’s scheduled 14:00 change from blue valve caps to white valve caps were reconciled. Warehouse record WR-88 confirms every required material was staged for the change. Cleaning record CR-41 shows the required cleaning was completed at 13:20. The mold specified for the white valve caps was installed, and the tooling sheet records that every required tooling check passed. Maintenance log ML-17 confirms that no required maintenance work remained open and that no defect affecting Line 4 remained unresolved. Quality inspector Mara Chen approved the first white-cap piece before release. At 13:36, assignment roster AR-62 listed seven trained operators assigned to Line 4 for the scheduled 14:00 change. Signed staffing specification SS-19 required eight operator positions for Line 4's scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing release policy, while the unchanged questions preserve the decision instructions and criteria. They maintain the Line 4 release decision, scheduled 14:00 change, and relevant readiness path. The two evidence spans are complete factual sentences. The counterfactual coherently changes the staffing requirement from six to eight positions while retaining seven assigned operators, without creating duplicate contradictory measurements. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, the release review for Line 4's scheduled 14:00 change from blue valve caps to white valve caps occurred at 13:50. At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change. The signed staffing specification for Line 4's scheduled 14:00 change required six positions. Every material required for the change was staged before release. The signed cleaning record showed that the required cleaning was completed at 13:20, before release. The mold specified for production of the white valve caps was installed before release, and every required tooling check had passed. At the release decision, no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. The quality inspector approved the first piece for the scheduled production of white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change."}, {"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change required six positions."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change.", "negative_left": "At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change.", "negative_right": "The signed staffing specification for Line 4's scheduled 14:00 change required eight positions.", "right": "The signed staffing specification for Line 4's scheduled 14:00 change required six positions."}, "verifier_independent_model": false}, "family": "scale-diverse-292-003", "id": "scale-diverse-292-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, the release review for Line 4's scheduled 14:00 change from blue valve caps to white valve caps occurred at 13:50. At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change. The signed staffing specification for Line 4's scheduled 14:00 change required six positions. Every material required for the change was staged before release. The signed cleaning record showed that the required cleaning was completed at 13:20, before release. The mold specified for production of the white valve caps was installed before release, and every required tooling check had passed. At the release decision, no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. The quality inspector approved the first piece for the scheduled production of white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing release policy, while the unchanged questions preserve the decision instructions and criteria. They maintain the Line 4 release decision, scheduled 14:00 change, and relevant readiness path. The two evidence spans are complete factual sentences. The counterfactual coherently changes the staffing requirement from six to eight positions while retaining seven assigned operators, without creating duplicate contradictory measurements. Neither context contains an explicit answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, the release review for Line 4's scheduled 14:00 change from blue valve caps to white valve caps occurred at 13:50. At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change. The signed staffing specification for Line 4's scheduled 14:00 change required six positions. Every material required for the change was staged before release. The signed cleaning record showed that the required cleaning was completed at 13:20, before release. The mold specified for production of the white valve caps was installed before release, and every required tooling check had passed. At the release decision, no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. The quality inspector approved the first piece for the scheduled production of white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change."}, {"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change required six positions."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change.", "negative_left": "At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change.", "negative_right": "The signed staffing specification for Line 4's scheduled 14:00 change required eight positions.", "right": "The signed staffing specification for Line 4's scheduled 14:00 change required six positions."}, "verifier_independent_model": false}, "family": "scale-diverse-292-003", "id": "scale-diverse-292-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, the release review for Line 4's scheduled 14:00 change from blue valve caps to white valve caps occurred at 13:50. At 13:42, the operational handoff roster assigned seven trained operators to Line 4 for the scheduled 14:00 change. The signed staffing specification for Line 4's scheduled 14:00 change required eight positions. Every material required for the change was staged before release. The signed cleaning record showed that the required cleaning was completed at 13:20, before release. The mold specified for production of the white valve caps was installed before release, and every required tooling check had passed. At the release decision, no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. The quality inspector approved the first piece for the scheduled production of white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing release policy, while the unchanged questions preserve the decision instructions and criteria. They maintain the same entity, production-change path, and 14:00 time binding. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the required staffing count from 9 to 13 without creating contradictory duplicate facts, and neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Concise field note: At Northstar Components, Line 4 is scheduled to change from blue valve caps to white valve caps at 14:00. The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change. The signed staffing specification requires exactly 9 positions for Line 4's scheduled 14:00 change. Before release, every material required for the change was staged, and the signed cleaning record confirmed that the required cleaning was complete. The specified mold for white valve caps was installed before release, and every required tooling check passed. At the release decision, no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. Before release, the quality inspector approved the first piece for the scheduled 14:00 production of white valve caps. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change."}, {"path": [], "text": "The signed staffing specification requires exactly 9 positions for Line 4's scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_left": "The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_right": "The signed staffing specification requires exactly 13 positions for Line 4's scheduled 14:00 change.", "right": "The signed staffing specification requires exactly 9 positions for Line 4's scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "scale-diverse-292-004", "id": "scale-diverse-292-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Concise field note: At Northstar Components, Line 4 is scheduled to change from blue valve caps to white valve caps at 14:00. The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change. The signed staffing specification requires exactly 9 positions for Line 4's scheduled 14:00 change. Before release, every material required for the change was staged, and the signed cleaning record confirmed that the required cleaning was complete. The specified mold for white valve caps was installed before release, and every required tooling check passed. At the release decision, no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. Before release, the quality inspector approved the first piece for the scheduled 14:00 production of white valve caps. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing release policy, while the unchanged questions preserve the decision instructions and criteria. They maintain the same entity, production-change path, and 14:00 time binding. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the required staffing count from 9 to 13 without creating contradictory duplicate facts, and neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Concise field note: At Northstar Components, Line 4 is scheduled to change from blue valve caps to white valve caps at 14:00. The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change. The signed staffing specification requires exactly 9 positions for Line 4's scheduled 14:00 change. Before release, every material required for the change was staged, and the signed cleaning record confirmed that the required cleaning was complete. The specified mold for white valve caps was installed before release, and every required tooling check passed. At the release decision, no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. Before release, the quality inspector approved the first piece for the scheduled 14:00 production of white valve caps. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change."}, {"path": [], "text": "The signed staffing specification requires exactly 9 positions for Line 4's scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_left": "The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_right": "The signed staffing specification requires exactly 13 positions for Line 4's scheduled 14:00 change.", "right": "The signed staffing specification requires exactly 9 positions for Line 4's scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "scale-diverse-292-004", "id": "scale-diverse-292-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Concise field note: At Northstar Components, Line 4 is scheduled to change from blue valve caps to white valve caps at 14:00. The finalized assignment roster records exactly 11 trained operators assigned to Line 4 for the scheduled 14:00 change. The signed staffing specification requires exactly 13 positions for Line 4's scheduled 14:00 change. Before release, every material required for the change was staged, and the signed cleaning record confirmed that the required cleaning was complete. The specified mold for white valve caps was installed before release, and every required tooling check passed. At the release decision, no required maintenance work remained open and no maintenance defect affecting Line 4 remained unresolved. Before release, the quality inspector approved the first piece for the scheduled 14:00 production of white valve caps. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scheduler’s release conditions and the unchanged request to apply the ordered delay-risk rubric, without adding policy exceptions or defaults. The Line 4, 07:00 decision, scope, and requested readiness/release determination remain fixed. The two focus spans are complete factual sentences. The counterfactual coherently replaces the absence of an authorized clearance with a 06:56 clearance and no later reactivation; this is consistent with the hold lasting until its first authorized clearance and with the 07:00 evidence. Neither context states a readiness level, release verdict, answer code, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"context\":\"Line 4 was reviewed chronologically for its scheduled 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.\",\"evidence\":[\"By 06:20, records show every required bottle, cap, label, and resin quantity staged for the run.\",\"At 06:28, the required tooling verification was recorded complete and passing.\",\"The 06:15 record shows required cleaning accepted and every scheduled staff member present.\",\"At 06:40, the quality inspector documented every required quality check as passed.\",\"At the 07:00 decision, the consolidated release log confirms every required resource available, no stated release condition with a confirmed failure, and no active quality hold.\",\"The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance.\",\"The complete audit record from 06:43 through the 07:00 release decision contains no authorized clearance of MH-317.\"],\"request\":\"Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "5"], "text": "The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance."}, {"path": ["evidence", "6"], "text": "The complete audit record from 06:43 through the 07:00 release decision contains no authorized clearance of MH-317."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance.", "negative_left": "The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance.", "negative_right": "The complete audit record shows an authorized clearance of MH-317 at 06:56 and contains no later reactivation through the 07:00 release decision.", "right": "The complete audit record from 06:43 through the 07:00 release decision contains no authorized clearance of MH-317."}, "verifier_independent_model": false}, "family": "scale-diverse-293-001", "id": "scale-diverse-293-001-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"context": "Line 4 was reviewed chronologically for its scheduled 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.", "evidence": ["By 06:20, records show every required bottle, cap, label, and resin quantity staged for the run.", "At 06:28, the required tooling verification was recorded complete and passing.", "The 06:15 record shows required cleaning accepted and every scheduled staff member present.", "At 06:40, the quality inspector documented every required quality check as passed.", "At the 07:00 decision, the consolidated release log confirms every required resource available, no stated release condition with a confirmed failure, and no active quality hold.", "The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance.", "The complete audit record from 06:43 through the 07:00 release decision contains no authorized clearance of MH-317."], "request": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scheduler’s release conditions and the unchanged request to apply the ordered delay-risk rubric, without adding policy exceptions or defaults. The Line 4, 07:00 decision, scope, and requested readiness/release determination remain fixed. The two focus spans are complete factual sentences. The counterfactual coherently replaces the absence of an authorized clearance with a 06:56 clearance and no later reactivation; this is consistent with the hold lasting until its first authorized clearance and with the 07:00 evidence. Neither context states a readiness level, release verdict, answer code, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"context\":\"Line 4 was reviewed chronologically for its scheduled 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.\",\"evidence\":[\"By 06:20, records show every required bottle, cap, label, and resin quantity staged for the run.\",\"At 06:28, the required tooling verification was recorded complete and passing.\",\"The 06:15 record shows required cleaning accepted and every scheduled staff member present.\",\"At 06:40, the quality inspector documented every required quality check as passed.\",\"At the 07:00 decision, the consolidated release log confirms every required resource available, no stated release condition with a confirmed failure, and no active quality hold.\",\"The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance.\",\"The complete audit record from 06:43 through the 07:00 release decision contains no authorized clearance of MH-317.\"],\"request\":\"Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "5"], "text": "The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance."}, {"path": ["evidence", "6"], "text": "The complete audit record from 06:43 through the 07:00 release decision contains no authorized clearance of MH-317."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance.", "negative_left": "The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance.", "negative_right": "The complete audit record shows an authorized clearance of MH-317 at 06:56 and contains no later reactivation through the 07:00 release decision.", "right": "The complete audit record from 06:43 through the 07:00 release decision contains no authorized clearance of MH-317."}, "verifier_independent_model": false}, "family": "scale-diverse-293-001", "id": "scale-diverse-293-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"context": "Line 4 was reviewed chronologically for its scheduled 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.", "evidence": ["By 06:20, records show every required bottle, cap, label, and resin quantity staged for the run.", "At 06:28, the required tooling verification was recorded complete and passing.", "The 06:15 record shows required cleaning accepted and every scheduled staff member present.", "At 06:40, the quality inspector documented every required quality check as passed.", "At the 07:00 decision, the consolidated release log confirms every required resource available, no stated release condition with a confirmed failure, and no active quality hold.", "The audit record for Line 4 identifies MH-317 as the only maintenance hold relevant to the 07:00 release decision and shows that it became active at 06:43 and would remain so until its first authorized clearance.", "The complete audit record shows an authorized clearance of MH-317 at 06:56 and contains no later reactivation through the 07:00 release decision."], "request": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scheduler’s governing release conditions, while the unchanged questions object preserves the ordered rubric and instructions. The request remains bound to Line 4, the 07:00 release decision, and the same decision criterion. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes WO-927 from open to closed; because MH-684 is stated to be active exactly while that work order is open, this does not create a contradictory duplicate status. Neither context contains an answer level, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"context\":\"Line 4 was scheduled for a 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.\",\"evidence\":[\"At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open.\",\"At the 07:00 release decision, work order WO-927 was open.\",\"By 07:00, the run record documented every required bottle, cap, label, and resin quantity as staged. It also documented completion of the required tooling verification, acceptance of required cleaning, presence of every required staff member, and passing results for every required quality check.\",\"At the 07:00 decision, every resource required for the run was available. No scheduler-stated release condition had a confirmed failure. The hold log recorded that no quality hold was active on Line 4.\"],\"request\":\"Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open."}, {"path": ["evidence", "1"], "text": "At the 07:00 release decision, work order WO-927 was open."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open.", "negative_left": "At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open.", "negative_right": "At the 07:00 release decision, work order WO-927 was closed.", "right": "At the 07:00 release decision, work order WO-927 was open."}, "verifier_independent_model": false}, "family": "scale-diverse-293-004", "id": "scale-diverse-293-004-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"context": "Line 4 was scheduled for a 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.", "evidence": ["At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open.", "At the 07:00 release decision, work order WO-927 was open.", "By 07:00, the run record documented every required bottle, cap, label, and resin quantity as staged. It also documented completion of the required tooling verification, acceptance of required cleaning, presence of every required staff member, and passing results for every required quality check.", "At the 07:00 decision, every resource required for the run was available. No scheduler-stated release condition had a confirmed failure. The hold log recorded that no quality hold was active on Line 4."], "request": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scheduler’s governing release conditions, while the unchanged questions object preserves the ordered rubric and instructions. The request remains bound to Line 4, the 07:00 release decision, and the same decision criterion. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes WO-927 from open to closed; because MH-684 is stated to be active exactly while that work order is open, this does not create a contradictory duplicate status. Neither context contains an answer level, output instruction, rule table, proposition identifier, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"context\":\"Line 4 was scheduled for a 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.\",\"evidence\":[\"At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open.\",\"At the 07:00 release decision, work order WO-927 was open.\",\"By 07:00, the run record documented every required bottle, cap, label, and resin quantity as staged. It also documented completion of the required tooling verification, acceptance of required cleaning, presence of every required staff member, and passing results for every required quality check.\",\"At the 07:00 decision, every resource required for the run was available. No scheduler-stated release condition had a confirmed failure. The hold log recorded that no quality hold was active on Line 4.\"],\"request\":\"Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open."}, {"path": ["evidence", "1"], "text": "At the 07:00 release decision, work order WO-927 was open."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open.", "negative_left": "At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open.", "negative_right": "At the 07:00 release decision, work order WO-927 was closed.", "right": "At the 07:00 release decision, work order WO-927 was open."}, "verifier_independent_model": false}, "family": "scale-diverse-293-004", "id": "scale-diverse-293-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"context": "Line 4 was scheduled for a 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.", "evidence": ["At 06:41, Line 4's maintenance register listed MH-684 as the only maintenance hold governing the 07:00 release decision and specified that MH-684 remains active exactly while work order WO-927 is open.", "At the 07:00 release decision, work order WO-927 was closed.", "By 07:00, the run record documented every required bottle, cap, label, and resin quantity as staged. It also documented completion of the required tooling verification, acceptance of required cleaning, presence of every required staff member, and passing results for every required quality check.", "At the 07:00 decision, every resource required for the run was available. No scheduler-stated release condition had a confirmed failure. The hold log recorded that no quality hold was active on Line 4."], "request": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original mandatory-gate policy and the decision scope for changeover L4 from Jar-A to Jar-B as of 06:52. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the staged-resin measurement from 864 kg to 776 kg, remains consistent with the unchanged 810 kg first-lot requirement, and creates no duplicate or contradictory measurement within that context. Neither context contains an answer code, explicit classifier instruction, rule table, proposition ID, or stated gold label; references to release approval, a nonblocking issue, and monitoring are natural operational facts.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Staging clerk\",\"text\":\"At 06:52, the verified staging log recorded 864 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Production planner\",\"text\":\"At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The 06:52 equipment inspection confirmed that the Jar-B mold was installed. Cleaning record C-441 was signed, and the attendance roster confirmed that all four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality had approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"The earlier stop-work hold was formally cleared before 06:52. The current status board showed no stop-work hold applying to changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery was documented in the issue log and could disrupt continued production. It was classified as a nonblocking issue that did not block line release, with materials support monitoring the delay.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the verified staging log recorded 864 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the verified staging log recorded 864 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "At 06:52, the verified staging log recorded 776 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B.", "right": "At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B."}, "verifier_independent_model": false}, "family": "scale-diverse-294-001", "id": "scale-diverse-294-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Staging clerk", "text": "At 06:52, the verified staging log recorded 864 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Production planner", "text": "At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B."}, {"speaker": "Line supervisor", "text": "The 06:52 equipment inspection confirmed that the Jar-B mold was installed. Cleaning record C-441 was signed, and the attendance roster confirmed that all four required operators were present."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality had approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "The earlier stop-work hold was formally cleared before 06:52. The current status board showed no stop-work hold applying to changeover L4."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery was documented in the issue log and could disrupt continued production. It was classified as a nonblocking issue that did not block line release, with materials support monitoring the delay."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original mandatory-gate policy and the decision scope for changeover L4 from Jar-A to Jar-B as of 06:52. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the staged-resin measurement from 864 kg to 776 kg, remains consistent with the unchanged 810 kg first-lot requirement, and creates no duplicate or contradictory measurement within that context. Neither context contains an answer code, explicit classifier instruction, rule table, proposition ID, or stated gold label; references to release approval, a nonblocking issue, and monitoring are natural operational facts.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Staging clerk\",\"text\":\"At 06:52, the verified staging log recorded 864 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Production planner\",\"text\":\"At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The 06:52 equipment inspection confirmed that the Jar-B mold was installed. Cleaning record C-441 was signed, and the attendance roster confirmed that all four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality had approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"The earlier stop-work hold was formally cleared before 06:52. The current status board showed no stop-work hold applying to changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery was documented in the issue log and could disrupt continued production. It was classified as a nonblocking issue that did not block line release, with materials support monitoring the delay.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the verified staging log recorded 864 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the verified staging log recorded 864 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "At 06:52, the verified staging log recorded 776 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B.", "right": "At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B."}, "verifier_independent_model": false}, "family": "scale-diverse-294-001", "id": "scale-diverse-294-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Staging clerk", "text": "At 06:52, the verified staging log recorded 776 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Production planner", "text": "At 06:52, the approved lot sheet specified 810 kilograms of resin for the first lot of changeover L4 from Jar-A to Jar-B."}, {"speaker": "Line supervisor", "text": "The 06:52 equipment inspection confirmed that the Jar-B mold was installed. Cleaning record C-441 was signed, and the attendance roster confirmed that all four required operators were present."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality had approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "The earlier stop-work hold was formally cleared before 06:52. The current status board showed no stop-work hold applying to changeover L4."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery was documented in the issue log and could disrupt continued production. It was classified as a nonblocking issue that did not block line release, with materials support monitoring the delay."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original readiness policy and preserve the L4 changeover, 06:52 decision time, and first-lot material path. The two focus-evidence spans are complete factual sentences. The counterfactual changes only staged resin from 704 kg to 651 kg while retaining the 683 kg requirement, creating no contradictory duplicate measurement. Neither context contains an answer code, rule table, proposition identifier, classifier instruction, or explicit gold-answer rationale; the use of natural policy terms such as “nonblocking” is permitted.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Batch control\",\"text\":\"At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin.\"},{\"speaker\":\"Inventory reconciliation\",\"text\":\"At 06:52, a reconciled hopper-and-bin inventory measured 704 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The 06:52 status check confirmed that the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for L4.\"},{\"speaker\":\"Maintenance and quality\",\"text\":\"By 06:52, maintenance had formally cleared the earlier stop-work hold, so none currently applied to L4. Quality had also approved the setup sample for release.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the delayed reserve-resin delivery was recorded as an issue that could disrupt continued production. The delivery delay itself was designated nonblocking and did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin."}, {"path": ["1", "text"], "text": "At 06:52, a reconciled hopper-and-bin inventory measured 704 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin.", "negative_left": "At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin.", "negative_right": "At 06:52, a reconciled hopper-and-bin inventory measured 651 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.", "right": "At 06:52, a reconciled hopper-and-bin inventory measured 704 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, "verifier_independent_model": false}, "family": "scale-diverse-294-002", "id": "scale-diverse-294-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Batch control", "text": "At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin."}, {"speaker": "Inventory reconciliation", "text": "At 06:52, a reconciled hopper-and-bin inventory measured 704 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Line supervisor", "text": "The 06:52 status check confirmed that the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for L4."}, {"speaker": "Maintenance and quality", "text": "By 06:52, maintenance had formally cleared the earlier stop-work hold, so none currently applied to L4. Quality had also approved the setup sample for release."}, {"speaker": "Materials coordinator", "text": "At 06:52, the delayed reserve-resin delivery was recorded as an issue that could disrupt continued production. The delivery delay itself was designated nonblocking and did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original readiness policy and preserve the L4 changeover, 06:52 decision time, and first-lot material path. The two focus-evidence spans are complete factual sentences. The counterfactual changes only staged resin from 704 kg to 651 kg while retaining the 683 kg requirement, creating no contradictory duplicate measurement. Neither context contains an answer code, rule table, proposition identifier, classifier instruction, or explicit gold-answer rationale; the use of natural policy terms such as “nonblocking” is permitted.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Batch control\",\"text\":\"At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin.\"},{\"speaker\":\"Inventory reconciliation\",\"text\":\"At 06:52, a reconciled hopper-and-bin inventory measured 704 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The 06:52 status check confirmed that the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for L4.\"},{\"speaker\":\"Maintenance and quality\",\"text\":\"By 06:52, maintenance had formally cleared the earlier stop-work hold, so none currently applied to L4. Quality had also approved the setup sample for release.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the delayed reserve-resin delivery was recorded as an issue that could disrupt continued production. The delivery delay itself was designated nonblocking and did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin."}, {"path": ["1", "text"], "text": "At 06:52, a reconciled hopper-and-bin inventory measured 704 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin.", "negative_left": "At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin.", "negative_right": "At 06:52, a reconciled hopper-and-bin inventory measured 651 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.", "right": "At 06:52, a reconciled hopper-and-bin inventory measured 704 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, "verifier_independent_model": false}, "family": "scale-diverse-294-002", "id": "scale-diverse-294-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Batch control", "text": "At 06:52, the approved batch sheet for changeover L4 from Jar-A to Jar-B specified that its first lot required 683 kilograms of resin."}, {"speaker": "Inventory reconciliation", "text": "At 06:52, a reconciled hopper-and-bin inventory measured 651 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Line supervisor", "text": "The 06:52 status check confirmed that the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for L4."}, {"speaker": "Maintenance and quality", "text": "By 06:52, maintenance had formally cleared the earlier stop-work hold, so none currently applied to L4. Quality had also approved the setup sample for release."}, {"speaker": "Materials coordinator", "text": "At 06:52, the delayed reserve-resin delivery was recorded as an issue that could disrupt continued production. The delivery delay itself was designated nonblocking and did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original readiness policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the decision’s changeover L4, Jar-A-to-Jar-B, first-lot, and 06:52 bindings. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the staged-resin observation from 204 kilograms to 171 kilograms while retaining the 186-kilogram requirement, creating no contradictory duplicate measurement or conflict with the unchanged facts. Neither context includes a score, proposition ID, rule table, output instruction, or explicit gold answer; references to a documented issue being nonblocking and monitored are permissible factual uses of natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production handoff\",\"text\":\"At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Staging record\",\"text\":\"At 06:52, the staging-scale record lists 204 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The 06:52 line-status check confirms that the Jar-B mold is installed, cleaning record C-441 is signed, and four required operators are present for changeover L4.\"},{\"speaker\":\"Quality and maintenance handoff\",\"text\":\"Quality approved the setup sample for release by 06:52. Maintenance formally cleared the earlier stop-work hold at 06:46, and the handoff shows no current hold on changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"The 06:52 handoff documents the delayed reserve-resin delivery as an issue that could disrupt continued production. It is designated nonblocking, does not block line release, and is assigned to materials support for monitoring.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "At 06:52, the staging-scale record lists 204 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B.", "negative_left": "At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B.", "negative_right": "At 06:52, the staging-scale record lists 171 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.", "right": "At 06:52, the staging-scale record lists 204 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, "verifier_independent_model": false}, "family": "scale-diverse-294-003", "id": "scale-diverse-294-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production handoff", "text": "At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B."}, {"speaker": "Staging record", "text": "At 06:52, the staging-scale record lists 204 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Line supervisor", "text": "The 06:52 line-status check confirms that the Jar-B mold is installed, cleaning record C-441 is signed, and four required operators are present for changeover L4."}, {"speaker": "Quality and maintenance handoff", "text": "Quality approved the setup sample for release by 06:52. Maintenance formally cleared the earlier stop-work hold at 06:46, and the handoff shows no current hold on changeover L4."}, {"speaker": "Materials coordinator", "text": "The 06:52 handoff documents the delayed reserve-resin delivery as an issue that could disrupt continued production. It is designated nonblocking, does not block line release, and is assigned to materials support for monitoring."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original readiness policy without adding exceptions, priorities, or missing-evidence defaults, and they preserve the decision’s changeover L4, Jar-A-to-Jar-B, first-lot, and 06:52 bindings. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the staged-resin observation from 204 kilograms to 171 kilograms while retaining the 186-kilogram requirement, creating no contradictory duplicate measurement or conflict with the unchanged facts. Neither context includes a score, proposition ID, rule table, output instruction, or explicit gold answer; references to a documented issue being nonblocking and monitored are permissible factual uses of natural policy terminology.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production handoff\",\"text\":\"At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Staging record\",\"text\":\"At 06:52, the staging-scale record lists 204 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The 06:52 line-status check confirms that the Jar-B mold is installed, cleaning record C-441 is signed, and four required operators are present for changeover L4.\"},{\"speaker\":\"Quality and maintenance handoff\",\"text\":\"Quality approved the setup sample for release by 06:52. Maintenance formally cleared the earlier stop-work hold at 06:46, and the handoff shows no current hold on changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"The 06:52 handoff documents the delayed reserve-resin delivery as an issue that could disrupt continued production. It is designated nonblocking, does not block line release, and is assigned to materials support for monitoring.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "At 06:52, the staging-scale record lists 204 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B.", "negative_left": "At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B.", "negative_right": "At 06:52, the staging-scale record lists 171 kilograms of resin staged for changeover L4 from Jar-A to Jar-B.", "right": "At 06:52, the staging-scale record lists 204 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, "verifier_independent_model": false}, "family": "scale-diverse-294-003", "id": "scale-diverse-294-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production handoff", "text": "At 06:52, the production handoff record lists 186 kilograms as the resin quantity required for the first lot of changeover L4 from Jar-A to Jar-B."}, {"speaker": "Staging record", "text": "At 06:52, the staging-scale record lists 171 kilograms of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Line supervisor", "text": "The 06:52 line-status check confirms that the Jar-B mold is installed, cleaning record C-441 is signed, and four required operators are present for changeover L4."}, {"speaker": "Quality and maintenance handoff", "text": "Quality approved the setup sample for release by 06:52. Maintenance formally cleared the earlier stop-work hold at 06:46, and the handoff shows no current hold on changeover L4."}, {"speaker": "Materials coordinator", "text": "The 06:52 handoff documents the delayed reserve-resin delivery as an issue that could disrupt continued production. It is designated nonblocking, does not block line release, and is assigned to materials support for monitoring."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete readiness policy, criteria, timestamp, and routing rules; neither context alters them or adds a governing exception. Both contexts remain bound to changeover L4 from Jar-A to Jar-B at 06:52. The two evidence spans are complete factual sentences. The counterfactual changes only the staged-resin measurement from 286 kg to 274 kg, which is coherent with the unchanged 280 kg first-lot requirement and creates no duplicate conflicting measurement. Neither context includes an answer code, rule table, proposition identifier, classifier instruction, or explicit gold-level selection; references to approval, holds, and a nonblocking issue are natural operational facts.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 286 kilograms.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The 06:52 floor check confirmed that the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"The earlier stop-work hold was formally cleared after the failed mold-clamp proximity switch was replaced and three dry cycles succeeded. No current stop-work hold applied to L4 at 06:52.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I approved the setup sample for release after verifying C-441.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"The delayed reserve-resin delivery was documented as an issue at 06:52. It could disrupt continued production and requires monitoring, but it did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 286 kilograms."}, {"path": ["1", "text"], "text": "The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 286 kilograms.", "negative_left": "At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 274 kilograms.", "negative_right": "The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin.", "right": "The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin."}, "verifier_independent_model": false}, "family": "scale-diverse-294-004", "id": "scale-diverse-294-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Materials clerk", "text": "At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 286 kilograms."}, {"speaker": "Production scheduler", "text": "The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin."}, {"speaker": "Line supervisor", "text": "The 06:52 floor check confirmed that the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for changeover L4."}, {"speaker": "Maintenance technician", "text": "The earlier stop-work hold was formally cleared after the failed mold-clamp proximity switch was replaced and three dry cycles succeeded. No current stop-work hold applied to L4 at 06:52."}, {"speaker": "Quality inspector", "text": "At 06:52, I approved the setup sample for release after verifying C-441."}, {"speaker": "Materials coordinator", "text": "The delayed reserve-resin delivery was documented as an issue at 06:52. It could disrupt continued production and requires monitoring, but it did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete readiness policy, criteria, timestamp, and routing rules; neither context alters them or adds a governing exception. Both contexts remain bound to changeover L4 from Jar-A to Jar-B at 06:52. The two evidence spans are complete factual sentences. The counterfactual changes only the staged-resin measurement from 286 kg to 274 kg, which is coherent with the unchanged 280 kg first-lot requirement and creates no duplicate conflicting measurement. Neither context includes an answer code, rule table, proposition identifier, classifier instruction, or explicit gold-level selection; references to approval, holds, and a nonblocking issue are natural operational facts.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 286 kilograms.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The 06:52 floor check confirmed that the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"The earlier stop-work hold was formally cleared after the failed mold-clamp proximity switch was replaced and three dry cycles succeeded. No current stop-work hold applied to L4 at 06:52.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I approved the setup sample for release after verifying C-441.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"The delayed reserve-resin delivery was documented as an issue at 06:52. It could disrupt continued production and requires monitoring, but it did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 286 kilograms."}, {"path": ["1", "text"], "text": "The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 286 kilograms.", "negative_left": "At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 274 kilograms.", "negative_right": "The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin.", "right": "The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin."}, "verifier_independent_model": false}, "family": "scale-diverse-294-004", "id": "scale-diverse-294-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Materials clerk", "text": "At 06:52, the staged resin for changeover L4 from Jar-A to Jar-B weighed 274 kilograms."}, {"speaker": "Production scheduler", "text": "The first lot of changeover L4 from Jar-A to Jar-B required 280 kilograms of resin."}, {"speaker": "Line supervisor", "text": "The 06:52 floor check confirmed that the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for changeover L4."}, {"speaker": "Maintenance technician", "text": "The earlier stop-work hold was formally cleared after the failed mold-clamp proximity switch was replaced and three dry cycles succeeded. No current stop-work hold applied to L4 at 06:52."}, {"speaker": "Quality inspector", "text": "At 06:52, I approved the setup sample for release after verifying C-441."}, {"speaker": "Materials coordinator", "text": "The delayed reserve-resin delivery was documented as an issue at 06:52. It could disrupt continued production and requires monitoring, but it did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the route criteria, instructions, request, stop entity, inspection time, and relevant decision scope. The two focus spans are complete factual sentences: one reports the inspection value and the other defines the condition-field coding without serving as route policy. The counterfactual changes only the recorded condition code from 31 to 68, remains consistent with the unchanged access, address, and recipient facts, and creates no duplicate contradictory measurement. Neither context states the eligible route or embeds classifier labels, proposition IDs, rule tables, or output instructions beyond the preserved task instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"At 14:12, the package inspection recorded code 31 in the binary condition field for the package assigned to stop 42.\",\"In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage.\",\"Dispatch verified the intended destination for stop 42 as 18 Vale Street before the driver arrived.\",\"At 14:12, the designated loading bay there was closed for construction, directly causing the access failure. Site management described the closure as temporary and expected the bay to reopen before the following morning.\",\"The recipient asked for another delivery attempt after 10 tomorrow and expressly declined collection from the depot. Depot operations confirmed that policy did not forbid another attempt for stop 42.\"],\"questions\":{\"decision\":{\"criteria\":{\"address_clarification\":\"Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.\",\"damage_review\":\"Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.\",\"depot_hold\":\"Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.\",\"redelivery\":\"Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt.\"},\"instructions\":\"Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.\",\"type\":\"choice\"}},\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:12, the package inspection recorded code 31 in the binary condition field for the package assigned to stop 42."}, {"path": ["evidence", "1"], "text": "In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12, the package inspection recorded code 31 in the binary condition field for the package assigned to stop 42.", "negative_left": "At 14:12, the package inspection recorded code 68 in the binary condition field for the package assigned to stop 42.", "negative_right": "In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage.", "right": "In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "scale-diverse-295-001", "id": "scale-diverse-295-001-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["At 14:12, the package inspection recorded code 31 in the binary condition field for the package assigned to stop 42.", "In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage.", "Dispatch verified the intended destination for stop 42 as 18 Vale Street before the driver arrived.", "At 14:12, the designated loading bay there was closed for construction, directly causing the access failure. Site management described the closure as temporary and expected the bay to reopen before the following morning.", "The recipient asked for another delivery attempt after 10 tomorrow and expressly declined collection from the depot. Depot operations confirmed that policy did not forbid another attempt for stop 42."], "questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the route criteria, instructions, request, stop entity, inspection time, and relevant decision scope. The two focus spans are complete factual sentences: one reports the inspection value and the other defines the condition-field coding without serving as route policy. The counterfactual changes only the recorded condition code from 31 to 68, remains consistent with the unchanged access, address, and recipient facts, and creates no duplicate contradictory measurement. Neither context states the eligible route or embeds classifier labels, proposition IDs, rule tables, or output instructions beyond the preserved task instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"At 14:12, the package inspection recorded code 31 in the binary condition field for the package assigned to stop 42.\",\"In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage.\",\"Dispatch verified the intended destination for stop 42 as 18 Vale Street before the driver arrived.\",\"At 14:12, the designated loading bay there was closed for construction, directly causing the access failure. Site management described the closure as temporary and expected the bay to reopen before the following morning.\",\"The recipient asked for another delivery attempt after 10 tomorrow and expressly declined collection from the depot. Depot operations confirmed that policy did not forbid another attempt for stop 42.\"],\"questions\":{\"decision\":{\"criteria\":{\"address_clarification\":\"Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.\",\"damage_review\":\"Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.\",\"depot_hold\":\"Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.\",\"redelivery\":\"Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt.\"},\"instructions\":\"Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.\",\"type\":\"choice\"}},\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:12, the package inspection recorded code 31 in the binary condition field for the package assigned to stop 42."}, {"path": ["evidence", "1"], "text": "In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12, the package inspection recorded code 31 in the binary condition field for the package assigned to stop 42.", "negative_left": "At 14:12, the package inspection recorded code 68 in the binary condition field for the package assigned to stop 42.", "negative_right": "In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage.", "right": "In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "scale-diverse-295-001", "id": "scale-diverse-295-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["At 14:12, the package inspection recorded code 68 in the binary condition field for the package assigned to stop 42.", "In the 14:12 package inspection, the binary condition field used code 31 for sound and code 68 for actual-or-suspected damage.", "Dispatch verified the intended destination for stop 42 as 18 Vale Street before the driver arrived.", "At 14:12, the designated loading bay there was closed for construction, directly causing the access failure. Site management described the closure as temporary and expected the bay to reopen before the following morning.", "The recipient asked for another delivery attempt after 10 tomorrow and expressly declined collection from the depot. Depot operations confirmed that policy did not forbid another attempt for stop 42."], "questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same route-policy instruction, request, stop 42 entity, 14:12 inspection binding, and evidence paths. The two focus spans are complete factual sentences: one reports the recorded condition code and the other states the applicable code mapping. Replacing Q7 with M3 is a coherent single-observation change and creates no duplicate or contradictory condition measurement. Neither context states an eligible route, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"The reconciliation file concerns stop 42 and records the intended destination as verified at 18 Vale Street, with no conflicting location details.\",\"The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code Q7.\",\"For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage.\",\"The driver and site manager agreed that a construction closure of the designated loading bay at 18 Vale Street caused the 14:12 access failure.\",\"The site manager confirmed that the bay closure was temporary and would end after the construction work.\",\"During the follow-up call, the recipient asked for another delivery attempt after 10 tomorrow.\",\"The recipient expressly declined collection from the depot and made no depot-collection request.\",\"Operations checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42.\"],\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "1"], "text": "The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code Q7."}, {"path": ["evidence", "2"], "text": "For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code Q7.", "negative_left": "The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code M3.", "negative_right": "For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage.", "right": "For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "scale-diverse-295-002", "id": "scale-diverse-295-002-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["The reconciliation file concerns stop 42 and records the intended destination as verified at 18 Vale Street, with no conflicting location details.", "The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code Q7.", "For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage.", "The driver and site manager agreed that a construction closure of the designated loading bay at 18 Vale Street caused the 14:12 access failure.", "The site manager confirmed that the bay closure was temporary and would end after the construction work.", "During the follow-up call, the recipient asked for another delivery attempt after 10 tomorrow.", "The recipient expressly declined collection from the depot and made no depot-collection request.", "Operations checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42."], "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same route-policy instruction, request, stop 42 entity, 14:12 inspection binding, and evidence paths. The two focus spans are complete factual sentences: one reports the recorded condition code and the other states the applicable code mapping. Replacing Q7 with M3 is a coherent single-observation change and creates no duplicate or contradictory condition measurement. Neither context states an eligible route, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"The reconciliation file concerns stop 42 and records the intended destination as verified at 18 Vale Street, with no conflicting location details.\",\"The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code Q7.\",\"For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage.\",\"The driver and site manager agreed that a construction closure of the designated loading bay at 18 Vale Street caused the 14:12 access failure.\",\"The site manager confirmed that the bay closure was temporary and would end after the construction work.\",\"During the follow-up call, the recipient asked for another delivery attempt after 10 tomorrow.\",\"The recipient expressly declined collection from the depot and made no depot-collection request.\",\"Operations checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42.\"],\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "1"], "text": "The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code Q7."}, {"path": ["evidence", "2"], "text": "For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code Q7.", "negative_left": "The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code M3.", "negative_right": "For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage.", "right": "For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "scale-diverse-295-002", "id": "scale-diverse-295-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["The reconciliation file concerns stop 42 and records the intended destination as verified at 18 Vale Street, with no conflicting location details.", "The binary condition field in the 14:12 inspection record for the package assigned to stop 42 contains code M3.", "For that 14:12 inspection's binary condition field, the complete code table maps Q7 to sound and M3 to actual-or-suspected damage.", "The driver and site manager agreed that a construction closure of the designated loading bay at 18 Vale Street caused the 14:12 access failure.", "The site manager confirmed that the bay closure was temporary and would end after the construction work.", "During the follow-up call, the recipient asked for another delivery attempt after 10 tomorrow.", "The recipient expressly declined collection from the depot and made no depot-collection request.", "Operations checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42."], "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original route-policy instruction, request scope, stop 42, terminal, and 14:12 measurement binding without adding policy exceptions or defaults. The two focus spans are complete factual sentences. The counterfactual changes only the stored condition code from Q7 to Q9; the unchanged legend makes that a suspected-damage observation, which does not contradict the access, address, recipient-request, or policy-check evidence. Neither context states the eligible route, supplies an answer code or rule table, or instructs the classifier which answer to choose.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"At 14:12, inspection terminal IX-604 stored condition code Q7 as the binary result for the package assigned to stop 42.\",\"The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage.\",\"Operational handoff: dispatch and the driver verified the intended destination for stop 42 as 18 Vale Street.\",\"At 14:12, construction barriers closing the designated loading bay at 18 Vale Street caused the access failure. The site supervisor confirmed that the barriers were temporary and scheduled for removal later that day.\",\"Customer service recorded the recipient’s request for another delivery attempt after 10 tomorrow. The recipient expressly declined collection from the depot.\",\"The operations lead checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42.\"],\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:12, inspection terminal IX-604 stored condition code Q7 as the binary result for the package assigned to stop 42."}, {"path": ["evidence", "1"], "text": "The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12, inspection terminal IX-604 stored condition code Q7 as the binary result for the package assigned to stop 42.", "negative_left": "At 14:12, inspection terminal IX-604 stored condition code Q9 as the binary result for the package assigned to stop 42.", "negative_right": "The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage.", "right": "The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "scale-diverse-295-003", "id": "scale-diverse-295-003-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["At 14:12, inspection terminal IX-604 stored condition code Q7 as the binary result for the package assigned to stop 42.", "The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage.", "Operational handoff: dispatch and the driver verified the intended destination for stop 42 as 18 Vale Street.", "At 14:12, construction barriers closing the designated loading bay at 18 Vale Street caused the access failure. The site supervisor confirmed that the barriers were temporary and scheduled for removal later that day.", "Customer service recorded the recipient’s request for another delivery attempt after 10 tomorrow. The recipient expressly declined collection from the depot.", "The operations lead checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42."], "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original route-policy instruction, request scope, stop 42, terminal, and 14:12 measurement binding without adding policy exceptions or defaults. The two focus spans are complete factual sentences. The counterfactual changes only the stored condition code from Q7 to Q9; the unchanged legend makes that a suspected-damage observation, which does not contradict the access, address, recipient-request, or policy-check evidence. Neither context states the eligible route, supplies an answer code or rule table, or instructs the classifier which answer to choose.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"At 14:12, inspection terminal IX-604 stored condition code Q7 as the binary result for the package assigned to stop 42.\",\"The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage.\",\"Operational handoff: dispatch and the driver verified the intended destination for stop 42 as 18 Vale Street.\",\"At 14:12, construction barriers closing the designated loading bay at 18 Vale Street caused the access failure. The site supervisor confirmed that the barriers were temporary and scheduled for removal later that day.\",\"Customer service recorded the recipient’s request for another delivery attempt after 10 tomorrow. The recipient expressly declined collection from the depot.\",\"The operations lead checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42.\"],\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:12, inspection terminal IX-604 stored condition code Q7 as the binary result for the package assigned to stop 42."}, {"path": ["evidence", "1"], "text": "The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12, inspection terminal IX-604 stored condition code Q7 as the binary result for the package assigned to stop 42.", "negative_left": "At 14:12, inspection terminal IX-604 stored condition code Q9 as the binary result for the package assigned to stop 42.", "negative_right": "The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage.", "right": "The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "scale-diverse-295-003", "id": "scale-diverse-295-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["At 14:12, inspection terminal IX-604 stored condition code Q9 as the binary result for the package assigned to stop 42.", "The condition-code legend loaded on inspection terminal IX-604 at 14:12 mapped Q7 exclusively to sound and Q9 exclusively to actual-or-suspected damage.", "Operational handoff: dispatch and the driver verified the intended destination for stop 42 as 18 Vale Street.", "At 14:12, construction barriers closing the designated loading bay at 18 Vale Street caused the access failure. The site supervisor confirmed that the barriers were temporary and scheduled for removal later that day.", "Customer service recorded the recipient’s request for another delivery attempt after 10 tomorrow. The recipient expressly declined collection from the depot.", "The operations lead checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42."], "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all route criteria and instructions, while no governing policy from the original state is lost or altered. Both contexts retain the same stop, package, inspection time, destination, request, and decision scope. The two focus spans are complete factual sentences: one reports the observed code and the other supplies the complete codebook interpretation. The counterfactual changes only the recorded code from 6B to 9F, which remains coherent with the codebook and does not duplicate or contradict another measurement. Neither context states the eligible route, supplies an answer label, or instructs the classifier which choice to make.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"At 14:12, inspection unit QN-73 recorded condition code 6B for the package assigned to stop 42.\",\"For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage.\",\"Dispatch and the recipient independently confirmed that the intended destination for stop 42 was 18 Vale Street.\",\"The driver and site supervisor confirmed that construction had closed the designated loading bay at 18 Vale Street, directly causing the 14:12 access failure. The supervisor described the closure as temporary and expected the bay to reopen tomorrow morning.\",\"The recipient requested another delivery attempt after 10 tomorrow and expressly declined collection from the depot.\",\"Depot operations checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42.\"],\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:12, inspection unit QN-73 recorded condition code 6B for the package assigned to stop 42."}, {"path": ["evidence", "1"], "text": "For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12, inspection unit QN-73 recorded condition code 6B for the package assigned to stop 42.", "negative_left": "At 14:12, inspection unit QN-73 recorded condition code 9F for the package assigned to stop 42.", "negative_right": "For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage.", "right": "For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "scale-diverse-295-004", "id": "scale-diverse-295-004-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["At 14:12, inspection unit QN-73 recorded condition code 6B for the package assigned to stop 42.", "For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage.", "Dispatch and the recipient independently confirmed that the intended destination for stop 42 was 18 Vale Street.", "The driver and site supervisor confirmed that construction had closed the designated loading bay at 18 Vale Street, directly causing the 14:12 access failure. The supervisor described the closure as temporary and expected the bay to reopen tomorrow morning.", "The recipient requested another delivery attempt after 10 tomorrow and expressly declined collection from the depot.", "Depot operations checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42."], "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all route criteria and instructions, while no governing policy from the original state is lost or altered. Both contexts retain the same stop, package, inspection time, destination, request, and decision scope. The two focus spans are complete factual sentences: one reports the observed code and the other supplies the complete codebook interpretation. The counterfactual changes only the recorded code from 6B to 9F, which remains coherent with the codebook and does not duplicate or contradict another measurement. Neither context states the eligible route, supplies an answer label, or instructs the classifier which choice to make.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"At 14:12, inspection unit QN-73 recorded condition code 6B for the package assigned to stop 42.\",\"For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage.\",\"Dispatch and the recipient independently confirmed that the intended destination for stop 42 was 18 Vale Street.\",\"The driver and site supervisor confirmed that construction had closed the designated loading bay at 18 Vale Street, directly causing the 14:12 access failure. The supervisor described the closure as temporary and expected the bay to reopen tomorrow morning.\",\"The recipient requested another delivery attempt after 10 tomorrow and expressly declined collection from the depot.\",\"Depot operations checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42.\"],\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 14:12, inspection unit QN-73 recorded condition code 6B for the package assigned to stop 42."}, {"path": ["evidence", "1"], "text": "For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12, inspection unit QN-73 recorded condition code 6B for the package assigned to stop 42.", "negative_left": "At 14:12, inspection unit QN-73 recorded condition code 9F for the package assigned to stop 42.", "negative_right": "For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage.", "right": "For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "scale-diverse-295-004", "id": "scale-diverse-295-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["At 14:12, inspection unit QN-73 recorded condition code 9F for the package assigned to stop 42.", "For the 14:12 inspection of the package assigned to stop 42, unit QN-73's complete binary codebook assigned 6B to sound and 9F to actual-or-suspected damage.", "Dispatch and the recipient independently confirmed that the intended destination for stop 42 was 18 Vale Street.", "The driver and site supervisor confirmed that construction had closed the designated loading bay at 18 Vale Street, directly causing the 14:12 access failure. The supervisor described the closure as temporary and expected the bay to reopen tomorrow morning.", "The recipient requested another delivery attempt after 10 tomorrow and expressly declined collection from the depot.", "Depot operations checked the applicable policy and confirmed that it did not forbid another delivery attempt for stop 42."], "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all route criteria and instructions, while both contexts retain the original policy application and synonym-equivalence instruction. The stop, package, destination, date/time frame, and requested route decision remain bound consistently. The two focus spans are complete factual sentences describing the recorded code and its contemporaneous legend. The counterfactual coherently changes only the legend’s meaning, making code 731 indicate damage rather than sound without creating a duplicate or conflicting package-condition assertion. Neither context supplies a route answer, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"At 13:50 on 17 September 2026, routing control verified that stop 42’s intended destination was 18 Vale Street.\",\"At 14:12, the driver documented that construction had closed the designated loading bay at 18 Vale Street, causing the access failure for stop 42. The site foreman said the bay would reopen the next morning, confirming the closure was temporary.\",\"At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42.\",\"The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “sound” and code 846 exclusively to “actual or suspected damage.”\",\"At 14:18, the recipient asked for another delivery attempt after 10 tomorrow and expressly declined collection from the depot.\",\"At 14:25, depot operations confirmed that no policy provision forbade another delivery attempt for stop 42.\"],\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42."}, {"path": ["evidence", "3"], "text": "The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “sound” and code 846 exclusively to “actual or suspected damage.”"}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42.", "negative_left": "At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42.", "negative_right": "The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “actual or suspected damage” and code 846 exclusively to “sound.”", "right": "The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “sound” and code 846 exclusively to “actual or suspected damage.”"}, "verifier_independent_model": false}, "family": "scale-diverse-295-005", "id": "scale-diverse-295-005-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["At 13:50 on 17 September 2026, routing control verified that stop 42’s intended destination was 18 Vale Street.", "At 14:12, the driver documented that construction had closed the designated loading bay at 18 Vale Street, causing the access failure for stop 42. The site foreman said the bay would reopen the next morning, confirming the closure was temporary.", "At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42.", "The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “sound” and code 846 exclusively to “actual or suspected damage.”", "At 14:18, the recipient asked for another delivery attempt after 10 tomorrow and expressly declined collection from the depot.", "At 14:25, depot operations confirmed that no policy provision forbade another delivery attempt for stop 42."], "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all route criteria and instructions, while both contexts retain the original policy application and synonym-equivalence instruction. The stop, package, destination, date/time frame, and requested route decision remain bound consistently. The two focus spans are complete factual sentences describing the recorded code and its contemporaneous legend. The counterfactual coherently changes only the legend’s meaning, making code 731 indicate damage rather than sound without creating a duplicate or conflicting package-condition assertion. Neither context supplies a route answer, answer code, rule table, proposition identifier, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"At 13:50 on 17 September 2026, routing control verified that stop 42’s intended destination was 18 Vale Street.\",\"At 14:12, the driver documented that construction had closed the designated loading bay at 18 Vale Street, causing the access failure for stop 42. The site foreman said the bay would reopen the next morning, confirming the closure was temporary.\",\"At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42.\",\"The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “sound” and code 846 exclusively to “actual or suspected damage.”\",\"At 14:18, the recipient asked for another delivery attempt after 10 tomorrow and expressly declined collection from the depot.\",\"At 14:25, depot operations confirmed that no policy provision forbade another delivery attempt for stop 42.\"],\"request\":\"Which single exception route is eligible now?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42."}, {"path": ["evidence", "3"], "text": "The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “sound” and code 846 exclusively to “actual or suspected damage.”"}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42.", "negative_left": "At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42.", "negative_right": "The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “actual or suspected damage” and code 846 exclusively to “sound.”", "right": "The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “sound” and code 846 exclusively to “actual or suspected damage.”"}, "verifier_independent_model": false}, "family": "scale-diverse-295-005", "id": "scale-diverse-295-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["At 13:50 on 17 September 2026, routing control verified that stop 42’s intended destination was 18 Vale Street.", "At 14:12, the driver documented that construction had closed the designated loading bay at 18 Vale Street, causing the access failure for stop 42. The site foreman said the bay would reopen the next morning, confirming the closure was temporary.", "At 14:12 on 17 September 2026, inspector Mara Venn entered code 731 in the binary condition field of the completed inspection record for the package assigned to stop 42.", "The contemporaneous legend for the 14:12 package inspection on 17 September 2026 mapped code 731 exclusively to “actual or suspected damage” and code 846 exclusively to “sound.”", "At 14:18, the recipient asked for another delivery attempt after 10 tomorrow and expressly declined collection from the depot.", "At 14:25, depot operations confirmed that no policy provision forbade another delivery attempt for stop 42."], "request": "Which single exception route is eligible now?"}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing routing policy while the unchanged questions preserve all criteria and instructions. The Stop 42 package, recorded-evidence scope, and decision path remain fixed; the only material counterfactual change is E731’s crush-depth measurement from 2.7 cm to 1.6 cm. The two focus-evidence spans are complete factual sentences, the changed measurement does not conflict with any unchanged assertion, and neither context states a case-specific answer, answer code, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Field recorder\",\"text\":\"As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731.\"},{\"speaker\":\"Inspection log\",\"text\":\"Inspection entry E731 reports a measured crush depth of 2.7 cm.\"},{\"speaker\":\"Package check\",\"text\":\"The seal and exterior show no leak or opening, and the handling check produced no rattling.\"},{\"speaker\":\"Route file\",\"text\":\"The delivery address is verified, usable, and free of conflicts. Complete, usable lift-gate access details are on file, and no vehicle restriction prevents service. This is the first attempt, so fewer than two attempts have failed. The recipient has confirmed availability tomorrow, making another attempt feasible.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731."}, {"path": ["1", "text"], "text": "Inspection entry E731 reports a measured crush depth of 2.7 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731.", "negative_left": "As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731.", "negative_right": "Inspection entry E731 reports a measured crush depth of 1.6 cm.", "right": "Inspection entry E731 reports a measured crush depth of 2.7 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-296-004", "id": "scale-diverse-296-004-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Field recorder", "text": "As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731."}, {"speaker": "Inspection log", "text": "Inspection entry E731 reports a measured crush depth of 2.7 cm."}, {"speaker": "Package check", "text": "The seal and exterior show no leak or opening, and the handling check produced no rattling."}, {"speaker": "Route file", "text": "The delivery address is verified, usable, and free of conflicts. Complete, usable lift-gate access details are on file, and no vehicle restriction prevents service. This is the first attempt, so fewer than two attempts have failed. The recipient has confirmed availability tomorrow, making another attempt feasible."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing routing policy while the unchanged questions preserve all criteria and instructions. The Stop 42 package, recorded-evidence scope, and decision path remain fixed; the only material counterfactual change is E731’s crush-depth measurement from 2.7 cm to 1.6 cm. The two focus-evidence spans are complete factual sentences, the changed measurement does not conflict with any unchanged assertion, and neither context states a case-specific answer, answer code, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Field recorder\",\"text\":\"As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731.\"},{\"speaker\":\"Inspection log\",\"text\":\"Inspection entry E731 reports a measured crush depth of 2.7 cm.\"},{\"speaker\":\"Package check\",\"text\":\"The seal and exterior show no leak or opening, and the handling check produced no rattling.\"},{\"speaker\":\"Route file\",\"text\":\"The delivery address is verified, usable, and free of conflicts. Complete, usable lift-gate access details are on file, and no vehicle restriction prevents service. This is the first attempt, so fewer than two attempts have failed. The recipient has confirmed availability tomorrow, making another attempt feasible.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731."}, {"path": ["1", "text"], "text": "Inspection entry E731 reports a measured crush depth of 2.7 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731.", "negative_left": "As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731.", "negative_right": "Inspection entry E731 reports a measured crush depth of 1.6 cm.", "right": "Inspection entry E731 reports a measured crush depth of 2.7 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-296-004", "id": "scale-diverse-296-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Field recorder", "text": "As of 2026-09-17 at 14:20 UTC, the complete recorded crush-depth evidence for the package assigned to Stop 42 consisted solely of inspection entry E731."}, {"speaker": "Inspection log", "text": "Inspection entry E731 reports a measured crush depth of 1.6 cm."}, {"speaker": "Package check", "text": "The seal and exterior show no leak or opening, and the handling check produced no rattling."}, {"speaker": "Route file", "text": "The delivery address is verified, usable, and free of conflicts. Complete, usable lift-gate access details are on file, and no vehicle restriction prevents service. This is the first attempt, so fewer than two attempts have failed. The recipient has confirmed availability tomorrow, making another attempt feasible."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing first-match routing policy from the original state, while the unchanged questions preserve all criteria and instructions. Stop 42 remains the relevant entity, the two evidence spans are complete factual sentences, and the counterfactual coherently changes only the recorded crush depth from 2.6 cm to 1.6 cm without conflicting measurements. Neither context contains a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Dispatch log\",\"text\":\"At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route.\"},{\"speaker\":\"Inspection log\",\"text\":\"At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 2.6 cm.\"},{\"speaker\":\"Inspection follow-up\",\"text\":\"At 08:42, the same inspection found the seal intact and recorded no leak, opening, or rattling.\"},{\"speaker\":\"Route desk\",\"text\":\"At 09:00, the address on file was confirmed as verified, usable, and free of conflicting entries. Complete lift-gate and entry instructions were present and usable, with no required access detail missing. No vehicle restriction applied. No delivery attempt had yet failed, and dispatch confirmed that another attempt was feasible.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route."}, {"path": ["1", "text"], "text": "At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 2.6 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route.", "negative_left": "At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route.", "negative_right": "At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 1.6 cm.", "right": "At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 2.6 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-296-005", "id": "scale-diverse-296-005-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Dispatch log", "text": "At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route."}, {"speaker": "Inspection log", "text": "At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 2.6 cm."}, {"speaker": "Inspection follow-up", "text": "At 08:42, the same inspection found the seal intact and recorded no leak, opening, or rattling."}, {"speaker": "Route desk", "text": "At 09:00, the address on file was confirmed as verified, usable, and free of conflicting entries. Complete lift-gate and entry instructions were present and usable, with no required access detail missing. No vehicle restriction applied. No delivery attempt had yet failed, and dispatch confirmed that another attempt was feasible."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing first-match routing policy from the original state, while the unchanged questions preserve all criteria and instructions. Stop 42 remains the relevant entity, the two evidence spans are complete factual sentences, and the counterfactual coherently changes only the recorded crush depth from 2.6 cm to 1.6 cm without conflicting measurements. Neither context contains a gold answer, output instruction, answer code, proposition ID, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Dispatch log\",\"text\":\"At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route.\"},{\"speaker\":\"Inspection log\",\"text\":\"At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 2.6 cm.\"},{\"speaker\":\"Inspection follow-up\",\"text\":\"At 08:42, the same inspection found the seal intact and recorded no leak, opening, or rattling.\"},{\"speaker\":\"Route desk\",\"text\":\"At 09:00, the address on file was confirmed as verified, usable, and free of conflicting entries. Complete lift-gate and entry instructions were present and usable, with no required access detail missing. No vehicle restriction applied. No delivery attempt had yet failed, and dispatch confirmed that another attempt was feasible.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route."}, {"path": ["1", "text"], "text": "At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 2.6 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route.", "negative_left": "At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route.", "negative_right": "At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 1.6 cm.", "right": "At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 2.6 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-296-005", "id": "scale-diverse-296-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Dispatch log", "text": "At 08:10 on 17 September 2026, dispatch record DR-581 assigned package QX-704 to Stop 42 for that day's route."}, {"speaker": "Inspection log", "text": "At 08:35 on 17 September 2026, inspection record IR-926 documented package QX-704's crush depth as 1.6 cm."}, {"speaker": "Inspection follow-up", "text": "At 08:42, the same inspection found the seal intact and recorded no leak, opening, or rattling."}, {"speaker": "Route desk", "text": "At 09:00, the address on file was confirmed as verified, usable, and free of conflicting entries. Complete lift-gate and entry instructions were present and usable, with no required access detail missing. No vehicle restriction applied. No delivery attempt had yet failed, and dispatch confirmed that another attempt was feasible."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing first-match routing policy when assessed together with the unchanged questions object. The route, stop, package, and date bindings remain consistent across the base and counterfactual contexts; the focus evidence consists of exactly two complete factual sentences; and the sole counterfactual change from a 2.6 cm crush depth to 1.6 cm creates no contradiction within its context. Neither context embeds a gold answer, output instruction, answer code, proposition ID, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Manifest custodian\",\"text\":\"The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42.\"},{\"speaker\":\"Inspection technician\",\"text\":\"The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 2.6 cm.\"},{\"speaker\":\"Evidence reconciler\",\"text\":\"The same inspection records that the seal is intact: the package has no leak or opening and does not rattle. Dispatch records show a verified, nonconflicting delivery address that is usable. Complete lift-gate access details are recorded and usable, with no required information missing. No delivery attempt has failed, no vehicle restriction prevents service, and the recipient’s confirmed availability tomorrow makes another attempt feasible.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42."}, {"path": ["1", "text"], "text": "The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 2.6 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42.", "negative_left": "The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42.", "negative_right": "The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 1.6 cm.", "right": "The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 2.6 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-296-006", "id": "scale-diverse-296-006-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Manifest custodian", "text": "The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42."}, {"speaker": "Inspection technician", "text": "The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 2.6 cm."}, {"speaker": "Evidence reconciler", "text": "The same inspection records that the seal is intact: the package has no leak or opening and does not rattle. Dispatch records show a verified, nonconflicting delivery address that is usable. Complete lift-gate access details are recorded and usable, with no required information missing. No delivery attempt has failed, no vehicle restriction prevents service, and the recipient’s confirmed availability tomorrow makes another attempt feasible."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing first-match routing policy when assessed together with the unchanged questions object. The route, stop, package, and date bindings remain consistent across the base and counterfactual contexts; the focus evidence consists of exactly two complete factual sentences; and the sole counterfactual change from a 2.6 cm crush depth to 1.6 cm creates no contradiction within its context. Neither context embeds a gold answer, output instruction, answer code, proposition ID, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Manifest custodian\",\"text\":\"The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42.\"},{\"speaker\":\"Inspection technician\",\"text\":\"The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 2.6 cm.\"},{\"speaker\":\"Evidence reconciler\",\"text\":\"The same inspection records that the seal is intact: the package has no leak or opening and does not rattle. Dispatch records show a verified, nonconflicting delivery address that is usable. Complete lift-gate access details are recorded and usable, with no required information missing. No delivery attempt has failed, no vehicle restriction prevents service, and the recipient’s confirmed availability tomorrow makes another attempt feasible.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42."}, {"path": ["1", "text"], "text": "The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 2.6 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42.", "negative_left": "The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42.", "negative_right": "The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 1.6 cm.", "right": "The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 2.6 cm."}, "verifier_independent_model": false}, "family": "scale-diverse-296-006", "id": "scale-diverse-296-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Manifest custodian", "text": "The finalized manifest for route R-619 on 2026-09-17 identifies package PK-734 as the sole package assigned to Stop 42."}, {"speaker": "Inspection technician", "text": "The calibrated depth gauge's inspection record dated 2026-09-17 records the crush depth of package PK-734 as 1.6 cm."}, {"speaker": "Evidence reconciler", "text": "The same inspection records that the seal is intact: the package has no leak or opening and does not rattle. Dispatch records show a verified, nonconflicting delivery address that is usable. Complete lift-gate access details are recorded and usable, with no required information missing. No delivery attempt has failed, no vehicle restriction prevents service, and the recipient’s confirmed availability tomorrow makes another attempt feasible."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same stop instruction, exception condition, parcel, stop, recipient-unavailability event, dispatcher, and routing-decision time. The only material change is the alternate-delivery request timestamp, from 15:46 to 15:54, which is coherent with the unchanged 15:50 decision and the record extending through 16:00; it creates no duplicate or contradictory request count. The two evidence spans are complete factual sentences. Neither context states the required yes/no output, a gold label, an answer code, or a label rationale; the repeated routing language is the governing policy itself.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Driver Joel’s return paperwork identified parcel FM-297 and Stop 18; the address was complete, the sealed carton was undamaged, and depot operations had secure hold space available. At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day. The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:46 that day. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day."}, {"path": [], "text": "The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:46 that day."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day.", "negative_left": "At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day.", "negative_right": "The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:54 that day.", "right": "The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:46 that day."}, "verifier_independent_model": false}, "family": "scale-diverse-297-001", "id": "scale-diverse-297-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Driver Joel’s return paperwork identified parcel FM-297 and Stop 18; the address was complete, the sealed carton was undamaged, and depot operations had secure hold space available. At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day. The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:46 that day. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same stop instruction, exception condition, parcel, stop, recipient-unavailability event, dispatcher, and routing-decision time. The only material change is the alternate-delivery request timestamp, from 15:46 to 15:54, which is coherent with the unchanged 15:50 decision and the record extending through 16:00; it creates no duplicate or contradictory request count. The two evidence spans are complete factual sentences. Neither context states the required yes/no output, a gold label, an answer code, or a label rationale; the repeated routing language is the governing policy itself.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Driver Joel’s return paperwork identified parcel FM-297 and Stop 18; the address was complete, the sealed carton was undamaged, and depot operations had secure hold space available. At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day. The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:46 that day. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day."}, {"path": [], "text": "The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:46 that day."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day.", "negative_left": "At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day.", "negative_right": "The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:54 that day.", "right": "The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:46 that day."}, "verifier_independent_model": false}, "family": "scale-diverse-297-001", "id": "scale-diverse-297-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Driver Joel’s return paperwork identified parcel FM-297 and Stop 18; the address was complete, the sealed carton was undamaged, and depot operations had secure hold space available. At the 15:42 Stop 18 delivery attempt for parcel FM-297 on 12 May 2026, the named recipient was unavailable, and dispatcher Mina recorded her routing decision at 15:50 that day. The complete lifetime event record for parcel FM-297 through 16:00 on 12 May 2026 lists exactly one alternate-delivery request, received at 15:54 that day. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged question preserves the governing exception rule, while both contexts retain the same parcel, recipient-unavailability event, routing decision, and relevant pre-decision request path. The two evidence spans are complete factual sentences; the counterfactual coherently replaces the alternate-delivery request with an exhaustive ledger assertion that all pre-decision requests were address confirmations, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Field note for parcel FM-297: At the 15:42 Stop 18 delivery attempt, the named recipient was unavailable. Driver Joel reported that nobody answered and no authorized neighbor was present. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant. The sealed intake ledger records receipt of a request categorized as alternate delivery for parcel FM-297 at 15:51 on 14 May 2026. Depot operations lead Priya confirmed that secure hold space was available, while intake staff found the address complete and the sealed carton undamaged.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant."}, {"path": [], "text": "The sealed intake ledger records receipt of a request categorized as alternate delivery for parcel FM-297 at 15:51 on 14 May 2026."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant.", "negative_left": "Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant.", "negative_right": "Every request for parcel FM-297 recorded in the sealed intake ledger before 16:08 on 14 May 2026 is categorized as address confirmation rather than alternate delivery.", "right": "The sealed intake ledger records receipt of a request categorized as alternate delivery for parcel FM-297 at 15:51 on 14 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-297-004", "id": "scale-diverse-297-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Field note for parcel FM-297: At the 15:42 Stop 18 delivery attempt, the named recipient was unavailable. Driver Joel reported that nobody answered and no authorized neighbor was present. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant. The sealed intake ledger records receipt of a request categorized as alternate delivery for parcel FM-297 at 15:51 on 14 May 2026. Depot operations lead Priya confirmed that secure hold space was available, while intake staff found the address complete and the sealed carton undamaged."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged question preserves the governing exception rule, while both contexts retain the same parcel, recipient-unavailability event, routing decision, and relevant pre-decision request path. The two evidence spans are complete factual sentences; the counterfactual coherently replaces the alternate-delivery request with an exhaustive ledger assertion that all pre-decision requests were address confirmations, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Field note for parcel FM-297: At the 15:42 Stop 18 delivery attempt, the named recipient was unavailable. Driver Joel reported that nobody answered and no authorized neighbor was present. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant. The sealed intake ledger records receipt of a request categorized as alternate delivery for parcel FM-297 at 15:51 on 14 May 2026. Depot operations lead Priya confirmed that secure hold space was available, while intake staff found the address complete and the sealed carton undamaged.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant."}, {"path": [], "text": "The sealed intake ledger records receipt of a request categorized as alternate delivery for parcel FM-297 at 15:51 on 14 May 2026."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant.", "negative_left": "Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant.", "negative_right": "Every request for parcel FM-297 recorded in the sealed intake ledger before 16:08 on 14 May 2026 is categorized as address confirmation rather than alternate delivery.", "right": "The sealed intake ledger records receipt of a request categorized as alternate delivery for parcel FM-297 at 15:51 on 14 May 2026."}, "verifier_independent_model": false}, "family": "scale-diverse-297-004", "id": "scale-diverse-297-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Field note for parcel FM-297: At the 15:42 Stop 18 delivery attempt, the named recipient was unavailable. Driver Joel reported that nobody answered and no authorized neighbor was present. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Dispatcher Mina recorded her routing decision for parcel FM-297 at 16:08 on 14 May 2026, and the sealed intake ledger enumerates every request for FM-297 received before that instant. Every request for parcel FM-297 recorded in the sealed intake ledger before 16:08 on 14 May 2026 is categorized as address confirmation rather than alternate delivery. Depot operations lead Priya confirmed that secure hold space was available, while intake staff found the address complete and the sealed carton undamaged."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the depot policy, dispatcher condition, and direct-handoff restriction when considered with the unchanged questions object. The request and the package, route, address, and incident bindings remain unchanged. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes Nia Calder from on-duty building staff to a visiting utility contractor without contradicting the body-camera statement, which identifies her only as a badge holder, and it introduces no duplicate conflicting measurement or count. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"building_staff_confirmed_move": "supported", "instruction_applies": "supported"}, "full_context_fact_states": {"base": {"building_staff_confirmed_move": "supported", "instruction_applies": "supported"}, "counterfactual": {"building_staff_confirmed_move": "refuted", "instruction_applies": "supported"}, "remove_left": {"building_staff_confirmed_move": "unknown"}, "remove_right": {"building_staff_confirmed_move": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"building_staff_confirmed_move": "unknown"}, "negative_pair": {"building_staff_confirmed_move": "refuted"}, "negative_sentence": {"building_staff_confirmed_move": "unknown"}, "positive_pair": {"building_staff_confirmed_move": "supported"}, "right": {"building_staff_confirmed_move": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express individual factual relationships. The focus concerns whether building staff confirmed the move, not a policy conclusion. The base and counter assignments can be realized while keeping the instruction applicable and changing only whether the required confirmation occurred. Policy evidence preserves the substantive state-originating depot rule and the dispatcher’s conditional instruction; instructions and criteria in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The rule requires both that the dispatcher’s conditional instruction applies and that building staff confirmed the specified move from Unit 4B. Under the preserved depot policy and instruction, this is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the building-staff confirmation entails that the dispatcher’s stated condition was not confirmed. The unchanged question explicitly requires the false outcome when that condition was not confirmed, so no additional ordinary-redelivery facts are necessary.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "instruction_applies", "statement": "For package ZX-418 on Route 12 at the unsuccessful-delivery incident for 8 Harbor Lane, Unit 4B, the route dispatcher’s conditional instruction applies."}, {"id": "building_staff_confirmed_move", "statement": "At the unsuccessful-delivery incident for package ZX-418, a building staff member confirmed to the delivery driver that the recipient had moved from Unit 4B."}], "base_state_json": "{\"context\":\"Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery.\",\"evidence\":[\"Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”\",\"The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B.\",\"At that incident, Nia Calder was an on-duty building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied.\",\"Field note: The Route 12 scan log identifies ZX-418, records arrival at 8 Harbor Lane, Unit 4B, and marks the attempt unsuccessful with no completion scan. The package was returned intact after the recipient could not be reached. The stop required direct handoff and did not permit leaving the package with another person.\"],\"request\":\"Under depot policy, should this exception be routed to address clarification rather than redelivery?\"}", "base_states": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "supported"}], "counter_states": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "refuted"}], "focus_atom": "building_staff_confirmed_move", "focus_evidence": [{"path": ["evidence", "1"], "text": "The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B."}, {"path": ["evidence", "2"], "text": "At that incident, Nia Calder was an on-duty building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied."}], "policy_evidence": [{"path": ["request"], "text": "Under depot policy, should this exception be routed to address clarification rather than redelivery?"}, {"path": ["context"], "text": "Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery."}, {"path": ["evidence", "0"], "text": "Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”"}], "rules": [{"justification": "The applicable dispatcher instruction’s stated address-conflict condition is confirmed, so depot policy routes the exception to address clarification.", "target": "true", "when": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "supported"}]}, {"justification": "The applicable dispatcher instruction requires confirmation by building staff that the recipient moved from Unit 4B; explicit refutation of that confirmation means the condition was not confirmed.", "target": "false", "when": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B.", "negative_left": "The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B.", "negative_right": "At that incident, Nia Calder was a visiting utility contractor rather than a building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied.", "right": "At that incident, Nia Calder was an on-duty building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied."}, "verifier_independent_model": false}, "family": "scale-diverse-298-004", "id": "scale-diverse-298-004-base", "input": {"questions": {"decision": {"criteria": {"false": "Route the incident for redelivery because the dispatcher’s address-conflict condition was not confirmed.", "true": "Route the incident to address clarification because the condition in the dispatcher’s instruction was confirmed."}, "instructions": "Answer yes if the dispatcher’s conditional address-conflict instruction was activated by confirmed evidence. Answer no if that condition was not confirmed or if the record instead supports ordinary redelivery.", "type": "noul"}}, "state": {"context": "Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery.", "evidence": ["Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”", "The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B.", "At that incident, Nia Calder was an on-duty building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied.", "Field note: The Route 12 scan log identifies ZX-418, records arrival at 8 Harbor Lane, Unit 4B, and marks the attempt unsuccessful with no completion scan. The package was returned intact after the recipient could not be reached. The stop required direct handoff and did not permit leaving the package with another person."], "request": "Under depot policy, should this exception be routed to address clarification rather than redelivery?"}}, "method": "c2d", "provenance": {"source_id": "diverse-298", "source_is_synthetic": true, "source_sha256": "d5c7c6aa5962d6762278e5ae977fe87f27888823aa0a70204e56364a3bea30d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the depot policy, dispatcher condition, and direct-handoff restriction when considered with the unchanged questions object. The request and the package, route, address, and incident bindings remain unchanged. The two focus-evidence spans are complete factual sentences. The counterfactual coherently changes Nia Calder from on-duty building staff to a visiting utility contractor without contradicting the body-camera statement, which identifies her only as a badge holder, and it introduces no duplicate conflicting measurement or count. Neither context contains a gold answer, output instruction, answer code, rule table, proposition ID, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"building_staff_confirmed_move": "refuted", "instruction_applies": "supported"}, "full_context_fact_states": {"base": {"building_staff_confirmed_move": "supported", "instruction_applies": "supported"}, "counterfactual": {"building_staff_confirmed_move": "refuted", "instruction_applies": "supported"}, "remove_left": {"building_staff_confirmed_move": "unknown"}, "remove_right": {"building_staff_confirmed_move": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"building_staff_confirmed_move": "unknown"}, "negative_pair": {"building_staff_confirmed_move": "refuted"}, "negative_sentence": {"building_staff_confirmed_move": "unknown"}, "positive_pair": {"building_staff_confirmed_move": "supported"}, "right": {"building_staff_confirmed_move": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express individual factual relationships. The focus concerns whether building staff confirmed the move, not a policy conclusion. The base and counter assignments can be realized while keeping the instruction applicable and changing only whether the required confirmation occurred. Policy evidence preserves the substantive state-originating depot rule and the dispatcher’s conditional instruction; instructions and criteria in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The rule requires both that the dispatcher’s conditional instruction applies and that building staff confirmed the specified move from Unit 4B. Under the preserved depot policy and instruction, this is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the building-staff confirmation entails that the dispatcher’s stated condition was not confirmed. The unchanged question explicitly requires the false outcome when that condition was not confirmed, so no additional ordinary-redelivery facts are necessary.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "instruction_applies", "statement": "For package ZX-418 on Route 12 at the unsuccessful-delivery incident for 8 Harbor Lane, Unit 4B, the route dispatcher’s conditional instruction applies."}, {"id": "building_staff_confirmed_move", "statement": "At the unsuccessful-delivery incident for package ZX-418, a building staff member confirmed to the delivery driver that the recipient had moved from Unit 4B."}], "base_state_json": "{\"context\":\"Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery.\",\"evidence\":[\"Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”\",\"The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B.\",\"At that incident, Nia Calder was an on-duty building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied.\",\"Field note: The Route 12 scan log identifies ZX-418, records arrival at 8 Harbor Lane, Unit 4B, and marks the attempt unsuccessful with no completion scan. The package was returned intact after the recipient could not be reached. The stop required direct handoff and did not permit leaving the package with another person.\"],\"request\":\"Under depot policy, should this exception be routed to address clarification rather than redelivery?\"}", "base_states": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "supported"}], "counter_states": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "refuted"}], "focus_atom": "building_staff_confirmed_move", "focus_evidence": [{"path": ["evidence", "1"], "text": "The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B."}, {"path": ["evidence", "2"], "text": "At that incident, Nia Calder was an on-duty building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied."}], "policy_evidence": [{"path": ["request"], "text": "Under depot policy, should this exception be routed to address clarification rather than redelivery?"}, {"path": ["context"], "text": "Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery."}, {"path": ["evidence", "0"], "text": "Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”"}], "rules": [{"justification": "The applicable dispatcher instruction’s stated address-conflict condition is confirmed, so depot policy routes the exception to address clarification.", "target": "true", "when": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "supported"}]}, {"justification": "The applicable dispatcher instruction requires confirmation by building staff that the recipient moved from Unit 4B; explicit refutation of that confirmation means the condition was not confirmed.", "target": "false", "when": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B.", "negative_left": "The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B.", "negative_right": "At that incident, Nia Calder was a visiting utility contractor rather than a building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied.", "right": "At that incident, Nia Calder was an on-duty building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied."}, "verifier_independent_model": false}, "family": "scale-diverse-298-004", "id": "scale-diverse-298-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Route the incident for redelivery because the dispatcher’s address-conflict condition was not confirmed.", "true": "Route the incident to address clarification because the condition in the dispatcher’s instruction was confirmed."}, "instructions": "Answer yes if the dispatcher’s conditional address-conflict instruction was activated by confirmed evidence. Answer no if that condition was not confirmed or if the record instead supports ordinary redelivery.", "type": "noul"}}, "state": {"context": "Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery.", "evidence": ["Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”", "The complete body-camera record of the unsuccessful-delivery incident for package ZX-418 at 8 Harbor Lane, Unit 4B, shows that badge holder Nia Calder was the only person who communicated information to the delivery driver and that she said the recipient had moved from Unit 4B.", "At that incident, Nia Calder was a visiting utility contractor rather than a building staff member, and the route dispatcher's conditional instruction for package ZX-418 on Route 12 applied.", "Field note: The Route 12 scan log identifies ZX-418, records arrival at 8 Harbor Lane, Unit 4B, and marks the attempt unsuccessful with no completion scan. The package was returned intact after the recipient could not be reached. The stop required direct handoff and did not permit leaving the package with another person."], "request": "Under depot policy, should this exception be routed to address clarification rather than redelivery?"}}, "method": "c2d", "provenance": {"source_id": "diverse-298", "source_is_synthetic": true, "source_sha256": "d5c7c6aa5962d6762278e5ae977fe87f27888823aa0a70204e56364a3bea30d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the ordered verification policy and the request concerning FX-184’s current depot custody and routing; the changed scan dates, record statuses, and rating-time observations are permissible case-observation changes rather than altered question bindings. The two focus spans are complete factual sentences, and the counterfactual coherently changes only CR-908 from reconciled to unresolved without conflicting with the exact-two-record inventory or other unchanged facts. Neither context includes a score, answer code, proposition identifier, rule table, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, including the universally scoped record-status and contradiction atoms. The focus atom is factual rather than policy-based. The base and counter assignments are jointly realizable with only a5 changing: an existing completion scan may either have been reconciled or remain unresolved. Empty policy_evidence is correct because the governing scoring criteria and instructions are entirely in the automatically retained questions object; no additional state-originating policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a later depot inbound scan tied to FX-184, physical label confirmation, reconciliation or voiding of every completion record, and absence of other contradictory evidence. This is sufficient for level 3.", "rule_index": 0, "sound": true}, {"reason": "The inbound scan and physical label confirmation establish depot custody. Refutation of the universal reconciliation atom entails that at least one completion record remains unvoided and unreconciled, excluding level 3 and satisfying the unresolved-conflict condition for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evidence record contains a delivery-completion scan whose parcel identifier is FX-184."}, {"id": "a2", "statement": "The depot inbound-scan record at issue has parcel identifier FX-184."}, {"id": "a3", "statement": "The timestamp of the depot inbound-scan record for FX-184 is later than the timestamp of the delivery-completion scan for FX-184."}, {"id": "a4", "statement": "The label on the carton physically inspected at the depot identifies the carton as FX-184."}, {"id": "a5", "statement": "As of the rating time, every delivery-completion scan or equivalent completion record for FX-184 has been voided or reconciled."}, {"id": "a6", "statement": "As of the rating time, no evidence other than a delivery-completion scan or equivalent completion record contradicts current depot custody of FX-184."}], "base_state_json": "\"The evidence ledger identifies DC-731 as a delivery-completion scan for parcel FX-184. At 2026-09-16 12:40 UTC, after that scan was logged, the depot system recorded an inbound scan whose parcel field was FX-184. During the receiving check, an operations lead physically inspected the carton and confirmed that its attached label identified it as FX-184. At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC. At the rating time of 2026-09-17 15:00 UTC, record CR-908 had been marked reconciled at 2026-09-17 11:26 UTC. A custody audit completed at the rating time found no evidence outside the delivery-completion or equivalent-completion category that contradicted current depot custody of FX-184.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC."}, {"path": [], "text": "At the rating time of 2026-09-17 15:00 UTC, record CR-908 had been marked reconciled at 2026-09-17 11:26 UTC."}], "policy_evidence": [], "rules": [{"justification": "The later depot inbound scan and physical label confirmation establish custody, while every completion record is voided or reconciled and no other contradictory evidence remains.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The later depot inbound scan and physical label confirmation establish custody, but refutation of universal reconciliation entails that at least one completion record remains unresolved and contradictory.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC.", "negative_left": "At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC.", "negative_right": "At the rating time of 2026-09-17 15:00 UTC, record CR-908 remained neither voided nor reconciled.", "right": "At the rating time of 2026-09-17 15:00 UTC, record CR-908 had been marked reconciled at 2026-09-17 11:26 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-300-001", "id": "scale-diverse-300-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsupported: No depot scan or physical confirmation exists, and reliable records instead show successful delivery or custody elsewhere.", "1 — Weakly supported: A driver note or customer report suggests depot return, but there is no depot inbound scan or physical package confirmation.", "2 — Strongly supported with a conflict: A depot inbound scan and physical label confirmation establish depot custody, but an unresolved delivery-completion scan or equivalent record still contradicts the claim.", "3 — Fully verified: The depot inbound scan and physical confirmation establish custody, and all completion records have been voided or reconciled with no remaining contradictory evidence."], "instructions": "Rate the evidence supporting the claim that FX-184 is currently at the depot and should be routed to depot hold rather than treated as completed. Apply the ordered verification levels below. A later inbound scan plus physical label confirmation establishes depot custody, but any contradictory completion scan prevents the highest level until corrected.", "type": "score"}}, "state": "The evidence ledger identifies DC-731 as a delivery-completion scan for parcel FX-184. At 2026-09-16 12:40 UTC, after that scan was logged, the depot system recorded an inbound scan whose parcel field was FX-184. During the receiving check, an operations lead physically inspected the carton and confirmed that its attached label identified it as FX-184. At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC. At the rating time of 2026-09-17 15:00 UTC, record CR-908 had been marked reconciled at 2026-09-17 11:26 UTC. A custody audit completed at the rating time found no evidence outside the delivery-completion or equivalent-completion category that contradicted current depot custody of FX-184."}, "method": "c2d", "provenance": {"source_id": "diverse-300", "source_is_synthetic": true, "source_sha256": "1e2508e0c80516c96c8e4568290fee85c6627ec2622e428e2eeceee6312eb181", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the ordered verification policy and the request concerning FX-184’s current depot custody and routing; the changed scan dates, record statuses, and rating-time observations are permissible case-observation changes rather than altered question bindings. The two focus spans are complete factual sentences, and the counterfactual coherently changes only CR-908 from reconciled to unresolved without conflicting with the exact-two-record inventory or other unchanged facts. Neither context includes a score, answer code, proposition identifier, rule table, output instruction, or explicit label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, including the universally scoped record-status and contradiction atoms. The focus atom is factual rather than policy-based. The base and counter assignments are jointly realizable with only a5 changing: an existing completion scan may either have been reconciled or remain unresolved. Empty policy_evidence is correct because the governing scoring criteria and instructions are entirely in the automatically retained questions object; no additional state-originating policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a later depot inbound scan tied to FX-184, physical label confirmation, reconciliation or voiding of every completion record, and absence of other contradictory evidence. This is sufficient for level 3.", "rule_index": 0, "sound": true}, {"reason": "The inbound scan and physical label confirmation establish depot custody. Refutation of the universal reconciliation atom entails that at least one completion record remains unvoided and unreconciled, excluding level 3 and satisfying the unresolved-conflict condition for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evidence record contains a delivery-completion scan whose parcel identifier is FX-184."}, {"id": "a2", "statement": "The depot inbound-scan record at issue has parcel identifier FX-184."}, {"id": "a3", "statement": "The timestamp of the depot inbound-scan record for FX-184 is later than the timestamp of the delivery-completion scan for FX-184."}, {"id": "a4", "statement": "The label on the carton physically inspected at the depot identifies the carton as FX-184."}, {"id": "a5", "statement": "As of the rating time, every delivery-completion scan or equivalent completion record for FX-184 has been voided or reconciled."}, {"id": "a6", "statement": "As of the rating time, no evidence other than a delivery-completion scan or equivalent completion record contradicts current depot custody of FX-184."}], "base_state_json": "\"The evidence ledger identifies DC-731 as a delivery-completion scan for parcel FX-184. At 2026-09-16 12:40 UTC, after that scan was logged, the depot system recorded an inbound scan whose parcel field was FX-184. During the receiving check, an operations lead physically inspected the carton and confirmed that its attached label identified it as FX-184. At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC. At the rating time of 2026-09-17 15:00 UTC, record CR-908 had been marked reconciled at 2026-09-17 11:26 UTC. A custody audit completed at the rating time found no evidence outside the delivery-completion or equivalent-completion category that contradicted current depot custody of FX-184.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC."}, {"path": [], "text": "At the rating time of 2026-09-17 15:00 UTC, record CR-908 had been marked reconciled at 2026-09-17 11:26 UTC."}], "policy_evidence": [], "rules": [{"justification": "The later depot inbound scan and physical label confirmation establish custody, while every completion record is voided or reconciled and no other contradictory evidence remains.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The later depot inbound scan and physical label confirmation establish custody, but refutation of universal reconciliation entails that at least one completion record remains unresolved and contradictory.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC.", "negative_left": "At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC.", "negative_right": "At the rating time of 2026-09-17 15:00 UTC, record CR-908 remained neither voided nor reconciled.", "right": "At the rating time of 2026-09-17 15:00 UTC, record CR-908 had been marked reconciled at 2026-09-17 11:26 UTC."}, "verifier_independent_model": false}, "family": "scale-diverse-300-001", "id": "scale-diverse-300-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsupported: No depot scan or physical confirmation exists, and reliable records instead show successful delivery or custody elsewhere.", "1 — Weakly supported: A driver note or customer report suggests depot return, but there is no depot inbound scan or physical package confirmation.", "2 — Strongly supported with a conflict: A depot inbound scan and physical label confirmation establish depot custody, but an unresolved delivery-completion scan or equivalent record still contradicts the claim.", "3 — Fully verified: The depot inbound scan and physical confirmation establish custody, and all completion records have been voided or reconciled with no remaining contradictory evidence."], "instructions": "Rate the evidence supporting the claim that FX-184 is currently at the depot and should be routed to depot hold rather than treated as completed. Apply the ordered verification levels below. A later inbound scan plus physical label confirmation establishes depot custody, but any contradictory completion scan prevents the highest level until corrected.", "type": "score"}}, "state": "The evidence ledger identifies DC-731 as a delivery-completion scan for parcel FX-184. At 2026-09-16 12:40 UTC, after that scan was logged, the depot system recorded an inbound scan whose parcel field was FX-184. During the receiving check, an operations lead physically inspected the carton and confirmed that its attached label identified it as FX-184. At the rating time of 2026-09-17 15:00 UTC, the complete FX-184 record inventory contained exactly two delivery-completion or equivalent completion records: scan DC-731, logged at 2026-09-16 09:12 UTC and marked voided at 2026-09-17 10:04 UTC, and record CR-908, logged at 2026-09-16 09:19 UTC. At the rating time of 2026-09-17 15:00 UTC, record CR-908 remained neither voided nor reconciled. A custody audit completed at the rating time found no evidence outside the delivery-completion or equivalent-completion category that contradicted current depot custody of FX-184."}, "method": "c2d", "provenance": {"source_id": "diverse-300", "source_is_synthetic": true, "source_sha256": "1e2508e0c80516c96c8e4568290fee85c6627ec2622e428e2eeceee6312eb181", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing policy. Entity, request-scope, and renewal bindings remain aligned, while dates are permissible changed observations. The evidence consists of two complete factual sentences: “The disputed September 14, 2026 annual renewal posted at 09:00 UTC.” and “The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC.” The counterfactual’s September 23 request is coherent with the September 14 renewal and does not duplicate or contradict any measurement. Neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "full_context_fact_states": {"base": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "counterfactual": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "remove_left": {"request_within_seven_days": "unknown"}, "remove_right": {"request_within_seven_days": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"request_within_seven_days": "unknown"}, "negative_pair": {"request_within_seven_days": "refuted"}, "negative_sentence": {"request_within_seven_days": "unknown"}, "positive_pair": {"request_within_seven_days": "supported"}, "right": {"request_within_seven_days": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single relevant factual relationship; none is a bundled final resolution classification. The focus is the factual timing relationship between the correction request and renewal. Base and counter assignments can be realized by changing only the request date from within the seven-day window to outside it while holding the remaining facts fixed. Empty policy_evidence is correct because all substantive governing rules are already preserved in the unchanged questions object; the original state supplies case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a recorded and valid offer, pre-renewal conditional acceptance tied to applying the credit, omission of the credit, and a correction request within seven calendar days. These facts satisfy the credit-and-annual-resolution rubric. Refutation of two settled charges also excludes duplicate-payment routing.", "rule_index": 0, "sound": true}, {"reason": "With the other eligibility facts fixed, a correction request explicitly outside the seven-calendar-day limit makes the credit ineligible. The conjunction also establishes each expressly requested fallback component—annual cancellation, movement to monthly, and the unused-term refund—and refutes the charge multiplicity required for duplicate-payment routing.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "offer_recorded", "statement": "A $24 retention offer was recorded on the subscriber's account for the disputed September 14 annual renewal."}, {"id": "offer_valid", "statement": "The $24 retention offer recorded on the subscriber's account was valid for the disputed September 14 annual renewal."}, {"id": "acceptance_before_renewal", "statement": "The subscriber accepted the recorded $24 retention offer before the disputed September 14 annual renewal posted."}, {"id": "acceptance_condition_credit", "statement": "The condition in the subscriber's acceptance of the recorded offer was application of the eligible $24 retention credit."}, {"id": "credit_omitted", "statement": "The $24 retention credit was omitted from invoice INV-8841 for the disputed September 14 annual renewal."}, {"id": "request_within_seven_days", "statement": "The subscriber's correction request in the current contact was made no later than seven calendar days after the disputed September 14 annual renewal."}, {"id": "fallback_cancel", "statement": "The subscriber expressly requested cancellation of the annual subscription if the $24 retention credit was ineligible."}, {"id": "fallback_monthly", "statement": "The subscriber expressly requested a move to a monthly subscription if the $24 retention credit was ineligible."}, {"id": "fallback_refund", "statement": "The subscriber expressly requested the permitted refund of the unused annual term if the $24 retention credit was ineligible."}, {"id": "two_settled_charges", "statement": "The subscriber's account shows at least two settled subscription charges for the disputed September 14 annual renewal."}], "base_state_json": "[{\"speaker\":\"Billing support agent\",\"text\":\"The account history records a $24 retention offer for the disputed annual renewal, and the offer was valid for annual renewals. The subscriber accepted it before the renewal posted, with acceptance conditioned on applying the eligible $24 retention credit. Invoice INV-8841 omitted that credit.\"},{\"speaker\":\"Billing support agent\",\"text\":\"The disputed September 14, 2026 annual renewal posted at 09:00 UTC.\"},{\"speaker\":\"Subscriber\",\"text\":\"The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC.\"},{\"speaker\":\"Subscriber\",\"text\":\"The subscriber expressly requested cancellation of annual service, a move to monthly, and a permitted refund of the unused annual term if the credit was ineligible.\"},{\"speaker\":\"Subscription operations analyst\",\"text\":\"The account shows one settled subscription charge for this renewal, not two. No adjustment, cancellation, monthly conversion, or refund has been completed.\"},{\"speaker\":\"Billing support agent\",\"text\":\"Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges.\"}]", "base_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "counter_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "focus_atom": "request_within_seven_days", "focus_evidence": [{"path": ["1", "text"], "text": "The disputed September 14, 2026 annual renewal posted at 09:00 UTC."}, {"path": ["2", "text"], "text": "The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC."}], "policy_evidence": [], "rules": [{"justification": "The omitted $24 credit is eligible because the valid recorded offer was accepted before renewal, the acceptance condition was application of that eligible credit, and correction was requested within seven calendar days. Refutation of two settled charges excludes duplicate-payment routing, and the actionable primary request excludes clarification and fallback.", "target": "apply_credit_keep_annual", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}, {"justification": "An explicitly late correction request makes the omitted credit ineligible under the seven-calendar-day requirement. The subscriber expressly supplied every component of the authorized fallback. Refutation of two settled charges excludes duplicate-payment routing, and the actionable fallback excludes clarification and none of the above.", "target": "execute_fallback_cancellation", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}]}, "verified_pair": {"left": "The disputed September 14, 2026 annual renewal posted at 09:00 UTC.", "negative_left": "The disputed September 14, 2026 annual renewal posted at 09:00 UTC.", "negative_right": "The subscriber's correction request in the current contact was logged on September 23, 2026, at 09:00 UTC.", "right": "The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-013-003", "id": "fast-41-diverse-013-003-base", "input": {"questions": {"decision": {"criteria": {"apply_credit_keep_annual": "Add the $24 credit and preserve the annual subscription when the offer and pre-renewal acceptance are verified and the correction request is within seven calendar days.", "execute_fallback_cancellation": "Cancel annual, move to monthly, and process the permitted unused-term refund only when the credit is ineligible and the subscriber expressly supplied this fallback.", "none_of_above": "Use only when the evidence satisfies none of the four resolution rubrics above.", "route_duplicate_payment_review": "Route for duplicate-payment investigation only when the account shows at least two settled subscription charges for the disputed renewal.", "seek_intent_clarification": "Ask the subscriber to clarify only when neither the primary conditional request nor any fallback identifies an actionable outcome."}, "instructions": "Choose the authorized resolution. Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges.", "type": "choice"}}, "state": [{"speaker": "Billing support agent", "text": "The account history records a $24 retention offer for the disputed annual renewal, and the offer was valid for annual renewals. The subscriber accepted it before the renewal posted, with acceptance conditioned on applying the eligible $24 retention credit. Invoice INV-8841 omitted that credit."}, {"speaker": "Billing support agent", "text": "The disputed September 14, 2026 annual renewal posted at 09:00 UTC."}, {"speaker": "Subscriber", "text": "The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC."}, {"speaker": "Subscriber", "text": "The subscriber expressly requested cancellation of annual service, a move to monthly, and a permitted refund of the unused annual term if the credit was ineligible."}, {"speaker": "Subscription operations analyst", "text": "The account shows one settled subscription charge for this renewal, not two. No adjustment, cancellation, monthly conversion, or refund has been completed."}, {"speaker": "Billing support agent", "text": "Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges."}]}, "method": "c2d", "provenance": {"source_id": "diverse-013", "source_is_synthetic": true, "source_sha256": "640e873611af9a131e8dfe8de5027d31b2e53121b8e09c71e14040fbe16c6d8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "apply_credit_keep_annual"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing policy. Entity, request-scope, and renewal bindings remain aligned, while dates are permissible changed observations. The evidence consists of two complete factual sentences: “The disputed September 14, 2026 annual renewal posted at 09:00 UTC.” and “The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC.” The counterfactual’s September 23 request is coherent with the September 14 renewal and does not duplicate or contradict any measurement. Neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "full_context_fact_states": {"base": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "counterfactual": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "remove_left": {"request_within_seven_days": "unknown"}, "remove_right": {"request_within_seven_days": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"request_within_seven_days": "unknown"}, "negative_pair": {"request_within_seven_days": "refuted"}, "negative_sentence": {"request_within_seven_days": "unknown"}, "positive_pair": {"request_within_seven_days": "supported"}, "right": {"request_within_seven_days": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single relevant factual relationship; none is a bundled final resolution classification. The focus is the factual timing relationship between the correction request and renewal. Base and counter assignments can be realized by changing only the request date from within the seven-day window to outside it while holding the remaining facts fixed. Empty policy_evidence is correct because all substantive governing rules are already preserved in the unchanged questions object; the original state supplies case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a recorded and valid offer, pre-renewal conditional acceptance tied to applying the credit, omission of the credit, and a correction request within seven calendar days. These facts satisfy the credit-and-annual-resolution rubric. Refutation of two settled charges also excludes duplicate-payment routing.", "rule_index": 0, "sound": true}, {"reason": "With the other eligibility facts fixed, a correction request explicitly outside the seven-calendar-day limit makes the credit ineligible. The conjunction also establishes each expressly requested fallback component—annual cancellation, movement to monthly, and the unused-term refund—and refutes the charge multiplicity required for duplicate-payment routing.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "offer_recorded", "statement": "A $24 retention offer was recorded on the subscriber's account for the disputed September 14 annual renewal."}, {"id": "offer_valid", "statement": "The $24 retention offer recorded on the subscriber's account was valid for the disputed September 14 annual renewal."}, {"id": "acceptance_before_renewal", "statement": "The subscriber accepted the recorded $24 retention offer before the disputed September 14 annual renewal posted."}, {"id": "acceptance_condition_credit", "statement": "The condition in the subscriber's acceptance of the recorded offer was application of the eligible $24 retention credit."}, {"id": "credit_omitted", "statement": "The $24 retention credit was omitted from invoice INV-8841 for the disputed September 14 annual renewal."}, {"id": "request_within_seven_days", "statement": "The subscriber's correction request in the current contact was made no later than seven calendar days after the disputed September 14 annual renewal."}, {"id": "fallback_cancel", "statement": "The subscriber expressly requested cancellation of the annual subscription if the $24 retention credit was ineligible."}, {"id": "fallback_monthly", "statement": "The subscriber expressly requested a move to a monthly subscription if the $24 retention credit was ineligible."}, {"id": "fallback_refund", "statement": "The subscriber expressly requested the permitted refund of the unused annual term if the $24 retention credit was ineligible."}, {"id": "two_settled_charges", "statement": "The subscriber's account shows at least two settled subscription charges for the disputed September 14 annual renewal."}], "base_state_json": "[{\"speaker\":\"Billing support agent\",\"text\":\"The account history records a $24 retention offer for the disputed annual renewal, and the offer was valid for annual renewals. The subscriber accepted it before the renewal posted, with acceptance conditioned on applying the eligible $24 retention credit. Invoice INV-8841 omitted that credit.\"},{\"speaker\":\"Billing support agent\",\"text\":\"The disputed September 14, 2026 annual renewal posted at 09:00 UTC.\"},{\"speaker\":\"Subscriber\",\"text\":\"The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC.\"},{\"speaker\":\"Subscriber\",\"text\":\"The subscriber expressly requested cancellation of annual service, a move to monthly, and a permitted refund of the unused annual term if the credit was ineligible.\"},{\"speaker\":\"Subscription operations analyst\",\"text\":\"The account shows one settled subscription charge for this renewal, not two. No adjustment, cancellation, monthly conversion, or refund has been completed.\"},{\"speaker\":\"Billing support agent\",\"text\":\"Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges.\"}]", "base_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "counter_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "focus_atom": "request_within_seven_days", "focus_evidence": [{"path": ["1", "text"], "text": "The disputed September 14, 2026 annual renewal posted at 09:00 UTC."}, {"path": ["2", "text"], "text": "The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC."}], "policy_evidence": [], "rules": [{"justification": "The omitted $24 credit is eligible because the valid recorded offer was accepted before renewal, the acceptance condition was application of that eligible credit, and correction was requested within seven calendar days. Refutation of two settled charges excludes duplicate-payment routing, and the actionable primary request excludes clarification and fallback.", "target": "apply_credit_keep_annual", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}, {"justification": "An explicitly late correction request makes the omitted credit ineligible under the seven-calendar-day requirement. The subscriber expressly supplied every component of the authorized fallback. Refutation of two settled charges excludes duplicate-payment routing, and the actionable fallback excludes clarification and none of the above.", "target": "execute_fallback_cancellation", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}]}, "verified_pair": {"left": "The disputed September 14, 2026 annual renewal posted at 09:00 UTC.", "negative_left": "The disputed September 14, 2026 annual renewal posted at 09:00 UTC.", "negative_right": "The subscriber's correction request in the current contact was logged on September 23, 2026, at 09:00 UTC.", "right": "The subscriber's correction request in the current contact was logged on September 20, 2026, at 09:00 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-013-003", "id": "fast-41-diverse-013-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"apply_credit_keep_annual": "Add the $24 credit and preserve the annual subscription when the offer and pre-renewal acceptance are verified and the correction request is within seven calendar days.", "execute_fallback_cancellation": "Cancel annual, move to monthly, and process the permitted unused-term refund only when the credit is ineligible and the subscriber expressly supplied this fallback.", "none_of_above": "Use only when the evidence satisfies none of the four resolution rubrics above.", "route_duplicate_payment_review": "Route for duplicate-payment investigation only when the account shows at least two settled subscription charges for the disputed renewal.", "seek_intent_clarification": "Ask the subscriber to clarify only when neither the primary conditional request nor any fallback identifies an actionable outcome."}, "instructions": "Choose the authorized resolution. Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges.", "type": "choice"}}, "state": [{"speaker": "Billing support agent", "text": "The account history records a $24 retention offer for the disputed annual renewal, and the offer was valid for annual renewals. The subscriber accepted it before the renewal posted, with acceptance conditioned on applying the eligible $24 retention credit. Invoice INV-8841 omitted that credit."}, {"speaker": "Billing support agent", "text": "The disputed September 14, 2026 annual renewal posted at 09:00 UTC."}, {"speaker": "Subscriber", "text": "The subscriber's correction request in the current contact was logged on September 23, 2026, at 09:00 UTC."}, {"speaker": "Subscriber", "text": "The subscriber expressly requested cancellation of annual service, a move to monthly, and a permitted refund of the unused annual term if the credit was ineligible."}, {"speaker": "Subscription operations analyst", "text": "The account shows one settled subscription charge for this renewal, not two. No adjustment, cancellation, monthly conversion, or refund has been completed."}, {"speaker": "Billing support agent", "text": "Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges."}]}, "method": "c2d", "provenance": {"source_id": "diverse-013", "source_is_synthetic": true, "source_sha256": "640e873611af9a131e8dfe8de5027d31b2e53121b8e09c71e14040fbe16c6d8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "execute_fallback_cancellation"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, and both contexts retain Mira Chen, unresolved billing issues, and latest-update bindings. Each evidence span is a complete factual sentence. The counterfactual changes only issue T’s timestamp, making issue R latest without creating internal contradictions. Neither context embeds an answer, label, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Mira Chen’s account has two verified billing issues, and both remain unresolved. Issue T concerns only correction of sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it is not about tax, a promotion, or a plan change. The invoice disputed in R was not later voided. Fewer than two settled payments cover the subscription period at issue, and the merchant has issued no refund. These two records comprise every verified unresolved billing issue currently listed for Mira. The latest timeline-update time recorded for billing issue T was 2026-09-17T14:00:00Z. The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z. Support retained both records as open after verification, with no additional verified unresolved billing issue identified.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The latest timeline-update time recorded for billing issue T was 2026-09-17T14:00:00Z."}, {"path": [], "text": "The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The latest timeline-update time recorded for billing issue T was 2026-09-17T14:00:00Z.", "negative_left": "The latest timeline-update time recorded for billing issue T was 2026-09-17T12:00:00Z.", "negative_right": "The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z.", "right": "The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z."}, "verifier_independent_model": false}, "family": "fast-41-diverse-014-009", "id": "fast-41-diverse-014-009-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Mira Chen’s account has two verified billing issues, and both remain unresolved. Issue T concerns only correction of sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it is not about tax, a promotion, or a plan change. The invoice disputed in R was not later voided. Fewer than two settled payments cover the subscription period at issue, and the merchant has issued no refund. These two records comprise every verified unresolved billing issue currently listed for Mira. The latest timeline-update time recorded for billing issue T was 2026-09-17T14:00:00Z. The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z. Support retained both records as open after verification, with no additional verified unresolved billing issue identified."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, and both contexts retain Mira Chen, unresolved billing issues, and latest-update bindings. Each evidence span is a complete factual sentence. The counterfactual changes only issue T’s timestamp, making issue R latest without creating internal contradictions. Neither context embeds an answer, label, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Mira Chen’s account has two verified billing issues, and both remain unresolved. Issue T concerns only correction of sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it is not about tax, a promotion, or a plan change. The invoice disputed in R was not later voided. Fewer than two settled payments cover the subscription period at issue, and the merchant has issued no refund. These two records comprise every verified unresolved billing issue currently listed for Mira. The latest timeline-update time recorded for billing issue T was 2026-09-17T14:00:00Z. The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z. Support retained both records as open after verification, with no additional verified unresolved billing issue identified.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The latest timeline-update time recorded for billing issue T was 2026-09-17T14:00:00Z."}, {"path": [], "text": "The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The latest timeline-update time recorded for billing issue T was 2026-09-17T14:00:00Z.", "negative_left": "The latest timeline-update time recorded for billing issue T was 2026-09-17T12:00:00Z.", "negative_right": "The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z.", "right": "The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z."}, "verifier_independent_model": false}, "family": "fast-41-diverse-014-009", "id": "fast-41-diverse-014-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Mira Chen’s account has two verified billing issues, and both remain unresolved. Issue T concerns only correction of sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it is not about tax, a promotion, or a plan change. The invoice disputed in R was not later voided. Fewer than two settled payments cover the subscription period at issue, and the merchant has issued no refund. These two records comprise every verified unresolved billing issue currently listed for Mira. The latest timeline-update time recorded for billing issue T was 2026-09-17T12:00:00Z. The latest timeline-update time recorded for billing issue R was 2026-09-17T13:00:00Z. Support retained both records as open after verification, with no additional verified unresolved billing issue identified."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mira Chen, billing issues T and R, and the latest-update time path; changed case observations are internally consistent. The two evidence spans are complete factual sentences and exactly support the base context. The counterfactual consistently changes only R’s latest timestamp to 17:00 UTC, with no contradictory duplicate measurement. No context embeds an answer code, rationale, proposition ID, rule table, or output instruction; the unchanged original questions preserve the governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Mira Chen’s account contains exactly two verified, unresolved billing issues: T and R. Issue T is limited to correcting sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it is not a tax, promotion, or plan-change matter, and its disputed invoice was not later voided. Fewer than two settled payments cover R’s subscription period, and the merchant has issued no refund for R. No other verified unresolved billing issue is recorded for Mira. The account ledger records the following timeline updates: The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC. The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 14:00 UTC. These records concern the latest updates associated with their respective issues.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC."}, {"path": [], "text": "The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 14:00 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC.", "negative_left": "The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC.", "negative_right": "The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 17:00 UTC.", "right": "The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 14:00 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-014-010", "id": "fast-41-diverse-014-010-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Mira Chen’s account contains exactly two verified, unresolved billing issues: T and R. Issue T is limited to correcting sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it is not a tax, promotion, or plan-change matter, and its disputed invoice was not later voided. Fewer than two settled payments cover R’s subscription period, and the merchant has issued no refund for R. No other verified unresolved billing issue is recorded for Mira. The account ledger records the following timeline updates: The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC. The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 14:00 UTC. These records concern the latest updates associated with their respective issues."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mira Chen, billing issues T and R, and the latest-update time path; changed case observations are internally consistent. The two evidence spans are complete factual sentences and exactly support the base context. The counterfactual consistently changes only R’s latest timestamp to 17:00 UTC, with no contradictory duplicate measurement. No context embeds an answer code, rationale, proposition ID, rule table, or output instruction; the unchanged original questions preserve the governing policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Mira Chen’s account contains exactly two verified, unresolved billing issues: T and R. Issue T is limited to correcting sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it is not a tax, promotion, or plan-change matter, and its disputed invoice was not later voided. Fewer than two settled payments cover R’s subscription period, and the merchant has issued no refund for R. No other verified unresolved billing issue is recorded for Mira. The account ledger records the following timeline updates: The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC. The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 14:00 UTC. These records concern the latest updates associated with their respective issues.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC."}, {"path": [], "text": "The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 14:00 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC.", "negative_left": "The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC.", "negative_right": "The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 17:00 UTC.", "right": "The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 14:00 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-014-010", "id": "fast-41-diverse-014-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Mira Chen’s account contains exactly two verified, unresolved billing issues: T and R. Issue T is limited to correcting sales tax on an invoice. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it is not a tax, promotion, or plan-change matter, and its disputed invoice was not later voided. Fewer than two settled payments cover R’s subscription period, and the merchant has issued no refund for R. No other verified unresolved billing issue is recorded for Mira. The account ledger records the following timeline updates: The latest timeline-update record for billing issue T was timestamped 18 March 2026 at 16:00 UTC. The latest timeline-update record for billing issue R was timestamped 18 March 2026 at 17:00 UTC. These records concern the latest updates associated with their respective issues."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the stated settlement policy, request, subscriber, renewal, and time bindings. The base evidence includes “Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references.” The base evidence includes “The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record.” The two evidence spans are complete factual sentences. The counterfactual changes only the ledger status and remains internally coherent with the unchanged policy and records. Neither context embeds an answer, code, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Case note: Billing support opened a duplicate-payment review for Mara Venn after her July 3 Pro plan renewal. Subscription operations examined the invoice, bank image, and settlement records. The invoice labels both listed charges as paid, while the bank image contains two entries and marks one as pending. The review concerns the $84 renewal and uses the settlement ledger as the controlling record when source descriptions differ. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries. Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references. The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record.\",\"request\":\"Verify whether this case qualifies for a duplicate-payment adjustment under the stated policy.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["context"], "text": "Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references."}, {"path": ["context"], "text": "The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record."}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references.", "negative_left": "Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references.", "negative_right": "The final settlement ledger records an unsettled entry for each of the two references listed in Mara Venn’s July 3 Pro renewal record.", "right": "The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record."}, "verifier_independent_model": false}, "family": "fast-41-diverse-015-005", "id": "fast-41-diverse-015-005-base", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Case note: Billing support opened a duplicate-payment review for Mara Venn after her July 3 Pro plan renewal. Subscription operations examined the invoice, bank image, and settlement records. The invoice labels both listed charges as paid, while the bank image contains two entries and marks one as pending. The review concerns the $84 renewal and uses the settlement ledger as the controlling record when source descriptions differ. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries. Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references. The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record.", "request": "Verify whether this case qualifies for a duplicate-payment adjustment under the stated policy."}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the stated settlement policy, request, subscriber, renewal, and time bindings. The base evidence includes “Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references.” The base evidence includes “The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record.” The two evidence spans are complete factual sentences. The counterfactual changes only the ledger status and remains internally coherent with the unchanged policy and records. Neither context embeds an answer, code, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Case note: Billing support opened a duplicate-payment review for Mara Venn after her July 3 Pro plan renewal. Subscription operations examined the invoice, bank image, and settlement records. The invoice labels both listed charges as paid, while the bank image contains two entries and marks one as pending. The review concerns the $84 renewal and uses the settlement ledger as the controlling record when source descriptions differ. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries. Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references. The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record.\",\"request\":\"Verify whether this case qualifies for a duplicate-payment adjustment under the stated policy.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["context"], "text": "Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references."}, {"path": ["context"], "text": "The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record."}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references.", "negative_left": "Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references.", "negative_right": "The final settlement ledger records an unsettled entry for each of the two references listed in Mara Venn’s July 3 Pro renewal record.", "right": "The final settlement ledger records a settled capture for each of the two references listed in Mara Venn’s July 3 Pro renewal record."}, "verifier_independent_model": false}, "family": "fast-41-diverse-015-005", "id": "fast-41-diverse-015-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Case note: Billing support opened a duplicate-payment review for Mara Venn after her July 3 Pro plan renewal. Subscription operations examined the invoice, bank image, and settlement records. The invoice labels both listed charges as paid, while the bank image contains two entries and marks one as pending. The review concerns the $84 renewal and uses the settlement ledger as the controlling record when source descriptions differ. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries. Mara Venn’s July 3 Pro renewal record identifies P441 and P442 as its two payment transaction references. The final settlement ledger records an unsettled entry for each of the two references listed in Mara Venn’s July 3 Pro renewal record.", "request": "Verify whether this case qualifies for a duplicate-payment adjustment under the stated policy."}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy, bindings, scope, and evidence format; the counterfactual changes only settlement facts coherently without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Subscriber Mara Venn disputes a July 3 Pro plan renewal. A billing support agent opened a duplicate-payment case, and a subscription operations analyst reviewed the records. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries. The invoice for the renewal labels both references as paid at $84 each. A bank screenshot shows two $84 entries, with one marked pending. The analyst’s case note identifies the renewal date, plan, subscriber, invoice, and ledger as belonging to the same review. The requested determination concerns eligibility for a duplicate-payment adjustment under the stated policy. The verified renewal record and settlement result are retained below as separate evidence entries.\",\"evidence\":[\"Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442.\",\"For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records a settled capture for each reference listed in the renewal record.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442."}, {"path": ["evidence", "1"], "text": "For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records a settled capture for each reference listed in the renewal record."}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442.", "negative_left": "Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442.", "negative_right": "For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records an unsettled entry for each reference listed in the renewal record.", "right": "For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records a settled capture for each reference listed in the renewal record."}, "verifier_independent_model": false}, "family": "fast-41-diverse-015-007", "id": "fast-41-diverse-015-007-base", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Subscriber Mara Venn disputes a July 3 Pro plan renewal. A billing support agent opened a duplicate-payment case, and a subscription operations analyst reviewed the records. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries. The invoice for the renewal labels both references as paid at $84 each. A bank screenshot shows two $84 entries, with one marked pending. The analyst’s case note identifies the renewal date, plan, subscriber, invoice, and ledger as belonging to the same review. The requested determination concerns eligibility for a duplicate-payment adjustment under the stated policy. The verified renewal record and settlement result are retained below as separate evidence entries.", "evidence": ["Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442.", "For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records a settled capture for each reference listed in the renewal record."]}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy, bindings, scope, and evidence format; the counterfactual changes only settlement facts coherently without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual set or disposition relationship, and the focus atom is factual rather than policy-based. The base and counter assignments are both realizable while holding a1 fixed and changing only a2. The policy evidence correctly cites the original state's eligibility condition and, importantly, its conflict-resolution hierarchy; no governing state-originated rule needed to interpret the retained question is omitted.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given a1, P441 and P442 are exactly the complete set of transaction references for the renewal. Given a2, both have final-ledger dispositions as settled captures. Therefore exactly two captures for that renewal settled, which is sufficient for true.", "rule_index": 0, "sound": true}, {"reason": "Given a1, there are exactly two transaction references associated with the renewal. Refutation of a2 entails that at least one of those two does not have a final-ledger disposition as a settled capture. Because a1 excludes additional associated transaction references, fewer than two captures settled, which is sufficient for false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The complete set of payment transaction references associated with Mara Venn’s July 3 Pro plan renewal is {P441, P442}."}, {"id": "a2", "statement": "Every payment transaction reference in {P441, P442} has the final settlement-ledger disposition of a settled capture."}], "base_state_json": "{\"context\":\"Subscriber Mara Venn disputes a July 3 Pro plan renewal. A billing support agent opened a duplicate-payment case, and a subscription operations analyst reviewed the records. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries. The invoice for the renewal labels both references as paid at $84 each. A bank screenshot shows two $84 entries, with one marked pending. The analyst’s case note identifies the renewal date, plan, subscriber, invoice, and ledger as belonging to the same review. The requested determination concerns eligibility for a duplicate-payment adjustment under the stated policy. The verified renewal record and settlement result are retained below as separate evidence entries.\",\"evidence\":[\"Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442.\",\"For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records a settled capture for each reference listed in the renewal record.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442."}, {"path": ["evidence", "1"], "text": "For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records a settled capture for each reference listed in the renewal record."}], "policy_evidence": [{"path": ["context"], "text": "Policy grants an adjustment only when two captures for the same renewal have both settled."}, {"path": ["context"], "text": "When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries."}], "rules": [{"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, and the final settlement ledger establishes a settled capture for each, so two captures for that renewal settled.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "P441 and P442 are exactly the transactions associated with the same July 3 renewal, but the final settlement ledger establishes that not every one is a settled capture; therefore fewer than two captures for that renewal settled.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442.", "negative_left": "Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442.", "negative_right": "For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records an unsettled entry for each reference listed in the renewal record.", "right": "For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records a settled capture for each reference listed in the renewal record."}, "verifier_independent_model": false}, "family": "fast-41-diverse-015-007", "id": "fast-41-diverse-015-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The case is not verified as eligible because the evidence, resolved under policy, shows fewer than two settled captures.", "true": "The case is verified as eligible for a duplicate-payment adjustment because two renewal captures settled."}, "instructions": "Answer yes only if the evidence verifies that two captures for the same renewal both settled. Answer no if fewer than two captures settled, applying the stated conflict-resolution rule.", "type": "noul"}}, "state": {"context": "Subscriber Mara Venn disputes a July 3 Pro plan renewal. A billing support agent opened a duplicate-payment case, and a subscription operations analyst reviewed the records. Policy grants an adjustment only when two captures for the same renewal have both settled. When records conflict, the final settlement ledger overrides invoice payment labels and pending bank entries. The invoice for the renewal labels both references as paid at $84 each. A bank screenshot shows two $84 entries, with one marked pending. The analyst’s case note identifies the renewal date, plan, subscriber, invoice, and ledger as belonging to the same review. The requested determination concerns eligibility for a duplicate-payment adjustment under the stated policy. The verified renewal record and settlement result are retained below as separate evidence entries.", "evidence": ["Mara Venn’s July 3 Pro plan renewal record lists exactly the payment transaction references P441 and P442.", "For Mara Venn’s July 3 Pro plan renewal, the final settlement ledger records an unsettled entry for each reference listed in the renewal record."]}}, "method": "c2d", "provenance": {"source_id": "diverse-015", "source_is_synthetic": true, "source_sha256": "231da71bea3da926cacaa55282011ec235083aba22fd18f3635c05b4af0efd70", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy remains unchanged in the verbatim questions and both contexts. Account SF-48219, StreamForge Plus, two charges, $84.00 USD, and September 12 remain bound identically. Evidence consists of exactly two complete factual sentences: “For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2.” and “The settlement ledger entry labeled SC-2 carries status POSTED.” The counterfactual coherently changes only SC-2 from POSTED to PENDING. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"A duplicate-payment review concerns the StreamForge Plus renewal associated with account SF-48219. The two disputed records are the only charges under review, each for $84.00 in USD and dated September 12.\"},{\"speaker\":\"Reconciliation analyst\",\"text\":\"For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2.\"},{\"speaker\":\"Reconciliation analyst\",\"text\":\"The settlement ledger entry labeled SC-2 carries status POSTED.\"},{\"speaker\":\"Case note\",\"text\":\"The file contains the invoice, account identifier, date, currency, amount, and both ledger entries for operations review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2."}, {"path": ["2", "text"], "text": "The settlement ledger entry labeled SC-2 carries status POSTED."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2.", "negative_left": "For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2.", "negative_right": "The settlement ledger entry labeled SC-2 carries status PENDING.", "right": "The settlement ledger entry labeled SC-2 carries status POSTED."}, "verifier_independent_model": false}, "family": "fast-41-diverse-016-002", "id": "fast-41-diverse-016-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "A duplicate-payment review concerns the StreamForge Plus renewal associated with account SF-48219. The two disputed records are the only charges under review, each for $84.00 in USD and dated September 12."}, {"speaker": "Reconciliation analyst", "text": "For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2."}, {"speaker": "Reconciliation analyst", "text": "The settlement ledger entry labeled SC-2 carries status POSTED."}, {"speaker": "Case note", "text": "The file contains the invoice, account identifier, date, currency, amount, and both ledger entries for operations review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy remains unchanged in the verbatim questions and both contexts. Account SF-48219, StreamForge Plus, two charges, $84.00 USD, and September 12 remain bound identically. Evidence consists of exactly two complete factual sentences: “For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2.” and “The settlement ledger entry labeled SC-2 carries status POSTED.” The counterfactual coherently changes only SC-2 from POSTED to PENDING. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"A duplicate-payment review concerns the StreamForge Plus renewal associated with account SF-48219. The two disputed records are the only charges under review, each for $84.00 in USD and dated September 12.\"},{\"speaker\":\"Reconciliation analyst\",\"text\":\"For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2.\"},{\"speaker\":\"Reconciliation analyst\",\"text\":\"The settlement ledger entry labeled SC-2 carries status POSTED.\"},{\"speaker\":\"Case note\",\"text\":\"The file contains the invoice, account identifier, date, currency, amount, and both ledger entries for operations review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2."}, {"path": ["2", "text"], "text": "The settlement ledger entry labeled SC-2 carries status POSTED."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2.", "negative_left": "For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2.", "negative_right": "The settlement ledger entry labeled SC-2 carries status PENDING.", "right": "The settlement ledger entry labeled SC-2 carries status POSTED."}, "verifier_independent_model": false}, "family": "fast-41-diverse-016-002", "id": "fast-41-diverse-016-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "A duplicate-payment review concerns the StreamForge Plus renewal associated with account SF-48219. The two disputed records are the only charges under review, each for $84.00 in USD and dated September 12."}, {"speaker": "Reconciliation analyst", "text": "For account SF-48219, exactly two disputed charges dated September 12 belong to the StreamForge Plus renewal; both are USD 84.00, the first record is POSTED, and the second charge is identified in the settlement ledger as entry SC-2."}, {"speaker": "Reconciliation analyst", "text": "The settlement ledger entry labeled SC-2 carries status PENDING."}, {"speaker": "Case note", "text": "The file contains the invoice, account identifier, date, currency, amount, and both ledger entries for operations review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The two evidence quotes are retained verbatim: “For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing.” and “The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7.” The counterfactual coherently changes only the second charge to code 8, while both contexts preserve the unchanged policy, bindings, and scope without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The review concerns two disputed StreamForge Plus renewal charges tied to account SF-48219, each dated September 12. Each charge is listed as $84.00 in USD. The first charge has a posted settlement record and its processor reference is present.\"},{\"speaker\":\"Ledger note\",\"text\":\"For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing.\"},{\"speaker\":\"Ledger note\",\"text\":\"The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7.\"},{\"speaker\":\"Case note\",\"text\":\"The two entries have distinct processor references, and the review file treats them as separate disputed charges. The account, amount, currency, date, and first-charge settlement record are available for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing."}, {"path": ["2", "text"], "text": "The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing.", "negative_left": "For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing.", "negative_right": "The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 8.", "right": "The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-016-009", "id": "fast-41-diverse-016-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The review concerns two disputed StreamForge Plus renewal charges tied to account SF-48219, each dated September 12. Each charge is listed as $84.00 in USD. The first charge has a posted settlement record and its processor reference is present."}, {"speaker": "Ledger note", "text": "For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing."}, {"speaker": "Ledger note", "text": "The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7."}, {"speaker": "Case note", "text": "The two entries have distinct processor references, and the review file treats them as separate disputed charges. The account, amount, currency, date, and first-charge settlement record are available for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The two evidence quotes are retained verbatim: “For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing.” and “The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7.” The counterfactual coherently changes only the second charge to code 8, while both contexts preserve the unchanged policy, bindings, and scope without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal statements over the explicit pair of disputed charges remain atomic. The focus atom concerns the second charge's settlement status, not a policy classification. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing completeness requirements are already preserved verbatim in original_input.questions, while the state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every required element: account, amount, currency, date, and posted status for both disputed charges. Under the explicit criteria, this is sufficient for the false outcome because no required evidence is missing.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes the identifying elements and the first charge's posted status while refuting that the second charge is posted rather than pending or processing. Therefore the required confirmation that both charges are posted is absent, which is sufficient for the true outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is associated with account SF-48219."}, {"id": "a2", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal has an amount of $84.00."}, {"id": "a3", "statement": "Each of the two disputed charges for the September 12 StreamForge Plus renewal is denominated in USD."}, {"id": "a4", "statement": "Each of the two disputed charges for the StreamForge Plus renewal has a transaction date of September 12."}, {"id": "a5", "statement": "The first disputed charge for the September 12 StreamForge Plus renewal has settlement status posted."}, {"id": "a6", "statement": "The second disputed charge for the September 12 StreamForge Plus renewal has settlement status posted rather than pending or processing."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The review concerns two disputed StreamForge Plus renewal charges tied to account SF-48219, each dated September 12. Each charge is listed as $84.00 in USD. The first charge has a posted settlement record and its processor reference is present.\"},{\"speaker\":\"Ledger note\",\"text\":\"For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing.\"},{\"speaker\":\"Ledger note\",\"text\":\"The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7.\"},{\"speaker\":\"Case note\",\"text\":\"The two entries have distinct processor references, and the review file treats them as separate disputed charges. The account, amount, currency, date, and first-charge settlement record are available for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing."}, {"path": ["2", "text"], "text": "The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7."}], "policy_evidence": [], "rules": [{"justification": "The evidence identifies the account, amount, currency, and date and confirms that both disputed charges are posted, so no required evidence is missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Although the account, amount, currency, date, and first charge's posted status are established, the second disputed charge is not posted rather than pending or processing; the required confirmation that both charges are posted is therefore absent.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing.", "negative_left": "For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing.", "negative_right": "The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 8.", "right": "The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-016-009", "id": "fast-41-diverse-016-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — all required evidence is present, so operations can decide duplicate-payment adjustment eligibility now.", "true": "Yes — required evidence is missing, so route the case for missing-information collection before deciding adjustment eligibility."}, "instructions": "Completeness task: Should this case be classified as incomplete and routed for missing-information collection before duplicate-payment adjustment eligibility is decided? Under policy, review is ready only when evidence identifies the account, amount, currency, date, and confirms that both disputed charges are posted rather than pending or processing. If any required element is missing, answer yes; otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The review concerns two disputed StreamForge Plus renewal charges tied to account SF-48219, each dated September 12. Each charge is listed as $84.00 in USD. The first charge has a posted settlement record and its processor reference is present."}, {"speaker": "Ledger note", "text": "For account SF-48219, the September 12 StreamForge Plus renewal ledger identifies both disputed charges as USD transactions of $84.00 each and defines settlement code 7 as posted and code 8 as pending or processing."}, {"speaker": "Ledger note", "text": "The second disputed charge for the September 12 StreamForge Plus renewal is recorded under settlement code 8."}, {"speaker": "Case note", "text": "The two entries have distinct processor references, and the review file treats them as separate disputed charges. The account, amount, currency, date, and first-charge settlement record are available for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-016", "source_is_synthetic": true, "source_sha256": "97204fa625c1b35978a8d0749c2001f21aac60a18a4483e865f758c686ff541d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question and governing policy, while retaining the invoice, subscription, requester, dates, prices, remaining-day calculation, and one-invoice scope. The required evidence consists of two complete factual sentences: \"The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request.\" and \"That event is classified as a Premium feature-use event.\" The counterfactual changes only that event classification to Basic, which coherently removes premium use without contradictory duplicate measurements or assertions. Neither context embeds an answer, label, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price was $120.00, while the applicable Basic annual price was $72.00. At the time of the request, 363 of 365 service days remained, and the requested correction covered exactly one paid invoice, R-841. The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That event is classified as a Premium feature-use event. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing recorded the requested change for analyst review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That event is classified as a Premium feature-use event."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That event is classified as a Basic feature-use event.", "right": "That event is classified as a Premium feature-use event."}, "verifier_independent_model": false}, "family": "fast-41-diverse-017-016", "id": "fast-41-diverse-017-016-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price was $120.00, while the applicable Basic annual price was $72.00. At the time of the request, 363 of 365 service days remained, and the requested correction covered exactly one paid invoice, R-841. The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That event is classified as a Premium feature-use event. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing recorded the requested change for analyst review."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question and governing policy, while retaining the invoice, subscription, requester, dates, prices, remaining-day calculation, and one-invoice scope. The required evidence consists of two complete factual sentences: \"The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request.\" and \"That event is classified as a Premium feature-use event.\" The counterfactual changes only that event classification to Basic, which coherently removes premium use without contradictory duplicate measurements or assertions. Neither context embeds an answer, label, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price was $120.00, while the applicable Basic annual price was $72.00. At the time of the request, 363 of 365 service days remained, and the requested correction covered exactly one paid invoice, R-841. The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That event is classified as a Premium feature-use event. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing recorded the requested change for analyst review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That event is classified as a Premium feature-use event."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That event is classified as a Basic feature-use event.", "right": "That event is classified as a Premium feature-use event."}, "verifier_independent_model": false}, "family": "fast-41-diverse-017-016", "id": "fast-41-diverse-017-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price was $120.00, while the applicable Basic annual price was $72.00. At the time of the request, 363 of 365 service days remained, and the requested correction covered exactly one paid invoice, R-841. The complete feature-use audit log for subscription S-204 contains exactly one event between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That event is classified as a Basic feature-use event. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing recorded the requested change for analyst review."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria, and both contexts retain the relevant policy without adding exceptions or defaults. Both contexts preserve the pickup-order, proposed-item, fulfillment-decision, and note timing bindings. The evidence consists of exactly two complete factual sentences: “At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417.” and “The pickup order’s original requested item is recorded with SKU PX-417.” The counterfactual coherently changes the requested SKU to PX-418, making the located PX-417 item a permitted substitution rather than creating contradictory duplicate measurements. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At 9:15, the picker’s fulfillment record lists a proposed item from the chilled-drinks shelf. At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417. The pickup order’s original requested item is recorded with SKU PX-417. The picker proposed that located item for this pickup order at 9:15. A customer note entered at 9:11 is the latest note effective before the decision; it permits the picker-located item and specifies its relevant attributes, all of which match the item record. The exception log contains no documented safety exception requiring supervisor authorization, no documented regulatory exception requiring supervisor authorization, and no documented price-limit exception requiring supervisor authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417."}, {"path": [], "text": "The pickup order’s original requested item is recorded with SKU PX-417."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417.", "negative_left": "At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417.", "negative_right": "The pickup order’s original requested item is recorded with SKU PX-418.", "right": "The pickup order’s original requested item is recorded with SKU PX-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-019-002", "id": "fast-41-diverse-019-002-base", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At 9:15, the picker’s fulfillment record lists a proposed item from the chilled-drinks shelf. At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417. The pickup order’s original requested item is recorded with SKU PX-417. The picker proposed that located item for this pickup order at 9:15. A customer note entered at 9:11 is the latest note effective before the decision; it permits the picker-located item and specifies its relevant attributes, all of which match the item record. The exception log contains no documented safety exception requiring supervisor authorization, no documented regulatory exception requiring supervisor authorization, and no documented price-limit exception requiring supervisor authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_exact_item"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria, and both contexts retain the relevant policy without adding exceptions or defaults. Both contexts preserve the pickup-order, proposed-item, fulfillment-decision, and note timing bindings. The evidence consists of exactly two complete factual sentences: “At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417.” and “The pickup order’s original requested item is recorded with SKU PX-417.” The counterfactual coherently changes the requested SKU to PX-418, making the located PX-417 item a permitted substitution rather than creating contradictory duplicate measurements. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At 9:15, the picker’s fulfillment record lists a proposed item from the chilled-drinks shelf. At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417. The pickup order’s original requested item is recorded with SKU PX-417. The picker proposed that located item for this pickup order at 9:15. A customer note entered at 9:11 is the latest note effective before the decision; it permits the picker-located item and specifies its relevant attributes, all of which match the item record. The exception log contains no documented safety exception requiring supervisor authorization, no documented regulatory exception requiring supervisor authorization, and no documented price-limit exception requiring supervisor authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417."}, {"path": [], "text": "The pickup order’s original requested item is recorded with SKU PX-417."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417.", "negative_left": "At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417.", "negative_right": "The pickup order’s original requested item is recorded with SKU PX-418.", "right": "The pickup order’s original requested item is recorded with SKU PX-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-019-002", "id": "fast-41-diverse-019-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At 9:15, the picker’s fulfillment record lists a proposed item from the chilled-drinks shelf. At the 9:15 fulfillment decision, the store picker’s scan record identifies the located item with SKU PX-417. The pickup order’s original requested item is recorded with SKU PX-418. The picker proposed that located item for this pickup order at 9:15. A customer note entered at 9:11 is the latest note effective before the decision; it permits the picker-located item and specifies its relevant attributes, all of which match the item record. The exception log contains no documented safety exception requiring supervisor authorization, no documented regulatory exception requiring supervisor authorization, and no documented price-limit exception requiring supervisor authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_proposed_substitute"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and the relevant pickup-order, item, SKU, and 9:15 decision bindings. The two evidence entries are complete factual sentences. The SKU change in the counterfactual is coherent and does not create contradictory duplicate assertions. Neither context embeds an answer, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At 8:52 on 14 March 2026, the picker located a beverage for the pickup order and proposed it at the 9:15 fulfillment decision. At 9:02, the customer’s latest effective note said that this located beverage was acceptable and specified attributes that all matched it. The fulfillment record contains no documented safety exception requiring supervisor authorization, no documented regulatory exception requiring such authorization, and no documented price-limit exception requiring such authorization. At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision. The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-482. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision."}, {"path": [], "text": "The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-482."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision.", "negative_left": "At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision.", "negative_right": "The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-483.", "right": "The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-482."}, "verifier_independent_model": false}, "family": "fast-41-diverse-019-006", "id": "fast-41-diverse-019-006-base", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At 8:52 on 14 March 2026, the picker located a beverage for the pickup order and proposed it at the 9:15 fulfillment decision. At 9:02, the customer’s latest effective note said that this located beverage was acceptable and specified attributes that all matched it. The fulfillment record contains no documented safety exception requiring supervisor authorization, no documented regulatory exception requiring such authorization, and no documented price-limit exception requiring such authorization. At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision. The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-482. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_exact_item"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and the relevant pickup-order, item, SKU, and 9:15 decision bindings. The two evidence entries are complete factual sentences. The SKU change in the counterfactual is coherent and does not create contradictory duplicate assertions. Neither context embeds an answer, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At 8:52 on 14 March 2026, the picker located a beverage for the pickup order and proposed it at the 9:15 fulfillment decision. At 9:02, the customer’s latest effective note said that this located beverage was acceptable and specified attributes that all matched it. The fulfillment record contains no documented safety exception requiring supervisor authorization, no documented regulatory exception requiring such authorization, and no documented price-limit exception requiring such authorization. At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision. The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-482. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision."}, {"path": [], "text": "The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-482."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision.", "negative_left": "At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision.", "negative_right": "The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-483.", "right": "The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-482."}, "verifier_independent_model": false}, "family": "fast-41-diverse-019-006", "id": "fast-41-diverse-019-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At 8:52 on 14 March 2026, the picker located a beverage for the pickup order and proposed it at the 9:15 fulfillment decision. At 9:02, the customer’s latest effective note said that this located beverage was acceptable and specified attributes that all matched it. The fulfillment record contains no documented safety exception requiring supervisor authorization, no documented regulatory exception requiring such authorization, and no documented price-limit exception requiring such authorization. At 9:15 on 14 March 2026, the store picker recorded SKU Q7M-482 for the item located for the fulfillment decision. The pickup order for the 9:15 fulfillment decision originally requested SKU Q7M-483. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_proposed_substitute"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing policy without added exceptions or defaults. Both contexts retain the 9:15 decision, order, picker, proposed-item, and note-chronology bindings while changing only case observations. Evidence consists of the exact factual quotes “At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4.”; “The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4.” The counterfactual coherently changes the requested SKU while retaining a distinct located SKU and contains no contradictory duplicate assertion. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4. The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4. The store picker proposed the located item for fulfillment at that decision. The latest timestamped customer note effective before 9:15 permits the picker-located item, and every item attribute stated in that note matches it. No documented safety, regulatory, or price-limit exception concerning the picker-located item requires supervisor authorization. The fulfillment record identifies the order, picker, proposed item, note chronology, and authorization review as complete. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4."}, {"path": [], "text": "The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4.", "negative_left": "At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4.", "negative_right": "The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU R2-K8.", "right": "The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-019-009", "id": "fast-41-diverse-019-009-base", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4. The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4. The store picker proposed the located item for fulfillment at that decision. The latest timestamped customer note effective before 9:15 permits the picker-located item, and every item attribute stated in that note matches it. No documented safety, regulatory, or price-limit exception concerning the picker-located item requires supervisor authorization. The fulfillment record identifies the order, picker, proposed item, note chronology, and authorization review as complete. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_exact_item"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing policy without added exceptions or defaults. Both contexts retain the 9:15 decision, order, picker, proposed-item, and note-chronology bindings while changing only case observations. Evidence consists of the exact factual quotes “At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4.”; “The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4.” The counterfactual coherently changes the requested SKU while retaining a distinct located SKU and contains no contradictory duplicate assertion. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4. The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4. The store picker proposed the located item for fulfillment at that decision. The latest timestamped customer note effective before 9:15 permits the picker-located item, and every item attribute stated in that note matches it. No documented safety, regulatory, or price-limit exception concerning the picker-located item requires supervisor authorization. The fulfillment record identifies the order, picker, proposed item, note chronology, and authorization review as complete. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4."}, {"path": [], "text": "The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4.", "negative_left": "At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4.", "negative_right": "The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU R2-K8.", "right": "The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU Q7-M4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-019-009", "id": "fast-41-diverse-019-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At the 9:15 fulfillment decision, the item located by the store picker bears SKU Q7-M4. The pickup order for the 9:15 fulfillment decision originally requested an item bearing SKU R2-K8. The store picker proposed the located item for fulfillment at that decision. The latest timestamped customer note effective before 9:15 permits the picker-located item, and every item attribute stated in that note matches it. No documented safety, regulatory, or price-limit exception concerning the picker-located item requires supervisor authorization. The fulfillment record identifies the order, picker, proposed item, note chronology, and authorization review as complete. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_proposed_substitute"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the full governing policy and choice criteria in both inputs. The 9:15 fulfillment-decision, pickup-order, picker, and item bindings remain intact while the SKU observations change. Both evidence spans are complete factual sentences: “For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472.” and “The pickup order's originally requested item was recorded under SKU-8472.” The counterfactual consistently makes SKU-8472 a substitute for SKU-3916 without contradictory duplicate measurements or assertions. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At 9:15, the store picker proposed the located item for fulfillment of the pickup order. For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472. The pickup order's originally requested item was recorded under SKU-8472. The latest timestamped customer note effective before 9:15 said, “Unsweetened soy is acceptable for this order,” and permitted the picker-located item. Every item attribute stated in that note matched the located item, including its unsweetened formulation, soy base, package size, and brand. The case file contained no documented safety exception concerning the located item requiring supervisor authorization, no documented regulatory exception requiring such authorization, and no documented price-limit exception requiring such authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472."}, {"path": [], "text": "The pickup order's originally requested item was recorded under SKU-8472."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472.", "negative_left": "For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472.", "negative_right": "The pickup order's originally requested item was recorded under SKU-3916.", "right": "The pickup order's originally requested item was recorded under SKU-8472."}, "verifier_independent_model": false}, "family": "fast-41-diverse-019-011", "id": "fast-41-diverse-019-011-base", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At 9:15, the store picker proposed the located item for fulfillment of the pickup order. For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472. The pickup order's originally requested item was recorded under SKU-8472. The latest timestamped customer note effective before 9:15 said, “Unsweetened soy is acceptable for this order,” and permitted the picker-located item. Every item attribute stated in that note matched the located item, including its unsweetened formulation, soy base, package size, and brand. The case file contained no documented safety exception concerning the located item requiring supervisor authorization, no documented regulatory exception requiring such authorization, and no documented price-limit exception requiring such authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_exact_item"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the full governing policy and choice criteria in both inputs. The 9:15 fulfillment-decision, pickup-order, picker, and item bindings remain intact while the SKU observations change. Both evidence spans are complete factual sentences: “For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472.” and “The pickup order's originally requested item was recorded under SKU-8472.” The counterfactual consistently makes SKU-8472 a substitute for SKU-3916 without contradictory duplicate measurements or assertions. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship; the universally quantified attribute-match atom remains a single relationship. The focus atom is the factual SKU-equality relation, not a policy conclusion. The base and counter assignments are jointly realizable while changing only whether the located SKU equals the originally requested SKU: both an exact item and a distinct permitted substitute can be proposed, match the latest note, and lack authorization exceptions. Policy evidence correctly preserves the two substantive rules originating in the original state; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the picker located the originally requested SKU, while refuting every documented exception category that could require supervisor review. No other policy outcome is triggered under these conditions.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a proposed replacement rather than the exact ordered item, permission from the latest effective note, matching of all attributes stated in that note, and absence of every supervisor-authorization exception. Because the proposal is permitted, the state's return-to-picking rule for barred proposals does not apply, and the rejection and customer-contact criteria are not met.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU on the item located by the store picker for the 9:15 fulfillment decision is equal to the SKU of the item originally requested in the pickup order."}, {"id": "a2", "statement": "The store picker proposed the located item for fulfillment of the pickup order at the 9:15 decision."}, {"id": "a3", "statement": "The latest timestamped customer note effective before the 9:15 fulfillment decision permits the picker-located item."}, {"id": "a4", "statement": "Every item attribute stated in the latest timestamped customer note effective before the 9:15 fulfillment decision matches the picker-located item."}, {"id": "a5", "statement": "A documented safety exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a6", "statement": "A documented regulatory exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}, {"id": "a7", "statement": "A documented price-limit exception concerning the picker-located item requires supervisor authorization for the 9:15 fulfillment decision."}], "base_state_json": "\"At 9:15, the store picker proposed the located item for fulfillment of the pickup order. For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472. The pickup order's originally requested item was recorded under SKU-8472. The latest timestamped customer note effective before 9:15 said, “Unsweetened soy is acceptable for this order,” and permitted the picker-located item. Every item attribute stated in that note matched the located item, including its unsweetened formulation, soy base, package size, and brand. The case file contained no documented safety exception concerning the located item requiring supervisor authorization, no documented regulatory exception requiring such authorization, and no documented price-limit exception requiring such authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472."}, {"path": [], "text": "The pickup order's originally requested item was recorded under SKU-8472."}], "policy_evidence": [{"path": [], "text": "Store policy says the latest timestamped note supersedes earlier notes."}, {"path": [], "text": "When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}], "rules": [{"justification": "Matching the picker scan SKU to the original order SKU establishes that the located item is the originally requested item, while the three refuted exception atoms exclude supervisor authorization requirements.", "target": "approve_exact_item", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "A proposed fulfillment item whose SKU differs from the ordered SKU is a replacement; approval is sufficient when the latest effective note permits it, every stated attribute matches, and no documented authorization exception competes.", "target": "approve_proposed_substitute", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472.", "negative_left": "For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472.", "negative_right": "The pickup order's originally requested item was recorded under SKU-3916.", "right": "The pickup order's originally requested item was recorded under SKU-8472."}, "verifier_independent_model": false}, "family": "fast-41-diverse-019-011", "id": "fast-41-diverse-019-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_exact_item": "Approve only when the picker located the originally requested item, so no substitution is being made.", "approve_proposed_substitute": "Approve only when the latest customer note permits the proposed replacement and all stated item attributes match.", "none_of_above": "Choose when policy requires an action not represented by any substantive option.", "reject_without_replacement": "Reject the item only when the latest note forbids every substitution, or when every customer-named alternative has been confirmed unavailable.", "route_customer_contact": "Send to the customer contact agent only when the latest note is ambiguous or conflicting and names no usable alternative.", "route_supervisor_review": "Send to the fulfillment supervisor only when there is a documented safety, regulatory, or price-limit exception requiring authorization."}, "instructions": "Choose the single correct route or decision under the stated policy. Apply the latest customer update and select none_of_above if the required action is not represented by a substantive option.", "type": "choice"}}, "state": "At 9:15, the store picker proposed the located item for fulfillment of the pickup order. For the 9:15 fulfillment decision, the store picker recorded the located item's SKU as SKU-8472. The pickup order's originally requested item was recorded under SKU-3916. The latest timestamped customer note effective before 9:15 said, “Unsweetened soy is acceptable for this order,” and permitted the picker-located item. Every item attribute stated in that note matched the located item, including its unsweetened formulation, soy base, package size, and brand. The case file contained no documented safety exception concerning the located item requiring supervisor authorization, no documented regulatory exception requiring such authorization, and no documented price-limit exception requiring such authorization. Store policy says the latest timestamped note supersedes earlier notes. When a proposal is barred but the latest note names an acceptable alternative with unknown availability, the order must return to picking to check that alternative; it must not yet be approved, rejected, escalated, or sent for customer contact."}, "method": "c2d", "provenance": {"source_id": "diverse-019", "source_is_synthetic": true, "source_sha256": "4f4b2714874485a6a0a17d556c22c39c786feb86a576c8771fed93844abc8fec", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_proposed_substitute"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the verbatim questions object and add no conflicting policy. Maya, the oat-milk substitution, and the decision context remain bound consistently. The evidence spans are complete factual sentences, including \"At the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \\\"unsweetened.\\\"\" and \"At the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \\\"unsweetened.\\\"\" The counterfactual coherently changes only the replacement’s recorded sweetness to \"sweetened.\" Neither context states a route, label, answer code, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified compliance and availability statements remain single relations rather than outcome classifications. The focus atom concerns the replacement’s sweetness designation, not policy. The base and counter assignments can differ only in sweetness compliance while all other facts remain fixed. Empty policy_evidence is correct because the governing precedence and routing rules are contained in the questions object, while the state supplies only case-specific observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an enabled toggle and compliance with the applicable note’s flavor and sweetness restrictions plus every other controlling restriction. They also exclude ambiguity, irreconcilability, and missing facts, so sending to picking is required.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a plausible replacement that violates the applicable item-specific sweetness restriction while satisfying the flavor and all other restrictions, with contact still possible. Ambiguity, irreconcilability, and missing facts are excluded, so customer contact is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The substitution toggle for Maya's ordered oat-milk line item is enabled at the substitution decision time."}, {"id": "a2", "statement": "The item-specific note applies to Maya's ordered oat-milk line item."}, {"id": "a3", "statement": "The applicable item-specific note permits only plain replacements for Maya's ordered oat-milk line item."}, {"id": "a4", "statement": "The sweetness designation recorded for the proposed replacement equals the sole sweetness designation permitted by the applicable item-specific note."}, {"id": "a5", "statement": "The flavor designation recorded for the proposed replacement equals the sole flavor designation permitted by the applicable item-specific note."}, {"id": "a6", "statement": "The proposed replacement satisfies every controlling restriction for Maya's ordered oat-milk line item other than its flavor and sweetness restrictions."}, {"id": "a7", "statement": "The store's substitution catalog identifies the proposed replacement as a plausible replacement for Maya's ordered oat-milk line item."}, {"id": "a8", "statement": "The customer-contact window for Maya's pickup order is open at the substitution decision time."}, {"id": "a9", "statement": "The governing records for Maya's ordered oat-milk line item are unambiguous at the substitution decision time."}, {"id": "a10", "statement": "The governing records for Maya's ordered oat-milk line item are reconcilable under the stated precedence policy."}, {"id": "a11", "statement": "Every fact required to route the proposed replacement is available at the substitution decision time."}], "base_state_json": "{\"case_note\":\"At 14:10 on 17 September 2026, Maya’s ordered oat-milk line item was being reviewed for substitution. The substitution toggle was enabled, and the applicable item-specific note applied to this line item and allowed only plain replacements. The proposed replacement matched the note’s sole permitted flavor designation and satisfied every other controlling restriction for the line item apart from the flavor and sweetness checks. The store catalog marked it as a plausible replacement. Maya’s pickup-order customer-contact window was open. The governing records were unambiguous and could be reconciled under the stated precedence policy, and every fact needed to route the proposal was available. The following two recorded observations were verified independently:\\n\\nAt the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \\\"unsweetened.\\\"\\n\\nAt the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \\\"unsweetened.\\\"\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["case_note"], "text": "At the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \"unsweetened.\""}, {"path": ["case_note"], "text": "At the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \"unsweetened.\""}], "policy_evidence": [], "rules": [{"justification": "The enabled toggle permits the proposal because it satisfies the applicable note's flavor and sweetness restrictions and every other controlling restriction; the records are unambiguous and reconcilable, and no required fact is missing.", "target": "send_to_picking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The catalog identifies an otherwise compliant proposal as plausible, but its recorded sweetness designation violates the clear applicable item-specific restriction while contact remains possible; ambiguity, irreconcilability, and missing required facts are excluded.", "target": "contact_customer_for_exception", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \"unsweetened.\"", "negative_left": "At the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \"unsweetened.\"", "negative_right": "At the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \"sweetened.\"", "right": "At the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \"unsweetened.\""}, "verifier_independent_model": false}, "family": "fast-41-diverse-020-005", "id": "fast-41-diverse-020-005-base", "input": {"questions": {"decision": {"criteria": {"contact_customer_for_exception": "Pause fulfillment and ask the pickup customer whether to accept this otherwise plausible replacement because it violates a clear controlling preference while contact is still possible.", "send_to_picking": "Approve the proposed replacement for picking because it satisfies every controlling item restriction.", "send_to_supervisor_review": "Escalate because the governing records cannot be reconciled using the stated precedence policy or necessary facts are unavailable."}, "instructions": "An item-specific note overrides a general checkout note. A substitution toggle permits only replacements satisfying applicable notes. Route to picking when the proposal meets all controlling restrictions. Route to customer contact when a plausible replacement violates a clear preference and the contact window remains open. Route to supervisor review only when controlling records are ambiguous, irreconcilable under the policy, or required facts are missing.", "type": "choice"}}, "state": {"case_note": "At 14:10 on 17 September 2026, Maya’s ordered oat-milk line item was being reviewed for substitution. The substitution toggle was enabled, and the applicable item-specific note applied to this line item and allowed only plain replacements. The proposed replacement matched the note’s sole permitted flavor designation and satisfied every other controlling restriction for the line item apart from the flavor and sweetness checks. The store catalog marked it as a plausible replacement. Maya’s pickup-order customer-contact window was open. The governing records were unambiguous and could be reconciled under the stated precedence policy, and every fact needed to route the proposal was available. The following two recorded observations were verified independently:\n\nAt the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \"unsweetened.\"\n\nAt the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \"unsweetened.\""}}, "method": "c2d", "provenance": {"source_id": "diverse-020", "source_is_synthetic": true, "source_sha256": "dbd1660f2d346102df74fed3269c0b8d0a77e59242a230ad55e6a6552e7f9753", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "send_to_picking"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the verbatim questions object and add no conflicting policy. Maya, the oat-milk substitution, and the decision context remain bound consistently. The evidence spans are complete factual sentences, including \"At the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \\\"unsweetened.\\\"\" and \"At the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \\\"unsweetened.\\\"\" The counterfactual coherently changes only the replacement’s recorded sweetness to \"sweetened.\" Neither context states a route, label, answer code, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified compliance and availability statements remain single relations rather than outcome classifications. The focus atom concerns the replacement’s sweetness designation, not policy. The base and counter assignments can differ only in sweetness compliance while all other facts remain fixed. Empty policy_evidence is correct because the governing precedence and routing rules are contained in the questions object, while the state supplies only case-specific observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an enabled toggle and compliance with the applicable note’s flavor and sweetness restrictions plus every other controlling restriction. They also exclude ambiguity, irreconcilability, and missing facts, so sending to picking is required.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a plausible replacement that violates the applicable item-specific sweetness restriction while satisfying the flavor and all other restrictions, with contact still possible. Ambiguity, irreconcilability, and missing facts are excluded, so customer contact is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The substitution toggle for Maya's ordered oat-milk line item is enabled at the substitution decision time."}, {"id": "a2", "statement": "The item-specific note applies to Maya's ordered oat-milk line item."}, {"id": "a3", "statement": "The applicable item-specific note permits only plain replacements for Maya's ordered oat-milk line item."}, {"id": "a4", "statement": "The sweetness designation recorded for the proposed replacement equals the sole sweetness designation permitted by the applicable item-specific note."}, {"id": "a5", "statement": "The flavor designation recorded for the proposed replacement equals the sole flavor designation permitted by the applicable item-specific note."}, {"id": "a6", "statement": "The proposed replacement satisfies every controlling restriction for Maya's ordered oat-milk line item other than its flavor and sweetness restrictions."}, {"id": "a7", "statement": "The store's substitution catalog identifies the proposed replacement as a plausible replacement for Maya's ordered oat-milk line item."}, {"id": "a8", "statement": "The customer-contact window for Maya's pickup order is open at the substitution decision time."}, {"id": "a9", "statement": "The governing records for Maya's ordered oat-milk line item are unambiguous at the substitution decision time."}, {"id": "a10", "statement": "The governing records for Maya's ordered oat-milk line item are reconcilable under the stated precedence policy."}, {"id": "a11", "statement": "Every fact required to route the proposed replacement is available at the substitution decision time."}], "base_state_json": "{\"case_note\":\"At 14:10 on 17 September 2026, Maya’s ordered oat-milk line item was being reviewed for substitution. The substitution toggle was enabled, and the applicable item-specific note applied to this line item and allowed only plain replacements. The proposed replacement matched the note’s sole permitted flavor designation and satisfied every other controlling restriction for the line item apart from the flavor and sweetness checks. The store catalog marked it as a plausible replacement. Maya’s pickup-order customer-contact window was open. The governing records were unambiguous and could be reconciled under the stated precedence policy, and every fact needed to route the proposal was available. The following two recorded observations were verified independently:\\n\\nAt the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \\\"unsweetened.\\\"\\n\\nAt the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \\\"unsweetened.\\\"\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["case_note"], "text": "At the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \"unsweetened.\""}, {"path": ["case_note"], "text": "At the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \"unsweetened.\""}], "policy_evidence": [], "rules": [{"justification": "The enabled toggle permits the proposal because it satisfies the applicable note's flavor and sweetness restrictions and every other controlling restriction; the records are unambiguous and reconcilable, and no required fact is missing.", "target": "send_to_picking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The catalog identifies an otherwise compliant proposal as plausible, but its recorded sweetness designation violates the clear applicable item-specific restriction while contact remains possible; ambiguity, irreconcilability, and missing required facts are excluded.", "target": "contact_customer_for_exception", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \"unsweetened.\"", "negative_left": "At the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \"unsweetened.\"", "negative_right": "At the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \"sweetened.\"", "right": "At the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \"unsweetened.\""}, "verifier_independent_model": false}, "family": "fast-41-diverse-020-005", "id": "fast-41-diverse-020-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"contact_customer_for_exception": "Pause fulfillment and ask the pickup customer whether to accept this otherwise plausible replacement because it violates a clear controlling preference while contact is still possible.", "send_to_picking": "Approve the proposed replacement for picking because it satisfies every controlling item restriction.", "send_to_supervisor_review": "Escalate because the governing records cannot be reconciled using the stated precedence policy or necessary facts are unavailable."}, "instructions": "An item-specific note overrides a general checkout note. A substitution toggle permits only replacements satisfying applicable notes. Route to picking when the proposal meets all controlling restrictions. Route to customer contact when a plausible replacement violates a clear preference and the contact window remains open. Route to supervisor review only when controlling records are ambiguous, irreconcilable under the policy, or required facts are missing.", "type": "choice"}}, "state": {"case_note": "At 14:10 on 17 September 2026, Maya’s ordered oat-milk line item was being reviewed for substitution. The substitution toggle was enabled, and the applicable item-specific note applied to this line item and allowed only plain replacements. The proposed replacement matched the note’s sole permitted flavor designation and satisfied every other controlling restriction for the line item apart from the flavor and sweetness checks. The store catalog marked it as a plausible replacement. Maya’s pickup-order customer-contact window was open. The governing records were unambiguous and could be reconciled under the stated precedence policy, and every fact needed to route the proposal was available. The following two recorded observations were verified independently:\n\nAt the 14:10 substitution decision time on 17 September 2026, the applicable item-specific note for Maya's ordered oat-milk line item permits only the sweetness designation \"unsweetened.\"\n\nAt the 14:10 substitution decision time on 17 September 2026, the proposed replacement for Maya's ordered oat-milk line item has the recorded sweetness designation \"sweetened.\""}}, "method": "c2d", "provenance": {"source_id": "diverse-020", "source_is_synthetic": true, "source_sha256": "dbd1660f2d346102df74fed3269c0b8d0a77e59242a230ad55e6a6552e7f9753", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "contact_customer_for_exception"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy bindings, contain two factual evidence sentences, and coherently change only the proposed replacement’s sweetness from “unsweetened” to “sweetened” without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified compliance and availability statements remain single relations rather than outcome classifications. The focus atom concerns the replacement’s sweetness designation, not policy. The base and counter assignments can differ only in sweetness compliance while all other facts remain fixed. Empty policy_evidence is correct because the governing precedence and routing rules are contained in the questions object, while the state supplies only case-specific observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an enabled toggle and compliance with the applicable note’s flavor and sweetness restrictions plus every other controlling restriction. They also exclude ambiguity, irreconcilability, and missing facts, so sending to picking is required.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a plausible replacement that violates the applicable item-specific sweetness restriction while satisfying the flavor and all other restrictions, with contact still possible. Ambiguity, irreconcilability, and missing facts are excluded, so customer contact is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The substitution toggle for Maya's ordered oat-milk line item is enabled at the substitution decision time."}, {"id": "a2", "statement": "The item-specific note applies to Maya's ordered oat-milk line item."}, {"id": "a3", "statement": "The applicable item-specific note permits only plain replacements for Maya's ordered oat-milk line item."}, {"id": "a4", "statement": "The sweetness designation recorded for the proposed replacement equals the sole sweetness designation permitted by the applicable item-specific note."}, {"id": "a5", "statement": "The flavor designation recorded for the proposed replacement equals the sole flavor designation permitted by the applicable item-specific note."}, {"id": "a6", "statement": "The proposed replacement satisfies every controlling restriction for Maya's ordered oat-milk line item other than its flavor and sweetness restrictions."}, {"id": "a7", "statement": "The store's substitution catalog identifies the proposed replacement as a plausible replacement for Maya's ordered oat-milk line item."}, {"id": "a8", "statement": "The customer-contact window for Maya's pickup order is open at the substitution decision time."}, {"id": "a9", "statement": "The governing records for Maya's ordered oat-milk line item are unambiguous at the substitution decision time."}, {"id": "a10", "statement": "The governing records for Maya's ordered oat-milk line item are reconcilable under the stated precedence policy."}, {"id": "a11", "statement": "Every fact required to route the proposed replacement is available at the substitution decision time."}], "base_state_json": "{\"context\":\"Case note: At the substitution decision time, Maya's ordered oat-milk line item has its substitution toggle enabled, and the item-specific note applies to that line item. The note permits only plain replacements. The proposed replacement's recorded flavor designation matches the sole flavor designation permitted by that note. Apart from flavor and sweetness restrictions, the replacement satisfies every controlling restriction for the line item. The store catalog identifies it as a plausible replacement. Maya's pickup-order customer-contact window is open. The governing records are unambiguous and can be reconciled under the stated precedence policy, and every fact needed to route the proposal is available. The order remains associated with Maya and the same proposed replacement throughout the review.\",\"evidence\":[\"At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'.\",\"At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'unsweetened'.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'."}, {"path": ["evidence", "1"], "text": "At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'unsweetened'."}], "policy_evidence": [], "rules": [{"justification": "The enabled toggle permits the proposal because it satisfies the applicable note's flavor and sweetness restrictions and every other controlling restriction; the records are unambiguous and reconcilable, and no required fact is missing.", "target": "send_to_picking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The catalog identifies an otherwise compliant proposal as plausible, but its recorded sweetness designation violates the clear applicable item-specific restriction while contact remains possible; ambiguity, irreconcilability, and missing required facts are excluded.", "target": "contact_customer_for_exception", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'.", "negative_left": "At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'.", "negative_right": "At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'sweetened'.", "right": "At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'unsweetened'."}, "verifier_independent_model": false}, "family": "fast-41-diverse-020-008", "id": "fast-41-diverse-020-008-base", "input": {"questions": {"decision": {"criteria": {"contact_customer_for_exception": "Pause fulfillment and ask the pickup customer whether to accept this otherwise plausible replacement because it violates a clear controlling preference while contact is still possible.", "send_to_picking": "Approve the proposed replacement for picking because it satisfies every controlling item restriction.", "send_to_supervisor_review": "Escalate because the governing records cannot be reconciled using the stated precedence policy or necessary facts are unavailable."}, "instructions": "An item-specific note overrides a general checkout note. A substitution toggle permits only replacements satisfying applicable notes. Route to picking when the proposal meets all controlling restrictions. Route to customer contact when a plausible replacement violates a clear preference and the contact window remains open. Route to supervisor review only when controlling records are ambiguous, irreconcilable under the policy, or required facts are missing.", "type": "choice"}}, "state": {"context": "Case note: At the substitution decision time, Maya's ordered oat-milk line item has its substitution toggle enabled, and the item-specific note applies to that line item. The note permits only plain replacements. The proposed replacement's recorded flavor designation matches the sole flavor designation permitted by that note. Apart from flavor and sweetness restrictions, the replacement satisfies every controlling restriction for the line item. The store catalog identifies it as a plausible replacement. Maya's pickup-order customer-contact window is open. The governing records are unambiguous and can be reconciled under the stated precedence policy, and every fact needed to route the proposal is available. The order remains associated with Maya and the same proposed replacement throughout the review.", "evidence": ["At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'.", "At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'unsweetened'."]}}, "method": "c2d", "provenance": {"source_id": "diverse-020", "source_is_synthetic": true, "source_sha256": "dbd1660f2d346102df74fed3269c0b8d0a77e59242a230ad55e6a6552e7f9753", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "send_to_picking"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy bindings, contain two factual evidence sentences, and coherently change only the proposed replacement’s sweetness from “unsweetened” to “sweetened” without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified compliance and availability statements remain single relations rather than outcome classifications. The focus atom concerns the replacement’s sweetness designation, not policy. The base and counter assignments can differ only in sweetness compliance while all other facts remain fixed. Empty policy_evidence is correct because the governing precedence and routing rules are contained in the questions object, while the state supplies only case-specific observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an enabled toggle and compliance with the applicable note’s flavor and sweetness restrictions plus every other controlling restriction. They also exclude ambiguity, irreconcilability, and missing facts, so sending to picking is required.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a plausible replacement that violates the applicable item-specific sweetness restriction while satisfying the flavor and all other restrictions, with contact still possible. Ambiguity, irreconcilability, and missing facts are excluded, so customer contact is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The substitution toggle for Maya's ordered oat-milk line item is enabled at the substitution decision time."}, {"id": "a2", "statement": "The item-specific note applies to Maya's ordered oat-milk line item."}, {"id": "a3", "statement": "The applicable item-specific note permits only plain replacements for Maya's ordered oat-milk line item."}, {"id": "a4", "statement": "The sweetness designation recorded for the proposed replacement equals the sole sweetness designation permitted by the applicable item-specific note."}, {"id": "a5", "statement": "The flavor designation recorded for the proposed replacement equals the sole flavor designation permitted by the applicable item-specific note."}, {"id": "a6", "statement": "The proposed replacement satisfies every controlling restriction for Maya's ordered oat-milk line item other than its flavor and sweetness restrictions."}, {"id": "a7", "statement": "The store's substitution catalog identifies the proposed replacement as a plausible replacement for Maya's ordered oat-milk line item."}, {"id": "a8", "statement": "The customer-contact window for Maya's pickup order is open at the substitution decision time."}, {"id": "a9", "statement": "The governing records for Maya's ordered oat-milk line item are unambiguous at the substitution decision time."}, {"id": "a10", "statement": "The governing records for Maya's ordered oat-milk line item are reconcilable under the stated precedence policy."}, {"id": "a11", "statement": "Every fact required to route the proposed replacement is available at the substitution decision time."}], "base_state_json": "{\"context\":\"Case note: At the substitution decision time, Maya's ordered oat-milk line item has its substitution toggle enabled, and the item-specific note applies to that line item. The note permits only plain replacements. The proposed replacement's recorded flavor designation matches the sole flavor designation permitted by that note. Apart from flavor and sweetness restrictions, the replacement satisfies every controlling restriction for the line item. The store catalog identifies it as a plausible replacement. Maya's pickup-order customer-contact window is open. The governing records are unambiguous and can be reconciled under the stated precedence policy, and every fact needed to route the proposal is available. The order remains associated with Maya and the same proposed replacement throughout the review.\",\"evidence\":[\"At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'.\",\"At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'unsweetened'.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'."}, {"path": ["evidence", "1"], "text": "At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'unsweetened'."}], "policy_evidence": [], "rules": [{"justification": "The enabled toggle permits the proposal because it satisfies the applicable note's flavor and sweetness restrictions and every other controlling restriction; the records are unambiguous and reconcilable, and no required fact is missing.", "target": "send_to_picking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The catalog identifies an otherwise compliant proposal as plausible, but its recorded sweetness designation violates the clear applicable item-specific restriction while contact remains possible; ambiguity, irreconcilability, and missing required facts are excluded.", "target": "contact_customer_for_exception", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'.", "negative_left": "At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'.", "negative_right": "At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'sweetened'.", "right": "At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'unsweetened'."}, "verifier_independent_model": false}, "family": "fast-41-diverse-020-008", "id": "fast-41-diverse-020-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"contact_customer_for_exception": "Pause fulfillment and ask the pickup customer whether to accept this otherwise plausible replacement because it violates a clear controlling preference while contact is still possible.", "send_to_picking": "Approve the proposed replacement for picking because it satisfies every controlling item restriction.", "send_to_supervisor_review": "Escalate because the governing records cannot be reconciled using the stated precedence policy or necessary facts are unavailable."}, "instructions": "An item-specific note overrides a general checkout note. A substitution toggle permits only replacements satisfying applicable notes. Route to picking when the proposal meets all controlling restrictions. Route to customer contact when a plausible replacement violates a clear preference and the contact window remains open. Route to supervisor review only when controlling records are ambiguous, irreconcilable under the policy, or required facts are missing.", "type": "choice"}}, "state": {"context": "Case note: At the substitution decision time, Maya's ordered oat-milk line item has its substitution toggle enabled, and the item-specific note applies to that line item. The note permits only plain replacements. The proposed replacement's recorded flavor designation matches the sole flavor designation permitted by that note. Apart from flavor and sweetness restrictions, the replacement satisfies every controlling restriction for the line item. The store catalog identifies it as a plausible replacement. Maya's pickup-order customer-contact window is open. The governing records are unambiguous and can be reconciled under the stated precedence policy, and every fact needed to route the proposal is available. The order remains associated with Maya and the same proposed replacement throughout the review.", "evidence": ["At the substitution decision time, the applicable item-specific note for Maya's ordered oat-milk line item permits exactly the sweetness designation 'unsweetened'.", "At the substitution decision time, the sweetness designation recorded for the proposed replacement for Maya's ordered oat-milk line item is 'sweetened'."]}}, "method": "c2d", "provenance": {"source_id": "diverse-020", "source_is_synthetic": true, "source_sha256": "dbd1660f2d346102df74fed3269c0b8d0a77e59242a230ad55e6a6552e7f9753", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "contact_customer_for_exception"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy, and neither context alters it. Both contexts remain bound to Maya’s oat-milk substitution request and decision setting. The two evidence spans are complete factual sentences and retain the required exact wording. The counterfactual changes only sweetness and consistently preserves the applicable unsweetened restriction and all other facts. Neither context contains a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified compliance and availability statements remain single relations rather than outcome classifications. The focus atom concerns the replacement’s sweetness designation, not policy. The base and counter assignments can differ only in sweetness compliance while all other facts remain fixed. Empty policy_evidence is correct because the governing precedence and routing rules are contained in the questions object, while the state supplies only case-specific observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an enabled toggle and compliance with the applicable note’s flavor and sweetness restrictions plus every other controlling restriction. They also exclude ambiguity, irreconcilability, and missing facts, so sending to picking is required.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a plausible replacement that violates the applicable item-specific sweetness restriction while satisfying the flavor and all other restrictions, with contact still possible. Ambiguity, irreconcilability, and missing facts are excluded, so customer contact is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The substitution toggle for Maya's ordered oat-milk line item is enabled at the substitution decision time."}, {"id": "a2", "statement": "The item-specific note applies to Maya's ordered oat-milk line item."}, {"id": "a3", "statement": "The applicable item-specific note permits only plain replacements for Maya's ordered oat-milk line item."}, {"id": "a4", "statement": "The sweetness designation recorded for the proposed replacement equals the sole sweetness designation permitted by the applicable item-specific note."}, {"id": "a5", "statement": "The flavor designation recorded for the proposed replacement equals the sole flavor designation permitted by the applicable item-specific note."}, {"id": "a6", "statement": "The proposed replacement satisfies every controlling restriction for Maya's ordered oat-milk line item other than its flavor and sweetness restrictions."}, {"id": "a7", "statement": "The store's substitution catalog identifies the proposed replacement as a plausible replacement for Maya's ordered oat-milk line item."}, {"id": "a8", "statement": "The customer-contact window for Maya's pickup order is open at the substitution decision time."}, {"id": "a9", "statement": "The governing records for Maya's ordered oat-milk line item are unambiguous at the substitution decision time."}, {"id": "a10", "statement": "The governing records for Maya's ordered oat-milk line item are reconcilable under the stated precedence policy."}, {"id": "a11", "statement": "Every fact required to route the proposed replacement is available at the substitution decision time."}], "base_state_json": "{\"context\":\"Case note — 2026-09-17, 14:00 UTC: Maya’s ordered oat-milk line item is being reviewed while the requested product is unavailable. The substitution toggle is enabled, and the item-specific note applies to this line item. That note permits only plain replacements. The proposed replacement is the same size and brand, and its recorded flavor designation matches the sole flavor designation permitted by the applicable note. It satisfies every other controlling restriction. The store’s substitution catalog identifies it as a plausible replacement. Maya’s pickup-order contact window remains open. The governing records are unambiguous and reconcilable under the stated precedence policy, and every fact required for routing is available. At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is unsweetened. At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation. If the proposed replacement’s recorded sweetness designation were instead different from the designation permitted by the note, all other recorded circumstances would remain unchanged.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is unsweetened."}, {"path": ["context"], "text": "At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation."}], "policy_evidence": [], "rules": [{"justification": "The enabled toggle permits the proposal because it satisfies the applicable note's flavor and sweetness restrictions and every other controlling restriction; the records are unambiguous and reconcilable, and no required fact is missing.", "target": "send_to_picking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The catalog identifies an otherwise compliant proposal as plausible, but its recorded sweetness designation violates the clear applicable item-specific restriction while contact remains possible; ambiguity, irreconcilability, and missing required facts are excluded.", "target": "contact_customer_for_exception", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is unsweetened.", "negative_left": "At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is sweetened.", "negative_right": "At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation.", "right": "At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation."}, "verifier_independent_model": false}, "family": "fast-41-diverse-020-010", "id": "fast-41-diverse-020-010-base", "input": {"questions": {"decision": {"criteria": {"contact_customer_for_exception": "Pause fulfillment and ask the pickup customer whether to accept this otherwise plausible replacement because it violates a clear controlling preference while contact is still possible.", "send_to_picking": "Approve the proposed replacement for picking because it satisfies every controlling item restriction.", "send_to_supervisor_review": "Escalate because the governing records cannot be reconciled using the stated precedence policy or necessary facts are unavailable."}, "instructions": "An item-specific note overrides a general checkout note. A substitution toggle permits only replacements satisfying applicable notes. Route to picking when the proposal meets all controlling restrictions. Route to customer contact when a plausible replacement violates a clear preference and the contact window remains open. Route to supervisor review only when controlling records are ambiguous, irreconcilable under the policy, or required facts are missing.", "type": "choice"}}, "state": {"context": "Case note — 2026-09-17, 14:00 UTC: Maya’s ordered oat-milk line item is being reviewed while the requested product is unavailable. The substitution toggle is enabled, and the item-specific note applies to this line item. That note permits only plain replacements. The proposed replacement is the same size and brand, and its recorded flavor designation matches the sole flavor designation permitted by the applicable note. It satisfies every other controlling restriction. The store’s substitution catalog identifies it as a plausible replacement. Maya’s pickup-order contact window remains open. The governing records are unambiguous and reconcilable under the stated precedence policy, and every fact required for routing is available. At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is unsweetened. At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation. If the proposed replacement’s recorded sweetness designation were instead different from the designation permitted by the note, all other recorded circumstances would remain unchanged."}}, "method": "c2d", "provenance": {"source_id": "diverse-020", "source_is_synthetic": true, "source_sha256": "dbd1660f2d346102df74fed3269c0b8d0a77e59242a230ad55e6a6552e7f9753", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "send_to_picking"}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy, and neither context alters it. Both contexts remain bound to Maya’s oat-milk substitution request and decision setting. The two evidence spans are complete factual sentences and retain the required exact wording. The counterfactual changes only sweetness and consistently preserves the applicable unsweetened restriction and all other facts. Neither context contains a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified compliance and availability statements remain single relations rather than outcome classifications. The focus atom concerns the replacement’s sweetness designation, not policy. The base and counter assignments can differ only in sweetness compliance while all other facts remain fixed. Empty policy_evidence is correct because the governing precedence and routing rules are contained in the questions object, while the state supplies only case-specific observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an enabled toggle and compliance with the applicable note’s flavor and sweetness restrictions plus every other controlling restriction. They also exclude ambiguity, irreconcilability, and missing facts, so sending to picking is required.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a plausible replacement that violates the applicable item-specific sweetness restriction while satisfying the flavor and all other restrictions, with contact still possible. Ambiguity, irreconcilability, and missing facts are excluded, so customer contact is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The substitution toggle for Maya's ordered oat-milk line item is enabled at the substitution decision time."}, {"id": "a2", "statement": "The item-specific note applies to Maya's ordered oat-milk line item."}, {"id": "a3", "statement": "The applicable item-specific note permits only plain replacements for Maya's ordered oat-milk line item."}, {"id": "a4", "statement": "The sweetness designation recorded for the proposed replacement equals the sole sweetness designation permitted by the applicable item-specific note."}, {"id": "a5", "statement": "The flavor designation recorded for the proposed replacement equals the sole flavor designation permitted by the applicable item-specific note."}, {"id": "a6", "statement": "The proposed replacement satisfies every controlling restriction for Maya's ordered oat-milk line item other than its flavor and sweetness restrictions."}, {"id": "a7", "statement": "The store's substitution catalog identifies the proposed replacement as a plausible replacement for Maya's ordered oat-milk line item."}, {"id": "a8", "statement": "The customer-contact window for Maya's pickup order is open at the substitution decision time."}, {"id": "a9", "statement": "The governing records for Maya's ordered oat-milk line item are unambiguous at the substitution decision time."}, {"id": "a10", "statement": "The governing records for Maya's ordered oat-milk line item are reconcilable under the stated precedence policy."}, {"id": "a11", "statement": "Every fact required to route the proposed replacement is available at the substitution decision time."}], "base_state_json": "{\"context\":\"Case note — 2026-09-17, 14:00 UTC: Maya’s ordered oat-milk line item is being reviewed while the requested product is unavailable. The substitution toggle is enabled, and the item-specific note applies to this line item. That note permits only plain replacements. The proposed replacement is the same size and brand, and its recorded flavor designation matches the sole flavor designation permitted by the applicable note. It satisfies every other controlling restriction. The store’s substitution catalog identifies it as a plausible replacement. Maya’s pickup-order contact window remains open. The governing records are unambiguous and reconcilable under the stated precedence policy, and every fact required for routing is available. At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is unsweetened. At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation. If the proposed replacement’s recorded sweetness designation were instead different from the designation permitted by the note, all other recorded circumstances would remain unchanged.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is unsweetened."}, {"path": ["context"], "text": "At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation."}], "policy_evidence": [], "rules": [{"justification": "The enabled toggle permits the proposal because it satisfies the applicable note's flavor and sweetness restrictions and every other controlling restriction; the records are unambiguous and reconcilable, and no required fact is missing.", "target": "send_to_picking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The catalog identifies an otherwise compliant proposal as plausible, but its recorded sweetness designation violates the clear applicable item-specific restriction while contact remains possible; ambiguity, irreconcilability, and missing required facts are excluded.", "target": "contact_customer_for_exception", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is unsweetened.", "negative_left": "At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is sweetened.", "negative_right": "At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation.", "right": "At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation."}, "verifier_independent_model": false}, "family": "fast-41-diverse-020-010", "id": "fast-41-diverse-020-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"contact_customer_for_exception": "Pause fulfillment and ask the pickup customer whether to accept this otherwise plausible replacement because it violates a clear controlling preference while contact is still possible.", "send_to_picking": "Approve the proposed replacement for picking because it satisfies every controlling item restriction.", "send_to_supervisor_review": "Escalate because the governing records cannot be reconciled using the stated precedence policy or necessary facts are unavailable."}, "instructions": "An item-specific note overrides a general checkout note. A substitution toggle permits only replacements satisfying applicable notes. Route to picking when the proposal meets all controlling restrictions. Route to customer contact when a plausible replacement violates a clear preference and the contact window remains open. Route to supervisor review only when controlling records are ambiguous, irreconcilable under the policy, or required facts are missing.", "type": "choice"}}, "state": {"context": "Case note — 2026-09-17, 14:00 UTC: Maya’s ordered oat-milk line item is being reviewed while the requested product is unavailable. The substitution toggle is enabled, and the item-specific note applies to this line item. That note permits only plain replacements. The proposed replacement is the same size and brand, and its recorded flavor designation matches the sole flavor designation permitted by the applicable note. It satisfies every other controlling restriction. The store’s substitution catalog identifies it as a plausible replacement. Maya’s pickup-order contact window remains open. The governing records are unambiguous and reconcilable under the stated precedence policy, and every fact required for routing is available. At the 2026-09-17 14:00 UTC substitution decision for Maya's ordered oat-milk line item, the proposed replacement's recorded sweetness designation is sweetened. At the 2026-09-17 14:00 UTC substitution decision, the applicable item-specific note for Maya's ordered oat-milk line item permits unsweetened as its sole sweetness designation. If the proposed replacement’s recorded sweetness designation were instead different from the designation permitted by the note, all other recorded circumstances would remain unchanged."}}, "method": "c2d", "provenance": {"source_id": "diverse-020", "source_is_synthetic": true, "source_sha256": "dbd1660f2d346102df74fed3269c0b8d0a77e59242a230ad55e6a6552e7f9753", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "contact_customer_for_exception"}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all relevant product, approval, waiver, and timing bindings. The focus evidence consists of two complete factual sentences: “In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512.” and “On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”” The counterfactual coherently changes only the package-label fact while preserving the barcode linkage and other observations. Neither context states a gold answer, answer code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”\"},{\"speaker\":\"Store picker\",\"text\":\"The requested North Mill creamy peanut butter is unavailable. The proposed substitute is North Mill smooth peanut butter. Shelf records identify both products as North Mill, and the proposed jar's net quantity is exactly 16 ounces.\"},{\"speaker\":\"Customer contact agent\",\"text\":\"A call and text were sent, but the customer did not respond, so no waiver was received.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"The pickup window closes in 12 minutes; review the documented package information for approval.\"},{\"speaker\":\"Pickup-order record\",\"text\":\"In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512.\"},{\"speaker\":\"Package record\",\"text\":\"On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["4", "text"], "text": "In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512."}, {"path": ["5", "text"], "text": "On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”"}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512.", "negative_left": "In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512.", "negative_right": "On 2026-09-17, the package bearing barcode 074231806512 does not bear a label reading “no added sugar.”", "right": "On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-021-001", "id": "fast-41-diverse-021-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Store picker", "text": "The requested North Mill creamy peanut butter is unavailable. The proposed substitute is North Mill smooth peanut butter. Shelf records identify both products as North Mill, and the proposed jar's net quantity is exactly 16 ounces."}, {"speaker": "Customer contact agent", "text": "A call and text were sent, but the customer did not respond, so no waiver was received."}, {"speaker": "Fulfillment supervisor", "text": "The pickup window closes in 12 minutes; review the documented package information for approval."}, {"speaker": "Pickup-order record", "text": "In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512."}, {"speaker": "Package record", "text": "On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”"}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all relevant product, approval, waiver, and timing bindings. The focus evidence consists of two complete factual sentences: “In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512.” and “On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”” The counterfactual coherently changes only the package-label fact while preserving the barcode linkage and other observations. Neither context states a gold answer, answer code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”\"},{\"speaker\":\"Store picker\",\"text\":\"The requested North Mill creamy peanut butter is unavailable. The proposed substitute is North Mill smooth peanut butter. Shelf records identify both products as North Mill, and the proposed jar's net quantity is exactly 16 ounces.\"},{\"speaker\":\"Customer contact agent\",\"text\":\"A call and text were sent, but the customer did not respond, so no waiver was received.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"The pickup window closes in 12 minutes; review the documented package information for approval.\"},{\"speaker\":\"Pickup-order record\",\"text\":\"In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512.\"},{\"speaker\":\"Package record\",\"text\":\"On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["4", "text"], "text": "In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512."}, {"path": ["5", "text"], "text": "On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”"}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512.", "negative_left": "In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512.", "negative_right": "On 2026-09-17, the package bearing barcode 074231806512 does not bear a label reading “no added sugar.”", "right": "On 2026-09-17, the package bearing barcode 074231806512 bears a label reading “no added sugar.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-021-001", "id": "fast-41-diverse-021-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Store picker", "text": "The requested North Mill creamy peanut butter is unavailable. The proposed substitute is North Mill smooth peanut butter. Shelf records identify both products as North Mill, and the proposed jar's net quantity is exactly 16 ounces."}, {"speaker": "Customer contact agent", "text": "A call and text were sent, but the customer did not respond, so no waiver was received."}, {"speaker": "Fulfillment supervisor", "text": "The pickup window closes in 12 minutes; review the documented package information for approval."}, {"speaker": "Pickup-order record", "text": "In the 2026-09-17 pickup-order record, the proposed North Mill smooth peanut butter substitution is assigned package barcode 074231806512."}, {"speaker": "Package record", "text": "On 2026-09-17, the package bearing barcode 074231806512 does not bear a label reading “no added sugar.”"}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, and scope. Both contexts retain the proposed North Mill substitution, package barcode, customer-waiver status, and pickup-window timing. The two focus spans are complete factual sentences and retain the exact evidence quotes “no added sugar” and “contains added sugar.” The counterfactual changes only the package wording and remains consistent with the same barcode and unchanged observations. Neither context states a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”\"},{\"speaker\":\"Store picker\",\"text\":\"The requested North Mill creamy peanut butter is unavailable. The proposed North Mill smooth peanut butter is from the same brand, and its jar is marked with a net quantity of exactly 16 ounces.\"},{\"speaker\":\"Store picker\",\"text\":\"The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516.\"},{\"speaker\":\"Store picker\",\"text\":\"The package bearing barcode 074321908516 displays the exact label text “no added sugar.”\"},{\"speaker\":\"Customer contact agent\",\"text\":\"Phone and text outreach received no customer response, so no additional waiver or changed instruction was recorded.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"The replacement is under review before the pickup window closes. If the printed wording on that same package were changed, all other recorded product and contact observations would remain unchanged.\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["2", "text"], "text": "The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516."}, {"path": ["3", "text"], "text": "The package bearing barcode 074321908516 displays the exact label text “no added sugar.”"}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516.", "negative_left": "The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516.", "negative_right": "The package bearing barcode 074321908516 displays the exact label text “contains added sugar.”", "right": "The package bearing barcode 074321908516 displays the exact label text “no added sugar.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-021-004", "id": "fast-41-diverse-021-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Store picker", "text": "The requested North Mill creamy peanut butter is unavailable. The proposed North Mill smooth peanut butter is from the same brand, and its jar is marked with a net quantity of exactly 16 ounces."}, {"speaker": "Store picker", "text": "The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516."}, {"speaker": "Store picker", "text": "The package bearing barcode 074321908516 displays the exact label text “no added sugar.”"}, {"speaker": "Customer contact agent", "text": "Phone and text outreach received no customer response, so no additional waiver or changed instruction was recorded."}, {"speaker": "Fulfillment supervisor", "text": "The replacement is under review before the pickup window closes. If the printed wording on that same package were changed, all other recorded product and contact observations would remain unchanged."}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, and scope. Both contexts retain the proposed North Mill substitution, package barcode, customer-waiver status, and pickup-window timing. The two focus spans are complete factual sentences and retain the exact evidence quotes “no added sugar” and “contains added sugar.” The counterfactual changes only the package wording and remains consistent with the same barcode and unchanged observations. Neither context states a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”\"},{\"speaker\":\"Store picker\",\"text\":\"The requested North Mill creamy peanut butter is unavailable. The proposed North Mill smooth peanut butter is from the same brand, and its jar is marked with a net quantity of exactly 16 ounces.\"},{\"speaker\":\"Store picker\",\"text\":\"The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516.\"},{\"speaker\":\"Store picker\",\"text\":\"The package bearing barcode 074321908516 displays the exact label text “no added sugar.”\"},{\"speaker\":\"Customer contact agent\",\"text\":\"Phone and text outreach received no customer response, so no additional waiver or changed instruction was recorded.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"The replacement is under review before the pickup window closes. If the printed wording on that same package were changed, all other recorded product and contact observations would remain unchanged.\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["2", "text"], "text": "The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516."}, {"path": ["3", "text"], "text": "The package bearing barcode 074321908516 displays the exact label text “no added sugar.”"}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516.", "negative_left": "The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516.", "negative_right": "The package bearing barcode 074321908516 displays the exact label text “contains added sugar.”", "right": "The package bearing barcode 074321908516 displays the exact label text “no added sugar.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-021-004", "id": "fast-41-diverse-021-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Store picker", "text": "The requested North Mill creamy peanut butter is unavailable. The proposed North Mill smooth peanut butter is from the same brand, and its jar is marked with a net quantity of exactly 16 ounces."}, {"speaker": "Store picker", "text": "The proposed North Mill smooth peanut butter substitution in this pickup order is the product package bearing barcode 074321908516."}, {"speaker": "Store picker", "text": "The package bearing barcode 074321908516 displays the exact label text “contains added sugar.”"}, {"speaker": "Customer contact agent", "text": "Phone and text outreach received no customer response, so no additional waiver or changed instruction was recorded."}, {"speaker": "Fulfillment supervisor", "text": "The replacement is under review before the pickup window closes. If the printed wording on that same package were changed, all other recorded product and contact observations would remain unchanged."}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and substitution bindings, and the evidence remains factual and complete: “The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827.” and “The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.””", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”\"},{\"speaker\":\"Store picker\",\"text\":\"The requested item is North Mill creamy peanut butter, and the proposed substitution is North Mill smooth peanut butter; both are identified as North Mill.\"},{\"speaker\":\"Store picker\",\"text\":\"The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces.\"},{\"speaker\":\"Package inspection record\",\"text\":\"The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827.\"},{\"speaker\":\"Package inspection record\",\"text\":\"The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.”\"},{\"speaker\":\"Customer contact agent\",\"text\":\"Calls and a text received no response, and no waiver was recorded before the pickup window review.\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["3", "text"], "text": "The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827."}, {"path": ["4", "text"], "text": "The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.”"}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827.", "negative_left": "The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827.", "negative_right": "The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “contains added sugar.”", "right": "The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-021-005", "id": "fast-41-diverse-021-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Store picker", "text": "The requested item is North Mill creamy peanut butter, and the proposed substitution is North Mill smooth peanut butter; both are identified as North Mill."}, {"speaker": "Store picker", "text": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"speaker": "Package inspection record", "text": "The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827."}, {"speaker": "Package inspection record", "text": "The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.”"}, {"speaker": "Customer contact agent", "text": "Calls and a text received no response, and no waiver was recorded before the pickup window review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and substitution bindings, and the evidence remains factual and complete: “The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827.” and “The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.””", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”\"},{\"speaker\":\"Store picker\",\"text\":\"The requested item is North Mill creamy peanut butter, and the proposed substitution is North Mill smooth peanut butter; both are identified as North Mill.\"},{\"speaker\":\"Store picker\",\"text\":\"The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces.\"},{\"speaker\":\"Package inspection record\",\"text\":\"The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827.\"},{\"speaker\":\"Package inspection record\",\"text\":\"The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.”\"},{\"speaker\":\"Customer contact agent\",\"text\":\"Calls and a text received no response, and no waiver was recorded before the pickup window review.\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["3", "text"], "text": "The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827."}, {"path": ["4", "text"], "text": "The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.”"}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827.", "negative_left": "The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827.", "negative_right": "The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “contains added sugar.”", "right": "The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “no added sugar.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-021-005", "id": "fast-41-diverse-021-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Store picker", "text": "The requested item is North Mill creamy peanut butter, and the proposed substitution is North Mill smooth peanut butter; both are identified as North Mill."}, {"speaker": "Store picker", "text": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"speaker": "Package inspection record", "text": "The proposed North Mill smooth peanut butter substitution in this pickup order bears lot code NM-4827."}, {"speaker": "Package inspection record", "text": "The package bearing lot code NM-4827 was recorded on 2026-09-17 as labeled “contains added sugar.”"}, {"speaker": "Customer contact agent", "text": "Calls and a text received no response, and no waiver was recorded before the pickup window review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and governing policy remain preserved in both full inputs. Lena, the proposed substitution, and the order-item scope remain bound consistently. The evidence consists of two complete factual sentences: “On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros.” and “On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order.” The counterfactual coherently changes the replacement price to 22.50 euros without duplicate contradictions. Neither context embeds a gold answer, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Case note: Lena’s pickup order concerns one unavailable can of BrightStart infant formula, and the replacement record identifies R4 as the proposed substitute for Order Item Q7. The unavailable item and R4 are both BrightStart products, and each is labeled Stage 2. On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros. On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review. The item and replacement records were checked against the order note before routing.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros."}, {"path": [], "text": "On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros.", "negative_left": "On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros.", "negative_right": "On 17 September 2026, proposed replacement R4 was recorded at a price of 22.50 euros for Order Item Q7’s order.", "right": "On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order."}, "verifier_independent_model": false}, "family": "fast-41-diverse-022-003", "id": "fast-41-diverse-022-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Case note: Lena’s pickup order concerns one unavailable can of BrightStart infant formula, and the replacement record identifies R4 as the proposed substitute for Order Item Q7. The unavailable item and R4 are both BrightStart products, and each is labeled Stage 2. On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros. On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review. The item and replacement records were checked against the order note before routing."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and governing policy remain preserved in both full inputs. Lena, the proposed substitution, and the order-item scope remain bound consistently. The evidence consists of two complete factual sentences: “On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros.” and “On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order.” The counterfactual coherently changes the replacement price to 22.50 euros without duplicate contradictions. Neither context embeds a gold answer, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Case note: Lena’s pickup order concerns one unavailable can of BrightStart infant formula, and the replacement record identifies R4 as the proposed substitute for Order Item Q7. The unavailable item and R4 are both BrightStart products, and each is labeled Stage 2. On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros. On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review. The item and replacement records were checked against the order note before routing.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros."}, {"path": [], "text": "On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros.", "negative_left": "On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros.", "negative_right": "On 17 September 2026, proposed replacement R4 was recorded at a price of 22.50 euros for Order Item Q7’s order.", "right": "On 17 September 2026, proposed replacement R4 was recorded at a price of 21.50 euros for Order Item Q7’s order."}, "verifier_independent_model": false}, "family": "fast-41-diverse-022-003", "id": "fast-41-diverse-022-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Case note: Lena’s pickup order concerns one unavailable can of BrightStart infant formula, and the replacement record identifies R4 as the proposed substitute for Order Item Q7. The unavailable item and R4 are both BrightStart products, and each is labeled Stage 2. On 17 September 2026, Lena’s unavailable infant-formula order item, Order Item Q7, was sold at a recorded price of 20.00 euros. On 17 September 2026, proposed replacement R4 was recorded at a price of 22.50 euros for Order Item Q7’s order. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review. The item and replacement records were checked against the order note before routing."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions, substitution policy, routing scope, Lena, and the proposed replacement decision. The evidence consists of exactly two complete factual sentences: “On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00.” and “On 17 September 2026, the proposed replacement for that item was priced at $21.00.” The counterfactual coherently changes the replacement price from $21.00 to $23.00 without duplicating or contradicting measurements within that context. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Case note — 17 September 2026: Lena’s pickup order lists one can of BrightStart infant formula, and the unavailable item is the Stage 2 version. The picker located a BrightStart Stage 2 can as a proposed replacement. On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00. The replacement can is reserved for this order and has the same labeled brand and stage as the unavailable product. On 17 September 2026, the proposed replacement for that item was priced at $21.00. Lena’s order note governs substitutions. “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00."}, {"path": [], "text": "On 17 September 2026, the proposed replacement for that item was priced at $21.00."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00.", "negative_left": "On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00.", "negative_right": "On 17 September 2026, the proposed replacement for that item was priced at $23.00.", "right": "On 17 September 2026, the proposed replacement for that item was priced at $21.00."}, "verifier_independent_model": false}, "family": "fast-41-diverse-022-009", "id": "fast-41-diverse-022-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Case note — 17 September 2026: Lena’s pickup order lists one can of BrightStart infant formula, and the unavailable item is the Stage 2 version. The picker located a BrightStart Stage 2 can as a proposed replacement. On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00. The replacement can is reserved for this order and has the same labeled brand and stage as the unavailable product. On 17 September 2026, the proposed replacement for that item was priced at $21.00. Lena’s order note governs substitutions. “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions, substitution policy, routing scope, Lena, and the proposed replacement decision. The evidence consists of exactly two complete factual sentences: “On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00.” and “On 17 September 2026, the proposed replacement for that item was priced at $21.00.” The counterfactual coherently changes the replacement price from $21.00 to $23.00 without duplicating or contradicting measurements within that context. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Case note — 17 September 2026: Lena’s pickup order lists one can of BrightStart infant formula, and the unavailable item is the Stage 2 version. The picker located a BrightStart Stage 2 can as a proposed replacement. On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00. The replacement can is reserved for this order and has the same labeled brand and stage as the unavailable product. On 17 September 2026, the proposed replacement for that item was priced at $21.00. Lena’s order note governs substitutions. “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00."}, {"path": [], "text": "On 17 September 2026, the proposed replacement for that item was priced at $21.00."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00.", "negative_left": "On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00.", "negative_right": "On 17 September 2026, the proposed replacement for that item was priced at $23.00.", "right": "On 17 September 2026, the proposed replacement for that item was priced at $21.00."}, "verifier_independent_model": false}, "family": "fast-41-diverse-022-009", "id": "fast-41-diverse-022-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Case note — 17 September 2026: Lena’s pickup order lists one can of BrightStart infant formula, and the unavailable item is the Stage 2 version. The picker located a BrightStart Stage 2 can as a proposed replacement. On 17 September 2026, Lena’s unavailable ordered infant-formula item was priced at $20.00. The replacement can is reserved for this order and has the same labeled brand and stage as the unavailable product. On 17 September 2026, the proposed replacement for that item was priced at $23.00. Lena’s order note governs substitutions. “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy preserved: the unchanged questions object and both contexts retain all governing constraints. Question bindings preserved: both contexts identify the 32 oz Harvest Oat request and the proposed 28 oz Mill Lane oat milk. Evidence is two factual sentences: “The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183.” and “The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”” Counterfactual coherent: the same identified carton changes from 0 g to 7 g added sugars without creating duplicate measurements. No answer leakage: neither context contains labels, answer codes, rule tables, proposition IDs, rationale, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and its label identifies it as oat milk and unsweetened.\"},{\"speaker\":\"Inventory record\",\"text\":\"The requested Harvest Oat carton is 32 oz, while the proposed Mill Lane carton is 28 oz. The proposed quantity falls within the permitted inclusive range of 28–36 oz.\"},{\"speaker\":\"Verification record\",\"text\":\"The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183.\"},{\"speaker\":\"Verification record\",\"text\":\"The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”\"},{\"speaker\":\"Customer contact agent\",\"text\":\"A digital coupon applied only to the unavailable Harvest Oat carton, and the pickup window ends at 6:00 p.m.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["3", "text"], "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183."}, {"path": ["4", "text"], "text": "The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”"}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183.", "negative_left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183.", "negative_right": "The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 7 g.”", "right": "The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-002", "id": "fast-41-diverse-024-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and its label identifies it as oat milk and unsweetened."}, {"speaker": "Inventory record", "text": "The requested Harvest Oat carton is 32 oz, while the proposed Mill Lane carton is 28 oz. The proposed quantity falls within the permitted inclusive range of 28–36 oz."}, {"speaker": "Verification record", "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183."}, {"speaker": "Verification record", "text": "The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”"}, {"speaker": "Customer contact agent", "text": "A digital coupon applied only to the unavailable Harvest Oat carton, and the pickup window ends at 6:00 p.m."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy preserved: the unchanged questions object and both contexts retain all governing constraints. Question bindings preserved: both contexts identify the 32 oz Harvest Oat request and the proposed 28 oz Mill Lane oat milk. Evidence is two factual sentences: “The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183.” and “The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”” Counterfactual coherent: the same identified carton changes from 0 g to 7 g added sugars without creating duplicate measurements. No answer leakage: neither context contains labels, answer codes, rule tables, proposition IDs, rationale, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and its label identifies it as oat milk and unsweetened.\"},{\"speaker\":\"Inventory record\",\"text\":\"The requested Harvest Oat carton is 32 oz, while the proposed Mill Lane carton is 28 oz. The proposed quantity falls within the permitted inclusive range of 28–36 oz.\"},{\"speaker\":\"Verification record\",\"text\":\"The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183.\"},{\"speaker\":\"Verification record\",\"text\":\"The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”\"},{\"speaker\":\"Customer contact agent\",\"text\":\"A digital coupon applied only to the unavailable Harvest Oat carton, and the pickup window ends at 6:00 p.m.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["3", "text"], "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183."}, {"path": ["4", "text"], "text": "The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”"}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183.", "negative_left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183.", "negative_right": "The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 7 g.”", "right": "The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 0 g.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-002", "id": "fast-41-diverse-024-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and its label identifies it as oat milk and unsweetened."}, {"speaker": "Inventory record", "text": "The requested Harvest Oat carton is 32 oz, while the proposed Mill Lane carton is 28 oz. The proposed quantity falls within the permitted inclusive range of 28–36 oz."}, {"speaker": "Verification record", "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears barcode 740183."}, {"speaker": "Verification record", "text": "The sealed carton bearing barcode 740183 has a nutrition panel that prints “Added sugars: 7 g.”"}, {"speaker": "Customer contact agent", "text": "A digital coupon applied only to the unavailable Harvest Oat carton, and the pickup window ends at 6:00 p.m."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing current-order policy. The 32 oz Harvest Oat request and 28 oz Mill Lane replacement bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar.\" and \"The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar.\" In the counterfactual, 5 g total sugar and 3 g naturally occurring sugar coherently imply 2 g added sugar without duplicate measurements. Neither context contains a gold answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and its package identifies it as oat milk labeled unsweetened. The proposed quantity is within the order’s permitted range, but it is not equal to the requested carton’s 32 oz quantity.\"},{\"speaker\":\"Quality record\",\"text\":\"The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar. The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar."}, {"path": ["2", "text"], "text": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar.", "negative_left": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 5 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar.", "negative_right": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar.", "right": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-008", "id": "fast-41-diverse-024-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and its package identifies it as oat milk labeled unsweetened. The proposed quantity is within the order’s permitted range, but it is not equal to the requested carton’s 32 oz quantity."}, {"speaker": "Quality record", "text": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar. The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing current-order policy. The 32 oz Harvest Oat request and 28 oz Mill Lane replacement bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar.\" and \"The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar.\" In the counterfactual, 5 g total sugar and 3 g naturally occurring sugar coherently imply 2 g added sugar without duplicate measurements. Neither context contains a gold answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and its package identifies it as oat milk labeled unsweetened. The proposed quantity is within the order’s permitted range, but it is not equal to the requested carton’s 32 oz quantity.\"},{\"speaker\":\"Quality record\",\"text\":\"The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar. The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar."}, {"path": ["2", "text"], "text": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar.", "negative_left": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 5 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar.", "negative_right": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar.", "right": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-008", "id": "fast-41-diverse-024-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and its package identifies it as oat milk labeled unsweetened. The proposed quantity is within the order’s permitted range, but it is not equal to the requested carton’s 32 oz quantity."}, {"speaker": "Quality record", "text": "The proposed 28 oz Mill Lane oat milk’s tested serving contains 5 g total sugar, with total sugar defined as naturally occurring sugar plus added sugar. The proposed 28 oz Mill Lane oat milk’s tested serving contains 3 g naturally occurring oat sugar."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the verbatim questions object supplies the criteria and both contexts retain the operative constraints and irrelevance rule. Question bindings remain fixed to the 32 oz Harvest Oat request and proposed 28 oz Mill Lane oat milk. Evidence consists of two complete factual sentences: “The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482.” and “The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars.” The counterfactual coherently changes only the laboratory result while retaining the same lot and item. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable, so the current pickup order needs a substitute. The proposed Mill Lane carton is 28 oz and is oat milk labeled unsweetened.\"},{\"speaker\":\"Quality record\",\"text\":\"The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482.\"},{\"speaker\":\"Quality record\",\"text\":\"The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars.\"},{\"speaker\":\"Order review\",\"text\":\"The proposed carton's 28 oz net quantity falls within the customer's inclusive 28–36 oz range, but it differs from the requested Harvest Oat carton's 32 oz quantity.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482."}, {"path": ["3", "text"], "text": "The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482.", "negative_left": "The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482.", "negative_right": "The laboratory record for lot ML-482 reports 4 grams of total sugar, all attributable to added cane sugar.", "right": "The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-013", "id": "fast-41-diverse-024-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable, so the current pickup order needs a substitute. The proposed Mill Lane carton is 28 oz and is oat milk labeled unsweetened."}, {"speaker": "Quality record", "text": "The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482."}, {"speaker": "Quality record", "text": "The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars."}, {"speaker": "Order review", "text": "The proposed carton's 28 oz net quantity falls within the customer's inclusive 28–36 oz range, but it differs from the requested Harvest Oat carton's 32 oz quantity."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the verbatim questions object supplies the criteria and both contexts retain the operative constraints and irrelevance rule. Question bindings remain fixed to the 32 oz Harvest Oat request and proposed 28 oz Mill Lane oat milk. Evidence consists of two complete factual sentences: “The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482.” and “The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars.” The counterfactual coherently changes only the laboratory result while retaining the same lot and item. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable, so the current pickup order needs a substitute. The proposed Mill Lane carton is 28 oz and is oat milk labeled unsweetened.\"},{\"speaker\":\"Quality record\",\"text\":\"The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482.\"},{\"speaker\":\"Quality record\",\"text\":\"The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars.\"},{\"speaker\":\"Order review\",\"text\":\"The proposed carton's 28 oz net quantity falls within the customer's inclusive 28–36 oz range, but it differs from the requested Harvest Oat carton's 32 oz quantity.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482."}, {"path": ["3", "text"], "text": "The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482.", "negative_left": "The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482.", "negative_right": "The laboratory record for lot ML-482 reports 4 grams of total sugar, all attributable to added cane sugar.", "right": "The laboratory record for lot ML-482 reports 0 grams of total sugar, all attributable to naturally occurring oat sugars."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-013", "id": "fast-41-diverse-024-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable, so the current pickup order needs a substitute. The proposed Mill Lane carton is 28 oz and is oat milk labeled unsweetened."}, {"speaker": "Quality record", "text": "The current pickup order's proposed 28 oz Mill Lane carton bears lot code ML-482."}, {"speaker": "Quality record", "text": "The laboratory record for lot ML-482 reports 4 grams of total sugar, all attributable to added cane sugar."}, {"speaker": "Order review", "text": "The proposed carton's 28 oz net quantity falls within the customer's inclusive 28–36 oz range, but it differs from the requested Harvest Oat carton's 32 oz quantity."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policy, preserve all item and order bindings, and contain no answer labels or classifier instructions. The two evidence quotes are complete factual sentences, and changing the assay from 0 grams to 3 grams is coherent without creating contradictory duplicate measurements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton for the current pickup order. It is oat milk, is labeled unsweetened, falls within the permitted inclusive 28–36 oz range, and its quantity is not equal to the requested carton's 32 oz quantity.\"},{\"speaker\":\"Inventory record\",\"text\":\"For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741.\"},{\"speaker\":\"Assay record\",\"text\":\"The certified assay for carton ML-7741 reports an added-sugar quantity of 0 grams.\"},{\"speaker\":\"Customer contact agent\",\"text\":\"The customer requested no contact about this item. A digital coupon applied only to the unavailable Harvest Oat carton, and the pickup window ends at 6:00 p.m.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741."}, {"path": ["3", "text"], "text": "The certified assay for carton ML-7741 reports an added-sugar quantity of 0 grams."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741.", "negative_left": "For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741.", "negative_right": "The certified assay for carton ML-7741 reports an added-sugar quantity of 3 grams.", "right": "The certified assay for carton ML-7741 reports an added-sugar quantity of 0 grams."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-016", "id": "fast-41-diverse-024-016-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton for the current pickup order. It is oat milk, is labeled unsweetened, falls within the permitted inclusive 28–36 oz range, and its quantity is not equal to the requested carton's 32 oz quantity."}, {"speaker": "Inventory record", "text": "For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741."}, {"speaker": "Assay record", "text": "The certified assay for carton ML-7741 reports an added-sugar quantity of 0 grams."}, {"speaker": "Customer contact agent", "text": "The customer requested no contact about this item. A digital coupon applied only to the unavailable Harvest Oat carton, and the pickup window ends at 6:00 p.m."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policy, preserve all item and order bindings, and contain no answer labels or classifier instructions. The two evidence quotes are complete factual sentences, and changing the assay from 0 grams to 3 grams is coherent without creating contradictory duplicate measurements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton for the current pickup order. It is oat milk, is labeled unsweetened, falls within the permitted inclusive 28–36 oz range, and its quantity is not equal to the requested carton's 32 oz quantity.\"},{\"speaker\":\"Inventory record\",\"text\":\"For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741.\"},{\"speaker\":\"Assay record\",\"text\":\"The certified assay for carton ML-7741 reports an added-sugar quantity of 0 grams.\"},{\"speaker\":\"Customer contact agent\",\"text\":\"The customer requested no contact about this item. A digital coupon applied only to the unavailable Harvest Oat carton, and the pickup window ends at 6:00 p.m.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741."}, {"path": ["3", "text"], "text": "The certified assay for carton ML-7741 reports an added-sugar quantity of 0 grams."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741.", "negative_left": "For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741.", "negative_right": "The certified assay for carton ML-7741 reports an added-sugar quantity of 3 grams.", "right": "The certified assay for carton ML-7741 reports an added-sugar quantity of 0 grams."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-016", "id": "fast-41-diverse-024-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton for the current pickup order. It is oat milk, is labeled unsweetened, falls within the permitted inclusive 28–36 oz range, and its quantity is not equal to the requested carton's 32 oz quantity."}, {"speaker": "Inventory record", "text": "For the current pickup order, the proposed 28 oz Mill Lane oat milk carton bears inventory identifier ML-7741."}, {"speaker": "Assay record", "text": "The certified assay for carton ML-7741 reports an added-sugar quantity of 3 grams."}, {"speaker": "Customer contact agent", "text": "The customer requested no contact about this item. A digital coupon applied only to the unavailable Harvest Oat carton, and the pickup window ends at 6:00 p.m."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, and output scope, and both contexts retain the applicable customer and supervisor constraints. The requested item and proposed replacement remain bound to the same 32 oz, current-pickup-order substitution. The evidence consists of two complete factual sentences: “The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417.” and “Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams.” The counterfactual changes only the nutrition record’s sugar measurement from 0 grams to 3 grams without creating contradictory duplicate assertions. Neither context embeds an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and it is oat milk labeled unsweetened. Its net quantity is within the permitted 28–36 oz range but is not equal to the requested 32 oz quantity.\"},{\"speaker\":\"Inventory record\",\"text\":\"The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417.\"},{\"speaker\":\"Nutrition record\",\"text\":\"Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417."}, {"path": ["3", "text"], "text": "Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417.", "negative_left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417.", "negative_right": "Nutrition record MR-417, issued for the current pickup order, lists added sugar at 3 grams.", "right": "Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-019", "id": "fast-41-diverse-024-019-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and it is oat milk labeled unsweetened. Its net quantity is within the permitted 28–36 oz range but is not equal to the requested 32 oz quantity."}, {"speaker": "Inventory record", "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417."}, {"speaker": "Nutrition record", "text": "Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, and output scope, and both contexts retain the applicable customer and supervisor constraints. The requested item and proposed replacement remain bound to the same 32 oz, current-pickup-order substitution. The evidence consists of two complete factual sentences: “The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417.” and “Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams.” The counterfactual changes only the nutrition record’s sugar measurement from 0 grams to 3 grams without creating contradictory duplicate assertions. Neither context embeds an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and it is oat milk labeled unsweetened. Its net quantity is within the permitted 28–36 oz range but is not equal to the requested 32 oz quantity.\"},{\"speaker\":\"Inventory record\",\"text\":\"The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417.\"},{\"speaker\":\"Nutrition record\",\"text\":\"Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417."}, {"path": ["3", "text"], "text": "Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417.", "negative_left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417.", "negative_right": "Nutrition record MR-417, issued for the current pickup order, lists added sugar at 3 grams.", "right": "Nutrition record MR-417, issued for the current pickup order, lists added sugar at 0 grams."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-019", "id": "fast-41-diverse-024-019-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton, and it is oat milk labeled unsweetened. Its net quantity is within the permitted 28–36 oz range but is not equal to the requested 32 oz quantity."}, {"speaker": "Inventory record", "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears nutrition-record identifier MR-417."}, {"speaker": "Nutrition record", "text": "Nutrition record MR-417, issued for the current pickup order, lists added sugar at 3 grams."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings, the counterfactual coherently changes only the laboratory result, no answer is embedded, and the evidence consists of the exact factual sentences “The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771.” and “A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton. It is oat milk, and its package label says unsweetened. Its net quantity is within the permitted inclusive range, but it is not equal to the requested carton's 32 oz net quantity.\"},{\"speaker\":\"Inventory record\",\"text\":\"The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771.\"},{\"speaker\":\"Laboratory record\",\"text\":\"A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771."}, {"path": ["3", "text"], "text": "A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771.", "negative_left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771.", "negative_right": "A laboratory report for product code ML-28-OAT-771 records 3 grams of added sugar per carton.", "right": "A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-020", "id": "fast-41-diverse-024-020-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton. It is oat milk, and its package label says unsweetened. Its net quantity is within the permitted inclusive range, but it is not equal to the requested carton's 32 oz net quantity."}, {"speaker": "Inventory record", "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771."}, {"speaker": "Laboratory record", "text": "A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings, the counterfactual coherently changes only the laboratory result, no answer is embedded, and the evidence consists of the exact factual sentences “The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771.” and “A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual product attribute rather than a policy judgment. The base and counter assignments differ only on A3 and are realizable: an unsweetened 28 oz oat milk can have either zero or nonzero added sugar while the other listed facts remain fixed. The policy evidence cites the original state and preserves the state-originating current-order constraints and the instruction that history, coupon eligibility, and timing are irrelevant; question-originating scoring rules need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 being refuted entails that the proposed item does not have zero grams of added sugar, violating the mandatory no-added-sugar requirement. A single current-order requirement violation is sufficient for level 0, regardless of the other satisfied conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes oat milk, unsweetened status, zero added sugar, and an allowed size, while A5 being refuted excludes the requested 32 oz match. Thus every mandatory requirement is met, level 0 is excluded, and level 2 is excluded, making level 1 sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The proposed 28 oz Mill Lane carton for the current pickup order contains oat milk."}, {"id": "A2", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order is labeled unsweetened."}, {"id": "A3", "statement": "The proposed 28 oz Mill Lane oat milk for the current pickup order has zero grams of added sugar."}, {"id": "A4", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity is within the current order's permitted inclusive range of 28–36 oz."}, {"id": "A5", "statement": "The proposed 28 oz Mill Lane oat milk's net quantity equals the requested Harvest Oat carton's 32 oz net quantity."}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item.\"},{\"speaker\":\"Store picker\",\"text\":\"Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton. It is oat milk, and its package label says unsweetened. Its net quantity is within the permitted inclusive range, but it is not equal to the requested carton's 32 oz net quantity.\"},{\"speaker\":\"Inventory record\",\"text\":\"The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771.\"},{\"speaker\":\"Laboratory record\",\"text\":\"A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771."}, {"path": ["3", "text"], "text": "A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton."}], "policy_evidence": [{"path": ["0", "text"], "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"path": ["3", "text"], "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}], "rules": [{"justification": "The proposed item meets the oat-milk, unsweetened, and permitted-size requirements and does not match the requested 32 oz size, but it explicitly violates the current order's no-added-sugar requirement; one violated current-order requirement is sufficient for level 0.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The proposed item satisfies the oat-milk, unsweetened, no-added-sugar, and permitted-size requirements, while its size does not match the requested 32 oz size; therefore it is acceptable rather than a best match.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771.", "negative_left": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771.", "negative_right": "A laboratory report for product code ML-28-OAT-771 records 3 grams of added sugar per carton.", "right": "A laboratory report for product code ML-28-OAT-771 records 0 grams of added sugar per carton."}, "verifier_independent_model": false}, "family": "fast-41-diverse-024-020", "id": "fast-41-diverse-024-020-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: Reject the substitution because it violates at least one current-order requirement, such as oat-milk type, unsweetened status, no-added-sugar requirement, or the permitted 28–36 oz size range.", "1 — Acceptable: Approve the substitution because it satisfies every current-order requirement, but it differs from the requested item in size or another nonmandatory attribute allowed by the note.", "2 — Best match: Approve as the highest-suitability replacement because it satisfies every current-order requirement and matches the requested 32 oz size; a brand difference is allowed by the note."], "instructions": "Rate the proposed 28 oz Mill Lane oat milk’s substitution suitability. Apply only the customer’s current-order constraints and the proposed item’s attributes. Select exactly one level using the definitions below.", "type": "score"}}, "state": [{"speaker": "Pickup customer", "text": "For the 32 oz Harvest Oat unsweetened oat milk: any brand is fine, no added sugar, and a carton from 28–36 oz is acceptable. Do not contact me about this item."}, {"speaker": "Store picker", "text": "Harvest Oat is unavailable. The proposed replacement is a 28 oz Mill Lane carton. It is oat milk, and its package label says unsweetened. Its net quantity is within the permitted inclusive range, but it is not equal to the requested carton's 32 oz net quantity."}, {"speaker": "Inventory record", "text": "The proposed 28 oz Mill Lane oat milk for the current pickup order bears product code ML-28-OAT-771."}, {"speaker": "Laboratory record", "text": "A laboratory report for product code ML-28-OAT-771 records 3 grams of added sugar per carton."}, {"speaker": "Fulfillment supervisor", "text": "Current-order notes and proposed-item attributes control suitability; purchase history, coupon eligibility, and pickup timing do not change whether this replacement fits."}]}, "method": "c2d", "provenance": {"source_id": "diverse-024", "source_is_synthetic": true, "source_sha256": "833bdbb97e4f1463049c3db2a443e9f14904752c06e3d10221aea867f1398fc9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy, criteria, and instructions. Both contexts retain the K-184, claim-period, and balance bindings. The focus evidence contains two complete factual sentences. The counterfactual changes authorization consistently without creating contradictory measurements or assertions. Neither context states a gold answer, code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"For order K-184, the claim period opened with exactly 2,000 points in the member's loyalty account.\"},{\"speaker\":\"Ledger review\",\"text\":\"During the claim period, exactly 450 eligible points were posted for order K-184.\"},{\"speaker\":\"Ledger review\",\"text\":\"Exactly one reward-redemption transaction was posted during the claim period, and its point amount was exactly 700.\"},{\"speaker\":\"Ledger review\",\"text\":\"The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period.\"},{\"speaker\":\"Member record\",\"text\":\"The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as authorized.\"},{\"speaker\":\"Balance review\",\"text\":\"The actual loyalty-account balance at the end of the claim period was exactly 1,750 points.\"},{\"speaker\":\"Policy note\",\"text\":\"Policy awards one point per eligible merchandise dollar; gift cards earn none.\"},{\"speaker\":\"Policy note\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period."}, {"path": ["4", "text"], "text": "The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as authorized."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period.", "negative_left": "The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period.", "negative_right": "The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as unauthorized.", "right": "The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as authorized."}, "verifier_independent_model": false}, "family": "fast-41-diverse-029-020", "id": "fast-41-diverse-029-020-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Case note", "text": "For order K-184, the claim period opened with exactly 2,000 points in the member's loyalty account."}, {"speaker": "Ledger review", "text": "During the claim period, exactly 450 eligible points were posted for order K-184."}, {"speaker": "Ledger review", "text": "Exactly one reward-redemption transaction was posted during the claim period, and its point amount was exactly 700."}, {"speaker": "Ledger review", "text": "The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period."}, {"speaker": "Member record", "text": "The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as authorized."}, {"speaker": "Balance review", "text": "The actual loyalty-account balance at the end of the claim period was exactly 1,750 points."}, {"speaker": "Policy note", "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"speaker": "Policy note", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy, criteria, and instructions. Both contexts retain the K-184, claim-period, and balance bindings. The focus evidence contains two complete factual sentences. The counterfactual changes authorization consistently without creating contradictory measurements or assertions. Neither context states a gold answer, code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"For order K-184, the claim period opened with exactly 2,000 points in the member's loyalty account.\"},{\"speaker\":\"Ledger review\",\"text\":\"During the claim period, exactly 450 eligible points were posted for order K-184.\"},{\"speaker\":\"Ledger review\",\"text\":\"Exactly one reward-redemption transaction was posted during the claim period, and its point amount was exactly 700.\"},{\"speaker\":\"Ledger review\",\"text\":\"The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period.\"},{\"speaker\":\"Member record\",\"text\":\"The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as authorized.\"},{\"speaker\":\"Balance review\",\"text\":\"The actual loyalty-account balance at the end of the claim period was exactly 1,750 points.\"},{\"speaker\":\"Policy note\",\"text\":\"Policy awards one point per eligible merchandise dollar; gift cards earn none.\"},{\"speaker\":\"Policy note\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period."}, {"path": ["4", "text"], "text": "The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as authorized."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period.", "negative_left": "The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period.", "negative_right": "The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as unauthorized.", "right": "The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as authorized."}, "verifier_independent_model": false}, "family": "fast-41-diverse-029-020", "id": "fast-41-diverse-029-020-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Case note", "text": "For order K-184, the claim period opened with exactly 2,000 points in the member's loyalty account."}, {"speaker": "Ledger review", "text": "During the claim period, exactly 450 eligible points were posted for order K-184."}, {"speaker": "Ledger review", "text": "Exactly one reward-redemption transaction was posted during the claim period, and its point amount was exactly 700."}, {"speaker": "Ledger review", "text": "The account ledger for order K-184 lists reward-redemption transaction R-700 as posted during the claim period."}, {"speaker": "Member record", "text": "The member's signed redemption record for the sole posted transaction during that claim period marks that transaction as unauthorized."}, {"speaker": "Balance review", "text": "The actual loyalty-account balance at the end of the claim period was exactly 1,750 points."}, {"speaker": "Policy note", "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"speaker": "Policy note", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and retain the loyalty member’s June 8 redemption scope. The evidence includes “The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7.” The evidence also includes “The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward.” The counterfactual’s 2,401-point deduction is a coherent changed observation for the $7 reward and does not duplicate a conflicting measurement within that context. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards program analyst reviewed the loyalty member’s June 8 redemption after the account claim was forwarded for audit. The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7. The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The order record links the reward to the June 8 transaction, and the ledger shows no refund, transfer, expiration, manual adjustment, or other points activity that day. The analyst retained the order record, redemption entry, and ledger record together for review.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7.", "negative_left": "The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7.", "negative_right": "The redemption record for the loyalty member’s June 8 order records a points deduction of 2,401 points for its applied reward.", "right": "The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-001", "id": "fast-41-diverse-030-001-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards program analyst reviewed the loyalty member’s June 8 redemption after the account claim was forwarded for audit. The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7. The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The order record links the reward to the June 8 transaction, and the ledger shows no refund, transfer, expiration, manual adjustment, or other points activity that day. The analyst retained the order record, redemption entry, and ledger record together for review."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and retain the loyalty member’s June 8 redemption scope. The evidence includes “The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7.” The evidence also includes “The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward.” The counterfactual’s 2,401-point deduction is a coherent changed observation for the $7 reward and does not duplicate a conflicting measurement within that context. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards program analyst reviewed the loyalty member’s June 8 redemption after the account claim was forwarded for audit. The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7. The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The order record links the reward to the June 8 transaction, and the ledger shows no refund, transfer, expiration, manual adjustment, or other points activity that day. The analyst retained the order record, redemption entry, and ledger record together for review.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7.", "negative_left": "The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7.", "negative_right": "The redemption record for the loyalty member’s June 8 order records a points deduction of 2,401 points for its applied reward.", "right": "The redemption record for the loyalty member’s June 8 order records a points deduction of 2,400 points for its applied reward."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-001", "id": "fast-41-diverse-030-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards program analyst reviewed the loyalty member’s June 8 redemption after the account claim was forwarded for audit. The supplied program policy records a redemption rate of 200 points per dollar, and the reward applied to the loyalty member’s June 8 order has a value of $7. The redemption record for the loyalty member’s June 8 order records a points deduction of 2,401 points for its applied reward. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The order record links the reward to the June 8 transaction, and the ledger shows no refund, transfer, expiration, manual adjustment, or other points activity that day. The analyst retained the order record, redemption entry, and ledger record together for review."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the June 8 redemption scope, policy rate, and terminology clarification; the changed reward and deduction values are permitted observations. The evidence consists of two complete factual sentences: “For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption.” and “The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points.” The counterfactual changes only the recorded deduction to 5,601 points and remains consistent with the $17 reward and 200-points-per-dollar policy. Neither context contains a gold answer, answer code, label rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A loyalty member asked a rewards analyst to review the redemption attached to the June 8 order after an account inquiry. For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption. The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The ledger shows no other points adjustment on June 8, and the redemption record identifies the listed deduction as the charge for this reward. The analyst retained the order date, reward value, policy rate, and recorded deduction as the controlling entries.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption.", "negative_left": "For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption.", "negative_right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 5,601 points.", "right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-002", "id": "fast-41-diverse-030-002-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A loyalty member asked a rewards analyst to review the redemption attached to the June 8 order after an account inquiry. For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption. The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The ledger shows no other points adjustment on June 8, and the redemption record identifies the listed deduction as the charge for this reward. The analyst retained the order date, reward value, policy rate, and recorded deduction as the controlling entries."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the June 8 redemption scope, policy rate, and terminology clarification; the changed reward and deduction values are permitted observations. The evidence consists of two complete factual sentences: “For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption.” and “The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points.” The counterfactual changes only the recorded deduction to 5,601 points and remains consistent with the $17 reward and 200-points-per-dollar policy. Neither context contains a gold answer, answer code, label rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A loyalty member asked a rewards analyst to review the redemption attached to the June 8 order after an account inquiry. For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption. The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The ledger shows no other points adjustment on June 8, and the redemption record identifies the listed deduction as the charge for this reward. The analyst retained the order date, reward value, policy rate, and recorded deduction as the controlling entries.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption.", "negative_left": "For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption.", "negative_right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 5,601 points.", "right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 4,200 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-002", "id": "fast-41-diverse-030-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A loyalty member asked a rewards analyst to review the redemption attached to the June 8 order after an account inquiry. For the loyalty member’s June 8 order, the applied reward had a value of $17, and the supplied program policy assigned 200 points per dollar to that redemption. The redemption record for the loyalty member’s June 8 order shows a deduction of 5,601 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The ledger shows no other points adjustment on June 8, and the redemption record identifies the listed deduction as the charge for this reward. The analyst retained the order date, reward value, policy rate, and recorded deduction as the controlling entries."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and both contexts preserve the governing discrepancy policy. The loyalty-member redemption, June 8 order, and points-deduction bindings remain unchanged. The evidence consists of two complete factual sentences: “For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount.” and “For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points.” The counterfactual changes only the deduction from 5,500 to 5,900 and introduces no contradiction. Neither context embeds a label, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards analyst reviewed the loyalty member’s June 8 redemption after the account was flagged for an unusual points entry. The case file identifies one reward redemption, associates it with the June 8 order, and shows no other points adjustments that day. The supplied program policy states that redemption costs 200 points per $1 of reward value. The redemption ledger separately records the policy-required amount and the points removed from the member’s account. For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount. For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points. In the member’s wording, “credits vanished” refers to points deducted during redemption. The analyst retained the order date, reward linkage, and ledger entries for review.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount."}, {"path": [], "text": "For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount.", "negative_left": "For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount.", "negative_right": "For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,900 points.", "right": "For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-003", "id": "fast-41-diverse-030-003-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards analyst reviewed the loyalty member’s June 8 redemption after the account was flagged for an unusual points entry. The case file identifies one reward redemption, associates it with the June 8 order, and shows no other points adjustments that day. The supplied program policy states that redemption costs 200 points per $1 of reward value. The redemption ledger separately records the policy-required amount and the points removed from the member’s account. For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount. For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points. In the member’s wording, “credits vanished” refers to points deducted during redemption. The analyst retained the order date, reward linkage, and ledger entries for review."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and both contexts preserve the governing discrepancy policy. The loyalty-member redemption, June 8 order, and points-deduction bindings remain unchanged. The evidence consists of two complete factual sentences: “For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount.” and “For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points.” The counterfactual changes only the deduction from 5,500 to 5,900 and introduces no contradiction. Neither context embeds a label, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards analyst reviewed the loyalty member’s June 8 redemption after the account was flagged for an unusual points entry. The case file identifies one reward redemption, associates it with the June 8 order, and shows no other points adjustments that day. The supplied program policy states that redemption costs 200 points per $1 of reward value. The redemption ledger separately records the policy-required amount and the points removed from the member’s account. For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount. For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points. In the member’s wording, “credits vanished” refers to points deducted during redemption. The analyst retained the order date, reward linkage, and ledger entries for review.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount."}, {"path": [], "text": "For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount.", "negative_left": "For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount.", "negative_right": "For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,900 points.", "right": "For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,500 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-003", "id": "fast-41-diverse-030-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards analyst reviewed the loyalty member’s June 8 redemption after the account was flagged for an unusual points entry. The case file identifies one reward redemption, associates it with the June 8 order, and shows no other points adjustments that day. The supplied program policy states that redemption costs 200 points per $1 of reward value. The redemption ledger separately records the policy-required amount and the points removed from the member’s account. For the redemption of the reward applied to the loyalty member’s June 8 order, the supplied program policy records 4,800 points as the required amount. For the redemption of the reward applied to the loyalty member’s June 8 order, the transaction record shows a deduction of 5,900 points. In the member’s wording, “credits vanished” refers to points deducted during redemption. The analyst retained the order date, reward linkage, and ledger entries for review."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and criteria, while both contexts retain the loyalty-member June 8 redemption scope. The entity, redemption path, and time binding remain unchanged despite altered observations. The evidence contains exactly two complete factual sentences. The 3,501-point counterfactual is consistent with a single deduction for the stated $10 reward and 200-points-per-dollar rate. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"Case note: A loyalty member’s June 8 order was reviewed after a rewards-account inquiry. The order record identifies one applied reward, and the account ledger shows a single points deduction associated with that redemption. The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1. The redemption record for the loyalty member’s June 8 order shows that 2,800 points were deducted. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The ledger contains no other points adjustment for the member on June 8, and the review concerns only this order’s redemption entry.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order shows that 2,800 points were deducted."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1.", "negative_left": "The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1.", "negative_right": "The redemption record for the loyalty member’s June 8 order shows that 3,501 points were deducted.", "right": "The redemption record for the loyalty member’s June 8 order shows that 2,800 points were deducted."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-007", "id": "fast-41-diverse-030-007-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "Case note: A loyalty member’s June 8 order was reviewed after a rewards-account inquiry. The order record identifies one applied reward, and the account ledger shows a single points deduction associated with that redemption. The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1. The redemption record for the loyalty member’s June 8 order shows that 2,800 points were deducted. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The ledger contains no other points adjustment for the member on June 8, and the review concerns only this order’s redemption entry."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and criteria, while both contexts retain the loyalty-member June 8 redemption scope. The entity, redemption path, and time binding remain unchanged despite altered observations. The evidence contains exactly two complete factual sentences. The 3,501-point counterfactual is consistent with a single deduction for the stated $10 reward and 200-points-per-dollar rate. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"Case note: A loyalty member’s June 8 order was reviewed after a rewards-account inquiry. The order record identifies one applied reward, and the account ledger shows a single points deduction associated with that redemption. The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1. The redemption record for the loyalty member’s June 8 order shows that 2,800 points were deducted. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The ledger contains no other points adjustment for the member on June 8, and the review concerns only this order’s redemption entry.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order shows that 2,800 points were deducted."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1.", "negative_left": "The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1.", "negative_right": "The redemption record for the loyalty member’s June 8 order shows that 3,501 points were deducted.", "right": "The redemption record for the loyalty member’s June 8 order shows that 2,800 points were deducted."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-007", "id": "fast-41-diverse-030-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "Case note: A loyalty member’s June 8 order was reviewed after a rewards-account inquiry. The order record identifies one applied reward, and the account ledger shows a single points deduction associated with that redemption. The June 8 order record shows a reward value of $10, and the supplied program policy records a redemption rate of 200 points per $1. The redemption record for the loyalty member’s June 8 order shows that 3,501 points were deducted. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The ledger contains no other points adjustment for the member on June 8, and the review concerns only this order’s redemption entry."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy and bindings; both contexts coherently vary only recorded observations, and the evidence quotes are factual sentences: “The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1.” and “The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards analyst reviewed the loyalty member’s June 8 redemption after the member reported that credits had vanished. The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1. The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The account ledger shows no other adjustments that day, and the record identifies this deduction as the points charged for the applied reward. The analyst is comparing the recorded charge with the policy-required charge for that redemption.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1.", "negative_left": "The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1.", "negative_right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 6,700 points.", "right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-008", "id": "fast-41-diverse-030-008-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards analyst reviewed the loyalty member’s June 8 redemption after the member reported that credits had vanished. The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1. The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The account ledger shows no other adjustments that day, and the record identifies this deduction as the points charged for the applied reward. The analyst is comparing the recorded charge with the policy-required charge for that redemption."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy and bindings; both contexts coherently vary only recorded observations, and the evidence quotes are factual sentences: “The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1.” and “The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards analyst reviewed the loyalty member’s June 8 redemption after the member reported that credits had vanished. The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1. The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The account ledger shows no other adjustments that day, and the record identifies this deduction as the points charged for the applied reward. The analyst is comparing the recorded charge with the policy-required charge for that redemption.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1.", "negative_left": "The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1.", "negative_right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 6,700 points.", "right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 5,800 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-008", "id": "fast-41-diverse-030-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards analyst reviewed the loyalty member’s June 8 redemption after the member reported that credits had vanished. The supplied program policy records the June 8 order’s applied reward value as $25 and its redemption rate as 200 points per $1. The redemption record for the loyalty member’s June 8 order shows a deduction of 6,700 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The account ledger shows no other adjustments that day, and the record identifies this deduction as the points charged for the applied reward. The analyst is comparing the recorded charge with the policy-required charge for that redemption."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and scoring criteria, while both contexts retain the June 8 redemption scope and entity bindings. The evidence consists of two complete factual sentences: “The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points.” and “The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points.” The counterfactual consistently changes the deduction to 5,000 points without contradiction. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards-program analyst reviewed a loyalty member’s June 8 order after a redemption concern. The account identifies one applied reward redemption, and no unrelated points adjustments or corrections were posted that day. The supplied program policy states that redemption costs 200 points per $1 of reward value. The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points. The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points. The transaction record distinguishes this event from a purchase, transfer, expiration, or account correction. In the member’s wording, “credits vanished” refers to points deducted during redemption. For a separate audit variant, the same order, policy record, reward identity, and absence of other adjustments remain unchanged, while the redemption record shows a deduction of 5,000 points instead.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points.", "negative_left": "The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points.", "negative_right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 5,000 points.", "right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-011", "id": "fast-41-diverse-030-011-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards-program analyst reviewed a loyalty member’s June 8 order after a redemption concern. The account identifies one applied reward redemption, and no unrelated points adjustments or corrections were posted that day. The supplied program policy states that redemption costs 200 points per $1 of reward value. The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points. The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points. The transaction record distinguishes this event from a purchase, transfer, expiration, or account correction. In the member’s wording, “credits vanished” refers to points deducted during redemption. For a separate audit variant, the same order, policy record, reward identity, and absence of other adjustments remain unchanged, while the redemption record shows a deduction of 5,000 points instead."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and scoring criteria, while both contexts retain the June 8 redemption scope and entity bindings. The evidence consists of two complete factual sentences: “The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points.” and “The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points.” The counterfactual consistently changes the deduction to 5,000 points without contradiction. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards-program analyst reviewed a loyalty member’s June 8 order after a redemption concern. The account identifies one applied reward redemption, and no unrelated points adjustments or corrections were posted that day. The supplied program policy states that redemption costs 200 points per $1 of reward value. The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points. The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points. The transaction record distinguishes this event from a purchase, transfer, expiration, or account correction. In the member’s wording, “credits vanished” refers to points deducted during redemption. For a separate audit variant, the same order, policy record, reward identity, and absence of other adjustments remain unchanged, while the redemption record shows a deduction of 5,000 points instead.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points."}, {"path": [], "text": "The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points.", "negative_left": "The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points.", "negative_right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 5,000 points.", "right": "The redemption record for the loyalty member’s June 8 order shows a deduction of 3,000 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-011", "id": "fast-41-diverse-030-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards-program analyst reviewed a loyalty member’s June 8 order after a redemption concern. The account identifies one applied reward redemption, and no unrelated points adjustments or corrections were posted that day. The supplied program policy states that redemption costs 200 points per $1 of reward value. The supplied program policy records that the applied reward on the loyalty member’s June 8 order required 2,400 points. The redemption record for the loyalty member’s June 8 order shows a deduction of 5,000 points. The transaction record distinguishes this event from a purchase, transfer, expiration, or account correction. In the member’s wording, “credits vanished” refers to points deducted during redemption. For a separate audit variant, the same order, policy record, reward identity, and absence of other adjustments remain unchanged, while the redemption record shows a deduction of 5,000 points instead."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and the loyalty member’s June 8 order and redemption bindings, while only observations change. The evidence consists of exactly two complete factual sentences: \"The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1.\" and \"The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points.\" The base measurements are coherent, and the counterfactual’s $12 reward, 200-points-per-dollar policy, and 3,601-point deduction are also coherent without contradictory duplicates. Neither context contains a gold answer, answer code, rule table, proposition IDs, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards operations analyst reviewed the loyalty member’s June 8 order after a redemption inquiry. The order record and ledger identify one applied reward and one corresponding redemption entry, linked by the same transaction reference. The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1. The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The analyst found no other reward redemptions, reversals, transfers, or account adjustments associated with the order or date, and confirmed that the supplied policy version was applicable when the order was processed.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1."}, {"path": [], "text": "The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1.", "negative_left": "The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1.", "negative_right": "The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,601 points.", "right": "The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-017", "id": "fast-41-diverse-030-017-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards operations analyst reviewed the loyalty member’s June 8 order after a redemption inquiry. The order record and ledger identify one applied reward and one corresponding redemption entry, linked by the same transaction reference. The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1. The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The analyst found no other reward redemptions, reversals, transfers, or account adjustments associated with the order or date, and confirmed that the supplied policy version was applicable when the order was processed."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and the loyalty member’s June 8 order and redemption bindings, while only observations change. The evidence consists of exactly two complete factual sentences: \"The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1.\" and \"The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points.\" The base measurements are coherent, and the counterfactual’s $12 reward, 200-points-per-dollar policy, and 3,601-point deduction are also coherent without contradictory duplicates. Neither context contains a gold answer, answer code, rule table, proposition IDs, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards operations analyst reviewed the loyalty member’s June 8 order after a redemption inquiry. The order record and ledger identify one applied reward and one corresponding redemption entry, linked by the same transaction reference. The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1. The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The analyst found no other reward redemptions, reversals, transfers, or account adjustments associated with the order or date, and confirmed that the supplied policy version was applicable when the order was processed.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1."}, {"path": [], "text": "The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1.", "negative_left": "The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1.", "negative_right": "The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,601 points.", "right": "The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,000 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-017", "id": "fast-41-diverse-030-017-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards operations analyst reviewed the loyalty member’s June 8 order after a redemption inquiry. The order record and ledger identify one applied reward and one corresponding redemption entry, linked by the same transaction reference. The applied reward on the loyalty member’s June 8 order has a recorded value of $12, and the supplied program policy records a redemption rate of 200 points per $1. The redemption record for the reward applied to the loyalty member’s June 8 order shows a deduction of 3,601 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The analyst found no other reward redemptions, reversals, transfers, or account adjustments associated with the order or date, and confirmed that the supplied policy version was applicable when the order was processed."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, and instructions, while both contexts retain the supplied 200-points-per-dollar policy. The loyalty member, June 8 order, redemption, and points-deduction bindings remain unchanged despite altered observations. The evidence consists of two complete factual sentences: \"For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption.\" and \"The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points.\" The counterfactual coherently changes the deduction to 4,050 points without contradictory duplicate measurements. Neither context includes a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards-program analyst reviewed the loyalty member’s June 8 redemption after customer service escalated the account inquiry. For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption. The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The order record links the reward and points transaction to the same June 8 purchase. The analyst found no other points adjustments, redemptions, refunds, or account corrections that day. The audit record preserves the reward value, policy-required amount, and recorded deduction as separate fields. The claim was forwarded for discrepancy review after the member reported the account change to customer service.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption."}, {"path": [], "text": "The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption.", "negative_left": "For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption.", "negative_right": "The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 4,050 points.", "right": "The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-020", "id": "fast-41-diverse-030-020-base", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards-program analyst reviewed the loyalty member’s June 8 redemption after customer service escalated the account inquiry. For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption. The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The order record links the reward and points transaction to the same June 8 purchase. The analyst found no other points adjustments, redemptions, refunds, or account corrections that day. The audit record preserves the reward value, policy-required amount, and recorded deduction as separate fields. The claim was forwarded for discrepancy review after the member reported the account change to customer service."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, and instructions, while both contexts retain the supplied 200-points-per-dollar policy. The loyalty member, June 8 order, redemption, and points-deduction bindings remain unchanged despite altered observations. The evidence consists of two complete factual sentences: \"For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption.\" and \"The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points.\" The counterfactual coherently changes the deduction to 4,050 points without contradictory duplicate measurements. Neither context includes a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual inequality concerning the redemption deduction, and A2 is a factual relation rather than a policy classification. The base assignment is realizable with an excess from 1 through 1,000 points, while the counter assignment is realizable with an excess from 1,001 through 5,000 points; only A2 changes status. The policy evidence correctly preserves the state-originating conversion rate and the clarification binding “credits vanished” to redemption points deducted. The scoring ranges themselves remain preserved automatically in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 entails an excess of at least 1 point, while A2 entails an excess of at most 1,000 points. Together they are sufficient for score 1; A3 is redundant but consistent.", "rule_index": 0, "sound": true}, {"reason": "Refuting A2 entails an excess greater than 1,000 points, and supporting A3 entails an excess no greater than 5,000 points. This is sufficient for score 2; A1 is redundant.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is at least 1 point greater than the points amount required by the supplied program policy."}, {"id": "A2", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 1,000 points greater than the points amount required by the supplied program policy."}, {"id": "A3", "statement": "For the redemption of the reward applied to the loyalty member’s June 8 order, the recorded points deduction is no more than 5,000 points greater than the points amount required by the supplied program policy."}], "base_state_json": "\"A rewards-program analyst reviewed the loyalty member’s June 8 redemption after customer service escalated the account inquiry. For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption. The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The order record links the reward and points transaction to the same June 8 purchase. The analyst found no other points adjustments, redemptions, refunds, or account corrections that day. The audit record preserves the reward value, policy-required amount, and recorded deduction as separate fields. The claim was forwarded for discrepancy review after the member reported the account change to customer service.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption."}, {"path": [], "text": "The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points."}], "policy_evidence": [{"path": [], "text": "The supplied program policy states that redemption costs 200 points per $1 of reward value."}, {"path": [], "text": "In the member’s wording, “credits vanished” refers to points deducted during redemption."}], "rules": [{"justification": "A recorded deduction at least 1 point and no more than 1,000 points above the policy-required amount satisfies the inclusive low-discrepancy range.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "Refutation of the 1,000-point upper bound entails that the recorded deduction is more than 1,000 points above the required amount; the supported 5,000-point upper bound places it within the inclusive moderate-discrepancy range.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}]}]}, "verified_pair": {"left": "For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption.", "negative_left": "For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption.", "negative_right": "The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 4,050 points.", "right": "The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 3,400 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-030-020", "id": "fast-41-diverse-030-020-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No discrepancy: the recorded deduction equals the policy-required amount.", "1 — Low discrepancy: the recorded deduction is 1 to 1,000 points above the required amount, inclusive.", "2 — Moderate discrepancy: the recorded deduction is 1,001 to 5,000 points above the required amount, inclusive.", "3 — High discrepancy: the recorded deduction is more than 5,000 points above the required amount."], "instructions": "Verify the redemption discrepancy, then select the level matching the number of points deducted above the policy-required amount. Treat paraphrased terms in the claim according to the clarification in the state.", "type": "score"}}, "state": "A rewards-program analyst reviewed the loyalty member’s June 8 redemption after customer service escalated the account inquiry. For the loyalty member’s June 8 order, the applied reward is valued at $13, so the supplied program policy requires 2,600 points for that redemption. The recorded points deduction for the reward applied to the loyalty member’s June 8 order is 4,050 points. The supplied program policy states that redemption costs 200 points per $1 of reward value. In the member’s wording, “credits vanished” refers to points deducted during redemption. The order record links the reward and points transaction to the same June 8 purchase. The analyst found no other points adjustments, redemptions, refunds, or account corrections that day. The audit record preserves the reward value, policy-required amount, and recorded deduction as separate fields. The claim was forwarded for discrepancy review after the member reported the account change to customer service."}, "method": "c2d", "provenance": {"source_id": "diverse-030", "source_is_synthetic": true, "source_sha256": "c2a00a0bc94ad29974929fbf7c01ea811a2a0a2633a1db596427cf7ca847bb97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision criteria, and both contexts retain the same handoff, actors, scrubber, and timing bindings. The evidence contains exactly two complete factual sentences: \"The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.\" and \"Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff.\" The counterfactual changes only the exception register's scope observation, creating a coherent unresolved conflict without contradictory duplicate measurements. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At the 19:00 handoff, outgoing shift lead Mara briefed incoming shift lead Dev about scrubber S-2. The master scope record was updated at 18:45. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff. Inez had not reconciled the records before the readiness rating. S-2 had an unresolved alarm, with current status and telemetry available, but the appropriate owner had not accepted the alarm before handoff. No temporary evidence exception covered the missing acceptance. The applicable scope of every other handoff item was established at 19:00. The master scope record and the exception register were the complete set of records governing whether S-2 was in scope for the 19:00 handoff. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-001", "id": "fast-41-diverse-032-001-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At the 19:00 handoff, outgoing shift lead Mara briefed incoming shift lead Dev about scrubber S-2. The master scope record was updated at 18:45. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff. Inez had not reconciled the records before the readiness rating. S-2 had an unresolved alarm, with current status and telemetry available, but the appropriate owner had not accepted the alarm before handoff. No temporary evidence exception covered the missing acceptance. The applicable scope of every other handoff item was established at 19:00. The master scope record and the exception register were the complete set of records governing whether S-2 was in scope for the 19:00 handoff. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision criteria, and both contexts retain the same handoff, actors, scrubber, and timing bindings. The evidence contains exactly two complete factual sentences: \"The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.\" and \"Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff.\" The counterfactual changes only the exception register's scope observation, creating a coherent unresolved conflict without contradictory duplicate measurements. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At the 19:00 handoff, outgoing shift lead Mara briefed incoming shift lead Dev about scrubber S-2. The master scope record was updated at 18:45. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff. Inez had not reconciled the records before the readiness rating. S-2 had an unresolved alarm, with current status and telemetry available, but the appropriate owner had not accepted the alarm before handoff. No temporary evidence exception covered the missing acceptance. The applicable scope of every other handoff item was established at 19:00. The master scope record and the exception register were the complete set of records governing whether S-2 was in scope for the 19:00 handoff. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-001", "id": "fast-41-diverse-032-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At the 19:00 handoff, outgoing shift lead Mara briefed incoming shift lead Dev about scrubber S-2. The master scope record was updated at 18:45. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. Operations coordinator Inez approved the exception register at 18:50, and it lists scrubber S-2 as out of scope for the 19:00 handoff. Inez had not reconciled the records before the readiness rating. S-2 had an unresolved alarm, with current status and telemetry available, but the appropriate owner had not accepted the alarm before handoff. No temporary evidence exception covered the missing acceptance. The applicable scope of every other handoff item was established at 19:00. The master scope record and the exception register were the complete set of records governing whether S-2 was in scope for the 19:00 handoff. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing criteria, and both contexts retain the reconciliation policy without invented exceptions. The 19:00 handoff, S-2, Mara, Dev, and Inez bindings remain intact. The evidence consists of two complete factual sentences: \"The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.\" \"The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff.\" The counterfactual coherently changes only the exception register's scope assertion while retaining the unresolved conflict and unreconciled status. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, outgoing shift lead Mara prepares the handoff for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. The master scope record and the exception register are the complete records governing S-2's scope for this handoff. S-2 has an unresolved alarm at 19:00, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Operations coordinator Inez has not reconciled the records before the readiness rating. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The status and telemetry for S-2 are timestamped, but the alarm remains open pending owner acceptance.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-009", "id": "fast-41-diverse-032-009-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, outgoing shift lead Mara prepares the handoff for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. The master scope record and the exception register are the complete records governing S-2's scope for this handoff. S-2 has an unresolved alarm at 19:00, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Operations coordinator Inez has not reconciled the records before the readiness rating. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The status and telemetry for S-2 are timestamped, but the alarm remains open pending owner acceptance."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing criteria, and both contexts retain the reconciliation policy without invented exceptions. The 19:00 handoff, S-2, Mara, Dev, and Inez bindings remain intact. The evidence consists of two complete factual sentences: \"The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.\" \"The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff.\" The counterfactual coherently changes only the exception register's scope assertion while retaining the unresolved conflict and unreconciled status. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, outgoing shift lead Mara prepares the handoff for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. The master scope record and the exception register are the complete records governing S-2's scope for this handoff. S-2 has an unresolved alarm at 19:00, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Operations coordinator Inez has not reconciled the records before the readiness rating. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The status and telemetry for S-2 are timestamped, but the alarm remains open pending owner acceptance.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-009", "id": "fast-41-diverse-032-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, outgoing shift lead Mara prepares the handoff for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register updated at 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff. The master scope record and the exception register are the complete records governing S-2's scope for this handoff. S-2 has an unresolved alarm at 19:00, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Operations coordinator Inez has not reconciled the records before the readiness rating. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The status and telemetry for S-2 are timestamped, but the alarm remains open pending owner acceptance."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings, and the counterfactual coherently changes one scope observation without contradictory duplicates; required evidence quotes are “The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.” and “The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff.” in the base context, with the second quote changed to “The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff.” in the counterfactual.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, outgoing shift lead Mara records the handoff for incoming shift lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. These two records are the complete governing record set for S-2’s scope. Inez did not reconcile them before the readiness rating. S-2 has an unresolved alarm at handoff, and the appropriate owner did not accept it before 19:00. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Status and telemetry for S-2 are recorded, but the unresolved alarm remains open. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-010", "id": "fast-41-diverse-032-010-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, outgoing shift lead Mara records the handoff for incoming shift lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. These two records are the complete governing record set for S-2’s scope. Inez did not reconcile them before the readiness rating. S-2 has an unresolved alarm at handoff, and the appropriate owner did not accept it before 19:00. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Status and telemetry for S-2 are recorded, but the unresolved alarm remains open. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings, and the counterfactual coherently changes one scope observation without contradictory duplicates; required evidence quotes are “The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.” and “The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff.” in the base context, with the second quote changed to “The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff.” in the counterfactual.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, outgoing shift lead Mara records the handoff for incoming shift lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. These two records are the complete governing record set for S-2’s scope. Inez did not reconcile them before the readiness rating. S-2 has an unresolved alarm at handoff, and the appropriate owner did not accept it before 19:00. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Status and telemetry for S-2 are recorded, but the unresolved alarm remains open. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-010", "id": "fast-41-diverse-032-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, outgoing shift lead Mara records the handoff for incoming shift lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register dated 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff. These two records are the complete governing record set for S-2’s scope. Inez did not reconcile them before the readiness rating. S-2 has an unresolved alarm at handoff, and the appropriate owner did not accept it before 19:00. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Status and telemetry for S-2 are recorded, but the unresolved alarm remains open. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and all question bindings. The evidence consists of two complete factual sentences: \"The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.\" \"The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff.\" The counterfactual coherently changes only the exception register's scope assertion while retaining the unresolved conflict. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At the 19:00 handoff, outgoing lead Mara prepared the record for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff. Both documents are the complete governing record set for S-2's scope. Scrubber S-2 has an unresolved alarm at handoff, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. Inez did not reconcile the records before the readiness rating. The applicable scope of every handoff item other than S-2 is established at 19:00. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The handoff log records current status and telemetry for S-2, but no owner acceptance or exception approval for the alarm.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-017", "id": "fast-41-diverse-032-017-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At the 19:00 handoff, outgoing lead Mara prepared the record for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff. Both documents are the complete governing record set for S-2's scope. Scrubber S-2 has an unresolved alarm at handoff, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. Inez did not reconcile the records before the readiness rating. The applicable scope of every handoff item other than S-2 is established at 19:00. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The handoff log records current status and telemetry for S-2, but no owner acceptance or exception approval for the alarm."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and all question bindings. The evidence consists of two complete factual sentences: \"The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.\" \"The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff.\" The counterfactual coherently changes only the exception register's scope assertion while retaining the unresolved conflict. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At the 19:00 handoff, outgoing lead Mara prepared the record for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff. Both documents are the complete governing record set for S-2's scope. Scrubber S-2 has an unresolved alarm at handoff, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. Inez did not reconcile the records before the readiness rating. The applicable scope of every handoff item other than S-2 is established at 19:00. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The handoff log records current status and telemetry for S-2, but no owner acceptance or exception approval for the alarm.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-017", "id": "fast-41-diverse-032-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At the 19:00 handoff, outgoing lead Mara prepared the record for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register updated at 18:47 lists scrubber S-2 as out of scope for the 19:00 handoff. Both documents are the complete governing record set for S-2's scope. Scrubber S-2 has an unresolved alarm at handoff, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. Inez did not reconcile the records before the readiness rating. The applicable scope of every handoff item other than S-2 is established at 19:00. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The handoff log records current status and telemetry for S-2, but no owner acceptance or exception approval for the alarm."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing reconciliation policy and unchanged question bindings for the 19:00 handoff, Mara, Dev, S-2, and Inez. The base context’s two scope records agree, so its unperformed reconciliation statement is unnecessary but not contradictory. The counterfactual changes the exception-register observation to an agreeing in-scope entry without creating contradictory duplicate measurements or assertions. The evidence spans are complete factual sentences: \"The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.\" \"The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff.\" Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At the 19:00 handoff, outgoing lead Mara is preparing the record for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. The two records are the complete governing record set for S-2’s scope. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. Inez did not perform that reconciliation before the rating. S-2 has an unresolved alarm at handoff, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Current status and telemetry for S-2 are present, but the owner-acceptance gap remains documented.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-020", "id": "fast-41-diverse-032-020-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At the 19:00 handoff, outgoing lead Mara is preparing the record for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. The two records are the complete governing record set for S-2’s scope. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. Inez did not perform that reconciliation before the rating. S-2 has an unresolved alarm at handoff, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Current status and telemetry for S-2 are present, but the owner-acceptance gap remains documented."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing reconciliation policy and unchanged question bindings for the 19:00 handoff, Mara, Dev, S-2, and Inez. The base context’s two scope records agree, so its unperformed reconciliation statement is unnecessary but not contradictory. The counterfactual changes the exception-register observation to an agreeing in-scope entry without creating contradictory duplicate measurements or assertions. The evidence spans are complete factual sentences: \"The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.\" \"The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff.\" Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At the 19:00 handoff, outgoing lead Mara is preparing the record for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff. The two records are the complete governing record set for S-2’s scope. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. Inez did not perform that reconciliation before the rating. S-2 has an unresolved alarm at handoff, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Current status and telemetry for S-2 are present, but the owner-acceptance gap remains documented.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-020", "id": "fast-41-diverse-032-020-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At the 19:00 handoff, outgoing lead Mara is preparing the record for incoming lead Dev. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register finalized at 18:50 lists scrubber S-2 as out of scope for the 19:00 handoff. The two records are the complete governing record set for S-2’s scope. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. Inez did not perform that reconciliation before the rating. S-2 has an unresolved alarm at handoff, and the appropriate owner has not accepted it. No temporary evidence exception covers the missing acceptance. The applicable scope of every handoff item other than S-2 is established at 19:00. Current status and telemetry for S-2 are present, but the owner-acceptance gap remains documented."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and unchanged question instructions are preserved. Entity, path, and time bindings remain unchanged. Evidence is retained as two complete factual sentences: “The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.” “The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff.” The counterfactual coherently changes only the exception register’s scope assertion and creates an unreconciled conflict. Neither context embeds an answer, code, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, shift lead Mara prepared the handoff for Dev, using the records available at the readiness cutoff. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff. The log identifies those two records as the complete governing record set for S-2’s scope. No reconciliation by operations coordinator Inez occurred before the readiness rating. S-2 has an unresolved alarm at handoff; the appropriate owner had not accepted it, and no temporary evidence exception covers that missing acceptance. Status and telemetry are available, while all required scope determinations for the other handoff items are established at 19:00. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The remaining checklist entries contain no additional scope records for S-2.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-026", "id": "fast-41-diverse-032-026-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, shift lead Mara prepared the handoff for Dev, using the records available at the readiness cutoff. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff. The log identifies those two records as the complete governing record set for S-2’s scope. No reconciliation by operations coordinator Inez occurred before the readiness rating. S-2 has an unresolved alarm at handoff; the appropriate owner had not accepted it, and no temporary evidence exception covers that missing acceptance. Status and telemetry are available, while all required scope determinations for the other handoff items are established at 19:00. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The remaining checklist entries contain no additional scope records for S-2."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and unchanged question instructions are preserved. Entity, path, and time bindings remain unchanged. Evidence is retained as two complete factual sentences: “The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.” “The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff.” The counterfactual coherently changes only the exception register’s scope assertion and creates an unreconciled conflict. Neither context embeds an answer, code, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, shift lead Mara prepared the handoff for Dev, using the records available at the readiness cutoff. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff. The log identifies those two records as the complete governing record set for S-2’s scope. No reconciliation by operations coordinator Inez occurred before the readiness rating. S-2 has an unresolved alarm at handoff; the appropriate owner had not accepted it, and no temporary evidence exception covers that missing acceptance. Status and telemetry are available, while all required scope determinations for the other handoff items are established at 19:00. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The remaining checklist entries contain no additional scope records for S-2.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-41-diverse-032-026", "id": "fast-41-diverse-032-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, shift lead Mara prepared the handoff for Dev, using the records available at the readiness cutoff. The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register finalized at 18:52 lists scrubber S-2 as out of scope for the 19:00 handoff. The log identifies those two records as the complete governing record set for S-2’s scope. No reconciliation by operations coordinator Inez occurred before the readiness rating. S-2 has an unresolved alarm at handoff; the appropriate owner had not accepted it, and no temporary evidence exception covers that missing acceptance. Status and telemetry are available, while all required scope determinations for the other handoff items are established at 19:00. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence. The remaining checklist entries contain no additional scope records for S-2."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, bindings, request, and exact evidence quotes: Outgoing shift lead Nia: “The 05:44 board export is attached. Chiller alert C-7 remains. Marco has C-7. Eli, did you acknowledge that alert?” Incoming shift lead Eli: “Yes, I did. I accept the shift.” Operations coordinator Marco provides no acknowledgment in the record. Each evidence array has two complete factual sentences, and the counterfactual coherently changes only Marco’s acknowledgment while preserving all other observations.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"case_note\":\"The reviewed handoff is the 06:00 shift change. Status-board export E is attached, and its generation timestamp is 05:44, within the preceding 30-minute window. The unresolved-item set at that change contains exactly C-7. C-7 is an equipment alert. Marco serves as the operations coordinator for this shift change. The incoming lead explicitly accepts the 06:00 shift. The following two record excerpts are retained as verified evidence.\",\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\",\"evidence\":[\"In the evaluated handoff record, Marco is recorded as the owner of C-7.\",\"In the evaluated handoff record, Marco explicitly acknowledges C-7.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the evaluated handoff record, Marco is recorded as the owner of C-7."}, {"path": ["evidence", "1"], "text": "In the evaluated handoff record, Marco explicitly acknowledges C-7."}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the evaluated handoff record, Marco is recorded as the owner of C-7.", "negative_left": "In the evaluated handoff record, Marco is recorded as the owner of C-7.", "negative_right": "In the evaluated handoff record, Marco does not explicitly acknowledge C-7.", "right": "In the evaluated handoff record, Marco explicitly acknowledges C-7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-033-009", "id": "fast-41-diverse-033-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"case_note": "The reviewed handoff is the 06:00 shift change. Status-board export E is attached, and its generation timestamp is 05:44, within the preceding 30-minute window. The unresolved-item set at that change contains exactly C-7. C-7 is an equipment alert. Marco serves as the operations coordinator for this shift change. The incoming lead explicitly accepts the 06:00 shift. The following two record excerpts are retained as verified evidence.", "context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["In the evaluated handoff record, Marco is recorded as the owner of C-7.", "In the evaluated handoff record, Marco explicitly acknowledges C-7."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, bindings, request, and exact evidence quotes: Outgoing shift lead Nia: “The 05:44 board export is attached. Chiller alert C-7 remains. Marco has C-7. Eli, did you acknowledge that alert?” Incoming shift lead Eli: “Yes, I did. I accept the shift.” Operations coordinator Marco provides no acknowledgment in the record. Each evidence array has two complete factual sentences, and the counterfactual coherently changes only Marco’s acknowledgment while preserving all other observations.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"case_note\":\"The reviewed handoff is the 06:00 shift change. Status-board export E is attached, and its generation timestamp is 05:44, within the preceding 30-minute window. The unresolved-item set at that change contains exactly C-7. C-7 is an equipment alert. Marco serves as the operations coordinator for this shift change. The incoming lead explicitly accepts the 06:00 shift. The following two record excerpts are retained as verified evidence.\",\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\",\"evidence\":[\"In the evaluated handoff record, Marco is recorded as the owner of C-7.\",\"In the evaluated handoff record, Marco explicitly acknowledges C-7.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the evaluated handoff record, Marco is recorded as the owner of C-7."}, {"path": ["evidence", "1"], "text": "In the evaluated handoff record, Marco explicitly acknowledges C-7."}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the evaluated handoff record, Marco is recorded as the owner of C-7.", "negative_left": "In the evaluated handoff record, Marco is recorded as the owner of C-7.", "negative_right": "In the evaluated handoff record, Marco does not explicitly acknowledge C-7.", "right": "In the evaluated handoff record, Marco explicitly acknowledges C-7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-033-009", "id": "fast-41-diverse-033-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"case_note": "The reviewed handoff is the 06:00 shift change. Status-board export E is attached, and its generation timestamp is 05:44, within the preceding 30-minute window. The unresolved-item set at that change contains exactly C-7. C-7 is an equipment alert. Marco serves as the operations coordinator for this shift change. The incoming lead explicitly accepts the 06:00 shift. The following two record excerpts are retained as verified evidence.", "context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["In the evaluated handoff record, Marco is recorded as the owner of C-7.", "In the evaluated handoff record, Marco does not explicitly acknowledge C-7."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and question bindings, and each has two complete factual evidence sentences. The counterfactual coherently changes only the acknowledgment observation without contradictory duplicates. Neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready. Export E is attached to the evaluated handoff and was generated at 05:44. C-7 is the only item unresolved at the shift change, and it is an equipment alert. Marco is the operations coordinator, and the incoming lead explicitly accepts the shift.\",\"evidence\":[\"In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7.\",\"The evaluated handoff record contains an explicit acknowledgment of C-7 by Marco.\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7."}, {"path": ["evidence", "1"], "text": "The evaluated handoff record contains an explicit acknowledgment of C-7 by Marco."}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7.", "negative_left": "In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7.", "negative_right": "The evaluated handoff record contains no explicit acknowledgment of C-7 by Marco.", "right": "The evaluated handoff record contains an explicit acknowledgment of C-7 by Marco."}, "verifier_independent_model": false}, "family": "fast-41-diverse-033-010", "id": "fast-41-diverse-033-010-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready. Export E is attached to the evaluated handoff and was generated at 05:44. C-7 is the only item unresolved at the shift change, and it is an equipment alert. Marco is the operations coordinator, and the incoming lead explicitly accepts the shift.", "evidence": ["In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7.", "The evaluated handoff record contains an explicit acknowledgment of C-7 by Marco."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and question bindings, and each has two complete factual evidence sentences. The counterfactual coherently changes only the acknowledgment observation without contradictory duplicates. Neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready. Export E is attached to the evaluated handoff and was generated at 05:44. C-7 is the only item unresolved at the shift change, and it is an equipment alert. Marco is the operations coordinator, and the incoming lead explicitly accepts the shift.\",\"evidence\":[\"In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7.\",\"The evaluated handoff record contains an explicit acknowledgment of C-7 by Marco.\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7."}, {"path": ["evidence", "1"], "text": "The evaluated handoff record contains an explicit acknowledgment of C-7 by Marco."}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7.", "negative_left": "In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7.", "negative_right": "The evaluated handoff record contains no explicit acknowledgment of C-7 by Marco.", "right": "The evaluated handoff record contains an explicit acknowledgment of C-7 by Marco."}, "verifier_independent_model": false}, "family": "fast-41-diverse-033-010", "id": "fast-41-diverse-033-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready. Export E is attached to the evaluated handoff and was generated at 05:44. C-7 is the only item unresolved at the shift change, and it is an equipment alert. Marco is the operations coordinator, and the incoming lead explicitly accepts the shift.", "evidence": ["In the evaluated 06:00 handoff, Marco is recorded as the owner of unresolved equipment alert C-7.", "The evaluated handoff record contains no explicit acknowledgment of C-7 by Marco."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the original rubric, and both contexts retain the same routing and readiness policy. Both contexts keep the exact handoff, C-7, 06:00, and requested routing and acknowledgment bindings. The two focus spans are complete factual sentences: “In the evaluated handoff record, Marco is recorded as the owner of C-7.” and “The evaluated handoff record states that Marco explicitly acknowledged C-7.” The counterfactual coherently changes Marco’s acknowledgment to a decline while leaving ownership, timing, routing, export, and incoming acceptance consistent. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"evidence\":[\"In the evaluated handoff record, Marco is recorded as the owner of C-7.\",\"The evaluated handoff record states that Marco explicitly acknowledged C-7.\",\"Status-board export E is attached to the evaluated handoff and was generated at 05:44.\",\"The evaluated handoff occurs at the 06:00 shift change, and C-7 is the only unresolved item.\",\"C-7 is an equipment alert. Marco is the operations coordinator for the 06:00 shift change.\",\"The incoming lead for the 06:00 shift explicitly accepts that shift.\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the evaluated handoff record, Marco is recorded as the owner of C-7."}, {"path": ["evidence", "1"], "text": "The evaluated handoff record states that Marco explicitly acknowledged C-7."}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the evaluated handoff record, Marco is recorded as the owner of C-7.", "negative_left": "In the evaluated handoff record, Marco is recorded as the owner of C-7.", "negative_right": "The evaluated handoff record states that Marco explicitly declined to acknowledge C-7.", "right": "The evaluated handoff record states that Marco explicitly acknowledged C-7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-033-013", "id": "fast-41-diverse-033-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["In the evaluated handoff record, Marco is recorded as the owner of C-7.", "The evaluated handoff record states that Marco explicitly acknowledged C-7.", "Status-board export E is attached to the evaluated handoff and was generated at 05:44.", "The evaluated handoff occurs at the 06:00 shift change, and C-7 is the only unresolved item.", "C-7 is an equipment alert. Marco is the operations coordinator for the 06:00 shift change.", "The incoming lead for the 06:00 shift explicitly accepts that shift."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the original rubric, and both contexts retain the same routing and readiness policy. Both contexts keep the exact handoff, C-7, 06:00, and requested routing and acknowledgment bindings. The two focus spans are complete factual sentences: “In the evaluated handoff record, Marco is recorded as the owner of C-7.” and “The evaluated handoff record states that Marco explicitly acknowledged C-7.” The counterfactual coherently changes Marco’s acknowledgment to a decline while leaving ownership, timing, routing, export, and incoming acceptance consistent. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the exact-set statement for unresolved items. The focus atom concerns the factual acknowledgment relationship rather than a policy classification. The base and counter assignments can be realized with only that acknowledgment changing. Policy evidence preserves the substantive routing and readiness rubric from the original state; retaining the state request as additional evidence is unnecessary but does not make preservation incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail every Ready requirement: an export generated at 05:44 is attached to the 06:00 handoff, hence is from the prior 30 minutes; C-7 is the only unresolved item; it is an equipment alert owned by Marco, the operations coordinator; its recorded owner explicitly acknowledges it; and the incoming lead accepts the shift.", "rule_index": 0, "sound": true}, {"reason": "The recorded owner’s explicit acknowledgment is refuted. Because acknowledgment by every unresolved item’s owner is mandatory for Ready, this condition is sufficient for a false/No decision. With incoming acceptance present, the Conditional exception does not override that result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evaluated handoff occurs at the 06:00 shift change."}, {"id": "a2", "statement": "Status-board export E is attached to the evaluated 06:00 handoff."}, {"id": "a3", "statement": "Status-board export E was generated at 05:44."}, {"id": "a4", "statement": "The set of items unresolved at the 06:00 shift change is exactly {C-7}."}, {"id": "a5", "statement": "C-7 is an equipment alert."}, {"id": "a6", "statement": "Marco is the operations coordinator for the 06:00 shift change."}, {"id": "a7", "statement": "Marco is recorded as the owner of C-7 in the evaluated handoff."}, {"id": "a8", "statement": "The recorded owner of C-7 explicitly acknowledges C-7 in the evaluated handoff record."}, {"id": "a9", "statement": "The incoming lead for the 06:00 shift explicitly accepts that shift."}], "base_state_json": "{\"context\":\"At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.\",\"evidence\":[\"In the evaluated handoff record, Marco is recorded as the owner of C-7.\",\"The evaluated handoff record states that Marco explicitly acknowledged C-7.\",\"Status-board export E is attached to the evaluated handoff and was generated at 05:44.\",\"The evaluated handoff occurs at the 06:00 shift change, and C-7 is the only unresolved item.\",\"C-7 is an equipment alert. Marco is the operations coordinator for the 06:00 shift change.\",\"The incoming lead for the 06:00 shift explicitly accepts that shift.\"],\"request\":\"Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the evaluated handoff record, Marco is recorded as the owner of C-7."}, {"path": ["evidence", "1"], "text": "The evaluated handoff record states that Marco explicitly acknowledged C-7."}], "policy_evidence": [{"path": ["context"], "text": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready."}, {"path": ["request"], "text": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}], "rules": [{"justification": "Export E is attached and its 05:44 generation time is within the 30 minutes preceding 06:00. C-7 is the only unresolved item, is an equipment alert, and is recorded as owned by the operations coordinator, so it is correctly routed. Its recorded owner explicitly acknowledges it, and the incoming lead accepts the shift. Thus all Ready requirements are satisfied.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Although the timely export is attached, C-7 is the only unresolved item and is correctly routed to the operations coordinator, and the incoming lead accepts the shift, C-7's recorded owner does not explicitly acknowledge it. A required Ready condition is therefore missing.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the evaluated handoff record, Marco is recorded as the owner of C-7.", "negative_left": "In the evaluated handoff record, Marco is recorded as the owner of C-7.", "negative_right": "The evaluated handoff record states that Marco explicitly declined to acknowledge C-7.", "right": "The evaluated handoff record states that Marco explicitly acknowledged C-7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-033-013", "id": "fast-41-diverse-033-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one Ready requirement is missing; assign the applicable lower readiness level under the rubric.", "true": "Yes — all Ready requirements are satisfied, including explicit acknowledgment by each unresolved item's owner."}, "instructions": "Resolve references using speaker attribution, then apply the stated rubric. Answer yes only if the handoff qualifies as Ready.", "type": "noul"}}, "state": {"context": "At the 06:00 shift change, unresolved equipment alerts must be routed to the operations coordinator. A handoff is Ready only if a status-board export from the prior 30 minutes is attached, every unresolved item has the correct owner and that owner explicitly acknowledges it, and the incoming lead accepts the shift. It is Conditional only when all evidence is present but incoming acceptance is pending; otherwise it is Not ready.", "evidence": ["In the evaluated handoff record, Marco is recorded as the owner of C-7.", "The evaluated handoff record states that Marco explicitly declined to acknowledge C-7.", "Status-board export E is attached to the evaluated handoff and was generated at 05:44.", "The evaluated handoff occurs at the 06:00 shift change, and C-7 is the only unresolved item.", "C-7 is an equipment alert. Marco is the operations coordinator for the 06:00 shift change.", "The incoming lead for the 06:00 shift explicitly accepts that shift."], "request": "Is this handoff Ready? Also determine whether C-7 was correctly routed and whether the required owner acknowledgment is present."}}, "method": "c2d", "provenance": {"source_id": "diverse-033", "source_is_synthetic": true, "source_sha256": "c545ca237db1691a1e8a3ffcbf509a91adcbe908e56ac2cf3c6e02353f4e0139", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, and both contexts add no conflicting policy. The Line 4 pump, handoff scope, production-item scope, and timing bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.\" and \"Mara Chen made the 17:48 acknowledgment on ticket M-882.\" The counterfactual coherently changes the acknowledgment maker to Luis Ortega without creating contradictory duplicate assertions. Neither full context embeds an answer, code, rule table, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Outgoing shift lead\",\"text\":\"The production handoff covers the Line 4 pump and a completed filter swap. The pump was stable at 6.1 mm/s at 17:42, with timestamped trend snapshot TS-144 attached; the filter swap has timestamped work order W-219 with a 17:30 completion photo. The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item. Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item. The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item. The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48. Mara Chen made the 17:48 acknowledgment on ticket M-882.\"},{\"speaker\":\"Incoming shift lead\",\"text\":\"I will monitor the Line 4 trend until the inspection report arrives. Cafeteria freezer and parking repaint notices are facilities announcements outside the production handoff scope.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["1", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"path": ["1", "text"], "text": "Mara Chen made the 17:48 acknowledgment on ticket M-882."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_right": "Luis Ortega made the 17:48 acknowledgment on ticket M-882.", "right": "Mara Chen made the 17:48 acknowledgment on ticket M-882."}, "verifier_independent_model": false}, "family": "fast-41-diverse-034-002", "id": "fast-41-diverse-034-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Outgoing shift lead", "text": "The production handoff covers the Line 4 pump and a completed filter swap. The pump was stable at 6.1 mm/s at 17:42, with timestamped trend snapshot TS-144 attached; the filter swap has timestamped work order W-219 with a 17:30 completion photo. The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"speaker": "Operations coordinator", "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item. Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item. The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item. The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48. Mara Chen made the 17:48 acknowledgment on ticket M-882."}, {"speaker": "Incoming shift lead", "text": "I will monitor the Line 4 trend until the inspection report arrives. Cafeteria freezer and parking repaint notices are facilities announcements outside the production handoff scope."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, and both contexts add no conflicting policy. The Line 4 pump, handoff scope, production-item scope, and timing bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.\" and \"Mara Chen made the 17:48 acknowledgment on ticket M-882.\" The counterfactual coherently changes the acknowledgment maker to Luis Ortega without creating contradictory duplicate assertions. Neither full context embeds an answer, code, rule table, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Outgoing shift lead\",\"text\":\"The production handoff covers the Line 4 pump and a completed filter swap. The pump was stable at 6.1 mm/s at 17:42, with timestamped trend snapshot TS-144 attached; the filter swap has timestamped work order W-219 with a 17:30 completion photo. The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item. Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item. The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item. The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48. Mara Chen made the 17:48 acknowledgment on ticket M-882.\"},{\"speaker\":\"Incoming shift lead\",\"text\":\"I will monitor the Line 4 trend until the inspection report arrives. Cafeteria freezer and parking repaint notices are facilities announcements outside the production handoff scope.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["1", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"path": ["1", "text"], "text": "Mara Chen made the 17:48 acknowledgment on ticket M-882."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_right": "Luis Ortega made the 17:48 acknowledgment on ticket M-882.", "right": "Mara Chen made the 17:48 acknowledgment on ticket M-882."}, "verifier_independent_model": false}, "family": "fast-41-diverse-034-002", "id": "fast-41-diverse-034-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Outgoing shift lead", "text": "The production handoff covers the Line 4 pump and a completed filter swap. The pump was stable at 6.1 mm/s at 17:42, with timestamped trend snapshot TS-144 attached; the filter swap has timestamped work order W-219 with a 17:30 completion photo. The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"speaker": "Operations coordinator", "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item. Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item. The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item. The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48. Luis Ortega made the 17:48 acknowledgment on ticket M-882."}, {"speaker": "Incoming shift lead", "text": "I will monitor the Line 4 trend until the inspection report arrives. Cafeteria freezer and parking repaint notices are facilities announcements outside the production handoff scope."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the question and policy, remain coherent because only the acknowledgment actor changes, and contain no answer leakage; the evidence quotes are “Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.” and “The 17:48 acknowledgment on ticket M-882 was made by Mara Chen.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Outgoing shift lead\",\"text\":\"The shift-handoff scope contains two production items: the Line 4 pump vibration and a completed filter-swap job. The vibration status was stable at 6.1 mm/s at 17:42, and timestamped trend snapshot TS-144 is attached. The filter-swap job has a current completion status and a timestamped 17:30 completion photo in work order W-219. The Line 4 pump vibration is the only unresolved production item in scope.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item. Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Incoming shift lead\",\"text\":\"The handoff log contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48. The 17:48 acknowledgment on ticket M-882 was made by Mara Chen. Facilities notices about a cafeteria freezer alarm and tomorrow’s parking repaint are outside the production-item scope.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["1", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"path": ["3", "text"], "text": "The 17:48 acknowledgment on ticket M-882 was made by Mara Chen."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_right": "The 17:48 acknowledgment on ticket M-882 was made by Diego Ruiz.", "right": "The 17:48 acknowledgment on ticket M-882 was made by Mara Chen."}, "verifier_independent_model": false}, "family": "fast-41-diverse-034-007", "id": "fast-41-diverse-034-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Outgoing shift lead", "text": "The shift-handoff scope contains two production items: the Line 4 pump vibration and a completed filter-swap job. The vibration status was stable at 6.1 mm/s at 17:42, and timestamped trend snapshot TS-144 is attached. The filter-swap job has a current completion status and a timestamped 17:30 completion photo in work order W-219. The Line 4 pump vibration is the only unresolved production item in scope."}, {"speaker": "Operations coordinator", "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item. Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"speaker": "Incoming shift lead", "text": "The handoff log contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"speaker": "Operations coordinator", "text": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48. The 17:48 acknowledgment on ticket M-882 was made by Mara Chen. Facilities notices about a cafeteria freezer alarm and tomorrow’s parking repaint are outside the production-item scope."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the question and policy, remain coherent because only the acknowledgment actor changes, and contain no answer leakage; the evidence quotes are “Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.” and “The 17:48 acknowledgment on ticket M-882 was made by Mara Chen.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Outgoing shift lead\",\"text\":\"The shift-handoff scope contains two production items: the Line 4 pump vibration and a completed filter-swap job. The vibration status was stable at 6.1 mm/s at 17:42, and timestamped trend snapshot TS-144 is attached. The filter-swap job has a current completion status and a timestamped 17:30 completion photo in work order W-219. The Line 4 pump vibration is the only unresolved production item in scope.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item. Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Incoming shift lead\",\"text\":\"The handoff log contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48. The 17:48 acknowledgment on ticket M-882 was made by Mara Chen. Facilities notices about a cafeteria freezer alarm and tomorrow’s parking repaint are outside the production-item scope.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["1", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"path": ["3", "text"], "text": "The 17:48 acknowledgment on ticket M-882 was made by Mara Chen."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_right": "The 17:48 acknowledgment on ticket M-882 was made by Diego Ruiz.", "right": "The 17:48 acknowledgment on ticket M-882 was made by Mara Chen."}, "verifier_independent_model": false}, "family": "fast-41-diverse-034-007", "id": "fast-41-diverse-034-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Outgoing shift lead", "text": "The shift-handoff scope contains two production items: the Line 4 pump vibration and a completed filter-swap job. The vibration status was stable at 6.1 mm/s at 17:42, and timestamped trend snapshot TS-144 is attached. The filter-swap job has a current completion status and a timestamped 17:30 completion photo in work order W-219. The Line 4 pump vibration is the only unresolved production item in scope."}, {"speaker": "Operations coordinator", "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item. Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"speaker": "Incoming shift lead", "text": "The handoff log contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"speaker": "Operations coordinator", "text": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48. The 17:48 acknowledgment on ticket M-882 was made by Diego Ruiz. Facilities notices about a cafeteria freezer alarm and tomorrow’s parking repaint are outside the production-item scope."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and unchanged question scope. The handoff, unresolved item, incoming lead, owner, and scoring-time bindings remain intact. The evidence contains exactly two complete factual sentences: \"At scoring time, the handoff lists unresolved item U as a refrigeration alarm.\" and \"At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team.\" The counterfactual changes routing to Engineering without introducing contradictory duplicate measurements or assertions. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship rather than a policy conclusion or bundled final classification. The focus atom concerns the factual routing of item U. The base and counter assignments are jointly realizable under the state policy and differ only in whether U is routed to Facilities. The policy evidence correctly preserves the state-originating requirements needed to interpret the unchanged question, including the artifact, routing, paraphrase, and refrigeration-owner rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "With exactly one unresolved item, the conjunction establishes a timestamped status artifact, correct routing to Facilities for the refrigeration alarm, and an accurate own-word paraphrase of both status and next action. The sole owner's acknowledgment is explicitly refuted, so score 1 is entailed by the ordered rubric.", "rule_index": 0, "sound": true}, {"reason": "Facilities is the responsible owner for the sole refrigeration alarm, while routing to Facilities is explicitly refuted. Thus the unresolved item is routed to the wrong owner or no owner, which is sufficient for score 0 regardless of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At scoring time, the number of unresolved operational items in the handoff is exactly one."}, {"id": "a2", "statement": "The sole unresolved operational item U is a refrigeration alarm."}, {"id": "a3", "statement": "A status artifact for unresolved item U is attached to the handoff."}, {"id": "a4", "statement": "The attached status artifact for unresolved item U has a timestamp."}, {"id": "a5", "statement": "At scoring time, unresolved item U is routed to the Facilities team."}, {"id": "a6", "statement": "Incoming shift lead Dev's paraphrase accurately represents the status of unresolved item U."}, {"id": "a7", "statement": "Incoming shift lead Dev expresses the status of unresolved item U in his own words."}, {"id": "a8", "statement": "Incoming shift lead Dev's paraphrase accurately represents the next action for unresolved item U."}, {"id": "a9", "statement": "Incoming shift lead Dev expresses the next action for unresolved item U in his own words."}, {"id": "a10", "statement": "The owner of unresolved item U has acknowledged receipt of the item."}], "base_state_json": "\"At scoring time, the handoff lists unresolved item U as a refrigeration alarm. A timestamped status artifact for U is attached, and U is the only unresolved operational item in the handoff. Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count. Facilities handles refrigeration alarms. At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team. Incoming shift lead Dev accurately explains U’s current condition in his own words and separately describes the next action in his own words. His account matches the attached record. The owner has not acknowledged receipt of U by the scoring deadline. Operations coordinator Lena reviews the handoff before the shift change, and no other unresolved operational item is recorded.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At scoring time, the handoff lists unresolved item U as a refrigeration alarm."}, {"path": [], "text": "At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team."}], "policy_evidence": [{"path": [], "text": "Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count."}, {"path": [], "text": "Facilities handles refrigeration alarms."}], "rules": [{"justification": "There is exactly one unresolved item; it has the required timestamped status artifact, is correctly routed to Facilities under the refrigeration-alarm policy, and has an accurate own-word paraphrase of both status and next action. The owner's explicitly absent acknowledgment therefore makes the handoff minimally ready.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}, {"justification": "The sole unresolved item is a refrigeration alarm, for which Facilities is the responsible owner, but the item is explicitly not routed to Facilities. This wrong routing is sufficient for a not-ready score despite the other documented requirements.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At scoring time, the handoff lists unresolved item U as a refrigeration alarm.", "negative_left": "At scoring time, the handoff lists unresolved item U as a refrigeration alarm.", "negative_right": "At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Engineering team.", "right": "At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team."}, "verifier_independent_model": false}, "family": "fast-41-diverse-035-026", "id": "fast-41-diverse-035-026-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one unresolved operational item lacks a required status artifact, is routed to the wrong or no owner, or lacks an accurate own-word paraphrase from the incoming lead.", "1 — Minimally ready: Every unresolved item has the required artifact, correct owner, and valid paraphrase, but at least one owner has not acknowledged receipt.", "2 — Provisionally ready: All required fields and owner acknowledgments are present, but at least one unresolved item lacks a follow-up time or completion checkpoint.", "3 — Ready: Required evidence, routing, paraphrases, acknowledgments, and follow-up checkpoints are complete; only minor nonoperational clarification remains.", "4 — Fully ready: Every item is fully documented, correctly routed, paraphrased, acknowledged, and scheduled, with no clarification or correction needed."], "instructions": "Rate the handoff’s readiness using the ordered rubric. Check the routing, required status evidence, and whether the incoming lead supplied a valid paraphrase.", "type": "score"}}, "state": "At scoring time, the handoff lists unresolved item U as a refrigeration alarm. A timestamped status artifact for U is attached, and U is the only unresolved operational item in the handoff. Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count. Facilities handles refrigeration alarms. At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team. Incoming shift lead Dev accurately explains U’s current condition in his own words and separately describes the next action in his own words. His account matches the attached record. The owner has not acknowledged receipt of U by the scoring deadline. Operations coordinator Lena reviews the handoff before the shift change, and no other unresolved operational item is recorded."}, "method": "c2d", "provenance": {"source_id": "diverse-035", "source_is_synthetic": true, "source_sha256": "ef2504c9a18e9a4c12f01c98be90595f7a44fb099e598d65aa259ad0418ac996", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and unchanged question scope. The handoff, unresolved item, incoming lead, owner, and scoring-time bindings remain intact. The evidence contains exactly two complete factual sentences: \"At scoring time, the handoff lists unresolved item U as a refrigeration alarm.\" and \"At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team.\" The counterfactual changes routing to Engineering without introducing contradictory duplicate measurements or assertions. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship rather than a policy conclusion or bundled final classification. The focus atom concerns the factual routing of item U. The base and counter assignments are jointly realizable under the state policy and differ only in whether U is routed to Facilities. The policy evidence correctly preserves the state-originating requirements needed to interpret the unchanged question, including the artifact, routing, paraphrase, and refrigeration-owner rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "With exactly one unresolved item, the conjunction establishes a timestamped status artifact, correct routing to Facilities for the refrigeration alarm, and an accurate own-word paraphrase of both status and next action. The sole owner's acknowledgment is explicitly refuted, so score 1 is entailed by the ordered rubric.", "rule_index": 0, "sound": true}, {"reason": "Facilities is the responsible owner for the sole refrigeration alarm, while routing to Facilities is explicitly refuted. Thus the unresolved item is routed to the wrong owner or no owner, which is sufficient for score 0 regardless of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At scoring time, the number of unresolved operational items in the handoff is exactly one."}, {"id": "a2", "statement": "The sole unresolved operational item U is a refrigeration alarm."}, {"id": "a3", "statement": "A status artifact for unresolved item U is attached to the handoff."}, {"id": "a4", "statement": "The attached status artifact for unresolved item U has a timestamp."}, {"id": "a5", "statement": "At scoring time, unresolved item U is routed to the Facilities team."}, {"id": "a6", "statement": "Incoming shift lead Dev's paraphrase accurately represents the status of unresolved item U."}, {"id": "a7", "statement": "Incoming shift lead Dev expresses the status of unresolved item U in his own words."}, {"id": "a8", "statement": "Incoming shift lead Dev's paraphrase accurately represents the next action for unresolved item U."}, {"id": "a9", "statement": "Incoming shift lead Dev expresses the next action for unresolved item U in his own words."}, {"id": "a10", "statement": "The owner of unresolved item U has acknowledged receipt of the item."}], "base_state_json": "\"At scoring time, the handoff lists unresolved item U as a refrigeration alarm. A timestamped status artifact for U is attached, and U is the only unresolved operational item in the handoff. Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count. Facilities handles refrigeration alarms. At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team. Incoming shift lead Dev accurately explains U’s current condition in his own words and separately describes the next action in his own words. His account matches the attached record. The owner has not acknowledged receipt of U by the scoring deadline. Operations coordinator Lena reviews the handoff before the shift change, and no other unresolved operational item is recorded.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At scoring time, the handoff lists unresolved item U as a refrigeration alarm."}, {"path": [], "text": "At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team."}], "policy_evidence": [{"path": [], "text": "Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count."}, {"path": [], "text": "Facilities handles refrigeration alarms."}], "rules": [{"justification": "There is exactly one unresolved item; it has the required timestamped status artifact, is correctly routed to Facilities under the refrigeration-alarm policy, and has an accurate own-word paraphrase of both status and next action. The owner's explicitly absent acknowledgment therefore makes the handoff minimally ready.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}, {"justification": "The sole unresolved item is a refrigeration alarm, for which Facilities is the responsible owner, but the item is explicitly not routed to Facilities. This wrong routing is sufficient for a not-ready score despite the other documented requirements.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At scoring time, the handoff lists unresolved item U as a refrigeration alarm.", "negative_left": "At scoring time, the handoff lists unresolved item U as a refrigeration alarm.", "negative_right": "At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Engineering team.", "right": "At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Facilities team."}, "verifier_independent_model": false}, "family": "fast-41-diverse-035-026", "id": "fast-41-diverse-035-026-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one unresolved operational item lacks a required status artifact, is routed to the wrong or no owner, or lacks an accurate own-word paraphrase from the incoming lead.", "1 — Minimally ready: Every unresolved item has the required artifact, correct owner, and valid paraphrase, but at least one owner has not acknowledged receipt.", "2 — Provisionally ready: All required fields and owner acknowledgments are present, but at least one unresolved item lacks a follow-up time or completion checkpoint.", "3 — Ready: Required evidence, routing, paraphrases, acknowledgments, and follow-up checkpoints are complete; only minor nonoperational clarification remains.", "4 — Fully ready: Every item is fully documented, correctly routed, paraphrased, acknowledged, and scheduled, with no clarification or correction needed."], "instructions": "Rate the handoff’s readiness using the ordered rubric. Check the routing, required status evidence, and whether the incoming lead supplied a valid paraphrase.", "type": "score"}}, "state": "At scoring time, the handoff lists unresolved item U as a refrigeration alarm. A timestamped status artifact for U is attached, and U is the only unresolved operational item in the handoff. Policy requires each unresolved item to include a timestamped status artifact, routing to the responsible owner, and an incoming shift lead’s own-word paraphrase of the status and next action; copied wording does not count. Facilities handles refrigeration alarms. At scoring time, the handoff routes every unresolved refrigeration alarm exclusively to the Engineering team. Incoming shift lead Dev accurately explains U’s current condition in his own words and separately describes the next action in his own words. His account matches the attached record. The owner has not acknowledged receipt of U by the scoring deadline. Operations coordinator Lena reviews the handoff before the shift change, and no other unresolved operational item is recorded."}, "method": "c2d", "provenance": {"source_id": "diverse-035", "source_is_synthetic": true, "source_sha256": "ef2504c9a18e9a4c12f01c98be90595f7a44fb099e598d65aa259ad0418ac996", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve every governing policy. The 18:00 packaging handoff, unresolved-item scope, and request bindings remain unchanged. The evidence spans are complete factual sentences: “At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.” “At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field.” “At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.” “At 18:00, Taylor Kim is identified in the Filler 2 pressure alarm's owner field.” The counterfactual is coherent because Taylor Kim is distinct from Jordan Lee and replaces him as the identified alarm owner without duplicate contradictions. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The unresolved handoff list contains only the Conveyor 4 restart check and the Filler 2 pressure alarm. The former is routine, has a current status artifact and next action, and its owner is serving as Incoming shift lead. The latter is the only critical alarm, has a current status artifact, stated next action, attached diagnostic result, and identified owner. By 18:00, the Incoming shift lead has not acknowledged the handoff. The outgoing lead signed at 17:55.\",\"evidence\":[\"At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.\",\"At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator."}, {"path": ["evidence", "1"], "text": "At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.", "negative_left": "At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.", "negative_right": "At 18:00, Taylor Kim is identified in the Filler 2 pressure alarm's owner field.", "right": "At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-010", "id": "fast-41-diverse-036-010-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The unresolved handoff list contains only the Conveyor 4 restart check and the Filler 2 pressure alarm. The former is routine, has a current status artifact and next action, and its owner is serving as Incoming shift lead. The latter is the only critical alarm, has a current status artifact, stated next action, attached diagnostic result, and identified owner. By 18:00, the Incoming shift lead has not acknowledged the handoff. The outgoing lead signed at 17:55.", "evidence": ["At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.", "At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve every governing policy. The 18:00 packaging handoff, unresolved-item scope, and request bindings remain unchanged. The evidence spans are complete factual sentences: “At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.” “At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field.” “At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.” “At 18:00, Taylor Kim is identified in the Filler 2 pressure alarm's owner field.” The counterfactual is coherent because Taylor Kim is distinct from Jordan Lee and replaces him as the identified alarm owner without duplicate contradictions. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The unresolved handoff list contains only the Conveyor 4 restart check and the Filler 2 pressure alarm. The former is routine, has a current status artifact and next action, and its owner is serving as Incoming shift lead. The latter is the only critical alarm, has a current status artifact, stated next action, attached diagnostic result, and identified owner. By 18:00, the Incoming shift lead has not acknowledged the handoff. The outgoing lead signed at 17:55.\",\"evidence\":[\"At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.\",\"At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator."}, {"path": ["evidence", "1"], "text": "At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.", "negative_left": "At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.", "negative_right": "At 18:00, Taylor Kim is identified in the Filler 2 pressure alarm's owner field.", "right": "At 18:00, Jordan Lee is identified in the Filler 2 pressure alarm's owner field."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-010", "id": "fast-41-diverse-036-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The unresolved handoff list contains only the Conveyor 4 restart check and the Filler 2 pressure alarm. The former is routine, has a current status artifact and next action, and its owner is serving as Incoming shift lead. The latter is the only critical alarm, has a current status artifact, stated next action, attached diagnostic result, and identified owner. By 18:00, the Incoming shift lead has not acknowledged the handoff. The outgoing lead signed at 17:55.", "evidence": ["At 18:00, Jordan Lee, who is distinct from Taylor Kim, is serving as the Operations coordinator.", "At 18:00, Taylor Kim is identified in the Filler 2 pressure alarm's owner field."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy, and the unchanged questions object preserves all rubric criteria. Entity, path, time, and request-scope bindings remain consistent. The evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual coherently changes Jordan Lee's role without creating contradictory duplicate assertions. Neither context contains a gold answer, answer code, rationale, proposition identifier, or classifier instruction. Retained focus evidence: \"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.\" \"At 18:00, Jordan Lee is serving as the Operations coordinator.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.\",\"At 18:00, Jordan Lee is serving as the Operations coordinator.\",\"At 18:00, the unresolved handoff list contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm on that list.\",\"At 18:00, the Conveyor 4 restart check is a routine process follow-up.\",\"At 18:00, both unresolved items have current status artifacts and stated next actions.\",\"At 18:00, the Conveyor 4 restart check's owner field identifies the Incoming shift lead.\",\"At 18:00, the Filler 2 pressure alarm has an attached diagnostic result.\",\"By 18:00, the Incoming shift lead has not acknowledged the handoff.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee."}, {"path": ["evidence", "1"], "text": "At 18:00, Jordan Lee is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_right": "At 18:00, Jordan Lee is serving only as the Maintenance coordinator.", "right": "At 18:00, Jordan Lee is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-012", "id": "fast-41-diverse-036-012-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "At 18:00, Jordan Lee is serving as the Operations coordinator.", "At 18:00, the unresolved handoff list contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.", "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm on that list.", "At 18:00, the Conveyor 4 restart check is a routine process follow-up.", "At 18:00, both unresolved items have current status artifacts and stated next actions.", "At 18:00, the Conveyor 4 restart check's owner field identifies the Incoming shift lead.", "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result.", "By 18:00, the Incoming shift lead has not acknowledged the handoff."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy, and the unchanged questions object preserves all rubric criteria. Entity, path, time, and request-scope bindings remain consistent. The evidence spans are complete factual sentences rather than instructions or policy definitions. The counterfactual coherently changes Jordan Lee's role without creating contradictory duplicate assertions. Neither context contains a gold answer, answer code, rationale, proposition identifier, or classifier instruction. Retained focus evidence: \"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.\" \"At 18:00, Jordan Lee is serving as the Operations coordinator.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.\",\"At 18:00, Jordan Lee is serving as the Operations coordinator.\",\"At 18:00, the unresolved handoff list contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm on that list.\",\"At 18:00, the Conveyor 4 restart check is a routine process follow-up.\",\"At 18:00, both unresolved items have current status artifacts and stated next actions.\",\"At 18:00, the Conveyor 4 restart check's owner field identifies the Incoming shift lead.\",\"At 18:00, the Filler 2 pressure alarm has an attached diagnostic result.\",\"By 18:00, the Incoming shift lead has not acknowledged the handoff.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee."}, {"path": ["evidence", "1"], "text": "At 18:00, Jordan Lee is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_right": "At 18:00, Jordan Lee is serving only as the Maintenance coordinator.", "right": "At 18:00, Jordan Lee is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-012", "id": "fast-41-diverse-036-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "At 18:00, Jordan Lee is serving only as the Maintenance coordinator.", "At 18:00, the unresolved handoff list contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.", "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm on that list.", "At 18:00, the Conveyor 4 restart check is a routine process follow-up.", "At 18:00, both unresolved items have current status artifacts and stated next actions.", "At 18:00, the Conveyor 4 restart check's owner field identifies the Incoming shift lead.", "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result.", "By 18:00, the Incoming shift lead has not acknowledged the handoff."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions object preserves the complete policy, rubric, instructions, and request scope. Both contexts retain the packaging handoff, 18:00 shift-change, readiness, and unresolved-work bindings. The two focus evidence spans are complete factual sentences: \"At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.\" and \"At 18:00, the sole person serving as Operations coordinator is Dana Ruiz.\" The counterfactual coherently changes the Operations coordinator to Lee Morgan while retaining Dana Ruiz as the alarm owner. No context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"The unresolved handoff register contains exactly two items: the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"The Filler 2 pressure alarm is the only critical equipment alarm on the register.\",\"The Conveyor 4 restart check is a routine process follow-up.\",\"A current status artifact and a stated next action are recorded for each unresolved item.\",\"The Conveyor 4 restart check's owner field identifies the person serving as Incoming shift lead.\",\"At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.\",\"The Filler 2 pressure alarm has an attached diagnostic result.\",\"At 18:00, the sole person serving as Operations coordinator is Dana Ruiz.\",\"By 18:00, the Incoming shift lead had not acknowledged the handoff.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "5"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz."}, {"path": ["evidence", "7"], "text": "At 18:00, the sole person serving as Operations coordinator is Dana Ruiz."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.", "negative_right": "At 18:00, the sole person serving as Operations coordinator is Lee Morgan.", "right": "At 18:00, the sole person serving as Operations coordinator is Dana Ruiz."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-013", "id": "fast-41-diverse-036-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["The unresolved handoff register contains exactly two items: the Conveyor 4 restart check and the Filler 2 pressure alarm.", "The Filler 2 pressure alarm is the only critical equipment alarm on the register.", "The Conveyor 4 restart check is a routine process follow-up.", "A current status artifact and a stated next action are recorded for each unresolved item.", "The Conveyor 4 restart check's owner field identifies the person serving as Incoming shift lead.", "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.", "The Filler 2 pressure alarm has an attached diagnostic result.", "At 18:00, the sole person serving as Operations coordinator is Dana Ruiz.", "By 18:00, the Incoming shift lead had not acknowledged the handoff."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions object preserves the complete policy, rubric, instructions, and request scope. Both contexts retain the packaging handoff, 18:00 shift-change, readiness, and unresolved-work bindings. The two focus evidence spans are complete factual sentences: \"At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.\" and \"At 18:00, the sole person serving as Operations coordinator is Dana Ruiz.\" The counterfactual coherently changes the Operations coordinator to Lee Morgan while retaining Dana Ruiz as the alarm owner. No context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"The unresolved handoff register contains exactly two items: the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"The Filler 2 pressure alarm is the only critical equipment alarm on the register.\",\"The Conveyor 4 restart check is a routine process follow-up.\",\"A current status artifact and a stated next action are recorded for each unresolved item.\",\"The Conveyor 4 restart check's owner field identifies the person serving as Incoming shift lead.\",\"At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.\",\"The Filler 2 pressure alarm has an attached diagnostic result.\",\"At 18:00, the sole person serving as Operations coordinator is Dana Ruiz.\",\"By 18:00, the Incoming shift lead had not acknowledged the handoff.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "5"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz."}, {"path": ["evidence", "7"], "text": "At 18:00, the sole person serving as Operations coordinator is Dana Ruiz."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.", "negative_right": "At 18:00, the sole person serving as Operations coordinator is Lee Morgan.", "right": "At 18:00, the sole person serving as Operations coordinator is Dana Ruiz."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-013", "id": "fast-41-diverse-036-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["The unresolved handoff register contains exactly two items: the Conveyor 4 restart check and the Filler 2 pressure alarm.", "The Filler 2 pressure alarm is the only critical equipment alarm on the register.", "The Conveyor 4 restart check is a routine process follow-up.", "A current status artifact and a stated next action are recorded for each unresolved item.", "The Conveyor 4 restart check's owner field identifies the person serving as Incoming shift lead.", "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Ruiz.", "The Filler 2 pressure alarm has an attached diagnostic result.", "At 18:00, the sole person serving as Operations coordinator is Lee Morgan.", "By 18:00, the Incoming shift lead had not acknowledged the handoff."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy, and the unchanged questions object preserves all rubric rules. Entity, path, and 18:00 bindings remain unchanged. The evidence spans are complete factual sentences: “At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.” and “At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role.” The counterfactual coherently changes Jordan Lee’s role to Maintenance coordinator, making the alarm misrouted without creating contradictory duplicate assertions. Neither context embeds a gold answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The unresolved handoff register contains only the Conveyor 4 restart check and the Filler 2 pressure alarm. The alarm is the sole critical equipment alarm; the Conveyor item is routine. Both entries have current status artifacts and stated next actions. The alarm has an attached diagnostic result. The Conveyor owner field names the Incoming shift lead. At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee. At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role. By 18:00, the Incoming shift lead has not acknowledged the handoff.\",\"policy\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["context"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee."}, {"path": ["context"], "text": "At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_right": "At 18:00, Jordan Lee is serving in the Maintenance coordinator role, and the shift record assigns each person exactly one coordinator role.", "right": "At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-015", "id": "fast-41-diverse-036-015-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The unresolved handoff register contains only the Conveyor 4 restart check and the Filler 2 pressure alarm. The alarm is the sole critical equipment alarm; the Conveyor item is routine. Both entries have current status artifacts and stated next actions. The alarm has an attached diagnostic result. The Conveyor owner field names the Incoming shift lead. At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee. At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role. By 18:00, the Incoming shift lead has not acknowledged the handoff.", "policy": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy, and the unchanged questions object preserves all rubric rules. Entity, path, and 18:00 bindings remain unchanged. The evidence spans are complete factual sentences: “At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.” and “At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role.” The counterfactual coherently changes Jordan Lee’s role to Maintenance coordinator, making the alarm misrouted without creating contradictory duplicate assertions. Neither context embeds a gold answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The unresolved handoff register contains only the Conveyor 4 restart check and the Filler 2 pressure alarm. The alarm is the sole critical equipment alarm; the Conveyor item is routine. Both entries have current status artifacts and stated next actions. The alarm has an attached diagnostic result. The Conveyor owner field names the Incoming shift lead. At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee. At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role. By 18:00, the Incoming shift lead has not acknowledged the handoff.\",\"policy\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["context"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee."}, {"path": ["context"], "text": "At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_right": "At 18:00, Jordan Lee is serving in the Maintenance coordinator role, and the shift record assigns each person exactly one coordinator role.", "right": "At 18:00, Jordan Lee is serving in the Operations coordinator role, and the shift record assigns each person exactly one coordinator role."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-015", "id": "fast-41-diverse-036-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The unresolved handoff register contains only the Conveyor 4 restart check and the Filler 2 pressure alarm. The alarm is the sole critical equipment alarm; the Conveyor item is routine. Both entries have current status artifacts and stated next actions. The alarm has an attached diagnostic result. The Conveyor owner field names the Incoming shift lead. At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee. At 18:00, Jordan Lee is serving in the Maintenance coordinator role, and the shift record assigns each person exactly one coordinator role. By 18:00, the Incoming shift lead has not acknowledged the handoff.", "policy": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is unchanged in both contexts, and the original questions object is preserved verbatim. The focus evidence is retained as two complete factual sentences: \"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.\" and \"At 18:00, Jordan Lee is serving as the Operations coordinator.\" Entity, path, and time bindings remain unchanged. The counterfactual coherently changes only Jordan Lee's coordinator status without creating contradictory duplicate assertions. Neither context embeds a score, answer code, rationale, proposition ID, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"The unresolved handoff list contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"The Filler 2 pressure alarm is the only critical equipment alarm on that list.\",\"The Conveyor 4 restart check is a routine process follow-up.\",\"Both unresolved items have current status artifacts and stated next actions, and the pressure alarm has an attached diagnostic result.\",\"The Conveyor 4 restart check's owner field identifies Priya Shah, who is serving as the Incoming shift lead.\",\"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.\",\"At 18:00, Jordan Lee is serving as the Operations coordinator.\",\"The Incoming shift lead has not yet acknowledged the handoff.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "5"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee."}, {"path": ["evidence", "6"], "text": "At 18:00, Jordan Lee is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_right": "At 18:00, Jordan Lee is not serving as the Operations coordinator.", "right": "At 18:00, Jordan Lee is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-016", "id": "fast-41-diverse-036-016-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["The unresolved handoff list contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.", "The Filler 2 pressure alarm is the only critical equipment alarm on that list.", "The Conveyor 4 restart check is a routine process follow-up.", "Both unresolved items have current status artifacts and stated next actions, and the pressure alarm has an attached diagnostic result.", "The Conveyor 4 restart check's owner field identifies Priya Shah, who is serving as the Incoming shift lead.", "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "At 18:00, Jordan Lee is serving as the Operations coordinator.", "The Incoming shift lead has not yet acknowledged the handoff."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is unchanged in both contexts, and the original questions object is preserved verbatim. The focus evidence is retained as two complete factual sentences: \"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.\" and \"At 18:00, Jordan Lee is serving as the Operations coordinator.\" Entity, path, and time bindings remain unchanged. The counterfactual coherently changes only Jordan Lee's coordinator status without creating contradictory duplicate assertions. Neither context embeds a score, answer code, rationale, proposition ID, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"The unresolved handoff list contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"The Filler 2 pressure alarm is the only critical equipment alarm on that list.\",\"The Conveyor 4 restart check is a routine process follow-up.\",\"Both unresolved items have current status artifacts and stated next actions, and the pressure alarm has an attached diagnostic result.\",\"The Conveyor 4 restart check's owner field identifies Priya Shah, who is serving as the Incoming shift lead.\",\"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.\",\"At 18:00, Jordan Lee is serving as the Operations coordinator.\",\"The Incoming shift lead has not yet acknowledged the handoff.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "5"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee."}, {"path": ["evidence", "6"], "text": "At 18:00, Jordan Lee is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "negative_right": "At 18:00, Jordan Lee is not serving as the Operations coordinator.", "right": "At 18:00, Jordan Lee is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-016", "id": "fast-41-diverse-036-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["The unresolved handoff list contains exactly the Conveyor 4 restart check and the Filler 2 pressure alarm.", "The Filler 2 pressure alarm is the only critical equipment alarm on that list.", "The Conveyor 4 restart check is a routine process follow-up.", "Both unresolved items have current status artifacts and stated next actions, and the pressure alarm has an attached diagnostic result.", "The Conveyor 4 restart check's owner field identifies Priya Shah, who is serving as the Incoming shift lead.", "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Lee.", "At 18:00, Jordan Lee is not serving as the Operations coordinator.", "The Incoming shift lead has not yet acknowledged the handoff."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, unchanged questions, 18:00 timing, handoff scope, and readiness request. Each evidence list contains two complete factual sentences: base evidence quotes “At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.” and “At 18:00, Jordan's sole assigned role is Operations coordinator.” Counterfactual evidence quotes “At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.” and “At 18:00, Jordan's sole assigned role is Maintenance coordinator.” The counterfactual coherently changes Jordan's role and creates a wrong-owner observation without contradictory duplicate measurements. Neither context includes a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The handoff register contains exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The restart check is routine, has a current status artifact and next action, and its owner field names the person serving as Incoming shift lead. The pressure alarm is the only critical equipment alarm, has a current status artifact, a stated next action, and an attached diagnostic result. The incoming shift lead has not acknowledged the handoff by 18:00.\",\"evidence\":[\"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.\",\"At 18:00, Jordan's sole assigned role is Operations coordinator.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan."}, {"path": ["evidence", "1"], "text": "At 18:00, Jordan's sole assigned role is Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.", "negative_right": "At 18:00, Jordan's sole assigned role is Maintenance coordinator.", "right": "At 18:00, Jordan's sole assigned role is Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-020", "id": "fast-41-diverse-036-020-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The handoff register contains exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The restart check is routine, has a current status artifact and next action, and its owner field names the person serving as Incoming shift lead. The pressure alarm is the only critical equipment alarm, has a current status artifact, a stated next action, and an attached diagnostic result. The incoming shift lead has not acknowledged the handoff by 18:00.", "evidence": ["At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.", "At 18:00, Jordan's sole assigned role is Operations coordinator."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, unchanged questions, 18:00 timing, handoff scope, and readiness request. Each evidence list contains two complete factual sentences: base evidence quotes “At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.” and “At 18:00, Jordan's sole assigned role is Operations coordinator.” Counterfactual evidence quotes “At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.” and “At 18:00, Jordan's sole assigned role is Maintenance coordinator.” The counterfactual coherently changes Jordan's role and creates a wrong-owner observation without contradictory duplicate measurements. Neither context includes a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The handoff register contains exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The restart check is routine, has a current status artifact and next action, and its owner field names the person serving as Incoming shift lead. The pressure alarm is the only critical equipment alarm, has a current status artifact, a stated next action, and an attached diagnostic result. The incoming shift lead has not acknowledged the handoff by 18:00.\",\"evidence\":[\"At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.\",\"At 18:00, Jordan's sole assigned role is Operations coordinator.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan."}, {"path": ["evidence", "1"], "text": "At 18:00, Jordan's sole assigned role is Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.", "negative_right": "At 18:00, Jordan's sole assigned role is Maintenance coordinator.", "right": "At 18:00, Jordan's sole assigned role is Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-020", "id": "fast-41-diverse-036-020-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The handoff register contains exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The restart check is routine, has a current status artifact and next action, and its owner field names the person serving as Incoming shift lead. The pressure alarm is the only critical equipment alarm, has a current status artifact, a stated next action, and an attached diagnostic result. The incoming shift lead has not acknowledged the handoff by 18:00.", "evidence": ["At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan.", "At 18:00, Jordan's sole assigned role is Maintenance coordinator."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, policy, and bindings; the evidence consists of the complete factual sentences “At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee.” and “At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator.”; the counterfactual coherently changes only the Filler 2 owner to Taylor Chen, and neither context includes answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"case_note\":\"The handoff register lists exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. A current artifact and stated next action are recorded for each. The Conveyor 4 item is marked routine, and its owner field names the person serving as Incoming shift lead. The Filler 2 item is marked the sole critical equipment alarm and has an attached diagnostic result. At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee. At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator. The incoming shift lead has not acknowledged the handoff by 18:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["case_note"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee."}, {"path": ["case_note"], "text": "At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Taylor Chen.", "negative_right": "At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator.", "right": "At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-021", "id": "fast-41-diverse-036-021-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"case_note": "The handoff register lists exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. A current artifact and stated next action are recorded for each. The Conveyor 4 item is marked routine, and its owner field names the person serving as Incoming shift lead. The Filler 2 item is marked the sole critical equipment alarm and has an attached diagnostic result. At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee. At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator. The incoming shift lead has not acknowledged the handoff by 18:00.", "context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, policy, and bindings; the evidence consists of the complete factual sentences “At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee.” and “At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator.”; the counterfactual coherently changes only the Filler 2 owner to Taylor Chen, and neither context includes answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"case_note\":\"The handoff register lists exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. A current artifact and stated next action are recorded for each. The Conveyor 4 item is marked routine, and its owner field names the person serving as Incoming shift lead. The Filler 2 item is marked the sole critical equipment alarm and has an attached diagnostic result. At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee. At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator. The incoming shift lead has not acknowledged the handoff by 18:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["case_note"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee."}, {"path": ["case_note"], "text": "At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Jordan Lee.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Taylor Chen.", "negative_right": "At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator.", "right": "At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-021", "id": "fast-41-diverse-036-021-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"case_note": "The handoff register lists exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. A current artifact and stated next action are recorded for each. The Conveyor 4 item is marked routine, and its owner field names the person serving as Incoming shift lead. The Filler 2 item is marked the sole critical equipment alarm and has an attached diagnostic result. At 18:00, the Filler 2 pressure alarm's owner field identifies exactly one person, Taylor Chen. At 18:00, Jordan Lee is serving as the Operations coordinator, and Jordan Lee and Taylor Chen are different people; Taylor Chen is not serving as the Operations coordinator. The incoming shift lead has not acknowledged the handoff by 18:00.", "context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions are verbatim and both contexts retain every stated policy. Question bindings are preserved for the packaging handoff, 18:00 time, listed items, and observation paths. Evidence consists of two complete factual sentences: \"At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee.\" and \"At 18:00, Jordan Lee is serving as the Operations coordinator.\" The counterfactual coherently changes Jordan Lee’s role without adding contradictory duplicate assertions. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"observations\":\"The handoff register lists exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The pressure alarm is the only critical equipment alarm; the Conveyor 4 item is a routine process follow-up. Both records contain current status artifacts and stated next actions. The Conveyor 4 owner field names the person serving as Incoming shift lead. At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee. At 18:00, Jordan Lee is serving as the Operations coordinator. A diagnostic result is attached to the pressure-alarm record. The incoming shift lead has not acknowledged the handoff by 18:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["observations"], "text": "At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee."}, {"path": ["observations"], "text": "At 18:00, Jordan Lee is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee.", "negative_left": "At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee.", "negative_right": "At 18:00, Jordan Lee is serving as the Packaging supervisor rather than the Operations coordinator.", "right": "At 18:00, Jordan Lee is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-028", "id": "fast-41-diverse-036-028-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "observations": "The handoff register lists exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The pressure alarm is the only critical equipment alarm; the Conveyor 4 item is a routine process follow-up. Both records contain current status artifacts and stated next actions. The Conveyor 4 owner field names the person serving as Incoming shift lead. At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee. At 18:00, Jordan Lee is serving as the Operations coordinator. A diagnostic result is attached to the pressure-alarm record. The incoming shift lead has not acknowledged the handoff by 18:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions are verbatim and both contexts retain every stated policy. Question bindings are preserved for the packaging handoff, 18:00 time, listed items, and observation paths. Evidence consists of two complete factual sentences: \"At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee.\" and \"At 18:00, Jordan Lee is serving as the Operations coordinator.\" The counterfactual coherently changes Jordan Lee’s role without adding contradictory duplicate assertions. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"observations\":\"The handoff register lists exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The pressure alarm is the only critical equipment alarm; the Conveyor 4 item is a routine process follow-up. Both records contain current status artifacts and stated next actions. The Conveyor 4 owner field names the person serving as Incoming shift lead. At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee. At 18:00, Jordan Lee is serving as the Operations coordinator. A diagnostic result is attached to the pressure-alarm record. The incoming shift lead has not acknowledged the handoff by 18:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["observations"], "text": "At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee."}, {"path": ["observations"], "text": "At 18:00, Jordan Lee is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee.", "negative_left": "At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee.", "negative_right": "At 18:00, Jordan Lee is serving as the Packaging supervisor rather than the Operations coordinator.", "right": "At 18:00, Jordan Lee is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-028", "id": "fast-41-diverse-036-028-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "observations": "The handoff register lists exactly two unresolved items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The pressure alarm is the only critical equipment alarm; the Conveyor 4 item is a routine process follow-up. Both records contain current status artifacts and stated next actions. The Conveyor 4 owner field names the person serving as Incoming shift lead. At 18:00, the owner field on the Filler 2 pressure alarm identifies Jordan Lee. At 18:00, Jordan Lee is serving as the Packaging supervisor rather than the Operations coordinator. A diagnostic result is attached to the pressure-alarm record. The incoming shift lead has not acknowledged the handoff by 18:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy, and the original questions preserve the rubric and instructions. The 18:00 handoff scope and record bindings remain unchanged. The evidence consists of two complete factual sentences: \"At 18:00, handoff record H-27 is the Filler 2 pressure alarm record.\" and \"At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner.\" The counterfactual coherently changes H-27's owner while retaining the diagnostic result and other facts. Neither context states a gold score, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The complete unresolved handoff register contains exactly two items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The pressure alarm is the only critical equipment alarm; the Conveyor 4 check is a routine process follow-up. Both items have current status artifacts and stated next actions. At 18:00, handoff record H-27 is the Filler 2 pressure alarm record. At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner. A diagnostic result is attached to H-27. The Conveyor 4 record identifies the person serving as the Incoming shift lead in its owner field. The incoming shift lead has not acknowledged the handoff by 18:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["context"], "text": "At 18:00, handoff record H-27 is the Filler 2 pressure alarm record."}, {"path": ["context"], "text": "At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, handoff record H-27 is the Filler 2 pressure alarm record.", "negative_left": "At 18:00, handoff record H-27 is the Filler 2 pressure alarm record.", "negative_right": "At 18:00, record H-27 lists the person serving as the Incoming shift lead, rather than the Operations coordinator, as its owner.", "right": "At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-029", "id": "fast-41-diverse-036-029-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The complete unresolved handoff register contains exactly two items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The pressure alarm is the only critical equipment alarm; the Conveyor 4 check is a routine process follow-up. Both items have current status artifacts and stated next actions. At 18:00, handoff record H-27 is the Filler 2 pressure alarm record. At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner. A diagnostic result is attached to H-27. The Conveyor 4 record identifies the person serving as the Incoming shift lead in its owner field. The incoming shift lead has not acknowledged the handoff by 18:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy, and the original questions preserve the rubric and instructions. The 18:00 handoff scope and record bindings remain unchanged. The evidence consists of two complete factual sentences: \"At 18:00, handoff record H-27 is the Filler 2 pressure alarm record.\" and \"At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner.\" The counterfactual coherently changes H-27's owner while retaining the diagnostic result and other facts. Neither context states a gold score, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The complete unresolved handoff register contains exactly two items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The pressure alarm is the only critical equipment alarm; the Conveyor 4 check is a routine process follow-up. Both items have current status artifacts and stated next actions. At 18:00, handoff record H-27 is the Filler 2 pressure alarm record. At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner. A diagnostic result is attached to H-27. The Conveyor 4 record identifies the person serving as the Incoming shift lead in its owner field. The incoming shift lead has not acknowledged the handoff by 18:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["context"], "text": "At 18:00, handoff record H-27 is the Filler 2 pressure alarm record."}, {"path": ["context"], "text": "At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, handoff record H-27 is the Filler 2 pressure alarm record.", "negative_left": "At 18:00, handoff record H-27 is the Filler 2 pressure alarm record.", "negative_right": "At 18:00, record H-27 lists the person serving as the Incoming shift lead, rather than the Operations coordinator, as its owner.", "right": "At 18:00, record H-27 lists the person serving as the Operations coordinator as its owner."}, "verifier_independent_model": false}, "family": "fast-41-diverse-036-029", "id": "fast-41-diverse-036-029-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready. The complete unresolved handoff register contains exactly two items: the Conveyor 4 restart check and the Filler 2 pressure alarm. The pressure alarm is the only critical equipment alarm; the Conveyor 4 check is a routine process follow-up. Both items have current status artifacts and stated next actions. At 18:00, handoff record H-27 is the Filler 2 pressure alarm record. At 18:00, record H-27 lists the person serving as the Incoming shift lead, rather than the Operations coordinator, as its owner. A diagnostic result is attached to H-27. The Conveyor 4 record identifies the person serving as the Incoming shift lead in its owner field. The incoming shift lead has not acknowledged the handoff by 18:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts repeat it without alteration. Both contexts retain Mira Chen, the three-session loop, October 8, the final-session timing, and approval-before-confirmation bindings. Each context supplies exactly two complete factual evidence sentences. Base evidence: \"The sole written approval for the final session's after-hours schedule was issued by Jordan Lee.\" \"Jordan Lee holds a hiring-manager role rather than a coordinator role.\" Counterfactual evidence: \"The sole written approval for the final session's after-hours schedule was issued by Jordan Lee.\" \"Jordan Lee holds a coordinator role rather than a hiring-manager role.\" The counterfactual changes only Jordan Lee's role and remains internally consistent with the unchanged facts and policy. Neither context states a classification, answer code, rationale label, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Case note: Mira Chen confirmed the proposed three-session video loop on October 8. Every interviewer scheduled for the loop also confirmed, and logistics for every session are complete. The final session ends after 6:00 p.m. in the interviewer’s local time, and that timing is documented in the loop’s calendar record. The record contains exactly one written approval for the final session’s after-hours schedule. The sole written approval for the final session's after-hours schedule was issued by Jordan Lee. Jordan Lee holds a hiring-manager role rather than a coordinator role. That approval was issued before Mira confirmed the loop. The documented timing could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was issued by Jordan Lee."}, {"path": [], "text": "Jordan Lee holds a hiring-manager role rather than a coordinator role."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was issued by Jordan Lee.", "negative_left": "The sole written approval for the final session's after-hours schedule was issued by Jordan Lee.", "negative_right": "Jordan Lee holds a coordinator role rather than a hiring-manager role.", "right": "Jordan Lee holds a hiring-manager role rather than a coordinator role."}, "verifier_independent_model": false}, "family": "fast-41-diverse-037-006", "id": "fast-41-diverse-037-006-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Case note: Mira Chen confirmed the proposed three-session video loop on October 8. Every interviewer scheduled for the loop also confirmed, and logistics for every session are complete. The final session ends after 6:00 p.m. in the interviewer’s local time, and that timing is documented in the loop’s calendar record. The record contains exactly one written approval for the final session’s after-hours schedule. The sole written approval for the final session's after-hours schedule was issued by Jordan Lee. Jordan Lee holds a hiring-manager role rather than a coordinator role. That approval was issued before Mira confirmed the loop. The documented timing could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts repeat it without alteration. Both contexts retain Mira Chen, the three-session loop, October 8, the final-session timing, and approval-before-confirmation bindings. Each context supplies exactly two complete factual evidence sentences. Base evidence: \"The sole written approval for the final session's after-hours schedule was issued by Jordan Lee.\" \"Jordan Lee holds a hiring-manager role rather than a coordinator role.\" Counterfactual evidence: \"The sole written approval for the final session's after-hours schedule was issued by Jordan Lee.\" \"Jordan Lee holds a coordinator role rather than a hiring-manager role.\" The counterfactual changes only Jordan Lee's role and remains internally consistent with the unchanged facts and policy. Neither context states a classification, answer code, rationale label, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Case note: Mira Chen confirmed the proposed three-session video loop on October 8. Every interviewer scheduled for the loop also confirmed, and logistics for every session are complete. The final session ends after 6:00 p.m. in the interviewer’s local time, and that timing is documented in the loop’s calendar record. The record contains exactly one written approval for the final session’s after-hours schedule. The sole written approval for the final session's after-hours schedule was issued by Jordan Lee. Jordan Lee holds a hiring-manager role rather than a coordinator role. That approval was issued before Mira confirmed the loop. The documented timing could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was issued by Jordan Lee."}, {"path": [], "text": "Jordan Lee holds a hiring-manager role rather than a coordinator role."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was issued by Jordan Lee.", "negative_left": "The sole written approval for the final session's after-hours schedule was issued by Jordan Lee.", "negative_right": "Jordan Lee holds a coordinator role rather than a hiring-manager role.", "right": "Jordan Lee holds a hiring-manager role rather than a coordinator role."}, "verifier_independent_model": false}, "family": "fast-41-diverse-037-006", "id": "fast-41-diverse-037-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Case note: Mira Chen confirmed the proposed three-session video loop on October 8. Every interviewer scheduled for the loop also confirmed, and logistics for every session are complete. The final session ends after 6:00 p.m. in the interviewer’s local time, and that timing is documented in the loop’s calendar record. The record contains exactly one written approval for the final session’s after-hours schedule. The sole written approval for the final session's after-hours schedule was issued by Jordan Lee. Jordan Lee holds a coordinator role rather than a hiring-manager role. That approval was issued before Mira confirmed the loop. The documented timing could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions and the governing policy about readiness, scope, and prior written hiring-manager approval. The Mira Chen, October 8, three-session loop, final-session, and approval bindings remain consistent. The evidence consists of two complete factual sentences: \"The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz.\" and \"Dana Ortiz holds a hiring-manager role on October 8.\" The counterfactual coherently changes Dana Ortiz's role while retaining the after-hours session and prior approval facts. Neither context embeds a gold answer, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"The coordinator's October 8 case note concerns Mira Chen's proposed three-session video loop. Mira Chen confirmed the proposal, and every interviewer scheduled for it also confirmed. Logistics for every session are complete, including the calendar entries and video links. The final session ends after 6:00 p.m. in the interviewer's local time, and that timing appears in the calendar record. The documented timing could disrupt the loop. The record contains exactly one written approval for the final session's after-hours schedule. The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz. Dana Ortiz holds a hiring-manager role on October 8. That sole written approval was issued before confirmation of the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz."}, {"path": [], "text": "Dana Ortiz holds a hiring-manager role on October 8."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz.", "negative_left": "The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz.", "negative_right": "Dana Ortiz does not hold a hiring-manager role on October 8.", "right": "Dana Ortiz holds a hiring-manager role on October 8."}, "verifier_independent_model": false}, "family": "fast-41-diverse-037-010", "id": "fast-41-diverse-037-010-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "The coordinator's October 8 case note concerns Mira Chen's proposed three-session video loop. Mira Chen confirmed the proposal, and every interviewer scheduled for it also confirmed. Logistics for every session are complete, including the calendar entries and video links. The final session ends after 6:00 p.m. in the interviewer's local time, and that timing appears in the calendar record. The documented timing could disrupt the loop. The record contains exactly one written approval for the final session's after-hours schedule. The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz. Dana Ortiz holds a hiring-manager role on October 8. That sole written approval was issued before confirmation of the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions and the governing policy about readiness, scope, and prior written hiring-manager approval. The Mira Chen, October 8, three-session loop, final-session, and approval bindings remain consistent. The evidence consists of two complete factual sentences: \"The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz.\" and \"Dana Ortiz holds a hiring-manager role on October 8.\" The counterfactual coherently changes Dana Ortiz's role while retaining the after-hours session and prior approval facts. Neither context embeds a gold answer, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"The coordinator's October 8 case note concerns Mira Chen's proposed three-session video loop. Mira Chen confirmed the proposal, and every interviewer scheduled for it also confirmed. Logistics for every session are complete, including the calendar entries and video links. The final session ends after 6:00 p.m. in the interviewer's local time, and that timing appears in the calendar record. The documented timing could disrupt the loop. The record contains exactly one written approval for the final session's after-hours schedule. The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz. Dana Ortiz holds a hiring-manager role on October 8. That sole written approval was issued before confirmation of the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz."}, {"path": [], "text": "Dana Ortiz holds a hiring-manager role on October 8."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz.", "negative_left": "The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz.", "negative_right": "Dana Ortiz does not hold a hiring-manager role on October 8.", "right": "Dana Ortiz holds a hiring-manager role on October 8."}, "verifier_independent_model": false}, "family": "fast-41-diverse-037-010", "id": "fast-41-diverse-037-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "The coordinator's October 8 case note concerns Mira Chen's proposed three-session video loop. Mira Chen confirmed the proposal, and every interviewer scheduled for it also confirmed. Logistics for every session are complete, including the calendar entries and video links. The final session ends after 6:00 p.m. in the interviewer's local time, and that timing appears in the calendar record. The documented timing could disrupt the loop. The record contains exactly one written approval for the final session's after-hours schedule. The grantor identified in the sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop is Dana Ortiz. Dana Ortiz does not hold a hiring-manager role on October 8. That sole written approval was issued before confirmation of the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions and governing policy. The bindings for Mira Chen, the three-session loop, October 8, and the after-hours final session remain intact. The evidence consists of two complete factual sentences: “The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor.” and “Dana Ortiz holds a hiring-manager role.” The counterfactual coherently changes Dana Ortiz’s role without duplicating or contradicting measurements. Neither context contains a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"A coordinator is reviewing Mira Chen’s proposed three-session video loop for October 8. Mira confirmed the loop, every scheduled interviewer confirmed, and logistics for every session are complete. The final session ends after 6:00 p.m. in the interviewer’s local time, and that timing appears in the loop’s calendar record as a possible disruption. The record contains exactly one written approval for the final session’s after-hours schedule, issued before Mira confirmed the loop. The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor. Dana Ortiz holds a hiring-manager role. The confirmation log, logistics checklist, calendar entry, and approval document are retained together.\\n\\nPolicy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete.\\nThis policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor."}, {"path": [], "text": "Dana Ortiz holds a hiring-manager role."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor.", "negative_left": "The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor.", "negative_right": "Dana Ortiz holds a senior-recruiter role, not a hiring-manager role.", "right": "Dana Ortiz holds a hiring-manager role."}, "verifier_independent_model": false}, "family": "fast-41-diverse-037-011", "id": "fast-41-diverse-037-011-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "A coordinator is reviewing Mira Chen’s proposed three-session video loop for October 8. Mira confirmed the loop, every scheduled interviewer confirmed, and logistics for every session are complete. The final session ends after 6:00 p.m. in the interviewer’s local time, and that timing appears in the loop’s calendar record as a possible disruption. The record contains exactly one written approval for the final session’s after-hours schedule, issued before Mira confirmed the loop. The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor. Dana Ortiz holds a hiring-manager role. The confirmation log, logistics checklist, calendar entry, and approval document are retained together.\n\nPolicy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete.\nThis policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions and governing policy. The bindings for Mira Chen, the three-session loop, October 8, and the after-hours final session remain intact. The evidence consists of two complete factual sentences: “The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor.” and “Dana Ortiz holds a hiring-manager role.” The counterfactual coherently changes Dana Ortiz’s role without duplicating or contradicting measurements. Neither context contains a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"A coordinator is reviewing Mira Chen’s proposed three-session video loop for October 8. Mira confirmed the loop, every scheduled interviewer confirmed, and logistics for every session are complete. The final session ends after 6:00 p.m. in the interviewer’s local time, and that timing appears in the loop’s calendar record as a possible disruption. The record contains exactly one written approval for the final session’s after-hours schedule, issued before Mira confirmed the loop. The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor. Dana Ortiz holds a hiring-manager role. The confirmation log, logistics checklist, calendar entry, and approval document are retained together.\\n\\nPolicy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete.\\nThis policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor."}, {"path": [], "text": "Dana Ortiz holds a hiring-manager role."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor.", "negative_left": "The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor.", "negative_right": "Dana Ortiz holds a senior-recruiter role, not a hiring-manager role.", "right": "Dana Ortiz holds a hiring-manager role."}, "verifier_independent_model": false}, "family": "fast-41-diverse-037-011", "id": "fast-41-diverse-037-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "A coordinator is reviewing Mira Chen’s proposed three-session video loop for October 8. Mira confirmed the loop, every scheduled interviewer confirmed, and logistics for every session are complete. The final session ends after 6:00 p.m. in the interviewer’s local time, and that timing appears in the loop’s calendar record as a possible disruption. The record contains exactly one written approval for the final session’s after-hours schedule, issued before Mira confirmed the loop. The sole written approval for the final session's after-hours schedule identifies Dana Ortiz as its grantor. Dana Ortiz holds a senior-recruiter role, not a hiring-manager role. The confirmation log, logistics checklist, calendar entry, and approval document are retained together.\n\nPolicy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete.\nThis policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the verbatim questions, all policy criteria, and the October 13 Nadia Cole/Priya Shah/current-loop bindings. The two focus evidence spans are complete factual sentences: “The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response.” and “The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17.” The counterfactual coherently changes only Priya Shah’s status from C17 to C18, with no contradictory duplicate assertion, and neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"Nadia Cole has confirmed the October 13, 2026 interview loop, and the hiring manager has confirmed the same current loop. Priya Shah is listed as a required interviewer, while every other required interviewer has confirmed this loop. Priya’s scheduled block is 2:00–3:00 p.m., and the calendar record shows no conflict then.\"},{\"speaker\":\"Scheduling auditor\",\"text\":\"The audit log identifies the relevant response as authenticated and timestamps it at 09:12 UTC.\"},{\"speaker\":\"Scheduling system\",\"text\":\"The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response.\"},{\"speaker\":\"Scheduling system\",\"text\":\"The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"The record concerns current-loop required-participant coverage; optional attendees do not affect that coverage.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["2", "text"], "text": "The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response."}, {"path": ["3", "text"], "text": "The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response.", "negative_left": "The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response.", "negative_right": "The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C18.", "right": "The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-039-006", "id": "fast-41-diverse-039-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Recruiting coordinator", "text": "Nadia Cole has confirmed the October 13, 2026 interview loop, and the hiring manager has confirmed the same current loop. Priya Shah is listed as a required interviewer, while every other required interviewer has confirmed this loop. Priya’s scheduled block is 2:00–3:00 p.m., and the calendar record shows no conflict then."}, {"speaker": "Scheduling auditor", "text": "The audit log identifies the relevant response as authenticated and timestamps it at 09:12 UTC."}, {"speaker": "Scheduling system", "text": "The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response."}, {"speaker": "Scheduling system", "text": "The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17."}, {"speaker": "Recruiting coordinator", "text": "The record concerns current-loop required-participant coverage; optional attendees do not affect that coverage."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the verbatim questions, all policy criteria, and the October 13 Nadia Cole/Priya Shah/current-loop bindings. The two focus evidence spans are complete factual sentences: “The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response.” and “The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17.” The counterfactual coherently changes only Priya Shah’s status from C17 to C18, with no contradictory duplicate assertion, and neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"Nadia Cole has confirmed the October 13, 2026 interview loop, and the hiring manager has confirmed the same current loop. Priya Shah is listed as a required interviewer, while every other required interviewer has confirmed this loop. Priya’s scheduled block is 2:00–3:00 p.m., and the calendar record shows no conflict then.\"},{\"speaker\":\"Scheduling auditor\",\"text\":\"The audit log identifies the relevant response as authenticated and timestamps it at 09:12 UTC.\"},{\"speaker\":\"Scheduling system\",\"text\":\"The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response.\"},{\"speaker\":\"Scheduling system\",\"text\":\"The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"The record concerns current-loop required-participant coverage; optional attendees do not affect that coverage.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["2", "text"], "text": "The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response."}, {"path": ["3", "text"], "text": "The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response.", "negative_left": "The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response.", "negative_right": "The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C18.", "right": "The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-039-006", "id": "fast-41-diverse-039-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Recruiting coordinator", "text": "Nadia Cole has confirmed the October 13, 2026 interview loop, and the hiring manager has confirmed the same current loop. Priya Shah is listed as a required interviewer, while every other required interviewer has confirmed this loop. Priya’s scheduled block is 2:00–3:00 p.m., and the calendar record shows no conflict then."}, {"speaker": "Scheduling auditor", "text": "The audit log identifies the relevant response as authenticated and timestamps it at 09:12 UTC."}, {"speaker": "Scheduling system", "text": "The scheduling portal for Nadia Cole's October 13, 2026 interview loop uses status code C17 exclusively for a participant's confirmed response and status code C18 exclusively for a participant's unconfirmed response."}, {"speaker": "Scheduling system", "text": "The authenticated scheduling-portal response submitted by Priya Shah for Nadia Cole's October 13, 2026 interview loop at 09:12 UTC carries status code C18."}, {"speaker": "Recruiting coordinator", "text": "The record concerns current-loop required-participant coverage; optional attendees do not affect that coverage."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, and both contexts retain Nadia Cole, the October 13, 2026 loop, and Priya Shah bindings. The two complete factual evidence quotes are: “For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer.” and “The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4.” The counterfactual changes only Priya Shah’s record from Q4 to Q9, which coherently indicates nonconfirmation despite no calendar conflict. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"Nadia Cole is the candidate for the October 13, 2026 interview loop, and the candidate has confirmed that current loop. The hiring manager has confirmed it as well.\"},{\"speaker\":\"Scheduling record\",\"text\":\"Priya Shah is a required interviewer for Nadia's October 13 loop. Every required interviewer other than Priya Shah has confirmed the current loop.\"},{\"speaker\":\"Calendar administrator\",\"text\":\"Priya Shah has no calendar conflict during her scheduled October 13 interview block.\"},{\"speaker\":\"Records specialist\",\"text\":\"For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer.\"},{\"speaker\":\"Records specialist\",\"text\":\"The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["3", "text"], "text": "For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer."}, {"path": ["4", "text"], "text": "The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer.", "negative_left": "For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer.", "negative_right": "The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q9.", "right": "The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-039-007", "id": "fast-41-diverse-039-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Recruiting coordinator", "text": "Nadia Cole is the candidate for the October 13, 2026 interview loop, and the candidate has confirmed that current loop. The hiring manager has confirmed it as well."}, {"speaker": "Scheduling record", "text": "Priya Shah is a required interviewer for Nadia's October 13 loop. Every required interviewer other than Priya Shah has confirmed the current loop."}, {"speaker": "Calendar administrator", "text": "Priya Shah has no calendar conflict during her scheduled October 13 interview block."}, {"speaker": "Records specialist", "text": "For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer."}, {"speaker": "Records specialist", "text": "The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, and both contexts retain Nadia Cole, the October 13, 2026 loop, and Priya Shah bindings. The two complete factual evidence quotes are: “For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer.” and “The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4.” The counterfactual changes only Priya Shah’s record from Q4 to Q9, which coherently indicates nonconfirmation despite no calendar conflict. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"Nadia Cole is the candidate for the October 13, 2026 interview loop, and the candidate has confirmed that current loop. The hiring manager has confirmed it as well.\"},{\"speaker\":\"Scheduling record\",\"text\":\"Priya Shah is a required interviewer for Nadia's October 13 loop. Every required interviewer other than Priya Shah has confirmed the current loop.\"},{\"speaker\":\"Calendar administrator\",\"text\":\"Priya Shah has no calendar conflict during her scheduled October 13 interview block.\"},{\"speaker\":\"Records specialist\",\"text\":\"For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer.\"},{\"speaker\":\"Records specialist\",\"text\":\"The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["3", "text"], "text": "For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer."}, {"path": ["4", "text"], "text": "The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer.", "negative_left": "For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer.", "negative_right": "The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q9.", "right": "The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-039-007", "id": "fast-41-diverse-039-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Recruiting coordinator", "text": "Nadia Cole is the candidate for the October 13, 2026 interview loop, and the candidate has confirmed that current loop. The hiring manager has confirmed it as well."}, {"speaker": "Scheduling record", "text": "Priya Shah is a required interviewer for Nadia's October 13 loop. Every required interviewer other than Priya Shah has confirmed the current loop."}, {"speaker": "Calendar administrator", "text": "Priya Shah has no calendar conflict during her scheduled October 13 interview block."}, {"speaker": "Records specialist", "text": "For Nadia Cole's October 13, 2026 interview loop, record codes Q4 and Q9 mean, respectively, confirmed by the named interviewer and not confirmed by the named interviewer."}, {"speaker": "Records specialist", "text": "The October 13, 2026 interview-loop record for Nadia Cole assigns Priya Shah record code Q9."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy and criteria. Both contexts retain Nadia Cole, October 13, 2026, the interview loop, and the required-interviewer scope. The two evidence spans are complete factual sentences: “Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster.” “Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop.” The counterfactual changes the roster-confirmation observation without creating a contradictory duplicate assertion. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"Nadia Cole and the hiring manager have each confirmed the October 13, 2026 interview loop. Priya Shah is a required interviewer, and every required interviewer other than Priya Shah has confirmed the loop. Priya has no calendar conflict during her scheduled October 13 interview block.\"},{\"speaker\":\"Scheduling auditor\",\"text\":\"Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster.\"},{\"speaker\":\"Roster record\",\"text\":\"Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["1", "text"], "text": "Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster."}, {"path": ["2", "text"], "text": "Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster.", "negative_left": "Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster.", "negative_right": "No member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop.", "right": "Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop."}, "verifier_independent_model": false}, "family": "fast-41-diverse-039-013", "id": "fast-41-diverse-039-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Recruiting coordinator", "text": "Nadia Cole and the hiring manager have each confirmed the October 13, 2026 interview loop. Priya Shah is a required interviewer, and every required interviewer other than Priya Shah has confirmed the loop. Priya has no calendar conflict during her scheduled October 13 interview block."}, {"speaker": "Scheduling auditor", "text": "Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster."}, {"speaker": "Roster record", "text": "Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy and criteria. Both contexts retain Nadia Cole, October 13, 2026, the interview loop, and the required-interviewer scope. The two evidence spans are complete factual sentences: “Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster.” “Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop.” The counterfactual changes the roster-confirmation observation without creating a contradictory duplicate assertion. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A4 is an allowed universal fact over the explicit set of required interviewers other than Priya. A6 is factual rather than a policy conclusion. The base and counter assignments are realizable with only Priya's confirmation status changing. Empty policy_evidence is correct because all governing criteria and exceptions are contained in the retained questions object, while the state supplies only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes confirmation of the current October 13 loop by the candidate, hiring manager, Priya Shah, and every other required interviewer. This is sufficient for the true outcome. Priya's lack of conflict is redundant but does not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the candidate, hiring manager, and all required interviewers except Priya have confirmed; Priya is required, has not confirmed, and has no calendar conflict. Thus Priya is the sole gap, so the stated policy requires the false outcome: medium risk and routing to Priya.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Candidate Nadia Cole has confirmed the October 13, 2026 interview loop."}, {"id": "A2", "statement": "The hiring manager for Nadia Cole's October 13, 2026 interview loop has confirmed that loop."}, {"id": "A3", "statement": "Priya Shah is a required interviewer for Nadia Cole's October 13, 2026 interview loop."}, {"id": "A4", "statement": "Every required interviewer other than Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}, {"id": "A5", "statement": "Priya Shah has no calendar conflict during her scheduled October 13, 2026 interview block."}, {"id": "A6", "statement": "Priya Shah has confirmed Nadia Cole's October 13, 2026 interview loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"Nadia Cole and the hiring manager have each confirmed the October 13, 2026 interview loop. Priya Shah is a required interviewer, and every required interviewer other than Priya Shah has confirmed the loop. Priya has no calendar conflict during her scheduled October 13 interview block.\"},{\"speaker\":\"Scheduling auditor\",\"text\":\"Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster.\"},{\"speaker\":\"Roster record\",\"text\":\"Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}], "focus_atom": "A6", "focus_evidence": [{"path": ["1", "text"], "text": "Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster."}, {"path": ["2", "text"], "text": "Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop."}], "policy_evidence": [], "rules": [{"justification": "Nadia Cole, the hiring manager, Priya Shah, and every other required interviewer have confirmed the current October 13, 2026 loop, so no participant follow-up is needed.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}]}, {"justification": "The sole confirmation gap is required interviewer Priya Shah, and she has no calendar conflict; the policy therefore requires medium-risk classification and routing the missing response to Priya Shah.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster.", "negative_left": "Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster.", "negative_right": "No member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop.", "right": "Every member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop."}, "verifier_independent_model": false}, "family": "fast-41-diverse-039-013", "id": "fast-41-diverse-039-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not mark it ready; classify it as medium risk and route the missing current-loop confirmation to required interviewer Priya Shah.", "true": "Yes — mark the loop ready to confirm at low risk, with no participant follow-up needed."}, "instructions": "Decide whether the October 13 loop should be marked ready to confirm at low scheduling risk without further routing. A loop is ready only when the candidate, hiring manager, and every required interviewer have confirmed the current loop. A free calendar, tentative hold, or confirmation for an earlier interview does not count. Optional participants are irrelevant. If the sole gap is an unconfirmed required interviewer with no calendar conflict, classify risk as medium and route the missing response to that interviewer.", "type": "noul"}}, "state": [{"speaker": "Recruiting coordinator", "text": "Nadia Cole and the hiring manager have each confirmed the October 13, 2026 interview loop. Priya Shah is a required interviewer, and every required interviewer other than Priya Shah has confirmed the loop. Priya has no calendar conflict during her scheduled October 13 interview block."}, {"speaker": "Scheduling auditor", "text": "Priya Shah is the sole required interviewer on Nadia Cole's October 13, 2026 interview-loop roster."}, {"speaker": "Roster record", "text": "No member of Nadia Cole's October 13, 2026 interview-loop roster has confirmed the loop."}]}, "method": "c2d", "provenance": {"source_id": "diverse-039", "source_is_synthetic": true, "source_sha256": "f5964e5081ce7e12090d385cccae7bfee5355483301100e6659c3269efd88d1b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question, governing policy, Thursday loop, participants, and meeting times. The evidence consists of the complete factual sentences “The candidate’s only stated Thursday availability is one continuous window beginning at 10:00.” and “The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes.” The counterfactual’s 75-minute window runs from 10:00 to 11:15 and covers both assigned slots without contradiction. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"calendar_clear": "supported", "candidate_coverage": "supported", "leo_coverage": "supported", "manager_link_recorded": "supported", "portfolio_link_recorded": "supported", "priya_coverage": "supported"}, "full_context_fact_states": {"base": {"calendar_clear": "supported", "candidate_coverage": "supported", "leo_coverage": "supported", "manager_link_recorded": "supported", "portfolio_link_recorded": "supported", "priya_coverage": "supported"}, "counterfactual": {"calendar_clear": "supported", "candidate_coverage": "refuted", "leo_coverage": "supported", "manager_link_recorded": "supported", "portfolio_link_recorded": "supported", "priya_coverage": "supported"}, "remove_left": {"candidate_coverage": "unknown"}, "remove_right": {"candidate_coverage": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"candidate_coverage": "unknown"}, "negative_pair": {"candidate_coverage": "refuted"}, "negative_sentence": {"candidate_coverage": "unknown"}, "positive_pair": {"candidate_coverage": "supported"}, "right": {"candidate_coverage": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; quantification over the two explicit slots does not make candidate_coverage or calendar_clear an impermissible bundle. The focus is factual rather than a policy conclusion. Base and counter assignments are realizable with only the candidate’s coverage changing. The policy evidence accurately preserves the substantive readiness rule originating in the state. Calendar and routing requirements need not be repeated there because they are already retained verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes full-slot coverage for the candidate, Leo, and Priya, no calendar conflicts, and recorded links for both interviews. These are all required conditions, so it is sufficient for Yes.", "rule_index": 0, "sound": true}, {"reason": "Refutation of candidate_coverage entails that the candidate’s availability fails to fully cover at least one assigned slot. That missing or unsuitable required availability is independently sufficient for No, while the other literals consistently establish that all remaining requirements are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "candidate_coverage", "statement": "The candidate’s stated availability fully covers each assigned Thursday interview slot: the 10:00–10:45 portfolio review with Leo and the 11:00–11:30 manager discussion with Priya."}, {"id": "leo_coverage", "statement": "Leo’s confirmation or stated availability fully covers his assigned Thursday 10:00–10:45 portfolio-review slot."}, {"id": "priya_coverage", "statement": "Priya’s confirmation or stated availability fully covers her assigned Thursday 11:00–11:30 manager-discussion slot."}, {"id": "calendar_clear", "statement": "The calendars show no conflicts for any participant during either assigned Thursday slot: 10:00–10:45 or 11:00–11:30."}, {"id": "portfolio_link_recorded", "statement": "The video link for the Thursday 10:00–10:45 portfolio review is recorded."}, {"id": "manager_link_recorded", "statement": "The video link for the Thursday 11:00–11:30 manager discussion is recorded."}], "base_state_json": "\"Mina’s Thursday interview loop has a portfolio review with Leo from 10:00–10:45 and a manager discussion with Priya from 11:00–11:30. The candidate’s only stated Thursday availability is one continuous window beginning at 10:00. The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes. Leo confirmed that he is available through 10:45 on Thursday, covering the portfolio-review slot. Priya accepted the 11:00 invitation, covering the manager-discussion slot. The shared calendars show no conflicts for the candidate, Leo, or Priya during either assigned meeting. The portfolio-review video link is recorded, and the manager-discussion video link is recorded. Mina has assigned both meetings to the named participants and has not changed their times. Policy says a loop is ready when each participant’s confirmation or stated availability covers the full assigned slot and links are present. Such a loop is low scheduling risk.\"", "base_states": [{"atom_id": "candidate_coverage", "state": "supported"}, {"atom_id": "leo_coverage", "state": "supported"}, {"atom_id": "priya_coverage", "state": "supported"}, {"atom_id": "calendar_clear", "state": "supported"}, {"atom_id": "portfolio_link_recorded", "state": "supported"}, {"atom_id": "manager_link_recorded", "state": "supported"}], "counter_states": [{"atom_id": "candidate_coverage", "state": "refuted"}, {"atom_id": "leo_coverage", "state": "supported"}, {"atom_id": "priya_coverage", "state": "supported"}, {"atom_id": "calendar_clear", "state": "supported"}, {"atom_id": "portfolio_link_recorded", "state": "supported"}, {"atom_id": "manager_link_recorded", "state": "supported"}], "focus_atom": "candidate_coverage", "focus_evidence": [{"path": [], "text": "The candidate’s only stated Thursday availability is one continuous window beginning at 10:00."}, {"path": [], "text": "The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when each participant’s confirmation or stated availability covers the full assigned slot and links are present. Such a loop is low scheduling risk."}], "rules": [{"justification": "Every participant is covered for the full assigned slot, the calendars show no conflicts, and both required video links are recorded.", "target": "true", "when": [{"atom_id": "candidate_coverage", "state": "supported"}, {"atom_id": "leo_coverage", "state": "supported"}, {"atom_id": "priya_coverage", "state": "supported"}, {"atom_id": "calendar_clear", "state": "supported"}, {"atom_id": "portfolio_link_recorded", "state": "supported"}, {"atom_id": "manager_link_recorded", "state": "supported"}]}, {"justification": "The candidate is not covered for at least one assigned slot, so a required participant confirmation or availability window is unsuitable even though every other requirement is satisfied.", "target": "false", "when": [{"atom_id": "candidate_coverage", "state": "refuted"}, {"atom_id": "leo_coverage", "state": "supported"}, {"atom_id": "priya_coverage", "state": "supported"}, {"atom_id": "calendar_clear", "state": "supported"}, {"atom_id": "portfolio_link_recorded", "state": "supported"}, {"atom_id": "manager_link_recorded", "state": "supported"}]}]}, "verified_pair": {"left": "The candidate’s only stated Thursday availability is one continuous window beginning at 10:00.", "negative_left": "The candidate’s only stated Thursday availability is one continuous window beginning at 10:00.", "negative_right": "The candidate’s only stated Thursday availability is one continuous window lasting exactly 75 minutes.", "right": "The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes."}, "verifier_independent_model": false}, "family": "fast-41-diverse-040-002", "id": "fast-41-diverse-040-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required confirmation, availability window, conflict check, or video link is missing or unsuitable, so the case must be routed before confirmation.", "true": "Yes — the loop satisfies every confirmation, calendar, and link requirement and is ready to confirm with low scheduling risk."}, "instructions": "Decide whether the interview loop is ready to confirm and has low scheduling risk. Treat a participant’s paraphrased availability range as confirmation when it fully covers the assigned slot. Answer Yes only if every participant is covered, calendars show no conflicts, and all links are recorded; otherwise answer No and route the issue to the participant responsible for the missing input or conflict.", "type": "noul"}}, "state": "Mina’s Thursday interview loop has a portfolio review with Leo from 10:00–10:45 and a manager discussion with Priya from 11:00–11:30. The candidate’s only stated Thursday availability is one continuous window beginning at 10:00. The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes. Leo confirmed that he is available through 10:45 on Thursday, covering the portfolio-review slot. Priya accepted the 11:00 invitation, covering the manager-discussion slot. The shared calendars show no conflicts for the candidate, Leo, or Priya during either assigned meeting. The portfolio-review video link is recorded, and the manager-discussion video link is recorded. Mina has assigned both meetings to the named participants and has not changed their times. Policy says a loop is ready when each participant’s confirmation or stated availability covers the full assigned slot and links are present. Such a loop is low scheduling risk."}, "method": "c2d", "provenance": {"source_id": "diverse-040", "source_is_synthetic": true, "source_sha256": "842af8f1ad82c75c55e0fd58ff66e6bd53dc2b1f58341510a65f2ded881eba8e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question, governing policy, Thursday loop, participants, and meeting times. The evidence consists of the complete factual sentences “The candidate’s only stated Thursday availability is one continuous window beginning at 10:00.” and “The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes.” The counterfactual’s 75-minute window runs from 10:00 to 11:15 and covers both assigned slots without contradiction. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"calendar_clear": "supported", "candidate_coverage": "refuted", "leo_coverage": "supported", "manager_link_recorded": "supported", "portfolio_link_recorded": "supported", "priya_coverage": "supported"}, "full_context_fact_states": {"base": {"calendar_clear": "supported", "candidate_coverage": "supported", "leo_coverage": "supported", "manager_link_recorded": "supported", "portfolio_link_recorded": "supported", "priya_coverage": "supported"}, "counterfactual": {"calendar_clear": "supported", "candidate_coverage": "refuted", "leo_coverage": "supported", "manager_link_recorded": "supported", "portfolio_link_recorded": "supported", "priya_coverage": "supported"}, "remove_left": {"candidate_coverage": "unknown"}, "remove_right": {"candidate_coverage": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"candidate_coverage": "unknown"}, "negative_pair": {"candidate_coverage": "refuted"}, "negative_sentence": {"candidate_coverage": "unknown"}, "positive_pair": {"candidate_coverage": "supported"}, "right": {"candidate_coverage": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; quantification over the two explicit slots does not make candidate_coverage or calendar_clear an impermissible bundle. The focus is factual rather than a policy conclusion. Base and counter assignments are realizable with only the candidate’s coverage changing. The policy evidence accurately preserves the substantive readiness rule originating in the state. Calendar and routing requirements need not be repeated there because they are already retained verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes full-slot coverage for the candidate, Leo, and Priya, no calendar conflicts, and recorded links for both interviews. These are all required conditions, so it is sufficient for Yes.", "rule_index": 0, "sound": true}, {"reason": "Refutation of candidate_coverage entails that the candidate’s availability fails to fully cover at least one assigned slot. That missing or unsuitable required availability is independently sufficient for No, while the other literals consistently establish that all remaining requirements are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "candidate_coverage", "statement": "The candidate’s stated availability fully covers each assigned Thursday interview slot: the 10:00–10:45 portfolio review with Leo and the 11:00–11:30 manager discussion with Priya."}, {"id": "leo_coverage", "statement": "Leo’s confirmation or stated availability fully covers his assigned Thursday 10:00–10:45 portfolio-review slot."}, {"id": "priya_coverage", "statement": "Priya’s confirmation or stated availability fully covers her assigned Thursday 11:00–11:30 manager-discussion slot."}, {"id": "calendar_clear", "statement": "The calendars show no conflicts for any participant during either assigned Thursday slot: 10:00–10:45 or 11:00–11:30."}, {"id": "portfolio_link_recorded", "statement": "The video link for the Thursday 10:00–10:45 portfolio review is recorded."}, {"id": "manager_link_recorded", "statement": "The video link for the Thursday 11:00–11:30 manager discussion is recorded."}], "base_state_json": "\"Mina’s Thursday interview loop has a portfolio review with Leo from 10:00–10:45 and a manager discussion with Priya from 11:00–11:30. The candidate’s only stated Thursday availability is one continuous window beginning at 10:00. The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes. Leo confirmed that he is available through 10:45 on Thursday, covering the portfolio-review slot. Priya accepted the 11:00 invitation, covering the manager-discussion slot. The shared calendars show no conflicts for the candidate, Leo, or Priya during either assigned meeting. The portfolio-review video link is recorded, and the manager-discussion video link is recorded. Mina has assigned both meetings to the named participants and has not changed their times. Policy says a loop is ready when each participant’s confirmation or stated availability covers the full assigned slot and links are present. Such a loop is low scheduling risk.\"", "base_states": [{"atom_id": "candidate_coverage", "state": "supported"}, {"atom_id": "leo_coverage", "state": "supported"}, {"atom_id": "priya_coverage", "state": "supported"}, {"atom_id": "calendar_clear", "state": "supported"}, {"atom_id": "portfolio_link_recorded", "state": "supported"}, {"atom_id": "manager_link_recorded", "state": "supported"}], "counter_states": [{"atom_id": "candidate_coverage", "state": "refuted"}, {"atom_id": "leo_coverage", "state": "supported"}, {"atom_id": "priya_coverage", "state": "supported"}, {"atom_id": "calendar_clear", "state": "supported"}, {"atom_id": "portfolio_link_recorded", "state": "supported"}, {"atom_id": "manager_link_recorded", "state": "supported"}], "focus_atom": "candidate_coverage", "focus_evidence": [{"path": [], "text": "The candidate’s only stated Thursday availability is one continuous window beginning at 10:00."}, {"path": [], "text": "The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when each participant’s confirmation or stated availability covers the full assigned slot and links are present. Such a loop is low scheduling risk."}], "rules": [{"justification": "Every participant is covered for the full assigned slot, the calendars show no conflicts, and both required video links are recorded.", "target": "true", "when": [{"atom_id": "candidate_coverage", "state": "supported"}, {"atom_id": "leo_coverage", "state": "supported"}, {"atom_id": "priya_coverage", "state": "supported"}, {"atom_id": "calendar_clear", "state": "supported"}, {"atom_id": "portfolio_link_recorded", "state": "supported"}, {"atom_id": "manager_link_recorded", "state": "supported"}]}, {"justification": "The candidate is not covered for at least one assigned slot, so a required participant confirmation or availability window is unsuitable even though every other requirement is satisfied.", "target": "false", "when": [{"atom_id": "candidate_coverage", "state": "refuted"}, {"atom_id": "leo_coverage", "state": "supported"}, {"atom_id": "priya_coverage", "state": "supported"}, {"atom_id": "calendar_clear", "state": "supported"}, {"atom_id": "portfolio_link_recorded", "state": "supported"}, {"atom_id": "manager_link_recorded", "state": "supported"}]}]}, "verified_pair": {"left": "The candidate’s only stated Thursday availability is one continuous window beginning at 10:00.", "negative_left": "The candidate’s only stated Thursday availability is one continuous window beginning at 10:00.", "negative_right": "The candidate’s only stated Thursday availability is one continuous window lasting exactly 75 minutes.", "right": "The candidate’s only stated Thursday availability is one continuous window lasting exactly 90 minutes."}, "verifier_independent_model": false}, "family": "fast-41-diverse-040-002", "id": "fast-41-diverse-040-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required confirmation, availability window, conflict check, or video link is missing or unsuitable, so the case must be routed before confirmation.", "true": "Yes — the loop satisfies every confirmation, calendar, and link requirement and is ready to confirm with low scheduling risk."}, "instructions": "Decide whether the interview loop is ready to confirm and has low scheduling risk. Treat a participant’s paraphrased availability range as confirmation when it fully covers the assigned slot. Answer Yes only if every participant is covered, calendars show no conflicts, and all links are recorded; otherwise answer No and route the issue to the participant responsible for the missing input or conflict.", "type": "noul"}}, "state": "Mina’s Thursday interview loop has a portfolio review with Leo from 10:00–10:45 and a manager discussion with Priya from 11:00–11:30. The candidate’s only stated Thursday availability is one continuous window beginning at 10:00. The candidate’s only stated Thursday availability is one continuous window lasting exactly 75 minutes. Leo confirmed that he is available through 10:45 on Thursday, covering the portfolio-review slot. Priya accepted the 11:00 invitation, covering the manager-discussion slot. The shared calendars show no conflicts for the candidate, Leo, or Priya during either assigned meeting. The portfolio-review video link is recorded, and the manager-discussion video link is recorded. Mina has assigned both meetings to the named participants and has not changed their times. Policy says a loop is ready when each participant’s confirmation or stated availability covers the full assigned slot and links are present. Such a loop is low scheduling risk."}, "method": "c2d", "provenance": {"source_id": "diverse-040", "source_is_synthetic": true, "source_sha256": "842af8f1ad82c75c55e0fd58ff66e6bd53dc2b1f58341510a65f2ded881eba8e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy, and both contexts add no conflicting policy. Both contexts retain the October 8, 10:00–11:30 a.m. loop and the candidate, Priya Shah, and Mateo Ruiz bindings. The two evidence spans are complete factual sentences: “The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.” “The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set.” The counterfactual coherently changes the confirmation count to candidate and Mateo Ruiz only, leaving Priya Shah unconfirmed without contradicting other statements. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set and universal no-conflict/no-decline propositions remain atomic. The focus atom concerns Priya's written confirmation, not a policy conclusion. Base and counter assignments are realizable with only that confirmation fact changing. Empty policy_evidence is correct because the governing scheduling rules are entirely in the questions object, which is automatically retained; the state contributes case observations rather than additional policy needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact required-participant set is established; every member explicitly confirmed; and known conflicts and declines are excluded. This is sufficient for Low risk under the preserved question policy.", "rule_index": 0, "sound": true}, {"reason": "Given the exact required-participant set, the candidate and hiring manager confirmed while Priya's confirmation is refuted, so exactly one required confirmation is missing. Known conflicts and declines are excluded, making Moderate risk sufficient. The table may validly abstain on cases where missing confirmation is merely unknown.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"id": "a2", "statement": "The candidate provided explicit written confirmation for the full proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a3", "statement": "Hiring manager Mateo Ruiz provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a4", "statement": "Interviewer Priya Shah provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a5", "statement": "No required participant has a known conflict with the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a6", "statement": "No required participant has a known decline for the proposed October 8, 10:00–11:30 a.m. loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"Case note for the proposed October 8, 10:00–11:30 a.m. loop. The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"The candidate sent written confirmation covering the entire 90-minute loop. Mateo Ruiz separately confirmed the proposed time and his participation. The scheduling review found no known calendar conflict affecting any required participant, and no required participant has recorded a decline.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"The appointment remains tied to the stated date, time, and participant list; the records concern this loop only, rather than an alternate meeting or a partial attendance window.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["0", "text"], "text": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"path": ["2", "text"], "text": "The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set."}], "policy_evidence": [], "rules": [{"justification": "Exactly all required participants explicitly confirmed the proposed loop. No required participant has a known conflict or decline, which also excludes any recorded calendar conflict and the competing High-risk triggers.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Priya Shah is exactly the one required participant lacking explicit confirmation. No required participant has a known conflict or decline, so neither High-risk trigger applies and the Moderate-risk criterion governs.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.", "negative_left": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.", "negative_right": "The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and no other required participant, with no confirmation from anyone outside the required participant set.", "right": "The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set."}, "verifier_independent_model": false}, "family": "fast-41-diverse-042-010", "id": "fast-41-diverse-042-010-base", "input": {"questions": {"decision": {"criteria": ["Low risk: Ready to confirm because the candidate and every required interviewer or hiring manager explicitly confirmed, with no recorded calendar conflict.", "Moderate risk: Not ready to confirm because exactly one required participant lacks explicit confirmation, but no participant has a known conflict; route follow-up to the unconfirmed participant.", "High risk: Not ready to confirm because a required participant has a known conflict or decline, or because two or more required confirmations are missing; escalate to the recruiting coordinator and hiring manager for replanning."], "instructions": "Rate scheduling risk and readiness under this policy: every required participant must provide an explicit written confirmation; an open calendar is not confirmation. If exactly one required confirmation is missing and no conflict is known, route follow-up to that participant. Select the single best level.", "type": "score"}}, "state": [{"speaker": "Recruiting coordinator", "text": "Case note for the proposed October 8, 10:00–11:30 a.m. loop. The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"speaker": "Recruiting coordinator", "text": "The candidate sent written confirmation covering the entire 90-minute loop. Mateo Ruiz separately confirmed the proposed time and his participation. The scheduling review found no known calendar conflict affecting any required participant, and no required participant has recorded a decline."}, {"speaker": "Recruiting coordinator", "text": "The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set."}, {"speaker": "Recruiting coordinator", "text": "The appointment remains tied to the stated date, time, and participant list; the records concern this loop only, rather than an alternate meeting or a partial attendance window."}]}, "method": "c2d", "provenance": {"source_id": "diverse-042", "source_is_synthetic": true, "source_sha256": "ecee670d01b33980a9f3e996c8338710a94c55ed2c82ca837c4c5b6102412664", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy, and both contexts add no conflicting policy. Both contexts retain the October 8, 10:00–11:30 a.m. loop and the candidate, Priya Shah, and Mateo Ruiz bindings. The two evidence spans are complete factual sentences: “The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.” “The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set.” The counterfactual coherently changes the confirmation count to candidate and Mateo Ruiz only, leaving Priya Shah unconfirmed without contradicting other statements. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set and universal no-conflict/no-decline propositions remain atomic. The focus atom concerns Priya's written confirmation, not a policy conclusion. Base and counter assignments are realizable with only that confirmation fact changing. Empty policy_evidence is correct because the governing scheduling rules are entirely in the questions object, which is automatically retained; the state contributes case observations rather than additional policy needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact required-participant set is established; every member explicitly confirmed; and known conflicts and declines are excluded. This is sufficient for Low risk under the preserved question policy.", "rule_index": 0, "sound": true}, {"reason": "Given the exact required-participant set, the candidate and hiring manager confirmed while Priya's confirmation is refuted, so exactly one required confirmation is missing. Known conflicts and declines are excluded, making Moderate risk sufficient. The table may validly abstain on cases where missing confirmation is merely unknown.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"id": "a2", "statement": "The candidate provided explicit written confirmation for the full proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a3", "statement": "Hiring manager Mateo Ruiz provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a4", "statement": "Interviewer Priya Shah provided explicit written confirmation for the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a5", "statement": "No required participant has a known conflict with the proposed October 8, 10:00–11:30 a.m. loop."}, {"id": "a6", "statement": "No required participant has a known decline for the proposed October 8, 10:00–11:30 a.m. loop."}], "base_state_json": "[{\"speaker\":\"Recruiting coordinator\",\"text\":\"Case note for the proposed October 8, 10:00–11:30 a.m. loop. The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"The candidate sent written confirmation covering the entire 90-minute loop. Mateo Ruiz separately confirmed the proposed time and his participation. The scheduling review found no known calendar conflict affecting any required participant, and no required participant has recorded a decline.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set.\"},{\"speaker\":\"Recruiting coordinator\",\"text\":\"The appointment remains tied to the stated date, time, and participant list; the records concern this loop only, rather than an alternate meeting or a partial attendance window.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["0", "text"], "text": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"path": ["2", "text"], "text": "The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set."}], "policy_evidence": [], "rules": [{"justification": "Exactly all required participants explicitly confirmed the proposed loop. No required participant has a known conflict or decline, which also excludes any recorded calendar conflict and the competing High-risk triggers.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "Priya Shah is exactly the one required participant lacking explicit confirmation. No required participant has a known conflict or decline, so neither High-risk trigger applies and the Moderate-risk criterion governs.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.", "negative_left": "The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz.", "negative_right": "The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and no other required participant, with no confirmation from anyone outside the required participant set.", "right": "The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and one other required participant, with no confirmation from anyone outside the required participant set."}, "verifier_independent_model": false}, "family": "fast-41-diverse-042-010", "id": "fast-41-diverse-042-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["Low risk: Ready to confirm because the candidate and every required interviewer or hiring manager explicitly confirmed, with no recorded calendar conflict.", "Moderate risk: Not ready to confirm because exactly one required participant lacks explicit confirmation, but no participant has a known conflict; route follow-up to the unconfirmed participant.", "High risk: Not ready to confirm because a required participant has a known conflict or decline, or because two or more required confirmations are missing; escalate to the recruiting coordinator and hiring manager for replanning."], "instructions": "Rate scheduling risk and readiness under this policy: every required participant must provide an explicit written confirmation; an open calendar is not confirmation. If exactly one required confirmation is missing and no conflict is known, route follow-up to that participant. Select the single best level.", "type": "score"}}, "state": [{"speaker": "Recruiting coordinator", "text": "Case note for the proposed October 8, 10:00–11:30 a.m. loop. The set of required participants for the proposed October 8, 10:00–11:30 a.m. loop is exactly the candidate, interviewer Priya Shah, and hiring manager Mateo Ruiz."}, {"speaker": "Recruiting coordinator", "text": "The candidate sent written confirmation covering the entire 90-minute loop. Mateo Ruiz separately confirmed the proposed time and his participation. The scheduling review found no known calendar conflict affecting any required participant, and no required participant has recorded a decline."}, {"speaker": "Recruiting coordinator", "text": "The confirmation record for the proposed October 8, 10:00–11:30 a.m. loop lists explicit written confirmations from the candidate, Mateo Ruiz, and no other required participant, with no confirmation from anyone outside the required participant set."}, {"speaker": "Recruiting coordinator", "text": "The appointment remains tied to the stated date, time, and participant list; the records concern this loop only, rather than an alternate meeting or a partial attendance window."}]}, "method": "c2d", "provenance": {"source_id": "diverse-042", "source_is_synthetic": true, "source_sha256": "ecee670d01b33980a9f3e996c8338710a94c55ed2c82ca837c4c5b6102412664", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and question bindings, use complete factual evidence sentences, remain coherent after the attachment change, and contain no leaked answer or classifier instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns whether quantified schedule evidence is attached rather than a policy conclusion. The base and counter assignments are realizable with only that attachment fact changing. The policy evidence accurately preserves the impact thresholds, owner bindings, and readiness rule originating in the original state; rules already contained in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Eight added workdays and $12,000 establish moderate impact and therefore delivery-lead review. The conjunction also establishes a rationale plus attached quantified schedule and resource evidence, which is sufficient for readiness.", "rule_index": 0, "sound": true}, {"reason": "Eight added workdays and $12,000 establish moderate impact and therefore delivery-lead review. Refutation of attached quantified schedule evidence establishes that a required readiness component is missing, so the request is not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Ivo’s audit-export change request includes a rationale explaining its need."}, {"id": "a2", "statement": "Ivo’s audit-export change request adds eight workdays."}, {"id": "a3", "statement": "Ivo’s audit-export change request requires $12,000 in contractor support."}, {"id": "a4", "statement": "Ivo’s audit-export change request has attached quantified schedule evidence."}, {"id": "a5", "statement": "Ivo’s audit-export change request has attached quantified resource evidence."}], "base_state_json": "{\"context\":\"Case note: Mara is triaging Ivo’s audit-export change request. Ivo’s submission explains that the export is needed for a new contractual reporting clause. The request adds eight workdays and requires $12,000 in contractor support. A separate cost sheet quantifies the contractor requirement. Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000. Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes. A request is ready only with a rationale and attached, quantified schedule and resource evidence. The submission manifest and its associated file record identify the materials supplied for review.\",\"evidence\":[\"S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission.\",\"The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as included among the attached files.\",\"Ivo’s rationale identifies a new contractual reporting clause as the need for the export.\",\"The request specifies eight added workdays and $12,000 in contractor support.\",\"The attached cost sheet quantifies the contractor-support requirement.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission."}, {"path": ["evidence", "1"], "text": "The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as included among the attached files."}], "policy_evidence": [{"path": ["context"], "text": "Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000."}, {"path": ["context"], "text": "Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes."}, {"path": ["context"], "text": "A request is ready only with a rationale and attached, quantified schedule and resource evidence."}], "rules": [{"justification": "Eight added workdays and $12,000 are moderate impact, which is reviewed by the delivery lead. The rationale and both forms of attached quantified evidence are present, so the request is ready.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Eight added workdays and $12,000 are moderate impact, which is reviewed by the delivery lead. Although the rationale and quantified resource evidence are present, required attached quantified schedule evidence is missing, so the request is not ready.", "target": "delivery_lead_not_ready_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission.", "negative_left": "S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission.", "negative_right": "The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as excluded from the attached files.", "right": "The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as included among the attached files."}, "verifier_independent_model": false}, "family": "fast-41-diverse-043-005", "id": "fast-41-diverse-043-005-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_moderate": "Delivery lead review; not ready; moderate impact. Use only when impact is 3–10 days or $5,001–$25,000 but required rationale or attached quantified evidence is missing.", "delivery_lead_ready_moderate": "Delivery lead review; ready; moderate impact. Use only when impact is 3–10 days or $5,001–$25,000 and both rationale and attached quantified evidence are present.", "none_of_above": "Use only when the correct owner, readiness, and severity combination is not represented by any other option.", "project_manager_not_ready_low": "Project manager review; not ready; low impact. Use only when impact is at most two days and $5,000 but required rationale or attached quantified evidence is missing.", "project_manager_ready_low": "Project manager review; ready; low impact. Use only when the request has the required rationale and evidence and adds at most two days and $5,000.", "project_sponsor_ready_severe": "Project sponsor review; ready; severe impact. Use only when the request has the required rationale and evidence and exceeds ten days or $25,000."}, "instructions": "Select the single option that correctly identifies the review owner, readiness, and delivery-impact severity. Resolve pronouns and references using their surrounding context. The categories in the options are exact combinations; choose none_of_above only if no listed combination applies.", "type": "choice"}}, "state": {"context": "Case note: Mara is triaging Ivo’s audit-export change request. Ivo’s submission explains that the export is needed for a new contractual reporting clause. The request adds eight workdays and requires $12,000 in contractor support. A separate cost sheet quantifies the contractor requirement. Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000. Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes. A request is ready only with a rationale and attached, quantified schedule and resource evidence. The submission manifest and its associated file record identify the materials supplied for review.", "evidence": ["S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission.", "The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as included among the attached files.", "Ivo’s rationale identifies a new contractual reporting clause as the need for the export.", "The request specifies eight added workdays and $12,000 in contractor support.", "The attached cost sheet quantifies the contractor-support requirement."]}}, "method": "c2d", "provenance": {"source_id": "diverse-043", "source_is_synthetic": true, "source_sha256": "1e08b3d5ffbc4028fbb9e3181f5514c4bafda4745f42969ea60a2dd3f1360142", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and question bindings, use complete factual evidence sentences, remain coherent after the attachment change, and contain no leaked answer or classifier instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns whether quantified schedule evidence is attached rather than a policy conclusion. The base and counter assignments are realizable with only that attachment fact changing. The policy evidence accurately preserves the impact thresholds, owner bindings, and readiness rule originating in the original state; rules already contained in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Eight added workdays and $12,000 establish moderate impact and therefore delivery-lead review. The conjunction also establishes a rationale plus attached quantified schedule and resource evidence, which is sufficient for readiness.", "rule_index": 0, "sound": true}, {"reason": "Eight added workdays and $12,000 establish moderate impact and therefore delivery-lead review. Refutation of attached quantified schedule evidence establishes that a required readiness component is missing, so the request is not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Ivo’s audit-export change request includes a rationale explaining its need."}, {"id": "a2", "statement": "Ivo’s audit-export change request adds eight workdays."}, {"id": "a3", "statement": "Ivo’s audit-export change request requires $12,000 in contractor support."}, {"id": "a4", "statement": "Ivo’s audit-export change request has attached quantified schedule evidence."}, {"id": "a5", "statement": "Ivo’s audit-export change request has attached quantified resource evidence."}], "base_state_json": "{\"context\":\"Case note: Mara is triaging Ivo’s audit-export change request. Ivo’s submission explains that the export is needed for a new contractual reporting clause. The request adds eight workdays and requires $12,000 in contractor support. A separate cost sheet quantifies the contractor requirement. Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000. Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes. A request is ready only with a rationale and attached, quantified schedule and resource evidence. The submission manifest and its associated file record identify the materials supplied for review.\",\"evidence\":[\"S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission.\",\"The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as included among the attached files.\",\"Ivo’s rationale identifies a new contractual reporting clause as the need for the export.\",\"The request specifies eight added workdays and $12,000 in contractor support.\",\"The attached cost sheet quantifies the contractor-support requirement.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission."}, {"path": ["evidence", "1"], "text": "The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as included among the attached files."}], "policy_evidence": [{"path": ["context"], "text": "Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000."}, {"path": ["context"], "text": "Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes."}, {"path": ["context"], "text": "A request is ready only with a rationale and attached, quantified schedule and resource evidence."}], "rules": [{"justification": "Eight added workdays and $12,000 are moderate impact, which is reviewed by the delivery lead. The rationale and both forms of attached quantified evidence are present, so the request is ready.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Eight added workdays and $12,000 are moderate impact, which is reviewed by the delivery lead. Although the rationale and quantified resource evidence are present, required attached quantified schedule evidence is missing, so the request is not ready.", "target": "delivery_lead_not_ready_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission.", "negative_left": "S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission.", "negative_right": "The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as excluded from the attached files.", "right": "The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as included among the attached files."}, "verifier_independent_model": false}, "family": "fast-41-diverse-043-005", "id": "fast-41-diverse-043-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_moderate": "Delivery lead review; not ready; moderate impact. Use only when impact is 3–10 days or $5,001–$25,000 but required rationale or attached quantified evidence is missing.", "delivery_lead_ready_moderate": "Delivery lead review; ready; moderate impact. Use only when impact is 3–10 days or $5,001–$25,000 and both rationale and attached quantified evidence are present.", "none_of_above": "Use only when the correct owner, readiness, and severity combination is not represented by any other option.", "project_manager_not_ready_low": "Project manager review; not ready; low impact. Use only when impact is at most two days and $5,000 but required rationale or attached quantified evidence is missing.", "project_manager_ready_low": "Project manager review; ready; low impact. Use only when the request has the required rationale and evidence and adds at most two days and $5,000.", "project_sponsor_ready_severe": "Project sponsor review; ready; severe impact. Use only when the request has the required rationale and evidence and exceeds ten days or $25,000."}, "instructions": "Select the single option that correctly identifies the review owner, readiness, and delivery-impact severity. Resolve pronouns and references using their surrounding context. The categories in the options are exact combinations; choose none_of_above only if no listed combination applies.", "type": "choice"}}, "state": {"context": "Case note: Mara is triaging Ivo’s audit-export change request. Ivo’s submission explains that the export is needed for a new contractual reporting clause. The request adds eight workdays and requires $12,000 in contractor support. A separate cost sheet quantifies the contractor requirement. Low impact means 0–2 added workdays and no more than $5,000; moderate means 3–10 days or $5,001–$25,000; severe means over 10 days or over $25,000. Mara reviews low changes, delivery lead Nia reviews moderate changes, and project sponsor Omar reviews severe changes. A request is ready only with a rationale and attached, quantified schedule and resource evidence. The submission manifest and its associated file record identify the materials supplied for review.", "evidence": ["S-417 is the sole quantified schedule-evidence spreadsheet associated with Ivo’s 2026-09-17 audit-export submission.", "The 2026-09-17 submission manifest for Ivo’s audit-export change request records S-417 as excluded from the attached files.", "Ivo’s rationale identifies a new contractual reporting clause as the need for the export.", "The request specifies eight added workdays and $12,000 in contractor support.", "The attached cost sheet quantifies the contractor-support requirement."]}}, "method": "c2d", "provenance": {"source_id": "diverse-043", "source_is_synthetic": true, "source_sha256": "1e08b3d5ffbc4028fbb9e3181f5514c4bafda4745f42969ea60a2dd3f1360142", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_not_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both context policy entries preserve the governing policy. The internal-beta entity and review-scope bindings remain unchanged while dates may vary as observations. The two evidence spans are complete factual sentences: \"The currently scheduled internal-beta date is 4 March 2027.\" and \"The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days.\" The counterfactual moves the date from 4 to 9 March, a coherent five-business-day change under the stated calendar. Neither context embeds an answer, code, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Schedule record\",\"text\":\"The currently scheduled internal-beta date is 4 March 2027.\"},{\"speaker\":\"Proposed schedule record\",\"text\":\"The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days.\"},{\"speaker\":\"Change requester\",\"text\":\"The internal-beta change request explains that authentication defects require retesting after the contractor departs. It includes a signed schedule estimate.\"},{\"speaker\":\"Dependency register\",\"text\":\"Every dependency owner affected by the internal-beta change request has provided written acknowledgement.\"},{\"speaker\":\"Budget review\",\"text\":\"The internal-beta change request makes no budget change.\"},{\"speaker\":\"Scope review\",\"text\":\"The internal-beta change request makes no scope-baseline change.\"},{\"speaker\":\"Contract review\",\"text\":\"The internal-beta change request causes no contractual milestone miss.\"},{\"speaker\":\"Policy\",\"text\":\"Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The currently scheduled internal-beta date is 4 March 2027."}, {"path": ["1", "text"], "text": "The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The currently scheduled internal-beta date is 4 March 2027.", "negative_left": "The currently scheduled internal-beta date is 4 March 2027.", "negative_right": "The proposed internal-beta date is 9 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days.", "right": "The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-001", "id": "fast-41-diverse-044-001-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Schedule record", "text": "The currently scheduled internal-beta date is 4 March 2027."}, {"speaker": "Proposed schedule record", "text": "The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days."}, {"speaker": "Change requester", "text": "The internal-beta change request explains that authentication defects require retesting after the contractor departs. It includes a signed schedule estimate."}, {"speaker": "Dependency register", "text": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement."}, {"speaker": "Budget review", "text": "The internal-beta change request makes no budget change."}, {"speaker": "Scope review", "text": "The internal-beta change request makes no scope-baseline change."}, {"speaker": "Contract review", "text": "The internal-beta change request causes no contractual milestone miss."}, {"speaker": "Policy", "text": "Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both context policy entries preserve the governing policy. The internal-beta entity and review-scope bindings remain unchanged while dates may vary as observations. The two evidence spans are complete factual sentences: \"The currently scheduled internal-beta date is 4 March 2027.\" and \"The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days.\" The counterfactual moves the date from 4 to 9 March, a coherent five-business-day change under the stated calendar. Neither context embeds an answer, code, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Schedule record\",\"text\":\"The currently scheduled internal-beta date is 4 March 2027.\"},{\"speaker\":\"Proposed schedule record\",\"text\":\"The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days.\"},{\"speaker\":\"Change requester\",\"text\":\"The internal-beta change request explains that authentication defects require retesting after the contractor departs. It includes a signed schedule estimate.\"},{\"speaker\":\"Dependency register\",\"text\":\"Every dependency owner affected by the internal-beta change request has provided written acknowledgement.\"},{\"speaker\":\"Budget review\",\"text\":\"The internal-beta change request makes no budget change.\"},{\"speaker\":\"Scope review\",\"text\":\"The internal-beta change request makes no scope-baseline change.\"},{\"speaker\":\"Contract review\",\"text\":\"The internal-beta change request causes no contractual milestone miss.\"},{\"speaker\":\"Policy\",\"text\":\"Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The currently scheduled internal-beta date is 4 March 2027."}, {"path": ["1", "text"], "text": "The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The currently scheduled internal-beta date is 4 March 2027.", "negative_left": "The currently scheduled internal-beta date is 4 March 2027.", "negative_right": "The proposed internal-beta date is 9 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days.", "right": "The proposed internal-beta date is 12 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-001", "id": "fast-41-diverse-044-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Schedule record", "text": "The currently scheduled internal-beta date is 4 March 2027."}, {"speaker": "Proposed schedule record", "text": "The proposed internal-beta date is 9 March 2027, and the applicable project calendar marks 5, 6, 7, 8, 9, 10, 11, and 12 March 2027 as business days."}, {"speaker": "Change requester", "text": "The internal-beta change request explains that authentication defects require retesting after the contractor departs. It includes a signed schedule estimate."}, {"speaker": "Dependency register", "text": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement."}, {"speaker": "Budget review", "text": "The internal-beta change request makes no budget change."}, {"speaker": "Scope review", "text": "The internal-beta change request makes no scope-baseline change."}, {"speaker": "Contract review", "text": "The internal-beta change request causes no contractual milestone miss."}, {"speaker": "Policy", "text": "Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The exact evidence sentences are “The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47.” and “Interval IB-2026-09-17-A spans six business days between its endpoints.” Both contexts preserve the request and policy bindings, and the four-day counterfactual is coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47. The request explains that contractor departure requires authentication retesting.\"},{\"speaker\":\"Delivery lead\",\"text\":\"Interval IB-2026-09-17-A spans six business days between its endpoints. I attached a signed schedule estimate, and each affected dependency owner has provided written acknowledgement.\"},{\"speaker\":\"Project manager\",\"text\":\"The internal-beta change request leaves the budget unchanged and does not alter the scope baseline. No contractual milestone will be missed.\"},{\"speaker\":\"Change requester\",\"text\":\"The request is for an internal-beta schedule adjustment, and its rationale is included in the change record. The dependency acknowledgements cover every owner affected by the request.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47."}, {"path": ["1", "text"], "text": "Interval IB-2026-09-17-A spans six business days between its endpoints."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47.", "negative_left": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47.", "negative_right": "Interval IB-2026-09-17-A spans four business days between its endpoints.", "right": "Interval IB-2026-09-17-A spans six business days between its endpoints."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-003", "id": "fast-41-diverse-044-003-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47. The request explains that contractor departure requires authentication retesting."}, {"speaker": "Delivery lead", "text": "Interval IB-2026-09-17-A spans six business days between its endpoints. I attached a signed schedule estimate, and each affected dependency owner has provided written acknowledgement."}, {"speaker": "Project manager", "text": "The internal-beta change request leaves the budget unchanged and does not alter the scope baseline. No contractual milestone will be missed."}, {"speaker": "Change requester", "text": "The request is for an internal-beta schedule adjustment, and its rationale is included in the change record. The dependency acknowledgements cover every owner affected by the request."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The exact evidence sentences are “The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47.” and “Interval IB-2026-09-17-A spans six business days between its endpoints.” Both contexts preserve the request and policy bindings, and the four-day counterfactual is coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47. The request explains that contractor departure requires authentication retesting.\"},{\"speaker\":\"Delivery lead\",\"text\":\"Interval IB-2026-09-17-A spans six business days between its endpoints. I attached a signed schedule estimate, and each affected dependency owner has provided written acknowledgement.\"},{\"speaker\":\"Project manager\",\"text\":\"The internal-beta change request leaves the budget unchanged and does not alter the scope baseline. No contractual milestone will be missed.\"},{\"speaker\":\"Change requester\",\"text\":\"The request is for an internal-beta schedule adjustment, and its rationale is included in the change record. The dependency acknowledgements cover every owner affected by the request.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47."}, {"path": ["1", "text"], "text": "Interval IB-2026-09-17-A spans six business days between its endpoints."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47.", "negative_left": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47.", "negative_right": "Interval IB-2026-09-17-A spans four business days between its endpoints.", "right": "Interval IB-2026-09-17-A spans six business days between its endpoints."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-003", "id": "fast-41-diverse-044-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date is recorded as interval IB-2026-09-17-A in change record IB-CR-47. The request explains that contractor departure requires authentication retesting."}, {"speaker": "Delivery lead", "text": "Interval IB-2026-09-17-A spans four business days between its endpoints. I attached a signed schedule estimate, and each affected dependency owner has provided written acknowledgement."}, {"speaker": "Project manager", "text": "The internal-beta change request leaves the budget unchanged and does not alter the scope baseline. No contractual milestone will be missed."}, {"speaker": "Change requester", "text": "The request is for an internal-beta schedule adjustment, and its rationale is included in the change record. The dependency acknowledgements cover every owner affected by the request."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and bindings, the evidence consists of two complete factual sentences—\"The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar.\" and \"The release calendar records interval IB-Delta-17 as lasting seven business days.\"—and the four-day counterfactual is explicitly distinguished from the seven-day base observation without leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar.\"},{\"speaker\":\"Release calendar\",\"text\":\"The release calendar records interval IB-Delta-17 as lasting seven business days.\"},{\"speaker\":\"Change requester\",\"text\":\"The request includes a written rationale explaining the move and a signed schedule estimate.\"},{\"speaker\":\"Dependency register\",\"text\":\"Every dependency owner affected by the request has provided written acknowledgement.\"},{\"speaker\":\"Finance and scope records\",\"text\":\"The request includes no budget change and no scope-baseline change.\"},{\"speaker\":\"Contract register\",\"text\":\"The request causes no contractual milestone miss.\"},{\"speaker\":\"Counterfactual release calendar\",\"text\":\"Under the counterfactual calendar, the release calendar records interval IB-Delta-17 as lasting four business days; all other recorded facts remain unchanged.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar."}, {"path": ["1", "text"], "text": "The release calendar records interval IB-Delta-17 as lasting seven business days."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar.", "negative_left": "The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar.", "negative_right": "The release calendar records interval IB-Delta-17 as lasting four business days.", "right": "The release calendar records interval IB-Delta-17 as lasting seven business days."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-005", "id": "fast-41-diverse-044-005-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Release coordinator", "text": "The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar."}, {"speaker": "Release calendar", "text": "The release calendar records interval IB-Delta-17 as lasting seven business days."}, {"speaker": "Change requester", "text": "The request includes a written rationale explaining the move and a signed schedule estimate."}, {"speaker": "Dependency register", "text": "Every dependency owner affected by the request has provided written acknowledgement."}, {"speaker": "Finance and scope records", "text": "The request includes no budget change and no scope-baseline change."}, {"speaker": "Contract register", "text": "The request causes no contractual milestone miss."}, {"speaker": "Counterfactual release calendar", "text": "Under the counterfactual calendar, the release calendar records interval IB-Delta-17 as lasting four business days; all other recorded facts remain unchanged."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and bindings, the evidence consists of two complete factual sentences—\"The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar.\" and \"The release calendar records interval IB-Delta-17 as lasting seven business days.\"—and the four-day counterfactual is explicitly distinguished from the seven-day base observation without leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar.\"},{\"speaker\":\"Release calendar\",\"text\":\"The release calendar records interval IB-Delta-17 as lasting seven business days.\"},{\"speaker\":\"Change requester\",\"text\":\"The request includes a written rationale explaining the move and a signed schedule estimate.\"},{\"speaker\":\"Dependency register\",\"text\":\"Every dependency owner affected by the request has provided written acknowledgement.\"},{\"speaker\":\"Finance and scope records\",\"text\":\"The request includes no budget change and no scope-baseline change.\"},{\"speaker\":\"Contract register\",\"text\":\"The request causes no contractual milestone miss.\"},{\"speaker\":\"Counterfactual release calendar\",\"text\":\"Under the counterfactual calendar, the release calendar records interval IB-Delta-17 as lasting four business days; all other recorded facts remain unchanged.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar."}, {"path": ["1", "text"], "text": "The release calendar records interval IB-Delta-17 as lasting seven business days."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar.", "negative_left": "The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar.", "negative_right": "The release calendar records interval IB-Delta-17 as lasting four business days.", "right": "The release calendar records interval IB-Delta-17 as lasting seven business days."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-005", "id": "fast-41-diverse-044-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Release coordinator", "text": "The internal-beta change request concerns moving from the currently scheduled internal-beta date to the proposed internal-beta date, an interval labeled IB-Delta-17 in the release calendar."}, {"speaker": "Release calendar", "text": "The release calendar records interval IB-Delta-17 as lasting four business days."}, {"speaker": "Change requester", "text": "The request includes a written rationale explaining the move and a signed schedule estimate."}, {"speaker": "Dependency register", "text": "Every dependency owner affected by the request has provided written acknowledgement."}, {"speaker": "Finance and scope records", "text": "The request includes no budget change and no scope-baseline change."}, {"speaker": "Contract register", "text": "The request causes no contractual milestone miss."}, {"speaker": "Counterfactual release calendar", "text": "Under the counterfactual calendar, the release calendar records interval IB-Delta-17 as lasting four business days; all other recorded facts remain unchanged."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions are verbatim, and both contexts preserve the governing policy without added exceptions. The internal-beta request and its date relationships remain bound to the same entity and review question. The evidence consists of the complete factual sentences “The currently scheduled internal-beta date is 4 May 2027.” and “The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027.” The counterfactual coherently changes the current date to 8 May 2027 without contradictory measurements within that context. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[\"The currently scheduled internal-beta date is 4 May 2027.\",\"The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027.\",\"The change request explains that authentication defects require retesting after the contractor departs.\",\"A signed schedule estimate is attached to the request.\",\"Every dependency owner affected by the request has provided written acknowledgement.\",\"The request does not change the budget or scope baseline.\",\"The project manager confirmed that no contractual milestone will be missed.\",\"Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.\"]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0"], "text": "The currently scheduled internal-beta date is 4 May 2027."}, {"path": ["1"], "text": "The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The currently scheduled internal-beta date is 4 May 2027.", "negative_left": "The currently scheduled internal-beta date is 8 May 2027.", "negative_right": "The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027.", "right": "The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-011", "id": "fast-41-diverse-044-011-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": ["The currently scheduled internal-beta date is 4 May 2027.", "The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027.", "The change request explains that authentication defects require retesting after the contractor departs.", "A signed schedule estimate is attached to the request.", "Every dependency owner affected by the request has provided written acknowledgement.", "The request does not change the budget or scope baseline.", "The project manager confirmed that no contractual milestone will be missed.", "Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests."]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions are verbatim, and both contexts preserve the governing policy without added exceptions. The internal-beta request and its date relationships remain bound to the same entity and review question. The evidence consists of the complete factual sentences “The currently scheduled internal-beta date is 4 May 2027.” and “The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027.” The counterfactual coherently changes the current date to 8 May 2027 without contradictory measurements within that context. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[\"The currently scheduled internal-beta date is 4 May 2027.\",\"The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027.\",\"The change request explains that authentication defects require retesting after the contractor departs.\",\"A signed schedule estimate is attached to the request.\",\"Every dependency owner affected by the request has provided written acknowledgement.\",\"The request does not change the budget or scope baseline.\",\"The project manager confirmed that no contractual milestone will be missed.\",\"Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.\"]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0"], "text": "The currently scheduled internal-beta date is 4 May 2027."}, {"path": ["1"], "text": "The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The currently scheduled internal-beta date is 4 May 2027.", "negative_left": "The currently scheduled internal-beta date is 8 May 2027.", "negative_right": "The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027.", "right": "The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-011", "id": "fast-41-diverse-044-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": ["The currently scheduled internal-beta date is 8 May 2027.", "The proposed internal-beta date is 13 May 2027, and the project calendar records 4 May 2027 as seven business days before 13 May 2027 and 8 May 2027 as three business days before 13 May 2027.", "The change request explains that authentication defects require retesting after the contractor departs.", "A signed schedule estimate is attached to the request.", "Every dependency owner affected by the request has provided written acknowledgement.", "The request does not change the budget or scope baseline.", "The project manager confirmed that no contractual milestone will be missed.", "Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests."]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve the governing policy, criteria, and thresholds, while both contexts add no conflicting policy. Both contexts retain the internal-beta request and IB-204 schedule binding without introducing contradictory duplicate assertions. The two focus spans are complete factual sentences and contain no policy or instruction. The counterfactual changes only the interval from six to five business days, so it remains coherent with the other unchanged facts. Neither context contains a gold answer, answer code, label rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move. The request file includes a rationale explaining that authentication defects require retesting after a contractor departure.\"},{\"speaker\":\"Delivery lead\",\"text\":\"The authorized business-day calendar lists the interval in schedule record IB-204 as six business days. A signed schedule estimate is attached, and every dependency owner affected by the request has provided written acknowledgement.\"},{\"speaker\":\"Project manager\",\"text\":\"The request changes neither the budget nor the scope baseline. The affected handoff can move without causing a contractual milestone miss.\"},{\"speaker\":\"Project sponsor\",\"text\":\"The submission is complete for review, and the promotional conference demonstration does not create a contractual milestone.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move."}, {"path": ["1", "text"], "text": "The authorized business-day calendar lists the interval in schedule record IB-204 as six business days."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move.", "negative_left": "The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move.", "negative_right": "The authorized business-day calendar lists the interval in schedule record IB-204 as five business days.", "right": "The authorized business-day calendar lists the interval in schedule record IB-204 as six business days."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-012", "id": "fast-41-diverse-044-012-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move. The request file includes a rationale explaining that authentication defects require retesting after a contractor departure."}, {"speaker": "Delivery lead", "text": "The authorized business-day calendar lists the interval in schedule record IB-204 as six business days. A signed schedule estimate is attached, and every dependency owner affected by the request has provided written acknowledgement."}, {"speaker": "Project manager", "text": "The request changes neither the budget nor the scope baseline. The affected handoff can move without causing a contractual milestone miss."}, {"speaker": "Project sponsor", "text": "The submission is complete for review, and the promotional conference demonstration does not create a contractual milestone."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve the governing policy, criteria, and thresholds, while both contexts add no conflicting policy. Both contexts retain the internal-beta request and IB-204 schedule binding without introducing contradictory duplicate assertions. The two focus spans are complete factual sentences and contain no policy or instruction. The counterfactual changes only the interval from six to five business days, so it remains coherent with the other unchanged facts. Neither context contains a gold answer, answer code, label rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move. The request file includes a rationale explaining that authentication defects require retesting after a contractor departure.\"},{\"speaker\":\"Delivery lead\",\"text\":\"The authorized business-day calendar lists the interval in schedule record IB-204 as six business days. A signed schedule estimate is attached, and every dependency owner affected by the request has provided written acknowledgement.\"},{\"speaker\":\"Project manager\",\"text\":\"The request changes neither the budget nor the scope baseline. The affected handoff can move without causing a contractual milestone miss.\"},{\"speaker\":\"Project sponsor\",\"text\":\"The submission is complete for review, and the promotional conference demonstration does not create a contractual milestone.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move."}, {"path": ["1", "text"], "text": "The authorized business-day calendar lists the interval in schedule record IB-204 as six business days."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move.", "negative_left": "The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move.", "negative_right": "The authorized business-day calendar lists the interval in schedule record IB-204 as five business days.", "right": "The authorized business-day calendar lists the interval in schedule record IB-204 as six business days."}, "verifier_independent_model": false}, "family": "fast-41-diverse-044-012", "id": "fast-41-diverse-044-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta change request identifies schedule record IB-204 as containing the currently scheduled and proposed dates for the requested move. The request file includes a rationale explaining that authentication defects require retesting after a contractor departure."}, {"speaker": "Delivery lead", "text": "The authorized business-day calendar lists the interval in schedule record IB-204 as five business days. A signed schedule estimate is attached, and every dependency owner affected by the request has provided written acknowledgement."}, {"speaker": "Project manager", "text": "The request changes neither the budget nor the scope baseline. The affected handoff can move without causing a contractual milestone miss."}, {"speaker": "Project sponsor", "text": "The submission is complete for review, and the promotional conference demonstration does not create a contractual milestone."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and question bindings remain unchanged in both full inputs. Base evidence is exactly two complete factual sentences: \"In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial.\" and \"The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial.\" Counterfactual evidence is exactly two complete factual sentences: \"In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial.\" and \"The attached January 2026 pilot report for Mara's advanced-filtering request records 29% fewer filtering errors during that trial.\" The counterfactual coherently changes the report measurement and creates an explicit paraphrase mismatch without contradictory duplicate assertions. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Case note — 14 March 2026: Mara submitted a request to add advanced filtering to the reporting module. The request explains why the enhancement is sought and states delivery effects of five additional weeks and $60,000. It is logged as a scope change, and the project sponsor is the designated recipient for that type of request. In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial. The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial. The request contains a rationale, a quantified impact, and an attached pilot report for review. The delivery lead supplied the schedule and cost estimates. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial."}, {"path": [], "text": "The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial.", "negative_left": "In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial.", "negative_right": "The attached January 2026 pilot report for Mara's advanced-filtering request records 29% fewer filtering errors during that trial.", "right": "The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial."}, "verifier_independent_model": false}, "family": "fast-41-diverse-045-003", "id": "fast-41-diverse-045-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Case note — 14 March 2026: Mara submitted a request to add advanced filtering to the reporting module. The request explains why the enhancement is sought and states delivery effects of five additional weeks and $60,000. It is logged as a scope change, and the project sponsor is the designated recipient for that type of request. In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial. The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial. The request contains a rationale, a quantified impact, and an attached pilot report for review. The delivery lead supplied the schedule and cost estimates. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and question bindings remain unchanged in both full inputs. Base evidence is exactly two complete factual sentences: \"In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial.\" and \"The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial.\" Counterfactual evidence is exactly two complete factual sentences: \"In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial.\" and \"The attached January 2026 pilot report for Mara's advanced-filtering request records 29% fewer filtering errors during that trial.\" The counterfactual coherently changes the report measurement and creates an explicit paraphrase mismatch without contradictory duplicate assertions. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Case note — 14 March 2026: Mara submitted a request to add advanced filtering to the reporting module. The request explains why the enhancement is sought and states delivery effects of five additional weeks and $60,000. It is logged as a scope change, and the project sponsor is the designated recipient for that type of request. In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial. The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial. The request contains a rationale, a quantified impact, and an attached pilot report for review. The delivery lead supplied the schedule and cost estimates. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial."}, {"path": [], "text": "The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial.", "negative_left": "In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial.", "negative_right": "The attached January 2026 pilot report for Mara's advanced-filtering request records 29% fewer filtering errors during that trial.", "right": "The attached January 2026 pilot report for Mara's advanced-filtering request records 37% fewer filtering errors during that trial."}, "verifier_independent_model": false}, "family": "fast-41-diverse-045-003", "id": "fast-41-diverse-045-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Case note — 14 March 2026: Mara submitted a request to add advanced filtering to the reporting module. The request explains why the enhancement is sought and states delivery effects of five additional weeks and $60,000. It is logged as a scope change, and the project sponsor is the designated recipient for that type of request. In the 14 March 2026 version of Mara's advanced-filtering request, the evidence paraphrase states that the pilot report recorded 37% fewer filtering errors during the January 2026 trial. The attached January 2026 pilot report for Mara's advanced-filtering request records 29% fewer filtering errors during that trial. The request contains a rationale, a quantified impact, and an attached pilot report for review. The delivery lead supplied the schedule and cost estimates. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and the unchanged questions object preserves all question criteria. Entity, request, module, and decision bindings remain aligned with the original question. The two evidence spans are complete factual sentences: “In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent.” and “The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time.” The counterfactual coherently changes the report measurement to 24 percent, making the unchanged 31-percent paraphrase inaccurate without creating contradictory duplicate measurements. Neither context embeds an answer, answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara submitted a request to add advanced filtering to the reporting module, explaining the user need and listing five additional weeks and $60,000 in delivery impact. In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent. The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time. The request therefore contains a rationale and quantified impact, and the change alters the module's scope. The project manager is reviewing the request under the governing policy. The delivery lead's five-week schedule estimate and $60,000 cost estimate remain the only applicable delivery figures. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent."}, {"path": [], "text": "The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent.", "negative_left": "In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent.", "negative_right": "The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 24 percent reduction in report-generation time.", "right": "The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-045-012", "id": "fast-41-diverse-045-012-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara submitted a request to add advanced filtering to the reporting module, explaining the user need and listing five additional weeks and $60,000 in delivery impact. In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent. The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time. The request therefore contains a rationale and quantified impact, and the change alters the module's scope. The project manager is reviewing the request under the governing policy. The delivery lead's five-week schedule estimate and $60,000 cost estimate remain the only applicable delivery figures. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and the unchanged questions object preserves all question criteria. Entity, request, module, and decision bindings remain aligned with the original question. The two evidence spans are complete factual sentences: “In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent.” and “The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time.” The counterfactual coherently changes the report measurement to 24 percent, making the unchanged 31-percent paraphrase inaccurate without creating contradictory duplicate measurements. Neither context embeds an answer, answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara submitted a request to add advanced filtering to the reporting module, explaining the user need and listing five additional weeks and $60,000 in delivery impact. In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent. The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time. The request therefore contains a rationale and quantified impact, and the change alters the module's scope. The project manager is reviewing the request under the governing policy. The delivery lead's five-week schedule estimate and $60,000 cost estimate remain the only applicable delivery figures. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent."}, {"path": [], "text": "The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent.", "negative_left": "In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent.", "negative_right": "The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 24 percent reduction in report-generation time.", "right": "The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 31 percent reduction in report-generation time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-045-012", "id": "fast-41-diverse-045-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara submitted a request to add advanced filtering to the reporting module, explaining the user need and listing five additional weeks and $60,000 in delivery impact. In Mara's advanced-filtering request dated 14 March 2026, her paraphrase states that the attached pilot reduced report-generation time by 31 percent. The attached pilot report for Mara's advanced-filtering request, finalized 12 March 2026, records a 24 percent reduction in report-generation time. The request therefore contains a rationale and quantified impact, and the change alters the module's scope. The project manager is reviewing the request under the governing policy. The delivery lead's five-week schedule estimate and $60,000 cost estimate remain the only applicable delivery figures. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy and bindings; the evidence is factual and complete: “The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026.” and “The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026.”; the counterfactual changes only the schedule date to 29 June 2026 without contradiction, and neither context leaks an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The change packet explains that encryption and key rotation are required to address the payment export's field-level security gap. It records an acceptance-date movement of exactly 10 business days, an added cost of exactly 5% of the $300,000 baseline, and an assignment of two engineers.\"},{\"speaker\":\"Contract administrator\",\"text\":\"The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026.\"},{\"speaker\":\"Delivery lead\",\"text\":\"The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026. The schedule, cost, staffing, and rationale entries are attached to the review packet.\"},{\"speaker\":\"Project manager\",\"text\":\"The packet is complete for review, with the quantified schedule, cost, and resource entries cross-referenced to the request and its supporting rationale.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026."}, {"path": ["2", "text"], "text": "The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026.", "negative_left": "The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026.", "negative_right": "The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 29 June 2026.", "right": "The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-047-010", "id": "fast-41-diverse-047-010-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The change packet explains that encryption and key rotation are required to address the payment export's field-level security gap. It records an acceptance-date movement of exactly 10 business days, an added cost of exactly 5% of the $300,000 baseline, and an assignment of two engineers."}, {"speaker": "Contract administrator", "text": "The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026."}, {"speaker": "Delivery lead", "text": "The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026. The schedule, cost, staffing, and rationale entries are attached to the review packet."}, {"speaker": "Project manager", "text": "The packet is complete for review, with the quantified schedule, cost, and resource entries cross-referenced to the request and its supporting rationale."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy and bindings; the evidence is factual and complete: “The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026.” and “The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026.”; the counterfactual changes only the schedule date to 29 June 2026 without contradiction, and neither context leaks an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The change packet explains that encryption and key rotation are required to address the payment export's field-level security gap. It records an acceptance-date movement of exactly 10 business days, an added cost of exactly 5% of the $300,000 baseline, and an assignment of two engineers.\"},{\"speaker\":\"Contract administrator\",\"text\":\"The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026.\"},{\"speaker\":\"Delivery lead\",\"text\":\"The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026. The schedule, cost, staffing, and rationale entries are attached to the review packet.\"},{\"speaker\":\"Project manager\",\"text\":\"The packet is complete for review, with the quantified schedule, cost, and resource entries cross-referenced to the request and its supporting rationale.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026."}, {"path": ["2", "text"], "text": "The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026.", "negative_left": "The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026.", "negative_right": "The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 29 June 2026.", "right": "The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 1 July 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-047-010", "id": "fast-41-diverse-047-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The change packet explains that encryption and key rotation are required to address the payment export's field-level security gap. It records an acceptance-date movement of exactly 10 business days, an added cost of exactly 5% of the $300,000 baseline, and an assignment of two engineers."}, {"speaker": "Contract administrator", "text": "The signed contract for the encryption-and-key-rotation change request identifies the security-readiness milestone as its sole contractual milestone and requires completion by 30 June 2026."}, {"speaker": "Delivery lead", "text": "The approved project schedule for the encryption-and-key-rotation change request records completion of the security-readiness milestone on 29 June 2026. The schedule, cost, staffing, and rationale entries are attached to the review packet."}, {"speaker": "Project manager", "text": "The packet is complete for review, with the quantified schedule, cost, and resource entries cross-referenced to the request and its supporting rationale."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The two exact evidence spans are “The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver.” and “A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026.”; the counterfactual coherently removes that signature without duplicate measurements, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\":\"Nimbus payroll-export package case note, 17 September 2026. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing. The submitted test report records 20 of 20 tests passed, and the defect register records zero open Severity 1 or 2 defects. The submitted acceptance evidence contains a signed Quality reviewer test attestation dated 14 September 2026. The defect log records no remaining cosmetic defects. The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver. A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026. The packet is being reviewed for final acceptance and administrative routing.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver."}, {"path": ["context"], "text": "A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver.", "negative_left": "The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver.", "negative_right": "A page in that packet bears no signature from Dana Ortiz and the date 17 September 2026.", "right": "A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-056-003", "id": "fast-41-diverse-056-003-base", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "Nimbus payroll-export package case note, 17 September 2026. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing. The submitted test report records 20 of 20 tests passed, and the defect register records zero open Severity 1 or 2 defects. The submitted acceptance evidence contains a signed Quality reviewer test attestation dated 14 September 2026. The defect log records no remaining cosmetic defects. The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver. A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026. The packet is being reviewed for final acceptance and administrative routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_high"}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The two exact evidence spans are “The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver.” and “A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026.”; the counterfactual coherently removes that signature without duplicate measurements, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\":\"Nimbus payroll-export package case note, 17 September 2026. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing. The submitted test report records 20 of 20 tests passed, and the defect register records zero open Severity 1 or 2 defects. The submitted acceptance evidence contains a signed Quality reviewer test attestation dated 14 September 2026. The defect log records no remaining cosmetic defects. The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver. A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026. The packet is being reviewed for final acceptance and administrative routing.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver."}, {"path": ["context"], "text": "A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver.", "negative_left": "The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver.", "negative_right": "A page in that packet bears no signature from Dana Ortiz and the date 17 September 2026.", "right": "A page in that packet bears Dana Ortiz’s handwritten signature and the date 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-056-003", "id": "fast-41-diverse-056-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "Nimbus payroll-export package case note, 17 September 2026. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing. The submitted test report records 20 of 20 tests passed, and the defect register records zero open Severity 1 or 2 defects. The submitted acceptance evidence contains a signed Quality reviewer test attestation dated 14 September 2026. The defect log records no remaining cosmetic defects. The 17 September 2026 Nimbus payroll-export acceptance packet contains a document titled “Business approval” that names Dana Ortiz as the designated Business approver. A page in that packet bears no signature from Dana Ortiz and the date 17 September 2026. The packet is being reviewed for final acceptance and administrative routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "hold_near_complete_business_approver"}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the policy and Nimbus package bindings without answer leakage; the two complete factual evidence sentences are “The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature.” and “The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package.”; the counterfactual coherently changes Jordan Lee's role to Quality reviewer, making the written record non-Business approval without contradictory duplicate assertions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\":\"The 2026 Nimbus payroll-export acceptance packet includes a submitted test report showing 20 of 20 tests passed, a defect register showing zero open Severity 1 or 2 defects, and a signed Quality reviewer test attestation. The packet's cosmetic-review log records no remaining cosmetic defects. The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature. The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature."}, {"path": ["context"], "text": "The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature.", "negative_left": "The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature.", "negative_right": "The archived personnel register identifies Jordan Lee as the Quality reviewer rather than the Business approver for the Nimbus payroll-export package.", "right": "The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package."}, "verifier_independent_model": false}, "family": "fast-41-diverse-056-012", "id": "fast-41-diverse-056-012-base", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "The 2026 Nimbus payroll-export acceptance packet includes a submitted test report showing 20 of 20 tests passed, a defect register showing zero open Severity 1 or 2 defects, and a signed Quality reviewer test attestation. The packet's cosmetic-review log records no remaining cosmetic defects. The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature. The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_high"}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the policy and Nimbus package bindings without answer leakage; the two complete factual evidence sentences are “The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature.” and “The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package.”; the counterfactual coherently changes Jordan Lee's role to Quality reviewer, making the written record non-Business approval without contradictory duplicate assertions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\":\"The 2026 Nimbus payroll-export acceptance packet includes a submitted test report showing 20 of 20 tests passed, a defect register showing zero open Severity 1 or 2 defects, and a signed Quality reviewer test attestation. The packet's cosmetic-review log records no remaining cosmetic defects. The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature. The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature."}, {"path": ["context"], "text": "The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature.", "negative_left": "The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature.", "negative_right": "The archived personnel register identifies Jordan Lee as the Quality reviewer rather than the Business approver for the Nimbus payroll-export package.", "right": "The archived personnel register identifies Jordan Lee as the Business approver for the Nimbus payroll-export package."}, "verifier_independent_model": false}, "family": "fast-41-diverse-056-012", "id": "fast-41-diverse-056-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "The 2026 Nimbus payroll-export acceptance packet includes a submitted test report showing 20 of 20 tests passed, a defect register showing zero open Severity 1 or 2 defects, and a signed Quality reviewer test attestation. The packet's cosmetic-review log records no remaining cosmetic defects. The 2026 acceptance packet for the Nimbus payroll-export package contains exactly one written approval record bearing Jordan Lee's signature. The archived personnel register identifies Jordan Lee as the Quality reviewer rather than the Business approver for the Nimbus payroll-export package. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "hold_near_complete_business_approver"}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions preserve the governing policy and bindings, both evidence spans are complete factual sentences, the counterfactual changes only approval status consistently, and neither context contains prohibited answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14.\"},{\"speaker\":\"Approval record\",\"text\":\"The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as approved.\"},{\"speaker\":\"Test report\",\"text\":\"The delivery packet records all 12 acceptance tests as passed, with the test results linked to the submitted work package.\"},{\"speaker\":\"Quality review\",\"text\":\"Quality reviewer Priya Shah signed QR-17 after checking that the evidence matched WP-17.\"},{\"speaker\":\"Evidence index\",\"text\":\"The technical evidence index is complete, including the implementation notes, test report, and review attachments.\"},{\"speaker\":\"Defect log\",\"text\":\"The defect log was closed with no remaining defects assigned to WP-17.\"},{\"speaker\":\"Policy\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Policy\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["0", "text"], "text": "For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14."}, {"path": ["1", "text"], "text": "The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as approved."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14.", "negative_left": "For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14.", "negative_right": "The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as not approved.", "right": "The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as approved."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-001", "id": "fast-41-diverse-057-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14."}, {"speaker": "Approval record", "text": "The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as approved."}, {"speaker": "Test report", "text": "The delivery packet records all 12 acceptance tests as passed, with the test results linked to the submitted work package."}, {"speaker": "Quality review", "text": "Quality reviewer Priya Shah signed QR-17 after checking that the evidence matched WP-17."}, {"speaker": "Evidence index", "text": "The technical evidence index is complete, including the implementation notes, test report, and review attachments."}, {"speaker": "Defect log", "text": "The defect log was closed with no remaining defects assigned to WP-17."}, {"speaker": "Policy", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Policy", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions preserve the governing policy and bindings, both evidence spans are complete factual sentences, the counterfactual changes only approval status consistently, and neither context contains prohibited answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14.\"},{\"speaker\":\"Approval record\",\"text\":\"The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as approved.\"},{\"speaker\":\"Test report\",\"text\":\"The delivery packet records all 12 acceptance tests as passed, with the test results linked to the submitted work package.\"},{\"speaker\":\"Quality review\",\"text\":\"Quality reviewer Priya Shah signed QR-17 after checking that the evidence matched WP-17.\"},{\"speaker\":\"Evidence index\",\"text\":\"The technical evidence index is complete, including the implementation notes, test report, and review attachments.\"},{\"speaker\":\"Defect log\",\"text\":\"The defect log was closed with no remaining defects assigned to WP-17.\"},{\"speaker\":\"Policy\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Policy\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["0", "text"], "text": "For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14."}, {"path": ["1", "text"], "text": "The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as approved."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14.", "negative_left": "For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14.", "negative_right": "The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as not approved.", "right": "The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as approved."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-001", "id": "fast-41-diverse-057-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "For WP-17, Morgan Lee is the assigned Business approver as of 2026-09-14."}, {"speaker": "Approval record", "text": "The approval record for WP-17, filed on 2026-09-14, lists Morgan Lee's decision as not approved."}, {"speaker": "Test report", "text": "The delivery packet records all 12 acceptance tests as passed, with the test results linked to the submitted work package."}, {"speaker": "Quality review", "text": "Quality reviewer Priya Shah signed QR-17 after checking that the evidence matched WP-17."}, {"speaker": "Evidence index", "text": "The technical evidence index is complete, including the implementation notes, test report, and review attachments."}, {"speaker": "Defect log", "text": "The defect log was closed with no remaining defects assigned to WP-17."}, {"speaker": "Policy", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Policy", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the WP-17 scope, policy, and bindings; the two evidence quotes are complete factual sentences. The counterfactual changes approval to explicit denial without contradiction, and neither context embeds a gold answer or classifier label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Delivery owner\",\"text\":\"The WP-17 readiness packet records that all 12 acceptance tests passed. Its technical evidence set is complete, and the defect register shows zero remaining defects.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"I checked the WP-17 verification packet, confirmed that it matches the work product, and recorded Quality reviewer sign-off.\"},{\"speaker\":\"Case note\",\"text\":\"As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz.\"},{\"speaker\":\"Case note\",\"text\":\"On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly approved WP-17.\"},{\"speaker\":\"Coordinator\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Coordinator\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["2", "text"], "text": "As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz."}, {"path": ["3", "text"], "text": "On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly approved WP-17."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz.", "negative_left": "As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz.", "negative_right": "On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly denied approval for WP-17.", "right": "On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly approved WP-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-003", "id": "fast-41-diverse-057-003-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Delivery owner", "text": "The WP-17 readiness packet records that all 12 acceptance tests passed. Its technical evidence set is complete, and the defect register shows zero remaining defects."}, {"speaker": "Quality reviewer", "text": "I checked the WP-17 verification packet, confirmed that it matches the work product, and recorded Quality reviewer sign-off."}, {"speaker": "Case note", "text": "As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz."}, {"speaker": "Case note", "text": "On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly approved WP-17."}, {"speaker": "Coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the WP-17 scope, policy, and bindings; the two evidence quotes are complete factual sentences. The counterfactual changes approval to explicit denial without contradiction, and neither context embeds a gold answer or classifier label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Delivery owner\",\"text\":\"The WP-17 readiness packet records that all 12 acceptance tests passed. Its technical evidence set is complete, and the defect register shows zero remaining defects.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"I checked the WP-17 verification packet, confirmed that it matches the work product, and recorded Quality reviewer sign-off.\"},{\"speaker\":\"Case note\",\"text\":\"As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz.\"},{\"speaker\":\"Case note\",\"text\":\"On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly approved WP-17.\"},{\"speaker\":\"Coordinator\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Coordinator\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["2", "text"], "text": "As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz."}, {"path": ["3", "text"], "text": "On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly approved WP-17."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz.", "negative_left": "As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz.", "negative_right": "On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly denied approval for WP-17.", "right": "On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly approved WP-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-003", "id": "fast-41-diverse-057-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Delivery owner", "text": "The WP-17 readiness packet records that all 12 acceptance tests passed. Its technical evidence set is complete, and the defect register shows zero remaining defects."}, {"speaker": "Quality reviewer", "text": "I checked the WP-17 verification packet, confirmed that it matches the work product, and recorded Quality reviewer sign-off."}, {"speaker": "Case note", "text": "As of 2026-08-14, WP-17's assigned Business approver was Dana Ruiz."}, {"speaker": "Case note", "text": "On 2026-08-14, the signed decision record for WP-17 states that Dana Ruiz explicitly denied approval for WP-17."}, {"speaker": "Coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and WP-17 bindings; the exact evidence quotes are “WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204.” and “Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17.”; the counterfactual coherently changes approval to refusal without contradictions or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Delivery owner\",\"text\":\"WP-17's release file contains the final test report, which records all 12 acceptance tests as passed. The defect ledger is closed, and the technical evidence packet has been marked complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"I checked the evidence packet against the implementation and signed the quality review record for WP-17.\"},{\"speaker\":\"Records coordinator\",\"text\":\"WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204.\"},{\"speaker\":\"Records coordinator\",\"text\":\"Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17.\"},{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["2", "text"], "text": "WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204."}, {"path": ["3", "text"], "text": "Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204.", "negative_left": "WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204.", "negative_right": "Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit refusal to approve WP-17.", "right": "Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-004", "id": "fast-41-diverse-057-004-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Delivery owner", "text": "WP-17's release file contains the final test report, which records all 12 acceptance tests as passed. The defect ledger is closed, and the technical evidence packet has been marked complete."}, {"speaker": "Quality reviewer", "text": "I checked the evidence packet against the implementation and signed the quality review record for WP-17."}, {"speaker": "Records coordinator", "text": "WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204."}, {"speaker": "Records coordinator", "text": "Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17."}, {"speaker": "Project coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and WP-17 bindings; the exact evidence quotes are “WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204.” and “Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17.”; the counterfactual coherently changes approval to refusal without contradictions or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Delivery owner\",\"text\":\"WP-17's release file contains the final test report, which records all 12 acceptance tests as passed. The defect ledger is closed, and the technical evidence packet has been marked complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"I checked the evidence packet against the implementation and signed the quality review record for WP-17.\"},{\"speaker\":\"Records coordinator\",\"text\":\"WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204.\"},{\"speaker\":\"Records coordinator\",\"text\":\"Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17.\"},{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["2", "text"], "text": "WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204."}, {"path": ["3", "text"], "text": "Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204.", "negative_left": "WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204.", "negative_right": "Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit refusal to approve WP-17.", "right": "Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit approval of WP-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-004", "id": "fast-41-diverse-057-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Delivery owner", "text": "WP-17's release file contains the final test report, which records all 12 acceptance tests as passed. The defect ledger is closed, and the technical evidence packet has been marked complete."}, {"speaker": "Quality reviewer", "text": "I checked the evidence packet against the implementation and signed the quality review record for WP-17."}, {"speaker": "Records coordinator", "text": "WP-17's assigned Business approver is Dana Ortiz, and the approval record for WP-17 is document BA-204."}, {"speaker": "Records coordinator", "text": "Document BA-204, timestamped 14 March 2026 at 09:20 UTC, records Dana Ortiz's explicit refusal to approve WP-17."}, {"speaker": "Project coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both contexts, while the unchanged original questions retain the decision instructions and criteria. Question bindings remain WP-17, the 14 March 2026 date, Dana Lee, and the Business approver path. The evidence contains two complete factual sentences: \"The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver.\" and \"The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved.\" The counterfactual coherently changes the approval status from approved to pending without contradictory duplicate assertions. Neither context embeds a gold answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"WP-17's acceptance packet contains complete technical evidence, and its test report records all 12 acceptance tests as passed. The defect register shows zero remaining defects.\"},{\"speaker\":\"Assignment record\",\"text\":\"The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver.\"},{\"speaker\":\"Approval ledger\",\"text\":\"The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The Quality reviewer checked the supporting materials and signed the review record for WP-17.\"},{\"speaker\":\"Policy record\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Policy record\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["1", "text"], "text": "The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver."}, {"path": ["2", "text"], "text": "The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver.", "negative_left": "The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver.", "negative_right": "The 14 March 2026 approval ledger for WP-17 records Dana Lee with status pending.", "right": "The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-008", "id": "fast-41-diverse-057-008-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "WP-17's acceptance packet contains complete technical evidence, and its test report records all 12 acceptance tests as passed. The defect register shows zero remaining defects."}, {"speaker": "Assignment record", "text": "The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver."}, {"speaker": "Approval ledger", "text": "The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved."}, {"speaker": "Quality reviewer", "text": "The Quality reviewer checked the supporting materials and signed the review record for WP-17."}, {"speaker": "Policy record", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Policy record", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both contexts, while the unchanged original questions retain the decision instructions and criteria. Question bindings remain WP-17, the 14 March 2026 date, Dana Lee, and the Business approver path. The evidence contains two complete factual sentences: \"The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver.\" and \"The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved.\" The counterfactual coherently changes the approval status from approved to pending without contradictory duplicate assertions. Neither context embeds a gold answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"WP-17's acceptance packet contains complete technical evidence, and its test report records all 12 acceptance tests as passed. The defect register shows zero remaining defects.\"},{\"speaker\":\"Assignment record\",\"text\":\"The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver.\"},{\"speaker\":\"Approval ledger\",\"text\":\"The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The Quality reviewer checked the supporting materials and signed the review record for WP-17.\"},{\"speaker\":\"Policy record\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Policy record\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["1", "text"], "text": "The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver."}, {"path": ["2", "text"], "text": "The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver.", "negative_left": "The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver.", "negative_right": "The 14 March 2026 approval ledger for WP-17 records Dana Lee with status pending.", "right": "The 14 March 2026 approval ledger for WP-17 records Dana Lee with status approved."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-008", "id": "fast-41-diverse-057-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "WP-17's acceptance packet contains complete technical evidence, and its test report records all 12 acceptance tests as passed. The defect register shows zero remaining defects."}, {"speaker": "Assignment record", "text": "The 14 March 2026 assignment record names Dana Lee as WP-17's assigned Business approver."}, {"speaker": "Approval ledger", "text": "The 14 March 2026 approval ledger for WP-17 records Dana Lee with status pending."}, {"speaker": "Quality reviewer", "text": "The Quality reviewer checked the supporting materials and signed the review record for WP-17."}, {"speaker": "Policy record", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Policy record", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and the WP-17, acceptance-now, and Complete bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"The assigned Business approver for WP-17 is Dana Lee, employee ID B-482.\" and \"The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval.\" The counterfactual changes approval to rejection without creating contradictory duplicate assertions. Neither context embeds a gold answer, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"WP-17 is under review for the current release. The assigned Business approver for WP-17 is Dana Lee, employee ID B-482.\"},{\"speaker\":\"Validation lead\",\"text\":\"The final validation report records all 12 acceptance tests as passed, and the technical evidence package is complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"Quality review QR-17 is signed off, and the defect register shows zero remaining defects.\"},{\"speaker\":\"Decision record\",\"text\":\"The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval.\"},{\"speaker\":\"Release coordinator\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Policy note\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["0", "text"], "text": "The assigned Business approver for WP-17 is Dana Lee, employee ID B-482."}, {"path": ["3", "text"], "text": "The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "The assigned Business approver for WP-17 is Dana Lee, employee ID B-482.", "negative_left": "The assigned Business approver for WP-17 is Dana Lee, employee ID B-482.", "negative_right": "The WP-17 decision log dated 2026-09-15 records Dana Lee's rejection of WP-17.", "right": "The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-011", "id": "fast-41-diverse-057-011-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "WP-17 is under review for the current release. The assigned Business approver for WP-17 is Dana Lee, employee ID B-482."}, {"speaker": "Validation lead", "text": "The final validation report records all 12 acceptance tests as passed, and the technical evidence package is complete."}, {"speaker": "Quality reviewer", "text": "Quality review QR-17 is signed off, and the defect register shows zero remaining defects."}, {"speaker": "Decision record", "text": "The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval."}, {"speaker": "Release coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Policy note", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and the WP-17, acceptance-now, and Complete bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"The assigned Business approver for WP-17 is Dana Lee, employee ID B-482.\" and \"The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval.\" The counterfactual changes approval to rejection without creating contradictory duplicate assertions. Neither context embeds a gold answer, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"WP-17 is under review for the current release. The assigned Business approver for WP-17 is Dana Lee, employee ID B-482.\"},{\"speaker\":\"Validation lead\",\"text\":\"The final validation report records all 12 acceptance tests as passed, and the technical evidence package is complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"Quality review QR-17 is signed off, and the defect register shows zero remaining defects.\"},{\"speaker\":\"Decision record\",\"text\":\"The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval.\"},{\"speaker\":\"Release coordinator\",\"text\":\"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Policy note\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["0", "text"], "text": "The assigned Business approver for WP-17 is Dana Lee, employee ID B-482."}, {"path": ["3", "text"], "text": "The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "The assigned Business approver for WP-17 is Dana Lee, employee ID B-482.", "negative_left": "The assigned Business approver for WP-17 is Dana Lee, employee ID B-482.", "negative_right": "The WP-17 decision log dated 2026-09-15 records Dana Lee's rejection of WP-17.", "right": "The WP-17 decision log dated 2026-09-15 records Dana Lee's explicit approval."}, "verifier_independent_model": false}, "family": "fast-41-diverse-057-011", "id": "fast-41-diverse-057-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "WP-17 is under review for the current release. The assigned Business approver for WP-17 is Dana Lee, employee ID B-482."}, {"speaker": "Validation lead", "text": "The final validation report records all 12 acceptance tests as passed, and the technical evidence package is complete."}, {"speaker": "Quality reviewer", "text": "Quality review QR-17 is signed off, and the defect register shows zero remaining defects."}, {"speaker": "Decision record", "text": "The WP-17 decision log dated 2026-09-15 records Dana Lee's rejection of WP-17."}, {"speaker": "Release coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Policy note", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the acceptance, testing, sign-off, conditional-intent, and quality-rating policy. The Atlas package, acceptance decision, and High-rating bindings remain unchanged. The two evidence spans are complete factual sentences: \"The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time.\" \"Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time.\" The counterfactual changes Jordan Vale's role consistently and introduces no contradictory duplicate measurement or assertion. Neither context embeds a gold answer, answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"At the acceptance decision time, all 12 tests specified for the Atlas reporting work package had passing status, and the review record noted no remaining defect. Finance had approved the Atlas totals. The Delivery owner had expressed intent to submit the package for acceptance if Finance approved the totals, and that condition was satisfied. A dated signed approval record for the Atlas totals existed. The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time. Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time. No person other than the signer recorded on the dated approval of the Atlas totals had provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time.\\n\\nAcceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met.\\nRate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time."}, {"path": [], "text": "Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time.", "negative_left": "The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time.", "negative_right": "Jordan Vale did not hold the Business approver role for the Atlas reporting work package at the acceptance decision time.", "right": "Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-058-002", "id": "fast-41-diverse-058-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "At the acceptance decision time, all 12 tests specified for the Atlas reporting work package had passing status, and the review record noted no remaining defect. Finance had approved the Atlas totals. The Delivery owner had expressed intent to submit the package for acceptance if Finance approved the totals, and that condition was satisfied. A dated signed approval record for the Atlas totals existed. The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time. Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time. No person other than the signer recorded on the dated approval of the Atlas totals had provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time.\n\nAcceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met.\nRate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the acceptance, testing, sign-off, conditional-intent, and quality-rating policy. The Atlas package, acceptance decision, and High-rating bindings remain unchanged. The two evidence spans are complete factual sentences: \"The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time.\" \"Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time.\" The counterfactual changes Jordan Vale's role consistently and introduces no contradictory duplicate measurement or assertion. Neither context embeds a gold answer, answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"At the acceptance decision time, all 12 tests specified for the Atlas reporting work package had passing status, and the review record noted no remaining defect. Finance had approved the Atlas totals. The Delivery owner had expressed intent to submit the package for acceptance if Finance approved the totals, and that condition was satisfied. A dated signed approval record for the Atlas totals existed. The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time. Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time. No person other than the signer recorded on the dated approval of the Atlas totals had provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time.\\n\\nAcceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met.\\nRate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time."}, {"path": [], "text": "Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time.", "negative_left": "The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time.", "negative_right": "Jordan Vale did not hold the Business approver role for the Atlas reporting work package at the acceptance decision time.", "right": "Jordan Vale held the Business approver role for the Atlas reporting work package at the acceptance decision time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-058-002", "id": "fast-41-diverse-058-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "At the acceptance decision time, all 12 tests specified for the Atlas reporting work package had passing status, and the review record noted no remaining defect. Finance had approved the Atlas totals. The Delivery owner had expressed intent to submit the package for acceptance if Finance approved the totals, and that condition was satisfied. A dated signed approval record for the Atlas totals existed. The dated approval record for the Atlas totals lists Jordan Vale as its signer at the acceptance decision time. Jordan Vale did not hold the Business approver role for the Atlas reporting work package at the acceptance decision time. No person other than the signer recorded on the dated approval of the Atlas totals had provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time.\n\nAcceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met.\nRate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, policy, Orion package, final-build scope, and routing request; all evidence entries are complete factual sentences, and the counterfactual changes only the test-status observation without contradiction or embedded answer instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A2's quantification over the explicit set of 20 tests does not make it a bundle. A1 is a factual test-result proposition rather than a policy conclusion. The base and counter assignments are realizable with only A1 changing: the base can have one or more recorded failures, while the counter can have all recorded results be passes, with current-build approval and no evidence or documentation issues in both. Policy evidence preserves the state-originated final-build acceptance requirement, role routing, and newer-build supersession rule. Rules appearing only in the retained questions object need not be duplicated. The rule table may validly abstain on moderate-completion scenarios.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A1 entails that at least one mandatory test failed on the final build. The explicit criterion assigns level 0 whenever any mandatory final-build test fails, regardless of the other atoms.", "rule_index": 0, "sound": true}, {"reason": "Refuted A1 excludes every failed mandatory test, while supported A2 establishes that every mandatory test has a pass-or-fail result; together these entail that all 20 passed. A3 establishes approval naming the final build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions are sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At least one of the 20 mandatory tests failed on the build submitted as final for the Orion payroll-export package."}, {"id": "A2", "statement": "Each of the 20 mandatory tests has a recorded pass-or-fail result on the build submitted as final for the Orion payroll-export package."}, {"id": "A3", "statement": "The business approval names the build submitted as final for the Orion payroll-export package."}, {"id": "A4", "statement": "An unresolved evidence issue remains for the build submitted as final for the Orion payroll-export package."}, {"id": "A5", "statement": "An unresolved documentation issue remains for the build submitted as final for the Orion payroll-export package."}], "base_state_json": "{\"context\":\"Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds. The review concerns the Orion payroll-export package, whose submitted final build has a complete set of twenty recorded test-result entries. The business approval on file names that submitted final build. The evidence register shows no unresolved evidence issue for it, and the documentation register shows no unresolved documentation issue. The latest-build record is the controlling record for this review.\",\"evidence\":[\"On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue.\",\"For the Orion payroll-export package, the result status Tally-Blue denotes a failed test.\",\"The business approval on file names the build submitted as final for the Orion payroll-export package.\",\"The evidence register lists no unresolved evidence issue for the build submitted as final for the Orion payroll-export package.\",\"The documentation register lists no unresolved documentation issue for the build submitted as final for the Orion payroll-export package.\"],\"request\":\"Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue."}, {"path": ["evidence", "1"], "text": "For the Orion payroll-export package, the result status Tally-Blue denotes a failed test."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build."}, {"path": ["context"], "text": "The delivery owner fixes product defects; the business approver renews outdated approvals."}, {"path": ["context"], "text": "Evidence for a newer build supersedes evidence for earlier builds."}, {"path": ["request"], "text": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}], "rules": [{"justification": "A supported A1 establishes a failed mandatory test on the final build. The low-completion criterion therefore requires non-acceptance, and the failed-test blocker is routed to the delivery owner because that role fixes product defects.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}]}, {"justification": "A2 establishes a pass-or-fail result for every mandatory test, while refuted A1 excludes any failed result, so all 20 mandatory tests pass on the final build. A3 supplies approval naming that build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions satisfy high completion.", "target": "2", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue.", "negative_left": "On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue.", "negative_right": "For the Orion payroll-export package, the result status Tally-Blue denotes a passed test.", "right": "For the Orion payroll-export package, the result status Tally-Blue denotes a failed test."}, "verifier_independent_model": false}, "family": "fast-41-diverse-059-022", "id": "fast-41-diverse-059-022-base", "input": {"questions": {"decision": {"criteria": ["0 — Low completion: Do not accept. At least one mandatory final-build test fails or required final-build approval is missing; route each blocker to its responsible role.", "1 — Moderate completion: Conditionally accept only when all mandatory final-build tests pass and approval exists, but minor non-blocking documentation corrections remain.", "2 — High completion: Accept fully. Every mandatory test passes on the final build, the business approval names that build, and no unresolved evidence or documentation issues remain."], "instructions": "Apply the latest-build rule and the acceptance criteria. Choose exactly one ordered level. Identify acceptance status and route any failed test or missing current-build approval.", "type": "score"}}, "state": {"context": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds. The review concerns the Orion payroll-export package, whose submitted final build has a complete set of twenty recorded test-result entries. The business approval on file names that submitted final build. The evidence register shows no unresolved evidence issue for it, and the documentation register shows no unresolved documentation issue. The latest-build record is the controlling record for this review.", "evidence": ["On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue.", "For the Orion payroll-export package, the result status Tally-Blue denotes a failed test.", "The business approval on file names the build submitted as final for the Orion payroll-export package.", "The evidence register lists no unresolved evidence issue for the build submitted as final for the Orion payroll-export package.", "The documentation register lists no unresolved documentation issue for the build submitted as final for the Orion payroll-export package."], "request": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}}, "method": "c2d", "provenance": {"source_id": "diverse-059", "source_is_synthetic": true, "source_sha256": "f8f6806cfdd18388512db0984c3eb1386d711a91bbc8e11806a11e30573c54e8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, policy, Orion package, final-build scope, and routing request; all evidence entries are complete factual sentences, and the counterfactual changes only the test-status observation without contradiction or embedded answer instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; A2's quantification over the explicit set of 20 tests does not make it a bundle. A1 is a factual test-result proposition rather than a policy conclusion. The base and counter assignments are realizable with only A1 changing: the base can have one or more recorded failures, while the counter can have all recorded results be passes, with current-build approval and no evidence or documentation issues in both. Policy evidence preserves the state-originated final-build acceptance requirement, role routing, and newer-build supersession rule. Rules appearing only in the retained questions object need not be duplicated. The rule table may validly abstain on moderate-completion scenarios.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A1 entails that at least one mandatory test failed on the final build. The explicit criterion assigns level 0 whenever any mandatory final-build test fails, regardless of the other atoms.", "rule_index": 0, "sound": true}, {"reason": "Refuted A1 excludes every failed mandatory test, while supported A2 establishes that every mandatory test has a pass-or-fail result; together these entail that all 20 passed. A3 establishes approval naming the final build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions are sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At least one of the 20 mandatory tests failed on the build submitted as final for the Orion payroll-export package."}, {"id": "A2", "statement": "Each of the 20 mandatory tests has a recorded pass-or-fail result on the build submitted as final for the Orion payroll-export package."}, {"id": "A3", "statement": "The business approval names the build submitted as final for the Orion payroll-export package."}, {"id": "A4", "statement": "An unresolved evidence issue remains for the build submitted as final for the Orion payroll-export package."}, {"id": "A5", "statement": "An unresolved documentation issue remains for the build submitted as final for the Orion payroll-export package."}], "base_state_json": "{\"context\":\"Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds. The review concerns the Orion payroll-export package, whose submitted final build has a complete set of twenty recorded test-result entries. The business approval on file names that submitted final build. The evidence register shows no unresolved evidence issue for it, and the documentation register shows no unresolved documentation issue. The latest-build record is the controlling record for this review.\",\"evidence\":[\"On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue.\",\"For the Orion payroll-export package, the result status Tally-Blue denotes a failed test.\",\"The business approval on file names the build submitted as final for the Orion payroll-export package.\",\"The evidence register lists no unresolved evidence issue for the build submitted as final for the Orion payroll-export package.\",\"The documentation register lists no unresolved documentation issue for the build submitted as final for the Orion payroll-export package.\"],\"request\":\"Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue."}, {"path": ["evidence", "1"], "text": "For the Orion payroll-export package, the result status Tally-Blue denotes a failed test."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build."}, {"path": ["context"], "text": "The delivery owner fixes product defects; the business approver renews outdated approvals."}, {"path": ["context"], "text": "Evidence for a newer build supersedes evidence for earlier builds."}, {"path": ["request"], "text": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}], "rules": [{"justification": "A supported A1 establishes a failed mandatory test on the final build. The low-completion criterion therefore requires non-acceptance, and the failed-test blocker is routed to the delivery owner because that role fixes product defects.", "target": "0", "when": [{"atom_id": "A1", "state": "supported"}]}, {"justification": "A2 establishes a pass-or-fail result for every mandatory test, while refuted A1 excludes any failed result, so all 20 mandatory tests pass on the final build. A3 supplies approval naming that build, and refuted A4 and A5 exclude unresolved evidence and documentation issues. These conditions satisfy high completion.", "target": "2", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue.", "negative_left": "On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue.", "negative_right": "For the Orion payroll-export package, the result status Tally-Blue denotes a passed test.", "right": "For the Orion payroll-export package, the result status Tally-Blue denotes a failed test."}, "verifier_independent_model": false}, "family": "fast-41-diverse-059-022", "id": "fast-41-diverse-059-022-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low completion: Do not accept. At least one mandatory final-build test fails or required final-build approval is missing; route each blocker to its responsible role.", "1 — Moderate completion: Conditionally accept only when all mandatory final-build tests pass and approval exists, but minor non-blocking documentation corrections remain.", "2 — High completion: Accept fully. Every mandatory test passes on the final build, the business approval names that build, and no unresolved evidence or documentation issues remain."], "instructions": "Apply the latest-build rule and the acceptance criteria. Choose exactly one ordered level. Identify acceptance status and route any failed test or missing current-build approval.", "type": "score"}}, "state": {"context": "Acceptance requires all 20 tests to pass on the final build and a business approval naming that same build. The delivery owner fixes product defects; the business approver renews outdated approvals. Evidence for a newer build supersedes evidence for earlier builds. The review concerns the Orion payroll-export package, whose submitted final build has a complete set of twenty recorded test-result entries. The business approval on file names that submitted final build. The evidence register shows no unresolved evidence issue for it, and the documentation register shows no unresolved documentation issue. The latest-build record is the controlling record for this review.", "evidence": ["On the build submitted as final for the Orion payroll-export package, each of the 20 mandatory tests has the recorded result status Tally-Blue.", "For the Orion payroll-export package, the result status Tally-Blue denotes a passed test.", "The business approval on file names the build submitted as final for the Orion payroll-export package.", "The evidence register lists no unresolved evidence issue for the build submitted as final for the Orion payroll-export package.", "The documentation register lists no unresolved documentation issue for the build submitted as final for the Orion payroll-export package."], "request": "Select the completion-quality level, state whether to accept the deliverable, and route each unresolved item to the responsible role."}}, "method": "c2d", "provenance": {"source_id": "diverse-059", "source_is_synthetic": true, "source_sha256": "f8f6806cfdd18388512db0984c3eb1386d711a91bbc8e11806a11e30573c54e8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings without answer leakage; the evidence quotes are complete factual sentences: “The approval scope recorded for business-approval artifact BA-1 is release package R-17.” and “The build 4.2 deliverable is release package R-17, and it is not release package R-18.”; the counterfactual coherently changes BA-1’s scope to R-18 while retaining the R-17 deliverable distinction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The verified run record for build 4.2 shows the specified CSV-export test failing. Each of the other 19 specified tests has a verified passing result. QA approval remains unsigned because no rerun of the failed test is attached.\"},{\"speaker\":\"Business approver\",\"text\":\"Submitted business-approval artifact BA-1 bears my signature. The approval scope recorded for business-approval artifact BA-1 is release package R-17. The build 4.2 deliverable is release package R-17, and it is not release package R-18. No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable.\"},{\"speaker\":\"Delivery owner\",\"text\":\"I am responsible for correcting or rerunning the CSV-export test and supplying the resulting evidence for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "The approval scope recorded for business-approval artifact BA-1 is release package R-17."}, {"path": ["2", "text"], "text": "The build 4.2 deliverable is release package R-17, and it is not release package R-18."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The approval scope recorded for business-approval artifact BA-1 is release package R-17.", "negative_left": "The approval scope recorded for business-approval artifact BA-1 is release package R-18.", "negative_right": "The build 4.2 deliverable is release package R-17, and it is not release package R-18.", "right": "The build 4.2 deliverable is release package R-17, and it is not release package R-18."}, "verifier_independent_model": false}, "family": "fast-41-diverse-060-009", "id": "fast-41-diverse-060-009-base", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Project coordinator", "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}, {"speaker": "Quality reviewer", "text": "The verified run record for build 4.2 shows the specified CSV-export test failing. Each of the other 19 specified tests has a verified passing result. QA approval remains unsigned because no rerun of the failed test is attached."}, {"speaker": "Business approver", "text": "Submitted business-approval artifact BA-1 bears my signature. The approval scope recorded for business-approval artifact BA-1 is release package R-17. The build 4.2 deliverable is release package R-17, and it is not release package R-18. No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}, {"speaker": "Delivery owner", "text": "I am responsible for correcting or rerunning the CSV-export test and supplying the resulting evidence for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings without answer leakage; the evidence quotes are complete factual sentences: “The approval scope recorded for business-approval artifact BA-1 is release package R-17.” and “The build 4.2 deliverable is release package R-17, and it is not release package R-18.”; the counterfactual coherently changes BA-1’s scope to R-18 while retaining the R-17 deliverable distinction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The verified run record for build 4.2 shows the specified CSV-export test failing. Each of the other 19 specified tests has a verified passing result. QA approval remains unsigned because no rerun of the failed test is attached.\"},{\"speaker\":\"Business approver\",\"text\":\"Submitted business-approval artifact BA-1 bears my signature. The approval scope recorded for business-approval artifact BA-1 is release package R-17. The build 4.2 deliverable is release package R-17, and it is not release package R-18. No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable.\"},{\"speaker\":\"Delivery owner\",\"text\":\"I am responsible for correcting or rerunning the CSV-export test and supplying the resulting evidence for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "The approval scope recorded for business-approval artifact BA-1 is release package R-17."}, {"path": ["2", "text"], "text": "The build 4.2 deliverable is release package R-17, and it is not release package R-18."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The approval scope recorded for business-approval artifact BA-1 is release package R-17.", "negative_left": "The approval scope recorded for business-approval artifact BA-1 is release package R-18.", "negative_right": "The build 4.2 deliverable is release package R-17, and it is not release package R-18.", "right": "The build 4.2 deliverable is release package R-17, and it is not release package R-18."}, "verifier_independent_model": false}, "family": "fast-41-diverse-060-009", "id": "fast-41-diverse-060-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Project coordinator", "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}, {"speaker": "Quality reviewer", "text": "The verified run record for build 4.2 shows the specified CSV-export test failing. Each of the other 19 specified tests has a verified passing result. QA approval remains unsigned because no rerun of the failed test is attached."}, {"speaker": "Business approver", "text": "Submitted business-approval artifact BA-1 bears my signature. The approval scope recorded for business-approval artifact BA-1 is release package R-18. The build 4.2 deliverable is release package R-17, and it is not release package R-18. No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}, {"speaker": "Delivery owner", "text": "I am responsible for correcting or rerunning the CSV-export test and supplying the resulting evidence for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and question bindings; the two evidence spans are complete factual sentences—“The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42.” and “The build 4.2 deliverable is registered under the unique deliverable identifier D-42.”—and the counterfactual’s D-41 versus D-42 identifier mismatch is coherent without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Test record\",\"text\":\"For build 4.2, the specified CSV-export test has a verified failing result. Each of the other 19 specified tests for build 4.2 has a verified passing result.\"},{\"speaker\":\"Business-approval record\",\"text\":\"Submitted business-approval artifact BA-1 bears the business approver's signature.\"},{\"speaker\":\"Identifier records\",\"text\":\"The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42. The build 4.2 deliverable is registered under the unique deliverable identifier D-42.\"},{\"speaker\":\"Submission audit\",\"text\":\"No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable. The Quality reviewer has not signed QA approval because no verified rerun of the CSV-export test is attached.\"},{\"speaker\":\"Governing policy\",\"text\":\"Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42."}, {"path": ["2", "text"], "text": "The build 4.2 deliverable is registered under the unique deliverable identifier D-42."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42.", "negative_left": "The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-41.", "negative_right": "The build 4.2 deliverable is registered under the unique deliverable identifier D-42.", "right": "The build 4.2 deliverable is registered under the unique deliverable identifier D-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-060-011", "id": "fast-41-diverse-060-011-base", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Test record", "text": "For build 4.2, the specified CSV-export test has a verified failing result. Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"speaker": "Business-approval record", "text": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"speaker": "Identifier records", "text": "The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42. The build 4.2 deliverable is registered under the unique deliverable identifier D-42."}, {"speaker": "Submission audit", "text": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable. The Quality reviewer has not signed QA approval because no verified rerun of the CSV-export test is attached."}, {"speaker": "Governing policy", "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and question bindings; the two evidence spans are complete factual sentences—“The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42.” and “The build 4.2 deliverable is registered under the unique deliverable identifier D-42.”—and the counterfactual’s D-41 versus D-42 identifier mismatch is coherent without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Test record\",\"text\":\"For build 4.2, the specified CSV-export test has a verified failing result. Each of the other 19 specified tests for build 4.2 has a verified passing result.\"},{\"speaker\":\"Business-approval record\",\"text\":\"Submitted business-approval artifact BA-1 bears the business approver's signature.\"},{\"speaker\":\"Identifier records\",\"text\":\"The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42. The build 4.2 deliverable is registered under the unique deliverable identifier D-42.\"},{\"speaker\":\"Submission audit\",\"text\":\"No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable. The Quality reviewer has not signed QA approval because no verified rerun of the CSV-export test is attached.\"},{\"speaker\":\"Governing policy\",\"text\":\"Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42."}, {"path": ["2", "text"], "text": "The build 4.2 deliverable is registered under the unique deliverable identifier D-42."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-42.", "negative_left": "The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-41.", "negative_right": "The build 4.2 deliverable is registered under the unique deliverable identifier D-42.", "right": "The build 4.2 deliverable is registered under the unique deliverable identifier D-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-060-011", "id": "fast-41-diverse-060-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Test record", "text": "For build 4.2, the specified CSV-export test has a verified failing result. Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"speaker": "Business-approval record", "text": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"speaker": "Identifier records", "text": "The submitted business-approval artifact BA-1 lists its approved deliverable identifier as D-41. The build 4.2 deliverable is registered under the unique deliverable identifier D-42."}, {"speaker": "Submission audit", "text": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable. The Quality reviewer has not signed QA approval because no verified rerun of the CSV-export test is attached."}, {"speaker": "Governing policy", "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions retain the rubric and both contexts retain the governing acceptance conditions. Question bindings for build 4.2, its 20 tests, and the relevant approval and test paths remain unchanged. The evidence consists of two complete factual sentences: “The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42.” and “The build 4.2 deliverable is identified in the release register as deliverable D42.” The counterfactual is coherent because D42 is the recorded approval scope while D41 is the separately asserted build 4.2 identifier, indicating a scope mismatch. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Release review concerns build 4.2 and its 20 specified tests. The CSV-export test has a verified failing result; the other 19 specified tests each have verified passing results. Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"},{\"speaker\":\"Delivery owner\",\"text\":\"The submission includes business-approval artifact BA-1, signed by the business approver. No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The test record confirms the CSV-export failure and the 19 passing results. A rerun has not been attached, so QA approval remains pending.\"},{\"speaker\":\"Release register\",\"text\":\"The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42.\"},{\"speaker\":\"Approval record\",\"text\":\"The build 4.2 deliverable is identified in the release register as deliverable D42.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["3", "text"], "text": "The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42."}, {"path": ["4", "text"], "text": "The build 4.2 deliverable is identified in the release register as deliverable D42."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42.", "negative_left": "The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42.", "negative_right": "The build 4.2 deliverable is identified in the release register as deliverable D41.", "right": "The build 4.2 deliverable is identified in the release register as deliverable D42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-060-024", "id": "fast-41-diverse-060-024-base", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Project coordinator", "text": "Release review concerns build 4.2 and its 20 specified tests. The CSV-export test has a verified failing result; the other 19 specified tests each have verified passing results. Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}, {"speaker": "Delivery owner", "text": "The submission includes business-approval artifact BA-1, signed by the business approver. No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}, {"speaker": "Quality reviewer", "text": "The test record confirms the CSV-export failure and the 19 passing results. A rerun has not been attached, so QA approval remains pending."}, {"speaker": "Release register", "text": "The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42."}, {"speaker": "Approval record", "text": "The build 4.2 deliverable is identified in the release register as deliverable D42."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions retain the rubric and both contexts retain the governing acceptance conditions. Question bindings for build 4.2, its 20 tests, and the relevant approval and test paths remain unchanged. The evidence consists of two complete factual sentences: “The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42.” and “The build 4.2 deliverable is identified in the release register as deliverable D42.” The counterfactual is coherent because D42 is the recorded approval scope while D41 is the separately asserted build 4.2 identifier, indicating a scope mismatch. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a2 and a5 remain atomic despite quantifying over explicit sets. The focus a4 is a factual approval-scope relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a4 changing: BA-1 can remain signed and no other artifact can cover build 4.2 whether BA-1 itself does or does not. Policy evidence correctly cites only substantive policy from the original state; all rubric details and routing instructions in the questions object are automatically retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one failed test—the policy-designated noncritical CSV-export test—verified passes for all other 19 tests, and a signed business approval scoped to build 4.2. This is sufficient for score 1 under the unchanged question, including the stated remediation and QA-review routing.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a4 establishes that BA-1 does not cover build 4.2, while a5 excludes every other submitted business-approval artifact covering that deliverable. Thus the required business approval is missing, which independently requires score 0 and routing to the responsible business-approval role.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For build 4.2, the specified CSV-export test has a verified failing result."}, {"id": "a2", "statement": "Each of the other 19 specified tests for build 4.2 has a verified passing result."}, {"id": "a3", "statement": "Submitted business-approval artifact BA-1 bears the business approver's signature."}, {"id": "a4", "statement": "The approval scope of business-approval artifact BA-1 corresponds to the build 4.2 deliverable."}, {"id": "a5", "statement": "No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}], "base_state_json": "[{\"speaker\":\"Project coordinator\",\"text\":\"Release review concerns build 4.2 and its 20 specified tests. The CSV-export test has a verified failing result; the other 19 specified tests each have verified passing results. Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete.\"},{\"speaker\":\"Delivery owner\",\"text\":\"The submission includes business-approval artifact BA-1, signed by the business approver. No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The test record confirms the CSV-export failure and the 19 passing results. A rerun has not been attached, so QA approval remains pending.\"},{\"speaker\":\"Release register\",\"text\":\"The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42.\"},{\"speaker\":\"Approval record\",\"text\":\"The build 4.2 deliverable is identified in the release register as deliverable D42.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["3", "text"], "text": "The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42."}, {"path": ["4", "text"], "text": "The build 4.2 deliverable is identified in the release register as deliverable D42."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}], "rules": [{"justification": "The policy-designated noncritical CSV-export test is the sole failed test, every other specified test has a verified pass, and a signed business approval applies to build 4.2. Conditional acceptance therefore applies; the Delivery owner must correct or rerun the CSV-export test, and the Quality reviewer must then verify the evidence and provide QA approval.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "BA-1 does not approve the build 4.2 deliverable and no other submitted business-approval artifact does, so the required business approval is missing. The deliverable must be rejected and the missing approval routed to the responsible business-approval role.", "target": "0", "when": [{"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42.", "negative_left": "The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42.", "negative_right": "The build 4.2 deliverable is identified in the release register as deliverable D41.", "right": "The build 4.2 deliverable is identified in the release register as deliverable D42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-060-024", "id": "fast-41-diverse-060-024-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not acceptable: Reject when any critical test fails or is unverified, business approval is missing, more than one test is failed or unverified, or the evidence cannot establish that the remaining criteria are complete. Route each defect or missing approval to its responsible role.", "1 — Conditionally acceptable: Exactly one noncritical test is failed or unverified, all other specified tests have verified passes, and business approval is signed. The Delivery owner must correct or rerun the test, after which the Quality reviewer must verify the evidence and provide QA approval.", "2 — Fully acceptable: All 20 tests have verified passes for build 4.2, and both signed QA test approval and signed business approval are present; no defect remediation or approval follow-up remains."], "instructions": "Rate overall deliverable completion quality using the ordered rubric. Resolve conflicting evidence by relying on test-run metadata rather than the document title or unsupported correction claim. Also identify the required routing in the reason.", "type": "score"}}, "state": [{"speaker": "Project coordinator", "text": "Release review concerns build 4.2 and its 20 specified tests. The CSV-export test has a verified failing result; the other 19 specified tests each have verified passing results. Acceptance requires all 20 specified tests to pass on build 4.2, a signed QA test approval, and business approval. The CSV-export test is noncritical. Policy permits conditional acceptance when exactly one noncritical test remains failed or unverified and all other tests plus business approval are complete."}, {"speaker": "Delivery owner", "text": "The submission includes business-approval artifact BA-1, signed by the business approver. No submitted business-approval artifact other than BA-1 has an approval scope corresponding to the build 4.2 deliverable."}, {"speaker": "Quality reviewer", "text": "The test record confirms the CSV-export failure and the 19 passing results. A rerun has not been attached, so QA approval remains pending."}, {"speaker": "Release register", "text": "The approval scope recorded on business-approval artifact BA-1 names deliverable identifier D42."}, {"speaker": "Approval record", "text": "The build 4.2 deliverable is identified in the release register as deliverable D41."}]}, "method": "c2d", "provenance": {"source_id": "diverse-060", "source_is_synthetic": true, "source_sha256": "d3a26eb2c7d88fea03cba613ea57cd1808db379ae813bc2e617a3850cf3e7003", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policy, preserve Miko’s case scope and bindings, and contain exactly two factual evidence sentences; the counterfactual changes only the answer-key code without creating an internal contradiction or stating a classification answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"case_note\":\"Miko’s five-item area-and-perimeter check was fully documented on 2026-09-17. A response is recorded for each item, and the answer key supplies an entry for all five. Items 1 and 5 are designated diagnostic items; their recorded responses are incorrect, while items 2 and 4 are correct. The work shown for diagnostic items 1 and 5 uses different incorrect procedures, so the errors do not reflect one repeated method. Diagnostic designations and shown work are available for every diagnostic item. Miko’s lesson record is available and confirms attendance at the relevant area-versus-perimeter lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Determine mastery, route, and ordered intervention urgency using the supplied rubric.\",\"verified_evidence\":[\"In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3.\",\"The answer key issued for that same five-item check on 2026-09-17 specifies response code Q7 for item 3.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["verified_evidence", "0"], "text": "In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3."}, {"path": ["verified_evidence", "1"], "text": "The answer key issued for that same five-item check on 2026-09-17 specifies response code Q7 for item 3."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3.", "negative_left": "In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3.", "negative_right": "The answer key issued for that same five-item check on 2026-09-17 specifies response code Q8 for item 3.", "right": "The answer key issued for that same five-item check on 2026-09-17 specifies response code Q7 for item 3."}, "verifier_independent_model": false}, "family": "fast-41-diverse-061-003", "id": "fast-41-diverse-061-003-base", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"case_note": "Miko’s five-item area-and-perimeter check was fully documented on 2026-09-17. A response is recorded for each item, and the answer key supplies an entry for all five. Items 1 and 5 are designated diagnostic items; their recorded responses are incorrect, while items 2 and 4 are correct. The work shown for diagnostic items 1 and 5 uses different incorrect procedures, so the errors do not reflect one repeated method. Diagnostic designations and shown work are available for every diagnostic item. Miko’s lesson record is available and confirms attendance at the relevant area-versus-perimeter lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Determine mastery, route, and ordered intervention urgency using the supplied rubric.", "verified_evidence": ["In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3.", "The answer key issued for that same five-item check on 2026-09-17 specifies response code Q7 for item 3."]}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "B — Not yet mastered; practice set; U1"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policy, preserve Miko’s case scope and bindings, and contain exactly two factual evidence sentences; the counterfactual changes only the answer-key code without creating an internal contradiction or stating a classification answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"case_note\":\"Miko’s five-item area-and-perimeter check was fully documented on 2026-09-17. A response is recorded for each item, and the answer key supplies an entry for all five. Items 1 and 5 are designated diagnostic items; their recorded responses are incorrect, while items 2 and 4 are correct. The work shown for diagnostic items 1 and 5 uses different incorrect procedures, so the errors do not reflect one repeated method. Diagnostic designations and shown work are available for every diagnostic item. Miko’s lesson record is available and confirms attendance at the relevant area-versus-perimeter lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Determine mastery, route, and ordered intervention urgency using the supplied rubric.\",\"verified_evidence\":[\"In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3.\",\"The answer key issued for that same five-item check on 2026-09-17 specifies response code Q7 for item 3.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["verified_evidence", "0"], "text": "In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3."}, {"path": ["verified_evidence", "1"], "text": "The answer key issued for that same five-item check on 2026-09-17 specifies response code Q7 for item 3."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3.", "negative_left": "In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3.", "negative_right": "The answer key issued for that same five-item check on 2026-09-17 specifies response code Q8 for item 3.", "right": "The answer key issued for that same five-item check on 2026-09-17 specifies response code Q7 for item 3."}, "verifier_independent_model": false}, "family": "fast-41-diverse-061-003", "id": "fast-41-diverse-061-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"case_note": "Miko’s five-item area-and-perimeter check was fully documented on 2026-09-17. A response is recorded for each item, and the answer key supplies an entry for all five. Items 1 and 5 are designated diagnostic items; their recorded responses are incorrect, while items 2 and 4 are correct. The work shown for diagnostic items 1 and 5 uses different incorrect procedures, so the errors do not reflect one repeated method. Diagnostic designations and shown work are available for every diagnostic item. Miko’s lesson record is available and confirms attendance at the relevant area-versus-perimeter lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Determine mastery, route, and ordered intervention urgency using the supplied rubric.", "verified_evidence": ["In the five-item check recorded on 2026-09-17, Miko entered response code Q7 for item 3.", "The answer key issued for that same five-item check on 2026-09-17 specifies response code Q8 for item 3."]}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Not mastered; tutor remediation; U3"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and the contexts add no conflicting rules. Both contexts retain Miko, the five-item check, the route-and-urgency request, and all relevant item bindings. The focus evidence contains two complete factual sentences in each context. Changing item 3's key from “Lumen.” to “Neris.” creates no contradiction with the other unchanged facts. Neither context states a gold decision, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Miko's five-item area-and-perimeter check has a recorded response and answer-key entry for every item. Items 1 and 5 are the diagnostic items; Miko's responses to 1 and 5 differ from their key entries, while responses to 2 and 4 agree with their keys. The shown work for diagnostic item 1 uses a different method from the shown work for diagnostic item 5. Diagnostic designations and shown work are recorded for all five items. Miko's lesson record is available and confirms attendance at the relevant lesson. The instructor must select one route and urgency from the supplied rubric.\",\"evidence\":[\"Item 1 response does not match its key; item 2 response matches its key; item 4 response matches its key; item 5 response does not match its key.\",\"Items 1 and 5 are diagnostic, and their recorded shown-work methods differ.\",\"A response, answer-key entry, diagnostic designation, and shown work are recorded for every item; Miko's lesson record is available.\",\"For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”\",\"The answer key for item 3 in the five-item check lists “Lumen.”\"],\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "3"], "text": "For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”"}, {"path": ["evidence", "4"], "text": "The answer key for item 3 in the five-item check lists “Lumen.”"}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”", "negative_left": "For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”", "negative_right": "The answer key for item 3 in the five-item check lists “Neris.”", "right": "The answer key for item 3 in the five-item check lists “Lumen.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-061-008", "id": "fast-41-diverse-061-008-base", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Miko's five-item area-and-perimeter check has a recorded response and answer-key entry for every item. Items 1 and 5 are the diagnostic items; Miko's responses to 1 and 5 differ from their key entries, while responses to 2 and 4 agree with their keys. The shown work for diagnostic item 1 uses a different method from the shown work for diagnostic item 5. Diagnostic designations and shown work are recorded for all five items. Miko's lesson record is available and confirms attendance at the relevant lesson. The instructor must select one route and urgency from the supplied rubric.", "evidence": ["Item 1 response does not match its key; item 2 response matches its key; item 4 response matches its key; item 5 response does not match its key.", "Items 1 and 5 are diagnostic, and their recorded shown-work methods differ.", "A response, answer-key entry, diagnostic designation, and shown work are recorded for every item; Miko's lesson record is available.", "For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”", "The answer key for item 3 in the five-item check lists “Lumen.”"], "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "B — Not yet mastered; practice set; U1"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and the contexts add no conflicting rules. Both contexts retain Miko, the five-item check, the route-and-urgency request, and all relevant item bindings. The focus evidence contains two complete factual sentences in each context. Changing item 3's key from “Lumen.” to “Neris.” creates no contradiction with the other unchanged facts. Neither context states a gold decision, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Miko's five-item area-and-perimeter check has a recorded response and answer-key entry for every item. Items 1 and 5 are the diagnostic items; Miko's responses to 1 and 5 differ from their key entries, while responses to 2 and 4 agree with their keys. The shown work for diagnostic item 1 uses a different method from the shown work for diagnostic item 5. Diagnostic designations and shown work are recorded for all five items. Miko's lesson record is available and confirms attendance at the relevant lesson. The instructor must select one route and urgency from the supplied rubric.\",\"evidence\":[\"Item 1 response does not match its key; item 2 response matches its key; item 4 response matches its key; item 5 response does not match its key.\",\"Items 1 and 5 are diagnostic, and their recorded shown-work methods differ.\",\"A response, answer-key entry, diagnostic designation, and shown work are recorded for every item; Miko's lesson record is available.\",\"For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”\",\"The answer key for item 3 in the five-item check lists “Lumen.”\"],\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "3"], "text": "For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”"}, {"path": ["evidence", "4"], "text": "The answer key for item 3 in the five-item check lists “Lumen.”"}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”", "negative_left": "For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”", "negative_right": "The answer key for item 3 in the five-item check lists “Neris.”", "right": "The answer key for item 3 in the five-item check lists “Lumen.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-061-008", "id": "fast-41-diverse-061-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Miko's five-item area-and-perimeter check has a recorded response and answer-key entry for every item. Items 1 and 5 are the diagnostic items; Miko's responses to 1 and 5 differ from their key entries, while responses to 2 and 4 agree with their keys. The shown work for diagnostic item 1 uses a different method from the shown work for diagnostic item 5. Diagnostic designations and shown work are recorded for all five items. Miko's lesson record is available and confirms attendance at the relevant lesson. The instructor must select one route and urgency from the supplied rubric.", "evidence": ["Item 1 response does not match its key; item 2 response matches its key; item 4 response matches its key; item 5 response does not match its key.", "Items 1 and 5 are diagnostic, and their recorded shown-work methods differ.", "A response, answer-key entry, diagnostic designation, and shown work are recorded for every item; Miko's lesson record is available.", "For Miko's response to item 3 in the five-item check, the recorded answer is “Lumen.”", "The answer key for item 3 in the five-item check lists “Neris.”"], "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Not mastered; tutor remediation; U3"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions, policy, Miko, item paths, date, and request bindings. The evidence spans are complete factual sentences: “In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36.” and “For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3.” The counterfactual coherently changes only item 3's answer key while retaining the recorded response and other assertions. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"Case note dated 14 March 2026: Miko completed a five-item area-and-perimeter check. In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36. For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3. The recorded responses and answer-key entries are present for all five items. Responses to items 1 and 5 differ from their corresponding key entries, while responses to items 2 and 4 match theirs. Items 1 and 5 are the designated diagnostic items, and shown work is recorded for each diagnostic item. The methods shown on those two items are different. Miko's lesson record is available and confirms attendance at the area-versus-perimeter lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.\",\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["context"], "text": "In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36."}, {"path": ["context"], "text": "For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36.", "negative_left": "In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36.", "negative_right": "For the five-item check dated 14 March 2026, the answer key specifies 41 for item 3.", "right": "For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3."}, "verifier_independent_model": false}, "family": "fast-41-diverse-061-010", "id": "fast-41-diverse-061-010-base", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "Case note dated 14 March 2026: Miko completed a five-item area-and-perimeter check. In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36. For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3. The recorded responses and answer-key entries are present for all five items. Responses to items 1 and 5 differ from their corresponding key entries, while responses to items 2 and 4 match theirs. Items 1 and 5 are the designated diagnostic items, and shown work is recorded for each diagnostic item. The methods shown on those two items are different. Miko's lesson record is available and confirms attendance at the area-versus-perimeter lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.", "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "B — Not yet mastered; practice set; U1"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions, policy, Miko, item paths, date, and request bindings. The evidence spans are complete factual sentences: “In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36.” and “For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3.” The counterfactual coherently changes only item 3's answer key while retaining the recorded response and other assertions. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"Case note dated 14 March 2026: Miko completed a five-item area-and-perimeter check. In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36. For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3. The recorded responses and answer-key entries are present for all five items. Responses to items 1 and 5 differ from their corresponding key entries, while responses to items 2 and 4 match theirs. Items 1 and 5 are the designated diagnostic items, and shown work is recorded for each diagnostic item. The methods shown on those two items are different. Miko's lesson record is available and confirms attendance at the area-versus-perimeter lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.\",\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["context"], "text": "In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36."}, {"path": ["context"], "text": "For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36.", "negative_left": "In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36.", "negative_right": "For the five-item check dated 14 March 2026, the answer key specifies 41 for item 3.", "right": "For the five-item check dated 14 March 2026, the answer key specifies 36 for item 3."}, "verifier_independent_model": false}, "family": "fast-41-diverse-061-010", "id": "fast-41-diverse-061-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "Case note dated 14 March 2026: Miko completed a five-item area-and-perimeter check. In the five-item check dated 14 March 2026, Miko's recorded response to item 3 is 36. For the five-item check dated 14 March 2026, the answer key specifies 41 for item 3. The recorded responses and answer-key entries are present for all five items. Responses to items 1 and 5 differ from their corresponding key entries, while responses to items 2 and 4 match theirs. Items 1 and 5 are the designated diagnostic items, and shown work is recorded for each diagnostic item. The methods shown on those two items are different. Miko's lesson record is available and confirms attendance at the area-versus-perimeter lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip.", "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Not mastered; tutor remediation; U3"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The evidence spans are complete factual sentences: “Miko's recorded response to item 3 is \"14\".” and “The answer key for item 3 specifies \"14\" as the answer.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"Case note: Miko completed a five-item area-and-perimeter check. The recorded responses for items 1 and 5 do not match their keyed answers, while the responses for items 2 and 4 do. Miko's recorded response to item 3 is \\\"14\\\". The answer key for item 3 specifies \\\"14\\\" as the answer. The answer key contains an entry for each of the five items, and a response is recorded for each one. Items 1 and 5 are the designated diagnostic items. Shown work is available for every diagnostic item; the two diagnostic errors use different methods. Miko's lesson record is available, confirming attendance at the relevant lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Determine mastery, route, and ordered intervention urgency using the supplied rubric.\",\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["context"], "text": "Miko's recorded response to item 3 is \"14\"."}, {"path": ["context"], "text": "The answer key for item 3 specifies \"14\" as the answer."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "Miko's recorded response to item 3 is \"14\".", "negative_left": "Miko's recorded response to item 3 is \"14\".", "negative_right": "The answer key for item 3 specifies \"19\" as the answer.", "right": "The answer key for item 3 specifies \"14\" as the answer."}, "verifier_independent_model": false}, "family": "fast-41-diverse-061-012", "id": "fast-41-diverse-061-012-base", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "Case note: Miko completed a five-item area-and-perimeter check. The recorded responses for items 1 and 5 do not match their keyed answers, while the responses for items 2 and 4 do. Miko's recorded response to item 3 is \"14\". The answer key for item 3 specifies \"14\" as the answer. The answer key contains an entry for each of the five items, and a response is recorded for each one. Items 1 and 5 are the designated diagnostic items. Shown work is available for every diagnostic item; the two diagnostic errors use different methods. Miko's lesson record is available, confirming attendance at the relevant lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Determine mastery, route, and ordered intervention urgency using the supplied rubric.", "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "B — Not yet mastered; practice set; U1"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The evidence spans are complete factual sentences: “Miko's recorded response to item 3 is \"14\".” and “The answer key for item 3 specifies \"14\" as the answer.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\":\"Case note: Miko completed a five-item area-and-perimeter check. The recorded responses for items 1 and 5 do not match their keyed answers, while the responses for items 2 and 4 do. Miko's recorded response to item 3 is \\\"14\\\". The answer key for item 3 specifies \\\"14\\\" as the answer. The answer key contains an entry for each of the five items, and a response is recorded for each one. Items 1 and 5 are the designated diagnostic items. Shown work is available for every diagnostic item; the two diagnostic errors use different methods. Miko's lesson record is available, confirming attendance at the relevant lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Determine mastery, route, and ordered intervention urgency using the supplied rubric.\",\"request\":\"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["context"], "text": "Miko's recorded response to item 3 is \"14\"."}, {"path": ["context"], "text": "The answer key for item 3 specifies \"14\" as the answer."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "Miko's recorded response to item 3 is \"14\".", "negative_left": "Miko's recorded response to item 3 is \"14\".", "negative_right": "The answer key for item 3 specifies \"19\" as the answer.", "right": "The answer key for item 3 specifies \"14\" as the answer."}, "verifier_independent_model": false}, "family": "fast-41-diverse-061-012", "id": "fast-41-diverse-061-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "Case note: Miko completed a five-item area-and-perimeter check. The recorded responses for items 1 and 5 do not match their keyed answers, while the responses for items 2 and 4 do. Miko's recorded response to item 3 is \"14\". The answer key for item 3 specifies \"19\" as the answer. The answer key contains an entry for each of the five items, and a response is recorded for each one. Items 1 and 5 are the designated diagnostic items. Shown work is available for every diagnostic item; the two diagnostic errors use different methods. Miko's lesson record is available, confirming attendance at the relevant lesson. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Determine mastery, route, and ordered intervention urgency using the supplied rubric.", "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Not mastered; tutor remediation; U3"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria, and both contexts retain the same policy without inventing exceptions. The quiz date, item, affirmative-version reference, and routing scope remain bound consistently. The required evidence sentences are: \"In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B.\" \"In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B.\" Both evidence spans are complete factual sentences. The counterfactual’s changed answer-key value coheres with the stated incorrect response and introduces no contradictory duplicate assertion. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a8 is one consistency relation over an explicit set of records rather than a bundle of policy requirements. The focus a4 is factual. The base and counter assignments can both be realized while changing only whether the item-4 response matches the affirmative-version answer; the response can remain incorrect in either case, preserving the 4/5 score and all other atom states. The policy evidence is an accurate citation from the original state. No additional state-derived rule is needed because the complete governing rubric remains available verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. The rule also establishes that item 4 is the negation probe, the selected response matches its affirmative-version answer, and all required records are present and consistent. These conditions sufficiently entail U2_negation_review and exclude EVIDENCE_HOLD and the competing scored outcomes.", "rule_index": 0, "sound": true}, {"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. With item 4 established as the negation probe and the affirmative-version match refuted, the U2 exception does not apply and U1_boundary_mastery is entailed. Record presence and consistency exclude EVIDENCE_HOLD.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The learner answered exactly four of the five quiz items correctly."}, {"id": "a2", "statement": "The learner answered quiz item 4 incorrectly."}, {"id": "a3", "statement": "Quiz item 4 is the negation probe."}, {"id": "a4", "statement": "The learner’s selected response on quiz item 4 matches the answer to the affirmative version of quiz item 4."}, {"id": "a5", "statement": "The learner’s quiz responses are present."}, {"id": "a6", "statement": "The answer key is present."}, {"id": "a7", "statement": "The lesson record is present."}, {"id": "a8", "statement": "The learner’s quiz responses, answer key, and lesson record contain no mutually inconsistent information relevant to scoring or routing this quiz."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"On 14 March 2026, the learner’s five-item quiz was scored against the answer key: four responses were correct, and quiz item 4 was the only incorrect item. The lesson materials identify item 4 as the negation probe. The learner’s quiz responses, answer key, and lesson record are present. A records check found no mutually inconsistent information relevant to scoring or routing this quiz.\"},{\"speaker\":\"Learner record\",\"text\":\"In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B.\"},{\"speaker\":\"Answer-key record\",\"text\":\"In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B.\"},{\"speaker\":\"Lesson record\",\"text\":\"The lesson record shows completion of the affirmative-claims activity and no completion entry for the NOT/EXCEPT contrast sort.\"},{\"speaker\":\"Policy record\",\"text\":\"Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B."}, {"path": ["2", "text"], "text": "In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B."}], "policy_evidence": [{"path": ["3", "text"], "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}], "rules": [{"justification": "The records are present and consistent, the score is 4/5, item 4 is necessarily the sole miss, item 4 is the negation probe, and the selected response matches its affirmative-version answer.", "target": "U2_negation_review", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The records are present and consistent, the score is 4/5, and item 4 is necessarily the sole miss; although it is the negation probe, the selected response does not match its affirmative-version answer.", "target": "U1_boundary_mastery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B.", "negative_left": "In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B.", "negative_right": "In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option C.", "right": "In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B."}, "verifier_independent_model": false}, "family": "fast-41-diverse-062-001", "id": "fast-41-diverse-062-001-base", "input": {"questions": {"decision": {"criteria": {"EVIDENCE_HOLD": "No urgency assigned: Responses, answer key, or lesson record are missing or inconsistent, so scoring or routing cannot yet be completed.", "U0_full_mastery": "Urgency 0: Score is 5/5. Verify mastery; assign no review or support role.", "U1_boundary_mastery": "Urgency 1: Score is 4/5, but the sole miss either is not the negation probe or does not match the affirmative-version answer. Verify mastery; assign an optional independent recap.", "U2_negation_review": "Urgency 2: Score is 4/5, the only miss is the negation probe, and the chosen response matches the affirmative-version answer. Do not verify mastery; route to the course instructor’s NOT/EXCEPT contrast sort before the next quiz.", "U3_tutor_reteach": "Urgency 3: Score is exactly 3/5, regardless of error type. Do not verify mastery; route to a learning support tutor for a full guided practice set before the next quiz.", "U4_intensive_support": "Urgency 4: Score is 0–2/5. Do not verify mastery; route first to the course instructor for reteaching and then to the learning support tutor for supervised practice."}, "instructions": "Select the single rubric option that correctly verifies mastery status, assigns the prescribed route, and gives the ordered intervention urgency. Lower urgency numbers mean less urgent intervention.", "type": "choice"}}, "state": [{"speaker": "Case note", "text": "On 14 March 2026, the learner’s five-item quiz was scored against the answer key: four responses were correct, and quiz item 4 was the only incorrect item. The lesson materials identify item 4 as the negation probe. The learner’s quiz responses, answer key, and lesson record are present. A records check found no mutually inconsistent information relevant to scoring or routing this quiz."}, {"speaker": "Learner record", "text": "In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B."}, {"speaker": "Answer-key record", "text": "In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B."}, {"speaker": "Lesson record", "text": "The lesson record shows completion of the affirmative-claims activity and no completion entry for the NOT/EXCEPT contrast sort."}, {"speaker": "Policy record", "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}]}, "method": "c2d", "provenance": {"source_id": "diverse-062", "source_is_synthetic": true, "source_sha256": "571895f9ad2bd8e1cdf4993a3c848224bd6a8603b7160d6c25bc7527551b1fa8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "U2_negation_review"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria, and both contexts retain the same policy without inventing exceptions. The quiz date, item, affirmative-version reference, and routing scope remain bound consistently. The required evidence sentences are: \"In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B.\" \"In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B.\" Both evidence spans are complete factual sentences. The counterfactual’s changed answer-key value coheres with the stated incorrect response and introduces no contradictory duplicate assertion. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a8 is one consistency relation over an explicit set of records rather than a bundle of policy requirements. The focus a4 is factual. The base and counter assignments can both be realized while changing only whether the item-4 response matches the affirmative-version answer; the response can remain incorrect in either case, preserving the 4/5 score and all other atom states. The policy evidence is an accurate citation from the original state. No additional state-derived rule is needed because the complete governing rubric remains available verbatim in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. The rule also establishes that item 4 is the negation probe, the selected response matches its affirmative-version answer, and all required records are present and consistent. These conditions sufficiently entail U2_negation_review and exclude EVIDENCE_HOLD and the competing scored outcomes.", "rule_index": 0, "sound": true}, {"reason": "A 4/5 score plus item 4 being incorrect entails that item 4 is the sole miss. With item 4 established as the negation probe and the affirmative-version match refuted, the U2 exception does not apply and U1_boundary_mastery is entailed. Record presence and consistency exclude EVIDENCE_HOLD.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The learner answered exactly four of the five quiz items correctly."}, {"id": "a2", "statement": "The learner answered quiz item 4 incorrectly."}, {"id": "a3", "statement": "Quiz item 4 is the negation probe."}, {"id": "a4", "statement": "The learner’s selected response on quiz item 4 matches the answer to the affirmative version of quiz item 4."}, {"id": "a5", "statement": "The learner’s quiz responses are present."}, {"id": "a6", "statement": "The answer key is present."}, {"id": "a7", "statement": "The lesson record is present."}, {"id": "a8", "statement": "The learner’s quiz responses, answer key, and lesson record contain no mutually inconsistent information relevant to scoring or routing this quiz."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"On 14 March 2026, the learner’s five-item quiz was scored against the answer key: four responses were correct, and quiz item 4 was the only incorrect item. The lesson materials identify item 4 as the negation probe. The learner’s quiz responses, answer key, and lesson record are present. A records check found no mutually inconsistent information relevant to scoring or routing this quiz.\"},{\"speaker\":\"Learner record\",\"text\":\"In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B.\"},{\"speaker\":\"Answer-key record\",\"text\":\"In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B.\"},{\"speaker\":\"Lesson record\",\"text\":\"The lesson record shows completion of the affirmative-claims activity and no completion entry for the NOT/EXCEPT contrast sort.\"},{\"speaker\":\"Policy record\",\"text\":\"Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B."}, {"path": ["2", "text"], "text": "In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B."}], "policy_evidence": [{"path": ["3", "text"], "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}], "rules": [{"justification": "The records are present and consistent, the score is 4/5, item 4 is necessarily the sole miss, item 4 is the negation probe, and the selected response matches its affirmative-version answer.", "target": "U2_negation_review", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The records are present and consistent, the score is 4/5, and item 4 is necessarily the sole miss; although it is the negation probe, the selected response does not match its affirmative-version answer.", "target": "U1_boundary_mastery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B.", "negative_left": "In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B.", "negative_right": "In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option C.", "right": "In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option B."}, "verifier_independent_model": false}, "family": "fast-41-diverse-062-001", "id": "fast-41-diverse-062-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"EVIDENCE_HOLD": "No urgency assigned: Responses, answer key, or lesson record are missing or inconsistent, so scoring or routing cannot yet be completed.", "U0_full_mastery": "Urgency 0: Score is 5/5. Verify mastery; assign no review or support role.", "U1_boundary_mastery": "Urgency 1: Score is 4/5, but the sole miss either is not the negation probe or does not match the affirmative-version answer. Verify mastery; assign an optional independent recap.", "U2_negation_review": "Urgency 2: Score is 4/5, the only miss is the negation probe, and the chosen response matches the affirmative-version answer. Do not verify mastery; route to the course instructor’s NOT/EXCEPT contrast sort before the next quiz.", "U3_tutor_reteach": "Urgency 3: Score is exactly 3/5, regardless of error type. Do not verify mastery; route to a learning support tutor for a full guided practice set before the next quiz.", "U4_intensive_support": "Urgency 4: Score is 0–2/5. Do not verify mastery; route first to the course instructor for reteaching and then to the learning support tutor for supervised practice."}, "instructions": "Select the single rubric option that correctly verifies mastery status, assigns the prescribed route, and gives the ordered intervention urgency. Lower urgency numbers mean less urgent intervention.", "type": "choice"}}, "state": [{"speaker": "Case note", "text": "On 14 March 2026, the learner’s five-item quiz was scored against the answer key: four responses were correct, and quiz item 4 was the only incorrect item. The lesson materials identify item 4 as the negation probe. The learner’s quiz responses, answer key, and lesson record are present. A records check found no mutually inconsistent information relevant to scoring or routing this quiz."}, {"speaker": "Learner record", "text": "In the learner’s recorded responses for the 14 March 2026 five-item quiz, the selected response for quiz item 4 was option B."}, {"speaker": "Answer-key record", "text": "In the answer key recorded for the 14 March 2026 five-item quiz, the answer for the affirmative version of quiz item 4 was option C."}, {"speaker": "Lesson record", "text": "The lesson record shows completion of the affirmative-claims activity and no completion entry for the NOT/EXCEPT contrast sort."}, {"speaker": "Policy record", "text": "Our rubric treats 4/5 as mastery unless the only error is the negation probe and the selected answer matches the affirmative version. That exception requires targeted review before the next quiz."}]}, "method": "c2d", "provenance": {"source_id": "diverse-062", "source_is_synthetic": true, "source_sha256": "571895f9ad2bd8e1cdf4993a3c848224bd6a8603b7160d6c25bc7527551b1fa8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "U1_boundary_mastery"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in the unchanged questions and both contexts. Mina, the five-item quiz, diagnostic items, and routing bindings remain unchanged. The evidence consists of exactly two complete factual sentences: \"On Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\" and \"Mina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\" The counterfactual coherently changes only diagnostic selections, leaving exactly items 3 and 5 correct and making all three diagnostic responses incorrect. Neither context states a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "full_context_fact_states": {"base": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "counterfactual": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "remove_left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "remove_right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "negative_pair": {"pattern_on_at_least_2_diagnostic_items": "refuted"}, "negative_sentence": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "positive_pair": {"pattern_on_at_least_2_diagnostic_items": "supported"}, "right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express single factual relationships, and the focus is the factual presence or absence of a response pattern rather than a policy conclusion. Both focus assignments are realizable while holding the 2/5 score fixed. The policy evidence preserves all substantive state-originating rules needed to apply the unchanged question; omitted details such as guided-practice performance are not needed for these rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly 2/5 correct necessarily fails the stated 4/5 mastery threshold. With the misconception pattern present on at least two diagnostic items, Mina’s condition is met, requiring the tutor’s Fraction-Strip Comparison Lab at urgency 3; therefore the proposed instructor conference and urgency 2 triage is incorrect.", "rule_index": 0, "sound": true}, {"reason": "Exactly 2/5 correct establishes nonmastery. Refutation of the at-least-two pattern proposition entails that Mina’s condition is not met, so the otherwise branch requires the instructor correction conference at urgency 2, making every component of the proposal correct.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "quiz_score_2_of_5", "statement": "Mina answered exactly 2 of the 5 items on this five-item fraction quiz correctly."}, {"id": "pattern_on_at_least_2_diagnostic_items", "statement": "Mina’s responses on this five-item fraction quiz exhibit the “larger denominator means larger fraction” pattern on at least 2 of diagnostic items 1, 2, and 4."}], "base_state_json": "\"Mina’s five-item fraction quiz was scored against the instructor’s key. The record shows exactly two correct responses: items 3 and 5. Items 1, 2, and 4 were designated diagnostic items for comparing fractions, and the response sheet preserved the listed-option order for those items.\\n\\nOn Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\\nMina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\\n\\nMina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”\\nPolicy defines mastery as 4/5 plus 2/3 diagnostic items correct.\\nNonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2.\\nUrgency runs from 1 (lowest) to 3 (highest).\"", "base_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}], "counter_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}], "focus_atom": "pattern_on_at_least_2_diagnostic_items", "focus_evidence": [{"path": [], "text": "On Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4."}, {"path": [], "text": "Mina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4."}], "policy_evidence": [{"path": [], "text": "Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”"}, {"path": [], "text": "Policy defines mastery as 4/5 plus 2/3 diagnostic items correct."}, {"path": [], "text": "Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2."}, {"path": [], "text": "Urgency runs from 1 (lowest) to 3 (highest)."}], "rules": [{"justification": "A score of 2/5 establishes nonmastery. Exhibiting the stated pattern on at least two diagnostic items meets Mina’s condition, so policy requires the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3 rather than the proposed instructor conference at urgency 2.", "target": "false", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}]}, {"justification": "A score of 2/5 establishes nonmastery. Explicit refutation of exhibiting the pattern on at least two diagnostic items establishes that Mina’s condition is not met, so policy requires the instructor correction conference at urgency 2, matching the proposal.", "target": "true", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}]}]}, "verified_pair": {"left": "On Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.", "negative_left": "On Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.", "negative_right": "Mina selected the first listed option for diagnostic item 1, the second listed option for diagnostic item 2, and the first listed option for diagnostic item 4.", "right": "Mina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-063-006", "id": "fast-41-diverse-063-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — Mina is not mastered, but she must be routed to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3.", "true": "Yes — the proposed not-mastered classification, instructor conference, and urgency 2 assignment are all correct."}, "instructions": "Decide whether this proposed triage is correct: mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Answer Yes or No according to the stated policy.", "type": "noul"}}, "state": "Mina’s five-item fraction quiz was scored against the instructor’s key. The record shows exactly two correct responses: items 3 and 5. Items 1, 2, and 4 were designated diagnostic items for comparing fractions, and the response sheet preserved the listed-option order for those items.\n\nOn Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\nMina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\n\nMina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”\nPolicy defines mastery as 4/5 plus 2/3 diagnostic items correct.\nNonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2.\nUrgency runs from 1 (lowest) to 3 (highest)."}, "method": "c2d", "provenance": {"source_id": "diverse-063", "source_is_synthetic": true, "source_sha256": "1ab2940f7e11265b1fc7b64fecf6943e6d0f6fc9c6402c436749bbc8483bd245", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in the unchanged questions and both contexts. Mina, the five-item quiz, diagnostic items, and routing bindings remain unchanged. The evidence consists of exactly two complete factual sentences: \"On Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\" and \"Mina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\" The counterfactual coherently changes only diagnostic selections, leaving exactly items 3 and 5 correct and making all three diagnostic responses incorrect. Neither context states a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "full_context_fact_states": {"base": {"pattern_on_at_least_2_diagnostic_items": "supported", "quiz_score_2_of_5": "supported"}, "counterfactual": {"pattern_on_at_least_2_diagnostic_items": "refuted", "quiz_score_2_of_5": "supported"}, "remove_left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "remove_right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "negative_pair": {"pattern_on_at_least_2_diagnostic_items": "refuted"}, "negative_sentence": {"pattern_on_at_least_2_diagnostic_items": "unknown"}, "positive_pair": {"pattern_on_at_least_2_diagnostic_items": "supported"}, "right": {"pattern_on_at_least_2_diagnostic_items": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express single factual relationships, and the focus is the factual presence or absence of a response pattern rather than a policy conclusion. Both focus assignments are realizable while holding the 2/5 score fixed. The policy evidence preserves all substantive state-originating rules needed to apply the unchanged question; omitted details such as guided-practice performance are not needed for these rules.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Exactly 2/5 correct necessarily fails the stated 4/5 mastery threshold. With the misconception pattern present on at least two diagnostic items, Mina’s condition is met, requiring the tutor’s Fraction-Strip Comparison Lab at urgency 3; therefore the proposed instructor conference and urgency 2 triage is incorrect.", "rule_index": 0, "sound": true}, {"reason": "Exactly 2/5 correct establishes nonmastery. Refutation of the at-least-two pattern proposition entails that Mina’s condition is not met, so the otherwise branch requires the instructor correction conference at urgency 2, making every component of the proposal correct.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "quiz_score_2_of_5", "statement": "Mina answered exactly 2 of the 5 items on this five-item fraction quiz correctly."}, {"id": "pattern_on_at_least_2_diagnostic_items", "statement": "Mina’s responses on this five-item fraction quiz exhibit the “larger denominator means larger fraction” pattern on at least 2 of diagnostic items 1, 2, and 4."}], "base_state_json": "\"Mina’s five-item fraction quiz was scored against the instructor’s key. The record shows exactly two correct responses: items 3 and 5. Items 1, 2, and 4 were designated diagnostic items for comparing fractions, and the response sheet preserved the listed-option order for those items.\\n\\nOn Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\\nMina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\\n\\nMina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”\\nPolicy defines mastery as 4/5 plus 2/3 diagnostic items correct.\\nNonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2.\\nUrgency runs from 1 (lowest) to 3 (highest).\"", "base_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}], "counter_states": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}], "focus_atom": "pattern_on_at_least_2_diagnostic_items", "focus_evidence": [{"path": [], "text": "On Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4."}, {"path": [], "text": "Mina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4."}], "policy_evidence": [{"path": [], "text": "Mina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”"}, {"path": [], "text": "Policy defines mastery as 4/5 plus 2/3 diagnostic items correct."}, {"path": [], "text": "Nonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2."}, {"path": [], "text": "Urgency runs from 1 (lowest) to 3 (highest)."}], "rules": [{"justification": "A score of 2/5 establishes nonmastery. Exhibiting the stated pattern on at least two diagnostic items meets Mina’s condition, so policy requires the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3 rather than the proposed instructor conference at urgency 2.", "target": "false", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "supported"}]}, {"justification": "A score of 2/5 establishes nonmastery. Explicit refutation of exhibiting the pattern on at least two diagnostic items establishes that Mina’s condition is not met, so policy requires the instructor correction conference at urgency 2, matching the proposal.", "target": "true", "when": [{"atom_id": "quiz_score_2_of_5", "state": "supported"}, {"atom_id": "pattern_on_at_least_2_diagnostic_items", "state": "refuted"}]}]}, "verified_pair": {"left": "On Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.", "negative_left": "On Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.", "negative_right": "Mina selected the first listed option for diagnostic item 1, the second listed option for diagnostic item 2, and the first listed option for diagnostic item 4.", "right": "Mina selected the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-063-006", "id": "fast-41-diverse-063-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — Mina is not mastered, but she must be routed to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3.", "true": "Yes — the proposed not-mastered classification, instructor conference, and urgency 2 assignment are all correct."}, "instructions": "Decide whether this proposed triage is correct: mark Mina as not mastered, route her to the course instructor’s correction conference, and assign urgency 2. Answer Yes or No according to the stated policy.", "type": "noul"}}, "state": "Mina’s five-item fraction quiz was scored against the instructor’s key. The record shows exactly two correct responses: items 3 and 5. Items 1, 2, and 4 were designated diagnostic items for comparing fractions, and the response sheet preserved the listed-option order for those items.\n\nOn Mina’s five-item fraction quiz, the fraction having the larger denominator was the second listed option for diagnostic item 1, the first listed option for diagnostic item 2, and the second listed option for diagnostic item 4.\nMina selected the first listed option for diagnostic item 1, the second listed option for diagnostic item 2, and the first listed option for diagnostic item 4.\n\nMina said, “If my quiz shows that pattern on at least two diagnostic items, send me to a tutor; otherwise, give me your correction conference.”\nPolicy defines mastery as 4/5 plus 2/3 diagnostic items correct.\nNonmastery with Mina’s condition met routes to the learning support tutor’s Fraction-Strip Comparison Lab at urgency 3; otherwise it routes to the instructor conference at urgency 2.\nUrgency runs from 1 (lowest) to 3 (highest)."}, "method": "c2d", "provenance": {"source_id": "diverse-063", "source_is_synthetic": true, "source_sha256": "1ab2940f7e11265b1fc7b64fecf6943e6d0f6fc9c6402c436749bbc8483bd245", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the verbatim questions and Friday-current/Tuesday-prior bindings, and the two focus spans are complete factual sentences: “The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.” and “The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.” The counterfactual changes only the Tuesday observation, remains coherent with the unchanged records, and contains no decision label or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Assessment record\",\"text\":\"The authoritative answer key records that the learner answered exactly 3 of 5 items correctly on Friday’s decimal-comparison quiz.\"},{\"speaker\":\"Assessment record\",\"text\":\"The two incorrect quiz responses were attributed to the same incorrect rule.\"},{\"speaker\":\"Review note\",\"text\":\"The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.\"},{\"speaker\":\"Lesson-record note\",\"text\":\"The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.\"},{\"speaker\":\"Course instructor\",\"text\":\"The Friday quiz is the current assessment, and Tuesday’s entry is the prior lesson record for this learner. Both records were reviewed for routing.\"},{\"speaker\":\"Learning support tutor\",\"text\":\"The decimal-grid comparison activity is available if the review calls for additional support.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}, {"path": ["3", "text"], "text": "The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.", "negative_left": "The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.", "negative_right": "The Tuesday prior lesson record identifies the learner's misconception as treating 0.5 as greater than 0.45 because 5 is greater than 45.", "right": "The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-065-003", "id": "fast-41-diverse-065-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Assessment record", "text": "The authoritative answer key records that the learner answered exactly 3 of 5 items correctly on Friday’s decimal-comparison quiz."}, {"speaker": "Assessment record", "text": "The two incorrect quiz responses were attributed to the same incorrect rule."}, {"speaker": "Review note", "text": "The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}, {"speaker": "Lesson-record note", "text": "The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}, {"speaker": "Course instructor", "text": "The Friday quiz is the current assessment, and Tuesday’s entry is the prior lesson record for this learner. Both records were reviewed for routing."}, {"speaker": "Learning support tutor", "text": "The decimal-grid comparison activity is available if the review calls for additional support."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the verbatim questions and Friday-current/Tuesday-prior bindings, and the two focus spans are complete factual sentences: “The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.” and “The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.” The counterfactual changes only the Tuesday observation, remains coherent with the unchanged records, and contains no decision label or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Assessment record\",\"text\":\"The authoritative answer key records that the learner answered exactly 3 of 5 items correctly on Friday’s decimal-comparison quiz.\"},{\"speaker\":\"Assessment record\",\"text\":\"The two incorrect quiz responses were attributed to the same incorrect rule.\"},{\"speaker\":\"Review note\",\"text\":\"The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.\"},{\"speaker\":\"Lesson-record note\",\"text\":\"The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.\"},{\"speaker\":\"Course instructor\",\"text\":\"The Friday quiz is the current assessment, and Tuesday’s entry is the prior lesson record for this learner. Both records were reviewed for routing.\"},{\"speaker\":\"Learning support tutor\",\"text\":\"The decimal-grid comparison activity is available if the review calls for additional support.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}, {"path": ["3", "text"], "text": "The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.", "negative_left": "The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5.", "negative_right": "The Tuesday prior lesson record identifies the learner's misconception as treating 0.5 as greater than 0.45 because 5 is greater than 45.", "right": "The Tuesday prior lesson record identifies the learner's misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-065-003", "id": "fast-41-diverse-065-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Assessment record", "text": "The authoritative answer key records that the learner answered exactly 3 of 5 items correctly on Friday’s decimal-comparison quiz."}, {"speaker": "Assessment record", "text": "The two incorrect quiz responses were attributed to the same incorrect rule."}, {"speaker": "Review note", "text": "The authoritative answer review identifies the Friday quiz's repeated misconception as treating 0.45 as greater than 0.5 because 45 is greater than 5."}, {"speaker": "Lesson-record note", "text": "The Tuesday prior lesson record identifies the learner's misconception as treating 0.5 as greater than 0.45 because 5 is greater than 45."}, {"speaker": "Course instructor", "text": "The Friday quiz is the current assessment, and Tuesday’s entry is the prior lesson record for this learner. Both records were reviewed for routing."}, {"speaker": "Learning support tutor", "text": "The decimal-grid comparison activity is available if the review calls for additional support."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy. Both contexts retain the Friday quiz, Tuesday record, learner, and support-route bindings. The two evidence spans are complete factual sentences. The counterfactual changes only Tuesday’s documented misconception and creates no contradiction. Neither context states a level, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Course instructor\",\"text\":\"The authoritative answer key records exactly 3 correct answers out of 5 on Friday’s decimal-comparison quiz. The two incorrect responses were 0.58 > 0.6 and 0.35 > 0.4; the other three responses match the key.\"},{\"speaker\":\"Learning support tutor\",\"text\":\"The two incorrect responses use the same incorrect comparison rule, as shown by the learner’s written explanation for each error.\"},{\"speaker\":\"Record note\",\"text\":\"The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number.\"},{\"speaker\":\"Record note\",\"text\":\"The learner’s Tuesday prior lesson record documents the misconception that more digits means a larger number.\"},{\"speaker\":\"Course instructor\",\"text\":\"Tuesday’s record is a four-item exit ticket from comparing tenths with hundredths. The instructor is comparing the dated records under the supplied routing policy, and the tutor has a decimal-grid comparison activity ready if support is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number."}, {"path": ["3", "text"], "text": "The learner’s Tuesday prior lesson record documents the misconception that more digits means a larger number."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number.", "negative_left": "The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number.", "negative_right": "The learner’s Tuesday prior lesson record documents the misconception that the first digit means a larger number.", "right": "The learner’s Tuesday prior lesson record documents the misconception that more digits means a larger number."}, "verifier_independent_model": false}, "family": "fast-41-diverse-065-004", "id": "fast-41-diverse-065-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Course instructor", "text": "The authoritative answer key records exactly 3 correct answers out of 5 on Friday’s decimal-comparison quiz. The two incorrect responses were 0.58 > 0.6 and 0.35 > 0.4; the other three responses match the key."}, {"speaker": "Learning support tutor", "text": "The two incorrect responses use the same incorrect comparison rule, as shown by the learner’s written explanation for each error."}, {"speaker": "Record note", "text": "The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number."}, {"speaker": "Record note", "text": "The learner’s Tuesday prior lesson record documents the misconception that more digits means a larger number."}, {"speaker": "Course instructor", "text": "Tuesday’s record is a four-item exit ticket from comparing tenths with hundredths. The instructor is comparing the dated records under the supplied routing policy, and the tutor has a decimal-grid comparison activity ready if support is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy. Both contexts retain the Friday quiz, Tuesday record, learner, and support-route bindings. The two evidence spans are complete factual sentences. The counterfactual changes only Tuesday’s documented misconception and creates no contradiction. Neither context states a level, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Course instructor\",\"text\":\"The authoritative answer key records exactly 3 correct answers out of 5 on Friday’s decimal-comparison quiz. The two incorrect responses were 0.58 > 0.6 and 0.35 > 0.4; the other three responses match the key.\"},{\"speaker\":\"Learning support tutor\",\"text\":\"The two incorrect responses use the same incorrect comparison rule, as shown by the learner’s written explanation for each error.\"},{\"speaker\":\"Record note\",\"text\":\"The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number.\"},{\"speaker\":\"Record note\",\"text\":\"The learner’s Tuesday prior lesson record documents the misconception that more digits means a larger number.\"},{\"speaker\":\"Course instructor\",\"text\":\"Tuesday’s record is a four-item exit ticket from comparing tenths with hundredths. The instructor is comparing the dated records under the supplied routing policy, and the tutor has a decimal-grid comparison activity ready if support is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number."}, {"path": ["3", "text"], "text": "The learner’s Tuesday prior lesson record documents the misconception that more digits means a larger number."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number.", "negative_left": "The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number.", "negative_right": "The learner’s Tuesday prior lesson record documents the misconception that the first digit means a larger number.", "right": "The learner’s Tuesday prior lesson record documents the misconception that more digits means a larger number."}, "verifier_independent_model": false}, "family": "fast-41-diverse-065-004", "id": "fast-41-diverse-065-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Course instructor", "text": "The authoritative answer key records exactly 3 correct answers out of 5 on Friday’s decimal-comparison quiz. The two incorrect responses were 0.58 > 0.6 and 0.35 > 0.4; the other three responses match the key."}, {"speaker": "Learning support tutor", "text": "The two incorrect responses use the same incorrect comparison rule, as shown by the learner’s written explanation for each error."}, {"speaker": "Record note", "text": "The learner’s Friday decimal-comparison quiz documents the misconception that more digits means a larger number."}, {"speaker": "Record note", "text": "The learner’s Tuesday prior lesson record documents the misconception that the first digit means a larger number."}, {"speaker": "Course instructor", "text": "Tuesday’s record is a four-item exit ticket from comparing tenths with hundredths. The instructor is comparing the dated records under the supplied routing policy, and the tutor has a decimal-grid comparison activity ready if support is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing routing policy. The focus evidence consists of two complete factual sentences: \"The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17.\" \"The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17.\" The counterfactual coherently changes only the Tuesday record from MC-17 to MC-23 without duplicating or contradicting measurements within that context. Neither context embeds a gold answer, rule table, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Course instructor\",\"text\":\"The Friday decimal-comparison quiz has five scored questions, and the authoritative answer key confirms that exactly three of the learner’s five answers are correct.\"},{\"speaker\":\"Learning support tutor\",\"text\":\"The two incorrect responses involving 0.58 and 0.6, and 0.35 and 0.4, were both explained by the learner’s rule that a decimal with more digits is larger. The other three responses were correct.\"},{\"speaker\":\"Records administrator\",\"text\":\"The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17.\"},{\"speaker\":\"Records administrator\",\"text\":\"The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17.\"},{\"speaker\":\"Course instructor\",\"text\":\"The quiz key is the controlling scoring record, and the learner is being reviewed for the applicable decimal-comparison support route.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17."}, {"path": ["3", "text"], "text": "The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17.", "negative_left": "The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17.", "negative_right": "The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-23.", "right": "The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-065-007", "id": "fast-41-diverse-065-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Course instructor", "text": "The Friday decimal-comparison quiz has five scored questions, and the authoritative answer key confirms that exactly three of the learner’s five answers are correct."}, {"speaker": "Learning support tutor", "text": "The two incorrect responses involving 0.58 and 0.6, and 0.35 and 0.4, were both explained by the learner’s rule that a decimal with more digits is larger. The other three responses were correct."}, {"speaker": "Records administrator", "text": "The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17."}, {"speaker": "Records administrator", "text": "The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17."}, {"speaker": "Course instructor", "text": "The quiz key is the controlling scoring record, and the learner is being reviewed for the applicable decimal-comparison support route."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing routing policy. The focus evidence consists of two complete factual sentences: \"The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17.\" \"The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17.\" The counterfactual coherently changes only the Tuesday record from MC-17 to MC-23 without duplicating or contradicting measurements within that context. Neither context embeds a gold answer, rule table, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Course instructor\",\"text\":\"The Friday decimal-comparison quiz has five scored questions, and the authoritative answer key confirms that exactly three of the learner’s five answers are correct.\"},{\"speaker\":\"Learning support tutor\",\"text\":\"The two incorrect responses involving 0.58 and 0.6, and 0.35 and 0.4, were both explained by the learner’s rule that a decimal with more digits is larger. The other three responses were correct.\"},{\"speaker\":\"Records administrator\",\"text\":\"The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17.\"},{\"speaker\":\"Records administrator\",\"text\":\"The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17.\"},{\"speaker\":\"Course instructor\",\"text\":\"The quiz key is the controlling scoring record, and the learner is being reviewed for the applicable decimal-comparison support route.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17."}, {"path": ["3", "text"], "text": "The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17.", "negative_left": "The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17.", "negative_right": "The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-23.", "right": "The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-065-007", "id": "fast-41-diverse-065-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Course instructor", "text": "The Friday decimal-comparison quiz has five scored questions, and the authoritative answer key confirms that exactly three of the learner’s five answers are correct."}, {"speaker": "Learning support tutor", "text": "The two incorrect responses involving 0.58 and 0.6, and 0.35 and 0.4, were both explained by the learner’s rule that a decimal with more digits is larger. The other three responses were correct."}, {"speaker": "Records administrator", "text": "The authoritative record for the learner’s Friday decimal-comparison quiz identifies the documented misconception as misconception code MC-17."}, {"speaker": "Records administrator", "text": "The learner’s Tuesday prior lesson record identifies the documented misconception as misconception code MC-23."}, {"speaker": "Course instructor", "text": "The quiz key is the controlling scoring record, and the learner is being reviewed for the applicable decimal-comparison support route."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and request bindings, provide exactly two complete factual evidence sentences, and contain no embedded answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns Leo’s recorded grade rather than a policy conclusion. The base and counter assignments can both occur under the policy while differing only in whether the grade is at least C+; file completeness, support, unconditional intent, and a below-75 placement score can remain fixed. The policy evidence preserves the substantive state-originating prerequisite/exception and future-dependent-request rules; observations about Leo’s particular case need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes an unconditional request, a complete file, at least a C+ grade, and written instructor support. Those facts are sufficient for approve_ready_registrar under the stated criterion. A placement score below 75 does not negate the C+-plus-support exception.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an unconditional request and refutes a grade of at least C+, so the request lacks the minimum C+ required by the denial criterion. The below-75 placement score also excludes qualification through the normal placement route. Instructor support does not cure the missing minimum grade.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Leo’s request to enter Applied Statistics is unconditional at Mina’s review."}, {"id": "a2", "statement": "Leo’s submitted standard file for the Applied Statistics request is complete at Mina’s review."}, {"id": "a3", "statement": "The Statistics I grade recorded for Leo ranks at least C+ on the registrar’s letter-grade scale."}, {"id": "a4", "statement": "Leo’s submitted file contains written instructor support for his admission to Applied Statistics."}, {"id": "a5", "statement": "Leo’s placement score for Applied Statistics is below 75 at Mina’s review."}], "base_state_json": "\"Mina’s case note records that Leo’s request to enter Applied Statistics is unconditional, and his standard submission was complete when she reviewed it. The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points. The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C+, and every score above 79 to a higher letter grade. His file also contains written instructor support for admission. His Applied Statistics placement score is recorded as below 75 at the time of review. The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support. Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points."}, {"path": [], "text": "The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C+, and every score above 79 to a higher letter grade."}], "policy_evidence": [{"path": [], "text": "The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support."}, {"path": [], "text": "Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}], "rules": [{"justification": "The request is unconditional and complete, and Leo has both the minimum Statistics I grade and written instructor support required for exception approval. His sub-75 placement score does not defeat the expressly permitted C+-plus-support exception.", "target": "approve_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The request is unconditional, but Leo lacks the minimum C+ and also has no qualifying placement score; written instructor support alone cannot satisfy either the normal prerequisite or the exception requirements.", "target": "deny_not_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points.", "negative_left": "The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points.", "negative_right": "The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C-, and every score above 79 to a higher letter grade.", "right": "The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C+, and every score above 79 to a higher letter grade."}, "verifier_independent_model": false}, "family": "fast-41-diverse-073-005", "id": "fast-41-diverse-073-005-base", "input": {"questions": {"decision": {"criteria": {"approve_ready_registrar": "Recommend yes, rate Ready, and let registrar staff finalize. Use only for an unconditional request with a complete file, at least a C+ in Statistics I, and written instructor support.", "deny_not_ready_registrar": "Recommend no, rate Not Ready, and let registrar staff close the request. Use only for an unconditional request that lacks either the minimum C+ or instructor support.", "none_of_above": "Make no yes/no recommendation, rate Provisional, and route to the course coordinator. Use for a request whose intent depends on a future event or condition that has not yet resolved."}, "instructions": "Select the single processing outcome that matches the evidence and policy. Apply the conditional-intent rule even if the ordinary exception criteria would otherwise support a decision.", "type": "choice"}}, "state": "Mina’s case note records that Leo’s request to enter Applied Statistics is unconditional, and his standard submission was complete when she reviewed it. The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points. The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C+, and every score above 79 to a higher letter grade. His file also contains written instructor support for admission. His Applied Statistics placement score is recorded as below 75 at the time of review. The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support. Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}, "method": "c2d", "provenance": {"source_id": "diverse-073", "source_is_synthetic": true, "source_sha256": "3d8b5af4be06b90ac301446c5958d7d599067d0a272006d010e26fd0847ef133", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "approve_ready_registrar"}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and request bindings, provide exactly two complete factual evidence sentences, and contain no embedded answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns Leo’s recorded grade rather than a policy conclusion. The base and counter assignments can both occur under the policy while differing only in whether the grade is at least C+; file completeness, support, unconditional intent, and a below-75 placement score can remain fixed. The policy evidence preserves the substantive state-originating prerequisite/exception and future-dependent-request rules; observations about Leo’s particular case need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes an unconditional request, a complete file, at least a C+ grade, and written instructor support. Those facts are sufficient for approve_ready_registrar under the stated criterion. A placement score below 75 does not negate the C+-plus-support exception.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an unconditional request and refutes a grade of at least C+, so the request lacks the minimum C+ required by the denial criterion. The below-75 placement score also excludes qualification through the normal placement route. Instructor support does not cure the missing minimum grade.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Leo’s request to enter Applied Statistics is unconditional at Mina’s review."}, {"id": "a2", "statement": "Leo’s submitted standard file for the Applied Statistics request is complete at Mina’s review."}, {"id": "a3", "statement": "The Statistics I grade recorded for Leo ranks at least C+ on the registrar’s letter-grade scale."}, {"id": "a4", "statement": "Leo’s submitted file contains written instructor support for his admission to Applied Statistics."}, {"id": "a5", "statement": "Leo’s placement score for Applied Statistics is below 75 at Mina’s review."}], "base_state_json": "\"Mina’s case note records that Leo’s request to enter Applied Statistics is unconditional, and his standard submission was complete when she reviewed it. The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points. The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C+, and every score above 79 to a higher letter grade. His file also contains written instructor support for admission. His Applied Statistics placement score is recorded as below 75 at the time of review. The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support. Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points."}, {"path": [], "text": "The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C+, and every score above 79 to a higher letter grade."}], "policy_evidence": [{"path": [], "text": "The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support."}, {"path": [], "text": "Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}], "rules": [{"justification": "The request is unconditional and complete, and Leo has both the minimum Statistics I grade and written instructor support required for exception approval. His sub-75 placement score does not defeat the expressly permitted C+-plus-support exception.", "target": "approve_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The request is unconditional, but Leo lacks the minimum C+ and also has no qualifying placement score; written instructor support alone cannot satisfy either the normal prerequisite or the exception requirements.", "target": "deny_not_ready_registrar", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points.", "negative_left": "The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points.", "negative_right": "The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C-, and every score above 79 to a higher letter grade.", "right": "The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C+, and every score above 79 to a higher letter grade."}, "verifier_independent_model": false}, "family": "fast-41-diverse-073-005", "id": "fast-41-diverse-073-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"approve_ready_registrar": "Recommend yes, rate Ready, and let registrar staff finalize. Use only for an unconditional request with a complete file, at least a C+ in Statistics I, and written instructor support.", "deny_not_ready_registrar": "Recommend no, rate Not Ready, and let registrar staff close the request. Use only for an unconditional request that lacks either the minimum C+ or instructor support.", "none_of_above": "Make no yes/no recommendation, rate Provisional, and route to the course coordinator. Use for a request whose intent depends on a future event or condition that has not yet resolved."}, "instructions": "Select the single processing outcome that matches the evidence and policy. Apply the conditional-intent rule even if the ordinary exception criteria would otherwise support a decision.", "type": "choice"}}, "state": "Mina’s case note records that Leo’s request to enter Applied Statistics is unconditional, and his standard submission was complete when she reviewed it. The registrar’s 14 June 2026 transcript record assigns Leo a Statistics I score of 77 points. The registrar’s 2026 letter-grade scale assigns scores of 75 through 79 to C-, and every score above 79 to a higher letter grade. His file also contains written instructor support for admission. His Applied Statistics placement score is recorded as below 75 at the time of review. The normal prerequisite is Statistics I with B or a placement score of 75. Exception policy permits approval with a C+ and written instructor support. Policy requires future-dependent requests to be held for the course coordinator, with no yes/no recommendation and readiness rated Provisional until the condition resolves."}, "method": "c2d", "provenance": {"source_id": "diverse-073", "source_is_synthetic": true, "source_sha256": "3d8b5af4be06b90ac301446c5958d7d599067d0a272006d010e26fd0847ef133", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "deny_not_ready_registrar"}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and policy, preserve Mira's request bindings, use complete factual evidence, and remain coherent without embedding an answer; the required quotes are \"Mira's August current-term placement result carries record identifier P-417.\" and \"Record P-417 reports 86 points, and the placement threshold is 80 points.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual placement-threshold result rather than a policy conclusion. The base and counter assignments differ only on that focus and can be realized by changing the placement score from at least 80 to below 80. The policy evidence preserves the substantive state-originating policy, including latest-update precedence; the unchanged questions automatically preserve the outcome criteria. Both proposed rules include the required outcome exclusions and routing conditions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes completeness, passing grade and placement thresholds, absence of current instructor support, and non-equivalent coursework. These conditions are sufficient for no recommendation, coordinator routing, and Moderate readiness.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completeness, failure of the placement threshold, and non-equivalent coursework as an unusual coordinator-routing condition. These conditions are sufficient for no recommendation, coordinator routing, and Low readiness.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira's final transcript is present in the file."}, {"id": "a2", "statement": "Mira's current-term placement result is present in the file."}, {"id": "a3", "statement": "Mira's latest instructor position is present in the file."}, {"id": "a4", "statement": "Mira's final-transcript grade for Applied Programming is B or better."}, {"id": "a5", "statement": "Applied Programming is related to the study for which Mira seeks entry."}, {"id": "a6", "statement": "Mira's August current-term placement result is at least 80."}, {"id": "a7", "statement": "Mira's latest instructor position supports entry."}, {"id": "a8", "statement": "Mira's Applied Programming coursework is non-equivalent."}], "base_state_json": "{\"context\":\"Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete.\",\"evidence\":[\"Mira's August current-term placement result carries record identifier P-417.\",\"Record P-417 reports 86 points, and the placement threshold is 80 points.\",\"Mira's final transcript is present and records a B+ in Applied Programming.\",\"Applied Programming is related to the study for which Mira seeks entry.\",\"The transcript classifies Mira's Applied Programming coursework as non-equivalent.\",\"Mira's current-term placement result and latest instructor position are also present in the file.\",\"The latest instructor position withdraws support for entry after a lab diagnostic; an earlier note had supported entry, and updates supersede earlier notes.\"],\"request\":\"Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["evidence", "0"], "text": "Mira's August current-term placement result carries record identifier P-417."}, {"path": ["evidence", "1"], "text": "Record P-417 reports 86 points, and the placement threshold is 80 points."}], "policy_evidence": [{"path": ["context"], "text": "Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete."}, {"path": ["evidence", "3"], "text": "updates supersede earlier notes."}, {"path": ["request"], "text": "Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence."}], "rules": [{"justification": "All required documents are present, both academic thresholds pass, current instructor support is absent, and non-equivalent coursework requires coordinator routing.", "target": "complete_no_coordinator_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All required documents are present, the placement threshold fails, and non-equivalent coursework supplies the unusual condition requiring coordinator routing.", "target": "complete_no_coordinator_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "Mira's August current-term placement result carries record identifier P-417.", "negative_left": "Mira's August current-term placement result carries record identifier P-417.", "negative_right": "Record P-417 reports 74 points, and the placement threshold is 80 points.", "right": "Record P-417 reports 86 points, and the placement threshold is 80 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-074-010", "id": "fast-41-diverse-074-010-base", "input": {"questions": {"decision": {"criteria": {"complete_no_coordinator_low": "Use when the file is complete but the grade or placement threshold fails, and an unusual routing condition exists: recommend no, route to the course coordinator, and rate Low.", "complete_no_coordinator_moderate": "Use when the file is complete and academic thresholds pass, but current instructor support is absent, with non-equivalent coursework or a changed position: recommend no, route to the course coordinator, and rate Moderate.", "complete_yes_registrar_high": "Use when the file is complete, every approval test passes, and no unusual routing condition exists: recommend yes, route to registrar staff, and rate High.", "incomplete_adviser_indeterminate": "Use when any required document is absent: mark incomplete, make no approval recommendation, route to the academic adviser, and rate Indeterminate."}, "instructions": "Choose the single option whose complete outcome rubric matches the policy and evidence as currently updated.", "type": "choice"}}, "state": {"context": "Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete.", "evidence": ["Mira's August current-term placement result carries record identifier P-417.", "Record P-417 reports 86 points, and the placement threshold is 80 points.", "Mira's final transcript is present and records a B+ in Applied Programming.", "Applied Programming is related to the study for which Mira seeks entry.", "The transcript classifies Mira's Applied Programming coursework as non-equivalent.", "Mira's current-term placement result and latest instructor position are also present in the file.", "The latest instructor position withdraws support for entry after a lab diagnostic; an earlier note had supported entry, and updates supersede earlier notes."], "request": "Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-074", "source_is_synthetic": true, "source_sha256": "f6b73c0c2df392cc17b3d3a55974caeaa07084ae400a30235bea57a187f39bcf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_no_coordinator_moderate"}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and policy, preserve Mira's request bindings, use complete factual evidence, and remain coherent without embedding an answer; the required quotes are \"Mira's August current-term placement result carries record identifier P-417.\" and \"Record P-417 reports 86 points, and the placement threshold is 80 points.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual placement-threshold result rather than a policy conclusion. The base and counter assignments differ only on that focus and can be realized by changing the placement score from at least 80 to below 80. The policy evidence preserves the substantive state-originating policy, including latest-update precedence; the unchanged questions automatically preserve the outcome criteria. Both proposed rules include the required outcome exclusions and routing conditions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes completeness, passing grade and placement thresholds, absence of current instructor support, and non-equivalent coursework. These conditions are sufficient for no recommendation, coordinator routing, and Moderate readiness.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completeness, failure of the placement threshold, and non-equivalent coursework as an unusual coordinator-routing condition. These conditions are sufficient for no recommendation, coordinator routing, and Low readiness.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira's final transcript is present in the file."}, {"id": "a2", "statement": "Mira's current-term placement result is present in the file."}, {"id": "a3", "statement": "Mira's latest instructor position is present in the file."}, {"id": "a4", "statement": "Mira's final-transcript grade for Applied Programming is B or better."}, {"id": "a5", "statement": "Applied Programming is related to the study for which Mira seeks entry."}, {"id": "a6", "statement": "Mira's August current-term placement result is at least 80."}, {"id": "a7", "statement": "Mira's latest instructor position supports entry."}, {"id": "a8", "statement": "Mira's Applied Programming coursework is non-equivalent."}], "base_state_json": "{\"context\":\"Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete.\",\"evidence\":[\"Mira's August current-term placement result carries record identifier P-417.\",\"Record P-417 reports 86 points, and the placement threshold is 80 points.\",\"Mira's final transcript is present and records a B+ in Applied Programming.\",\"Applied Programming is related to the study for which Mira seeks entry.\",\"The transcript classifies Mira's Applied Programming coursework as non-equivalent.\",\"Mira's current-term placement result and latest instructor position are also present in the file.\",\"The latest instructor position withdraws support for entry after a lab diagnostic; an earlier note had supported entry, and updates supersede earlier notes.\"],\"request\":\"Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["evidence", "0"], "text": "Mira's August current-term placement result carries record identifier P-417."}, {"path": ["evidence", "1"], "text": "Record P-417 reports 86 points, and the placement threshold is 80 points."}], "policy_evidence": [{"path": ["context"], "text": "Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete."}, {"path": ["evidence", "3"], "text": "updates supersede earlier notes."}, {"path": ["request"], "text": "Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence."}], "rules": [{"justification": "All required documents are present, both academic thresholds pass, current instructor support is absent, and non-equivalent coursework requires coordinator routing.", "target": "complete_no_coordinator_moderate", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All required documents are present, the placement threshold fails, and non-equivalent coursework supplies the unusual condition requiring coordinator routing.", "target": "complete_no_coordinator_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "Mira's August current-term placement result carries record identifier P-417.", "negative_left": "Mira's August current-term placement result carries record identifier P-417.", "negative_right": "Record P-417 reports 74 points, and the placement threshold is 80 points.", "right": "Record P-417 reports 86 points, and the placement threshold is 80 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-074-010", "id": "fast-41-diverse-074-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_no_coordinator_low": "Use when the file is complete but the grade or placement threshold fails, and an unusual routing condition exists: recommend no, route to the course coordinator, and rate Low.", "complete_no_coordinator_moderate": "Use when the file is complete and academic thresholds pass, but current instructor support is absent, with non-equivalent coursework or a changed position: recommend no, route to the course coordinator, and rate Moderate.", "complete_yes_registrar_high": "Use when the file is complete, every approval test passes, and no unusual routing condition exists: recommend yes, route to registrar staff, and rate High.", "incomplete_adviser_indeterminate": "Use when any required document is absent: mark incomplete, make no approval recommendation, route to the academic adviser, and rate Indeterminate."}, "instructions": "Choose the single option whose complete outcome rubric matches the policy and evidence as currently updated.", "type": "choice"}}, "state": {"context": "Policy: complete means a final transcript, current-term placement result, and latest instructor position are present. Approve only with B or better in related study, placement ≥80, and current support. Incomplete files go to the academic adviser; ordinary approvals to registrar staff; non-equivalent coursework or a changed instructor position goes to the course coordinator. Readiness is High if all approval tests pass, Moderate if complete and academic thresholds pass without current support, Low if complete but an academic threshold fails, and Indeterminate if incomplete.", "evidence": ["Mira's August current-term placement result carries record identifier P-417.", "Record P-417 reports 74 points, and the placement threshold is 80 points.", "Mira's final transcript is present and records a B+ in Applied Programming.", "Applied Programming is related to the study for which Mira seeks entry.", "The transcript classifies Mira's Applied Programming coursework as non-equivalent.", "Mira's current-term placement result and latest instructor position are also present in the file.", "The latest instructor position withdraws support for entry after a lab diagnostic; an earlier note had supported entry, and updates supersede earlier notes."], "request": "Determine completeness, yes/no recommendation, routing, and readiness using the latest evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-074", "source_is_synthetic": true, "source_sha256": "f6b73c0c2df392cc17b3d3a55974caeaa07084ae400a30235bea57a187f39bcf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_no_coordinator_low"}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the EDU 340, registrar-routing, recommendation, and review bindings; the evidence quotes are complete factual sentences: \"The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.\" and \"The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.\"; the counterfactual coherently changes the second report to 79 while retaining the reported maximum of 80, and neither context adds answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.\"},{\"speaker\":\"Case note\",\"text\":\"The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.\"},{\"speaker\":\"File inventory\",\"text\":\"The complete EDU 340 exception submission contains no quantitative placement-test score reports beyond the testing portal report and the dated testing-center export. The file summary lists 80 as the maximum reported value.\"},{\"speaker\":\"Verification log\",\"text\":\"The score in the testing portal report has verified status.\"},{\"speaker\":\"Instructor note\",\"text\":\"The submitted instructor note is signed and recommends admission.\"},{\"speaker\":\"Policy\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Policy\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}, {"path": ["1", "text"], "text": "The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.", "negative_left": "The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.", "negative_right": "The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 79 points.", "right": "The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-002", "id": "fast-41-diverse-075-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}, {"speaker": "Case note", "text": "The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}, {"speaker": "File inventory", "text": "The complete EDU 340 exception submission contains no quantitative placement-test score reports beyond the testing portal report and the dated testing-center export. The file summary lists 80 as the maximum reported value."}, {"speaker": "Verification log", "text": "The score in the testing portal report has verified status."}, {"speaker": "Instructor note", "text": "The submitted instructor note is signed and recommends admission."}, {"speaker": "Policy", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Policy", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the EDU 340, registrar-routing, recommendation, and review bindings; the evidence quotes are complete factual sentences: \"The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.\" and \"The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.\"; the counterfactual coherently changes the second report to 79 while retaining the reported maximum of 80, and neither context adds answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.\"},{\"speaker\":\"Case note\",\"text\":\"The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.\"},{\"speaker\":\"File inventory\",\"text\":\"The complete EDU 340 exception submission contains no quantitative placement-test score reports beyond the testing portal report and the dated testing-center export. The file summary lists 80 as the maximum reported value.\"},{\"speaker\":\"Verification log\",\"text\":\"The score in the testing portal report has verified status.\"},{\"speaker\":\"Instructor note\",\"text\":\"The submitted instructor note is signed and recommends admission.\"},{\"speaker\":\"Policy\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Policy\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}, {"path": ["1", "text"], "text": "The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.", "negative_left": "The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points.", "negative_right": "The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 79 points.", "right": "The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-002", "id": "fast-41-diverse-075-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The testing portal report dated 14 March 2026 records the student's EDU 340 quantitative placement-test score as 80 points."}, {"speaker": "Case note", "text": "The dated testing-center export issued 15 March 2026 records the student's EDU 340 quantitative placement-test score as 79 points."}, {"speaker": "File inventory", "text": "The complete EDU 340 exception submission contains no quantitative placement-test score reports beyond the testing portal report and the dated testing-center export. The file summary lists 80 as the maximum reported value."}, {"speaker": "Verification log", "text": "The score in the testing portal report has verified status."}, {"speaker": "Instructor note", "text": "The submitted instructor note is signed and recommends admission."}, {"speaker": "Policy", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Policy", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the EDU 340 scope, registrar-routing question, policies, and bindings. The two evidence quotes are complete factual sentences, and changing the export score to 79 coherently creates a boundary conflict without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The EDU 340 exception packet is complete. It contains only the identified portal report and dated testing-center export as quantitative placement-test records. The portal report has verified status. The instructor note is signed and recommends admission.\"},{\"speaker\":\"Testing portal record\",\"text\":\"The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80.\"},{\"speaker\":\"Testing-center export\",\"text\":\"The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 80.\"},{\"speaker\":\"Registrar policy\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Coordinator policy\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80."}, {"path": ["2", "text"], "text": "The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80.", "negative_left": "The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80.", "negative_right": "The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 79.", "right": "The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 80."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-004", "id": "fast-41-diverse-075-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The EDU 340 exception packet is complete. It contains only the identified portal report and dated testing-center export as quantitative placement-test records. The portal report has verified status. The instructor note is signed and recommends admission."}, {"speaker": "Testing portal record", "text": "The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80."}, {"speaker": "Testing-center export", "text": "The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 80."}, {"speaker": "Registrar policy", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Coordinator policy", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the EDU 340 scope, registrar-routing question, policies, and bindings. The two evidence quotes are complete factual sentences, and changing the export score to 79 coherently creates a boundary conflict without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The EDU 340 exception packet is complete. It contains only the identified portal report and dated testing-center export as quantitative placement-test records. The portal report has verified status. The instructor note is signed and recommends admission.\"},{\"speaker\":\"Testing portal record\",\"text\":\"The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80.\"},{\"speaker\":\"Testing-center export\",\"text\":\"The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 80.\"},{\"speaker\":\"Registrar policy\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Coordinator policy\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80."}, {"path": ["2", "text"], "text": "The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80.", "negative_left": "The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80.", "negative_right": "The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 79.", "right": "The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 80."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-004", "id": "fast-41-diverse-075-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The EDU 340 exception packet is complete. It contains only the identified portal report and dated testing-center export as quantitative placement-test records. The portal report has verified status. The instructor note is signed and recommends admission."}, {"speaker": "Testing portal record", "text": "The testing portal report dated 2026-04-03 records the student's EDU 340 quantitative placement-test score as 80."}, {"speaker": "Testing-center export", "text": "The testing-center export dated 2026-04-04 records the student's EDU 340 quantitative placement-test score as 79."}, {"speaker": "Registrar policy", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Coordinator policy", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policies, preserve the EDU 340, registrar, recommendation, and routing bindings, use two complete factual evidence sentences, and contain no embedded answer or labeling instructions; the counterfactual consistently changes the export score to 79 while keeping the stated maximum of 80.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The complete EDU 340 exception packet contains exactly two quantitative placement-test reports: the testing portal report and the dated testing-center export. The packet also contains a signed instructor note recommending admission. The portal report has verified status, and the maximum of the two reported scores is 80.\"},{\"speaker\":\"Testing record\",\"text\":\"For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80.\"},{\"speaker\":\"Testing record\",\"text\":\"For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 80.\"},{\"speaker\":\"Policy\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Policy\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80."}, {"path": ["2", "text"], "text": "For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80.", "negative_left": "For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80.", "negative_right": "For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 79.", "right": "For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 80."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-006", "id": "fast-41-diverse-075-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The complete EDU 340 exception packet contains exactly two quantitative placement-test reports: the testing portal report and the dated testing-center export. The packet also contains a signed instructor note recommending admission. The portal report has verified status, and the maximum of the two reported scores is 80."}, {"speaker": "Testing record", "text": "For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80."}, {"speaker": "Testing record", "text": "For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 80."}, {"speaker": "Policy", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Policy", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policies, preserve the EDU 340, registrar, recommendation, and routing bindings, use two complete factual evidence sentences, and contain no embedded answer or labeling instructions; the counterfactual consistently changes the export score to 79 while keeping the stated maximum of 80.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The complete EDU 340 exception packet contains exactly two quantitative placement-test reports: the testing portal report and the dated testing-center export. The packet also contains a signed instructor note recommending admission. The portal report has verified status, and the maximum of the two reported scores is 80.\"},{\"speaker\":\"Testing record\",\"text\":\"For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80.\"},{\"speaker\":\"Testing record\",\"text\":\"For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 80.\"},{\"speaker\":\"Policy\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Policy\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80."}, {"path": ["2", "text"], "text": "For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80.", "negative_left": "For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80.", "negative_right": "For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 79.", "right": "For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 80."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-006", "id": "fast-41-diverse-075-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The complete EDU 340 exception packet contains exactly two quantitative placement-test reports: the testing portal report and the dated testing-center export. The packet also contains a signed instructor note recommending admission. The portal report has verified status, and the maximum of the two reported scores is 80."}, {"speaker": "Testing record", "text": "For the student's EDU 340 quantitative placement test, the testing portal report dated 2026-08-14 records a score of 80."}, {"speaker": "Testing record", "text": "For the student's EDU 340 quantitative placement test, the testing-center export dated 2026-08-15 records a score of 79."}, {"speaker": "Policy", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Policy", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the EDU 340 scope, score threshold, recommendation requirement, conflict routing, registrar restriction, and readiness rubric. The question’s entity, request, and document/date bindings remain unchanged. The evidence quotes are complete factual sentences: \"The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026.\" and \"The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80.\" The counterfactual coherently changes only the export score to 79, creating a conflict that straddles 80. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The complete EDU 340 exception file contains only the testing portal report, the dated testing-center export, and the instructor note; no other quantitative placement-test score report is present. The portal entry has verified status. The instructor note is signed and recommends admission.\"},{\"speaker\":\"Testing portal report\",\"text\":\"The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026.\"},{\"speaker\":\"Testing-center export\",\"text\":\"The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80.\"},{\"speaker\":\"Routing policy\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Readiness policy\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026."}, {"path": ["2", "text"], "text": "The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026.", "negative_left": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026.", "negative_right": "The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 79.", "right": "The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-007", "id": "fast-41-diverse-075-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The complete EDU 340 exception file contains only the testing portal report, the dated testing-center export, and the instructor note; no other quantitative placement-test score report is present. The portal entry has verified status. The instructor note is signed and recommends admission."}, {"speaker": "Testing portal report", "text": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026."}, {"speaker": "Testing-center export", "text": "The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80."}, {"speaker": "Routing policy", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Readiness policy", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the EDU 340 scope, score threshold, recommendation requirement, conflict routing, registrar restriction, and readiness rubric. The question’s entity, request, and document/date bindings remain unchanged. The evidence quotes are complete factual sentences: \"The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026.\" and \"The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80.\" The counterfactual coherently changes only the export score to 79, creating a conflict that straddles 80. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The complete EDU 340 exception file contains only the testing portal report, the dated testing-center export, and the instructor note; no other quantitative placement-test score report is present. The portal entry has verified status. The instructor note is signed and recommends admission.\"},{\"speaker\":\"Testing portal report\",\"text\":\"The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026.\"},{\"speaker\":\"Testing-center export\",\"text\":\"The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80.\"},{\"speaker\":\"Routing policy\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Readiness policy\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026."}, {"path": ["2", "text"], "text": "The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026.", "negative_left": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026.", "negative_right": "The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 79.", "right": "The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 80."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-007", "id": "fast-41-diverse-075-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The complete EDU 340 exception file contains only the testing portal report, the dated testing-center export, and the instructor note; no other quantitative placement-test score report is present. The portal entry has verified status. The instructor note is signed and recommends admission."}, {"speaker": "Testing portal report", "text": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 on 14 March 2026."}, {"speaker": "Testing-center export", "text": "The testing-center export dated 15 March 2026 records the student's EDU 340 quantitative placement-test score as 79."}, {"speaker": "Routing policy", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Readiness policy", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and EDU 340/date bindings; the focus spans are exact complete factual sentences, the 79-versus-80 records form a coherent boundary conflict, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points.\"},{\"speaker\":\"Records clerk\",\"text\":\"The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 80 points.\"},{\"speaker\":\"Registrar\",\"text\":\"The complete exception file contains no quantitative placement-test score report beyond those two records. The comparison sheet lists 80 as the maximum of the reported scores, and the portal entry carries verified status.\"},{\"speaker\":\"Registrar\",\"text\":\"The instructor note submitted with the EDU 340 exception request is signed and recommends admission.\"},{\"speaker\":\"Policy file\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Policy file\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points."}, {"path": ["1", "text"], "text": "The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 80 points."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points.", "negative_left": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points.", "negative_right": "The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 79 points.", "right": "The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 80 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-009", "id": "fast-41-diverse-075-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points."}, {"speaker": "Records clerk", "text": "The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 80 points."}, {"speaker": "Registrar", "text": "The complete exception file contains no quantitative placement-test score report beyond those two records. The comparison sheet lists 80 as the maximum of the reported scores, and the portal entry carries verified status."}, {"speaker": "Registrar", "text": "The instructor note submitted with the EDU 340 exception request is signed and recommends admission."}, {"speaker": "Policy file", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Policy file", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and EDU 340/date bindings; the focus spans are exact complete factual sentences, the 79-versus-80 records form a coherent boundary conflict, and neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points.\"},{\"speaker\":\"Records clerk\",\"text\":\"The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 80 points.\"},{\"speaker\":\"Registrar\",\"text\":\"The complete exception file contains no quantitative placement-test score report beyond those two records. The comparison sheet lists 80 as the maximum of the reported scores, and the portal entry carries verified status.\"},{\"speaker\":\"Registrar\",\"text\":\"The instructor note submitted with the EDU 340 exception request is signed and recommends admission.\"},{\"speaker\":\"Policy file\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Policy file\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points."}, {"path": ["1", "text"], "text": "The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 80 points."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points.", "negative_left": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points.", "negative_right": "The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 79 points.", "right": "The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 80 points."}, "verifier_independent_model": false}, "family": "fast-41-diverse-075-009", "id": "fast-41-diverse-075-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "The testing portal report records the student's EDU 340 quantitative placement-test score as 80 points."}, {"speaker": "Records clerk", "text": "The testing-center export dated 2026-09-01 records the student's EDU 340 quantitative placement-test score as 79 points."}, {"speaker": "Registrar", "text": "The complete exception file contains no quantitative placement-test score report beyond those two records. The comparison sheet lists 80 as the maximum of the reported scores, and the portal entry carries verified status."}, {"speaker": "Registrar", "text": "The instructor note submitted with the EDU 340 exception request is signed and recommends admission."}, {"speaker": "Policy file", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Policy file", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both contexts, including documentation requirements, routing rules, and readiness categories. The student, courses, exception scope, placement report, and relevant time binding are preserved, while permitted case observations change. The evidence spans are complete factual sentences: \"On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17.\" and \"On 14 May 2026, document PR-17 contains no numerical placement-score value.\" The counterfactual's score of 82 is consistent with the report identifier, the threshold, and the absence of below-threshold evidence. Neither context embeds an answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet includes her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in an algebra course equivalent to MAT 120, and the coordinator's note states that Maya appears prepared. The packet contains no documented evidence that her placement result is below 75. On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17. On 14 May 2026, document PR-17 contains no numerical placement-score value. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17."}, {"path": [], "text": "On 14 May 2026, document PR-17 contains no numerical placement-score value."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17.", "negative_left": "On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17.", "negative_right": "On 14 May 2026, document PR-17 contains a numerical placement-score value of 82.", "right": "On 14 May 2026, document PR-17 contains no numerical placement-score value."}, "verifier_independent_model": false}, "family": "fast-41-diverse-076-002", "id": "fast-41-diverse-076-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet includes her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in an algebra course equivalent to MAT 120, and the coordinator's note states that Maya appears prepared. The packet contains no documented evidence that her placement result is below 75. On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17. On 14 May 2026, document PR-17 contains no numerical placement-score value. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both contexts, including documentation requirements, routing rules, and readiness categories. The student, courses, exception scope, placement report, and relevant time binding are preserved, while permitted case observations change. The evidence spans are complete factual sentences: \"On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17.\" and \"On 14 May 2026, document PR-17 contains no numerical placement-score value.\" The counterfactual's score of 82 is consistent with the report identifier, the threshold, and the absence of below-threshold evidence. Neither context embeds an answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet includes her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in an algebra course equivalent to MAT 120, and the coordinator's note states that Maya appears prepared. The packet contains no documented evidence that her placement result is below 75. On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17. On 14 May 2026, document PR-17 contains no numerical placement-score value. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17."}, {"path": [], "text": "On 14 May 2026, document PR-17 contains no numerical placement-score value."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17.", "negative_left": "On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17.", "negative_right": "On 14 May 2026, document PR-17 contains a numerical placement-score value of 82.", "right": "On 14 May 2026, document PR-17 contains no numerical placement-score value."}, "verifier_independent_model": false}, "family": "fast-41-diverse-076-002", "id": "fast-41-diverse-076-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet includes her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in an algebra course equivalent to MAT 120, and the coordinator's note states that Maya appears prepared. The packet contains no documented evidence that her placement result is below 75. On 14 May 2026, Maya Chen's official placement report was assigned document identifier PR-17. On 14 May 2026, document PR-17 contains a numerical placement-score value of 82. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question bindings and all governing policy. The two evidence spans are complete factual sentences, including the exact quotes “pending” and “82”. The counterfactual coherently changes R-17 from text “pending” to numeric “82” without duplicate or contradictory measurements. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet contains her transcript, the academic adviser's request form, and a course-coordinator note stating that she appears prepared. The transcript records a B+ in an algebra course equivalent to MAT 120. The packet contains no documented evidence that her placement result is below 75. Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17. Record R-17 displays the text “pending” and is typed as text. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17."}, {"path": [], "text": "Record R-17 displays the text “pending” and is typed as text."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17.", "negative_left": "Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17.", "negative_right": "Record R-17 displays the number “82” and is typed as numeric.", "right": "Record R-17 displays the text “pending” and is typed as text."}, "verifier_independent_model": false}, "family": "fast-41-diverse-076-005", "id": "fast-41-diverse-076-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet contains her transcript, the academic adviser's request form, and a course-coordinator note stating that she appears prepared. The transcript records a B+ in an algebra course equivalent to MAT 120. The packet contains no documented evidence that her placement result is below 75. Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17. Record R-17 displays the text “pending” and is typed as text. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question bindings and all governing policy. The two evidence spans are complete factual sentences, including the exact quotes “pending” and “82”. The counterfactual coherently changes R-17 from text “pending” to numeric “82” without duplicate or contradictory measurements. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet contains her transcript, the academic adviser's request form, and a course-coordinator note stating that she appears prepared. The transcript records a B+ in an algebra course equivalent to MAT 120. The packet contains no documented evidence that her placement result is below 75. Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17. Record R-17 displays the text “pending” and is typed as text. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17."}, {"path": [], "text": "Record R-17 displays the text “pending” and is typed as text."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17.", "negative_left": "Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17.", "negative_right": "Record R-17 displays the number “82” and is typed as numeric.", "right": "Record R-17 displays the text “pending” and is typed as text."}, "verifier_independent_model": false}, "family": "fast-41-diverse-076-005", "id": "fast-41-diverse-076-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet contains her transcript, the academic adviser's request form, and a course-coordinator note stating that she appears prepared. The transcript records a B+ in an algebra course equivalent to MAT 120. The packet contains no documented evidence that her placement result is below 75. Maya Chen's official placement report filed on 14 March 2026 identifies its placement-score cell as record R-17. Record R-17 displays the number “82” and is typed as numeric. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings; the two evidence spans are factual sentences, and the counterfactual is consistent because its unspecified numerical score need not be below 75. Neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet includes her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in an algebra course equivalent to MAT 120, and the note states that Maya appears prepared. Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026. Document PR-418 contains no numerical placement-score value. The packet contains no documented evidence that her placement result is below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026."}, {"path": [], "text": "Document PR-418 contains no numerical placement-score value."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026.", "negative_left": "Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026.", "negative_right": "Document PR-418 contains a numerical placement-score value.", "right": "Document PR-418 contains no numerical placement-score value."}, "verifier_independent_model": false}, "family": "fast-41-diverse-076-010", "id": "fast-41-diverse-076-010-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet includes her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in an algebra course equivalent to MAT 120, and the note states that Maya appears prepared. Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026. Document PR-418 contains no numerical placement-score value. The packet contains no documented evidence that her placement result is below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings; the two evidence spans are factual sentences, and the counterfactual is consistent because its unspecified numerical score need not be below 75. Neither context embeds an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the negative evidence claims in a8 and a9. The focus atom a8 is factual rather than policy-based. The base and counter assignments are jointly realizable while changing only a8: the counter can contain a numerical score of at least 75 while still containing no evidence of a below-75 result. Policy evidence correctly preserves the substantive rules originating in the original state; rules already contained in the unchanged questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all non-placement documentation, supportive evidence from the equivalent-course B+ and coordinator note, absence of the required numerical placement score, and no documented below-threshold result. Under the preserved policy, this is sufficient to withhold approval, classify Maya as Provisionally Ready, and route the incomplete packet to the academic adviser.", "rule_index": 0, "sound": true}, {"reason": "Refuting a8 establishes that the official placement report does contain a numerical score. Together with a9, that score cannot be documented as below 75. More importantly, the requested composite disposition cannot apply because there is no missing numerical score to obtain, so the packet need not be routed to the adviser for that purpose and the answer is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's request seeks an exception to the MAT 120 prerequisite for MAT 210."}, {"id": "a2", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains her transcript."}, {"id": "a3", "statement": "Maya Chen's transcript records a B+ grade in the cited algebra course."}, {"id": "a4", "statement": "The algebra course cited on Maya Chen's transcript is equivalent to MAT 120."}, {"id": "a5", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains the academic adviser's request form."}, {"id": "a6", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains a course-coordinator note."}, {"id": "a7", "statement": "The course-coordinator note in Maya Chen's submitted MAT 210 prerequisite-exception packet states that Maya Chen appears prepared."}, {"id": "a8", "statement": "Maya Chen's official placement report contains no numerical placement-score value."}, {"id": "a9", "statement": "Maya Chen's submitted MAT 210 prerequisite-exception packet contains no documented evidence that her placement result is below 75."}], "base_state_json": "\"Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet includes her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in an algebra course equivalent to MAT 120, and the note states that Maya appears prepared. Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026. Document PR-418 contains no numerical placement-score value. The packet contains no documented evidence that her placement result is below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds).\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026."}, {"path": [], "text": "Document PR-418 contains no numerical placement-score value."}], "policy_evidence": [{"path": [], "text": "Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75."}, {"path": [], "text": "Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator."}, {"path": [], "text": "Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}], "rules": [{"justification": "The transcript, adviser form, and coordinator note are present; the B+ in an equivalent course and the coordinator's statement provide supportive evidence; the required numerical placement value is missing; and the packet contains no documented placement evidence below 75 that would require Not Ready. The standard missing documentation therefore requires withholding approval, classifying Maya as Provisionally Ready, and routing the packet to the academic adviser.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Refuting the claim that Maya's official placement report contains no numerical value establishes that it contains a numerical placement-score value. Because all required documents are present and the packet contains no documented placement result below 75, the case is not Provisionally Ready due to a missing numerical score and should not be routed to the academic adviser to obtain that score. The queried composite disposition is therefore false.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026.", "negative_left": "Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026.", "negative_right": "Document PR-418 contains a numerical placement-score value.", "right": "Document PR-418 contains no numerical placement-score value."}, "verifier_independent_model": false}, "family": "fast-41-diverse-076-010", "id": "fast-41-diverse-076-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve or deny the exception now, assign a different readiness category, or route the case to the course coordinator or registrar despite the standard missing document.", "true": "Yes — withhold approval for now, classify the student as Provisionally Ready, and send the incomplete packet to the academic adviser to obtain the numerical placement score."}, "instructions": "Is the proper disposition to withhold a yes recommendation on the prerequisite exception, rate Maya Provisionally Ready, and route the packet to the academic adviser for the missing numerical placement score?", "type": "noul"}}, "state": "Maya Chen seeks an exception to the MAT 120 prerequisite for MAT 210. Her submitted exception packet includes her transcript, the academic adviser's request form, and a course-coordinator note. The transcript records a B+ in an algebra course equivalent to MAT 120, and the note states that Maya appears prepared. Maya Chen's official placement report for her MAT 210 prerequisite-exception request is document PR-418, filed on 14 March 2026. Document PR-418 contains a numerical placement-score value. The packet contains no documented evidence that her placement result is below 75. Course rules require a transcript, adviser form, instructor or coordinator note, and an official placement score of at least 75. Missing standard documentation goes to the academic adviser; unusual complete cases go to the coordinator. Readiness is ordered as Ready (complete and meets thresholds), Provisionally Ready (supportive evidence but required data missing), or Not Ready (documented evidence below thresholds)."}, "method": "c2d", "provenance": {"source_id": "diverse-076", "source_is_synthetic": true, "source_sha256": "9f957075bd891bb05bd1b05842d493e588098fd4d07857d14235058c0a69b1b3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all four governing policy rules. Entity, course, request, term, and routing bindings are unchanged. The evidence spans are complete factual sentences: \"Mina Cho's requested CHEM 220 term starts on 1 September 2026.\" and \"Mina Cho's official placement report is dated 14 March 2026.\" The counterfactual coherently changes only the report date to 14 March 2025, without contradictory duplicate claims within that context. Neither context embeds a gold answer, label code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":{\"case_note\":\"Mina Cho requests a prerequisite exception for CHEM 220. Her official transcript records D+ in CHEM 110, below the normal minimum of C. The exception file contains the official transcript, an official placement report, and a CHEM 220 instructor note. The placement report records a score of 82. Mina Cho's requested CHEM 220 term starts on 1 September 2026. Mina Cho's official placement report is dated 14 March 2026. The instructor note describes strong laboratory skills and supports enrollment. The academic adviser forwarded the file to the course coordinator.\",\"policy\":[\"The normal prerequisite is CHEM 110 with at least C.\",\"A placement score of 75 or higher substitutes only when dated within 12 months.\",\"The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel.\",\"A complete file requires an official transcript, dated placement report, and instructor note.\"]}}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context", "case_note"], "text": "Mina Cho's requested CHEM 220 term starts on 1 September 2026."}, {"path": ["context", "case_note"], "text": "Mina Cho's official placement report is dated 14 March 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's requested CHEM 220 term starts on 1 September 2026.", "negative_left": "Mina Cho's requested CHEM 220 term starts on 1 September 2026.", "negative_right": "Mina Cho's official placement report is dated 14 March 2025.", "right": "Mina Cho's official placement report is dated 14 March 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-004", "id": "fast-41-diverse-077-004-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": {"case_note": "Mina Cho requests a prerequisite exception for CHEM 220. Her official transcript records D+ in CHEM 110, below the normal minimum of C. The exception file contains the official transcript, an official placement report, and a CHEM 220 instructor note. The placement report records a score of 82. Mina Cho's requested CHEM 220 term starts on 1 September 2026. Mina Cho's official placement report is dated 14 March 2026. The instructor note describes strong laboratory skills and supports enrollment. The academic adviser forwarded the file to the course coordinator.", "policy": ["The normal prerequisite is CHEM 110 with at least C.", "A placement score of 75 or higher substitutes only when dated within 12 months.", "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel.", "A complete file requires an official transcript, dated placement report, and instructor note."]}}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all four governing policy rules. Entity, course, request, term, and routing bindings are unchanged. The evidence spans are complete factual sentences: \"Mina Cho's requested CHEM 220 term starts on 1 September 2026.\" and \"Mina Cho's official placement report is dated 14 March 2026.\" The counterfactual coherently changes only the report date to 14 March 2025, without contradictory duplicate claims within that context. Neither context embeds a gold answer, label code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":{\"case_note\":\"Mina Cho requests a prerequisite exception for CHEM 220. Her official transcript records D+ in CHEM 110, below the normal minimum of C. The exception file contains the official transcript, an official placement report, and a CHEM 220 instructor note. The placement report records a score of 82. Mina Cho's requested CHEM 220 term starts on 1 September 2026. Mina Cho's official placement report is dated 14 March 2026. The instructor note describes strong laboratory skills and supports enrollment. The academic adviser forwarded the file to the course coordinator.\",\"policy\":[\"The normal prerequisite is CHEM 110 with at least C.\",\"A placement score of 75 or higher substitutes only when dated within 12 months.\",\"The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel.\",\"A complete file requires an official transcript, dated placement report, and instructor note.\"]}}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context", "case_note"], "text": "Mina Cho's requested CHEM 220 term starts on 1 September 2026."}, {"path": ["context", "case_note"], "text": "Mina Cho's official placement report is dated 14 March 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's requested CHEM 220 term starts on 1 September 2026.", "negative_left": "Mina Cho's requested CHEM 220 term starts on 1 September 2026.", "negative_right": "Mina Cho's official placement report is dated 14 March 2025.", "right": "Mina Cho's official placement report is dated 14 March 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-004", "id": "fast-41-diverse-077-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": {"case_note": "Mina Cho requests a prerequisite exception for CHEM 220. Her official transcript records D+ in CHEM 110, below the normal minimum of C. The exception file contains the official transcript, an official placement report, and a CHEM 220 instructor note. The placement report records a score of 82. Mina Cho's requested CHEM 220 term starts on 1 September 2026. Mina Cho's official placement report is dated 14 March 2025. The instructor note describes strong laboratory skills and supports enrollment. The academic adviser forwarded the file to the course coordinator.", "policy": ["The normal prerequisite is CHEM 110 with at least C.", "A placement score of 75 or higher substitutes only when dated within 12 months.", "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel.", "A complete file requires an official transcript, dated placement report, and instructor note."]}}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing prerequisite, score, timing, scope, routing, and documentation policies alongside the unchanged questions. Mina Cho, CHEM 220, the exception request, and the relevant date bindings remain consistent. The two evidence spans are complete factual sentences: \"The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026.\" and \"Mina Cho's requested CHEM 220 term starts on December 1, 2026.\" The counterfactual changes only the requested term date and remains coherent with the placement-report date. Neither context embeds a gold answer, answer code, rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Case note: Mina Cho requests a CHEM 220 prerequisite exception. The official transcript in her exception file records D+ for CHEM 110, below the normal minimum. The file also contains an official placement report and a CHEM 220 instructor note; the report records a score of 82. The academic adviser forwarded all three documents to the course coordinator, and the report's date is stated in the file. The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026. Mina Cho's requested CHEM 220 term starts on December 1, 2026. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The instructor note describes strong laboratory skills and supports enrollment. The coordinator has received the complete packet for review.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026."}, {"path": ["context"], "text": "Mina Cho's requested CHEM 220 term starts on December 1, 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026.", "negative_left": "The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026.", "negative_right": "Mina Cho's requested CHEM 220 term starts on January 20, 2027.", "right": "Mina Cho's requested CHEM 220 term starts on December 1, 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-007", "id": "fast-41-diverse-077-007-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Case note: Mina Cho requests a CHEM 220 prerequisite exception. The official transcript in her exception file records D+ for CHEM 110, below the normal minimum. The file also contains an official placement report and a CHEM 220 instructor note; the report records a score of 82. The academic adviser forwarded all three documents to the course coordinator, and the report's date is stated in the file. The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026. Mina Cho's requested CHEM 220 term starts on December 1, 2026. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The instructor note describes strong laboratory skills and supports enrollment. The coordinator has received the complete packet for review."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing prerequisite, score, timing, scope, routing, and documentation policies alongside the unchanged questions. Mina Cho, CHEM 220, the exception request, and the relevant date bindings remain consistent. The two evidence spans are complete factual sentences: \"The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026.\" and \"Mina Cho's requested CHEM 220 term starts on December 1, 2026.\" The counterfactual changes only the requested term date and remains coherent with the placement-report date. Neither context embeds a gold answer, answer code, rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Case note: Mina Cho requests a CHEM 220 prerequisite exception. The official transcript in her exception file records D+ for CHEM 110, below the normal minimum. The file also contains an official placement report and a CHEM 220 instructor note; the report records a score of 82. The academic adviser forwarded all three documents to the course coordinator, and the report's date is stated in the file. The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026. Mina Cho's requested CHEM 220 term starts on December 1, 2026. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The instructor note describes strong laboratory skills and supports enrollment. The coordinator has received the complete packet for review.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026."}, {"path": ["context"], "text": "Mina Cho's requested CHEM 220 term starts on December 1, 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026.", "negative_left": "The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026.", "negative_right": "Mina Cho's requested CHEM 220 term starts on January 20, 2027.", "right": "Mina Cho's requested CHEM 220 term starts on December 1, 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-007", "id": "fast-41-diverse-077-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Case note: Mina Cho requests a CHEM 220 prerequisite exception. The official transcript in her exception file records D+ for CHEM 110, below the normal minimum. The file also contains an official placement report and a CHEM 220 instructor note; the report records a score of 82. The academic adviser forwarded all three documents to the course coordinator, and the report's date is stated in the file. The official placement report in Mina Cho's CHEM 220 exception file is dated January 15, 2026. Mina Cho's requested CHEM 220 term starts on January 20, 2027. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The instructor note describes strong laboratory skills and supports enrollment. The coordinator has received the complete packet for review."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, the evidence is factual and complete, the changed term date is coherent, and neither context leaks an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Mina Cho requests a CHEM 220 prerequisite exception. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The file contains all three official documents, and the instructor note supports enrollment. The transcript records a D+ in CHEM 110, while the placement report records a score of 82. The adviser forwarded the complete file to the coordinator. No additional authenticity concern is recorded.\",\"evidence\":[\"Mina Cho's official placement report is dated 14 February 2026.\",\"Mina Cho's requested CHEM 220 term starts on 1 February 2027.\",\"Mina Cho's CHEM 220 exception file contains an official transcript.\",\"Mina Cho's CHEM 220 exception file contains an official placement report.\",\"Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note.\",\"Mina Cho's grade in CHEM 110 is D+.\",\"The score stated in Mina Cho's official placement report is 82.\"]}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "0"], "text": "Mina Cho's official placement report is dated 14 February 2026."}, {"path": ["evidence", "1"], "text": "Mina Cho's requested CHEM 220 term starts on 1 February 2027."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated 14 February 2026.", "negative_left": "Mina Cho's official placement report is dated 14 February 2026.", "negative_right": "Mina Cho's requested CHEM 220 term starts on 1 February 2028.", "right": "Mina Cho's requested CHEM 220 term starts on 1 February 2027."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-008", "id": "fast-41-diverse-077-008-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Mina Cho requests a CHEM 220 prerequisite exception. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The file contains all three official documents, and the instructor note supports enrollment. The transcript records a D+ in CHEM 110, while the placement report records a score of 82. The adviser forwarded the complete file to the coordinator. No additional authenticity concern is recorded.", "evidence": ["Mina Cho's official placement report is dated 14 February 2026.", "Mina Cho's requested CHEM 220 term starts on 1 February 2027.", "Mina Cho's CHEM 220 exception file contains an official transcript.", "Mina Cho's CHEM 220 exception file contains an official placement report.", "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note.", "Mina Cho's grade in CHEM 110 is D+.", "The score stated in Mina Cho's official placement report is 82."]}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, the evidence is factual and complete, the changed term date is coherent, and neither context leaks an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Mina Cho requests a CHEM 220 prerequisite exception. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The file contains all three official documents, and the instructor note supports enrollment. The transcript records a D+ in CHEM 110, while the placement report records a score of 82. The adviser forwarded the complete file to the coordinator. No additional authenticity concern is recorded.\",\"evidence\":[\"Mina Cho's official placement report is dated 14 February 2026.\",\"Mina Cho's requested CHEM 220 term starts on 1 February 2027.\",\"Mina Cho's CHEM 220 exception file contains an official transcript.\",\"Mina Cho's CHEM 220 exception file contains an official placement report.\",\"Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note.\",\"Mina Cho's grade in CHEM 110 is D+.\",\"The score stated in Mina Cho's official placement report is 82.\"]}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "0"], "text": "Mina Cho's official placement report is dated 14 February 2026."}, {"path": ["evidence", "1"], "text": "Mina Cho's requested CHEM 220 term starts on 1 February 2027."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated 14 February 2026.", "negative_left": "Mina Cho's official placement report is dated 14 February 2026.", "negative_right": "Mina Cho's requested CHEM 220 term starts on 1 February 2028.", "right": "Mina Cho's requested CHEM 220 term starts on 1 February 2027."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-008", "id": "fast-41-diverse-077-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Mina Cho requests a CHEM 220 prerequisite exception. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The file contains all three official documents, and the instructor note supports enrollment. The transcript records a D+ in CHEM 110, while the placement report records a score of 82. The adviser forwarded the complete file to the coordinator. No additional authenticity concern is recorded.", "evidence": ["Mina Cho's official placement report is dated 14 February 2026.", "Mina Cho's requested CHEM 220 term starts on 1 February 2028.", "Mina Cho's CHEM 220 exception file contains an official transcript.", "Mina Cho's CHEM 220 exception file contains an official placement report.", "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note.", "Mina Cho's grade in CHEM 110 is D+.", "The score stated in Mina Cho's official placement report is 82."]}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and question bindings, and the two evidence spans are complete factual sentences. The counterfactual date change is coherent with the unchanged report date and introduces no contradictory duplicate claims. Neither context states a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Mina Cho’s CHEM 220 exception file contains an official transcript, an official placement report, and a CHEM 220 instructor note. The transcript records a D+ in CHEM 110, rather than the required passing grade. The placement report records a score of 82. The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026. Mina Cho's requested CHEM 220 term starts November 20, 2026. The instructor note supports enrollment based on Mina’s laboratory skills, and the adviser forwarded all three documents to the coordinator. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026."}, {"path": ["context"], "text": "Mina Cho's requested CHEM 220 term starts November 20, 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026.", "negative_left": "The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026.", "negative_right": "Mina Cho's requested CHEM 220 term starts November 20, 2027.", "right": "Mina Cho's requested CHEM 220 term starts November 20, 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-009", "id": "fast-41-diverse-077-009-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Mina Cho’s CHEM 220 exception file contains an official transcript, an official placement report, and a CHEM 220 instructor note. The transcript records a D+ in CHEM 110, rather than the required passing grade. The placement report records a score of 82. The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026. Mina Cho's requested CHEM 220 term starts November 20, 2026. The instructor note supports enrollment based on Mina’s laboratory skills, and the adviser forwarded all three documents to the coordinator. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and question bindings, and the two evidence spans are complete factual sentences. The counterfactual date change is coherent with the unchanged report date and introduces no contradictory duplicate claims. Neither context states a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Mina Cho’s CHEM 220 exception file contains an official transcript, an official placement report, and a CHEM 220 instructor note. The transcript records a D+ in CHEM 110, rather than the required passing grade. The placement report records a score of 82. The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026. Mina Cho's requested CHEM 220 term starts November 20, 2026. The instructor note supports enrollment based on Mina’s laboratory skills, and the adviser forwarded all three documents to the coordinator. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026."}, {"path": ["context"], "text": "Mina Cho's requested CHEM 220 term starts November 20, 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026.", "negative_left": "The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026.", "negative_right": "Mina Cho's requested CHEM 220 term starts November 20, 2027.", "right": "Mina Cho's requested CHEM 220 term starts November 20, 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-009", "id": "fast-41-diverse-077-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Mina Cho’s CHEM 220 exception file contains an official transcript, an official placement report, and a CHEM 220 instructor note. The transcript records a D+ in CHEM 110, rather than the required passing grade. The placement report records a score of 82. The official placement report in Mina Cho's CHEM 220 exception file is dated March 3, 2026. Mina Cho's requested CHEM 220 term starts November 20, 2027. The instructor note supports enrollment based on Mina’s laboratory skills, and the adviser forwarded all three documents to the coordinator. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and bindings remain preserved through the unchanged original questions and matching Mina Cho, CHEM 220, coordinator, and panel references. The evidence consists of exactly two complete factual sentences: “The date stated on Mina Cho's official placement report is 1 September 2025.” and “The start of Mina Cho's requested CHEM 220 term is 31 August 2026.” The counterfactual changes only the term-start date and remains internally coherent. Neither context embeds a gold answer, label rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The file contains Mina's official transcript, official placement report, and CHEM 220 instructor note. The transcript records a D+ in CHEM 110. The placement report records a score of 82 and states its date. The instructor note describes strong laboratory skills and supports enrollment. The adviser forwarded the complete file to the coordinator. The date stated on Mina Cho's official placement report is 1 September 2025. The start of Mina Cho's requested CHEM 220 term is 31 August 2026. A registrar staff member confirms that older-score cases require panel review.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "The date stated on Mina Cho's official placement report is 1 September 2025."}, {"path": ["context"], "text": "The start of Mina Cho's requested CHEM 220 term is 31 August 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "The date stated on Mina Cho's official placement report is 1 September 2025.", "negative_left": "The date stated on Mina Cho's official placement report is 1 September 2025.", "negative_right": "The start of Mina Cho's requested CHEM 220 term is 2 September 2026.", "right": "The start of Mina Cho's requested CHEM 220 term is 31 August 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-011", "id": "fast-41-diverse-077-011-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The file contains Mina's official transcript, official placement report, and CHEM 220 instructor note. The transcript records a D+ in CHEM 110. The placement report records a score of 82 and states its date. The instructor note describes strong laboratory skills and supports enrollment. The adviser forwarded the complete file to the coordinator. The date stated on Mina Cho's official placement report is 1 September 2025. The start of Mina Cho's requested CHEM 220 term is 31 August 2026. A registrar staff member confirms that older-score cases require panel review."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and bindings remain preserved through the unchanged original questions and matching Mina Cho, CHEM 220, coordinator, and panel references. The evidence consists of exactly two complete factual sentences: “The date stated on Mina Cho's official placement report is 1 September 2025.” and “The start of Mina Cho's requested CHEM 220 term is 31 August 2026.” The counterfactual changes only the term-start date and remains internally coherent. Neither context embeds a gold answer, label rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The file contains Mina's official transcript, official placement report, and CHEM 220 instructor note. The transcript records a D+ in CHEM 110. The placement report records a score of 82 and states its date. The instructor note describes strong laboratory skills and supports enrollment. The adviser forwarded the complete file to the coordinator. The date stated on Mina Cho's official placement report is 1 September 2025. The start of Mina Cho's requested CHEM 220 term is 31 August 2026. A registrar staff member confirms that older-score cases require panel review.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "The date stated on Mina Cho's official placement report is 1 September 2025."}, {"path": ["context"], "text": "The start of Mina Cho's requested CHEM 220 term is 31 August 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "The date stated on Mina Cho's official placement report is 1 September 2025.", "negative_left": "The date stated on Mina Cho's official placement report is 1 September 2025.", "negative_right": "The start of Mina Cho's requested CHEM 220 term is 2 September 2026.", "right": "The start of Mina Cho's requested CHEM 220 term is 31 August 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-011", "id": "fast-41-diverse-077-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. The file contains Mina's official transcript, official placement report, and CHEM 220 instructor note. The transcript records a D+ in CHEM 110. The placement report records a score of 82 and states its date. The instructor note describes strong laboratory skills and supports enrollment. The adviser forwarded the complete file to the coordinator. The date stated on Mina Cho's official placement report is 1 September 2025. The start of Mina Cho's requested CHEM 220 term is 2 September 2026. A registrar staff member confirms that older-score cases require panel review."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing prerequisite, timing, documentation, authority, and routing policies through the unchanged questions and context text. The request remains Mina Cho’s CHEM 220 prerequisite-exception request with unchanged entity and scope bindings. The base evidence retains exactly \"Mina Cho's official placement report states the date 2025-08-20.\" and \"Mina Cho's requested CHEM 220 term starts on 2026-05-01.\" The counterfactual evidence retains exactly \"Mina Cho's official placement report states the date 2025-08-20.\" and \"Mina Cho's requested CHEM 220 term starts on 2026-10-01.\" The counterfactual changes only the term date and remains internally coherent with the unchanged placement-report date and policy. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Case note: Mina Cho requests a prerequisite exception for CHEM 220. Her official transcript records D+ in CHEM 110, while the placement report records a score of 82. The CHEM 220 instructor note describes strong laboratory skills and supports enrollment. An adviser forwarded the transcript, placement report, and instructor note to the coordinator. Mina Cho's official placement report states the date 2025-08-20. Mina Cho's requested CHEM 220 term starts on 2026-05-01. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. Registrar staff will handle any case outside the coordinator's authority.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "Mina Cho's official placement report states the date 2025-08-20."}, {"path": ["context"], "text": "Mina Cho's requested CHEM 220 term starts on 2026-05-01."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report states the date 2025-08-20.", "negative_left": "Mina Cho's official placement report states the date 2025-08-20.", "negative_right": "Mina Cho's requested CHEM 220 term starts on 2026-10-01.", "right": "Mina Cho's requested CHEM 220 term starts on 2026-05-01."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-012", "id": "fast-41-diverse-077-012-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Case note: Mina Cho requests a prerequisite exception for CHEM 220. Her official transcript records D+ in CHEM 110, while the placement report records a score of 82. The CHEM 220 instructor note describes strong laboratory skills and supports enrollment. An adviser forwarded the transcript, placement report, and instructor note to the coordinator. Mina Cho's official placement report states the date 2025-08-20. Mina Cho's requested CHEM 220 term starts on 2026-05-01. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. Registrar staff will handle any case outside the coordinator's authority."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing prerequisite, timing, documentation, authority, and routing policies through the unchanged questions and context text. The request remains Mina Cho’s CHEM 220 prerequisite-exception request with unchanged entity and scope bindings. The base evidence retains exactly \"Mina Cho's official placement report states the date 2025-08-20.\" and \"Mina Cho's requested CHEM 220 term starts on 2026-05-01.\" The counterfactual evidence retains exactly \"Mina Cho's official placement report states the date 2025-08-20.\" and \"Mina Cho's requested CHEM 220 term starts on 2026-10-01.\" The counterfactual changes only the term date and remains internally coherent with the unchanged placement-report date and policy. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Case note: Mina Cho requests a prerequisite exception for CHEM 220. Her official transcript records D+ in CHEM 110, while the placement report records a score of 82. The CHEM 220 instructor note describes strong laboratory skills and supports enrollment. An adviser forwarded the transcript, placement report, and instructor note to the coordinator. Mina Cho's official placement report states the date 2025-08-20. Mina Cho's requested CHEM 220 term starts on 2026-05-01. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. Registrar staff will handle any case outside the coordinator's authority.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "Mina Cho's official placement report states the date 2025-08-20."}, {"path": ["context"], "text": "Mina Cho's requested CHEM 220 term starts on 2026-05-01."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report states the date 2025-08-20.", "negative_left": "Mina Cho's official placement report states the date 2025-08-20.", "negative_right": "Mina Cho's requested CHEM 220 term starts on 2026-10-01.", "right": "Mina Cho's requested CHEM 220 term starts on 2026-05-01."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-012", "id": "fast-41-diverse-077-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Case note: Mina Cho requests a prerequisite exception for CHEM 220. Her official transcript records D+ in CHEM 110, while the placement report records a score of 82. The CHEM 220 instructor note describes strong laboratory skills and supports enrollment. An adviser forwarded the transcript, placement report, and instructor note to the coordinator. Mina Cho's official placement report states the date 2025-08-20. Mina Cho's requested CHEM 220 term starts on 2026-10-01. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. Registrar staff will handle any case outside the coordinator's authority."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question policy, Mina/CHEM 220 bindings, and request scope. The exact evidence quotes are \"Mina Cho's official placement report is dated 10 January 2026.\" and \"Mina Cho's requested CHEM 220 term starts on 1 December 2026.\" The counterfactual changes only the term date to 1 February 2027, making the score timing expired without creating contradictory duplicate assertions or embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. Mina Cho's official placement report is dated 10 January 2026. Mina Cho's requested CHEM 220 term starts on 1 December 2026. The report records a placement score of 82. Her official transcript records a D+ in CHEM 110. The file contains the transcript, placement report, and a CHEM 220 instructor note. The instructor note describes strong laboratory skills and supports enrollment. An adviser forwarded the assembled file to the course coordinator, and registrar staff noted that expired placement-score cases require panel review.\",\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "Mina Cho's official placement report is dated 10 January 2026."}, {"path": ["context"], "text": "Mina Cho's requested CHEM 220 term starts on 1 December 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated 10 January 2026.", "negative_left": "Mina Cho's official placement report is dated 10 January 2026.", "negative_right": "Mina Cho's requested CHEM 220 term starts on 1 February 2027.", "right": "Mina Cho's requested CHEM 220 term starts on 1 December 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-015", "id": "fast-41-diverse-077-015-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. Mina Cho's official placement report is dated 10 January 2026. Mina Cho's requested CHEM 220 term starts on 1 December 2026. The report records a placement score of 82. Her official transcript records a D+ in CHEM 110. The file contains the transcript, placement report, and a CHEM 220 instructor note. The instructor note describes strong laboratory skills and supports enrollment. An adviser forwarded the assembled file to the course coordinator, and registrar staff noted that expired placement-score cases require panel review.", "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question policy, Mina/CHEM 220 bindings, and request scope. The exact evidence quotes are \"Mina Cho's official placement report is dated 10 January 2026.\" and \"Mina Cho's requested CHEM 220 term starts on 1 December 2026.\" The counterfactual changes only the term date to 1 February 2027, making the score timing expired without creating contradictory duplicate assertions or embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. Mina Cho's official placement report is dated 10 January 2026. Mina Cho's requested CHEM 220 term starts on 1 December 2026. The report records a placement score of 82. Her official transcript records a D+ in CHEM 110. The file contains the transcript, placement report, and a CHEM 220 instructor note. The instructor note describes strong laboratory skills and supports enrollment. An adviser forwarded the assembled file to the course coordinator, and registrar staff noted that expired placement-score cases require panel review.\",\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["context"], "text": "Mina Cho's official placement report is dated 10 January 2026."}, {"path": ["context"], "text": "Mina Cho's requested CHEM 220 term starts on 1 December 2026."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated 10 January 2026.", "negative_left": "Mina Cho's official placement report is dated 10 January 2026.", "negative_right": "Mina Cho's requested CHEM 220 term starts on 1 February 2027.", "right": "Mina Cho's requested CHEM 220 term starts on 1 December 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-077-015", "id": "fast-41-diverse-077-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note. Mina Cho's official placement report is dated 10 January 2026. Mina Cho's requested CHEM 220 term starts on 1 February 2027. The report records a placement score of 82. Her official transcript records a D+ in CHEM 110. The file contains the transcript, placement report, and a CHEM 220 instructor note. The instructor note describes strong laboratory skills and supports enrollment. An adviser forwarded the assembled file to the course coordinator, and registrar staff noted that expired placement-score cases require panel review.", "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy wording and the unchanged questions are preserved in both contexts. Maya Chen, Calculus I, packet MC-4821, and the relevant evidence bindings remain unchanged. The evidence spans are complete factual sentences: \"On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739.\" and \"On 14 August 2026, the signed note bearing reference I-739 had the disposition \\\"supports enrollment in Calculus I\\\" for the student identified in packet MC-4821.\" The counterfactual changes only the note disposition and does not create duplicate measurements or assertions within that context. Neither context states a gold answer, label rationale, proposition identifier, or output instruction. The contexts contain natural policy prose rather than an answer code or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar staff member\",\"text\":\"For Maya Chen's Calculus I exception request, packet MC-4821 contains her transcript and her Calculus I placement result.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The assessment ledger records Maya's placement score as 79.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"On 14 August 2026, the signed note bearing reference I-739 had the disposition \\\"supports enrollment in Calculus I\\\" for the student identified in packet MC-4821.\"},{\"speaker\":\"Course coordinator\",\"text\":\"The complete packet is being reviewed under the following policy: The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739."}, {"path": ["3", "text"], "text": "On 14 August 2026, the signed note bearing reference I-739 had the disposition \"supports enrollment in Calculus I\" for the student identified in packet MC-4821."}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739.", "negative_left": "On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739.", "negative_right": "On 14 August 2026, the signed note bearing reference I-739 had the disposition \"does not support enrollment in Calculus I\" for the student identified in packet MC-4821.", "right": "On 14 August 2026, the signed note bearing reference I-739 had the disposition \"supports enrollment in Calculus I\" for the student identified in packet MC-4821."}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-002", "id": "fast-41-diverse-078-002-base", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar staff member", "text": "For Maya Chen's Calculus I exception request, packet MC-4821 contains her transcript and her Calculus I placement result."}, {"speaker": "Academic adviser", "text": "The assessment ledger records Maya's placement score as 79."}, {"speaker": "Registrar staff member", "text": "On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739."}, {"speaker": "Registrar staff member", "text": "On 14 August 2026, the signed note bearing reference I-739 had the disposition \"supports enrollment in Calculus I\" for the student identified in packet MC-4821."}, {"speaker": "Course coordinator", "text": "The complete packet is being reviewed under the following policy: The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy wording and the unchanged questions are preserved in both contexts. Maya Chen, Calculus I, packet MC-4821, and the relevant evidence bindings remain unchanged. The evidence spans are complete factual sentences: \"On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739.\" and \"On 14 August 2026, the signed note bearing reference I-739 had the disposition \\\"supports enrollment in Calculus I\\\" for the student identified in packet MC-4821.\" The counterfactual changes only the note disposition and does not create duplicate measurements or assertions within that context. Neither context states a gold answer, label rationale, proposition identifier, or output instruction. The contexts contain natural policy prose rather than an answer code or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar staff member\",\"text\":\"For Maya Chen's Calculus I exception request, packet MC-4821 contains her transcript and her Calculus I placement result.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The assessment ledger records Maya's placement score as 79.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"On 14 August 2026, the signed note bearing reference I-739 had the disposition \\\"supports enrollment in Calculus I\\\" for the student identified in packet MC-4821.\"},{\"speaker\":\"Course coordinator\",\"text\":\"The complete packet is being reviewed under the following policy: The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739."}, {"path": ["3", "text"], "text": "On 14 August 2026, the signed note bearing reference I-739 had the disposition \"supports enrollment in Calculus I\" for the student identified in packet MC-4821."}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739.", "negative_left": "On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739.", "negative_right": "On 14 August 2026, the signed note bearing reference I-739 had the disposition \"does not support enrollment in Calculus I\" for the student identified in packet MC-4821.", "right": "On 14 August 2026, the signed note bearing reference I-739 had the disposition \"supports enrollment in Calculus I\" for the student identified in packet MC-4821."}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-002", "id": "fast-41-diverse-078-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar staff member", "text": "For Maya Chen's Calculus I exception request, packet MC-4821 contains her transcript and her Calculus I placement result."}, {"speaker": "Academic adviser", "text": "The assessment ledger records Maya's placement score as 79."}, {"speaker": "Registrar staff member", "text": "On 14 August 2026, Maya Chen's Calculus I readiness packet MC-4821 contained Professor Ibarra's signed instructor note bearing reference I-739."}, {"speaker": "Registrar staff member", "text": "On 14 August 2026, the signed note bearing reference I-739 had the disposition \"does not support enrollment in Calculus I\" for the student identified in packet MC-4821."}, {"speaker": "Course coordinator", "text": "The complete packet is being reviewed under the following policy: The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question and all governing policy, and they stay bound to Maya Chen’s Calculus I packet. The two focus spans are complete factual sentences. The counterfactual changes only the instructor’s enrollment stance, without creating contradictory duplicate facts. Neither context contains a gold answer, label code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar staff member\",\"text\":\"Maya Chen's Calculus I readiness packet contains her transcript and placement result. The recorded Calculus I placement score is 79, and the packet has been checked against Maya's student record.\"},{\"speaker\":\"File note\",\"text\":\"Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.\"},{\"speaker\":\"File note\",\"text\":\"The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should be enrolled in Calculus I.\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The packet is being reviewed as a complete Calculus I exception submission, and the recorded score remains 79.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra."}, {"path": ["2", "text"], "text": "The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should be enrolled in Calculus I."}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.", "negative_left": "Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.", "negative_right": "The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should not be enrolled in Calculus I.", "right": "The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should be enrolled in Calculus I."}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-006", "id": "fast-41-diverse-078-006-base", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar staff member", "text": "Maya Chen's Calculus I readiness packet contains her transcript and placement result. The recorded Calculus I placement score is 79, and the packet has been checked against Maya's student record."}, {"speaker": "File note", "text": "Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra."}, {"speaker": "File note", "text": "The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should be enrolled in Calculus I."}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}, {"speaker": "Academic adviser", "text": "The packet is being reviewed as a complete Calculus I exception submission, and the recorded score remains 79."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question and all governing policy, and they stay bound to Maya Chen’s Calculus I packet. The two focus spans are complete factual sentences. The counterfactual changes only the instructor’s enrollment stance, without creating contradictory duplicate facts. Neither context contains a gold answer, label code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar staff member\",\"text\":\"Maya Chen's Calculus I readiness packet contains her transcript and placement result. The recorded Calculus I placement score is 79, and the packet has been checked against Maya's student record.\"},{\"speaker\":\"File note\",\"text\":\"Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.\"},{\"speaker\":\"File note\",\"text\":\"The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should be enrolled in Calculus I.\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The packet is being reviewed as a complete Calculus I exception submission, and the recorded score remains 79.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra."}, {"path": ["2", "text"], "text": "The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should be enrolled in Calculus I."}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.", "negative_left": "Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.", "negative_right": "The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should not be enrolled in Calculus I.", "right": "The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should be enrolled in Calculus I."}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-006", "id": "fast-41-diverse-078-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar staff member", "text": "Maya Chen's Calculus I readiness packet contains her transcript and placement result. The recorded Calculus I placement score is 79, and the packet has been checked against Maya's student record."}, {"speaker": "File note", "text": "Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra."}, {"speaker": "File note", "text": "The signed instructor note in Maya Chen's Calculus I readiness packet states that Maya Chen should not be enrolled in Calculus I."}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}, {"speaker": "Academic adviser", "text": "The packet is being reviewed as a complete Calculus I exception submission, and the recorded score remains 79."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and both contexts preserve the governing policy. Both contexts retain Maya Chen, Calculus I, the exception request, and the relevant packet facts. The focus evidence contains two complete factual sentences. The counterfactual changes only instructor support and remains consistent with the score, documents, and policy. Neither context embeds labels, answer codes, rationale, proposition IDs, or classifier instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar staff member\",\"text\":\"As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"The note in that packet states, \\\"Maya Chen may enroll in Calculus I.\\\"\"},{\"speaker\":\"Academic adviser\",\"text\":\"The packet also contains Maya Chen's academic transcript and her Calculus I placement result. The recorded placement score is 79.\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"The three listed materials were checked as belonging to Maya Chen, and the placement result is recorded in the packet for the requested Calculus I exception.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra."}, {"path": ["1", "text"], "text": "The note in that packet states, \"Maya Chen may enroll in Calculus I.\""}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.", "negative_left": "As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.", "negative_right": "The note in that packet states, \"Maya Chen should not enroll in Calculus I.\"", "right": "The note in that packet states, \"Maya Chen may enroll in Calculus I.\""}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-011", "id": "fast-41-diverse-078-011-base", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar staff member", "text": "As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra."}, {"speaker": "Registrar staff member", "text": "The note in that packet states, \"Maya Chen may enroll in Calculus I.\""}, {"speaker": "Academic adviser", "text": "The packet also contains Maya Chen's academic transcript and her Calculus I placement result. The recorded placement score is 79."}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}, {"speaker": "Registrar staff member", "text": "The three listed materials were checked as belonging to Maya Chen, and the placement result is recorded in the packet for the requested Calculus I exception."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and both contexts preserve the governing policy. Both contexts retain Maya Chen, Calculus I, the exception request, and the relevant packet facts. The focus evidence contains two complete factual sentences. The counterfactual changes only instructor support and remains consistent with the score, documents, and policy. Neither context embeds labels, answer codes, rationale, proposition IDs, or classifier instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar staff member\",\"text\":\"As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"The note in that packet states, \\\"Maya Chen may enroll in Calculus I.\\\"\"},{\"speaker\":\"Academic adviser\",\"text\":\"The packet also contains Maya Chen's academic transcript and her Calculus I placement result. The recorded placement score is 79.\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"The three listed materials were checked as belonging to Maya Chen, and the placement result is recorded in the packet for the requested Calculus I exception.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra."}, {"path": ["1", "text"], "text": "The note in that packet states, \"Maya Chen may enroll in Calculus I.\""}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.", "negative_left": "As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra.", "negative_right": "The note in that packet states, \"Maya Chen should not enroll in Calculus I.\"", "right": "The note in that packet states, \"Maya Chen may enroll in Calculus I.\""}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-011", "id": "fast-41-diverse-078-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar staff member", "text": "As of 2026-09-17, Maya Chen's Calculus I readiness packet contains a signed instructor note authored by Professor Ibarra."}, {"speaker": "Registrar staff member", "text": "The note in that packet states, \"Maya Chen should not enroll in Calculus I.\""}, {"speaker": "Academic adviser", "text": "The packet also contains Maya Chen's academic transcript and her Calculus I placement result. The recorded placement score is 79."}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}, {"speaker": "Registrar staff member", "text": "The three listed materials were checked as belonging to Maya Chen, and the placement result is recorded in the packet for the requested Calculus I exception."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing criteria and instructions. Both contexts retain the coordinator’s policy without invented exceptions or defaults. Maya Chen, Calculus I, and the relevant packet facts remain consistently bound. The two focus evidence spans are complete factual sentences. The counterfactual changes only instructor support, so it remains coherent with the unchanged score and documents. Neither context embeds an answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"On 14 May 2026, Maya Chen submitted her Calculus I readiness packet for review. The packet includes her transcript and placement result. Her placement score for Calculus I is 79.\"},{\"speaker\":\"Verified record\",\"text\":\"As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name.\"},{\"speaker\":\"Verified record\",\"text\":\"As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should enroll in Calculus I.”\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name."}, {"path": ["2", "text"], "text": "As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should enroll in Calculus I.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name.", "negative_left": "As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name.", "negative_right": "As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should not enroll in Calculus I.”", "right": "As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should enroll in Calculus I.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-012", "id": "fast-41-diverse-078-012-base", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Case note", "text": "On 14 May 2026, Maya Chen submitted her Calculus I readiness packet for review. The packet includes her transcript and placement result. Her placement score for Calculus I is 79."}, {"speaker": "Verified record", "text": "As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name."}, {"speaker": "Verified record", "text": "As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should enroll in Calculus I.”"}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing criteria and instructions. Both contexts retain the coordinator’s policy without invented exceptions or defaults. Maya Chen, Calculus I, and the relevant packet facts remain consistently bound. The two focus evidence spans are complete factual sentences. The counterfactual changes only instructor support, so it remains coherent with the unchanged score and documents. Neither context embeds an answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"On 14 May 2026, Maya Chen submitted her Calculus I readiness packet for review. The packet includes her transcript and placement result. Her placement score for Calculus I is 79.\"},{\"speaker\":\"Verified record\",\"text\":\"As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name.\"},{\"speaker\":\"Verified record\",\"text\":\"As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should enroll in Calculus I.”\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name."}, {"path": ["2", "text"], "text": "As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should enroll in Calculus I.”"}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name.", "negative_left": "As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name.", "negative_right": "As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should not enroll in Calculus I.”", "right": "As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should enroll in Calculus I.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-012", "id": "fast-41-diverse-078-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Case note", "text": "On 14 May 2026, Maya Chen submitted her Calculus I readiness packet for review. The packet includes her transcript and placement result. Her placement score for Calculus I is 79."}, {"speaker": "Verified record", "text": "As of 14 May 2026, Maya Chen's Calculus I readiness packet contains a signed instructor note bearing Professor Ibarra's name."}, {"speaker": "Verified record", "text": "As of 14 May 2026, the signed instructor note in that packet states, “Maya Chen should not enroll in Calculus I.”"}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim and both contexts preserve its governing policy. Both contexts retain Maya Chen, Calculus I, the packet, the score, and the instructor-note bindings. The evidence consists of two complete factual sentences: \"On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet.\" and \"On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \\\"Maya Chen should enroll in Calculus I.\\\"\" The counterfactual coherently changes only the note's support statement. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar staff member\",\"text\":\"Maya Chen's Calculus I readiness packet contains her transcript and placement result. The placement result records a Calculus I score of 79.\"},{\"speaker\":\"Course coordinator\",\"text\":\"On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet.\"},{\"speaker\":\"Academic adviser\",\"text\":\"On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \\\"Maya Chen should enroll in Calculus I.\\\" The packet is being reviewed as a Calculus I exception request.\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet."}, {"path": ["2", "text"], "text": "On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \"Maya Chen should enroll in Calculus I.\""}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet.", "negative_left": "On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet.", "negative_right": "On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \"Maya Chen should not enroll in Calculus I.\"", "right": "On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \"Maya Chen should enroll in Calculus I.\""}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-013", "id": "fast-41-diverse-078-013-base", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar staff member", "text": "Maya Chen's Calculus I readiness packet contains her transcript and placement result. The placement result records a Calculus I score of 79."}, {"speaker": "Course coordinator", "text": "On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet."}, {"speaker": "Academic adviser", "text": "On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \"Maya Chen should enroll in Calculus I.\" The packet is being reviewed as a Calculus I exception request."}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim and both contexts preserve its governing policy. Both contexts retain Maya Chen, Calculus I, the packet, the score, and the instructor-note bindings. The evidence consists of two complete factual sentences: \"On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet.\" and \"On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \\\"Maya Chen should enroll in Calculus I.\\\"\" The counterfactual coherently changes only the note's support statement. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual issue of instructor support rather than a policy conclusion. The base and counter assignments can be realized with only that support relationship changing; presence of a signed note does not itself entail that the note supports enrollment. The cited state evidence accurately preserves the coordinator’s substantive rules, while all criteria, thresholds, routing instructions, and exceptions in the questions object are automatically retained. Both rules include enough conditions to exclude competing readiness levels.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three required documents, an exact score of 79 (therefore within 75–79 and below the direct-readiness threshold), and instructor support. Under the retained criteria, this is sufficient for target 1, Provisionally ready.", "rule_index": 0, "sound": true}, {"reason": "The conjunction explicitly refutes instructor support. The Not ready criterion is disjunctive and applies whenever the instructor does not support enrollment, even though the packet is complete and the score is 79. Thus it is sufficient for target 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Maya Chen's Calculus I readiness packet contains her transcript."}, {"id": "a2", "statement": "Maya Chen's Calculus I readiness packet contains her placement result."}, {"id": "a3", "statement": "Maya Chen's Calculus I readiness packet contains Professor Ibarra's signed instructor note."}, {"id": "a4", "statement": "Maya Chen's placement score for Calculus I is 79."}, {"id": "a5", "statement": "Professor Ibarra supports Maya Chen's enrollment in Calculus I."}], "base_state_json": "[{\"speaker\":\"Registrar staff member\",\"text\":\"Maya Chen's Calculus I readiness packet contains her transcript and placement result. The placement result records a Calculus I score of 79.\"},{\"speaker\":\"Course coordinator\",\"text\":\"On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet.\"},{\"speaker\":\"Academic adviser\",\"text\":\"On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \\\"Maya Chen should enroll in Calculus I.\\\" The packet is being reviewed as a Calculus I exception request.\"},{\"speaker\":\"Course coordinator\",\"text\":\"The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet."}, {"path": ["2", "text"], "text": "On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \"Maya Chen should enroll in Calculus I.\""}], "policy_evidence": [{"path": ["2", "text"], "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}], "rules": [{"justification": "All three required documents are present, the score is within 75–79, and the instructor supports Maya. This is Provisionally ready: recommend the exception and route the complete packet to the course coordinator.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Although the packet is complete and the score is 79, the instructor explicitly does not support Maya's enrollment. This is Not ready: do not recommend the exception.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet.", "negative_left": "On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet.", "negative_right": "On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \"Maya Chen should not enroll in Calculus I.\"", "right": "On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \"Maya Chen should enroll in Calculus I.\""}, "verifier_independent_model": false}, "family": "fast-41-diverse-078-013", "id": "fast-41-diverse-078-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: The packet is incomplete, the placement score is below 75, or the instructor does not support enrollment. Do not recommend the exception; route missing-document cases to registrar staff.", "Provisionally ready: The packet contains the transcript, placement result, and signed instructor note; the score is 75–79; and the note supports this student. Recommend the exception and route the complete packet to the course coordinator.", "Directly ready: The packet is complete and the placement score is 80 or higher. The student qualifies without a prerequisite exception; route the record for standard enrollment processing."], "instructions": "Rate Maya’s Calculus I readiness using the ordered levels. Also determine from the applicable level whether the packet is complete, whether to recommend the exception, and where to route it.", "type": "score"}}, "state": [{"speaker": "Registrar staff member", "text": "Maya Chen's Calculus I readiness packet contains her transcript and placement result. The placement result records a Calculus I score of 79."}, {"speaker": "Course coordinator", "text": "On 14 August 2026, Professor Ibarra authored and signed the instructor note in Maya Chen's Calculus I readiness packet."}, {"speaker": "Academic adviser", "text": "On 14 August 2026, the instructor note in Maya Chen's Calculus I readiness packet states, \"Maya Chen should not enroll in Calculus I.\" The packet is being reviewed as a Calculus I exception request."}, {"speaker": "Course coordinator", "text": "The rule requires a transcript, placement result, and signed instructor note. A score of 80 qualifies directly; 75–79 may receive an exception with instructor support. Complete exception packets come to me."}]}, "method": "c2d", "provenance": {"source_id": "diverse-078", "source_is_synthetic": true, "source_sha256": "c36d3a2227f0d81ba827b0cc15504599bcdeb6f6a8e923ab59c14b392db1ae6f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, governing policy, booking entities, request scope, and evidence path bindings. The two evidence spans are complete factual sentences: \"At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \\\"pending sync.\\\"\" and \"At 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \\\"confirmed.\\\"\" The counterfactual coherently changes only the attachment status to \"pending sync,\" without contradictory duplicate assertions. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Desk review note\",\"text\":\"The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, for a party of six. One attendee uses a wheelchair. Walk-ins are occupying Cedar during that interval. Maple and Birch are free throughout 2:00–4:00.\"},{\"speaker\":\"Audit log\",\"text\":\"At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \\\"pending sync.\\\"\\nAt 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \\\"confirmed.\\\"\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Booking assistant\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation. The desk review covers both bookings, any required supervisor handling, and suitable fallback rooms for the accessible group.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["1", "text"], "text": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \"pending sync.\""}, {"path": ["1", "text"], "text": "At 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \"confirmed.\""}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \"pending sync.\"", "negative_left": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \"pending sync.\"", "negative_right": "At 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \"pending sync.\"", "right": "At 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \"confirmed.\""}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-001", "id": "fast-41-diverse-085-001-base", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Desk review note", "text": "The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, for a party of six. One attendee uses a wheelchair. Walk-ins are occupying Cedar during that interval. Maple and Birch are free throughout 2:00–4:00."}, {"speaker": "Audit log", "text": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \"pending sync.\"\nAt 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \"confirmed.\""}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Booking assistant", "text": "Cedar has walk-ins, who must yield to a valid reservation. The desk review covers both bookings, any required supervisor handling, and suitable fallback rooms for the accessible group."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_valid_cedar_route_conflict"}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, governing policy, booking entities, request scope, and evidence path bindings. The two evidence spans are complete factual sentences: \"At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \\\"pending sync.\\\"\" and \"At 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \\\"confirmed.\\\"\" The counterfactual coherently changes only the attachment status to \"pending sync,\" without contradictory duplicate assertions. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Desk review note\",\"text\":\"The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, for a party of six. One attendee uses a wheelchair. Walk-ins are occupying Cedar during that interval. Maple and Birch are free throughout 2:00–4:00.\"},{\"speaker\":\"Audit log\",\"text\":\"At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \\\"pending sync.\\\"\\nAt 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \\\"confirmed.\\\"\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Booking assistant\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation. The desk review covers both bookings, any required supervisor handling, and suitable fallback rooms for the accessible group.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["1", "text"], "text": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \"pending sync.\""}, {"path": ["1", "text"], "text": "At 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \"confirmed.\""}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \"pending sync.\"", "negative_left": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \"pending sync.\"", "negative_right": "At 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \"pending sync.\"", "right": "At 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \"confirmed.\""}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-001", "id": "fast-41-diverse-085-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Desk review note", "text": "The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, for a party of six. One attendee uses a wheelchair. Walk-ins are occupying Cedar during that interval. Maple and Birch are free throughout 2:00–4:00."}, {"speaker": "Audit log", "text": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status \"pending sync.\"\nAt 2026-09-17 09:00 UTC, the audit attachment for booking R-219 records status \"pending sync.\""}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Booking assistant", "text": "Cedar has walk-ins, who must yield to a valid reservation. The desk review covers both bookings, any required supervisor handling, and suitable fallback rooms for the accessible group."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_reject_pending_booking"}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, and the unchanged original questions preserve the remaining policy criteria. Booking IDs, room, time, capacity, accessibility, and routing bindings remain applicable. The evidence spans are the complete factual sentences “At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync.” and “At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed.” The counterfactual coherently changes only the audit status to pending sync, eliminating the conflict without contradictory duplicate assertions. Neither context contains a gold answer, answer code, rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync.\"},{\"speaker\":\"Records clerk\",\"text\":\"At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed.\"},{\"speaker\":\"Booking coordinator\",\"text\":\"The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, with a party of six including one wheelchair user. Walk-ins are occupying Cedar during that interval.\"},{\"speaker\":\"Facilities clerk\",\"text\":\"Maple and Birch are free throughout 2:00–4:00. Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess. Cedar has walk-ins, who must yield to a valid reservation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["0", "text"], "text": "At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync."}, {"path": ["1", "text"], "text": "At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync.", "negative_left": "At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync.", "negative_right": "At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as pending sync.", "right": "At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-008", "id": "fast-41-diverse-085-008-base", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Booking clerk", "text": "At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync."}, {"speaker": "Records clerk", "text": "At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed."}, {"speaker": "Booking coordinator", "text": "The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, with a party of six including one wheelchair user. Walk-ins are occupying Cedar during that interval."}, {"speaker": "Facilities clerk", "text": "Maple and Birch are free throughout 2:00–4:00. Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess. Cedar has walk-ins, who must yield to a valid reservation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_valid_cedar_route_conflict"}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, and the unchanged original questions preserve the remaining policy criteria. Booking IDs, room, time, capacity, accessibility, and routing bindings remain applicable. The evidence spans are the complete factual sentences “At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync.” and “At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed.” The counterfactual coherently changes only the audit status to pending sync, eliminating the conflict without contradictory duplicate assertions. Neither context contains a gold answer, answer code, rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync.\"},{\"speaker\":\"Records clerk\",\"text\":\"At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed.\"},{\"speaker\":\"Booking coordinator\",\"text\":\"The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, with a party of six including one wheelchair user. Walk-ins are occupying Cedar during that interval.\"},{\"speaker\":\"Facilities clerk\",\"text\":\"Maple and Birch are free throughout 2:00–4:00. Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess. Cedar has walk-ins, who must yield to a valid reservation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["0", "text"], "text": "At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync."}, {"path": ["1", "text"], "text": "At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync.", "negative_left": "At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync.", "negative_right": "At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as pending sync.", "right": "At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as confirmed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-008", "id": "fast-41-diverse-085-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Booking clerk", "text": "At 09:00 UTC on 17 September 2026, the live ledger records booking R-219 as pending sync."}, {"speaker": "Records clerk", "text": "At 09:00 UTC on 17 September 2026, the audit attachment for booking R-219 records its status as pending sync."}, {"speaker": "Booking coordinator", "text": "The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, with a party of six including one wheelchair user. Walk-ins are occupying Cedar during that interval."}, {"speaker": "Facilities clerk", "text": "Maple and Birch are free throughout 2:00–4:00. Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess. Cedar has walk-ins, who must yield to a valid reservation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_reject_pending_booking"}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing policy without adding exceptions or defaults. The booking entities, review scope, and relevant time bindings remain aligned with the original question. The evidence consists of exactly two complete factual sentences: \"At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync.\" and \"At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed.\" The counterfactual coherently changes only the audit status to pending sync and removes the conflict. Neither context embeds an answer, option code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync.\"},{\"speaker\":\"Records clerk\",\"text\":\"At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed.\"},{\"speaker\":\"Booking assistant\",\"text\":\"The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, with a party of six; one attendee uses a wheelchair. Walk-ins are occupying Cedar during that interval. Maple is free throughout 2:00–4:00, and Birch is free throughout 2:00–4:00.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Booking assistant\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation. The room records and accessibility notes are being prepared for the booking review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["0", "text"], "text": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync."}, {"path": ["1", "text"], "text": "At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync.", "negative_left": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync.", "negative_right": "At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status pending sync.", "right": "At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-009", "id": "fast-41-diverse-085-009-base", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Records clerk", "text": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync."}, {"speaker": "Records clerk", "text": "At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed."}, {"speaker": "Booking assistant", "text": "The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, with a party of six; one attendee uses a wheelchair. Walk-ins are occupying Cedar during that interval. Maple is free throughout 2:00–4:00, and Birch is free throughout 2:00–4:00."}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Booking assistant", "text": "Cedar has walk-ins, who must yield to a valid reservation. The room records and accessibility notes are being prepared for the booking review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_valid_cedar_route_conflict"}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing policy without adding exceptions or defaults. The booking entities, review scope, and relevant time bindings remain aligned with the original question. The evidence consists of exactly two complete factual sentences: \"At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync.\" and \"At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed.\" The counterfactual coherently changes only the audit status to pending sync and removes the conflict. Neither context embeds an answer, option code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync.\"},{\"speaker\":\"Records clerk\",\"text\":\"At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed.\"},{\"speaker\":\"Booking assistant\",\"text\":\"The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, with a party of six; one attendee uses a wheelchair. Walk-ins are occupying Cedar during that interval. Maple is free throughout 2:00–4:00, and Birch is free throughout 2:00–4:00.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Booking assistant\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation. The room records and accessibility notes are being prepared for the booking review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["0", "text"], "text": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync."}, {"path": ["1", "text"], "text": "At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync.", "negative_left": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync.", "negative_right": "At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status pending sync.", "right": "At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status confirmed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-009", "id": "fast-41-diverse-085-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Records clerk", "text": "At 2026-09-17 09:00 UTC, the live ledger records booking R-219 with status pending sync."}, {"speaker": "Records clerk", "text": "At 2026-09-17 09:00 UTC, the audit attachment records booking R-219 with status pending sync."}, {"speaker": "Booking assistant", "text": "The live ledger lists booking R-214 as confirmed for Cedar from 2:00–4:00, with a party of six; one attendee uses a wheelchair. Walk-ins are occupying Cedar during that interval. Maple is free throughout 2:00–4:00, and Birch is free throughout 2:00–4:00."}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Booking assistant", "text": "Cedar has walk-ins, who must yield to a valid reservation. The room records and accessibility notes are being prepared for the booking review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_reject_pending_booking"}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve the governing policy and request scope. The unchanged original state preserves all R-214/R-219 entity, path, and time bindings. The evidence spans are complete factual sentences: “The live ledger records booking R-219 with a pending-sync status.” and “The audit attachment for booking R-219 records a confirmed status.” The counterfactual changes the audit attachment to pending-sync, coherently removing the status conflict. Neither context includes a gold answer, answer code, rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Booking desk\",\"text\":\"The live ledger lists booking R-214 as confirmed. Booking R-214 is for Cedar, covers 2:00–4:00, and has a party size of six people.\"},{\"speaker\":\"Accessibility coordinator\",\"text\":\"One attendee in the group for R-214 uses a wheelchair. Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval.\"},{\"speaker\":\"Booking desk\",\"text\":\"The live ledger records booking R-219 with a pending-sync status.\"},{\"speaker\":\"Audit clerk\",\"text\":\"The audit attachment for booking R-219 records a confirmed status.\"},{\"speaker\":\"Facilities report\",\"text\":\"Maple is free throughout 2:00–4:00. Birch is free throughout 2:00–4:00.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["2", "text"], "text": "The live ledger records booking R-219 with a pending-sync status."}, {"path": ["3", "text"], "text": "The audit attachment for booking R-219 records a confirmed status."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "The live ledger records booking R-219 with a pending-sync status.", "negative_left": "The live ledger records booking R-219 with a pending-sync status.", "negative_right": "The audit attachment for booking R-219 records a pending-sync status.", "right": "The audit attachment for booking R-219 records a confirmed status."}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-011", "id": "fast-41-diverse-085-011-base", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Booking desk", "text": "The live ledger lists booking R-214 as confirmed. Booking R-214 is for Cedar, covers 2:00–4:00, and has a party size of six people."}, {"speaker": "Accessibility coordinator", "text": "One attendee in the group for R-214 uses a wheelchair. Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"speaker": "Booking desk", "text": "The live ledger records booking R-219 with a pending-sync status."}, {"speaker": "Audit clerk", "text": "The audit attachment for booking R-219 records a confirmed status."}, {"speaker": "Facilities report", "text": "Maple is free throughout 2:00–4:00. Birch is free throughout 2:00–4:00."}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Branch supervisor", "text": "Cedar has walk-ins, who must yield to a valid reservation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_valid_cedar_route_conflict"}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve the governing policy and request scope. The unchanged original state preserves all R-214/R-219 entity, path, and time bindings. The evidence spans are complete factual sentences: “The live ledger records booking R-219 with a pending-sync status.” and “The audit attachment for booking R-219 records a confirmed status.” The counterfactual changes the audit attachment to pending-sync, coherently removing the status conflict. Neither context includes a gold answer, answer code, rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Booking desk\",\"text\":\"The live ledger lists booking R-214 as confirmed. Booking R-214 is for Cedar, covers 2:00–4:00, and has a party size of six people.\"},{\"speaker\":\"Accessibility coordinator\",\"text\":\"One attendee in the group for R-214 uses a wheelchair. Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval.\"},{\"speaker\":\"Booking desk\",\"text\":\"The live ledger records booking R-219 with a pending-sync status.\"},{\"speaker\":\"Audit clerk\",\"text\":\"The audit attachment for booking R-219 records a confirmed status.\"},{\"speaker\":\"Facilities report\",\"text\":\"Maple is free throughout 2:00–4:00. Birch is free throughout 2:00–4:00.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["2", "text"], "text": "The live ledger records booking R-219 with a pending-sync status."}, {"path": ["3", "text"], "text": "The audit attachment for booking R-219 records a confirmed status."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "The live ledger records booking R-219 with a pending-sync status.", "negative_left": "The live ledger records booking R-219 with a pending-sync status.", "negative_right": "The audit attachment for booking R-219 records a pending-sync status.", "right": "The audit attachment for booking R-219 records a confirmed status."}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-011", "id": "fast-41-diverse-085-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Booking desk", "text": "The live ledger lists booking R-214 as confirmed. Booking R-214 is for Cedar, covers 2:00–4:00, and has a party size of six people."}, {"speaker": "Accessibility coordinator", "text": "One attendee in the group for R-214 uses a wheelchair. Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"speaker": "Booking desk", "text": "The live ledger records booking R-219 with a pending-sync status."}, {"speaker": "Audit clerk", "text": "The audit attachment for booking R-219 records a pending-sync status."}, {"speaker": "Facilities report", "text": "Maple is free throughout 2:00–4:00. Birch is free throughout 2:00–4:00."}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Branch supervisor", "text": "Cedar has walk-ins, who must yield to a valid reservation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_reject_pending_booking"}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, provide two complete factual evidence sentences, and make a coherent status change without leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Booking desk\",\"text\":\"The afternoon schedule lists R-214 as confirmed in the live ledger. It is assigned to Cedar from 2:00–4:00 for six people, including one wheelchair user. Walk-ins are occupying Cedar during that interval. Maple and Birch are free throughout 2:00–4:00. The desk is also reviewing booking R-219 for the same organizer, with its room assignment held for review.\"},{\"speaker\":\"Booking record\",\"text\":\"The live-ledger status recorded for R-219 is pending sync.\"},{\"speaker\":\"Booking record\",\"text\":\"The audit-attachment status recorded for R-219 is confirmed.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Booking desk\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation. The desk will preserve the stated time and party size while reviewing the records.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["1", "text"], "text": "The live-ledger status recorded for R-219 is pending sync."}, {"path": ["2", "text"], "text": "The audit-attachment status recorded for R-219 is confirmed."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "The live-ledger status recorded for R-219 is pending sync.", "negative_left": "The live-ledger status recorded for R-219 is pending sync.", "negative_right": "The audit-attachment status recorded for R-219 is pending sync.", "right": "The audit-attachment status recorded for R-219 is confirmed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-014", "id": "fast-41-diverse-085-014-base", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Booking desk", "text": "The afternoon schedule lists R-214 as confirmed in the live ledger. It is assigned to Cedar from 2:00–4:00 for six people, including one wheelchair user. Walk-ins are occupying Cedar during that interval. Maple and Birch are free throughout 2:00–4:00. The desk is also reviewing booking R-219 for the same organizer, with its room assignment held for review."}, {"speaker": "Booking record", "text": "The live-ledger status recorded for R-219 is pending sync."}, {"speaker": "Booking record", "text": "The audit-attachment status recorded for R-219 is confirmed."}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Booking desk", "text": "Cedar has walk-ins, who must yield to a valid reservation. The desk will preserve the stated time and party size while reviewing the records."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_valid_cedar_route_conflict"}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, provide two complete factual evidence sentences, and make a coherent status change without leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual relationships, and the focus atom is the factual relation between two recorded statuses rather than a policy conclusion. The base and counter assignments are realizable by changing only R-219's audit status: in the base it differs from the pending live-ledger status, while in the counter it is also pending. The policy evidence preserves the substantive state-originating rules needed to interpret the unchanged question, including ledger validity, conflict routing, organizer limits, room accessibility/capacity, fallback ranking, and the walk-in rule. Both rules are sufficient for their respective option targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that R-214 is ledger-confirmed, covers Cedar for six people including a wheelchair user, and overlaps Cedar walk-ins; policy therefore makes R-214 valid and requires the walk-ins to yield. R-219 is pending in the live ledger with a differing audit status, satisfying the conflicting-sync condition for supervisor routing rather than rejection. Maple and Birch are free and step-free; Maple exactly fits six and therefore ranks above larger Birch, while Pine is unsuitable because it is not step-free.", "rule_index": 0, "sound": true}, {"reason": "With R-219 pending in the live ledger and the status-difference proposition refuted, its audit status is not different from the pending ledger status, so no conflicting-sync exception requires supervisor routing. Because live-ledger confirmation is necessary for validity, R-219 is rejected as pending. The remaining literals establish R-214's validity, the walk-ins' duty to yield, and Maple's priority over Birch based on exact capacity after access suitability.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The live ledger lists booking R-214 as confirmed."}, {"id": "a2", "statement": "Booking R-214 is for Cedar."}, {"id": "a3", "statement": "Booking R-214 has a party size of six people."}, {"id": "a4", "statement": "Booking R-214 covers 2:00–4:00."}, {"id": "a5", "statement": "One attendee in the group for R-214 uses a wheelchair."}, {"id": "a6", "statement": "Walk-ins are occupying Cedar during R-214's 2:00–4:00 interval."}, {"id": "a7", "statement": "Maple is free throughout 2:00–4:00."}, {"id": "a8", "statement": "Birch is free throughout 2:00–4:00."}, {"id": "a9", "statement": "The live ledger lists booking R-219 as pending sync."}, {"id": "a10", "statement": "The live-ledger status and audit-attachment status recorded for R-219 are different."}], "base_state_json": "[{\"speaker\":\"Booking desk\",\"text\":\"The afternoon schedule lists R-214 as confirmed in the live ledger. It is assigned to Cedar from 2:00–4:00 for six people, including one wheelchair user. Walk-ins are occupying Cedar during that interval. Maple and Birch are free throughout 2:00–4:00. The desk is also reviewing booking R-219 for the same organizer, with its room assignment held for review.\"},{\"speaker\":\"Booking record\",\"text\":\"The live-ledger status recorded for R-219 is pending sync.\"},{\"speaker\":\"Booking record\",\"text\":\"The audit-attachment status recorded for R-219 is confirmed.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess.\"},{\"speaker\":\"Booking desk\",\"text\":\"Cedar has walk-ins, who must yield to a valid reservation. The desk will preserve the stated time and party size while reviewing the records.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["1", "text"], "text": "The live-ledger status recorded for R-219 is pending sync."}, {"path": ["2", "text"], "text": "The audit-attachment status recorded for R-219 is confirmed."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"path": ["3", "text"], "text": "Cedar has walk-ins, who must yield to a valid reservation."}], "rules": [{"justification": "R-214 is valid because its live-ledger status is confirmed, so the Cedar walk-ins must yield. R-219 has pending-sync ledger status that conflicts with its audit status, so policy requires supervisor routing rather than rejection. For the six-person group requiring wheelchair access, free step-free Maple is exact-capacity and ranks ahead of free step-free Birch, while Pine is excluded because it is not step-free.", "target": "D_valid_cedar_route_conflict", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "R-214 remains valid and the Cedar walk-ins must yield. R-219's live ledger is pending sync, and explicit absence of a difference between the ledger and audit statuses removes the conflicting-sync condition that would require supervisor routing; because only a confirmed live-ledger status is valid, R-219 is rejected for its pending ledger status. Maple ranks above Birch because both are free and step-free, but Maple exactly fits six while Birch has excess capacity.", "target": "C_reject_pending_booking", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "The live-ledger status recorded for R-219 is pending sync.", "negative_left": "The live-ledger status recorded for R-219 is pending sync.", "negative_right": "The audit-attachment status recorded for R-219 is pending sync.", "right": "The audit-attachment status recorded for R-219 is confirmed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-085-014", "id": "fast-41-diverse-085-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_both_valid_choose_room": "Treat both R-214 and R-219 as valid, allow the organizer to choose Cedar or Pine, and rank Pine above Maple because it has more seats.", "B_cedar_over_capacity": "Reject R-214 because a group equal to the room capacity is too large, route R-219 to the supervisor, and rank Birch above Maple.", "C_reject_pending_booking": "Accept R-214 and require the walk-ins to yield, but reject R-219 solely because the ledger says pending; rank Maple above Birch.", "D_valid_cedar_route_conflict": "Accept R-214 and require the walk-ins to yield; treat R-219 as unresolved and route it to the branch supervisor; rank Maple first and Birch second as fallbacks, excluding Pine for the accessibility need.", "E_none_of_above": "Choose only if none of the other options gives the correct booking statuses, routing, walk-in action, and fallback ranking."}, "instructions": "Select the single option that correctly verifies both bookings, routes any unresolved conflict, applies the walk-in rule, and ranks suitable fallback rooms under the stated policy.", "type": "choice"}}, "state": [{"speaker": "Booking desk", "text": "The afternoon schedule lists R-214 as confirmed in the live ledger. It is assigned to Cedar from 2:00–4:00 for six people, including one wheelchair user. Walk-ins are occupying Cedar during that interval. Maple and Birch are free throughout 2:00–4:00. The desk is also reviewing booking R-219 for the same organizer, with its room assignment held for review."}, {"speaker": "Booking record", "text": "The live-ledger status recorded for R-219 is pending sync."}, {"speaker": "Booking record", "text": "The audit-attachment status recorded for R-219 is pending sync."}, {"speaker": "Branch supervisor", "text": "Policy: a booking is valid only when the live ledger says confirmed; conflicting sync evidence goes to the branch supervisor, not rejection. One organizer may hold one active room. Cedar and Maple are step-free and hold six; Birch is step-free and holds ten; Pine holds eight but is not step-free. Rank alternatives by required access, exact capacity, then least excess."}, {"speaker": "Booking desk", "text": "Cedar has walk-ins, who must yield to a valid reservation. The desk will preserve the stated time and party size while reviewing the records."}]}, "method": "c2d", "provenance": {"source_id": "diverse-085", "source_is_synthetic": true, "source_sha256": "20bda98bbae1f379647575a62205722f6a62d060912100d5e5a33b96d463d6b8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_reject_pending_booking"}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings without answer leakage; the evidence is exactly the two factual sentences “For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision.” and “For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision.”; the counterfactual coherently changes Birch to 3 unused seats and retains unequal counts.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the current decision time, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping Priya’s interval was submitted later, and Leo’s Cedar request overlaps Priya’s. Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his request before the decision.\"},{\"speaker\":\"Booking clerk\",\"text\":\"Leo’s group has no accessibility need for the 2–3 reassignment. Maple and Birch each have capacity for Leo’s group, and each has a display available. Leo requested a display.\"},{\"speaker\":\"Booking clerk\",\"text\":\"For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision.\"},{\"speaker\":\"Booking clerk\",\"text\":\"For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision.\"},{\"speaker\":\"Booking clerk\",\"text\":\"The unused-seat counts for Maple and Birch are unequal. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision."}, {"path": ["3", "text"], "text": "For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision.", "negative_left": "For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision.", "negative_right": "For Leo’s described 2–3 reassignment, Birch has exactly 3 unused seats at the current booking-validity decision.", "right": "For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-001", "id": "fast-41-diverse-088-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the current decision time, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping Priya’s interval was submitted later, and Leo’s Cedar request overlaps Priya’s. Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his request before the decision."}, {"speaker": "Booking clerk", "text": "Leo’s group has no accessibility need for the 2–3 reassignment. Maple and Birch each have capacity for Leo’s group, and each has a display available. Leo requested a display."}, {"speaker": "Booking clerk", "text": "For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision."}, {"speaker": "Booking clerk", "text": "For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision."}, {"speaker": "Booking clerk", "text": "The unused-seat counts for Maple and Birch are unequal. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings without answer leakage; the evidence is exactly the two factual sentences “For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision.” and “For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision.”; the counterfactual coherently changes Birch to 3 unused seats and retains unequal counts.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the current decision time, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping Priya’s interval was submitted later, and Leo’s Cedar request overlaps Priya’s. Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his request before the decision.\"},{\"speaker\":\"Booking clerk\",\"text\":\"Leo’s group has no accessibility need for the 2–3 reassignment. Maple and Birch each have capacity for Leo’s group, and each has a display available. Leo requested a display.\"},{\"speaker\":\"Booking clerk\",\"text\":\"For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision.\"},{\"speaker\":\"Booking clerk\",\"text\":\"For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision.\"},{\"speaker\":\"Booking clerk\",\"text\":\"The unused-seat counts for Maple and Birch are unequal. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision."}, {"path": ["3", "text"], "text": "For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision.", "negative_left": "For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision.", "negative_right": "For Leo’s described 2–3 reassignment, Birch has exactly 3 unused seats at the current booking-validity decision.", "right": "For Leo’s described 2–3 reassignment, Birch has exactly 7 unused seats at the current booking-validity decision."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-001", "id": "fast-41-diverse-088-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the current decision time, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping Priya’s interval was submitted later, and Leo’s Cedar request overlaps Priya’s. Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his request before the decision."}, {"speaker": "Booking clerk", "text": "Leo’s group has no accessibility need for the 2–3 reassignment. Maple and Birch each have capacity for Leo’s group, and each has a display available. Leo requested a display."}, {"speaker": "Booking clerk", "text": "For Leo’s described 2–3 reassignment, Maple has exactly 4 unused seats at the current booking-validity decision."}, {"speaker": "Booking clerk", "text": "For Leo’s described 2–3 reassignment, Birch has exactly 3 unused seats at the current booking-validity decision."}, {"speaker": "Booking clerk", "text": "The unused-seat counts for Maple and Birch are unequal. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policies and relevant request bindings. The two evidence spans are complete factual sentences. The counterfactual changes Birch’s count consistently from 7 to 3 without contradiction. Neither context adds answer labels, rationale, identifiers, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 decision, Priya’s request for Cedar from 2–3 was valid. Leo’s Cedar request covered the same interval, while every other valid overlapping Cedar request had been submitted later than Priya’s. Leo’s five-person group had exactly two no-shows in the preceding 30 days, and no branch supervisor had approved his request before the decision.\"},{\"speaker\":\"Booking clerk\",\"text\":\"Leo requested a display, had no accessibility need, and needed a 2–3 reassignment. Maple and Birch each had capacity for his five-person group and each had a display available. Their unused-seat counts for his group were unequal. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"},{\"speaker\":\"Capacity record\",\"text\":\"At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats.\"},{\"speaker\":\"Capacity record\",\"text\":\"At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 7 seats.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats."}, {"path": ["3", "text"], "text": "At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 7 seats."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats.", "negative_left": "At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats.", "negative_right": "At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 3 seats.", "right": "At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 7 seats."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-003", "id": "fast-41-diverse-088-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the 2026-09-17 decision, Priya’s request for Cedar from 2–3 was valid. Leo’s Cedar request covered the same interval, while every other valid overlapping Cedar request had been submitted later than Priya’s. Leo’s five-person group had exactly two no-shows in the preceding 30 days, and no branch supervisor had approved his request before the decision."}, {"speaker": "Booking clerk", "text": "Leo requested a display, had no accessibility need, and needed a 2–3 reassignment. Maple and Birch each had capacity for his five-person group and each had a display available. Their unused-seat counts for his group were unequal. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}, {"speaker": "Capacity record", "text": "At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats."}, {"speaker": "Capacity record", "text": "At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 7 seats."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policies and relevant request bindings. The two evidence spans are complete factual sentences. The counterfactual changes Birch’s count consistently from 7 to 3 without contradiction. Neither context adds answer labels, rationale, identifiers, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 decision, Priya’s request for Cedar from 2–3 was valid. Leo’s Cedar request covered the same interval, while every other valid overlapping Cedar request had been submitted later than Priya’s. Leo’s five-person group had exactly two no-shows in the preceding 30 days, and no branch supervisor had approved his request before the decision.\"},{\"speaker\":\"Booking clerk\",\"text\":\"Leo requested a display, had no accessibility need, and needed a 2–3 reassignment. Maple and Birch each had capacity for his five-person group and each had a display available. Their unused-seat counts for his group were unequal. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"},{\"speaker\":\"Capacity record\",\"text\":\"At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats.\"},{\"speaker\":\"Capacity record\",\"text\":\"At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 7 seats.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats."}, {"path": ["3", "text"], "text": "At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 7 seats."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats.", "negative_left": "At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats.", "negative_right": "At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 3 seats.", "right": "At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 7 seats."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-003", "id": "fast-41-diverse-088-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the 2026-09-17 decision, Priya’s request for Cedar from 2–3 was valid. Leo’s Cedar request covered the same interval, while every other valid overlapping Cedar request had been submitted later than Priya’s. Leo’s five-person group had exactly two no-shows in the preceding 30 days, and no branch supervisor had approved his request before the decision."}, {"speaker": "Booking clerk", "text": "Leo requested a display, had no accessibility need, and needed a 2–3 reassignment. Maple and Birch each had capacity for his five-person group and each had a display available. Their unused-seat counts for his group were unequal. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}, {"speaker": "Capacity record", "text": "At the 2026-09-17 booking-validity decision, Maple’s unused-seat count for Leo’s group was 4 seats."}, {"speaker": "Capacity record", "text": "At the 2026-09-17 booking-validity decision, Birch’s unused-seat count for Leo’s group was 3 seats."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria, and both contexts retain the policy statements. The Priya, Leo, Cedar, 2–3, review-time, and reassignment bindings remain consistent. The required evidence quotes are complete factual sentences: “At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment.” and “At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment.” The counterfactual coherently changes Birch’s unused-seat measurement from 9 to 4 without duplicating contradictory measurements within that context. Neither context states a label, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the current review, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping that interval was submitted later than Priya’s. Leo’s Cedar request overlaps Priya’s request. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Booking clerk\",\"text\":\"Leo’s group has exactly two no-shows in the 30 days before this decision, and no branch supervisor approved his request before the decision. His group has no accessibility need, requested a display, and both Maple and Birch have capacity and a display for the 2–3 reassignment. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment."}, {"path": ["3", "text"], "text": "At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment.", "negative_left": "At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment.", "negative_right": "At the 2026-09-17 booking review, Birch has 4 unused seats for Leo’s group during the described 2–3 reassignment.", "right": "At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-004", "id": "fast-41-diverse-088-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the current review, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping that interval was submitted later than Priya’s. Leo’s Cedar request overlaps Priya’s request. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Booking clerk", "text": "Leo’s group has exactly two no-shows in the 30 days before this decision, and no branch supervisor approved his request before the decision. His group has no accessibility need, requested a display, and both Maple and Birch have capacity and a display for the 2–3 reassignment. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}, {"speaker": "Booking clerk", "text": "At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment."}, {"speaker": "Booking clerk", "text": "At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria, and both contexts retain the policy statements. The Priya, Leo, Cedar, 2–3, review-time, and reassignment bindings remain consistent. The required evidence quotes are complete factual sentences: “At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment.” and “At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment.” The counterfactual coherently changes Birch’s unused-seat measurement from 9 to 4 without duplicating contradictory measurements within that context. Neither context states a label, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the current review, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping that interval was submitted later than Priya’s. Leo’s Cedar request overlaps Priya’s request. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Booking clerk\",\"text\":\"Leo’s group has exactly two no-shows in the 30 days before this decision, and no branch supervisor approved his request before the decision. His group has no accessibility need, requested a display, and both Maple and Birch have capacity and a display for the 2–3 reassignment. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment."}, {"path": ["3", "text"], "text": "At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment.", "negative_left": "At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment.", "negative_right": "At the 2026-09-17 booking review, Birch has 4 unused seats for Leo’s group during the described 2–3 reassignment.", "right": "At the 2026-09-17 booking review, Birch has 9 unused seats for Leo’s group during the described 2–3 reassignment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-004", "id": "fast-41-diverse-088-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the current review, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping that interval was submitted later than Priya’s. Leo’s Cedar request overlaps Priya’s request. Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Booking clerk", "text": "Leo’s group has exactly two no-shows in the 30 days before this decision, and no branch supervisor approved his request before the decision. His group has no accessibility need, requested a display, and both Maple and Birch have capacity and a display for the 2–3 reassignment. Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}, {"speaker": "Booking clerk", "text": "At the 2026-09-17 booking review, Maple has 7 unused seats for Leo’s group during the described 2–3 reassignment."}, {"speaker": "Booking clerk", "text": "At the 2026-09-17 booking review, Birch has 4 unused seats for Leo’s group during the described 2–3 reassignment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing policies, while preserving the booking, entity, path, and time bindings. Each evidence span is a complete factual sentence, and the counterfactual's Maple and Birch counts are noncontradictory. Neither context adds answer codes, rationale, rule tables, proposition IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking ledger\",\"text\":\"At the current decision time, Priya’s request for Cedar from 2–3 is valid, and every other valid Cedar request overlapping that interval was submitted later. Leo’s Cedar request overlaps Priya’s interval. Leo has exactly two no-shows during the preceding 30 days, and no branch supervisor approved his request before the decision, so his request is being reviewed as currently invalid. Leo’s group has no accessibility need for the reassignment. Both Maple and Birch can accommodate the group, and each has a display available for Leo’s booking; Leo did request a display. Their unused-seat counts for Leo’s group are unequal.\"},{\"speaker\":\"Policy record\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Policy record\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"},{\"speaker\":\"Verified observation\",\"text\":\"At the current booking-validity decision, Maple has 3 unused seats for Leo’s group.\"},{\"speaker\":\"Verified observation\",\"text\":\"At the current booking-validity decision, Birch has 5 unused seats for Leo’s group.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["3", "text"], "text": "At the current booking-validity decision, Maple has 3 unused seats for Leo’s group."}, {"path": ["4", "text"], "text": "At the current booking-validity decision, Birch has 5 unused seats for Leo’s group."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the current booking-validity decision, Maple has 3 unused seats for Leo’s group.", "negative_left": "At the current booking-validity decision, Maple has 7 unused seats for Leo’s group.", "negative_right": "At the current booking-validity decision, Birch has 5 unused seats for Leo’s group.", "right": "At the current booking-validity decision, Birch has 5 unused seats for Leo’s group."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-005", "id": "fast-41-diverse-088-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking ledger", "text": "At the current decision time, Priya’s request for Cedar from 2–3 is valid, and every other valid Cedar request overlapping that interval was submitted later. Leo’s Cedar request overlaps Priya’s interval. Leo has exactly two no-shows during the preceding 30 days, and no branch supervisor approved his request before the decision, so his request is being reviewed as currently invalid. Leo’s group has no accessibility need for the reassignment. Both Maple and Birch can accommodate the group, and each has a display available for Leo’s booking; Leo did request a display. Their unused-seat counts for Leo’s group are unequal."}, {"speaker": "Policy record", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Policy record", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}, {"speaker": "Verified observation", "text": "At the current booking-validity decision, Maple has 3 unused seats for Leo’s group."}, {"speaker": "Verified observation", "text": "At the current booking-validity decision, Birch has 5 unused seats for Leo’s group."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing policies, while preserving the booking, entity, path, and time bindings. Each evidence span is a complete factual sentence, and the counterfactual's Maple and Birch counts are noncontradictory. Neither context adds answer codes, rationale, rule tables, proposition IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking ledger\",\"text\":\"At the current decision time, Priya’s request for Cedar from 2–3 is valid, and every other valid Cedar request overlapping that interval was submitted later. Leo’s Cedar request overlaps Priya’s interval. Leo has exactly two no-shows during the preceding 30 days, and no branch supervisor approved his request before the decision, so his request is being reviewed as currently invalid. Leo’s group has no accessibility need for the reassignment. Both Maple and Birch can accommodate the group, and each has a display available for Leo’s booking; Leo did request a display. Their unused-seat counts for Leo’s group are unequal.\"},{\"speaker\":\"Policy record\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Policy record\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"},{\"speaker\":\"Verified observation\",\"text\":\"At the current booking-validity decision, Maple has 3 unused seats for Leo’s group.\"},{\"speaker\":\"Verified observation\",\"text\":\"At the current booking-validity decision, Birch has 5 unused seats for Leo’s group.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["3", "text"], "text": "At the current booking-validity decision, Maple has 3 unused seats for Leo’s group."}, {"path": ["4", "text"], "text": "At the current booking-validity decision, Birch has 5 unused seats for Leo’s group."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the current booking-validity decision, Maple has 3 unused seats for Leo’s group.", "negative_left": "At the current booking-validity decision, Maple has 7 unused seats for Leo’s group.", "negative_right": "At the current booking-validity decision, Birch has 5 unused seats for Leo’s group.", "right": "At the current booking-validity decision, Birch has 5 unused seats for Leo’s group."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-005", "id": "fast-41-diverse-088-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking ledger", "text": "At the current decision time, Priya’s request for Cedar from 2–3 is valid, and every other valid Cedar request overlapping that interval was submitted later. Leo’s Cedar request overlaps Priya’s interval. Leo has exactly two no-shows during the preceding 30 days, and no branch supervisor approved his request before the decision, so his request is being reviewed as currently invalid. Leo’s group has no accessibility need for the reassignment. Both Maple and Birch can accommodate the group, and each has a display available for Leo’s booking; Leo did request a display. Their unused-seat counts for Leo’s group are unequal."}, {"speaker": "Policy record", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Policy record", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}, {"speaker": "Verified observation", "text": "At the current booking-validity decision, Maple has 7 unused seats for Leo’s group."}, {"speaker": "Verified observation", "text": "At the current booking-validity decision, Birch has 5 unused seats for Leo’s group."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing criteria, and both contexts retain the applicable policy statements. Priya, Leo, Cedar, Maple, Birch, the 2–3 interval, and the current decision binding are preserved. The evidence consists of the complete factual sentences \"At the current booking-validity decision, Maple had 4 unused seats for Leo’s group.\" and \"At the current booking-validity decision, Birch had 7 unused seats for Leo’s group.\" The counterfactual coherently changes only Maple’s unused-seat count from 4 to 9, leaving Birch at 7 and avoiding duplicate contradictions. Neither context states a gold answer, answer code, rule table, rationale, proposition ID, or output instruction. Natural policy terminology and factual observations do not leak the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[\"At the current decision time, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping that interval was submitted after Priya’s. Leo’s Cedar request covers the same 2–3 interval. Leo recorded exactly two no-shows during the preceding 30 days, and no branch supervisor had approved his request before this decision. His group has no accessibility need for the reassignment. Maple and Birch each have capacity at least equal to Leo’s group size, and each has a display available; Leo requested a display.\",\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\",\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\",\"At the current booking-validity decision, Maple had 4 unused seats for Leo’s group.\",\"At the current booking-validity decision, Birch had 7 unused seats for Leo’s group.\",\"The two room counts were not equal, and both rooms met Leo’s stated capacity and display requirements.\"]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["3"], "text": "At the current booking-validity decision, Maple had 4 unused seats for Leo’s group."}, {"path": ["4"], "text": "At the current booking-validity decision, Birch had 7 unused seats for Leo’s group."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the current booking-validity decision, Maple had 4 unused seats for Leo’s group.", "negative_left": "At the current booking-validity decision, Maple had 9 unused seats for Leo’s group.", "negative_right": "At the current booking-validity decision, Birch had 7 unused seats for Leo’s group.", "right": "At the current booking-validity decision, Birch had 7 unused seats for Leo’s group."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-006", "id": "fast-41-diverse-088-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": ["At the current decision time, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping that interval was submitted after Priya’s. Leo’s Cedar request covers the same 2–3 interval. Leo recorded exactly two no-shows during the preceding 30 days, and no branch supervisor had approved his request before this decision. His group has no accessibility need for the reassignment. Maple and Birch each have capacity at least equal to Leo’s group size, and each has a display available; Leo requested a display.", "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.", "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.", "At the current booking-validity decision, Maple had 4 unused seats for Leo’s group.", "At the current booking-validity decision, Birch had 7 unused seats for Leo’s group.", "The two room counts were not equal, and both rooms met Leo’s stated capacity and display requirements."]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing criteria, and both contexts retain the applicable policy statements. Priya, Leo, Cedar, Maple, Birch, the 2–3 interval, and the current decision binding are preserved. The evidence consists of the complete factual sentences \"At the current booking-validity decision, Maple had 4 unused seats for Leo’s group.\" and \"At the current booking-validity decision, Birch had 7 unused seats for Leo’s group.\" The counterfactual coherently changes only Maple’s unused-seat count from 4 to 9, leaving Birch at 7 and avoiding duplicate contradictions. Neither context states a gold answer, answer code, rule table, rationale, proposition ID, or output instruction. Natural policy terminology and factual observations do not leak the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[\"At the current decision time, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping that interval was submitted after Priya’s. Leo’s Cedar request covers the same 2–3 interval. Leo recorded exactly two no-shows during the preceding 30 days, and no branch supervisor had approved his request before this decision. His group has no accessibility need for the reassignment. Maple and Birch each have capacity at least equal to Leo’s group size, and each has a display available; Leo requested a display.\",\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\",\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\",\"At the current booking-validity decision, Maple had 4 unused seats for Leo’s group.\",\"At the current booking-validity decision, Birch had 7 unused seats for Leo’s group.\",\"The two room counts were not equal, and both rooms met Leo’s stated capacity and display requirements.\"]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["3"], "text": "At the current booking-validity decision, Maple had 4 unused seats for Leo’s group."}, {"path": ["4"], "text": "At the current booking-validity decision, Birch had 7 unused seats for Leo’s group."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the current booking-validity decision, Maple had 4 unused seats for Leo’s group.", "negative_left": "At the current booking-validity decision, Maple had 9 unused seats for Leo’s group.", "negative_right": "At the current booking-validity decision, Birch had 7 unused seats for Leo’s group.", "right": "At the current booking-validity decision, Birch had 7 unused seats for Leo’s group."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-006", "id": "fast-41-diverse-088-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": ["At the current decision time, Priya’s Cedar request for 2–3 is valid. Every other valid Cedar request overlapping that interval was submitted after Priya’s. Leo’s Cedar request covers the same 2–3 interval. Leo recorded exactly two no-shows during the preceding 30 days, and no branch supervisor had approved his request before this decision. His group has no accessibility need for the reassignment. Maple and Birch each have capacity at least equal to Leo’s group size, and each has a display available; Leo requested a display.", "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.", "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.", "At the current booking-validity decision, Maple had 9 unused seats for Leo’s group.", "At the current booking-validity decision, Birch had 7 unused seats for Leo’s group.", "The two room counts were not equal, and both rooms met Leo’s stated capacity and display requirements."]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the governing rules are unchanged and the original questions object remains verbatim. Question bindings are preserved for Priya, Leo, Cedar, the 2–3 interval, the supervisor, Maple, and Birch. Evidence consists of two complete factual sentences: \"In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats.\" and \"In the same capacity audit and at the same decision time, Birch has 7 unused seats.\" The counterfactual is coherent because only Birch’s unused-seat count changes from 7 to 3, with no contradictory duplicate measurement. Neither generated context embeds an explicit answer, label, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the current decision time, Priya’s request for Cedar from 2–3 is valid. Leo’s Cedar request covers the same interval, while every other valid overlapping Cedar request was submitted later than Priya’s. Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his request before the decision.\"},{\"speaker\":\"Reassignment coordinator\",\"text\":\"Leo’s group has no accessibility need, requested a display, and both Maple and Birch have enough capacity for the group and a display available. The audit confirms their unused-seat counts are unequal.\"},{\"speaker\":\"Capacity auditor\",\"text\":\"In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats.\"},{\"speaker\":\"Capacity auditor\",\"text\":\"In the same capacity audit and at the same decision time, Birch has 7 unused seats.\"},{\"speaker\":\"Policy ledger\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Policy ledger\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats."}, {"path": ["3", "text"], "text": "In the same capacity audit and at the same decision time, Birch has 7 unused seats."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats.", "negative_left": "In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats.", "negative_right": "In the same capacity audit and at the same decision time, Birch has 3 unused seats.", "right": "In the same capacity audit and at the same decision time, Birch has 7 unused seats."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-008", "id": "fast-41-diverse-088-008-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the current decision time, Priya’s request for Cedar from 2–3 is valid. Leo’s Cedar request covers the same interval, while every other valid overlapping Cedar request was submitted later than Priya’s. Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his request before the decision."}, {"speaker": "Reassignment coordinator", "text": "Leo’s group has no accessibility need, requested a display, and both Maple and Birch have enough capacity for the group and a display available. The audit confirms their unused-seat counts are unequal."}, {"speaker": "Capacity auditor", "text": "In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats."}, {"speaker": "Capacity auditor", "text": "In the same capacity audit and at the same decision time, Birch has 7 unused seats."}, {"speaker": "Policy ledger", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Policy ledger", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the governing rules are unchanged and the original questions object remains verbatim. Question bindings are preserved for Priya, Leo, Cedar, the 2–3 interval, the supervisor, Maple, and Birch. Evidence consists of two complete factual sentences: \"In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats.\" and \"In the same capacity audit and at the same decision time, Birch has 7 unused seats.\" The counterfactual is coherent because only Birch’s unused-seat count changes from 7 to 3, with no contradictory duplicate measurement. Neither generated context embeds an explicit answer, label, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the current decision time, Priya’s request for Cedar from 2–3 is valid. Leo’s Cedar request covers the same interval, while every other valid overlapping Cedar request was submitted later than Priya’s. Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his request before the decision.\"},{\"speaker\":\"Reassignment coordinator\",\"text\":\"Leo’s group has no accessibility need, requested a display, and both Maple and Birch have enough capacity for the group and a display available. The audit confirms their unused-seat counts are unequal.\"},{\"speaker\":\"Capacity auditor\",\"text\":\"In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats.\"},{\"speaker\":\"Capacity auditor\",\"text\":\"In the same capacity audit and at the same decision time, Birch has 7 unused seats.\"},{\"speaker\":\"Policy ledger\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Policy ledger\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats."}, {"path": ["3", "text"], "text": "In the same capacity audit and at the same decision time, Birch has 7 unused seats."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats.", "negative_left": "In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats.", "negative_right": "In the same capacity audit and at the same decision time, Birch has 3 unused seats.", "right": "In the same capacity audit and at the same decision time, Birch has 7 unused seats."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-008", "id": "fast-41-diverse-088-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the current decision time, Priya’s request for Cedar from 2–3 is valid. Leo’s Cedar request covers the same interval, while every other valid overlapping Cedar request was submitted later than Priya’s. Leo has exactly two no-shows in the preceding 30 days, and no branch supervisor approved his request before the decision."}, {"speaker": "Reassignment coordinator", "text": "Leo’s group has no accessibility need, requested a display, and both Maple and Birch have enough capacity for the group and a display available. The audit confirms their unused-seat counts are unequal."}, {"speaker": "Capacity auditor", "text": "In the capacity audit for Leo’s group at the current booking decision time, Maple has 4 unused seats."}, {"speaker": "Capacity auditor", "text": "In the same capacity audit and at the same decision time, Birch has 3 unused seats."}, {"speaker": "Policy ledger", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Policy ledger", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because both contexts retain the governing policy and the original questions remain unchanged. Question bindings are preserved for Priya, Leo, Cedar, 2–3, the branch supervisor, Maple, and Birch. Evidence consists of two complete factual sentences: “At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4.” and “At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6.” The counterfactual coherently changes Maple’s count to 9 while retaining Birch’s count of 6 and their unequal-count assertion. Neither context embeds an answer code, gold answer, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 decision, Priya’s valid Cedar request covers 2–3, and every other valid Cedar request overlapping that interval was submitted later. Leo’s Cedar request overlaps Priya’s, but Leo has exactly two no-shows in the preceding 30 days and no branch supervisor approved his request before the decision.\"},{\"speaker\":\"Reassignment coordinator\",\"text\":\"Leo’s group has no accessibility need for the 2–3 reassignment. Maple and Birch each have capacity for the group, and both have a display available; Leo requested a display. Their unused-seat counts for Leo’s group are unequal.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6.\"},{\"speaker\":\"Policy record\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Policy record\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4."}, {"path": ["3", "text"], "text": "At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4.", "negative_left": "At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 9.", "negative_right": "At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6.", "right": "At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-009", "id": "fast-41-diverse-088-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the 2026-09-17 decision, Priya’s valid Cedar request covers 2–3, and every other valid Cedar request overlapping that interval was submitted later. Leo’s Cedar request overlaps Priya’s, but Leo has exactly two no-shows in the preceding 30 days and no branch supervisor approved his request before the decision."}, {"speaker": "Reassignment coordinator", "text": "Leo’s group has no accessibility need for the 2–3 reassignment. Maple and Birch each have capacity for the group, and both have a display available; Leo requested a display. Their unused-seat counts for Leo’s group are unequal."}, {"speaker": "Booking clerk", "text": "At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4."}, {"speaker": "Booking clerk", "text": "At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6."}, {"speaker": "Policy record", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Policy record", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because both contexts retain the governing policy and the original questions remain unchanged. Question bindings are preserved for Priya, Leo, Cedar, 2–3, the branch supervisor, Maple, and Birch. Evidence consists of two complete factual sentences: “At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4.” and “At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6.” The counterfactual coherently changes Maple’s count to 9 while retaining Birch’s count of 6 and their unequal-count assertion. Neither context embeds an answer code, gold answer, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 decision, Priya’s valid Cedar request covers 2–3, and every other valid Cedar request overlapping that interval was submitted later. Leo’s Cedar request overlaps Priya’s, but Leo has exactly two no-shows in the preceding 30 days and no branch supervisor approved his request before the decision.\"},{\"speaker\":\"Reassignment coordinator\",\"text\":\"Leo’s group has no accessibility need for the 2–3 reassignment. Maple and Birch each have capacity for the group, and both have a display available; Leo requested a display. Their unused-seat counts for Leo’s group are unequal.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4.\"},{\"speaker\":\"Booking clerk\",\"text\":\"At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6.\"},{\"speaker\":\"Policy record\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Policy record\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4."}, {"path": ["3", "text"], "text": "At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 4.", "negative_left": "At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 9.", "negative_right": "At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6.", "right": "At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6."}, "verifier_independent_model": false}, "family": "fast-41-diverse-088-009", "id": "fast-41-diverse-088-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Booking clerk", "text": "At the 2026-09-17 decision, Priya’s valid Cedar request covers 2–3, and every other valid Cedar request overlapping that interval was submitted later. Leo’s Cedar request overlaps Priya’s, but Leo has exactly two no-shows in the preceding 30 days and no branch supervisor approved his request before the decision."}, {"speaker": "Reassignment coordinator", "text": "Leo’s group has no accessibility need for the 2–3 reassignment. Maple and Birch each have capacity for the group, and both have a display available; Leo requested a display. Their unused-seat counts for Leo’s group are unequal."}, {"speaker": "Booking clerk", "text": "At the 2026-09-17 booking decision, Maple’s unused-seat count for Leo’s group is 9."}, {"speaker": "Booking clerk", "text": "At the 2026-09-17 booking decision, Birch’s unused-seat count for Leo’s group is 6."}, {"speaker": "Policy record", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Policy record", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing scoring policy and criteria. Both contexts retain Mina, Omar, Elm, Birch, and the relevant booking and ranking bindings. \"In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3.\" is a complete factual sentence. \"In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7.\" is a complete factual sentence. The counterfactual changes only Birch’s ranking position and does not create a contradictory duplicate assertion. Neither context contains a label, answer code, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation; the universally quantified log atoms are not bundled classifications. A4 is a factual ranking relation rather than a policy statement. The base and counter assignments differ only on A4 and are realizable: an assistant can make all other listed decisions identically while ranking Elm above Birch in one case and failing to do so in the other. Policy evidence correctly cites the substantive booking, capacity, ranking, verification, and escalation rules originating in the original state. Rules and scoring criteria already present in the retained questions object do not need to be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails every level-4 requirement: Mina is validated and assigned Cedar, Omar’s booking is voided, Elm is ranked above Birch, every disputed log is routed to the branch supervisor, and the charger note is not relied upon. The additional log-handling atoms do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "The conjunction has both booking outcomes correct, preserves the required disputed-log routing, and excludes reliance on the charger note, while refuting the required Elm-over-Birch ranking. Thus exactly the ranking secondary requirement is wrong or omitted, which is sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assistant validates Mina’s booking."}, {"id": "A2", "statement": "The assistant assigns Mina’s group to Cedar."}, {"id": "A3", "statement": "The assistant voids Omar’s booking."}, {"id": "A4", "statement": "In the assistant’s alternative-room ranking for Omar, Elm has a numerically lower rank position than Birch."}, {"id": "A5", "statement": "Every disputed log is routed by the assistant to the branch supervisor."}, {"id": "A6", "statement": "Every log routed by the assistant to the branch supervisor is disputed."}, {"id": "A7", "statement": "Every undisputed booking log is verified by the booking assistant."}, {"id": "A8", "statement": "The assistant relies on the charger note when making the booking decision."}], "base_state_json": "\"Case note, 14 May 2026: Mina’s five-person party checked in at 2:07 for a Cedar reservation running from 2:00 to 4:00; Cedar holds six people and is accessible. The assistant validated Mina’s booking and assigned her group to Cedar. Omar’s four-person reservation for Elm ran from 2:00 to 3:30, but he arrived at 2:18, so the assistant voided Omar’s booking. In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3. In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7. Birch holds eight people but is not accessible. Every disputed log is routed by the assistant to the branch supervisor, and every log routed by the assistant to the branch supervisor is disputed. Every undisputed booking log is verified by the booking assistant. The assistant made the decision from the reservation record and arrival times; the note that Omar returned a charger yesterday was not used. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": [], "text": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3."}, {"path": [], "text": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7."}], "policy_evidence": [{"path": [], "text": "Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed."}, {"path": [], "text": "Booking assistants verify logs; only disputed records go to the branch supervisor."}], "rules": [{"justification": "All required booking outcomes, the Elm-over-Birch ranking, log handling, and the charger-note exclusion are correct, satisfying the fully correct and complete criterion.", "target": "4", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Both bookings and log routing are correct, and the charger note is not used, but the Elm-over-Birch ranking is wrong; this is exactly one incorrect secondary requirement.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3.", "negative_left": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3.", "negative_right": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 2.", "right": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-089-011", "id": "fast-41-diverse-089-011-base", "input": {"questions": {"decision": {"criteria": ["0 — Fully incorrect: reverses or ignores the booking rules, such as removing Mina, honoring Omar’s canceled reservation, or assigning a room that cannot hold the group.", "1 — Mostly incorrect: recognizes only one relevant fact but reaches incorrect booking-validity conclusions or provides unsupported room assignment and routing decisions.", "2 — Partly correct: correctly resolves only one booking, or resolves both bookings but gives neither the required alternative ranking nor the correct route for a disputed record.", "3 — Mostly correct: correctly validates Mina’s booking and voids Omar’s booking, but omits or gets wrong exactly one secondary requirement—either ranking Elm above Birch or routing a log dispute to the branch supervisor.", "4 — Fully correct and complete: validates Mina in Cedar, voids Omar’s late booking, ranks Elm above Birch for Omar under the capacity rule, routes any disputed log to the branch supervisor, and does not rely on the charger-note distractor."], "instructions": "Score the library booking assistant’s decision for completeness and correctness using the five ordered levels. Verify both bookings, the alternative-room ranking, and staff routing. Treat unrelated facts as distractors.", "type": "score"}}, "state": "Case note, 14 May 2026: Mina’s five-person party checked in at 2:07 for a Cedar reservation running from 2:00 to 4:00; Cedar holds six people and is accessible. The assistant validated Mina’s booking and assigned her group to Cedar. Omar’s four-person reservation for Elm ran from 2:00 to 3:30, but he arrived at 2:18, so the assistant voided Omar’s booking. In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3. In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7. Birch holds eight people but is not accessible. Every disputed log is routed by the assistant to the branch supervisor, and every log routed by the assistant to the branch supervisor is disputed. Every undisputed booking log is verified by the booking assistant. The assistant made the decision from the reservation record and arrival times; the note that Omar returned a charger yesterday was not used. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor."}, "method": "c2d", "provenance": {"source_id": "diverse-089", "source_is_synthetic": true, "source_sha256": "09465f5cddfe8d4e40880c4b0b843bcad80d4656a8a7c38d24afa69f4f78a0e2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing scoring policy and criteria. Both contexts retain Mina, Omar, Elm, Birch, and the relevant booking and ranking bindings. \"In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3.\" is a complete factual sentence. \"In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7.\" is a complete factual sentence. The counterfactual changes only Birch’s ranking position and does not create a contradictory duplicate assertion. Neither context contains a label, answer code, proposition identifier, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation; the universally quantified log atoms are not bundled classifications. A4 is a factual ranking relation rather than a policy statement. The base and counter assignments differ only on A4 and are realizable: an assistant can make all other listed decisions identically while ranking Elm above Birch in one case and failing to do so in the other. Policy evidence correctly cites the substantive booking, capacity, ranking, verification, and escalation rules originating in the original state. Rules and scoring criteria already present in the retained questions object do not need to be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails every level-4 requirement: Mina is validated and assigned Cedar, Omar’s booking is voided, Elm is ranked above Birch, every disputed log is routed to the branch supervisor, and the charger note is not relied upon. The additional log-handling atoms do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "The conjunction has both booking outcomes correct, preserves the required disputed-log routing, and excludes reliance on the charger note, while refuting the required Elm-over-Birch ranking. Thus exactly the ranking secondary requirement is wrong or omitted, which is sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assistant validates Mina’s booking."}, {"id": "A2", "statement": "The assistant assigns Mina’s group to Cedar."}, {"id": "A3", "statement": "The assistant voids Omar’s booking."}, {"id": "A4", "statement": "In the assistant’s alternative-room ranking for Omar, Elm has a numerically lower rank position than Birch."}, {"id": "A5", "statement": "Every disputed log is routed by the assistant to the branch supervisor."}, {"id": "A6", "statement": "Every log routed by the assistant to the branch supervisor is disputed."}, {"id": "A7", "statement": "Every undisputed booking log is verified by the booking assistant."}, {"id": "A8", "statement": "The assistant relies on the charger note when making the booking decision."}], "base_state_json": "\"Case note, 14 May 2026: Mina’s five-person party checked in at 2:07 for a Cedar reservation running from 2:00 to 4:00; Cedar holds six people and is accessible. The assistant validated Mina’s booking and assigned her group to Cedar. Omar’s four-person reservation for Elm ran from 2:00 to 3:30, but he arrived at 2:18, so the assistant voided Omar’s booking. In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3. In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7. Birch holds eight people but is not accessible. Every disputed log is routed by the assistant to the branch supervisor, and every log routed by the assistant to the branch supervisor is disputed. Every undisputed booking log is verified by the booking assistant. The assistant made the decision from the reservation record and arrival times; the note that Omar returned a charger yesterday was not used. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": [], "text": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3."}, {"path": [], "text": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7."}], "policy_evidence": [{"path": [], "text": "Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed."}, {"path": [], "text": "Booking assistants verify logs; only disputed records go to the branch supervisor."}], "rules": [{"justification": "All required booking outcomes, the Elm-over-Birch ranking, log handling, and the charger-note exclusion are correct, satisfying the fully correct and complete criterion.", "target": "4", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "Both bookings and log routing are correct, and the charger note is not used, but the Elm-over-Birch ranking is wrong; this is exactly one incorrect secondary requirement.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3.", "negative_left": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3.", "negative_right": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 2.", "right": "In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-089-011", "id": "fast-41-diverse-089-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Fully incorrect: reverses or ignores the booking rules, such as removing Mina, honoring Omar’s canceled reservation, or assigning a room that cannot hold the group.", "1 — Mostly incorrect: recognizes only one relevant fact but reaches incorrect booking-validity conclusions or provides unsupported room assignment and routing decisions.", "2 — Partly correct: correctly resolves only one booking, or resolves both bookings but gives neither the required alternative ranking nor the correct route for a disputed record.", "3 — Mostly correct: correctly validates Mina’s booking and voids Omar’s booking, but omits or gets wrong exactly one secondary requirement—either ranking Elm above Birch or routing a log dispute to the branch supervisor.", "4 — Fully correct and complete: validates Mina in Cedar, voids Omar’s late booking, ranks Elm above Birch for Omar under the capacity rule, routes any disputed log to the branch supervisor, and does not rely on the charger-note distractor."], "instructions": "Score the library booking assistant’s decision for completeness and correctness using the five ordered levels. Verify both bookings, the alternative-room ranking, and staff routing. Treat unrelated facts as distractors.", "type": "score"}}, "state": "Case note, 14 May 2026: Mina’s five-person party checked in at 2:07 for a Cedar reservation running from 2:00 to 4:00; Cedar holds six people and is accessible. The assistant validated Mina’s booking and assigned her group to Cedar. Omar’s four-person reservation for Elm ran from 2:00 to 3:30, but he arrived at 2:18, so the assistant voided Omar’s booking. In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Elm has rank position 3. In the assistant’s alternative-room ranking for Omar dated 14 May 2026, Birch has rank position 2. Birch holds eight people but is not accessible. Every disputed log is routed by the assistant to the branch supervisor, and every log routed by the assistant to the branch supervisor is disputed. Every undisputed booking log is verified by the booking assistant. The assistant made the decision from the reservation record and arrival times; the note that Omar returned a charger yesterday was not used. Rules cancel bookings after 10 minutes, require groups of 2–capacity, and rank alternatives by smallest adequate capacity, then accessibility if needed. Booking assistants verify logs; only disputed records go to the branch supervisor."}, "method": "c2d", "provenance": {"source_id": "diverse-089", "source_is_synthetic": true, "source_sha256": "09465f5cddfe8d4e40880c4b0b843bcad80d4656a8a7c38d24afa69f4f78a0e2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, and the counterfactual evidence remains coherent: \"The booking ledger for Mara lists itinerary C under accessibility record IC-47.\" and \"Record IC-47 states that its assigned traveler requires stair use on the route.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\":\"Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station ledger records the following verified accessibility entries:\",\"verified_left\":\"The booking ledger for Mara lists itinerary C under accessibility record IC-47.\",\"verified_right\":\"Record IC-47 states that its assigned traveler can complete the route without stair use.\",\"timing_notes\":\"Option A arrives at 16:10, option B at 16:30, and option C at 17:20.\",\"policy_1\":\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"policy_2\":\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["verified_left"], "text": "The booking ledger for Mara lists itinerary C under accessibility record IC-47."}, {"path": ["verified_right"], "text": "Record IC-47 states that its assigned traveler can complete the route without stair use."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The booking ledger for Mara lists itinerary C under accessibility record IC-47.", "negative_left": "The booking ledger for Mara lists itinerary C under accessibility record IC-47.", "negative_right": "Record IC-47 states that its assigned traveler requires stair use on the route.", "right": "Record IC-47 states that its assigned traveler can complete the route without stair use."}, "verifier_independent_model": false}, "family": "fast-41-diverse-092-028", "id": "fast-41-diverse-092-028-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station ledger records the following verified accessibility entries:", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule.", "timing_notes": "Option A arrives at 16:10, option B at 16:30, and option C at 17:20.", "verified_left": "The booking ledger for Mara lists itinerary C under accessibility record IC-47.", "verified_right": "Record IC-47 states that its assigned traveler can complete the route without stair use."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, and the counterfactual evidence remains coherent: \"The booking ledger for Mara lists itinerary C under accessibility record IC-47.\" and \"Record IC-47 states that its assigned traveler requires stair use on the route.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\":\"Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station ledger records the following verified accessibility entries:\",\"verified_left\":\"The booking ledger for Mara lists itinerary C under accessibility record IC-47.\",\"verified_right\":\"Record IC-47 states that its assigned traveler can complete the route without stair use.\",\"timing_notes\":\"Option A arrives at 16:10, option B at 16:30, and option C at 17:20.\",\"policy_1\":\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"policy_2\":\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["verified_left"], "text": "The booking ledger for Mara lists itinerary C under accessibility record IC-47."}, {"path": ["verified_right"], "text": "Record IC-47 states that its assigned traveler can complete the route without stair use."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The booking ledger for Mara lists itinerary C under accessibility record IC-47.", "negative_left": "The booking ledger for Mara lists itinerary C under accessibility record IC-47.", "negative_right": "Record IC-47 states that its assigned traveler requires stair use on the route.", "right": "Record IC-47 states that its assigned traveler can complete the route without stair use."}, "verifier_independent_model": false}, "family": "fast-41-diverse-092-028", "id": "fast-41-diverse-092-028-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station ledger records the following verified accessibility entries:", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule.", "timing_notes": "Option A arrives at 16:10, option B at 16:30, and option C at 17:20.", "verified_left": "The booking ledger for Mara lists itinerary C under accessibility record IC-47.", "verified_right": "Record IC-47 states that its assigned traveler requires stair use on the route."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question and all governing constraints. Option A, C88, the route, and the relevant times remain bound consistently. The evidence contains exactly two complete factual sentences. The counterfactual changes only C88’s platform and consistently makes the connection same-platform. Neither context states a gold answer, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The replacement itinerary must use no more than one transfer and reach Seaborne by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, arriving at Bellford at 16:02, followed by Coastliner C88 from Bellford at 16:09, reaching Seaborne at 17:20. This is one transfer, and the scheduled connection is seven minutes.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"schedule verifier\",\"text\":\"For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.\"},{\"speaker\":\"schedule verifier\",\"text\":\"For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["4", "text"], "text": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3."}, {"path": ["5", "text"], "text": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.", "negative_left": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.", "negative_right": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 3.", "right": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-003", "id": "fast-41-diverse-093-003-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The replacement itinerary must use no more than one transfer and reach Seaborne by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, arriving at Bellford at 16:02, followed by Coastliner C88 from Bellford at 16:09, reaching Seaborne at 17:20. This is one transfer, and the scheduled connection is seven minutes."}, {"speaker": "train operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "schedule verifier", "text": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3."}, {"speaker": "schedule verifier", "text": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question and all governing constraints. Option A, C88, the route, and the relevant times remain bound consistently. The evidence contains exactly two complete factual sentences. The counterfactual changes only C88’s platform and consistently makes the connection same-platform. Neither context states a gold answer, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The replacement itinerary must use no more than one transfer and reach Seaborne by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, arriving at Bellford at 16:02, followed by Coastliner C88 from Bellford at 16:09, reaching Seaborne at 17:20. This is one transfer, and the scheduled connection is seven minutes.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"schedule verifier\",\"text\":\"For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.\"},{\"speaker\":\"schedule verifier\",\"text\":\"For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["4", "text"], "text": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3."}, {"path": ["5", "text"], "text": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.", "negative_left": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.", "negative_right": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 3.", "right": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-003", "id": "fast-41-diverse-093-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The replacement itinerary must use no more than one transfer and reach Seaborne by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, arriving at Bellford at 16:02, followed by Coastliner C88 from Bellford at 16:09, reaching Seaborne at 17:20. This is one transfer, and the scheduled connection is seven minutes."}, {"speaker": "train operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "schedule verifier", "text": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3."}, {"speaker": "schedule verifier", "text": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 3."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions, traveler constraints, Option A, C88, and connection rules. The evidence consists of the complete factual sentences \"At Bellford on Option A, R214 arrived at 16:02 on platform 3.\" and \"At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5.\" The counterfactual consistently changes C88 to platform 3, making the seven-minute same-platform connection compatible with the five-minute minimum. Neither context states a gold answer, answer code, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rebooking agent\",\"text\":\"The replacement itinerary is Option A: R214 leaves Alderwick at 15:10, reaches Bellford at 16:02, and connects to Coastliner C88 leaving at 16:09 for Seaborne, due at 17:20. It uses one transfer, and the scheduled arrival is before the traveler's 18:00 deadline.\"},{\"speaker\":\"operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"station record\",\"text\":\"At Bellford on Option A, R214 arrived at 16:02 on platform 3.\"},{\"speaker\":\"station record\",\"text\":\"At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5.\"},{\"speaker\":\"rebooking agent\",\"text\":\"The seven-minute interval is the scheduled connection between the two services.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford on Option A, R214 arrived at 16:02 on platform 3."}, {"path": ["4", "text"], "text": "At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford on Option A, R214 arrived at 16:02 on platform 3.", "negative_left": "At Bellford on Option A, R214 arrived at 16:02 on platform 3.", "negative_right": "At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 3.", "right": "At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-004", "id": "fast-41-diverse-093-004-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rebooking agent", "text": "The replacement itinerary is Option A: R214 leaves Alderwick at 15:10, reaches Bellford at 16:02, and connects to Coastliner C88 leaving at 16:09 for Seaborne, due at 17:20. It uses one transfer, and the scheduled arrival is before the traveler's 18:00 deadline."}, {"speaker": "operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "station record", "text": "At Bellford on Option A, R214 arrived at 16:02 on platform 3."}, {"speaker": "station record", "text": "At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5."}, {"speaker": "rebooking agent", "text": "The seven-minute interval is the scheduled connection between the two services."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions, traveler constraints, Option A, C88, and connection rules. The evidence consists of the complete factual sentences \"At Bellford on Option A, R214 arrived at 16:02 on platform 3.\" and \"At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5.\" The counterfactual consistently changes C88 to platform 3, making the seven-minute same-platform connection compatible with the five-minute minimum. Neither context states a gold answer, answer code, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rebooking agent\",\"text\":\"The replacement itinerary is Option A: R214 leaves Alderwick at 15:10, reaches Bellford at 16:02, and connects to Coastliner C88 leaving at 16:09 for Seaborne, due at 17:20. It uses one transfer, and the scheduled arrival is before the traveler's 18:00 deadline.\"},{\"speaker\":\"operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"station record\",\"text\":\"At Bellford on Option A, R214 arrived at 16:02 on platform 3.\"},{\"speaker\":\"station record\",\"text\":\"At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5.\"},{\"speaker\":\"rebooking agent\",\"text\":\"The seven-minute interval is the scheduled connection between the two services.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford on Option A, R214 arrived at 16:02 on platform 3."}, {"path": ["4", "text"], "text": "At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford on Option A, R214 arrived at 16:02 on platform 3.", "negative_left": "At Bellford on Option A, R214 arrived at 16:02 on platform 3.", "negative_right": "At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 3.", "right": "At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-004", "id": "fast-41-diverse-093-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rebooking agent", "text": "The replacement itinerary is Option A: R214 leaves Alderwick at 15:10, reaches Bellford at 16:02, and connects to Coastliner C88 leaving at 16:09 for Seaborne, due at 17:20. It uses one transfer, and the scheduled arrival is before the traveler's 18:00 deadline."}, {"speaker": "operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "station record", "text": "At Bellford on Option A, R214 arrived at 16:02 on platform 3."}, {"speaker": "station record", "text": "At Bellford on Option A, Coastliner C88 departed at 16:09 from platform 3."}, {"speaker": "rebooking agent", "text": "The seven-minute interval is the scheduled connection between the two services."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions and bindings for Option A, C88, the route, transfer count, and times; the evidence quotes are \"On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.\" and \"On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6.\"; the counterfactual consistently changes C88 to platform 3 without duplicate contradictions or an embedded answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"My replacement must use no more than one transfer and reach Seaborne by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10 to Bellford, followed by Coastliner C88 to Seaborne. It reaches Seaborne at 17:20, so the itinerary has one transfer and meets the requested arrival time.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"R214 is scheduled to arrive at Bellford at 16:02, and Coastliner C88 is scheduled to depart at 16:09, leaving seven minutes between the scheduled movements.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\\nCross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"verified timetable record\",\"text\":\"On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.\"},{\"speaker\":\"verified timetable record\",\"text\":\"On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"The requested ticket is being reviewed against the station's minimum connection rule before confirmation.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["4", "text"], "text": "On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["5", "text"], "text": "On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 3.", "right": "On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-007", "id": "fast-41-diverse-093-007-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My replacement must use no more than one transfer and reach Seaborne by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10 to Bellford, followed by Coastliner C88 to Seaborne. It reaches Seaborne at 17:20, so the itinerary has one transfer and meets the requested arrival time."}, {"speaker": "station rebooking agent", "text": "R214 is scheduled to arrive at Bellford at 16:02, and Coastliner C88 is scheduled to depart at 16:09, leaving seven minutes between the scheduled movements."}, {"speaker": "train operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00.\nCross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "verified timetable record", "text": "On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"speaker": "verified timetable record", "text": "On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6."}, {"speaker": "station rebooking agent", "text": "The requested ticket is being reviewed against the station's minimum connection rule before confirmation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions and bindings for Option A, C88, the route, transfer count, and times; the evidence quotes are \"On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.\" and \"On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6.\"; the counterfactual consistently changes C88 to platform 3 without duplicate contradictions or an embedded answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"My replacement must use no more than one transfer and reach Seaborne by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10 to Bellford, followed by Coastliner C88 to Seaborne. It reaches Seaborne at 17:20, so the itinerary has one transfer and meets the requested arrival time.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"R214 is scheduled to arrive at Bellford at 16:02, and Coastliner C88 is scheduled to depart at 16:09, leaving seven minutes between the scheduled movements.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\\nCross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"verified timetable record\",\"text\":\"On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.\"},{\"speaker\":\"verified timetable record\",\"text\":\"On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"The requested ticket is being reviewed against the station's minimum connection rule before confirmation.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["4", "text"], "text": "On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["5", "text"], "text": "On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 3.", "right": "On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 6."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-007", "id": "fast-41-diverse-093-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My replacement must use no more than one transfer and reach Seaborne by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10 to Bellford, followed by Coastliner C88 to Seaborne. It reaches Seaborne at 17:20, so the itinerary has one transfer and meets the requested arrival time."}, {"speaker": "station rebooking agent", "text": "R214 is scheduled to arrive at Bellford at 16:02, and Coastliner C88 is scheduled to depart at 16:09, leaving seven minutes between the scheduled movements."}, {"speaker": "train operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00.\nCross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "verified timetable record", "text": "On Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"speaker": "verified timetable record", "text": "On Option A, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 3."}, {"speaker": "station rebooking agent", "text": "The requested ticket is being reviewed against the station's minimum connection rule before confirmation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the traveler’s one-transfer, 18:00, Option A, and C88 bindings. The evidence consists of the two complete factual sentences “For Option A at Bellford, R214 arrives at 16:02 on platform 3.” and “For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7.” The counterfactual is coherent because both trains use platform 3, making the seven-minute interval compatible with the five-minute same-platform minimum. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The replacement must use no more than the one transfer allowed for my through-ticket, and I need to reach Seaborne by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, reaches Bellford at 16:02, and connects with Coastliner C88 at 16:09. C88 reaches Seaborne at 17:20.\"},{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"For Option A at Bellford, R214 arrives at 16:02 on platform 3.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"The scheduled interval between the two trains is seven minutes. Option B is a direct service arriving at 18:05, so it misses the requested deadline.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "For Option A at Bellford, R214 arrives at 16:02 on platform 3."}, {"path": ["4", "text"], "text": "For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A at Bellford, R214 arrives at 16:02 on platform 3.", "negative_left": "For Option A at Bellford, R214 arrives at 16:02 on platform 3.", "negative_right": "For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 3.", "right": "For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-008", "id": "fast-41-diverse-093-008-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The replacement must use no more than the one transfer allowed for my through-ticket, and I need to reach Seaborne by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, reaches Bellford at 16:02, and connects with Coastliner C88 at 16:09. C88 reaches Seaborne at 17:20."}, {"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "For Option A at Bellford, R214 arrives at 16:02 on platform 3."}, {"speaker": "station rebooking agent", "text": "For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "station rebooking agent", "text": "The scheduled interval between the two trains is seven minutes. Option B is a direct service arriving at 18:05, so it misses the requested deadline."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the traveler’s one-transfer, 18:00, Option A, and C88 bindings. The evidence consists of the two complete factual sentences “For Option A at Bellford, R214 arrives at 16:02 on platform 3.” and “For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7.” The counterfactual is coherent because both trains use platform 3, making the seven-minute interval compatible with the five-minute same-platform minimum. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The replacement must use no more than the one transfer allowed for my through-ticket, and I need to reach Seaborne by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, reaches Bellford at 16:02, and connects with Coastliner C88 at 16:09. C88 reaches Seaborne at 17:20.\"},{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"For Option A at Bellford, R214 arrives at 16:02 on platform 3.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"The scheduled interval between the two trains is seven minutes. Option B is a direct service arriving at 18:05, so it misses the requested deadline.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "For Option A at Bellford, R214 arrives at 16:02 on platform 3."}, {"path": ["4", "text"], "text": "For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A at Bellford, R214 arrives at 16:02 on platform 3.", "negative_left": "For Option A at Bellford, R214 arrives at 16:02 on platform 3.", "negative_right": "For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 3.", "right": "For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-008", "id": "fast-41-diverse-093-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The replacement must use no more than the one transfer allowed for my through-ticket, and I need to reach Seaborne by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, reaches Bellford at 16:02, and connects with Coastliner C88 at 16:09. C88 reaches Seaborne at 17:20."}, {"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "For Option A at Bellford, R214 arrives at 16:02 on platform 3."}, {"speaker": "station rebooking agent", "text": "For Option A at Bellford, Coastliner C88 departs at 16:09 from platform 3."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "station rebooking agent", "text": "The scheduled interval between the two trains is seven minutes. Option B is a direct service arriving at 18:05, so it misses the requested deadline."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings, and the counterfactual coherently changes the departure to platform 3; the required factual evidence is “For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.” and “For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The replacement itinerary must use no more than one transfer and reach Seaborne by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, reaches Bellford at 16:02, and continues on Coastliner C88 departing Bellford at 16:09. C88 reaches Seaborne at 17:20.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"The interval between the two Bellford schedule times is seven minutes.\"},{\"speaker\":\"rail traveler\",\"text\":\"For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.\"},{\"speaker\":\"rail traveler\",\"text\":\"For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5.\"},{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3."}, {"path": ["4", "text"], "text": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.", "negative_left": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.", "negative_right": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 3.", "right": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-009", "id": "fast-41-diverse-093-009-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The replacement itinerary must use no more than one transfer and reach Seaborne by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, reaches Bellford at 16:02, and continues on Coastliner C88 departing Bellford at 16:09. C88 reaches Seaborne at 17:20."}, {"speaker": "station rebooking agent", "text": "The interval between the two Bellford schedule times is seven minutes."}, {"speaker": "rail traveler", "text": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3."}, {"speaker": "rail traveler", "text": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}, {"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings, and the counterfactual coherently changes the departure to platform 3; the required factual evidence is “For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.” and “For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The replacement itinerary must use no more than one transfer and reach Seaborne by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, reaches Bellford at 16:02, and continues on Coastliner C88 departing Bellford at 16:09. C88 reaches Seaborne at 17:20.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"The interval between the two Bellford schedule times is seven minutes.\"},{\"speaker\":\"rail traveler\",\"text\":\"For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.\"},{\"speaker\":\"rail traveler\",\"text\":\"For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5.\"},{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3."}, {"path": ["4", "text"], "text": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.", "negative_left": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3.", "negative_right": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 3.", "right": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-009", "id": "fast-41-diverse-093-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The replacement itinerary must use no more than one transfer and reach Seaborne by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, reaches Bellford at 16:02, and continues on Coastliner C88 departing Bellford at 16:09. C88 reaches Seaborne at 17:20."}, {"speaker": "station rebooking agent", "text": "The interval between the two Bellford schedule times is seven minutes."}, {"speaker": "rail traveler", "text": "For Option A at Bellford, R214's scheduled 16:02 arrival is on platform 3."}, {"speaker": "rail traveler", "text": "For Option A at Bellford, Coastliner C88's scheduled 16:09 departure is from platform 3."}, {"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and bindings remain preserved through the unchanged questions and matching context details. The evidence quotes are complete factual sentences: \"For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4.\" and \"For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7.\" The counterfactual consistently changes C88 to platform 4, making the seven-minute connection same-platform without contradictory duplicate assertions. Neither context states a gold answer, answer code, rule table, proposition ID, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"booking record\",\"text\":\"The canceled Alderwick–Seaborne service is being replaced with Option A for the rail traveler. The requested itinerary allows one transfer and requires arrival at Seaborne by 18:00.\"},{\"speaker\":\"station schedule\",\"text\":\"Option A uses R214, leaving Alderwick at 15:10 and reaching Bellford at 16:02, followed by Coastliner C88 leaving Bellford at 16:09 and reaching Seaborne at 17:20.\"},{\"speaker\":\"platform record\",\"text\":\"For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4.\"},{\"speaker\":\"platform record\",\"text\":\"For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7.\"},{\"speaker\":\"station rule\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rule\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"booking agent\",\"text\":\"The itinerary provides seven minutes between the scheduled arrival and departure, and the traveler asks to be placed on Option A if it meets the stated conditions.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4."}, {"path": ["3", "text"], "text": "For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4.", "negative_left": "For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4.", "negative_right": "For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 4.", "right": "For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-010", "id": "fast-41-diverse-093-010-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "booking record", "text": "The canceled Alderwick–Seaborne service is being replaced with Option A for the rail traveler. The requested itinerary allows one transfer and requires arrival at Seaborne by 18:00."}, {"speaker": "station schedule", "text": "Option A uses R214, leaving Alderwick at 15:10 and reaching Bellford at 16:02, followed by Coastliner C88 leaving Bellford at 16:09 and reaching Seaborne at 17:20."}, {"speaker": "platform record", "text": "For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4."}, {"speaker": "platform record", "text": "For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7."}, {"speaker": "station rule", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rule", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "booking agent", "text": "The itinerary provides seven minutes between the scheduled arrival and departure, and the traveler asks to be placed on Option A if it meets the stated conditions."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and bindings remain preserved through the unchanged questions and matching context details. The evidence quotes are complete factual sentences: \"For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4.\" and \"For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7.\" The counterfactual consistently changes C88 to platform 4, making the seven-minute connection same-platform without contradictory duplicate assertions. Neither context states a gold answer, answer code, rule table, proposition ID, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"booking record\",\"text\":\"The canceled Alderwick–Seaborne service is being replaced with Option A for the rail traveler. The requested itinerary allows one transfer and requires arrival at Seaborne by 18:00.\"},{\"speaker\":\"station schedule\",\"text\":\"Option A uses R214, leaving Alderwick at 15:10 and reaching Bellford at 16:02, followed by Coastliner C88 leaving Bellford at 16:09 and reaching Seaborne at 17:20.\"},{\"speaker\":\"platform record\",\"text\":\"For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4.\"},{\"speaker\":\"platform record\",\"text\":\"For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7.\"},{\"speaker\":\"station rule\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rule\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"booking agent\",\"text\":\"The itinerary provides seven minutes between the scheduled arrival and departure, and the traveler asks to be placed on Option A if it meets the stated conditions.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4."}, {"path": ["3", "text"], "text": "For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4.", "negative_left": "For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4.", "negative_right": "For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 4.", "right": "For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-010", "id": "fast-41-diverse-093-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "booking record", "text": "The canceled Alderwick–Seaborne service is being replaced with Option A for the rail traveler. The requested itinerary allows one transfer and requires arrival at Seaborne by 18:00."}, {"speaker": "station schedule", "text": "Option A uses R214, leaving Alderwick at 15:10 and reaching Bellford at 16:02, followed by Coastliner C88 leaving Bellford at 16:09 and reaching Seaborne at 17:20."}, {"speaker": "platform record", "text": "For Option A at Bellford on the recorded itinerary, R214 arrives at 16:02 on platform 4."}, {"speaker": "platform record", "text": "For Option A at Bellford on the recorded itinerary, Coastliner C88 departs at 16:09 from platform 4."}, {"speaker": "station rule", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rule", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "booking agent", "text": "The itinerary provides seven minutes between the scheduled arrival and departure, and the traveler asks to be placed on Option A if it meets the stated conditions."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions retain the governing criteria and both contexts retain the connection policy. Question bindings remain Option A, C88, Bellford, the one-transfer limit, and the 18:00 deadline. The evidence contains two complete factual sentences: \"At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4.\" and \"At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7.\" The counterfactual coherently changes C88’s departure platform to 4 while retaining the seven-minute interval and all other compatible facts. Neither context embeds an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"schedule note\",\"text\":\"At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4.\"},{\"speaker\":\"schedule note\",\"text\":\"At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7.\"},{\"speaker\":\"traveler constraint\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking record\",\"text\":\"Option A uses R214 from Alderwick to Bellford and Coastliner C88 from Bellford to Seaborne.\"},{\"speaker\":\"timing record\",\"text\":\"R214 is scheduled to reach Bellford at 16:02, and C88 is scheduled to leave at 16:09, giving seven minutes between the services.\"},{\"speaker\":\"arrival record\",\"text\":\"Coastliner C88 is scheduled to arrive at Seaborne at 17:20.\"},{\"speaker\":\"policy\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["0", "text"], "text": "At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4."}, {"path": ["1", "text"], "text": "At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4.", "negative_left": "At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4.", "negative_right": "At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 4.", "right": "At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-011", "id": "fast-41-diverse-093-011-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "schedule note", "text": "At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4."}, {"speaker": "schedule note", "text": "At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7."}, {"speaker": "traveler constraint", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking record", "text": "Option A uses R214 from Alderwick to Bellford and Coastliner C88 from Bellford to Seaborne."}, {"speaker": "timing record", "text": "R214 is scheduled to reach Bellford at 16:02, and C88 is scheduled to leave at 16:09, giving seven minutes between the services."}, {"speaker": "arrival record", "text": "Coastliner C88 is scheduled to arrive at Seaborne at 17:20."}, {"speaker": "policy", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions retain the governing criteria and both contexts retain the connection policy. Question bindings remain Option A, C88, Bellford, the one-transfer limit, and the 18:00 deadline. The evidence contains two complete factual sentences: \"At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4.\" and \"At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7.\" The counterfactual coherently changes C88’s departure platform to 4 while retaining the seven-minute interval and all other compatible facts. Neither context embeds an answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"schedule note\",\"text\":\"At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4.\"},{\"speaker\":\"schedule note\",\"text\":\"At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7.\"},{\"speaker\":\"traveler constraint\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking record\",\"text\":\"Option A uses R214 from Alderwick to Bellford and Coastliner C88 from Bellford to Seaborne.\"},{\"speaker\":\"timing record\",\"text\":\"R214 is scheduled to reach Bellford at 16:02, and C88 is scheduled to leave at 16:09, giving seven minutes between the services.\"},{\"speaker\":\"arrival record\",\"text\":\"Coastliner C88 is scheduled to arrive at Seaborne at 17:20.\"},{\"speaker\":\"policy\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["0", "text"], "text": "At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4."}, {"path": ["1", "text"], "text": "At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4.", "negative_left": "At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4.", "negative_right": "At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 4.", "right": "At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-011", "id": "fast-41-diverse-093-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "schedule note", "text": "At Bellford on Option A, R214 is scheduled to arrive at 16:02 on platform 4."}, {"speaker": "schedule note", "text": "At Bellford on Option A, Coastliner C88 is scheduled to depart at 16:09 from platform 4."}, {"speaker": "traveler constraint", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking record", "text": "Option A uses R214 from Alderwick to Bellford and Coastliner C88 from Bellford to Seaborne."}, {"speaker": "timing record", "text": "R214 is scheduled to reach Bellford at 16:02, and C88 is scheduled to leave at 16:09, giving seven minutes between the services."}, {"speaker": "arrival record", "text": "Coastliner C88 is scheduled to arrive at Seaborne at 17:20."}, {"speaker": "policy", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and both contexts retain the stated connection rules. Question bindings for Option A, C88, the traveler, transfer count, arrival deadline, and timing are preserved. Evidence consists of the complete factual sentences “On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.” and “On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5.” The counterfactual is coherent because changing C88 to platform 3 makes the seven-minute connection same-platform while leaving all other facts consistent. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The canceled 14:40 Alderwick–Seaborne service is being replaced, and I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 at 15:10 from Alderwick, reaches Bellford at 16:02, continues on Coastliner C88 at 16:09, and is scheduled to arrive at Seaborne at 17:20. The itinerary has one transfer, and the interval between the Bellford services is seven minutes.\"},{\"speaker\":\"schedule clerk\",\"text\":\"On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.\"},{\"speaker\":\"schedule clerk\",\"text\":\"On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please assess Option A against my stated requirements.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["3", "text"], "text": "On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 3.", "right": "On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-012", "id": "fast-41-diverse-093-012-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The canceled 14:40 Alderwick–Seaborne service is being replaced, and I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 at 15:10 from Alderwick, reaches Bellford at 16:02, continues on Coastliner C88 at 16:09, and is scheduled to arrive at Seaborne at 17:20. The itinerary has one transfer, and the interval between the Bellford services is seven minutes."}, {"speaker": "schedule clerk", "text": "On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"speaker": "schedule clerk", "text": "On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5."}, {"speaker": "train operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please assess Option A against my stated requirements."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and both contexts retain the stated connection rules. Question bindings for Option A, C88, the traveler, transfer count, arrival deadline, and timing are preserved. Evidence consists of the complete factual sentences “On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.” and “On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5.” The counterfactual is coherent because changing C88 to platform 3 makes the seven-minute connection same-platform while leaving all other facts consistent. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The canceled 14:40 Alderwick–Seaborne service is being replaced, and I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 at 15:10 from Alderwick, reaches Bellford at 16:02, continues on Coastliner C88 at 16:09, and is scheduled to arrive at Seaborne at 17:20. The itinerary has one transfer, and the interval between the Bellford services is seven minutes.\"},{\"speaker\":\"schedule clerk\",\"text\":\"On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.\"},{\"speaker\":\"schedule clerk\",\"text\":\"On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please assess Option A against my stated requirements.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["3", "text"], "text": "On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 3.", "right": "On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-012", "id": "fast-41-diverse-093-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The canceled 14:40 Alderwick–Seaborne service is being replaced, and I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 at 15:10 from Alderwick, reaches Bellford at 16:02, continues on Coastliner C88 at 16:09, and is scheduled to arrive at Seaborne at 17:20. The itinerary has one transfer, and the interval between the Bellford services is seven minutes."}, {"speaker": "schedule clerk", "text": "On Option A's 14 June 2026 schedule, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"speaker": "schedule clerk", "text": "On Option A's 14 June 2026 schedule, Coastliner C88 is scheduled to depart from Bellford at 16:09 on platform 3."}, {"speaker": "train operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please assess Option A against my stated requirements."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the traveler constraints, Option A/C88 bindings, relevant times, and minimum-connection policy. The two evidence quotes are complete factual sentences, and the counterfactual consistently changes C88 to the same platform, making the seven-minute transfer compatible with the five-minute same-platform rule.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The replacement must use no more than one transfer, and arrival at Seaborne must be by 18:00.\"},{\"speaker\":\"rebooking agent\",\"text\":\"Option A uses R214 from Alderwick, departing at 15:10 and reaching Bellford at 16:02. Coastliner C88 leaves Bellford at 16:09 and reaches Seaborne at 17:20.\"},{\"speaker\":\"schedule clerk\",\"text\":\"The planned handoff at Bellford therefore allows exactly seven minutes between the two services.\"},{\"speaker\":\"rail record\",\"text\":\"Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.\"},{\"speaker\":\"rail record\",\"text\":\"Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 7.\"},{\"speaker\":\"operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["4", "text"], "text": "Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3.", "right": "Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-014", "id": "fast-41-diverse-093-014-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The replacement must use no more than one transfer, and arrival at Seaborne must be by 18:00."}, {"speaker": "rebooking agent", "text": "Option A uses R214 from Alderwick, departing at 15:10 and reaching Bellford at 16:02. Coastliner C88 leaves Bellford at 16:09 and reaches Seaborne at 17:20."}, {"speaker": "schedule clerk", "text": "The planned handoff at Bellford therefore allows exactly seven minutes between the two services."}, {"speaker": "rail record", "text": "Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"speaker": "rail record", "text": "Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 7."}, {"speaker": "operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the traveler constraints, Option A/C88 bindings, relevant times, and minimum-connection policy. The two evidence quotes are complete factual sentences, and the counterfactual consistently changes C88 to the same platform, making the seven-minute transfer compatible with the five-minute same-platform rule.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The replacement must use no more than one transfer, and arrival at Seaborne must be by 18:00.\"},{\"speaker\":\"rebooking agent\",\"text\":\"Option A uses R214 from Alderwick, departing at 15:10 and reaching Bellford at 16:02. Coastliner C88 leaves Bellford at 16:09 and reaches Seaborne at 17:20.\"},{\"speaker\":\"schedule clerk\",\"text\":\"The planned handoff at Bellford therefore allows exactly seven minutes between the two services.\"},{\"speaker\":\"rail record\",\"text\":\"Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.\"},{\"speaker\":\"rail record\",\"text\":\"Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 7.\"},{\"speaker\":\"operations coordinator\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["4", "text"], "text": "Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3.", "right": "Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-014", "id": "fast-41-diverse-093-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The replacement must use no more than one transfer, and arrival at Seaborne must be by 18:00."}, {"speaker": "rebooking agent", "text": "Option A uses R214 from Alderwick, departing at 15:10 and reaching Bellford at 16:02. Coastliner C88 leaves Bellford at 16:09 and reaches Seaborne at 17:20."}, {"speaker": "schedule clerk", "text": "The planned handoff at Bellford therefore allows exactly seven minutes between the two services."}, {"speaker": "rail record", "text": "Under Option A, R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"speaker": "rail record", "text": "Under Option A, Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3."}, {"speaker": "operations coordinator", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The contexts preserve the unchanged policy and bindings, and the required evidence is complete factual sentences: “Under Option A, R214 arrives at Bellford at 16:02 on platform 3.” and “Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7.” The counterfactual coherently changes C88 to platform 3, making the seven-minute same-platform connection satisfy the stated five-minute minimum without contradiction or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, reaching Bellford at 16:02, and Coastliner C88 from Bellford at 16:09, reaching Seaborne at 17:20.\"},{\"speaker\":\"schedule clerk\",\"text\":\"Under Option A, R214 arrives at Bellford at 16:02 on platform 3.\"},{\"speaker\":\"schedule clerk\",\"text\":\"Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7.\"},{\"speaker\":\"operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rebooking agent\",\"text\":\"The scheduled interchange is seven minutes, and Option A contains one transfer. Option B is a direct service arriving at 18:05.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please hold Option A while its connection is checked.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "Under Option A, R214 arrives at Bellford at 16:02 on platform 3."}, {"path": ["3", "text"], "text": "Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "Under Option A, R214 arrives at Bellford at 16:02 on platform 3.", "negative_left": "Under Option A, R214 arrives at Bellford at 16:02 on platform 3.", "negative_right": "Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 3.", "right": "Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-015", "id": "fast-41-diverse-093-015-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, reaching Bellford at 16:02, and Coastliner C88 from Bellford at 16:09, reaching Seaborne at 17:20."}, {"speaker": "schedule clerk", "text": "Under Option A, R214 arrives at Bellford at 16:02 on platform 3."}, {"speaker": "schedule clerk", "text": "Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7."}, {"speaker": "operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rebooking agent", "text": "The scheduled interchange is seven minutes, and Option A contains one transfer. Option B is a direct service arriving at 18:05."}, {"speaker": "rail traveler", "text": "Please hold Option A while its connection is checked."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The contexts preserve the unchanged policy and bindings, and the required evidence is complete factual sentences: “Under Option A, R214 arrives at Bellford at 16:02 on platform 3.” and “Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7.” The counterfactual coherently changes C88 to platform 3, making the seven-minute same-platform connection satisfy the stated five-minute minimum without contradiction or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, reaching Bellford at 16:02, and Coastliner C88 from Bellford at 16:09, reaching Seaborne at 17:20.\"},{\"speaker\":\"schedule clerk\",\"text\":\"Under Option A, R214 arrives at Bellford at 16:02 on platform 3.\"},{\"speaker\":\"schedule clerk\",\"text\":\"Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7.\"},{\"speaker\":\"operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rebooking agent\",\"text\":\"The scheduled interchange is seven minutes, and Option A contains one transfer. Option B is a direct service arriving at 18:05.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please hold Option A while its connection is checked.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "Under Option A, R214 arrives at Bellford at 16:02 on platform 3."}, {"path": ["3", "text"], "text": "Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "Under Option A, R214 arrives at Bellford at 16:02 on platform 3.", "negative_left": "Under Option A, R214 arrives at Bellford at 16:02 on platform 3.", "negative_right": "Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 3.", "right": "Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-015", "id": "fast-41-diverse-093-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, reaching Bellford at 16:02, and Coastliner C88 from Bellford at 16:09, reaching Seaborne at 17:20."}, {"speaker": "schedule clerk", "text": "Under Option A, R214 arrives at Bellford at 16:02 on platform 3."}, {"speaker": "schedule clerk", "text": "Under Option A, Coastliner C88 departs Bellford at 16:09 from platform 3."}, {"speaker": "operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rebooking agent", "text": "The scheduled interchange is seven minutes, and Option A contains one transfer. Option B is a direct service arriving at 18:05."}, {"speaker": "rail traveler", "text": "Please hold Option A while its connection is checked."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, bindings, criteria, and exact-minimum rule, while both contexts contain coherent factual observations and the counterfactual consistently changes C88 to the same platform without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The original Alderwick–Seaborne service was canceled. I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, reaching Bellford at 16:02, then Coastliner C88 at 16:09, reaching Seaborne at 17:20. It therefore has one transfer and meets the requested arrival time.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"The interval from R214's scheduled arrival at Bellford to C88's scheduled departure is seven minutes.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 5.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3."}, {"path": ["4", "text"], "text": "At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3.", "negative_left": "At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3.", "negative_right": "At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 3.", "right": "At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-016", "id": "fast-41-diverse-093-016-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The original Alderwick–Seaborne service was canceled. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, reaching Bellford at 16:02, then Coastliner C88 at 16:09, reaching Seaborne at 17:20. It therefore has one transfer and meets the requested arrival time."}, {"speaker": "station rebooking agent", "text": "The interval from R214's scheduled arrival at Bellford to C88's scheduled departure is seven minutes."}, {"speaker": "station rebooking agent", "text": "At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3."}, {"speaker": "station rebooking agent", "text": "At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 5."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, bindings, criteria, and exact-minimum rule, while both contexts contain coherent factual observations and the counterfactual consistently changes C88 to the same platform without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"The original Alderwick–Seaborne service was canceled. I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A uses R214 from Alderwick at 15:10, reaching Bellford at 16:02, then Coastliner C88 at 16:09, reaching Seaborne at 17:20. It therefore has one transfer and meets the requested arrival time.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"The interval from R214's scheduled arrival at Bellford to C88's scheduled departure is seven minutes.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 5.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3."}, {"path": ["4", "text"], "text": "At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3.", "negative_left": "At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3.", "negative_right": "At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 3.", "right": "At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-093-016", "id": "fast-41-diverse-093-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "The original Alderwick–Seaborne service was canceled. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A uses R214 from Alderwick at 15:10, reaching Bellford at 16:02, then Coastliner C88 at 16:09, reaching Seaborne at 17:20. It therefore has one transfer and meets the requested arrival time."}, {"speaker": "station rebooking agent", "text": "The interval from R214's scheduled arrival at Bellford to C88's scheduled departure is seven minutes."}, {"speaker": "station rebooking agent", "text": "At Bellford on the stated Option A itinerary, R214 is scheduled to arrive at 16:02 on platform 3."}, {"speaker": "station rebooking agent", "text": "At Bellford on the stated Option A itinerary, Coastliner C88 is scheduled to depart at 16:09 from platform 3."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, and each has two complete factual evidence sentences; the exact evidence includes “On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.” and, respectively, “On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10.” and “On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:45.” The counterfactual’s 18:45 arrival is consistent with Birch being later than Alder and does not embed an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"Case note, 14 March 2026: After the cancellation at Lanton, the traveler accepted the rebooking agent’s interpretation that the replacement journey could use at most one transfer and that no other hard restrictions applied. The traveler also confirmed no preference that would displace earliest arrival as the ordering priority. Birch and Alder were both operating from Lanton to Merrow, accepted the traveler’s flexible ticket with its cancellation endorsement, and each had an interchange meeting the published minimum. Each used no more than one transfer. Alder was scheduled to reach Merrow at 18:05, while Birch was scheduled later. Birch’s connection was 14 minutes. Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"evidence\":[\"On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.\",\"On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["evidence", "0"], "text": "On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30."}, {"path": ["evidence", "1"], "text": "On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.", "negative_left": "On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.", "negative_right": "On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:45.", "right": "On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10."}, "verifier_independent_model": false}, "family": "fast-41-diverse-095-009", "id": "fast-41-diverse-095-009-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "Case note, 14 March 2026: After the cancellation at Lanton, the traveler accepted the rebooking agent’s interpretation that the replacement journey could use at most one transfer and that no other hard restrictions applied. The traveler also confirmed no preference that would displace earliest arrival as the ordering priority. Birch and Alder were both operating from Lanton to Merrow, accepted the traveler’s flexible ticket with its cancellation endorsement, and each had an interchange meeting the published minimum. Each used no more than one transfer. Alder was scheduled to reach Merrow at 18:05, while Birch was scheduled later. Birch’s connection was 14 minutes. Birch’s 14-minute connection exceeds the published 12-minute minimum.", "evidence": ["On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.", "On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10."]}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, and each has two complete factual evidence sentences; the exact evidence includes “On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.” and, respectively, “On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10.” and “On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:45.” The counterfactual’s 18:45 arrival is consistent with Birch being later than Alder and does not embed an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"Case note, 14 March 2026: After the cancellation at Lanton, the traveler accepted the rebooking agent’s interpretation that the replacement journey could use at most one transfer and that no other hard restrictions applied. The traveler also confirmed no preference that would displace earliest arrival as the ordering priority. Birch and Alder were both operating from Lanton to Merrow, accepted the traveler’s flexible ticket with its cancellation endorsement, and each had an interchange meeting the published minimum. Each used no more than one transfer. Alder was scheduled to reach Merrow at 18:05, while Birch was scheduled later. Birch’s connection was 14 minutes. Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"evidence\":[\"On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.\",\"On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["evidence", "0"], "text": "On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30."}, {"path": ["evidence", "1"], "text": "On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.", "negative_left": "On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.", "negative_right": "On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:45.", "right": "On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:10."}, "verifier_independent_model": false}, "family": "fast-41-diverse-095-009", "id": "fast-41-diverse-095-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "Case note, 14 March 2026: After the cancellation at Lanton, the traveler accepted the rebooking agent’s interpretation that the replacement journey could use at most one transfer and that no other hard restrictions applied. The traveler also confirmed no preference that would displace earliest arrival as the ordering priority. Birch and Alder were both operating from Lanton to Merrow, accepted the traveler’s flexible ticket with its cancellation endorsement, and each had an interchange meeting the published minimum. Each used no more than one transfer. Alder was scheduled to reach Merrow at 18:05, while Birch was scheduled later. Birch’s connection was 14 minutes. Birch’s 14-minute connection exceeds the published 12-minute minimum.", "evidence": ["On 14 March 2026, the traveler confirmed the rebooking agent’s paraphrased latest permissible arrival time at Merrow for Birch’s replacement journey from Lanton as 18:30.", "On 14 March 2026, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:45."]}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric and Birch/Lanton–Merrow bindings; both evidence spans are complete factual sentences; changing Birch’s arrival from 18:20 to 18:45 remains coherent with Alder arriving earlier and meeting the 18:30 cutoff; neither context states a gold score, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"On 2026-09-17, a traveler at Lanton reviewed endorsed replacement routes to Merrow after a cancellation. On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time. The traveler also confirmed that one transfer was the maximum, that these were the only hard constraints, and that no priority displaced earliest-arrival ordering. Birch and Alder were both operating replacement routes, and the traveler’s flexible ticket with the cancellation endorsement was valid on each. Birch’s 14-minute connection exceeds the published 12-minute minimum. Alder’s interchange met that minimum. Each route stayed within the confirmed transfer limit, and Alder met the confirmed arrival cutoff. Alder was scheduled to arrive earlier than Birch. On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:20 local time.\",\"policy\":\"Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"request\":\"Record the case facts for applying the supplied route-suitability criteria.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["context"], "text": "On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time."}, {"path": ["context"], "text": "On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:20 local time."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time.", "negative_left": "On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time.", "negative_right": "On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:45 local time.", "right": "On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:20 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-095-013", "id": "fast-41-diverse-095-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "On 2026-09-17, a traveler at Lanton reviewed endorsed replacement routes to Merrow after a cancellation. On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time. The traveler also confirmed that one transfer was the maximum, that these were the only hard constraints, and that no priority displaced earliest-arrival ordering. Birch and Alder were both operating replacement routes, and the traveler’s flexible ticket with the cancellation endorsement was valid on each. Birch’s 14-minute connection exceeds the published 12-minute minimum. Alder’s interchange met that minimum. Each route stayed within the confirmed transfer limit, and Alder met the confirmed arrival cutoff. Alder was scheduled to arrive earlier than Birch. On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:20 local time.", "policy": "Birch’s 14-minute connection exceeds the published 12-minute minimum.", "request": "Record the case facts for applying the supplied route-suitability criteria."}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric and Birch/Lanton–Merrow bindings; both evidence spans are complete factual sentences; changing Birch’s arrival from 18:20 to 18:45 remains coherent with Alder arriving earlier and meeting the 18:30 cutoff; neither context states a gold score, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"On 2026-09-17, a traveler at Lanton reviewed endorsed replacement routes to Merrow after a cancellation. On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time. The traveler also confirmed that one transfer was the maximum, that these were the only hard constraints, and that no priority displaced earliest-arrival ordering. Birch and Alder were both operating replacement routes, and the traveler’s flexible ticket with the cancellation endorsement was valid on each. Birch’s 14-minute connection exceeds the published 12-minute minimum. Alder’s interchange met that minimum. Each route stayed within the confirmed transfer limit, and Alder met the confirmed arrival cutoff. Alder was scheduled to arrive earlier than Birch. On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:20 local time.\",\"policy\":\"Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"request\":\"Record the case facts for applying the supplied route-suitability criteria.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["context"], "text": "On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time."}, {"path": ["context"], "text": "On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:20 local time."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time.", "negative_left": "On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time.", "negative_right": "On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:45 local time.", "right": "On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:20 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-095-013", "id": "fast-41-diverse-095-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "On 2026-09-17, a traveler at Lanton reviewed endorsed replacement routes to Merrow after a cancellation. On 2026-09-17, the traveler confirmed that the agent’s paraphrase set the latest permissible arrival at Merrow for Birch’s replacement journey from Lanton at 18:30 local time. The traveler also confirmed that one transfer was the maximum, that these were the only hard constraints, and that no priority displaced earliest-arrival ordering. Birch and Alder were both operating replacement routes, and the traveler’s flexible ticket with the cancellation endorsement was valid on each. Birch’s 14-minute connection exceeds the published 12-minute minimum. Alder’s interchange met that minimum. Each route stayed within the confirmed transfer limit, and Alder met the confirmed arrival cutoff. Alder was scheduled to arrive earlier than Birch. On 2026-09-17, Birch’s scheduled arrival at Merrow for its replacement journey from Lanton was 18:45 local time.", "policy": "Birch’s 14-minute connection exceeds the published 12-minute minimum.", "request": "Record the case facts for applying the supplied route-suitability criteria."}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, remain coherent, and contain no answer leakage; the exact evidence is “The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30.” and “Birch’s scheduled arrival time at Merrow is 18:20.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"After a cancellation at Lanton, the traveler reviews endorsed replacement routes to Merrow with the rebooking agent. The agent’s paraphrase is accepted without qualification: The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30. The traveler likewise confirms that one transfer is the maximum acceptable number, that there are no other hard constraints, and that no special preference displaces earliest arrival as the ordering rule. Birch and Alder are both operating replacements, and the traveler’s flexible ticket with the cancellation endorsement is valid on each. Birch’s 14-minute connection exceeds the published 12-minute minimum. Birch uses one transfer. Birch’s scheduled arrival time at Merrow is 18:20. Alder uses one transfer, has a connection meeting the published 12-minute minimum, and arrives no later than the confirmed limit. Alder’s scheduled arrival at Merrow is earlier than Birch’s. The traveler supplies no additional restriction or preference.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["context"], "text": "The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30."}, {"path": ["context"], "text": "Birch’s scheduled arrival time at Merrow is 18:20."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30.", "negative_left": "The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30.", "negative_right": "Birch’s scheduled arrival time at Merrow is 18:40.", "right": "Birch’s scheduled arrival time at Merrow is 18:20."}, "verifier_independent_model": false}, "family": "fast-41-diverse-095-016", "id": "fast-41-diverse-095-016-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "After a cancellation at Lanton, the traveler reviews endorsed replacement routes to Merrow with the rebooking agent. The agent’s paraphrase is accepted without qualification: The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30. The traveler likewise confirms that one transfer is the maximum acceptable number, that there are no other hard constraints, and that no special preference displaces earliest arrival as the ordering rule. Birch and Alder are both operating replacements, and the traveler’s flexible ticket with the cancellation endorsement is valid on each. Birch’s 14-minute connection exceeds the published 12-minute minimum. Birch uses one transfer. Birch’s scheduled arrival time at Merrow is 18:20. Alder uses one transfer, has a connection meeting the published 12-minute minimum, and arrives no later than the confirmed limit. Alder’s scheduled arrival at Merrow is earlier than Birch’s. The traveler supplies no additional restriction or preference."}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, remain coherent, and contain no answer leakage; the exact evidence is “The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30.” and “Birch’s scheduled arrival time at Merrow is 18:20.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"After a cancellation at Lanton, the traveler reviews endorsed replacement routes to Merrow with the rebooking agent. The agent’s paraphrase is accepted without qualification: The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30. The traveler likewise confirms that one transfer is the maximum acceptable number, that there are no other hard constraints, and that no special preference displaces earliest arrival as the ordering rule. Birch and Alder are both operating replacements, and the traveler’s flexible ticket with the cancellation endorsement is valid on each. Birch’s 14-minute connection exceeds the published 12-minute minimum. Birch uses one transfer. Birch’s scheduled arrival time at Merrow is 18:20. Alder uses one transfer, has a connection meeting the published 12-minute minimum, and arrives no later than the confirmed limit. Alder’s scheduled arrival at Merrow is earlier than Birch’s. The traveler supplies no additional restriction or preference.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["context"], "text": "The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30."}, {"path": ["context"], "text": "Birch’s scheduled arrival time at Merrow is 18:20."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30.", "negative_left": "The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30.", "negative_right": "Birch’s scheduled arrival time at Merrow is 18:40.", "right": "Birch’s scheduled arrival time at Merrow is 18:20."}, "verifier_independent_model": false}, "family": "fast-41-diverse-095-016", "id": "fast-41-diverse-095-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "After a cancellation at Lanton, the traveler reviews endorsed replacement routes to Merrow with the rebooking agent. The agent’s paraphrase is accepted without qualification: The agent’s paraphrased latest permissible arrival time at Merrow, confirmed by the traveler, is 18:30. The traveler likewise confirms that one transfer is the maximum acceptable number, that there are no other hard constraints, and that no special preference displaces earliest arrival as the ordering rule. Birch and Alder are both operating replacements, and the traveler’s flexible ticket with the cancellation endorsement is valid on each. Birch’s 14-minute connection exceeds the published 12-minute minimum. Birch uses one transfer. Birch’s scheduled arrival time at Merrow is 18:40. Alder uses one transfer, has a connection meeting the published 12-minute minimum, and arrives no later than the confirmed limit. Alder’s scheduled arrival at Merrow is earlier than Birch’s. The traveler supplies no additional restriction or preference."}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing vehicle, timing, evidence, accommodation, and ranking policies. The Abbey, Harbor, Lakeside Lodge, Cedar, and time bindings remain unchanged. The evidence consists of two complete factual sentences: \"Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling.\" \"Cedar drives 150 km from Abbey to Harbor before Harbor fueling.\" The counterfactual changes only the first distance to 160 km and contains no contradictory duplicate assertion, although that measurement makes Cedar infeasible under the 300 km range. Neither context embeds a gold answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "refuted", "A17": "supported", "A18": "supported", "A19": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "refuted", "A17": "supported", "A18": "supported", "A19": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "refuted", "A17": "supported", "A18": "supported", "A19": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; the universally quantified leg-evidence atoms are permitted atomic forms. A9 is a factual distance/range relation rather than a policy conclusion. The base and counter assignments can be realized as two proposed itineraries differing in pre-fueling cumulative distance while retaining the other scheduling and route properties; the counter itinerary is simply infeasible. Policy evidence preserves the state-origin timing, vehicle, rest, fueling, and ranking rules needed with the unchanged question. Omitted route details are case observations intended to be replaced by synthetic observations, not missing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction covers Cedar’s visit and check-in timing, single check-in, required rests, both fuel-range segments, Harbor refueling, explicit distance/time/surface evidence, paved-road restriction, and travel-time feasibility. A16 refuted establishes that Pine lacks an explicitly documented permitted surface, making Pine infeasible; therefore the ranking rule places achievable Cedar above Pine.", "rule_index": 0, "sound": true}, {"reason": "Refuting A9 entails that Cedar’s pre-Harbor-fueling cumulative distance exceeds the inclusive 300 km range, which is independently sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Cedar schedules the Abbey visit to begin at 10:15."}, {"id": "A2", "statement": "Cedar schedules the Abbey visit to end at 10:45."}, {"id": "A3", "statement": "Cedar schedules the Harbor tour to begin at 12:45."}, {"id": "A4", "statement": "Cedar schedules the Harbor tour to end at 13:15."}, {"id": "A5", "statement": "Cedar schedules the Lakeside Lodge check-in no later than 15:35."}, {"id": "A6", "statement": "Cedar includes exactly one lodge check-in."}, {"id": "A7", "statement": "Cedar provides at least 15 non-driving minutes before its second driving leg."}, {"id": "A8", "statement": "Cedar provides at least 15 non-driving minutes before its third driving leg."}, {"id": "A9", "statement": "Cedar’s cumulative driving distance from its full-tank departure through arrival at Harbor, before Harbor fueling, is at most 300 km."}, {"id": "A10", "statement": "Cedar’s driving distance from Harbor to Lakeside Lodge is at most 300 km."}, {"id": "A11", "statement": "Cedar fuels at Harbor before beginning the Harbor-to-Lakeside Lodge driving leg."}, {"id": "A12", "statement": "Every Cedar driving leg has an explicitly supplied distance."}, {"id": "A13", "statement": "Every Cedar driving leg has an explicitly supplied driving time."}, {"id": "A14", "statement": "Every Cedar driving leg has an explicitly supplied surface designation."}, {"id": "A15", "statement": "Every Cedar driving leg uses a paved road."}, {"id": "A16", "statement": "Pine’s road from Harbor to Hill Lodge has an explicitly documented permitted surface."}, {"id": "A17", "statement": "Cedar’s departure-to-Abbey driving time is at most 2 hours 15 minutes."}, {"id": "A18", "statement": "Cedar’s Abbey-to-Harbor driving time is at most 2 hours."}, {"id": "A19", "statement": "Cedar’s Harbor-to-Lakeside Lodge driving time is at most 2 hours."}], "base_state_json": "[{\"speaker\":\"traveler\",\"text\":\"I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00.\"},{\"speaker\":\"coordinator\",\"text\":\"The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range.\"},{\"speaker\":\"planner\",\"text\":\"Cedar’s Abbey visit begins at 10:15 and ends at 10:45; its Harbor tour begins at 12:45 and ends at 13:15. Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling. Cedar drives 150 km from Abbey to Harbor before Harbor fueling.\"},{\"speaker\":\"planner\",\"text\":\"Cedar allows at least 15 non-driving minutes before both the second and third driving legs, and its first, second, and final driving legs take no more than 2 hours 15 minutes, 2 hours, and 2 hours. Every leg has an explicitly supplied distance, driving time, and surface, and every surface is paved. Cedar fuels at Harbor before driving to Lakeside Lodge; that leg is at most 300 km. It checks in exactly once at Lakeside Lodge no later than 15:35. Pine’s Harbor-to-Hill Lodge road has no explicitly documented permitted surface.\"},{\"speaker\":\"policy\",\"text\":\"I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00.\"},{\"speaker\":\"policy\",\"text\":\"The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range.\"},{\"speaker\":\"policy\",\"text\":\"Check-in closes at 15:35.\"},{\"speaker\":\"policy\",\"text\":\"Rank achievable plans above infeasible ones, then prefer earlier check-in.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "refuted"}, {"atom_id": "A17", "state": "supported"}, {"atom_id": "A18", "state": "supported"}, {"atom_id": "A19", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "refuted"}, {"atom_id": "A17", "state": "supported"}, {"atom_id": "A18", "state": "supported"}, {"atom_id": "A19", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["2", "text"], "text": "Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling."}, {"path": ["2", "text"], "text": "Cedar drives 150 km from Abbey to Harbor before Harbor fueling."}], "policy_evidence": [{"path": ["0", "text"], "text": "I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00."}, {"path": ["1", "text"], "text": "The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range."}, {"path": ["2", "text"], "text": "Check-in closes at 15:35."}, {"path": ["3", "text"], "text": "Rank achievable plans above infeasible ones, then prefer earlier check-in."}], "rules": [{"justification": "Cedar satisfies the visit and check-in boundaries, both required rests, both inclusive-range limits, the fueling condition, the paved-road restriction, and every explicit-evidence requirement. Pine lacks an explicitly documented permitted surface and is therefore infeasible, so the ranking rule places achievable Cedar above Pine.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "refuted"}, {"atom_id": "A17", "state": "supported"}, {"atom_id": "A18", "state": "supported"}, {"atom_id": "A19", "state": "supported"}]}, {"justification": "If Cedar’s cumulative pre-fueling distance exceeds the van’s inclusive 300 km full-tank range, Cedar violates a range requirement and cannot be selected.", "target": "false", "when": [{"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling.", "negative_left": "Cedar drives 160 km from its full-tank departure to Abbey before Harbor fueling.", "negative_right": "Cedar drives 150 km from Abbey to Harbor before Harbor fueling.", "right": "Cedar drives 150 km from Abbey to Harbor before Harbor fueling."}, "verifier_independent_model": false}, "family": "fast-41-diverse-106-012", "id": "fast-41-diverse-106-012-base", "input": {"questions": {"decision": {"criteria": {"false": "Cedar violates at least one schedule, rest, range, surface, accommodation, or explicit-evidence requirement, or Pine validly outranks it.", "true": "Cedar is fully supported and achievable, while Pine is infeasible under the explicit-evidence and vehicle rules, making Cedar top-ranked."}, "instructions": "Should Cedar be selected as the top-ranked plan? Answer yes only if Cedar meets every inclusive visit and check-in boundary, rest and vehicle constraint, and every leg has an explicitly supplied distance, driving time, and permitted surface. An undocumented fact cannot be inferred. Apply the traveler’s ranking rule.", "type": "noul"}}, "state": [{"speaker": "traveler", "text": "I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00."}, {"speaker": "coordinator", "text": "The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range."}, {"speaker": "planner", "text": "Cedar’s Abbey visit begins at 10:15 and ends at 10:45; its Harbor tour begins at 12:45 and ends at 13:15. Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling. Cedar drives 150 km from Abbey to Harbor before Harbor fueling."}, {"speaker": "planner", "text": "Cedar allows at least 15 non-driving minutes before both the second and third driving legs, and its first, second, and final driving legs take no more than 2 hours 15 minutes, 2 hours, and 2 hours. Every leg has an explicitly supplied distance, driving time, and surface, and every surface is paved. Cedar fuels at Harbor before driving to Lakeside Lodge; that leg is at most 300 km. It checks in exactly once at Lakeside Lodge no later than 15:35. Pine’s Harbor-to-Hill Lodge road has no explicitly documented permitted surface."}, {"speaker": "policy", "text": "I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00."}, {"speaker": "policy", "text": "The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range."}, {"speaker": "policy", "text": "Check-in closes at 15:35."}, {"speaker": "policy", "text": "Rank achievable plans above infeasible ones, then prefer earlier check-in."}]}, "method": "c2d", "provenance": {"source_id": "diverse-106", "source_is_synthetic": true, "source_sha256": "c5d919c8b3a90ed32a789e544dbfe231dae8696a3892babc004ba7a11fed2826", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing vehicle, timing, evidence, accommodation, and ranking policies. The Abbey, Harbor, Lakeside Lodge, Cedar, and time bindings remain unchanged. The evidence consists of two complete factual sentences: \"Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling.\" \"Cedar drives 150 km from Abbey to Harbor before Harbor fueling.\" The counterfactual changes only the first distance to 160 km and contains no contradictory duplicate assertion, although that measurement makes Cedar infeasible under the 300 km range. Neither context embeds a gold answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "refuted", "A17": "supported", "A18": "supported", "A19": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "refuted", "A17": "supported", "A18": "supported", "A19": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "refuted", "A17": "supported", "A18": "supported", "A19": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; the universally quantified leg-evidence atoms are permitted atomic forms. A9 is a factual distance/range relation rather than a policy conclusion. The base and counter assignments can be realized as two proposed itineraries differing in pre-fueling cumulative distance while retaining the other scheduling and route properties; the counter itinerary is simply infeasible. Policy evidence preserves the state-origin timing, vehicle, rest, fueling, and ranking rules needed with the unchanged question. Omitted route details are case observations intended to be replaced by synthetic observations, not missing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction covers Cedar’s visit and check-in timing, single check-in, required rests, both fuel-range segments, Harbor refueling, explicit distance/time/surface evidence, paved-road restriction, and travel-time feasibility. A16 refuted establishes that Pine lacks an explicitly documented permitted surface, making Pine infeasible; therefore the ranking rule places achievable Cedar above Pine.", "rule_index": 0, "sound": true}, {"reason": "Refuting A9 entails that Cedar’s pre-Harbor-fueling cumulative distance exceeds the inclusive 300 km range, which is independently sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Cedar schedules the Abbey visit to begin at 10:15."}, {"id": "A2", "statement": "Cedar schedules the Abbey visit to end at 10:45."}, {"id": "A3", "statement": "Cedar schedules the Harbor tour to begin at 12:45."}, {"id": "A4", "statement": "Cedar schedules the Harbor tour to end at 13:15."}, {"id": "A5", "statement": "Cedar schedules the Lakeside Lodge check-in no later than 15:35."}, {"id": "A6", "statement": "Cedar includes exactly one lodge check-in."}, {"id": "A7", "statement": "Cedar provides at least 15 non-driving minutes before its second driving leg."}, {"id": "A8", "statement": "Cedar provides at least 15 non-driving minutes before its third driving leg."}, {"id": "A9", "statement": "Cedar’s cumulative driving distance from its full-tank departure through arrival at Harbor, before Harbor fueling, is at most 300 km."}, {"id": "A10", "statement": "Cedar’s driving distance from Harbor to Lakeside Lodge is at most 300 km."}, {"id": "A11", "statement": "Cedar fuels at Harbor before beginning the Harbor-to-Lakeside Lodge driving leg."}, {"id": "A12", "statement": "Every Cedar driving leg has an explicitly supplied distance."}, {"id": "A13", "statement": "Every Cedar driving leg has an explicitly supplied driving time."}, {"id": "A14", "statement": "Every Cedar driving leg has an explicitly supplied surface designation."}, {"id": "A15", "statement": "Every Cedar driving leg uses a paved road."}, {"id": "A16", "statement": "Pine’s road from Harbor to Hill Lodge has an explicitly documented permitted surface."}, {"id": "A17", "statement": "Cedar’s departure-to-Abbey driving time is at most 2 hours 15 minutes."}, {"id": "A18", "statement": "Cedar’s Abbey-to-Harbor driving time is at most 2 hours."}, {"id": "A19", "statement": "Cedar’s Harbor-to-Lakeside Lodge driving time is at most 2 hours."}], "base_state_json": "[{\"speaker\":\"traveler\",\"text\":\"I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00.\"},{\"speaker\":\"coordinator\",\"text\":\"The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range.\"},{\"speaker\":\"planner\",\"text\":\"Cedar’s Abbey visit begins at 10:15 and ends at 10:45; its Harbor tour begins at 12:45 and ends at 13:15. Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling. Cedar drives 150 km from Abbey to Harbor before Harbor fueling.\"},{\"speaker\":\"planner\",\"text\":\"Cedar allows at least 15 non-driving minutes before both the second and third driving legs, and its first, second, and final driving legs take no more than 2 hours 15 minutes, 2 hours, and 2 hours. Every leg has an explicitly supplied distance, driving time, and surface, and every surface is paved. Cedar fuels at Harbor before driving to Lakeside Lodge; that leg is at most 300 km. It checks in exactly once at Lakeside Lodge no later than 15:35. Pine’s Harbor-to-Hill Lodge road has no explicitly documented permitted surface.\"},{\"speaker\":\"policy\",\"text\":\"I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00.\"},{\"speaker\":\"policy\",\"text\":\"The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range.\"},{\"speaker\":\"policy\",\"text\":\"Check-in closes at 15:35.\"},{\"speaker\":\"policy\",\"text\":\"Rank achievable plans above infeasible ones, then prefer earlier check-in.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "refuted"}, {"atom_id": "A17", "state": "supported"}, {"atom_id": "A18", "state": "supported"}, {"atom_id": "A19", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "refuted"}, {"atom_id": "A17", "state": "supported"}, {"atom_id": "A18", "state": "supported"}, {"atom_id": "A19", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["2", "text"], "text": "Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling."}, {"path": ["2", "text"], "text": "Cedar drives 150 km from Abbey to Harbor before Harbor fueling."}], "policy_evidence": [{"path": ["0", "text"], "text": "I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00."}, {"path": ["1", "text"], "text": "The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range."}, {"path": ["2", "text"], "text": "Check-in closes at 15:35."}, {"path": ["3", "text"], "text": "Rank achievable plans above infeasible ones, then prefer earlier check-in."}], "rules": [{"justification": "Cedar satisfies the visit and check-in boundaries, both required rests, both inclusive-range limits, the fueling condition, the paved-road restriction, and every explicit-evidence requirement. Pine lacks an explicitly documented permitted surface and is therefore infeasible, so the ranking rule places achievable Cedar above Pine.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "refuted"}, {"atom_id": "A17", "state": "supported"}, {"atom_id": "A18", "state": "supported"}, {"atom_id": "A19", "state": "supported"}]}, {"justification": "If Cedar’s cumulative pre-fueling distance exceeds the van’s inclusive 300 km full-tank range, Cedar violates a range requirement and cannot be selected.", "target": "false", "when": [{"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Cedar drives 140 km from its full-tank departure to Abbey before Harbor fueling.", "negative_left": "Cedar drives 160 km from its full-tank departure to Abbey before Harbor fueling.", "negative_right": "Cedar drives 150 km from Abbey to Harbor before Harbor fueling.", "right": "Cedar drives 150 km from Abbey to Harbor before Harbor fueling."}, "verifier_independent_model": false}, "family": "fast-41-diverse-106-012", "id": "fast-41-diverse-106-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Cedar violates at least one schedule, rest, range, surface, accommodation, or explicit-evidence requirement, or Pine validly outranks it.", "true": "Cedar is fully supported and achievable, while Pine is infeasible under the explicit-evidence and vehicle rules, making Cedar top-ranked."}, "instructions": "Should Cedar be selected as the top-ranked plan? Answer yes only if Cedar meets every inclusive visit and check-in boundary, rest and vehicle constraint, and every leg has an explicitly supplied distance, driving time, and permitted surface. An undocumented fact cannot be inferred. Apply the traveler’s ranking rule.", "type": "noul"}}, "state": [{"speaker": "traveler", "text": "I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00."}, {"speaker": "coordinator", "text": "The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range."}, {"speaker": "planner", "text": "Cedar’s Abbey visit begins at 10:15 and ends at 10:45; its Harbor tour begins at 12:45 and ends at 13:15. Cedar drives 160 km from its full-tank departure to Abbey before Harbor fueling. Cedar drives 150 km from Abbey to Harbor before Harbor fueling."}, {"speaker": "planner", "text": "Cedar allows at least 15 non-driving minutes before both the second and third driving legs, and its first, second, and final driving legs take no more than 2 hours 15 minutes, 2 hours, and 2 hours. Every leg has an explicitly supplied distance, driving time, and surface, and every surface is paved. Cedar fuels at Harbor before driving to Lakeside Lodge; that leg is at most 300 km. It checks in exactly once at Lakeside Lodge no later than 15:35. Pine’s Harbor-to-Hill Lodge road has no explicitly documented permitted surface."}, {"speaker": "policy", "text": "I need the 10:15–10:45 Abbey visit and 12:45–13:15 Harbor tour, followed by one lodge check-in. I leave at 08:00."}, {"speaker": "policy", "text": "The van starts full, has a 300 km inclusive range, and may use only paved roads. Before another driving leg, 15 non-driving minutes are required; visits count. Harbor fueling takes 20 minutes and restores full range."}, {"speaker": "policy", "text": "Check-in closes at 15:35."}, {"speaker": "policy", "text": "Rank achievable plans above infeasible ones, then prefer earlier check-in."}]}, "method": "c2d", "provenance": {"source_id": "diverse-106", "source_is_synthetic": true, "source_sha256": "c5d919c8b3a90ed32a789e544dbfe231dae8696a3892babc004ba7a11fed2826", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring policy, and both contexts retain every governing policy clause. Mara, Monday, Harbor City, the itinerary, and all relevant time bindings remain unchanged. The two evidence spans are complete factual sentences: \"At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7.\" and \"At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7.\" The counterfactual coherently changes the compact-car catalog code to CC-8 without duplicating contradictory measurements. Neither context states a gold answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday Harbor City rental reservation has a canonical vehicle class belonging to {SUV, compact car}. At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7. At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7. A compact car can be made available for her 08:00 pickup if she changes the reservation beforehand, and every listed driving time is confirmed for both the reserved vehicle and that replacement. The itinerary reaches Pine Gate at 10:00, then she must take the required 30-minute rest and attend the fixed 10:30–11:30 garden visit. She reaches Seaside Ferry at 13:00 for 13:00 boarding, rides the ferry for 45 minutes, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7."}, {"path": [], "text": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7.", "negative_left": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7.", "negative_right": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-8.", "right": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-014", "id": "fast-41-diverse-107-014-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday Harbor City rental reservation has a canonical vehicle class belonging to {SUV, compact car}. At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7. At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7. A compact car can be made available for her 08:00 pickup if she changes the reservation beforehand, and every listed driving time is confirmed for both the reserved vehicle and that replacement. The itinerary reaches Pine Gate at 10:00, then she must take the required 30-minute rest and attend the fixed 10:30–11:30 garden visit. She reaches Seaside Ferry at 13:00 for 13:00 boarding, rides the ferry for 45 minutes, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the scoring policy, and both contexts retain every governing policy clause. Mara, Monday, Harbor City, the itinerary, and all relevant time bindings remain unchanged. The two evidence spans are complete factual sentences: \"At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7.\" and \"At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7.\" The counterfactual coherently changes the compact-car catalog code to CC-8 without duplicating contradictory measurements. Neither context states a gold answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday Harbor City rental reservation has a canonical vehicle class belonging to {SUV, compact car}. At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7. At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7. A compact car can be made available for her 08:00 pickup if she changes the reservation beforehand, and every listed driving time is confirmed for both the reserved vehicle and that replacement. The itinerary reaches Pine Gate at 10:00, then she must take the required 30-minute rest and attend the fixed 10:30–11:30 garden visit. She reaches Seaside Ferry at 13:00 for 13:00 boarding, rides the ferry for 45 minutes, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7."}, {"path": [], "text": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7.", "negative_left": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7.", "negative_right": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-8.", "right": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-7."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-014", "id": "fast-41-diverse-107-014-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday Harbor City rental reservation has a canonical vehicle class belonging to {SUV, compact car}. At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is CC-7. At 08:00 on Monday, the rental catalog's canonical code for a compact car is CC-8. A compact car can be made available for her 08:00 pickup if she changes the reservation beforehand, and every listed driving time is confirmed for both the reserved vehicle and that replacement. The itinerary reaches Pine Gate at 10:00, then she must take the required 30-minute rest and attend the fixed 10:30–11:30 garden visit. She reaches Seaside Ferry at 13:00 for 13:00 boarding, rides the ferry for 45 minutes, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and scoring criteria, and both contexts retain the policy wording. Mara, Monday, the rental, schedule, ferry, and lodging bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14.\" \"At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14.\" The counterfactual's changed compact-car code is consistent with the reservation retaining HX-14 and does not create a duplicate contradiction. Neither context contains a gold answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday rental file records a Harbor City pickup at 08:00, with the canonical vehicle class listed as either SUV or compact car. At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14. At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14. A compact car can be made available for that pickup if Mara changes the reservation before pickup. Driving times are confirmed for both the reserved vehicle and the available replacement. The schedule places Mara at Pine Gate at 10:00, followed by a half-hour pause and a garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, with a 13:00 boarding, and the ferry ride lasts 45 minutes. She reaches Island Lodge at 14:45; check-in runs from 15:00 to 20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14."}, {"path": [], "text": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14.", "negative_left": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14.", "negative_right": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-19.", "right": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-028", "id": "fast-41-diverse-107-028-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday rental file records a Harbor City pickup at 08:00, with the canonical vehicle class listed as either SUV or compact car. At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14. At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14. A compact car can be made available for that pickup if Mara changes the reservation before pickup. Driving times are confirmed for both the reserved vehicle and the available replacement. The schedule places Mara at Pine Gate at 10:00, followed by a half-hour pause and a garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, with a 13:00 boarding, and the ferry ride lasts 45 minutes. She reaches Island Lodge at 14:45; check-in runs from 15:00 to 20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and scoring criteria, and both contexts retain the policy wording. Mara, Monday, the rental, schedule, ferry, and lodging bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14.\" \"At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14.\" The counterfactual's changed compact-car code is consistent with the reservation retaining HX-14 and does not create a duplicate contradiction. Neither context contains a gold answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday rental file records a Harbor City pickup at 08:00, with the canonical vehicle class listed as either SUV or compact car. At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14. At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14. A compact car can be made available for that pickup if Mara changes the reservation before pickup. Driving times are confirmed for both the reserved vehicle and the available replacement. The schedule places Mara at Pine Gate at 10:00, followed by a half-hour pause and a garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, with a 13:00 boarding, and the ferry ride lasts 45 minutes. She reaches Island Lodge at 14:45; check-in runs from 15:00 to 20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14."}, {"path": [], "text": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14.", "negative_left": "At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14.", "negative_right": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-19.", "right": "At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-14."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-028", "id": "fast-41-diverse-107-028-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday rental file records a Harbor City pickup at 08:00, with the canonical vehicle class listed as either SUV or compact car. At 08:00 on Monday, the canonical vehicle-class code recorded on Mara's Harbor City rental reservation is HX-14. At 08:00 on Monday, the rental catalog's canonical code for a compact car is HX-19. A compact car can be made available for that pickup if Mara changes the reservation before pickup. Driving times are confirmed for both the reserved vehicle and the available replacement. The schedule places Mara at Pine Gate at 10:00, followed by a half-hour pause and a garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, with a 13:00 boarding, and the ferry ride lasts 45 minutes. She reaches Island Lodge at 14:45; check-in runs from 15:00 to 20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, so the governing policy is preserved. Mara, Monday, Harbor City, the route, and all time bindings remain unchanged. The two evidence spans are complete factual sentences: “The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.” and “The rental catalog's canonical code for a compact car is HC-42.” The counterfactual coherently changes the compact-car catalog code to HC-57 without creating duplicate contradictory measurements. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation for a vehicle collected in Harbor City records a canonical class belonging to {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for her Monday 08:00 Harbor City pickup if she changes the reservation before pickup. Every listed driving time is confirmed for both the vehicle on her reservation and the available replacement compact car. Her planned pickup is at 08:00. She is scheduled to reach Pine Gate at 10:00, take a rest occupying 10:00–10:30, and attend the garden from 10:30–11:30. She is due at Seaside Ferry at 13:00, with the ferry ride lasting 13:00–13:45, and at Island Lodge at 14:45. take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars. Island Lodge’s Monday check-in interval for Mara is 15:00–20:00, and her room there is confirmed for Monday.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-033", "id": "fast-41-diverse-107-033-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation for a vehicle collected in Harbor City records a canonical class belonging to {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for her Monday 08:00 Harbor City pickup if she changes the reservation before pickup. Every listed driving time is confirmed for both the vehicle on her reservation and the available replacement compact car. Her planned pickup is at 08:00. She is scheduled to reach Pine Gate at 10:00, take a rest occupying 10:00–10:30, and attend the garden from 10:30–11:30. She is due at Seaside Ferry at 13:00, with the ferry ride lasting 13:00–13:45, and at Island Lodge at 14:45. take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars. Island Lodge’s Monday check-in interval for Mara is 15:00–20:00, and her room there is confirmed for Monday."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, so the governing policy is preserved. Mara, Monday, Harbor City, the route, and all time bindings remain unchanged. The two evidence spans are complete factual sentences: “The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.” and “The rental catalog's canonical code for a compact car is HC-42.” The counterfactual coherently changes the compact-car catalog code to HC-57 without creating duplicate contradictory measurements. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation for a vehicle collected in Harbor City records a canonical class belonging to {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for her Monday 08:00 Harbor City pickup if she changes the reservation before pickup. Every listed driving time is confirmed for both the vehicle on her reservation and the available replacement compact car. Her planned pickup is at 08:00. She is scheduled to reach Pine Gate at 10:00, take a rest occupying 10:00–10:30, and attend the garden from 10:30–11:30. She is due at Seaside Ferry at 13:00, with the ferry ride lasting 13:00–13:45, and at Island Lodge at 14:45. take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars. Island Lodge’s Monday check-in interval for Mara is 15:00–20:00, and her room there is confirmed for Monday.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-033", "id": "fast-41-diverse-107-033-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation for a vehicle collected in Harbor City records a canonical class belonging to {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-57. A compact car can be made available for her Monday 08:00 Harbor City pickup if she changes the reservation before pickup. Every listed driving time is confirmed for both the vehicle on her reservation and the available replacement compact car. Her planned pickup is at 08:00. She is scheduled to reach Pine Gate at 10:00, take a rest occupying 10:00–10:30, and attend the garden from 10:30–11:30. She is due at Seaside Ferry at 13:00, with the ferry ride lasting 13:00–13:45, and at Island Lodge at 14:45. take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars. Island Lodge’s Monday check-in interval for Mara is 15:00–20:00, and her room there is confirmed for Monday."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria, and both contexts retain the same entities, route, and time bindings. The evidence contains exactly two complete factual sentences: “Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42.” and “The rental catalog's canonical code for a compact car is HC-42.” The counterfactual changes only the catalog code and remains coherent because the reservation can be changed before pickup. Neither context includes a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday Harbor City rental file lists the canonical vehicle class as a member of {SUV, compact car}. Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for Mara’s Monday 08:00 Harbor City pickup if she changes the reservation before pickup. Every listed driving time is confirmed for both the vehicle on the reservation and the available replacement compact car. The itinerary records pickup at 08:00, arrival at Pine Gate at 10:00, and a rest interval from 10:00–10:30. Mara’s garden attendance is scheduled for 10:30–11:30, followed by arrival at Seaside Ferry at 13:00 and a ferry ride from 13:00–13:45. She is due at Island Lodge at 14:45; the Monday check-in interval is 15:00–20:00, and her room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42.", "negative_left": "Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-043", "id": "fast-41-diverse-107-043-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday Harbor City rental file lists the canonical vehicle class as a member of {SUV, compact car}. Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for Mara’s Monday 08:00 Harbor City pickup if she changes the reservation before pickup. Every listed driving time is confirmed for both the vehicle on the reservation and the available replacement compact car. The itinerary records pickup at 08:00, arrival at Pine Gate at 10:00, and a rest interval from 10:00–10:30. Mara’s garden attendance is scheduled for 10:30–11:30, followed by arrival at Seaside Ferry at 13:00 and a ferry ride from 13:00–13:45. She is due at Island Lodge at 14:45; the Monday check-in interval is 15:00–20:00, and her room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria, and both contexts retain the same entities, route, and time bindings. The evidence contains exactly two complete factual sentences: “Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42.” and “The rental catalog's canonical code for a compact car is HC-42.” The counterfactual changes only the catalog code and remains coherent because the reservation can be changed before pickup. Neither context includes a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday Harbor City rental file lists the canonical vehicle class as a member of {SUV, compact car}. Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for Mara’s Monday 08:00 Harbor City pickup if she changes the reservation before pickup. Every listed driving time is confirmed for both the vehicle on the reservation and the available replacement compact car. The itinerary records pickup at 08:00, arrival at Pine Gate at 10:00, and a rest interval from 10:00–10:30. Mara’s garden attendance is scheduled for 10:30–11:30, followed by arrival at Seaside Ferry at 13:00 and a ferry ride from 13:00–13:45. She is due at Island Lodge at 14:45; the Monday check-in interval is 15:00–20:00, and her room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42.", "negative_left": "Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-043", "id": "fast-41-diverse-107-043-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday Harbor City rental file lists the canonical vehicle class as a member of {SUV, compact car}. Mara's Monday 08:00 Harbor City rental reservation records canonical vehicle-class code HC-42. The rental catalog's canonical code for a compact car is HC-57. A compact car can be made available for Mara’s Monday 08:00 Harbor City pickup if she changes the reservation before pickup. Every listed driving time is confirmed for both the vehicle on the reservation and the available replacement compact car. The itinerary records pickup at 08:00, arrival at Pine Gate at 10:00, and a rest interval from 10:00–10:30. Mara’s garden attendance is scheduled for 10:30–11:30, followed by arrival at Seaside Ferry at 13:00 and a ferry ride from 13:00–13:45. She is due at Island Lodge at 14:45; the Monday check-in interval is 15:00–20:00, and her room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara, Monday, Harbor City, the itinerary bindings, vehicle restriction, timing requirements, lodging facts, and unchanged questions. The evidence contains exactly two complete factual sentences: “The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42.” and “The rental catalog's canonical code for a compact car is HX-42.” The counterfactual changes only the catalog code to HX-77 and does not create contradictory duplicate measurements or embed an answer, label, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation at the Harbor City rental desk records a canonical vehicle class in the set {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42. The rental catalog's canonical code for a compact car is HX-42. If Mara changes the reservation before pickup, a compact car can be ready for her Monday 08:00 Harbor City pickup. Driving durations are confirmed for both the reserved vehicle and the available replacement. She is scheduled to reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, with the ferry ride lasting until 13:45, then at Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HX-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42.", "negative_right": "The rental catalog's canonical code for a compact car is HX-77.", "right": "The rental catalog's canonical code for a compact car is HX-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-044", "id": "fast-41-diverse-107-044-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation at the Harbor City rental desk records a canonical vehicle class in the set {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42. The rental catalog's canonical code for a compact car is HX-42. If Mara changes the reservation before pickup, a compact car can be ready for her Monday 08:00 Harbor City pickup. Driving durations are confirmed for both the reserved vehicle and the available replacement. She is scheduled to reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, with the ferry ride lasting until 13:45, then at Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mara, Monday, Harbor City, the itinerary bindings, vehicle restriction, timing requirements, lodging facts, and unchanged questions. The evidence contains exactly two complete factual sentences: “The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42.” and “The rental catalog's canonical code for a compact car is HX-42.” The counterfactual changes only the catalog code to HX-77 and does not create contradictory duplicate measurements or embed an answer, label, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation at the Harbor City rental desk records a canonical vehicle class in the set {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42. The rental catalog's canonical code for a compact car is HX-42. If Mara changes the reservation before pickup, a compact car can be ready for her Monday 08:00 Harbor City pickup. Driving durations are confirmed for both the reserved vehicle and the available replacement. She is scheduled to reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, with the ferry ride lasting until 13:45, then at Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HX-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42.", "negative_right": "The rental catalog's canonical code for a compact car is HX-77.", "right": "The rental catalog's canonical code for a compact car is HX-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-044", "id": "fast-41-diverse-107-044-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation at the Harbor City rental desk records a canonical vehicle class in the set {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HX-42. The rental catalog's canonical code for a compact car is HX-77. If Mara changes the reservation before pickup, a compact car can be ready for her Monday 08:00 Harbor City pickup. Driving durations are confirmed for both the reserved vehicle and the available replacement. She is scheduled to reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, with the ferry ride lasting until 13:45, then at Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and Mara’s room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the original questions verbatim and preserve Mara, Monday, the route, timing, lodging, and vehicle-scope bindings. The base evidence quotes are complete factual sentences: “The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.” and “The rental catalog's canonical code for a compact car is HC-42.” The counterfactual changes only the catalog code to HC-57, producing no contradictory duplicate measurement or assertion. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation covers a Harbor City pickup at 08:00, and the recorded vehicle class is one of SUV or compact car. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be supplied if Mara changes the reservation before pickup. Driving durations are confirmed for both the reserved vehicle and that replacement. The schedule places Mara at Pine Gate at 10:00, followed by a half-hour pause and a garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, takes a 45-minute ferry ride, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in runs from 15:00 to 20:00, and Mara’s room is confirmed for Monday. The plan records these requirements: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-053", "id": "fast-41-diverse-107-053-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation covers a Harbor City pickup at 08:00, and the recorded vehicle class is one of SUV or compact car. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be supplied if Mara changes the reservation before pickup. Driving durations are confirmed for both the reserved vehicle and that replacement. The schedule places Mara at Pine Gate at 10:00, followed by a half-hour pause and a garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, takes a 45-minute ferry ride, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in runs from 15:00 to 20:00, and Mara’s room is confirmed for Monday. The plan records these requirements: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the original questions verbatim and preserve Mara, Monday, the route, timing, lodging, and vehicle-scope bindings. The base evidence quotes are complete factual sentences: “The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.” and “The rental catalog's canonical code for a compact car is HC-42.” The counterfactual changes only the catalog code to HC-57, producing no contradictory duplicate measurement or assertion. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation covers a Harbor City pickup at 08:00, and the recorded vehicle class is one of SUV or compact car. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be supplied if Mara changes the reservation before pickup. Driving durations are confirmed for both the reserved vehicle and that replacement. The schedule places Mara at Pine Gate at 10:00, followed by a half-hour pause and a garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, takes a 45-minute ferry ride, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in runs from 15:00 to 20:00, and Mara’s room is confirmed for Monday. The plan records these requirements: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-053", "id": "fast-41-diverse-107-053-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation covers a Harbor City pickup at 08:00, and the recorded vehicle class is one of SUV or compact car. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-57. A compact car can be supplied if Mara changes the reservation before pickup. Driving durations are confirmed for both the reserved vehicle and that replacement. The schedule places Mara at Pine Gate at 10:00, followed by a half-hour pause and a garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, takes a 45-minute ferry ride, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in runs from 15:00 to 20:00, and Mara’s room is confirmed for Monday. The plan records these requirements: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy and scoring criteria in both full inputs. Entity, route, timing, lodging, and request-scope bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4.\" and \"The rental catalog's canonical code for a compact car is VX-4.\" The counterfactual changes only the compact-car code to VX-9 and creates no contradictory duplicate measurement or assertion. Neither context embeds an answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation is recorded for a vehicle class in the set {SUV, compact car}, with pickup in Harbor City at 08:00. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4. The rental catalog's canonical code for a compact car is VX-4. A compact car can be made available for that Monday 08:00 pickup if Mara changes the reservation before pickup. All listed driving times are confirmed for both the reserved vehicle and the available replacement compact car. She is scheduled to reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed 10:30–11:30 garden visit. She is due at Seaside Ferry at 13:00, will ride the ferry for 45 minutes, and will reach Island Lodge at 14:45. The lodging check-in interval is 15:00–20:00, and her Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is VX-4."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4.", "negative_right": "The rental catalog's canonical code for a compact car is VX-9.", "right": "The rental catalog's canonical code for a compact car is VX-4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-054", "id": "fast-41-diverse-107-054-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation is recorded for a vehicle class in the set {SUV, compact car}, with pickup in Harbor City at 08:00. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4. The rental catalog's canonical code for a compact car is VX-4. A compact car can be made available for that Monday 08:00 pickup if Mara changes the reservation before pickup. All listed driving times are confirmed for both the reserved vehicle and the available replacement compact car. She is scheduled to reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed 10:30–11:30 garden visit. She is due at Seaside Ferry at 13:00, will ride the ferry for 45 minutes, and will reach Island Lodge at 14:45. The lodging check-in interval is 15:00–20:00, and her Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy and scoring criteria in both full inputs. Entity, route, timing, lodging, and request-scope bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4.\" and \"The rental catalog's canonical code for a compact car is VX-4.\" The counterfactual changes only the compact-car code to VX-9 and creates no contradictory duplicate measurement or assertion. Neither context embeds an answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation is recorded for a vehicle class in the set {SUV, compact car}, with pickup in Harbor City at 08:00. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4. The rental catalog's canonical code for a compact car is VX-4. A compact car can be made available for that Monday 08:00 pickup if Mara changes the reservation before pickup. All listed driving times are confirmed for both the reserved vehicle and the available replacement compact car. She is scheduled to reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed 10:30–11:30 garden visit. She is due at Seaside Ferry at 13:00, will ride the ferry for 45 minutes, and will reach Island Lodge at 14:45. The lodging check-in interval is 15:00–20:00, and her Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is VX-4."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4.", "negative_right": "The rental catalog's canonical code for a compact car is VX-9.", "right": "The rental catalog's canonical code for a compact car is VX-4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-054", "id": "fast-41-diverse-107-054-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation is recorded for a vehicle class in the set {SUV, compact car}, with pickup in Harbor City at 08:00. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is VX-4. The rental catalog's canonical code for a compact car is VX-9. A compact car can be made available for that Monday 08:00 pickup if Mara changes the reservation before pickup. All listed driving times are confirmed for both the reserved vehicle and the available replacement compact car. She is scheduled to reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed 10:30–11:30 garden visit. She is due at Seaside Ferry at 13:00, will ride the ferry for 45 minutes, and will reach Island Lodge at 14:45. The lodging check-in interval is 15:00–20:00, and her Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and scoring criteria, and both contexts retain Mara’s Monday Harbor City pickup, itinerary, and lodging bindings. The evidence spans are complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42.\" and \"The rental catalog's canonical code for a compact car on Monday is HC-42.\" The counterfactual coherently changes the compact-car catalog code to HC-57 while leaving the reserved code at HC-42, which can represent the reserved SUV. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation for a Harbor City pickup at 08:00 lists a canonical vehicle class belonging to {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42. The rental catalog's canonical code for a compact car on Monday is HC-42. A compact car can be made available for the 08:00 pickup if Mara changes the reservation before pickup. Driving-time confirmations are recorded for both the reserved vehicle and the available replacement compact car. Her schedule places arrival at Pine Gate at 10:00, followed by a half-hour rest and a garden attendance from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, rides for 45 minutes, and reaches Island Lodge at 14:45. Island Lodge records Monday check-in as 15:00–20:00 and confirms Mara's room. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car on Monday is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car on Monday is HC-57.", "right": "The rental catalog's canonical code for a compact car on Monday is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-055", "id": "fast-41-diverse-107-055-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation for a Harbor City pickup at 08:00 lists a canonical vehicle class belonging to {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42. The rental catalog's canonical code for a compact car on Monday is HC-42. A compact car can be made available for the 08:00 pickup if Mara changes the reservation before pickup. Driving-time confirmations are recorded for both the reserved vehicle and the available replacement compact car. Her schedule places arrival at Pine Gate at 10:00, followed by a half-hour rest and a garden attendance from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, rides for 45 minutes, and reaches Island Lodge at 14:45. Island Lodge records Monday check-in as 15:00–20:00 and confirms Mara's room. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and scoring criteria, and both contexts retain Mara’s Monday Harbor City pickup, itinerary, and lodging bindings. The evidence spans are complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42.\" and \"The rental catalog's canonical code for a compact car on Monday is HC-42.\" The counterfactual coherently changes the compact-car catalog code to HC-57 while leaving the reserved code at HC-42, which can represent the reserved SUV. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday reservation for a Harbor City pickup at 08:00 lists a canonical vehicle class belonging to {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42. The rental catalog's canonical code for a compact car on Monday is HC-42. A compact car can be made available for the 08:00 pickup if Mara changes the reservation before pickup. Driving-time confirmations are recorded for both the reserved vehicle and the available replacement compact car. Her schedule places arrival at Pine Gate at 10:00, followed by a half-hour rest and a garden attendance from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, rides for 45 minutes, and reaches Island Lodge at 14:45. Island Lodge records Monday check-in as 15:00–20:00 and confirms Mara's room. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car on Monday is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car on Monday is HC-57.", "right": "The rental catalog's canonical code for a compact car on Monday is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-055", "id": "fast-41-diverse-107-055-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday reservation for a Harbor City pickup at 08:00 lists a canonical vehicle class belonging to {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation at 08:00 is HC-42. The rental catalog's canonical code for a compact car on Monday is HC-57. A compact car can be made available for the 08:00 pickup if Mara changes the reservation before pickup. Driving-time confirmations are recorded for both the reserved vehicle and the available replacement compact car. Her schedule places arrival at Pine Gate at 10:00, followed by a half-hour rest and a garden attendance from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, rides for 45 minutes, and reaches Island Lodge at 14:45. Island Lodge records Monday check-in as 15:00–20:00 and confirms Mara's room. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain the relevant ferry, timing, rest, visit, and lodging constraints. The Mara, Monday, Harbor City, 08:00 pickup, itinerary, and policy bindings remain unchanged in both contexts. The evidence consists of two complete factual sentences: \"At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42.\" and \"The rental catalog's canonical code for a compact car is HC-42.\" The counterfactual coherently changes only the catalog code to HC-57 without creating duplicate measurements or contradictory assertions. Neither context embeds an answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday Harbor City reservation records a canonical vehicle class belonging to the set {SUV, compact car}. At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car is available for Mara’s Monday 08:00 Harbor City pickup if she changes her reservation before pickup. Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation, and every listed driving time is confirmed for the available replacement compact car. Her schedule places pickup at 08:00, arrival at Pine Gate at 10:00, and rest there from 10:00–10:30, followed by garden attendance from 10:30–11:30. She reaches Seaside Ferry at 13:00, takes the 45-minute ferry, and arrives at Island Lodge at 14:45. Check-in runs from 15:00–20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42.", "negative_left": "At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-057", "id": "fast-41-diverse-107-057-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday Harbor City reservation records a canonical vehicle class belonging to the set {SUV, compact car}. At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car is available for Mara’s Monday 08:00 Harbor City pickup if she changes her reservation before pickup. Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation, and every listed driving time is confirmed for the available replacement compact car. Her schedule places pickup at 08:00, arrival at Pine Gate at 10:00, and rest there from 10:00–10:30, followed by garden attendance from 10:30–11:30. She reaches Seaside Ferry at 13:00, takes the 45-minute ferry, and arrives at Island Lodge at 14:45. Check-in runs from 15:00–20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain the relevant ferry, timing, rest, visit, and lodging constraints. The Mara, Monday, Harbor City, 08:00 pickup, itinerary, and policy bindings remain unchanged in both contexts. The evidence consists of two complete factual sentences: \"At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42.\" and \"The rental catalog's canonical code for a compact car is HC-42.\" The counterfactual coherently changes only the catalog code to HC-57 without creating duplicate measurements or contradictory assertions. Neither context embeds an answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday Harbor City reservation records a canonical vehicle class belonging to the set {SUV, compact car}. At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car is available for Mara’s Monday 08:00 Harbor City pickup if she changes her reservation before pickup. Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation, and every listed driving time is confirmed for the available replacement compact car. Her schedule places pickup at 08:00, arrival at Pine Gate at 10:00, and rest there from 10:00–10:30, followed by garden attendance from 10:30–11:30. She reaches Seaside Ferry at 13:00, takes the 45-minute ferry, and arrives at Island Lodge at 14:45. Check-in runs from 15:00–20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42.", "negative_left": "At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-057", "id": "fast-41-diverse-107-057-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday Harbor City reservation records a canonical vehicle class belonging to the set {SUV, compact car}. At 08:00 on Monday, Mara's Harbor City rental reservation records canonical vehicle-class code HC-42. The rental catalog's canonical code for a compact car is HC-57. A compact car is available for Mara’s Monday 08:00 Harbor City pickup if she changes her reservation before pickup. Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation, and every listed driving time is confirmed for the available replacement compact car. Her schedule places pickup at 08:00, arrival at Pine Gate at 10:00, and rest there from 10:00–10:30, followed by garden attendance from 10:30–11:30. She reaches Seaside Ferry at 13:00, takes the 45-minute ferry, and arrives at Island Lodge at 14:45. Check-in runs from 15:00–20:00, and her room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing policy, scope, and Monday Mara itinerary bindings. The two evidence spans are complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.\" and \"The rental catalog's canonical code for a compact car is HC-42.\" The counterfactual changes only the catalog code to HC-57 and introduces no contradictory duplicate assertion. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"A Monday booking record for Mara places her Harbor City vehicle pickup at 08:00 and classifies the reserved vehicle within the set {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. If Mara changes the reservation before pickup, a compact car is available for the same Monday 08:00 Harbor City collection. Confirmed travel records show arrival at Pine Gate at 10:00, followed by a required half-hour pause, then the fixed garden attendance from 10:30 to 11:30. She is scheduled to reach Seaside Ferry at 13:00, take the 45-minute ferry, and arrive at Island Lodge at 14:45. All listed driving times are confirmed for both the reserved vehicle and the available replacement compact car. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and Mara's room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-064", "id": "fast-41-diverse-107-064-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "A Monday booking record for Mara places her Harbor City vehicle pickup at 08:00 and classifies the reserved vehicle within the set {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. If Mara changes the reservation before pickup, a compact car is available for the same Monday 08:00 Harbor City collection. Confirmed travel records show arrival at Pine Gate at 10:00, followed by a required half-hour pause, then the fixed garden attendance from 10:30 to 11:30. She is scheduled to reach Seaside Ferry at 13:00, take the 45-minute ferry, and arrive at Island Lodge at 14:45. All listed driving times are confirmed for both the reserved vehicle and the available replacement compact car. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and Mara's room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing policy, scope, and Monday Mara itinerary bindings. The two evidence spans are complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.\" and \"The rental catalog's canonical code for a compact car is HC-42.\" The counterfactual changes only the catalog code to HC-57 and introduces no contradictory duplicate assertion. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"A Monday booking record for Mara places her Harbor City vehicle pickup at 08:00 and classifies the reserved vehicle within the set {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. If Mara changes the reservation before pickup, a compact car is available for the same Monday 08:00 Harbor City collection. Confirmed travel records show arrival at Pine Gate at 10:00, followed by a required half-hour pause, then the fixed garden attendance from 10:30 to 11:30. She is scheduled to reach Seaside Ferry at 13:00, take the 45-minute ferry, and arrive at Island Lodge at 14:45. All listed driving times are confirmed for both the reserved vehicle and the available replacement compact car. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and Mara's room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-064", "id": "fast-41-diverse-107-064-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "A Monday booking record for Mara places her Harbor City vehicle pickup at 08:00 and classifies the reserved vehicle within the set {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-57. If Mara changes the reservation before pickup, a compact car is available for the same Monday 08:00 Harbor City collection. Confirmed travel records show arrival at Pine Gate at 10:00, followed by a required half-hour pause, then the fixed garden attendance from 10:30 to 11:30. She is scheduled to reach Seaside Ferry at 13:00, take the 45-minute ferry, and arrive at Island Lodge at 14:45. All listed driving times are confirmed for both the reserved vehicle and the available replacement compact car. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and Mara's room is confirmed for Monday. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria. Entity, route, date, and time bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14.\" and \"The rental catalog's canonical code for a compact car on Monday is VX-14.\" The counterfactual's changed compact-car code does not contradict any unchanged assertion. Neither context states a gold score, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Monday’s Harbor City reservation for Mara has a canonical vehicle class that is a member of the set {SUV, compact car}. At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14. The rental catalog's canonical code for a compact car on Monday is VX-14. A compact car is available for Mara’s Monday 08:00 Harbor City pickup if she changes her reservation before pickup. Every listed driving time is confirmed for both the reserved vehicle and the available replacement compact car. Mara plans to collect the vehicle at 08:00, arrive at Pine Gate at 10:00, and rest there from 10:00–10:30 before the garden visit from 10:30–11:30. She is scheduled to arrive at Seaside Ferry at 13:00, take the ferry from 13:00–13:45, and reach Island Lodge at 14:45. take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars. Island Lodge’s Monday check-in interval for Mara’s stay is 15:00–20:00, and her room is confirmed for Monday.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14."}, {"path": [], "text": "The rental catalog's canonical code for a compact car on Monday is VX-14."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14.", "negative_left": "At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14.", "negative_right": "The rental catalog's canonical code for a compact car on Monday is VX-19.", "right": "The rental catalog's canonical code for a compact car on Monday is VX-14."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-072", "id": "fast-41-diverse-107-072-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Monday’s Harbor City reservation for Mara has a canonical vehicle class that is a member of the set {SUV, compact car}. At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14. The rental catalog's canonical code for a compact car on Monday is VX-14. A compact car is available for Mara’s Monday 08:00 Harbor City pickup if she changes her reservation before pickup. Every listed driving time is confirmed for both the reserved vehicle and the available replacement compact car. Mara plans to collect the vehicle at 08:00, arrive at Pine Gate at 10:00, and rest there from 10:00–10:30 before the garden visit from 10:30–11:30. She is scheduled to arrive at Seaside Ferry at 13:00, take the ferry from 13:00–13:45, and reach Island Lodge at 14:45. take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars. Island Lodge’s Monday check-in interval for Mara’s stay is 15:00–20:00, and her room is confirmed for Monday."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria. Entity, route, date, and time bindings remain unchanged. The evidence contains exactly two complete factual sentences: \"At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14.\" and \"The rental catalog's canonical code for a compact car on Monday is VX-14.\" The counterfactual's changed compact-car code does not contradict any unchanged assertion. Neither context states a gold score, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Monday’s Harbor City reservation for Mara has a canonical vehicle class that is a member of the set {SUV, compact car}. At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14. The rental catalog's canonical code for a compact car on Monday is VX-14. A compact car is available for Mara’s Monday 08:00 Harbor City pickup if she changes her reservation before pickup. Every listed driving time is confirmed for both the reserved vehicle and the available replacement compact car. Mara plans to collect the vehicle at 08:00, arrive at Pine Gate at 10:00, and rest there from 10:00–10:30 before the garden visit from 10:30–11:30. She is scheduled to arrive at Seaside Ferry at 13:00, take the ferry from 13:00–13:45, and reach Island Lodge at 14:45. take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars. Island Lodge’s Monday check-in interval for Mara’s stay is 15:00–20:00, and her room is confirmed for Monday.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14."}, {"path": [], "text": "The rental catalog's canonical code for a compact car on Monday is VX-14."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14.", "negative_left": "At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14.", "negative_right": "The rental catalog's canonical code for a compact car on Monday is VX-19.", "right": "The rental catalog's canonical code for a compact car on Monday is VX-14."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-072", "id": "fast-41-diverse-107-072-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Monday’s Harbor City reservation for Mara has a canonical vehicle class that is a member of the set {SUV, compact car}. At 08:00 on Monday, Mara's Harbor City rental reservation record HCR-482 has canonical vehicle-class code VX-14. The rental catalog's canonical code for a compact car on Monday is VX-19. A compact car is available for Mara’s Monday 08:00 Harbor City pickup if she changes her reservation before pickup. Every listed driving time is confirmed for both the reserved vehicle and the available replacement compact car. Mara plans to collect the vehicle at 08:00, arrive at Pine Gate at 10:00, and rest there from 10:00–10:30 before the garden visit from 10:30–11:30. She is scheduled to arrive at Seaside Ferry at 13:00, take the ferry from 13:00–13:45, and reach Island Lodge at 14:45. take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars. Island Lodge’s Monday check-in interval for Mara’s stay is 15:00–20:00, and her room is confirmed for Monday."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, scope, and bindings, and both contexts repeat the applicable policy facts. The evidence contains exactly two complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.\" and \"The rental catalog's canonical code for a compact car is HC-42.\" The counterfactual coherently changes only the compact-car catalog code to HC-57, eliminating the code match while retaining the SUV reservation and available compact replacement. Neither context states a gold answer, answer code, rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara's Monday rental file for Harbor City lists a vehicle class drawn from {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for her 08:00 Harbor City pickup if she changes the reservation before pickup, and the listed driving times are confirmed for both the reserved vehicle and that replacement. She plans to collect the vehicle at 08:00, reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00 for 13:00 boarding, rides for 45 minutes, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and her Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-075", "id": "fast-41-diverse-107-075-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara's Monday rental file for Harbor City lists a vehicle class drawn from {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for her 08:00 Harbor City pickup if she changes the reservation before pickup, and the listed driving times are confirmed for both the reserved vehicle and that replacement. She plans to collect the vehicle at 08:00, reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00 for 13:00 boarding, rides for 45 minutes, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and her Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, scope, and bindings, and both contexts repeat the applicable policy facts. The evidence contains exactly two complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.\" and \"The rental catalog's canonical code for a compact car is HC-42.\" The counterfactual coherently changes only the compact-car catalog code to HC-57, eliminating the code match while retaining the SUV reservation and available compact replacement. Neither context states a gold answer, answer code, rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara's Monday rental file for Harbor City lists a vehicle class drawn from {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-42. A compact car can be made available for her 08:00 Harbor City pickup if she changes the reservation before pickup, and the listed driving times are confirmed for both the reserved vehicle and that replacement. She plans to collect the vehicle at 08:00, reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00 for 13:00 boarding, rides for 45 minutes, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and her Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42."}, {"path": [], "text": "The rental catalog's canonical code for a compact car is HC-42."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42.", "negative_right": "The rental catalog's canonical code for a compact car is HC-57.", "right": "The rental catalog's canonical code for a compact car is HC-42."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-075", "id": "fast-41-diverse-107-075-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara's Monday rental file for Harbor City lists a vehicle class drawn from {SUV, compact car}. The canonical vehicle-class code recorded on Mara's Monday Harbor City rental reservation is HC-42. The rental catalog's canonical code for a compact car is HC-57. A compact car can be made available for her 08:00 Harbor City pickup if she changes the reservation before pickup, and the listed driving times are confirmed for both the reserved vehicle and that replacement. She plans to collect the vehicle at 08:00, reach Pine Gate at 10:00, take the required 30-minute rest, and attend the fixed garden visit from 10:30 to 11:30. She is due at Seaside Ferry at 13:00 for 13:00 boarding, rides for 45 minutes, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Check-in is 15:00–20:00, and her Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the unchanged questions and governing policy. All question entity, path, date, and time bindings remain unchanged. The two evidence spans are complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14.\" and \"The rental catalog's canonical code for a compact car on Monday is AX-14.\" The counterfactual changes only the catalog code to AX-19, without creating contradictory duplicate measurements. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday travel file concerns reservation R-482, with pickup scheduled in Harbor City at 08:00. The reservation’s canonical vehicle class is one of {SUV, compact car}. The rental coordinator says a compact car can be supplied for that pickup if Mara changes the reservation beforehand, and driving durations are confirmed for both the reserved vehicle and the possible replacement. The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14. The rental catalog's canonical code for a compact car on Monday is AX-14. The route record places Mara at Pine Gate at 10:00, followed by a 10:00–10:30 pause and a garden attendance window from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, takes a 13:00–13:45 ferry ride, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Island Lodge offers check-in from 15:00–20:00, and Mara’s Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14."}, {"path": [], "text": "The rental catalog's canonical code for a compact car on Monday is AX-14."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14.", "negative_right": "The rental catalog's canonical code for a compact car on Monday is AX-19.", "right": "The rental catalog's canonical code for a compact car on Monday is AX-14."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-078", "id": "fast-41-diverse-107-078-base", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday travel file concerns reservation R-482, with pickup scheduled in Harbor City at 08:00. The reservation’s canonical vehicle class is one of {SUV, compact car}. The rental coordinator says a compact car can be supplied for that pickup if Mara changes the reservation beforehand, and driving durations are confirmed for both the reserved vehicle and the possible replacement. The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14. The rental catalog's canonical code for a compact car on Monday is AX-14. The route record places Mara at Pine Gate at 10:00, followed by a 10:00–10:30 pause and a garden attendance window from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, takes a 13:00–13:45 ferry ride, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Island Lodge offers check-in from 15:00–20:00, and Mara’s Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the unchanged questions and governing policy. All question entity, path, date, and time bindings remain unchanged. The two evidence spans are complete factual sentences: \"The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14.\" and \"The rental catalog's canonical code for a compact car on Monday is AX-14.\" The counterfactual changes only the catalog code to AX-19, without creating contradictory duplicate measurements. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified driving-time atoms remain atomic because they apply one relation over the listed legs. The focus atom is the factual vehicle-class code, not a policy conclusion. Base and counter assignments differ only on that focus and are jointly realizable: the reservation may be compact or non-compact while the other schedule and availability facts remain fixed. Policy evidence correctly preserves the state-originated required rest, fixed visit, boarding and check-in windows, and ferry vehicle restriction. Criteria and scoring instructions from the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given the exhaustive class restriction in a1, refuting compact status in a2 makes the reservation an SUV. The preserved ferry restriction rejects SUVs, while a3 and a5 establish an available compact correction with confirmed times. The remaining atoms establish that the rest, fixed visit, ferry boarding, travel, and lodging constraints work, so exactly one essential vehicle-class correction is required, entailing level 2.", "rule_index": 0, "sound": true}, {"reason": "Supported compact status entails that the booked vehicle is accepted by the ferry. The timing atoms establish compliance with the required rest, fixed visit, boarding, travel, and lodging requirements. Arrival at 14:45 before the 15:00 check-in opening creates only a 15-minute wait and requires no major itinerary change, excluding level 4 convenience while satisfying level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The canonical vehicle class on Mara's Monday Harbor City rental reservation is a member of the set {SUV, compact car}."}, {"id": "a2", "statement": "The canonical vehicle-class code on Mara's Monday Harbor City rental reservation equals the rental catalog's canonical code for a compact car."}, {"id": "a3", "statement": "A compact car is available for Mara's Monday 08:00 Harbor City pickup if she changes her reservation before pickup."}, {"id": "a4", "statement": "Every listed driving time is confirmed for the vehicle on Mara's Monday rental reservation."}, {"id": "a5", "statement": "Every listed driving time is confirmed for the available replacement compact car."}, {"id": "a6", "statement": "Mara's planned Monday vehicle pickup in Harbor City is at 08:00."}, {"id": "a7", "statement": "Mara's planned Monday arrival at Pine Gate is at 10:00."}, {"id": "a8", "statement": "Mara's planned Monday rest at Pine Gate occupies the interval 10:00–10:30."}, {"id": "a9", "statement": "Mara's planned Monday garden attendance occupies the interval 10:30–11:30."}, {"id": "a10", "statement": "Mara's planned Monday arrival at Seaside Ferry is at 13:00."}, {"id": "a11", "statement": "Mara's planned Monday ferry ride occupies the interval 13:00–13:45."}, {"id": "a12", "statement": "Mara's planned Monday arrival at Island Lodge is at 14:45."}, {"id": "a13", "statement": "Island Lodge's Monday check-in interval for Mara's stay is 15:00–20:00."}, {"id": "a14", "statement": "Mara's room at Island Lodge is confirmed for Monday."}], "base_state_json": "\"Mara’s Monday travel file concerns reservation R-482, with pickup scheduled in Harbor City at 08:00. The reservation’s canonical vehicle class is one of {SUV, compact car}. The rental coordinator says a compact car can be supplied for that pickup if Mara changes the reservation beforehand, and driving durations are confirmed for both the reserved vehicle and the possible replacement. The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14. The rental catalog's canonical code for a compact car on Monday is AX-14. The route record places Mara at Pine Gate at 10:00, followed by a 10:00–10:30 pause and a garden attendance window from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, takes a 13:00–13:45 ferry ride, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Island Lodge offers check-in from 15:00–20:00, and Mara’s Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14."}, {"path": [], "text": "The rental catalog's canonical code for a compact car on Monday is AX-14."}], "policy_evidence": [{"path": [], "text": "take the required 30-minute rest"}, {"path": [], "text": "attend the fixed 10:30–11:30 garden visit"}, {"path": [], "text": "13:00 boarding"}, {"path": [], "text": "Check-in is 15:00–20:00"}, {"path": [], "text": "the ferry does not accept SUVs; it accepts compact cars"}], "rules": [{"justification": "Because the reserved class is either SUV or compact and is explicitly not compact, it is an SUV, which the ferry rejects. One available pre-pickup change to the accepted compact class resolves that sole essential conflict; the compact's driving times, all fixed events, the required rest, and lodging then work, apart from waiting 15 minutes for check-in.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "The booked vehicle is a compact car accepted by the ferry, and its confirmed driving times satisfy the required rest, fixed visit, boarding, and lodging schedule. The only inconvenience is waiting from 14:45 until check-in opens at 15:00, which changes no vehicle class, route, fixed visit, travel date, or lodging.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14.", "negative_left": "The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14.", "negative_right": "The rental catalog's canonical code for a compact car on Monday is AX-19.", "right": "The rental catalog's canonical code for a compact car on Monday is AX-14."}, "verifier_independent_model": false}, "family": "fast-41-diverse-107-078", "id": "fast-41-diverse-107-078-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Impossible: The itinerary cannot be completed, and no supplied vehicle, schedule, route, or accommodation change can resolve the conflict.", "1 — Severely infeasible: Two or more major unresolved conflicts affect fixed windows, transport, required rests, or lodging, even though substantial replanning might make a different trip possible.", "2 — Feasible only after one major correction: A single essential vehicle, route, date, or accommodation change is required; once made, all supplied timings and constraints work.", "3 — Feasible with minor inconvenience: The itinerary works as planned except for a small adjustment that does not change the vehicle class, route, fixed visit, travel date, or lodging.", "4 — Fully feasible and convenient: Every leg, fixed window, rest, vehicle rule, and lodging requirement works exactly as booked, with no changes needed."], "instructions": "Rate the plan’s overall feasibility and convenience using only the supplied facts. Account for visit windows, driving and rest timing, lodging, and the vehicle restriction. Select exactly one level.", "type": "score"}}, "state": "Mara’s Monday travel file concerns reservation R-482, with pickup scheduled in Harbor City at 08:00. The reservation’s canonical vehicle class is one of {SUV, compact car}. The rental coordinator says a compact car can be supplied for that pickup if Mara changes the reservation beforehand, and driving durations are confirmed for both the reserved vehicle and the possible replacement. The canonical vehicle-class code recorded on Mara's Monday 08:00 Harbor City rental reservation R-482 is AX-14. The rental catalog's canonical code for a compact car on Monday is AX-19. The route record places Mara at Pine Gate at 10:00, followed by a 10:00–10:30 pause and a garden attendance window from 10:30 to 11:30. She is due at Seaside Ferry at 13:00, takes a 13:00–13:45 ferry ride, and reaches Island Lodge at 14:45. The ferry does not accept SUVs; it accepts compact cars. Island Lodge offers check-in from 15:00–20:00, and Mara’s Monday room is confirmed. Policy wording: take the required 30-minute rest; attend the fixed 10:30–11:30 garden visit; 13:00 boarding; Check-in is 15:00–20:00; the ferry does not accept SUVs; it accepts compact cars."}, "method": "c2d", "provenance": {"source_id": "diverse-107", "source_is_synthetic": true, "source_sha256": "e5484a83549c17f43e4873a2fc93885f89b0549093ada07f3235ea3e9608152f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and Friday Valley bindings. The required evidence quotes are complete factual sentences: “The Friday Valley plan records its winery arrival at 12:00 on Friday.” and “The booked Friday winery visit has a fixed start time of 12:15.” The counterfactual changes the arrival to 12:30 while consistently preserving the stated earliest-arrival constraint. Neither context contains a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":{\"observations\":[\"The traveler’s Friday itinerary has no confirmed AWD vehicle.\",\"An automatic compact rental vehicle is confirmed for Friday, and it is permitted on every road in the Friday Valley plan.\",\"The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window.\",\"The Friday Valley plan records its winery arrival at 12:00 on Friday.\",\"The booked Friday winery visit has a fixed start time of 12:15.\",\"Dispatch records show that the earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time.\",\"The lodging arrival margin before the 15:15 check-in closing time is between 0 and 20 minutes inclusive.\",\"Every continuous driving segment in the Friday Valley plan is at most 2 hours.\",\"A room is confirmed for the required Friday night, and a parking space is confirmed for the Friday Valley plan.\"],\"policy\":[\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\",\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"]}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["context", "observations", "3"], "text": "The Friday Valley plan records its winery arrival at 12:00 on Friday."}, {"path": ["context", "observations", "4"], "text": "The booked Friday winery visit has a fixed start time of 12:15."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The Friday Valley plan records its winery arrival at 12:00 on Friday.", "negative_left": "The Friday Valley plan records its winery arrival at 12:30 on Friday.", "negative_right": "The booked Friday winery visit has a fixed start time of 12:15.", "right": "The booked Friday winery visit has a fixed start time of 12:15."}, "verifier_independent_model": false}, "family": "fast-41-diverse-108-025", "id": "fast-41-diverse-108-025-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": {"observations": ["The traveler’s Friday itinerary has no confirmed AWD vehicle.", "An automatic compact rental vehicle is confirmed for Friday, and it is permitted on every road in the Friday Valley plan.", "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window.", "The Friday Valley plan records its winery arrival at 12:00 on Friday.", "The booked Friday winery visit has a fixed start time of 12:15.", "Dispatch records show that the earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time.", "The lodging arrival margin before the 15:15 check-in closing time is between 0 and 20 minutes inclusive.", "Every continuous driving segment in the Friday Valley plan is at most 2 hours.", "A room is confirmed for the required Friday night, and a parking space is confirmed for the Friday Valley plan."], "policy": ["An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock.", "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."]}}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and Friday Valley bindings. The required evidence quotes are complete factual sentences: “The Friday Valley plan records its winery arrival at 12:00 on Friday.” and “The booked Friday winery visit has a fixed start time of 12:15.” The counterfactual changes the arrival to 12:30 while consistently preserving the stated earliest-arrival constraint. Neither context contains a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":{\"observations\":[\"The traveler’s Friday itinerary has no confirmed AWD vehicle.\",\"An automatic compact rental vehicle is confirmed for Friday, and it is permitted on every road in the Friday Valley plan.\",\"The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window.\",\"The Friday Valley plan records its winery arrival at 12:00 on Friday.\",\"The booked Friday winery visit has a fixed start time of 12:15.\",\"Dispatch records show that the earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time.\",\"The lodging arrival margin before the 15:15 check-in closing time is between 0 and 20 minutes inclusive.\",\"Every continuous driving segment in the Friday Valley plan is at most 2 hours.\",\"A room is confirmed for the required Friday night, and a parking space is confirmed for the Friday Valley plan.\"],\"policy\":[\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\",\"Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels.\"]}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["context", "observations", "3"], "text": "The Friday Valley plan records its winery arrival at 12:00 on Friday."}, {"path": ["context", "observations", "4"], "text": "The booked Friday winery visit has a fixed start time of 12:15."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The Friday Valley plan records its winery arrival at 12:00 on Friday.", "negative_left": "The Friday Valley plan records its winery arrival at 12:30 on Friday.", "negative_right": "The booked Friday winery visit has a fixed start time of 12:15.", "right": "The booked Friday winery visit has a fixed start time of 12:15."}, "verifier_independent_model": false}, "family": "fast-41-diverse-108-025", "id": "fast-41-diverse-108-025-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": {"observations": ["The traveler’s Friday itinerary has no confirmed AWD vehicle.", "An automatic compact rental vehicle is confirmed for Friday, and it is permitted on every road in the Friday Valley plan.", "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window.", "The Friday Valley plan records its winery arrival at 12:30 on Friday.", "The booked Friday winery visit has a fixed start time of 12:15.", "Dispatch records show that the earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time.", "The lodging arrival margin before the 15:15 check-in closing time is between 0 and 20 minutes inclusive.", "Every continuous driving segment in the Friday Valley plan is at most 2 hours.", "A room is confirmed for the required Friday night, and a parking space is confirmed for the Friday Valley plan."], "policy": ["An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock.", "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."]}}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, instructions, scope, and bindings, while both contexts restate the same policy without exceptions. The evidence contains exactly two complete factual sentences: \"On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers.\" and \"On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers.\" The counterfactual coherently changes the van capacity to two without contradictory duplicate measurements or assertions. Neither context embeds an answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers. The van’s records document room for four suitcases and Elena’s rigid wheelchair, along with a ramp and wheelchair securement. Its pickup is scheduled for 18:35, and Elena requested hotel pickup by 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, after Elena’s requested deadline. The airport assistance attendant provides help only from baggage claim to the curb, rather than the complete journey to the hotel. Suitcase color and hotel membership are recorded but are not routing requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is two travelers.", "right": "On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-006", "id": "fast-41-diverse-109-006-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers. The van’s records document room for four suitcases and Elena’s rigid wheelchair, along with a ramp and wheelchair securement. Its pickup is scheduled for 18:35, and Elena requested hotel pickup by 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, after Elena’s requested deadline. The airport assistance attendant provides help only from baggage claim to the curb, rather than the complete journey to the hotel. Suitcase color and hotel membership are recorded but are not routing requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, instructions, scope, and bindings, while both contexts restate the same policy without exceptions. The evidence contains exactly two complete factual sentences: \"On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers.\" and \"On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers.\" The counterfactual coherently changes the van capacity to two without contradictory duplicate measurements or assertions. Neither context embeds an answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers. The van’s records document room for four suitcases and Elena’s rigid wheelchair, along with a ramp and wheelchair securement. Its pickup is scheduled for 18:35, and Elena requested hotel pickup by 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, after Elena’s requested deadline. The airport assistance attendant provides help only from baggage claim to the curb, rather than the complete journey to the hotel. Suitcase color and hotel membership are recorded but are not routing requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is two travelers.", "right": "On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-006", "id": "fast-41-diverse-109-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 14 June 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 14 June 2026, the 18:35 accessible van’s documented passenger capacity is two travelers. The van’s records document room for four suitcases and Elena’s rigid wheelchair, along with a ramp and wheelchair securement. Its pickup is scheduled for 18:35, and Elena requested hotel pickup by 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, after Elena’s requested deadline. The airport assistance attendant provides help only from baggage claim to the curb, rather than the complete journey to the hotel. Suitcase color and hotel membership are recorded but are not routing requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged original questions and governing policy. The original location, date, passenger, route, deadline, and option bindings remain unchanged. The evidence consists of two complete factual sentences: \"Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers.\" and \"The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers.\" The counterfactual changes only the van’s capacity to two travelers, which is coherent with the unchanged party size and other observations. Neither context contains a gold answer, answer code, rationale, proposition identifier, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 17 June 2026 at Northstar Airport, Elena is traveling to her hotel with companions and four suitcases, using a rigid wheelchair. Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers. The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers. Its record also confirms room for four suitcases, a rigid-wheelchair bay, a ramp, and wheelchair securement, and its pickup is scheduled before Elena’s requested 18:40 deadline. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant handles only the airport-to-curb portion and does not provide the complete hotel transfer. Suitcase color and hotel membership are irrelevant. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers."}, {"path": [], "text": "The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers.", "negative_left": "Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers.", "negative_right": "The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is two travelers.", "right": "The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-014", "id": "fast-41-diverse-109-014-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 17 June 2026 at Northstar Airport, Elena is traveling to her hotel with companions and four suitcases, using a rigid wheelchair. Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers. The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers. Its record also confirms room for four suitcases, a rigid-wheelchair bay, a ramp, and wheelchair securement, and its pickup is scheduled before Elena’s requested 18:40 deadline. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant handles only the airport-to-curb portion and does not provide the complete hotel transfer. Suitcase color and hotel membership are irrelevant. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged original questions and governing policy. The original location, date, passenger, route, deadline, and option bindings remain unchanged. The evidence consists of two complete factual sentences: \"Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers.\" and \"The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers.\" The counterfactual changes only the van’s capacity to two travelers, which is coherent with the unchanged party size and other observations. Neither context contains a gold answer, answer code, rationale, proposition identifier, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 17 June 2026 at Northstar Airport, Elena is traveling to her hotel with companions and four suitcases, using a rigid wheelchair. Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers. The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers. Its record also confirms room for four suitcases, a rigid-wheelchair bay, a ramp, and wheelchair securement, and its pickup is scheduled before Elena’s requested 18:40 deadline. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant handles only the airport-to-curb portion and does not provide the complete hotel transfer. Suitcase color and hotel membership are irrelevant. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers."}, {"path": [], "text": "The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers.", "negative_left": "Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers.", "negative_right": "The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is two travelers.", "right": "The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-014", "id": "fast-41-diverse-109-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 17 June 2026 at Northstar Airport, Elena is traveling to her hotel with companions and four suitcases, using a rigid wheelchair. Elena’s airport-to-hotel party on 17 June 2026 comprises exactly three travelers. The 18:35 accessible van’s documented passenger capacity for 17 June 2026 is two travelers. Its record also confirms room for four suitcases, a rigid-wheelchair bay, a ramp, and wheelchair securement, and its pickup is scheduled before Elena’s requested 18:40 deadline. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant handles only the airport-to-curb portion and does not provide the complete hotel transfer. Suitcase color and hotel membership are irrelevant. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and all choice criteria. Entity, route, deadline, and transfer bindings remain unchanged. The evidence spans are complete factual sentences: \"On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers.\" \"The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026.\" The counterfactual changes only the van’s capacity to two travelers, which coherently conflicts with the unchanged three-traveler party without duplicate measurements. Neither context states a gold answer, label code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Northstar Airport’s 17 September 2026 dispatch record concerns Elena, her rigid wheelchair, four suitcases, and a requested hotel pickup by 18:40. On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers. The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026. Its records also list capacity for four suitcases, wheelchair capacity, a ramp, and wheelchair securement, with pickup scheduled at 18:35. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after Elena’s deadline. The airport assistance attendant serves only from baggage claim to the curb, rather than the complete transfer to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers."}, {"path": [], "text": "The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers.", "negative_left": "On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers.", "negative_right": "The 18:35 accessible van’s documented passenger capacity was two travelers for Elena’s airport-to-hotel transfer on 17 September 2026.", "right": "The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-017", "id": "fast-41-diverse-109-017-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Northstar Airport’s 17 September 2026 dispatch record concerns Elena, her rigid wheelchair, four suitcases, and a requested hotel pickup by 18:40. On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers. The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026. Its records also list capacity for four suitcases, wheelchair capacity, a ramp, and wheelchair securement, with pickup scheduled at 18:35. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after Elena’s deadline. The airport assistance attendant serves only from baggage claim to the curb, rather than the complete transfer to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and all choice criteria. Entity, route, deadline, and transfer bindings remain unchanged. The evidence spans are complete factual sentences: \"On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers.\" \"The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026.\" The counterfactual changes only the van’s capacity to two travelers, which coherently conflicts with the unchanged three-traveler party without duplicate measurements. Neither context states a gold answer, label code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Northstar Airport’s 17 September 2026 dispatch record concerns Elena, her rigid wheelchair, four suitcases, and a requested hotel pickup by 18:40. On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers. The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026. Its records also list capacity for four suitcases, wheelchair capacity, a ramp, and wheelchair securement, with pickup scheduled at 18:35. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after Elena’s deadline. The airport assistance attendant serves only from baggage claim to the curb, rather than the complete transfer to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers."}, {"path": [], "text": "The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers.", "negative_left": "On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers.", "negative_right": "The 18:35 accessible van’s documented passenger capacity was two travelers for Elena’s airport-to-hotel transfer on 17 September 2026.", "right": "The 18:35 accessible van’s documented passenger capacity was four travelers for Elena’s airport-to-hotel transfer on 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-017", "id": "fast-41-diverse-109-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Northstar Airport’s 17 September 2026 dispatch record concerns Elena, her rigid wheelchair, four suitcases, and a requested hotel pickup by 18:40. On 17 September 2026, Elena’s entire airport-to-hotel party consisted of three travelers. The 18:35 accessible van’s documented passenger capacity was two travelers for Elena’s airport-to-hotel transfer on 17 September 2026. Its records also list capacity for four suitcases, wheelchair capacity, a ramp, and wheelchair securement, with pickup scheduled at 18:35. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after Elena’s deadline. The airport assistance attendant serves only from baggage claim to the curb, rather than the complete transfer to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, criteria, scope, entities, dates, and paths. The base context preserves Elena’s airport-to-hotel bindings and the counterfactual changes only the party-size observation. The two evidence spans are complete factual sentences: \"For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers.\" and \"For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers.\" The counterfactual’s four-person party is consistent with its other facts and creates no contradictory duplicate measurement. Neither context embeds an answer, label, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Northstar Airport transfer case note, 18 September 2026: For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers. For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers. The same van's documented suitcase capacity is at least four, and its records specify capacity for Elena's rigid wheelchair, a ramp, and wheelchair securement. Its pickup is scheduled for 18:35, no later than Elena's requested 18:40 hotel deadline. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant's service ends at the curb and does not provide the complete journey to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena cannot transfer from her rigid wheelchair or use steps; her companions and four suitcases travel with her.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers."}, {"path": [], "text": "For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers.", "negative_left": "For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers.", "negative_right": "For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of four travelers.", "right": "For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-025", "id": "fast-41-diverse-109-025-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Northstar Airport transfer case note, 18 September 2026: For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers. For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers. The same van's documented suitcase capacity is at least four, and its records specify capacity for Elena's rigid wheelchair, a ramp, and wheelchair securement. Its pickup is scheduled for 18:35, no later than Elena's requested 18:40 hotel deadline. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant's service ends at the curb and does not provide the complete journey to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena cannot transfer from her rigid wheelchair or use steps; her companions and four suitcases travel with her."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, criteria, scope, entities, dates, and paths. The base context preserves Elena’s airport-to-hotel bindings and the counterfactual changes only the party-size observation. The two evidence spans are complete factual sentences: \"For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers.\" and \"For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers.\" The counterfactual’s four-person party is consistent with its other facts and creates no contradictory duplicate measurement. Neither context embeds an answer, label, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Northstar Airport transfer case note, 18 September 2026: For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers. For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers. The same van's documented suitcase capacity is at least four, and its records specify capacity for Elena's rigid wheelchair, a ramp, and wheelchair securement. Its pickup is scheduled for 18:35, no later than Elena's requested 18:40 hotel deadline. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant's service ends at the curb and does not provide the complete journey to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena cannot transfer from her rigid wheelchair or use steps; her companions and four suitcases travel with her.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers."}, {"path": [], "text": "For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers.", "negative_left": "For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers.", "negative_right": "For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of four travelers.", "right": "For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-025", "id": "fast-41-diverse-109-025-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Northstar Airport transfer case note, 18 September 2026: For Elena's airport-to-hotel transfer on 18 September 2026, the documented passenger capacity of the 18:35 accessible van is exactly three travelers. For Elena's airport-to-hotel transfer on 18 September 2026, her entire party consists of four travelers. The same van's documented suitcase capacity is at least four, and its records specify capacity for Elena's rigid wheelchair, a ramp, and wheelchair securement. Its pickup is scheduled for 18:35, no later than Elena's requested 18:40 hotel deadline. The 18:25 standard sedan has no documented ramp. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant's service ends at the curb and does not provide the complete journey to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena cannot transfer from her rigid wheelchair or use steps; her companions and four suitcases travel with her."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy and criteria, and both contexts repeat the applicable policy without exceptions. Elena, the airport-to-hotel scope, 3 October 2026 date, and 18:40 deadline remain bound consistently. The evidence contains exactly two complete factual sentences: \"On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers.\" and \"The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers.\" The counterfactual coherently changes the van capacity to two travelers without creating duplicate measurements or contradictions. Neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers. The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers. Its transfer manifest lists room for four suitcases and a rigid wheelchair, and documents both a boarding ramp and wheelchair securement. The van is scheduled for 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond Elena’s requested deadline. The airport assistance attendant’s service ends at the curb and does not continue to the hotel. Suitcase color and hotel membership appear in the intake record but are irrelevant to routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers."}, {"path": [], "text": "The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers.", "negative_left": "On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers.", "negative_right": "The 18:35 accessible van’s capacity record for 3 October 2026 lists room for two travelers.", "right": "The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-060", "id": "fast-41-diverse-109-060-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers. The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers. Its transfer manifest lists room for four suitcases and a rigid wheelchair, and documents both a boarding ramp and wheelchair securement. The van is scheduled for 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond Elena’s requested deadline. The airport assistance attendant’s service ends at the curb and does not continue to the hotel. Suitcase color and hotel membership appear in the intake record but are irrelevant to routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy and criteria, and both contexts repeat the applicable policy without exceptions. Elena, the airport-to-hotel scope, 3 October 2026 date, and 18:40 deadline remain bound consistently. The evidence contains exactly two complete factual sentences: \"On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers.\" and \"The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers.\" The counterfactual coherently changes the van capacity to two travelers without creating duplicate measurements or contradictions. Neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers. The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers. Its transfer manifest lists room for four suitcases and a rigid wheelchair, and documents both a boarding ramp and wheelchair securement. The van is scheduled for 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond Elena’s requested deadline. The airport assistance attendant’s service ends at the curb and does not continue to the hotel. Suitcase color and hotel membership appear in the intake record but are irrelevant to routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers."}, {"path": [], "text": "The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers.", "negative_left": "On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers.", "negative_right": "The 18:35 accessible van’s capacity record for 3 October 2026 lists room for two travelers.", "right": "The 18:35 accessible van’s capacity record for 3 October 2026 lists room for three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-060", "id": "fast-41-diverse-109-060-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 3 October 2026, Elena’s complete airport-to-hotel party consists of three travelers. The 18:35 accessible van’s capacity record for 3 October 2026 lists room for two travelers. Its transfer manifest lists room for four suitcases and a rigid wheelchair, and documents both a boarding ramp and wheelchair securement. The van is scheduled for 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond Elena’s requested deadline. The airport assistance attendant’s service ends at the curb and does not continue to the hotel. Suitcase color and hotel membership appear in the intake record but are irrelevant to routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions are included verbatim, and both contexts retain the governing policy. Entity, airport-to-hotel path, and timing bindings remain unchanged; the changed party count is an allowed observation change. The two evidence spans are complete factual sentences: “On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.” and “On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers.” The counterfactual coherently changes the van capacity from four to three while the party remains four. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers. On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers. Elena is traveling with four suitcases and a rigid wheelchair, and the van’s records document room for at least four suitcases and the wheelchair. The same records document a ramp and wheelchair securement. Its scheduled pickup is 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its earlier time. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant offers help only between baggage claim and the curb, not the complete journey to the hotel. Blue suitcase labels and Gold hotel membership appear in the file but are irrelevant to routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers."}, {"path": [], "text": "On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.", "negative_left": "On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.", "negative_right": "On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is three travelers.", "right": "On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-070", "id": "fast-41-diverse-109-070-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers. On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers. Elena is traveling with four suitcases and a rigid wheelchair, and the van’s records document room for at least four suitcases and the wheelchair. The same records document a ramp and wheelchair securement. Its scheduled pickup is 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its earlier time. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant offers help only between baggage claim and the curb, not the complete journey to the hotel. Blue suitcase labels and Gold hotel membership appear in the file but are irrelevant to routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions are included verbatim, and both contexts retain the governing policy. Entity, airport-to-hotel path, and timing bindings remain unchanged; the changed party count is an allowed observation change. The two evidence spans are complete factual sentences: “On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.” and “On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers.” The counterfactual coherently changes the van capacity from four to three while the party remains four. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers. On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers. Elena is traveling with four suitcases and a rigid wheelchair, and the van’s records document room for at least four suitcases and the wheelchair. The same records document a ramp and wheelchair securement. Its scheduled pickup is 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its earlier time. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant offers help only between baggage claim and the curb, not the complete journey to the hotel. Blue suitcase labels and Gold hotel membership appear in the file but are irrelevant to routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers."}, {"path": [], "text": "On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.", "negative_left": "On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.", "negative_right": "On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is three travelers.", "right": "On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is four travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-070", "id": "fast-41-diverse-109-070-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 7 September 2026, Elena’s entire airport-to-hotel party comprises four travelers. On 7 September 2026, the 18:35 accessible van’s documented passenger capacity is three travelers. Elena is traveling with four suitcases and a rigid wheelchair, and the van’s records document room for at least four suitcases and the wheelchair. The same records document a ramp and wheelchair securement. Its scheduled pickup is 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its earlier time. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant offers help only between baggage claim and the curb, not the complete journey to the hotel. Blue suitcase labels and Gold hotel membership appear in the file but are irrelevant to routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, criteria, thresholds, and instructions. Entity, route, deadline, party-size, luggage, and accessibility bindings remain intact. The two focus evidence spans are complete factual sentences: \"On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.\" and \"On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers.\" The counterfactual coherently changes that van’s capacity to two without creating contradictory duplicate assertions within its context. Neither context embeds an answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At Northstar Airport, the dispatch log for 17 September 2026 records the following. On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers. The same van is documented to carry four suitcases and Elena’s rigid wheelchair; its equipment list includes a ramp and wheelchair securement. Its pickup is scheduled for 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Blue suitcase color and Gold hotel membership appear in the file but are not routing requirements.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is two travelers.", "right": "On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-075", "id": "fast-41-diverse-109-075-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At Northstar Airport, the dispatch log for 17 September 2026 records the following. On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers. The same van is documented to carry four suitcases and Elena’s rigid wheelchair; its equipment list includes a ramp and wheelchair securement. Its pickup is scheduled for 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Blue suitcase color and Gold hotel membership appear in the file but are not routing requirements."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, criteria, thresholds, and instructions. Entity, route, deadline, party-size, luggage, and accessibility bindings remain intact. The two focus evidence spans are complete factual sentences: \"On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.\" and \"On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers.\" The counterfactual coherently changes that van’s capacity to two without creating contradictory duplicate assertions within its context. Neither context embeds an answer, label rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At Northstar Airport, the dispatch log for 17 September 2026 records the following. On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers. The same van is documented to carry four suitcases and Elena’s rigid wheelchair; its equipment list includes a ramp and wheelchair securement. Its pickup is scheduled for 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Blue suitcase color and Gold hotel membership appear in the file but are not routing requirements.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is two travelers.", "right": "On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-075", "id": "fast-41-diverse-109-075-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At Northstar Airport, the dispatch log for 17 September 2026 records the following. On 17 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 17 September 2026, the documented passenger capacity of the 18:35 accessible van is two travelers. The same van is documented to carry four suitcases and Elena’s rigid wheelchair; its equipment list includes a ramp and wheelchair securement. Its pickup is scheduled for 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Blue suitcase color and Gold hotel membership appear in the file but are not routing requirements."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the governing transfer policy, with the original questions object preserving all question-level criteria. Changed party size and van capacity are observations, while Elena, airport-to-hotel scope, schedules, and the 18:40 deadline remain bound. The evidence consists of the complete factual sentences “On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.” and “On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers.” The counterfactual coherently makes van capacity insufficient for the four-person party without contradictory duplicate measurements. Neither context states a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers. The transfer manifest lists room for four suitcases and documented accommodation for Elena’s rigid wheelchair. It also records a boarding ramp and wheelchair securement for that journey. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. A standard sedan is listed for 18:25, but its file contains no documented ramp for Elena’s transfer. An accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s role ends at the curb and does not include transportation to the hotel. A blue luggage tag and Gold hotel membership appear in the records but have no routing significance. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers."}, {"path": [], "text": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.", "negative_right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers.", "right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-082", "id": "fast-41-diverse-109-082-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers. The transfer manifest lists room for four suitcases and documented accommodation for Elena’s rigid wheelchair. It also records a boarding ramp and wheelchair securement for that journey. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. A standard sedan is listed for 18:25, but its file contains no documented ramp for Elena’s transfer. An accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s role ends at the curb and does not include transportation to the hotel. A blue luggage tag and Gold hotel membership appear in the records but have no routing significance. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the governing transfer policy, with the original questions object preserving all question-level criteria. Changed party size and van capacity are observations, while Elena, airport-to-hotel scope, schedules, and the 18:40 deadline remain bound. The evidence consists of the complete factual sentences “On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.” and “On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers.” The counterfactual coherently makes van capacity insufficient for the four-person party without contradictory duplicate measurements. Neither context states a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers. The transfer manifest lists room for four suitcases and documented accommodation for Elena’s rigid wheelchair. It also records a boarding ramp and wheelchair securement for that journey. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. A standard sedan is listed for 18:25, but its file contains no documented ramp for Elena’s transfer. An accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s role ends at the curb and does not include transportation to the hotel. A blue luggage tag and Gold hotel membership appear in the records but have no routing significance. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers."}, {"path": [], "text": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers.", "negative_right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers.", "right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-082", "id": "fast-41-diverse-109-082-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises four travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is three travelers. The transfer manifest lists room for four suitcases and documented accommodation for Elena’s rigid wheelchair. It also records a boarding ramp and wheelchair securement for that journey. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. A standard sedan is listed for 18:25, but its file contains no documented ramp for Elena’s transfer. An accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s role ends at the curb and does not include transportation to the hotel. A blue luggage tag and Gold hotel membership appear in the records but have no routing significance. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and Elena’s airport-to-hotel, deadline, and option bindings. The evidence contains exactly two complete factual sentences: “On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers.” and “On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers.” The counterfactual coherently changes the van’s capacity to four without creating contradictory duplicate measurements or assertions. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers. The transfer manifest lists room for four suitcases and a rigid wheelchair, with a deployed ramp and documented wheelchair securement. The van is scheduled to leave at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its earlier departure. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s assignment ends at the curb after baggage claim and does not continue to the hotel. Elena’s blue suitcase labels and Gold hotel membership are administrative notes only. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers."}, {"path": [], "text": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers.", "negative_right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers.", "right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-083", "id": "fast-41-diverse-109-083-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers. The transfer manifest lists room for four suitcases and a rigid wheelchair, with a deployed ramp and documented wheelchair securement. The van is scheduled to leave at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its earlier departure. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s assignment ends at the curb after baggage claim and does not continue to the hotel. Elena’s blue suitcase labels and Gold hotel membership are administrative notes only. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and Elena’s airport-to-hotel, deadline, and option bindings. The evidence contains exactly two complete factual sentences: “On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers.” and “On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers.” The counterfactual coherently changes the van’s capacity to four without creating contradictory duplicate measurements or assertions. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers. The transfer manifest lists room for four suitcases and a rigid wheelchair, with a deployed ramp and documented wheelchair securement. The van is scheduled to leave at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its earlier departure. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s assignment ends at the curb after baggage claim and does not continue to the hotel. Elena’s blue suitcase labels and Gold hotel membership are administrative notes only. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers."}, {"path": [], "text": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers.", "negative_right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers.", "right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is five travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-083", "id": "fast-41-diverse-109-083-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises five travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van is four travelers. The transfer manifest lists room for four suitcases and a rigid wheelchair, with a deployed ramp and documented wheelchair securement. The van is scheduled to leave at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its earlier departure. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s assignment ends at the curb after baggage claim and does not continue to the hotel. Elena’s blue suitcase labels and Gold hotel membership are administrative notes only. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and all choice criteria. Both contexts retain Elena, the airport-to-hotel scope, the date, and the pickup deadline. The two evidence spans are complete factual sentences. The counterfactual changes only the van capacity and contains no internal contradiction. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is three travelers. The van’s dispatch record lists room for four suitcases and documented capacity for Elena’s rigid wheelchair. It also lists a ramp and wheelchair securement, with pickup at 18:35. Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its passenger listing. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. A luggage color note and a loyalty-program entry appear in the file but are not service requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is two travelers.", "right": "On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-088", "id": "fast-41-diverse-109-088-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is three travelers. The van’s dispatch record lists room for four suitcases and documented capacity for Elena’s rigid wheelchair. It also lists a ramp and wheelchair securement, with pickup at 18:35. Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its passenger listing. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. A luggage color note and a loyalty-program entry appear in the file but are not service requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and all choice criteria. Both contexts retain Elena, the airport-to-hotel scope, the date, and the pickup deadline. The two evidence spans are complete factual sentences. The counterfactual changes only the van capacity and contains no internal contradiction. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is three travelers. The van’s dispatch record lists room for four suitcases and documented capacity for Elena’s rigid wheelchair. It also lists a ramp and wheelchair securement, with pickup at 18:35. Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its passenger listing. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. A luggage color note and a loyalty-program entry appear in the file but are not service requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is two travelers.", "right": "On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-088", "id": "fast-41-diverse-109-088-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 22 September 2026, the documented passenger capacity recorded for the 18:35 accessible van is two travelers. The van’s dispatch record lists room for four suitcases and documented capacity for Elena’s rigid wheelchair. It also lists a ramp and wheelchair securement, with pickup at 18:35. Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer, despite its passenger listing. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. A luggage color note and a loyalty-program entry appear in the file but are not service requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both contexts, with the unchanged questions object retaining all criteria. Question bindings remain fixed for Elena’s airport-to-hotel route, 22 July 2026, the 18:40 deadline, and the 18:35 van. Evidence contains exactly two complete factual sentences: “On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers.” and “The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers.” The counterfactual coherently changes only the van’s documented passenger capacity from three to two, without internal duplicate contradictions. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers. The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers. Its manifest records room for four suitcases and Elena’s rigid wheelchair, and its equipment log confirms a loading ramp and wheelchair securement. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan’s inspection record shows no ramp for this transfer. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s service ends at the curb after baggage claim and does not continue to the hotel. Elena’s blue suitcase tags and Gold hotel membership appear in the intake record but are not transfer requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for two travelers.", "right": "The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-105", "id": "fast-41-diverse-109-105-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers. The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers. Its manifest records room for four suitcases and Elena’s rigid wheelchair, and its equipment log confirms a loading ramp and wheelchair securement. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan’s inspection record shows no ramp for this transfer. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s service ends at the curb after baggage claim and does not continue to the hotel. Elena’s blue suitcase tags and Gold hotel membership appear in the intake record but are not transfer requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both contexts, with the unchanged questions object retaining all criteria. Question bindings remain fixed for Elena’s airport-to-hotel route, 22 July 2026, the 18:40 deadline, and the 18:35 van. Evidence contains exactly two complete factual sentences: “On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers.” and “The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers.” The counterfactual coherently changes only the van’s documented passenger capacity from three to two, without internal duplicate contradictions. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers. The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers. Its manifest records room for four suitcases and Elena’s rigid wheelchair, and its equipment log confirms a loading ramp and wheelchair securement. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan’s inspection record shows no ramp for this transfer. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s service ends at the curb after baggage claim and does not continue to the hotel. Elena’s blue suitcase tags and Gold hotel membership appear in the intake record but are not transfer requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for two travelers.", "right": "The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-105", "id": "fast-41-diverse-109-105-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 July 2026, Elena’s entire airport-to-hotel party comprises three travelers. The 18:35 accessible van’s transfer record dated 22 July 2026 lists documented passenger capacity for two travelers. Its manifest records room for four suitcases and Elena’s rigid wheelchair, and its equipment log confirms a loading ramp and wheelchair securement. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan’s inspection record shows no ramp for this transfer. The accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s service ends at the curb after baggage claim and does not continue to the hotel. Elena’s blue suitcase tags and Gold hotel membership appear in the intake record but are not transfer requirements. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and all choice criteria. Both contexts retain Elena’s airport-to-hotel scope, party, luggage, wheelchair, deadline, and transfer-time bindings. The evidence contains two complete factual sentences: \"On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.\" and \"On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers.\" The counterfactual coherently changes the van capacity to two without duplicate or contradictory measurements. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers. The same ledger records room for four suitcases and documented accommodation for Elena’s rigid wheelchair, plus a ramp and wheelchair securement. Its scheduled pickup is 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The 18:50 accessible minibus is scheduled after the requested deadline. An airport assistance attendant handles Elena only from baggage claim to the curb, not the complete journey to her hotel. A luggage tag notes a different color, and a loyalty record notes hotel membership; neither affects dispatch. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for two travelers.", "right": "On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-110", "id": "fast-41-diverse-109-110-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers. The same ledger records room for four suitcases and documented accommodation for Elena’s rigid wheelchair, plus a ramp and wheelchair securement. Its scheduled pickup is 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The 18:50 accessible minibus is scheduled after the requested deadline. An airport assistance attendant handles Elena only from baggage claim to the curb, not the complete journey to her hotel. A luggage tag notes a different color, and a loyalty record notes hotel membership; neither affects dispatch. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and all choice criteria. Both contexts retain Elena’s airport-to-hotel scope, party, luggage, wheelchair, deadline, and transfer-time bindings. The evidence contains two complete factual sentences: \"On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.\" and \"On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers.\" The counterfactual coherently changes the van capacity to two without duplicate or contradictory measurements. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers. The same ledger records room for four suitcases and documented accommodation for Elena’s rigid wheelchair, plus a ramp and wheelchair securement. Its scheduled pickup is 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The 18:50 accessible minibus is scheduled after the requested deadline. An airport assistance attendant handles Elena only from baggage claim to the curb, not the complete journey to her hotel. A luggage tag notes a different color, and a loyalty record notes hotel membership; neither affects dispatch. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers."}, {"path": [], "text": "On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers.", "negative_right": "On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for two travelers.", "right": "On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for three travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-110", "id": "fast-41-diverse-109-110-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party comprises three travelers. On 22 September 2026, the dispatch ledger for the 18:35 accessible van records passenger capacity for two travelers. The same ledger records room for four suitcases and documented accommodation for Elena’s rigid wheelchair, plus a ramp and wheelchair securement. Its scheduled pickup is 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The 18:50 accessible minibus is scheduled after the requested deadline. An airport assistance attendant handles Elena only from baggage claim to the curb, not the complete journey to her hotel. A luggage tag notes a different color, and a loyalty record notes hotel membership; neither affects dispatch. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing transfer policy, and the unchanged original questions preserve all criteria and instructions. Elena, the airport-to-hotel route, and the relevant times remain correctly bound. The evidence contains exactly two complete factual sentences: “On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers.” “The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers.” The counterfactual coherently changes the van capacity to four without creating contradictory duplicate measurements or assertions. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers. The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers. A separate equipment record lists room for four suitcases and documented capacity for Elena’s rigid wheelchair, plus a ramp and wheelchair securement for the airport-to-hotel transfer. The van is scheduled to collect Elena at 18:35, while her requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. A dispatch note records blue suitcase covers and Gold hotel membership as incidental details. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers."}, {"path": [], "text": "The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers.", "negative_left": "On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers.", "negative_right": "The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as four travelers.", "right": "The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-124", "id": "fast-41-diverse-109-124-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers. The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers. A separate equipment record lists room for four suitcases and documented capacity for Elena’s rigid wheelchair, plus a ramp and wheelchair securement for the airport-to-hotel transfer. The van is scheduled to collect Elena at 18:35, while her requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. A dispatch note records blue suitcase covers and Gold hotel membership as incidental details. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing transfer policy, and the unchanged original questions preserve all criteria and instructions. Elena, the airport-to-hotel route, and the relevant times remain correctly bound. The evidence contains exactly two complete factual sentences: “On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers.” “The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers.” The counterfactual coherently changes the van capacity to four without creating contradictory duplicate measurements or assertions. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers. The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers. A separate equipment record lists room for four suitcases and documented capacity for Elena’s rigid wheelchair, plus a ramp and wheelchair securement for the airport-to-hotel transfer. The van is scheduled to collect Elena at 18:35, while her requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. A dispatch note records blue suitcase covers and Gold hotel membership as incidental details. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers."}, {"path": [], "text": "The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers.", "negative_left": "On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers.", "negative_right": "The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as four travelers.", "right": "The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as five travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-124", "id": "fast-41-diverse-109-124-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 14 September 2026, Elena’s entire airport-to-hotel party consists of five travelers. The 14 September 2026 capacity inspection records the 18:35 accessible van’s documented passenger capacity as four travelers. A separate equipment record lists room for four suitcases and documented capacity for Elena’s rigid wheelchair, plus a ramp and wheelchair securement for the airport-to-hotel transfer. The van is scheduled to collect Elena at 18:35, while her requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. The accessible minibus is scheduled for 18:50, beyond the requested deadline. The airport assistance attendant’s service ends at the curb and does not provide the complete journey to Elena’s hotel. A dispatch note records blue suitcase covers and Gold hotel membership as incidental details. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and the contexts repeat it without adding exceptions. Elena, the airport-to-hotel route, the 18:35 transfer, and the 18:40 deadline remain bound consistently, while party size may change as an observation. The evidence contains exactly two complete factual sentences: “On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers.” and “On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers.” The counterfactual coherently changes van capacity to four, which is insufficient for the five-person party, without contradictory duplicate measurements. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers. The dispatch record lists room for four suitcases and documented accommodation for Elena’s rigid wheelchair. It also records a deployable ramp and wheelchair securement for the airport-to-hotel journey. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. An accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s logged service ends at the curb and does not provide the complete trip to the hotel. Blue suitcase tags and Gold hotel membership appear in the intake record but do not affect routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers."}, {"path": [], "text": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers.", "negative_right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is four travelers.", "right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-125", "id": "fast-41-diverse-109-125-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers. The dispatch record lists room for four suitcases and documented accommodation for Elena’s rigid wheelchair. It also records a deployable ramp and wheelchair securement for the airport-to-hotel journey. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. An accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s logged service ends at the curb and does not provide the complete trip to the hotel. Blue suitcase tags and Gold hotel membership appear in the intake record but do not affect routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and the contexts repeat it without adding exceptions. Elena, the airport-to-hotel route, the 18:35 transfer, and the 18:40 deadline remain bound consistently, while party size may change as an observation. The evidence contains exactly two complete factual sentences: “On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers.” and “On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers.” The counterfactual coherently changes van capacity to four, which is insufficient for the five-person party, without contradictory duplicate measurements. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers. The dispatch record lists room for four suitcases and documented accommodation for Elena’s rigid wheelchair. It also records a deployable ramp and wheelchair securement for the airport-to-hotel journey. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. An accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s logged service ends at the curb and does not provide the complete trip to the hotel. Blue suitcase tags and Gold hotel membership appear in the intake record but do not affect routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers."}, {"path": [], "text": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers.", "negative_left": "On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers.", "negative_right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is four travelers.", "right": "On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is five travelers."}, "verifier_independent_model": false}, "family": "fast-41-diverse-109-125", "id": "fast-41-diverse-109-125-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "On 22 September 2026, Elena’s entire airport-to-hotel party for the 18:35 transfer comprises five travelers. On 22 September 2026, the documented passenger capacity of the 18:35 accessible van for Elena’s airport-to-hotel transfer is four travelers. The dispatch record lists room for four suitcases and documented accommodation for Elena’s rigid wheelchair. It also records a deployable ramp and wheelchair securement for the airport-to-hotel journey. The van is scheduled to depart at 18:35, while Elena’s requested hotel-pickup deadline is 18:40. The 18:25 standard sedan has no documented ramp for this transfer. An accessible minibus is scheduled for 18:50, after the requested deadline. The airport assistance attendant’s logged service ends at the curb and does not provide the complete trip to the hotel. Blue suitcase tags and Gold hotel membership appear in the intake record but do not affect routing. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, and both contexts preserve every governing policy. Airport, date, curb-ready time, entity, and single-routing-choice scope remain bound. The two focus spans are complete factual sentences: \"The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.\" and \"Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags.\" The counterfactual coherently changes Access B's capacity from 4 bags to 2 bags while the manifest remains at 3 bags. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The party must remain in a rigid power chair, and road transport is required after the airport assistance attendant brings them to the curb at 18:40.\",\"The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.\",\"Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags.\",\"Access A has a documented ramp, sufficient occupied-chair securements, passenger capacity, and luggage capacity; its pickup is at 18:50.\",\"Access B has a documented lift, sufficient occupied-chair securements and passenger capacity, and its pickup is at 18:42.\",\"The standard van requires a chair transfer. The 19:00 hotel shuttle lacks sufficient passenger and luggage capacity for the party. Assistance alone does not provide road transport.\"]}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags."}, {"path": ["evidence", "2"], "text": "Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.", "negative_left": "The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.", "negative_right": "Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 2 bags.", "right": "Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags."}, "verifier_independent_model": false}, "family": "fast-41-diverse-110-003", "id": "fast-41-diverse-110-003-base", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The party must remain in a rigid power chair, and road transport is required after the airport assistance attendant brings them to the curb at 18:40.", "The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.", "Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags.", "Access A has a documented ramp, sufficient occupied-chair securements, passenger capacity, and luggage capacity; its pickup is at 18:50.", "Access B has a documented lift, sufficient occupied-chair securements and passenger capacity, and its pickup is at 18:42.", "The standard van requires a chair transfer. The 19:00 hotel shuttle lacks sufficient passenger and luggage capacity for the party. Assistance alone does not provide road transport."]}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_b"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, and both contexts preserve every governing policy. Airport, date, curb-ready time, entity, and single-routing-choice scope remain bound. The two focus spans are complete factual sentences: \"The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.\" and \"Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags.\" The counterfactual coherently changes Access B's capacity from 4 bags to 2 bags while the manifest remains at 3 bags. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The party must remain in a rigid power chair, and road transport is required after the airport assistance attendant brings them to the curb at 18:40.\",\"The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.\",\"Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags.\",\"Access A has a documented ramp, sufficient occupied-chair securements, passenger capacity, and luggage capacity; its pickup is at 18:50.\",\"Access B has a documented lift, sufficient occupied-chair securements and passenger capacity, and its pickup is at 18:42.\",\"The standard van requires a chair transfer. The 19:00 hotel shuttle lacks sufficient passenger and luggage capacity for the party. Assistance alone does not provide road transport.\"]}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "1"], "text": "The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags."}, {"path": ["evidence", "2"], "text": "Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.", "negative_left": "The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.", "negative_right": "Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 2 bags.", "right": "Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 4 bags."}, "verifier_independent_model": false}, "family": "fast-41-diverse-110-003", "id": "fast-41-diverse-110-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The party must remain in a rigid power chair, and road transport is required after the airport assistance attendant brings them to the curb at 18:40.", "The arriving party's checked manifest, issued at Northstar Airport on 17 September 2026, lists 3 bags.", "Access B's capacity placard, valid for the 17 September 2026 dispatch, lists room for 2 bags.", "Access A has a documented ramp, sufficient occupied-chair securements, passenger capacity, and luggage capacity; its pickup is at 18:50.", "Access B has a documented lift, sufficient occupied-chair securements and passenger capacity, and its pickup is at 18:42.", "The standard van requires a chair transfer. The 19:00 hotel shuttle lacks sufficient passenger and luggage capacity for the party. Assistance alone does not provide road transport."]}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_a"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, request scope, and all entity, path, and time bindings. Both focus spans are complete factual sentences: \"Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.\" \"The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026.\" The counterfactual coherently changes the party's count to 5 bags, making it exceed Access B's stated limit without creating duplicate contradictions. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The party requires road transport after reaching the curb, and the passenger remains in her mobility device; the standard van cannot accommodate that arrangement.\",\"Access A's 18:50 record documents a ramp, adequate chair securements, passenger space, and luggage space, and it is after the 18:40 curb-ready time.\",\"Access B's 18:42 record documents a lift, adequate chair securements and passenger space, and its pickup is after 18:40 and earlier than Access A.\",\"Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.\",\"The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026.\",\"The 19:00 hotel shuttle does not have enough passenger or luggage capacity for the complete party.\",\"Airport assistance reaches only the curb, so it cannot complete the requested journey.\"]}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "3"], "text": "Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026."}, {"path": ["evidence", "4"], "text": "The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.", "negative_left": "Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.", "negative_right": "The arriving party checked 5 bags for the 18:42 Access B pickup on 17 September 2026.", "right": "The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-110-025", "id": "fast-41-diverse-110-025-base", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The party requires road transport after reaching the curb, and the passenger remains in her mobility device; the standard van cannot accommodate that arrangement.", "Access A's 18:50 record documents a ramp, adequate chair securements, passenger space, and luggage space, and it is after the 18:40 curb-ready time.", "Access B's 18:42 record documents a lift, adequate chair securements and passenger space, and its pickup is after 18:40 and earlier than Access A.", "Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.", "The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026.", "The 19:00 hotel shuttle does not have enough passenger or luggage capacity for the complete party.", "Airport assistance reaches only the curb, so it cannot complete the requested journey."]}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_b"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, request scope, and all entity, path, and time bindings. Both focus spans are complete factual sentences: \"Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.\" \"The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026.\" The counterfactual coherently changes the party's count to 5 bags, making it exceed Access B's stated limit without creating duplicate contradictions. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.\",\"evidence\":[\"The party requires road transport after reaching the curb, and the passenger remains in her mobility device; the standard van cannot accommodate that arrangement.\",\"Access A's 18:50 record documents a ramp, adequate chair securements, passenger space, and luggage space, and it is after the 18:40 curb-ready time.\",\"Access B's 18:42 record documents a lift, adequate chair securements and passenger space, and its pickup is after 18:40 and earlier than Access A.\",\"Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.\",\"The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026.\",\"The 19:00 hotel shuttle does not have enough passenger or luggage capacity for the complete party.\",\"Airport assistance reaches only the curb, so it cannot complete the requested journey.\"]}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["evidence", "3"], "text": "Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026."}, {"path": ["evidence", "4"], "text": "The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.", "negative_left": "Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.", "negative_right": "The arriving party checked 5 bags for the 18:42 Access B pickup on 17 September 2026.", "right": "The arriving party checked 4 bags for the 18:42 Access B pickup on 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-110-025", "id": "fast-41-diverse-110-025-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time.", "evidence": ["The party requires road transport after reaching the curb, and the passenger remains in her mobility device; the standard van cannot accommodate that arrangement.", "Access A's 18:50 record documents a ramp, adequate chair securements, passenger space, and luggage space, and it is after the 18:40 curb-ready time.", "Access B's 18:42 record documents a lift, adequate chair securements and passenger space, and its pickup is after 18:40 and earlier than Access A.", "Access B's dispatch record lists a luggage limit of 4 bags for the 18:42 pickup on 17 September 2026.", "The arriving party checked 5 bags for the 18:42 Access B pickup on 17 September 2026.", "The 19:00 hotel shuttle does not have enough passenger or luggage capacity for the complete party.", "Airport assistance reaches only the curb, so it cannot complete the requested journey."]}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_a"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing policy. The entity is LiftVan B, and the party, wheelchair, assistance, luggage, and 18:35 deadline bindings remain intact. The evidence consists of two complete factual sentences: \"The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases.\" \"The party manifest dated 17 September 2026 lists 4 suitcases for this party.\" The counterfactual coherently changes the manifest to six suitcases, which exceeds the stated five-suitcase capacity. Neither context contains an answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"dispatcher\",\"text\":\"The transfer log contains exactly sedan S, LiftVan B, and LiftVan C. The arriving party has a non-transfer passenger who remains in a rigid wheelchair, and the party needs boarding assistance.\"},{\"speaker\":\"dispatcher\",\"text\":\"LiftVan B is accessible for that passenger, with documented boarding equipment and wheelchair securement. It accommodates the party's traveler count and one rigid wheelchair, and its pickup ETA is 18:28.\"},{\"speaker\":\"assistance attendant\",\"text\":\"Boarding assistance for this party is booked at the curb for 18:25, no later than LiftVan B's pickup ETA. The requested pickup deadline is 18:35.\"},{\"speaker\":\"dispatcher\",\"text\":\"Sedan S's pickup ETA is 18:18, earlier than LiftVan B's, but Sedan S lacks explicit documented evidence of boarding equipment. LiftVan C's pickup ETA is 18:34, later than LiftVan B's.\"},{\"speaker\":\"inventory\",\"text\":\"The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases.\"},{\"speaker\":\"manifest\",\"text\":\"The party manifest dated 17 September 2026 lists 4 suitcases for this party.\"},{\"speaker\":\"policy\",\"text\":\"I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35.\"},{\"speaker\":\"policy\",\"text\":\"Policy routes non-transfer passengers to accessible dispatch.\"},{\"speaker\":\"policy\",\"text\":\"Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases."}, {"path": ["5", "text"], "text": "The party manifest dated 17 September 2026 lists 4 suitcases for this party."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases.", "negative_left": "The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases.", "negative_right": "The party manifest dated 17 September 2026 lists 6 suitcases for this party.", "right": "The party manifest dated 17 September 2026 lists 4 suitcases for this party."}, "verifier_independent_model": false}, "family": "fast-41-diverse-111-013", "id": "fast-41-diverse-111-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "dispatcher", "text": "The transfer log contains exactly sedan S, LiftVan B, and LiftVan C. The arriving party has a non-transfer passenger who remains in a rigid wheelchair, and the party needs boarding assistance."}, {"speaker": "dispatcher", "text": "LiftVan B is accessible for that passenger, with documented boarding equipment and wheelchair securement. It accommodates the party's traveler count and one rigid wheelchair, and its pickup ETA is 18:28."}, {"speaker": "assistance attendant", "text": "Boarding assistance for this party is booked at the curb for 18:25, no later than LiftVan B's pickup ETA. The requested pickup deadline is 18:35."}, {"speaker": "dispatcher", "text": "Sedan S's pickup ETA is 18:18, earlier than LiftVan B's, but Sedan S lacks explicit documented evidence of boarding equipment. LiftVan C's pickup ETA is 18:34, later than LiftVan B's."}, {"speaker": "inventory", "text": "The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases."}, {"speaker": "manifest", "text": "The party manifest dated 17 September 2026 lists 4 suitcases for this party."}, {"speaker": "policy", "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"speaker": "policy", "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"speaker": "policy", "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing policy. The entity is LiftVan B, and the party, wheelchair, assistance, luggage, and 18:35 deadline bindings remain intact. The evidence consists of two complete factual sentences: \"The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases.\" \"The party manifest dated 17 September 2026 lists 4 suitcases for this party.\" The counterfactual coherently changes the manifest to six suitcases, which exceeds the stated five-suitcase capacity. Neither context contains an answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relationships rather than bundled policy conclusions, and a6 is a factual capacity comparison. The base and counter assignments can be realized with only LiftVan B's suitcase sufficiency changing. The policy evidence preserves the substantive state-originating rules needed alongside the unchanged question: accessible dispatch for non-transfer passengers, all mandatory suitability checks, the pickup deadline, and earliest-ETA ranking among suitable vehicles. Extra state case observations need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes every LiftVan B suitability requirement, including sufficient party, chair, and luggage capacity, required equipment and securement, timely pickup, and booked assistance. It also excludes unlisted competitors, establishes that the earlier sedan fails a mandatory requirement, and establishes that LiftVan C is later. Therefore LiftVan B is the earliest suitable option.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a6 entails that LiftVan B lacks sufficient suitcase capacity. Luggage capacity is a mandatory suitability check, so this failure is sufficient for the false outcome regardless of ETA or other requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "LiftVan B is an accessible vehicle for the party's non-transfer passenger."}, {"id": "a2", "statement": "LiftVan B has explicit documented evidence of boarding equipment."}, {"id": "a3", "statement": "LiftVan B has explicit documented evidence of wheelchair securement."}, {"id": "a4", "statement": "LiftVan B's traveler capacity is at least the party's traveler count."}, {"id": "a5", "statement": "LiftVan B's rigid-wheelchair capacity is at least the party's rigid-wheelchair count."}, {"id": "a6", "statement": "LiftVan B's suitcase capacity is at least the party's suitcase count."}, {"id": "a7", "statement": "LiftVan B's pickup ETA is no later than 18:35."}, {"id": "a8", "statement": "Boarding assistance for this party is booked at the curb."}, {"id": "a9", "statement": "The booked curb-assistance time is no later than LiftVan B's pickup ETA."}, {"id": "a10", "statement": "The logged transfer options are exactly sedan S, LiftVan B, and LiftVan C."}, {"id": "a11", "statement": "Sedan S's pickup ETA is earlier than LiftVan B's pickup ETA."}, {"id": "a12", "statement": "LiftVan B's pickup ETA is earlier than LiftVan C's pickup ETA."}, {"id": "a13", "statement": "Sedan S lacks explicit documented evidence of boarding equipment."}], "base_state_json": "[{\"speaker\":\"dispatcher\",\"text\":\"The transfer log contains exactly sedan S, LiftVan B, and LiftVan C. The arriving party has a non-transfer passenger who remains in a rigid wheelchair, and the party needs boarding assistance.\"},{\"speaker\":\"dispatcher\",\"text\":\"LiftVan B is accessible for that passenger, with documented boarding equipment and wheelchair securement. It accommodates the party's traveler count and one rigid wheelchair, and its pickup ETA is 18:28.\"},{\"speaker\":\"assistance attendant\",\"text\":\"Boarding assistance for this party is booked at the curb for 18:25, no later than LiftVan B's pickup ETA. The requested pickup deadline is 18:35.\"},{\"speaker\":\"dispatcher\",\"text\":\"Sedan S's pickup ETA is 18:18, earlier than LiftVan B's, but Sedan S lacks explicit documented evidence of boarding equipment. LiftVan C's pickup ETA is 18:34, later than LiftVan B's.\"},{\"speaker\":\"inventory\",\"text\":\"The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases.\"},{\"speaker\":\"manifest\",\"text\":\"The party manifest dated 17 September 2026 lists 4 suitcases for this party.\"},{\"speaker\":\"policy\",\"text\":\"I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35.\"},{\"speaker\":\"policy\",\"text\":\"Policy routes non-transfer passengers to accessible dispatch.\"},{\"speaker\":\"policy\",\"text\":\"Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases."}, {"path": ["5", "text"], "text": "The party manifest dated 17 September 2026 lists 4 suitcases for this party."}], "policy_evidence": [{"path": ["0", "text"], "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"path": ["3", "text"], "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"path": ["3", "text"], "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}], "rules": [{"justification": "LiftVan B passes every documented suitability requirement. Sedan S is earlier but fails the explicit boarding-equipment-evidence requirement, while LiftVan C is later, so LiftVan B is the earliest suitable logged option.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A suitcase capacity below the party's suitcase count explicitly fails the required luggage suitability check, so LiftVan B cannot be selected regardless of its ETA.", "target": "false", "when": [{"atom_id": "a6", "state": "refuted"}]}]}, "verified_pair": {"left": "The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases.", "negative_left": "The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases.", "negative_right": "The party manifest dated 17 September 2026 lists 6 suitcases for this party.", "right": "The party manifest dated 17 September 2026 lists 4 suitcases for this party."}, "verifier_independent_model": false}, "family": "fast-41-diverse-111-013", "id": "fast-41-diverse-111-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — LiftVan B fails at least one required suitability check, so it cannot be selected despite its earlier ETA.", "true": "Yes — LiftVan B satisfies every documented suitability requirement and is the earliest suitable option."}, "instructions": "Decide yes or no: Under the stated policy, should LiftVan B be selected as the suitable, highest-ranked transfer for this party? Apply every suitability requirement before comparing pickup times.", "type": "noul"}}, "state": [{"speaker": "dispatcher", "text": "The transfer log contains exactly sedan S, LiftVan B, and LiftVan C. The arriving party has a non-transfer passenger who remains in a rigid wheelchair, and the party needs boarding assistance."}, {"speaker": "dispatcher", "text": "LiftVan B is accessible for that passenger, with documented boarding equipment and wheelchair securement. It accommodates the party's traveler count and one rigid wheelchair, and its pickup ETA is 18:28."}, {"speaker": "assistance attendant", "text": "Boarding assistance for this party is booked at the curb for 18:25, no later than LiftVan B's pickup ETA. The requested pickup deadline is 18:35."}, {"speaker": "dispatcher", "text": "Sedan S's pickup ETA is 18:18, earlier than LiftVan B's, but Sedan S lacks explicit documented evidence of boarding equipment. LiftVan C's pickup ETA is 18:34, later than LiftVan B's."}, {"speaker": "inventory", "text": "The inventory record dated 17 September 2026 lists LiftVan B's suitcase capacity as 5 suitcases."}, {"speaker": "manifest", "text": "The party manifest dated 17 September 2026 lists 6 suitcases for this party."}, {"speaker": "policy", "text": "I remain in my rigid wheelchair and cannot transfer to a seat. I need boarding help and want pickup by 18:35."}, {"speaker": "policy", "text": "Policy routes non-transfer passengers to accessible dispatch."}, {"speaker": "policy", "text": "Suitability requires explicit evidence for boarding equipment, securement, party, chair, luggage, pickup by 18:35, and booked assistance; rank suitable vehicles by earliest ETA."}]}, "method": "c2d", "provenance": {"source_id": "diverse-111", "source_is_synthetic": true, "source_sha256": "afbe748e2a6a403fc2c5854ce43ed0ee865cd55cbf69c45bdbb8371efb3b3ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing routing policy. The failure, pull-request scope, component identity, and time bindings remain compatible with the original question. The focus evidence contains exactly two complete factual sentences. The counterfactual changes only the registry classification and introduces no contradictory duplicate assertion. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"On 2026-09-16, a pull request changing documentation comments introduced a defect in the component named by the incident record. The defect was the single primary cause of the compile-job failure. The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component. The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as the build cache-invalidation component. The failure involved no test assertion, test setup failure, or flakiness; no tests started. Investigation found no CI worker fault, network fault, or hosted-service fault. Clearing the relevant cache allowed generation to complete and the compile job to pass, while the previous 20 runs succeeded. If the registry snapshot were unavailable, the other incident findings would remain unchanged. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component."}, {"path": [], "text": "The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as the build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component.", "negative_left": "The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component.", "negative_right": "The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as an application-source component.", "right": "The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as the build cache-invalidation component."}, "verifier_independent_model": false}, "family": "fast-41-diverse-127-003", "id": "fast-41-diverse-127-003-base", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "On 2026-09-16, a pull request changing documentation comments introduced a defect in the component named by the incident record. The defect was the single primary cause of the compile-job failure. The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component. The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as the build cache-invalidation component. The failure involved no test assertion, test setup failure, or flakiness; no tests started. Investigation found no CI worker fault, network fault, or hosted-service fault. Clearing the relevant cache allowed generation to complete and the compile job to pass, while the previous 20 runs succeeded. If the registry snapshot were unavailable, the other incident findings would remain unchanged. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing routing policy. The failure, pull-request scope, component identity, and time bindings remain compatible with the original question. The focus evidence contains exactly two complete factual sentences. The counterfactual changes only the registry classification and introduces no contradictory duplicate assertion. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"On 2026-09-16, a pull request changing documentation comments introduced a defect in the component named by the incident record. The defect was the single primary cause of the compile-job failure. The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component. The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as the build cache-invalidation component. The failure involved no test assertion, test setup failure, or flakiness; no tests started. Investigation found no CI worker fault, network fault, or hosted-service fault. Clearing the relevant cache allowed generation to complete and the compile job to pass, while the previous 20 runs succeeded. If the registry snapshot were unavailable, the other incident findings would remain unchanged. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component."}, {"path": [], "text": "The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as the build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component.", "negative_left": "The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component.", "negative_right": "The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as an application-source component.", "right": "The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as the build cache-invalidation component."}, "verifier_independent_model": false}, "family": "fast-41-diverse-127-003", "id": "fast-41-diverse-127-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "On 2026-09-16, a pull request changing documentation comments introduced a defect in the component named by the incident record. The defect was the single primary cause of the compile-job failure. The incident record issued on 2026-09-16 identifies component C-417, and the component registry classifies C-417 as exactly one of the build cache-invalidation component or an application-source component. The component registry snapshot timestamped 2026-09-16 14:00 UTC classifies component C-417 as an application-source component. The failure involved no test assertion, test setup failure, or flakiness; no tests started. Investigation found no CI worker fault, network fault, or hosted-service fault. Clearing the relevant cache allowed generation to complete and the compile job to pass, while the previous 20 runs succeeded. If the registry snapshot were unavailable, the other incident findings would remain unchanged. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "application_owner"}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing criteria and routing scope. Both contexts retain the routing policy and the compile-job triage subject without changing required bindings. The two evidence spans are complete factual sentences. The counterfactual consistently changes the registry component to C-418 while allowing C-417 to be the application-source alternative. Neither context states a gold answer, label code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure. The component registry snapshot dated 14 March 2026 identifies component C-417 as the build cache-invalidation component. The pull request introduced the defect in that incident component, and the defect was the sole primary cause of the compile-job failure. The incident component is exactly one of the registry's build cache-invalidation component or an application source component. Logs recorded no test assertion failure, test setup failure, or flaky behavior, and the test suite did not start. CI worker health was normal, with no network fault or hosted-service fault. Under a hypothetical registry update naming C-418 instead, the incident record and all other failure findings would remain unchanged. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure."}, {"path": [], "text": "The component registry snapshot dated 14 March 2026 identifies component C-417 as the build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure.", "negative_left": "Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure.", "negative_right": "The component registry snapshot dated 14 March 2026 identifies component C-418 as the build cache-invalidation component.", "right": "The component registry snapshot dated 14 March 2026 identifies component C-417 as the build cache-invalidation component."}, "verifier_independent_model": false}, "family": "fast-41-diverse-127-006", "id": "fast-41-diverse-127-006-base", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure. The component registry snapshot dated 14 March 2026 identifies component C-417 as the build cache-invalidation component. The pull request introduced the defect in that incident component, and the defect was the sole primary cause of the compile-job failure. The incident component is exactly one of the registry's build cache-invalidation component or an application source component. Logs recorded no test assertion failure, test setup failure, or flaky behavior, and the test suite did not start. CI worker health was normal, with no network fault or hosted-service fault. Under a hypothetical registry update naming C-418 instead, the incident record and all other failure findings would remain unchanged. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing criteria and routing scope. Both contexts retain the routing policy and the compile-job triage subject without changing required bindings. The two evidence spans are complete factual sentences. The counterfactual consistently changes the registry component to C-418 while allowing C-417 to be the application-source alternative. Neither context states a gold answer, label code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure. The component registry snapshot dated 14 March 2026 identifies component C-417 as the build cache-invalidation component. The pull request introduced the defect in that incident component, and the defect was the sole primary cause of the compile-job failure. The incident component is exactly one of the registry's build cache-invalidation component or an application source component. Logs recorded no test assertion failure, test setup failure, or flaky behavior, and the test suite did not start. CI worker health was normal, with no network fault or hosted-service fault. Under a hypothetical registry update naming C-418 instead, the incident record and all other failure findings would remain unchanged. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure."}, {"path": [], "text": "The component registry snapshot dated 14 March 2026 identifies component C-417 as the build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure.", "negative_left": "Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure.", "negative_right": "The component registry snapshot dated 14 March 2026 identifies component C-418 as the build cache-invalidation component.", "right": "The component registry snapshot dated 14 March 2026 identifies component C-417 as the build cache-invalidation component."}, "verifier_independent_model": false}, "family": "fast-41-diverse-127-006", "id": "fast-41-diverse-127-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "Incident record IR-208, filed on 14 March 2026, identifies component C-417 as the component involved in the compile-job failure. The component registry snapshot dated 14 March 2026 identifies component C-418 as the build cache-invalidation component. The pull request introduced the defect in that incident component, and the defect was the sole primary cause of the compile-job failure. The incident component is exactly one of the registry's build cache-invalidation component or an application source component. Logs recorded no test assertion failure, test setup failure, or flaky behavior, and the test suite did not start. CI worker health was normal, with no network fault or hosted-service fault. Under a hypothetical registry update naming C-418 instead, the incident record and all other failure findings would remain unchanged. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "application_owner"}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain compatible policy facts. The change, run, test, and time bindings remain unchanged. The evidence contains two complete factual sentences: “Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride.” and “Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride.” The counterfactual coherently changes only run 417's recorded failure identifier to ConfigParserTest.rejectsWhitespace. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"Change 8f31 is pending its final readiness review. Run 442 is the latest required run for change 8f31, and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history. policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in reproduction-search set S is a base-branch run, and every run in reproduction-search set S was executed in the same environment as run 442.\"},{\"speaker\":\"Release coordinator\",\"text\":\"Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Build engineer\",\"text\":\"Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["3", "text"], "text": "Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"path": ["4", "text"], "text": "Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride.", "negative_left": "Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride.", "negative_right": "Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsWhitespace.", "right": "Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "fast-41-diverse-129-005", "id": "fast-41-diverse-129-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "Change 8f31 is pending its final readiness review. Run 442 is the latest required run for change 8f31, and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Test automation engineer", "text": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history. policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Build engineer", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in reproduction-search set S is a base-branch run, and every run in reproduction-search set S was executed in the same environment as run 442."}, {"speaker": "Release coordinator", "text": "Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Build engineer", "text": "Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain compatible policy facts. The change, run, test, and time bindings remain unchanged. The evidence contains two complete factual sentences: “Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride.” and “Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride.” The counterfactual coherently changes only run 417's recorded failure identifier to ConfigParserTest.rejectsWhitespace. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"Change 8f31 is pending its final readiness review. Run 442 is the latest required run for change 8f31, and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history. policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in reproduction-search set S is a base-branch run, and every run in reproduction-search set S was executed in the same environment as run 442.\"},{\"speaker\":\"Release coordinator\",\"text\":\"Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Build engineer\",\"text\":\"Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["3", "text"], "text": "Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"path": ["4", "text"], "text": "Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride.", "negative_left": "Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride.", "negative_right": "Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsWhitespace.", "right": "Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "fast-41-diverse-129-005", "id": "fast-41-diverse-129-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "Change 8f31 is pending its final readiness review. Run 442 is the latest required run for change 8f31, and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Test automation engineer", "text": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history. policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Build engineer", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in reproduction-search set S is a base-branch run, and every run in reproduction-search set S was executed in the same environment as run 442."}, {"speaker": "Release coordinator", "text": "Reproduction-search set S contains exactly run 417, and run 442's sole failure has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Build engineer", "text": "Run 417's failure records contain exactly test-case identifier ConfigParserTest.rejectsWhitespace."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy and bindings. Both contexts retain the relevant exemption and reproduction-policy facts. The question entity, run, environment, branch, and temporal bindings remain unchanged. The two evidence spans are complete factual sentences. The counterfactual changes only run 731’s failure identifier and remains consistent with the other assertions. Neither context states an answer, code, rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Change 8f31 is awaiting a merge-readiness decision. Run 442 is the latest required run for change 8f31, and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride. That test failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history.\"},{\"speaker\":\"Policy record\",\"text\":\"policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in reproduction-search set S is a base-branch run, and every run in reproduction-search set S was executed in the same environment as run 442.\"},{\"speaker\":\"Search auditor\",\"text\":\"Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Search auditor\",\"text\":\"Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Release manager\",\"text\":\"Run 442 completed its required build and test execution; the earlier interrupted attempt was not the required result.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["3", "text"], "text": "Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"path": ["4", "text"], "text": "Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride.", "negative_left": "Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride.", "negative_right": "Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.acceptsEmptyOverride.", "right": "Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "fast-41-diverse-129-006", "id": "fast-41-diverse-129-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Change 8f31 is awaiting a merge-readiness decision. Run 442 is the latest required run for change 8f31, and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride. That test failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"speaker": "Policy record", "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Build engineer", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in reproduction-search set S is a base-branch run, and every run in reproduction-search set S was executed in the same environment as run 442."}, {"speaker": "Search auditor", "text": "Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Search auditor", "text": "Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Release manager", "text": "Run 442 completed its required build and test execution; the earlier interrupted attempt was not the required result."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy and bindings. Both contexts retain the relevant exemption and reproduction-policy facts. The question entity, run, environment, branch, and temporal bindings remain unchanged. The two evidence spans are complete factual sentences. The counterfactual changes only run 731’s failure identifier and remains consistent with the other assertions. Neither context states an answer, code, rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Change 8f31 is awaiting a merge-readiness decision. Run 442 is the latest required run for change 8f31, and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride. That test failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history.\"},{\"speaker\":\"Policy record\",\"text\":\"policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in reproduction-search set S is a base-branch run, and every run in reproduction-search set S was executed in the same environment as run 442.\"},{\"speaker\":\"Search auditor\",\"text\":\"Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Search auditor\",\"text\":\"Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Release manager\",\"text\":\"Run 442 completed its required build and test execution; the earlier interrupted attempt was not the required result.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["3", "text"], "text": "Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"path": ["4", "text"], "text": "Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.rejectsEmptyOverride."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride.", "negative_left": "Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride.", "negative_right": "Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.acceptsEmptyOverride.", "right": "Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, "verifier_independent_model": false}, "family": "fast-41-diverse-129-006", "id": "fast-41-diverse-129-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Change 8f31 is awaiting a merge-readiness decision. Run 442 is the latest required run for change 8f31, and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride. That test failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"speaker": "Policy record", "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Build engineer", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in reproduction-search set S is a base-branch run, and every run in reproduction-search set S was executed in the same environment as run 442."}, {"speaker": "Search auditor", "text": "Reproduction-search set S consists of exactly run 731, and the sole failure in run 442 has test-case identifier ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Search auditor", "text": "Run 731 has exactly one failure record, and that record has test-case identifier ConfigParserTest.acceptsEmptyOverride."}, {"speaker": "Release manager", "text": "Run 442 completed its required build and test execution; the earlier interrupted attempt was not the required result."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and both contexts retain it without invented exceptions. Change 8f31, run 442, the failure test, base branch, same environment, and temporal bindings remain fixed. The two evidence spans are complete factual sentences: \"The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride.\" and \"A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S.\" The counterfactual coherently removes the S reproduction without contradicting its other observations. Neither context states a gold answer, label code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"The release record concerns change 8f31, and run 442 is its latest required run. Run 442 completed in the standard CI environment and produced exactly one failure.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride. In the 20 base-branch runs used for the historical review, this test failed exactly once. policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every one was executed in that same environment.\"},{\"speaker\":\"Release manager\",\"text\":\"A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S. The records for run 442 and S were finalized after the earlier infrastructure interruption.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["1", "text"], "text": "The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride."}, {"path": ["3", "text"], "text": "A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride.", "negative_left": "The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride.", "negative_right": "No failure record for test case ConfigParserTest.rejectsEmptyOverride appears in any run belonging to reproduction-search set S.", "right": "A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S."}, "verifier_independent_model": false}, "family": "fast-41-diverse-129-009", "id": "fast-41-diverse-129-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "The release record concerns change 8f31, and run 442 is its latest required run. Run 442 completed in the standard CI environment and produced exactly one failure."}, {"speaker": "Test automation engineer", "text": "The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride. In the 20 base-branch runs used for the historical review, this test failed exactly once. policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Build engineer", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every one was executed in that same environment."}, {"speaker": "Release manager", "text": "A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S. The records for run 442 and S were finalized after the earlier infrastructure interruption."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and both contexts retain it without invented exceptions. Change 8f31, run 442, the failure test, base branch, same environment, and temporal bindings remain fixed. The two evidence spans are complete factual sentences: \"The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride.\" and \"A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S.\" The counterfactual coherently removes the S reproduction without contradicting its other observations. Neither context states a gold answer, label code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"The release record concerns change 8f31, and run 442 is its latest required run. Run 442 completed in the standard CI environment and produced exactly one failure.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride. In the 20 base-branch runs used for the historical review, this test failed exactly once. policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every one was executed in that same environment.\"},{\"speaker\":\"Release manager\",\"text\":\"A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S. The records for run 442 and S were finalized after the earlier infrastructure interruption.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["1", "text"], "text": "The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride."}, {"path": ["3", "text"], "text": "A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride.", "negative_left": "The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride.", "negative_right": "No failure record for test case ConfigParserTest.rejectsEmptyOverride appears in any run belonging to reproduction-search set S.", "right": "A failure record for test case ConfigParserTest.rejectsEmptyOverride appears in run 517, which belongs to reproduction-search set S."}, "verifier_independent_model": false}, "family": "fast-41-diverse-129-009", "id": "fast-41-diverse-129-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "The release record concerns change 8f31, and run 442 is its latest required run. Run 442 completed in the standard CI environment and produced exactly one failure."}, {"speaker": "Test automation engineer", "text": "The sole failure recorded for run 442 is the test case ConfigParserTest.rejectsEmptyOverride. In the 20 base-branch runs used for the historical review, this test failed exactly once. policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Build engineer", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every one was executed in that same environment."}, {"speaker": "Release manager", "text": "No failure record for test case ConfigParserTest.rejectsEmptyOverride appears in any run belonging to reproduction-search set S. The records for run 442 and S were finalized after the earlier infrastructure interruption."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing readiness policy, and both contexts retain the relevant policy statement without adding exceptions. Change 8f31, run 442, the latest-run relation, and the reproduction-search path remain bound correctly. The evidence consists of two complete factual sentences: \"Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442.\" and \"Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19.\" The counterfactual coherently changes run 442's identifier to PL-64 while leaving the other observations unchanged. Neither context embeds an answer, label, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"Change 8f31 is under review, and run 442 is the latest required run. It completed in the prescribed environment and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Test analyst\",\"text\":\"ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history. Repository policy states: policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in the same environment as run 442.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442.\"},{\"speaker\":\"Test analyst\",\"text\":\"Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["3", "text"], "text": "Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442."}, {"path": ["4", "text"], "text": "Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442.", "negative_left": "Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442.", "negative_right": "Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier PL-64.", "right": "Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19."}, "verifier_independent_model": false}, "family": "fast-41-diverse-129-014", "id": "fast-41-diverse-129-014-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "Change 8f31 is under review, and run 442 is the latest required run. It completed in the prescribed environment and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Test analyst", "text": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history. Repository policy states: policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Build engineer", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in the same environment as run 442."}, {"speaker": "Build engineer", "text": "Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442."}, {"speaker": "Test analyst", "text": "Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing readiness policy, and both contexts retain the relevant policy statement without adding exceptions. Change 8f31, run 442, the latest-run relation, and the reproduction-search path remain bound correctly. The evidence consists of two complete factual sentences: \"Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442.\" and \"Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19.\" The counterfactual coherently changes run 442's identifier to PL-64 while leaving the other observations unchanged. Neither context embeds an answer, label, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicitly quantified set-coverage atoms remain atomic. The focus a7 is factual rather than a policy conclusion. The supported and refuted focus assignments are both realizable while the other atoms remain fixed: an exhaustive search set can respectively contain or lack a matching failure record without changing its membership/scope properties or the 1-of-20 history. Policy evidence correctly preserves the state-originated flaky threshold; the latest-run and same-environment reproduction requirements are already retained in the questions and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that run 442 is the latest required run and has exactly one failure. That failure is not exempt because it occurred in only 1 of the relevant 20 base-branch runs, below the stated threshold of 2. A supported a7, together with a5 and a6, establishes a matching failure on a base-branch run in the same environment, so the reproduction exception is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same sole, non-exempt failure. Refuted a7 means no run in S has a failure record for that test, while a4 makes S exhaustive over recorded same-environment base-branch runs. Thus no matching base-branch reproduction is recorded, so the negative decision follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Run 442 is the latest required run for change 8f31."}, {"id": "a2", "statement": "The failure set of run 442 contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"id": "a3", "statement": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history."}, {"id": "a4", "statement": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442."}, {"id": "a5", "statement": "Every run in reproduction-search set S is a base-branch run."}, {"id": "a6", "statement": "Every run in reproduction-search set S was executed in the same environment as run 442."}, {"id": "a7", "statement": "At least one run in reproduction-search set S has a failure record with the same test-case identifier as the sole failure in run 442."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"Change 8f31 is under review, and run 442 is the latest required run. It completed in the prescribed environment and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride.\"},{\"speaker\":\"Test analyst\",\"text\":\"ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history. Repository policy states: policy grants a flaky exemption only with at least 2 failures in those 20 runs.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in the same environment as run 442.\"},{\"speaker\":\"Build engineer\",\"text\":\"Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442.\"},{\"speaker\":\"Test analyst\",\"text\":\"Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["3", "text"], "text": "Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442."}, {"path": ["4", "text"], "text": "Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19."}], "policy_evidence": [{"path": ["2", "text"], "text": "policy grants a flaky exemption only with at least 2 failures in those 20 runs."}], "rules": [{"justification": "The latest required run has one failure, which does not meet the at-least-2-in-20 flaky threshold, but the joined failure records establish a matching base-branch reproduction in the same environment.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest required run has one failure that is below the flaky threshold, and the exhaustive same-environment base-branch search contains no matching failure record.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442.", "negative_left": "Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442.", "negative_right": "Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier PL-64.", "right": "Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier TQ-19."}, "verifier_independent_model": false}, "family": "fast-41-diverse-129-014", "id": "fast-41-diverse-129-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the latest required run contains at least one failure that neither satisfies the flaky exemption nor has a matching base-branch reproduction.", "true": "Yes — the latest required run passes, or all its failures satisfy a policy exemption or have a matching base-branch reproduction."}, "instructions": "Readiness decision: Is change 8f31 safe to merge now? Under repository policy, a change is safe only if the latest required run passes, or every failure either qualifies for the stated flaky exemption or is reproduced on the base branch in the same environment. Apply the latest temporal update rather than relying only on the initial run.", "type": "noul"}}, "state": [{"speaker": "Release coordinator", "text": "Change 8f31 is under review, and run 442 is the latest required run. It completed in the prescribed environment and its failure set contains exactly ConfigParserTest.rejectsEmptyOverride."}, {"speaker": "Test analyst", "text": "ConfigParserTest.rejectsEmptyOverride failed in exactly 1 of the 20 base-branch runs used for the flaky-exemption history. Repository policy states: policy grants a flaky exemption only with at least 2 failures in those 20 runs."}, {"speaker": "Build engineer", "text": "Reproduction-search set S contains every recorded base-branch run executed in the same environment as run 442. Every run in S is a base-branch run, and every run in S was executed in the same environment as run 442."}, {"speaker": "Build engineer", "text": "Reproduction-search set S consists exactly of base-branch run 781, which was executed in the same environment as run 442."}, {"speaker": "Test analyst", "text": "Run 781 has exactly one failure record, for test-case TQ-19, and run 442 has exactly one failure record, with identifier PL-64."}]}, "method": "c2d", "provenance": {"source_id": "diverse-129", "source_is_synthetic": true, "source_sha256": "78821996530d466d009fc1858f719b9725e8c5dd6be86a426e0daeb2e6477eb9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and matching Mira, test, input, and changed-commit bindings. The two evidence spans are complete factual sentences. The counterfactual changes only the signed-sheet total, making the expectation conflict with the requirement without creating contradictory duplicate measurements. Neither context contains an answer label, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including permissible quantified relationships over explicit execution sets. The focus atom concerns the documented total for a specific input, not routing policy. The base and counter assignments are both realizable: with a4 supported, refuting a1 in the counter entails that the prescribed value is $10.49, while the base permits $10.50; all other facts can remain fixed. Empty policy evidence is correct because the governing routing rules are already preserved in the unchanged questions object, and no additional state-originated policy is required.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish repeated application output of $10.49 where both the documented requirement and test expect $10.50. They also establish that the application change selected floor rounding, compilation and packaging did not fail, and no disruption invalidated the results. This is sufficient for application-owner routing under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Together, a1 refuted and a4 supported entail that the uniquely prescribed total among {$10.49, $10.50} is $10.49. The application repeatedly produces that compliant value while the test expects $10.50, establishing a conflicting test expectation. The remaining conditions exclude build and infrastructure failures, so routing away from the application owner is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the invoice input exercised by `InvoiceTotalTest.rounds_tax` at Mira's changed commit, the total prescribed by the documented rounding requirement is $10.50."}, {"id": "a2", "statement": "Each of the three CI executions and the local execution at Mira's changed commit observed an application-produced invoice total of $10.49 for the invoice input exercised by `InvoiceTotalTest.rounds_tax`."}, {"id": "a3", "statement": "At Mira's changed commit, `InvoiceTotalTest.rounds_tax` expects $10.50 for its invoice input."}, {"id": "a4", "statement": "For the invoice input exercised by `InvoiceTotalTest.rounds_tax`, the documented rounding requirement prescribes exactly one total from the set {$10.49, $10.50}."}, {"id": "a5", "statement": "Mira's application change set the invoice-rounding method from half-up to floor."}, {"id": "a6", "statement": "Compilation completed for each of the three CI jobs at Mira's changed commit."}, {"id": "a7", "statement": "No packaging step failed in any of the three CI jobs at Mira's changed commit."}, {"id": "a8", "statement": "No runner or service disruption prevented any of the three CI executions or the local execution at Mira's changed commit from producing a valid test result."}], "base_state_json": "\"Mira’s changed commit switched the invoice-rounding method from half-up to floor. At that commit, `InvoiceTotalTest.rounds_tax` expected $10.50, while each of three CI executions and one local execution recorded an application-produced invoice total of $10.49 for the same input. The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}. That signed sheet lists $10.50 as the total for the invoice input. Compilation completed in all three CI jobs, and no packaging step failed. Every execution produced a valid test result; no runner or service disruption prevented the CI or local run from completing. The identical assertion appeared across the CI executions and local run. The test code, dependency lockfile, and runner image were unchanged at the changed commit.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}."}, {"path": [], "text": "That signed sheet lists $10.50 as the total for the invoice input."}], "policy_evidence": [], "rules": [{"justification": "The changed application reproducibly produces $10.49 while the documented requirement and test expectation are $10.50. Completed compilation, absence of packaging failure, and valid results exclude build and infrastructure routing; the matching expectation excludes test routing.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Because the requirement prescribes exactly one of $10.49 and $10.50, refuting that it prescribes $10.50 establishes that it prescribes $10.49. The application reproducibly produces that compliant total, while the test expects $10.50, so the expectation conflicts with the requirement and the failure routes to test rather than application.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}.", "negative_left": "The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}.", "negative_right": "That signed sheet lists $10.49 as the total for the invoice input.", "right": "That signed sheet lists $10.50 as the total for the invoice input."}, "verifier_independent_model": false}, "family": "fast-41-diverse-130-013", "id": "fast-41-diverse-130-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route the failure to the test, build, or infrastructure owner instead of the application owner.", "true": "Yes — route the CI failure to the application owner because the application change reproducibly violates the documented rounding requirement."}, "instructions": "Answer yes if Chen should route the failure to the application owner. Route to application when the changed application behavior reproducibly violates the documented requirement. Route to test only if the expectation conflicts with the requirement, to build if compilation or packaging fails, and to infrastructure if runner or service disruption prevents a valid test result. A recovered warning does not override completed, reproducible test evidence.", "type": "noul"}}, "state": "Mira’s changed commit switched the invoice-rounding method from half-up to floor. At that commit, `InvoiceTotalTest.rounds_tax` expected $10.50, while each of three CI executions and one local execution recorded an application-produced invoice total of $10.49 for the same input. The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}. That signed sheet lists $10.50 as the total for the invoice input. Compilation completed in all three CI jobs, and no packaging step failed. Every execution produced a valid test result; no runner or service disruption prevented the CI or local run from completing. The identical assertion appeared across the CI executions and local run. The test code, dependency lockfile, and runner image were unchanged at the changed commit."}, "method": "c2d", "provenance": {"source_id": "diverse-130", "source_is_synthetic": true, "source_sha256": "a959e24a39b18c5d8ac39ea291d3ecca9a7848af8bf1c3f607500d52f7172ab0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and matching Mira, test, input, and changed-commit bindings. The two evidence spans are complete factual sentences. The counterfactual changes only the signed-sheet total, making the expectation conflict with the requirement without creating contradictory duplicate measurements. Neither context contains an answer label, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including permissible quantified relationships over explicit execution sets. The focus atom concerns the documented total for a specific input, not routing policy. The base and counter assignments are both realizable: with a4 supported, refuting a1 in the counter entails that the prescribed value is $10.49, while the base permits $10.50; all other facts can remain fixed. Empty policy evidence is correct because the governing routing rules are already preserved in the unchanged questions object, and no additional state-originated policy is required.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish repeated application output of $10.49 where both the documented requirement and test expect $10.50. They also establish that the application change selected floor rounding, compilation and packaging did not fail, and no disruption invalidated the results. This is sufficient for application-owner routing under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Together, a1 refuted and a4 supported entail that the uniquely prescribed total among {$10.49, $10.50} is $10.49. The application repeatedly produces that compliant value while the test expects $10.50, establishing a conflicting test expectation. The remaining conditions exclude build and infrastructure failures, so routing away from the application owner is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the invoice input exercised by `InvoiceTotalTest.rounds_tax` at Mira's changed commit, the total prescribed by the documented rounding requirement is $10.50."}, {"id": "a2", "statement": "Each of the three CI executions and the local execution at Mira's changed commit observed an application-produced invoice total of $10.49 for the invoice input exercised by `InvoiceTotalTest.rounds_tax`."}, {"id": "a3", "statement": "At Mira's changed commit, `InvoiceTotalTest.rounds_tax` expects $10.50 for its invoice input."}, {"id": "a4", "statement": "For the invoice input exercised by `InvoiceTotalTest.rounds_tax`, the documented rounding requirement prescribes exactly one total from the set {$10.49, $10.50}."}, {"id": "a5", "statement": "Mira's application change set the invoice-rounding method from half-up to floor."}, {"id": "a6", "statement": "Compilation completed for each of the three CI jobs at Mira's changed commit."}, {"id": "a7", "statement": "No packaging step failed in any of the three CI jobs at Mira's changed commit."}, {"id": "a8", "statement": "No runner or service disruption prevented any of the three CI executions or the local execution at Mira's changed commit from producing a valid test result."}], "base_state_json": "\"Mira’s changed commit switched the invoice-rounding method from half-up to floor. At that commit, `InvoiceTotalTest.rounds_tax` expected $10.50, while each of three CI executions and one local execution recorded an application-produced invoice total of $10.49 for the same input. The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}. That signed sheet lists $10.50 as the total for the invoice input. Compilation completed in all three CI jobs, and no packaging step failed. Every execution produced a valid test result; no runner or service disruption prevented the CI or local run from completing. The identical assertion appeared across the CI executions and local run. The test code, dependency lockfile, and runner image were unchanged at the changed commit.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}."}, {"path": [], "text": "That signed sheet lists $10.50 as the total for the invoice input."}], "policy_evidence": [], "rules": [{"justification": "The changed application reproducibly produces $10.49 while the documented requirement and test expectation are $10.50. Completed compilation, absence of packaging failure, and valid results exclude build and infrastructure routing; the matching expectation excludes test routing.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Because the requirement prescribes exactly one of $10.49 and $10.50, refuting that it prescribes $10.50 establishes that it prescribes $10.49. The application reproducibly produces that compliant total, while the test expects $10.50, so the expectation conflicts with the requirement and the failure routes to test rather than application.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}.", "negative_left": "The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}.", "negative_right": "That signed sheet lists $10.49 as the total for the invoice input.", "right": "That signed sheet lists $10.50 as the total for the invoice input."}, "verifier_independent_model": false}, "family": "fast-41-diverse-130-013", "id": "fast-41-diverse-130-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route the failure to the test, build, or infrastructure owner instead of the application owner.", "true": "Yes — route the CI failure to the application owner because the application change reproducibly violates the documented rounding requirement."}, "instructions": "Answer yes if Chen should route the failure to the application owner. Route to application when the changed application behavior reproducibly violates the documented requirement. Route to test only if the expectation conflicts with the requirement, to build if compilation or packaging fails, and to infrastructure if runner or service disruption prevents a valid test result. A recovered warning does not override completed, reproducible test evidence.", "type": "noul"}}, "state": "Mira’s changed commit switched the invoice-rounding method from half-up to floor. At that commit, `InvoiceTotalTest.rounds_tax` expected $10.50, while each of three CI executions and one local execution recorded an application-produced invoice total of $10.49 for the same input. The signed sheet for Mira's changed commit records the documented rounding requirement for the invoice input exercised by `InvoiceTotalTest.rounds_tax` and specifies exactly one total from {$10.49, $10.50}. That signed sheet lists $10.49 as the total for the invoice input. Compilation completed in all three CI jobs, and no packaging step failed. Every execution produced a valid test result; no runner or service disruption prevented the CI or local run from completing. The identical assertion appeared across the CI executions and local run. The test code, dependency lockfile, and runner image were unchanged at the changed commit."}, "method": "c2d", "provenance": {"source_id": "diverse-130", "source_is_synthetic": true, "source_sha256": "a959e24a39b18c5d8ac39ea291d3ecca9a7848af8bf1c3f607500d52f7172ab0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions and artifact policy are preserved, including “Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient.” The request and parser-timeout, infrastructure, and merge-routing bindings remain unchanged. The focus spans are complete factual sentences: “Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run.” and “The parser timeout being classified occurred in CI run 8841.” The counterfactual coherently places the artifact in CI run 8842 while the timeout occurred in CI run 8841, without contradictory duplicate assertions. Neither context includes a label, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including permitted universal claims over explicit evidence sets. The focus concerns whether RH-17 pertains to the classified CI run, not a policy conclusion. The base and counter assignments are jointly realizable while changing only that relationship: RH-17 can concern the classified run in the base and a different failed job in the counter. The policy evidence correctly preserves the substantive repository rule from the original state; the remaining rubric is automatically retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the submitted runner-health artifact pertains to the classified run and attributes its failure to infrastructure; all other mandatory evidence is present and directly supportive; and all material competing causes are excluded. This is sufficient for level 2.", "rule_index": 0, "sound": true}, {"reason": "RH-17 is the sole submitted runner-health artifact but is established not to pertain to the classified run. Therefore the required runner-health evidence for that run is missing, which mandates level 0 regardless of other suggestive evidence.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Runner-health artifact RH-17 pertains to the CI run containing the parser timeout being classified."}, {"id": "a2", "statement": "Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause."}, {"id": "a3", "statement": "Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17."}, {"id": "a4", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present."}, {"id": "a5", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause."}, {"id": "a6", "statement": "The available evidence rules out every material competing cause of the parser timeout being classified."}, {"id": "a7", "statement": "The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one."}], "base_state_json": "{\"context\":\"A code review is assessing whether a parser timeout can be attributed to infrastructure for merge routing. Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause. Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17. Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present. Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause. The available evidence rules out every material competing cause of the parser timeout being classified. The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one.\",\"evidence\":[\"Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run.\",\"The parser timeout being classified occurred in CI run 8841.\",\"Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient.\"],\"request\":\"Rate how completely the claimed infrastructure cause is verified for merge routing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run."}, {"path": ["evidence", "1"], "text": "The parser timeout being classified occurred in CI run 8841."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."}], "rules": [{"justification": "RH-17 is present, pertains to the run being classified, and directly attributes that run's failure to infrastructure. All other mandatory evidence is present and directly supportive, and every material competing cause is ruled out.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "RH-17 is the sole submitted runner-health artifact, but it does not pertain to the run being classified. The mandatory runner-health artifact for the relevant infrastructure attribution is therefore missing, which controls the result.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run.", "negative_left": "Runner-health artifact RH-17 was collected from CI run 8842, and from no other CI run.", "negative_right": "The parser timeout being classified occurred in CI run 8841.", "right": "The parser timeout being classified occurred in CI run 8841."}, "verifier_independent_model": false}, "family": "fast-41-diverse-131-013", "id": "fast-41-diverse-131-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"context": "A code review is assessing whether a parser timeout can be attributed to infrastructure for merge routing. Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause. Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17. Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present. Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause. The available evidence rules out every material competing cause of the parser timeout being classified. The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one.", "evidence": ["Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run.", "The parser timeout being classified occurred in CI run 8841.", "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."], "request": "Rate how completely the claimed infrastructure cause is verified for merge routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-131", "source_is_synthetic": true, "source_sha256": "69e57342498529b071bbca7c99a798ef59cafacc702cdd17ca2f15328e12e0aa", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions and artifact policy are preserved, including “Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient.” The request and parser-timeout, infrastructure, and merge-routing bindings remain unchanged. The focus spans are complete factual sentences: “Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run.” and “The parser timeout being classified occurred in CI run 8841.” The counterfactual coherently places the artifact in CI run 8842 while the timeout occurred in CI run 8841, without contradictory duplicate assertions. Neither context includes a label, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including permitted universal claims over explicit evidence sets. The focus concerns whether RH-17 pertains to the classified CI run, not a policy conclusion. The base and counter assignments are jointly realizable while changing only that relationship: RH-17 can concern the classified run in the base and a different failed job in the counter. The policy evidence correctly preserves the substantive repository rule from the original state; the remaining rubric is automatically retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the submitted runner-health artifact pertains to the classified run and attributes its failure to infrastructure; all other mandatory evidence is present and directly supportive; and all material competing causes are excluded. This is sufficient for level 2.", "rule_index": 0, "sound": true}, {"reason": "RH-17 is the sole submitted runner-health artifact but is established not to pertain to the classified run. Therefore the required runner-health evidence for that run is missing, which mandates level 0 regardless of other suggestive evidence.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Runner-health artifact RH-17 pertains to the CI run containing the parser timeout being classified."}, {"id": "a2", "statement": "Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause."}, {"id": "a3", "statement": "Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17."}, {"id": "a4", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present."}, {"id": "a5", "statement": "Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause."}, {"id": "a6", "statement": "The available evidence rules out every material competing cause of the parser timeout being classified."}, {"id": "a7", "statement": "The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one."}], "base_state_json": "{\"context\":\"A code review is assessing whether a parser timeout can be attributed to infrastructure for merge routing. Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause. Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17. Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present. Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause. The available evidence rules out every material competing cause of the parser timeout being classified. The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one.\",\"evidence\":[\"Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run.\",\"The parser timeout being classified occurred in CI run 8841.\",\"Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient.\"],\"request\":\"Rate how completely the claimed infrastructure cause is verified for merge routing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run."}, {"path": ["evidence", "1"], "text": "The parser timeout being classified occurred in CI run 8841."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."}], "rules": [{"justification": "RH-17 is present, pertains to the run being classified, and directly attributes that run's failure to infrastructure. All other mandatory evidence is present and directly supportive, and every material competing cause is ruled out.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "RH-17 is the sole submitted runner-health artifact, but it does not pertain to the run being classified. The mandatory runner-health artifact for the relevant infrastructure attribution is therefore missing, which controls the result.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Runner-health artifact RH-17 was collected from CI run 8841, and from no other CI run.", "negative_left": "Runner-health artifact RH-17 was collected from CI run 8842, and from no other CI run.", "negative_right": "The parser timeout being classified occurred in CI run 8841.", "right": "The parser timeout being classified occurred in CI run 8841."}, "verifier_independent_model": false}, "family": "fast-41-diverse-131-013", "id": "fast-41-diverse-131-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unverified: Mandatory evidence for the claimed owner is missing, or available evidence leaves a plausible competing owner; do not route or merge on the claim.", "1 — Partially verified: All mandatory evidence is present and supports the claimed owner, but at least one material ambiguity remains; route tentatively and require follow-up before merge.", "2 — Fully verified: All mandatory evidence is present, directly supports the claimed owner, and rules out material competing causes; route confidently and permit merge on this issue."], "instructions": "Use the ordered completeness rubric. Missing mandatory evidence controls the result even when other evidence is suggestive. Select exactly one level.", "type": "score"}}, "state": {"context": "A code review is assessing whether a parser timeout can be attributed to infrastructure for merge routing. Runner-health artifact RH-17 was submitted as evidence for the claimed infrastructure cause. Runner-health artifact RH-17 reports that an infrastructure malfunction caused the failed job recorded in RH-17. Every mandatory evidence item other than runner-health artifact RH-17 for the claimed infrastructure cause is present. Every mandatory evidence item other than runner-health artifact RH-17 directly supports the claimed infrastructure cause. The available evidence rules out every material competing cause of the parser timeout being classified. The number of runner-health artifacts submitted as evidence for the claimed infrastructure cause is exactly one.", "evidence": ["Runner-health artifact RH-17 was collected from CI run 8842, and from no other CI run.", "The parser timeout being classified occurred in CI run 8841.", "Repository policy requires that artifact to verify an infrastructure attribution; a passing rerun alone is not sufficient."], "request": "Rate how completely the claimed infrastructure cause is verified for merge routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-131", "source_is_synthetic": true, "source_sha256": "69e57342498529b071bbca7c99a798ef59cafacc702cdd17ca2f15328e12e0aa", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision criteria, and both contexts repeat the same policy without exceptions or output guidance. The PR, test path, and time bindings remain unchanged. The evidence consists of two complete factual sentences: \"PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time.\" \"Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731.\" The counterfactual's WF-732 assertion concerns main runs and does not contradict the current PR run's WF-731 signature. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"PR 482's current Windows cli_integration run has a test-failure status and is within PR 482's affected test scope. It is the only current in-scope required-test failure.\"},{\"speaker\":\"Test analyst\",\"text\":\"PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time.\"},{\"speaker\":\"Test analyst\",\"text\":\"Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731.\"},{\"speaker\":\"Build engineer\",\"text\":\"The current failure signature is non-assertion. PR 482 changes no public interface.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time."}, {"path": ["2", "text"], "text": "Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time.", "negative_left": "PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time.", "negative_right": "Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-732.", "right": "Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731."}, "verifier_independent_model": false}, "family": "fast-41-diverse-132-006", "id": "fast-41-diverse-132-006-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Release coordinator", "text": "PR 482's current Windows cli_integration run has a test-failure status and is within PR 482's affected test scope. It is the only current in-scope required-test failure."}, {"speaker": "Test analyst", "text": "PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time."}, {"speaker": "Test analyst", "text": "Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731."}, {"speaker": "Build engineer", "text": "The current failure signature is non-assertion. PR 482 changes no public interface."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision criteria, and both contexts repeat the same policy without exceptions or output guidance. The PR, test path, and time bindings remain unchanged. The evidence consists of two complete factual sentences: \"PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time.\" \"Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731.\" The counterfactual's WF-732 assertion concerns main runs and does not contradict the current PR run's WF-731 signature. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"PR 482's current Windows cli_integration run has a test-failure status and is within PR 482's affected test scope. It is the only current in-scope required-test failure.\"},{\"speaker\":\"Test analyst\",\"text\":\"PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time.\"},{\"speaker\":\"Test analyst\",\"text\":\"Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731.\"},{\"speaker\":\"Build engineer\",\"text\":\"The current failure signature is non-assertion. PR 482 changes no public interface.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time."}, {"path": ["2", "text"], "text": "Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time.", "negative_left": "PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time.", "negative_right": "Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-732.", "right": "Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-731."}, "verifier_independent_model": false}, "family": "fast-41-diverse-132-006", "id": "fast-41-diverse-132-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Release coordinator", "text": "PR 482's current Windows cli_integration run has a test-failure status and is within PR 482's affected test scope. It is the only current in-scope required-test failure."}, {"speaker": "Test analyst", "text": "PR 482's current Windows cli_integration run, completed at 14:00 UTC on 2026-09-17, recorded failure signature code WF-731, and at least one Windows cli_integration run on main completed before that time."}, {"speaker": "Test analyst", "text": "Every Windows cli_integration run on main completed before 14:00 UTC on 2026-09-17 recorded failure signature code WF-732."}, {"speaker": "Build engineer", "text": "The current failure signature is non-assertion. PR 482 changes no public interface."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, contain two complete factual evidence sentences, and the counterfactual consistently changes only the prior-main signature without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"PR 482's current Windows cli_integration run is marked test-failure, and cli_integration is within the affected test scope for PR 482.\"},{\"speaker\":\"Test engineer\",\"text\":\"The recorded failure is a non-assertion failure. It is the only current in-scope required-test failure for this PR; no other such failure is recorded.\"},{\"speaker\":\"Test engineer\",\"text\":\"At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184.\"},{\"speaker\":\"Test engineer\",\"text\":\"At 2026-09-16T08:00Z, a Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-184.\"},{\"speaker\":\"Release coordinator\",\"text\":\"PR 482 makes no public-interface change.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184."}, {"path": ["3", "text"], "text": "At 2026-09-16T08:00Z, a Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-184."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184.", "negative_left": "At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184.", "negative_right": "At 2026-09-16T08:00Z, the only Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-927.", "right": "At 2026-09-16T08:00Z, a Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-184."}, "verifier_independent_model": false}, "family": "fast-41-diverse-132-010", "id": "fast-41-diverse-132-010-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Release coordinator", "text": "PR 482's current Windows cli_integration run is marked test-failure, and cli_integration is within the affected test scope for PR 482."}, {"speaker": "Test engineer", "text": "The recorded failure is a non-assertion failure. It is the only current in-scope required-test failure for this PR; no other such failure is recorded."}, {"speaker": "Test engineer", "text": "At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184."}, {"speaker": "Test engineer", "text": "At 2026-09-16T08:00Z, a Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-184."}, {"speaker": "Release coordinator", "text": "PR 482 makes no public-interface change."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, contain two complete factual evidence sentences, and the counterfactual consistently changes only the prior-main signature without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"PR 482's current Windows cli_integration run is marked test-failure, and cli_integration is within the affected test scope for PR 482.\"},{\"speaker\":\"Test engineer\",\"text\":\"The recorded failure is a non-assertion failure. It is the only current in-scope required-test failure for this PR; no other such failure is recorded.\"},{\"speaker\":\"Test engineer\",\"text\":\"At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184.\"},{\"speaker\":\"Test engineer\",\"text\":\"At 2026-09-16T08:00Z, a Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-184.\"},{\"speaker\":\"Release coordinator\",\"text\":\"PR 482 makes no public-interface change.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184."}, {"path": ["3", "text"], "text": "At 2026-09-16T08:00Z, a Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-184."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184.", "negative_left": "At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184.", "negative_right": "At 2026-09-16T08:00Z, the only Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-927.", "right": "At 2026-09-16T08:00Z, a Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-184."}, "verifier_independent_model": false}, "family": "fast-41-diverse-132-010", "id": "fast-41-diverse-132-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Release coordinator", "text": "PR 482's current Windows cli_integration run is marked test-failure, and cli_integration is within the affected test scope for PR 482."}, {"speaker": "Test engineer", "text": "The recorded failure is a non-assertion failure. It is the only current in-scope required-test failure for this PR; no other such failure is recorded."}, {"speaker": "Test engineer", "text": "At 2026-09-17T10:00Z, the failure signature recorded for PR 482's current Windows cli_integration run is SIG-184."}, {"speaker": "Test engineer", "text": "At 2026-09-16T08:00Z, the only Windows cli_integration run on main completed before that PR 482 run recorded failure signature SIG-927."}, {"speaker": "Release coordinator", "text": "PR 482 makes no public-interface change."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision criteria. The PR, test, and date bindings remain consistent in both contexts. The two evidence spans are complete factual sentences: \"The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417.\" \"The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417.\" Changing the main signature to WS-902 makes the counterfactual internally coherent and removes the same-signature relationship without creating duplicate measurements. Neither context embeds a gold answer, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Test automation engineer\",\"text\":\"PR 482's current Windows cli_integration run is marked failed and is within the affected test scope. It is the sole current in-scope required-test failure; the other required checks completed successfully.\"},{\"speaker\":\"Build engineer\",\"text\":\"The recorded failure is a non-assertion event. PR 482 changes ParserCache internals only and makes no public-interface change.\"},{\"speaker\":\"Release coordinator\",\"text\":\"The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417."}, {"path": ["3", "text"], "text": "The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417.", "negative_left": "The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417.", "negative_right": "The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-902.", "right": "The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-132-012", "id": "fast-41-diverse-132-012-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Test automation engineer", "text": "PR 482's current Windows cli_integration run is marked failed and is within the affected test scope. It is the sole current in-scope required-test failure; the other required checks completed successfully."}, {"speaker": "Build engineer", "text": "The recorded failure is a non-assertion event. PR 482 changes ParserCache internals only and makes no public-interface change."}, {"speaker": "Release coordinator", "text": "The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417."}, {"speaker": "Test automation engineer", "text": "The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision criteria. The PR, test, and date bindings remain consistent in both contexts. The two evidence spans are complete factual sentences: \"The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417.\" \"The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417.\" Changing the main signature to WS-902 makes the counterfactual internally coherent and removes the same-signature relationship without creating duplicate measurements. Neither context embeds a gold answer, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Test automation engineer\",\"text\":\"PR 482's current Windows cli_integration run is marked failed and is within the affected test scope. It is the sole current in-scope required-test failure; the other required checks completed successfully.\"},{\"speaker\":\"Build engineer\",\"text\":\"The recorded failure is a non-assertion event. PR 482 changes ParserCache internals only and makes no public-interface change.\"},{\"speaker\":\"Release coordinator\",\"text\":\"The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["2", "text"], "text": "The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417."}, {"path": ["3", "text"], "text": "The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417.", "negative_left": "The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417.", "negative_right": "The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-902.", "right": "The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-132-012", "id": "fast-41-diverse-132-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Test automation engineer", "text": "PR 482's current Windows cli_integration run is marked failed and is within the affected test scope. It is the sole current in-scope required-test failure; the other required checks completed successfully."}, {"speaker": "Build engineer", "text": "The recorded failure is a non-assertion event. PR 482 changes ParserCache internals only and makes no public-interface change."}, {"speaker": "Release coordinator", "text": "The failure signature recorded for PR 482's current Windows cli_integration run on 2026-09-17 is code WS-417."}, {"speaker": "Test automation engineer", "text": "The only Windows cli_integration run on main completed before PR 482's current run on 2026-09-17 recorded failure signature code WS-902."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim, the same merge policy, and all PR, scope, test-path, and timing bindings. The two evidence quotes are complete factual sentences. The counterfactual changes only main run M-204's signature to FS-914, creating no internal contradiction or duplicate measurement conflict. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"PR 482's current Windows cli_integration run is marked test-failure and is within the affected test scope. The run's failure signature is non-assertion, and every current in-scope required-test failure for PR 482 is this one Windows run.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"Main run M-204 has failure signature FS-731.\"},{\"speaker\":\"Developer\",\"text\":\"PR 482 changes ParserCache internals only and makes no public-interface change. Linux parser checks passed, while compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204."}, {"path": ["2", "text"], "text": "Main run M-204 has failure signature FS-731."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204.", "negative_left": "PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204.", "negative_right": "Main run M-204 has failure signature FS-914.", "right": "Main run M-204 has failure signature FS-731."}, "verifier_independent_model": false}, "family": "fast-41-diverse-132-015", "id": "fast-41-diverse-132-015-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Release coordinator", "text": "PR 482's current Windows cli_integration run is marked test-failure and is within the affected test scope. The run's failure signature is non-assertion, and every current in-scope required-test failure for PR 482 is this one Windows run."}, {"speaker": "Test automation engineer", "text": "PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204."}, {"speaker": "Test automation engineer", "text": "Main run M-204 has failure signature FS-731."}, {"speaker": "Developer", "text": "PR 482 changes ParserCache internals only and makes no public-interface change. Linux parser checks passed, while compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim, the same merge policy, and all PR, scope, test-path, and timing bindings. The two evidence quotes are complete factual sentences. The counterfactual changes only main run M-204's signature to FS-914, creating no internal contradiction or duplicate measurement conflict. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Release coordinator\",\"text\":\"PR 482's current Windows cli_integration run is marked test-failure and is within the affected test scope. The run's failure signature is non-assertion, and every current in-scope required-test failure for PR 482 is this one Windows run.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"Main run M-204 has failure signature FS-731.\"},{\"speaker\":\"Developer\",\"text\":\"PR 482 changes ParserCache internals only and makes no public-interface change. Linux parser checks passed, while compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204."}, {"path": ["2", "text"], "text": "Main run M-204 has failure signature FS-731."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204.", "negative_left": "PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204.", "negative_right": "Main run M-204 has failure signature FS-914.", "right": "Main run M-204 has failure signature FS-731."}, "verifier_independent_model": false}, "family": "fast-41-diverse-132-015", "id": "fast-41-diverse-132-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Release coordinator", "text": "PR 482's current Windows cli_integration run is marked test-failure and is within the affected test scope. The run's failure signature is non-assertion, and every current in-scope required-test failure for PR 482 is this one Windows run."}, {"speaker": "Test automation engineer", "text": "PR 482's current Windows cli_integration run has failure signature FS-731, and the only Windows cli_integration run on main completed before that run is main run M-204."}, {"speaker": "Test automation engineer", "text": "Main run M-204 has failure signature FS-914."}, {"speaker": "Developer", "text": "PR 482 changes ParserCache internals only and makes no public-interface change. Linux parser checks passed, while compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged rubric, checkout-api and 14:12 bindings, contain two complete factual evidence sentences, and introduce no contradictions, answer leakage, or invented policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal bucket-rate atoms remain atomic. The focus a1 is factual rather than a policy conclusion. Both assignments are realizable with only a1 changing: one context can supply at least 10 qualifying consecutive buckets, while another can explicitly supply fewer than 10, with all supplied rates still between 5% and 20% and the latest rate at least 1%. The retained question already preserves the main rubric, and policy_evidence appropriately preserves the state-origin runbook rule concerning complete consecutive buckets and boundary handling.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 guarantees at least 10 consecutive complete buckets, while a2 and a3 place every supplied complete bucket in the inclusive-5%-to-exclusive-20% band. This is sufficient for SEV-2 and excludes SEV-1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that fewer than 10 consecutive complete buckets are supplied, and a4 excludes Noise by establishing a latest rate of at least 1%. Under the exhaustive rubric, these conditions are sufficient for Observe.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 14:12, the supplied evidence contains a sequence of at least 10 consecutive complete one-minute buckets for checkout-api."}, {"id": "a2", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5%."}, {"id": "a3", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate below 20%."}, {"id": "a4", "statement": "At 14:12, the latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}], "base_state_json": "[{\"speaker\":\"Monitoring note\",\"text\":\"At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records.\"},{\"speaker\":\"Monitoring note\",\"text\":\"For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are consecutive one-minute timestamps.\"},{\"speaker\":\"Monitoring note\",\"text\":\"Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5% and below 20%. The latest user-visible HTTP 5xx rate for checkout-api is at least 1%.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band.\"},{\"speaker\":\"Service note\",\"text\":\"Payment and database dependencies report healthy status in both regions. The only deployment noted today was a catalog-ui release at 13:40; checkout-api had no deployment or configuration change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records."}, {"path": ["1", "text"], "text": "For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are consecutive one-minute timestamps."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}], "rules": [{"justification": "At least 10 consecutive complete buckets exist, and every supplied complete bucket is in the inclusive-5%-to-exclusive-20% band required for SEV-2.", "target": "sev2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "The latest rate is at least 1%, while the evidence does not contain 10 consecutive complete buckets, which is sufficient for Observe.", "target": "observe", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records.", "negative_left": "At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records.", "negative_right": "For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are not consecutive one-minute timestamps.", "right": "For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are consecutive one-minute timestamps."}, "verifier_independent_model": false}, "family": "fast-41-diverse-134-008", "id": "fast-41-diverse-134-008-base", "input": {"questions": {"decision": {"criteria": {"noise": "Classify as Noise because the latest user-visible error rate is below 1%.", "none_of_above": "Choose None of the above because the supplied evidence cannot be classified under the stated rubric.", "observe": "Classify as Observe because at least 1% errors are present but fewer than 10 consecutive complete buckets exist.", "sev1": "Classify as SEV-1 because at least 10 consecutive complete buckets are at least 20%.", "sev2": "Classify as SEV-2 because at least 10 consecutive complete buckets are at least 5% but below 20%.", "sev3": "Classify as SEV-3 because at least 10 consecutive complete buckets are at least 1% but below 5%."}, "instructions": "Rate the alert under this exhaustive rubric: Noise if the latest rate is below 1%; Observe if it is at least 1% but fewer than 10 consecutive complete buckets exist; SEV-3 if 10 or more consecutive buckets are at least 1% but below 5%; SEV-2 if 10 or more are at least 5% but below 20%; SEV-1 if 10 or more are at least 20%; choose None of the above only if the supplied evidence cannot fit any rubric.", "type": "choice"}}, "state": [{"speaker": "Monitoring note", "text": "At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records."}, {"speaker": "Monitoring note", "text": "For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are consecutive one-minute timestamps."}, {"speaker": "Monitoring note", "text": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5% and below 20%. The latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}, {"speaker": "Incident coordinator", "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}, {"speaker": "Service note", "text": "Payment and database dependencies report healthy status in both regions. The only deployment noted today was a catalog-ui release at 13:40; checkout-api had no deployment or configuration change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-134", "source_is_synthetic": true, "source_sha256": "ad1d4e1484cdb6a3a2e5104e486414ba4c97d04018b99270162dffeb80a0d8f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sev2"}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged rubric, checkout-api and 14:12 bindings, contain two complete factual evidence sentences, and introduce no contradictions, answer leakage, or invented policy.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal bucket-rate atoms remain atomic. The focus a1 is factual rather than a policy conclusion. Both assignments are realizable with only a1 changing: one context can supply at least 10 qualifying consecutive buckets, while another can explicitly supply fewer than 10, with all supplied rates still between 5% and 20% and the latest rate at least 1%. The retained question already preserves the main rubric, and policy_evidence appropriately preserves the state-origin runbook rule concerning complete consecutive buckets and boundary handling.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 guarantees at least 10 consecutive complete buckets, while a2 and a3 place every supplied complete bucket in the inclusive-5%-to-exclusive-20% band. This is sufficient for SEV-2 and excludes SEV-1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that fewer than 10 consecutive complete buckets are supplied, and a4 excludes Noise by establishing a latest rate of at least 1%. Under the exhaustive rubric, these conditions are sufficient for Observe.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 14:12, the supplied evidence contains a sequence of at least 10 consecutive complete one-minute buckets for checkout-api."}, {"id": "a2", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5%."}, {"id": "a3", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate below 20%."}, {"id": "a4", "statement": "At 14:12, the latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}], "base_state_json": "[{\"speaker\":\"Monitoring note\",\"text\":\"At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records.\"},{\"speaker\":\"Monitoring note\",\"text\":\"For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are consecutive one-minute timestamps.\"},{\"speaker\":\"Monitoring note\",\"text\":\"Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5% and below 20%. The latest user-visible HTTP 5xx rate for checkout-api is at least 1%.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band.\"},{\"speaker\":\"Service note\",\"text\":\"Payment and database dependencies report healthy status in both regions. The only deployment noted today was a catalog-ui release at 13:40; checkout-api had no deployment or configuration change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records."}, {"path": ["1", "text"], "text": "For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are consecutive one-minute timestamps."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}], "rules": [{"justification": "At least 10 consecutive complete buckets exist, and every supplied complete bucket is in the inclusive-5%-to-exclusive-20% band required for SEV-2.", "target": "sev2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "The latest rate is at least 1%, while the evidence does not contain 10 consecutive complete buckets, which is sufficient for Observe.", "target": "observe", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records.", "negative_left": "At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records.", "negative_right": "For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are not consecutive one-minute timestamps.", "right": "For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are consecutive one-minute timestamps."}, "verifier_independent_model": false}, "family": "fast-41-diverse-134-008", "id": "fast-41-diverse-134-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"noise": "Classify as Noise because the latest user-visible error rate is below 1%.", "none_of_above": "Choose None of the above because the supplied evidence cannot be classified under the stated rubric.", "observe": "Classify as Observe because at least 1% errors are present but fewer than 10 consecutive complete buckets exist.", "sev1": "Classify as SEV-1 because at least 10 consecutive complete buckets are at least 20%.", "sev2": "Classify as SEV-2 because at least 10 consecutive complete buckets are at least 5% but below 20%.", "sev3": "Classify as SEV-3 because at least 10 consecutive complete buckets are at least 1% but below 5%."}, "instructions": "Rate the alert under this exhaustive rubric: Noise if the latest rate is below 1%; Observe if it is at least 1% but fewer than 10 consecutive complete buckets exist; SEV-3 if 10 or more consecutive buckets are at least 1% but below 5%; SEV-2 if 10 or more are at least 5% but below 20%; SEV-1 if 10 or more are at least 20%; choose None of the above only if the supplied evidence cannot fit any rubric.", "type": "choice"}}, "state": [{"speaker": "Monitoring note", "text": "At 14:12, the supplied evidence contains exactly ten complete one-minute checkout-api bucket records."}, {"speaker": "Monitoring note", "text": "For the complete one-minute checkout-api bucket records supplied at 14:12, the timestamps of all records are not consecutive one-minute timestamps."}, {"speaker": "Monitoring note", "text": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5% and below 20%. The latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}, {"speaker": "Incident coordinator", "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}, {"speaker": "Service note", "text": "Payment and database dependencies report healthy status in both regions. The only deployment noted today was a catalog-ui release at 13:40; checkout-api had no deployment or configuration change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-134", "source_is_synthetic": true, "source_sha256": "ad1d4e1484cdb6a3a2e5104e486414ba4c97d04018b99270162dffeb80a0d8f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "observe"}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric and both contexts retain its applicable runbook policy. Entity, path, and time bindings remain checkout-api at 14:12 using complete one-minute buckets. The required evidence quotes are complete factual sentences: “At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api.” and “At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets.” The counterfactual consistently changes the run count to either 8 or 9 without contradictory duplicate measurements. Neither context states a classification, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal bucket-rate atoms remain atomic. The focus a1 is factual rather than a policy conclusion. Both assignments are realizable with only a1 changing: one context can supply at least 10 qualifying consecutive buckets, while another can explicitly supply fewer than 10, with all supplied rates still between 5% and 20% and the latest rate at least 1%. The retained question already preserves the main rubric, and policy_evidence appropriately preserves the state-origin runbook rule concerning complete consecutive buckets and boundary handling.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 guarantees at least 10 consecutive complete buckets, while a2 and a3 place every supplied complete bucket in the inclusive-5%-to-exclusive-20% band. This is sufficient for SEV-2 and excludes SEV-1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that fewer than 10 consecutive complete buckets are supplied, and a4 excludes Noise by establishing a latest rate of at least 1%. Under the exhaustive rubric, these conditions are sufficient for Observe.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 14:12, the supplied evidence contains a sequence of at least 10 consecutive complete one-minute buckets for checkout-api."}, {"id": "a2", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5%."}, {"id": "a3", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate below 20%."}, {"id": "a4", "statement": "At 14:12, the latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}], "base_state_json": "[{\"speaker\":\"Monitoring record\",\"text\":\"At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api.\"},{\"speaker\":\"Monitoring record\",\"text\":\"At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets.\"},{\"speaker\":\"Metrics analyst\",\"text\":\"Every complete one-minute checkout-api bucket in the supplied evidence at 14:12 has a user-visible HTTP 5xx rate between 5.0% and 5.2%, and the latest rate is 5.2%. No complete bucket reaches 20%.\"},{\"speaker\":\"Service owner\",\"text\":\"Checkout-api had no deployment or configuration change today, and payment and database dependencies report healthy status in both regions.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api."}, {"path": ["1", "text"], "text": "At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}], "rules": [{"justification": "At least 10 consecutive complete buckets exist, and every supplied complete bucket is in the inclusive-5%-to-exclusive-20% band required for SEV-2.", "target": "sev2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "The latest rate is at least 1%, while the evidence does not contain 10 consecutive complete buckets, which is sufficient for Observe.", "target": "observe", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api.", "negative_left": "At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api.", "negative_right": "At 14:12, the evidence records that run as containing either 8 or 9 complete one-minute buckets.", "right": "At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets."}, "verifier_independent_model": false}, "family": "fast-41-diverse-134-011", "id": "fast-41-diverse-134-011-base", "input": {"questions": {"decision": {"criteria": {"noise": "Classify as Noise because the latest user-visible error rate is below 1%.", "none_of_above": "Choose None of the above because the supplied evidence cannot be classified under the stated rubric.", "observe": "Classify as Observe because at least 1% errors are present but fewer than 10 consecutive complete buckets exist.", "sev1": "Classify as SEV-1 because at least 10 consecutive complete buckets are at least 20%.", "sev2": "Classify as SEV-2 because at least 10 consecutive complete buckets are at least 5% but below 20%.", "sev3": "Classify as SEV-3 because at least 10 consecutive complete buckets are at least 1% but below 5%."}, "instructions": "Rate the alert under this exhaustive rubric: Noise if the latest rate is below 1%; Observe if it is at least 1% but fewer than 10 consecutive complete buckets exist; SEV-3 if 10 or more consecutive buckets are at least 1% but below 5%; SEV-2 if 10 or more are at least 5% but below 20%; SEV-1 if 10 or more are at least 20%; choose None of the above only if the supplied evidence cannot fit any rubric.", "type": "choice"}}, "state": [{"speaker": "Monitoring record", "text": "At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api."}, {"speaker": "Monitoring record", "text": "At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets."}, {"speaker": "Metrics analyst", "text": "Every complete one-minute checkout-api bucket in the supplied evidence at 14:12 has a user-visible HTTP 5xx rate between 5.0% and 5.2%, and the latest rate is 5.2%. No complete bucket reaches 20%."}, {"speaker": "Service owner", "text": "Checkout-api had no deployment or configuration change today, and payment and database dependencies report healthy status in both regions."}, {"speaker": "Incident coordinator", "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}]}, "method": "c2d", "provenance": {"source_id": "diverse-134", "source_is_synthetic": true, "source_sha256": "ad1d4e1484cdb6a3a2e5104e486414ba4c97d04018b99270162dffeb80a0d8f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sev2"}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric and both contexts retain its applicable runbook policy. Entity, path, and time bindings remain checkout-api at 14:12 using complete one-minute buckets. The required evidence quotes are complete factual sentences: “At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api.” and “At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets.” The counterfactual consistently changes the run count to either 8 or 9 without contradictory duplicate measurements. Neither context states a classification, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal bucket-rate atoms remain atomic. The focus a1 is factual rather than a policy conclusion. Both assignments are realizable with only a1 changing: one context can supply at least 10 qualifying consecutive buckets, while another can explicitly supply fewer than 10, with all supplied rates still between 5% and 20% and the latest rate at least 1%. The retained question already preserves the main rubric, and policy_evidence appropriately preserves the state-origin runbook rule concerning complete consecutive buckets and boundary handling.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 guarantees at least 10 consecutive complete buckets, while a2 and a3 place every supplied complete bucket in the inclusive-5%-to-exclusive-20% band. This is sufficient for SEV-2 and excludes SEV-1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that fewer than 10 consecutive complete buckets are supplied, and a4 excludes Noise by establishing a latest rate of at least 1%. Under the exhaustive rubric, these conditions are sufficient for Observe.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 14:12, the supplied evidence contains a sequence of at least 10 consecutive complete one-minute buckets for checkout-api."}, {"id": "a2", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate of at least 5%."}, {"id": "a3", "statement": "Every complete one-minute checkout-api bucket in the evidence supplied at 14:12 has a user-visible HTTP 5xx rate below 20%."}, {"id": "a4", "statement": "At 14:12, the latest user-visible HTTP 5xx rate for checkout-api is at least 1%."}], "base_state_json": "[{\"speaker\":\"Monitoring record\",\"text\":\"At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api.\"},{\"speaker\":\"Monitoring record\",\"text\":\"At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets.\"},{\"speaker\":\"Metrics analyst\",\"text\":\"Every complete one-minute checkout-api bucket in the supplied evidence at 14:12 has a user-visible HTTP 5xx rate between 5.0% and 5.2%, and the latest rate is 5.2%. No complete bucket reaches 20%.\"},{\"speaker\":\"Service owner\",\"text\":\"Checkout-api had no deployment or configuration change today, and payment and database dependencies report healthy status in both regions.\"},{\"speaker\":\"Incident coordinator\",\"text\":\"Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api."}, {"path": ["1", "text"], "text": "At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}], "rules": [{"justification": "At least 10 consecutive complete buckets exist, and every supplied complete bucket is in the inclusive-5%-to-exclusive-20% band required for SEV-2.", "target": "sev2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "The latest rate is at least 1%, while the evidence does not contain 10 consecutive complete buckets, which is sufficient for Observe.", "target": "observe", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api.", "negative_left": "At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api.", "negative_right": "At 14:12, the evidence records that run as containing either 8 or 9 complete one-minute buckets.", "right": "At 14:12, the evidence records that run as containing either 10 or 11 complete one-minute buckets."}, "verifier_independent_model": false}, "family": "fast-41-diverse-134-011", "id": "fast-41-diverse-134-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"noise": "Classify as Noise because the latest user-visible error rate is below 1%.", "none_of_above": "Choose None of the above because the supplied evidence cannot be classified under the stated rubric.", "observe": "Classify as Observe because at least 1% errors are present but fewer than 10 consecutive complete buckets exist.", "sev1": "Classify as SEV-1 because at least 10 consecutive complete buckets are at least 20%.", "sev2": "Classify as SEV-2 because at least 10 consecutive complete buckets are at least 5% but below 20%.", "sev3": "Classify as SEV-3 because at least 10 consecutive complete buckets are at least 1% but below 5%."}, "instructions": "Rate the alert under this exhaustive rubric: Noise if the latest rate is below 1%; Observe if it is at least 1% but fewer than 10 consecutive complete buckets exist; SEV-3 if 10 or more consecutive buckets are at least 1% but below 5%; SEV-2 if 10 or more are at least 5% but below 20%; SEV-1 if 10 or more are at least 20%; choose None of the above only if the supplied evidence cannot fit any rubric.", "type": "choice"}}, "state": [{"speaker": "Monitoring record", "text": "At 14:12, the supplied evidence contains exactly one consecutive run of complete one-minute buckets for checkout-api."}, {"speaker": "Monitoring record", "text": "At 14:12, the evidence records that run as containing either 8 or 9 complete one-minute buckets."}, {"speaker": "Metrics analyst", "text": "Every complete one-minute checkout-api bucket in the supplied evidence at 14:12 has a user-visible HTTP 5xx rate between 5.0% and 5.2%, and the latest rate is 5.2%. No complete bucket reaches 20%."}, {"speaker": "Service owner", "text": "Checkout-api had no deployment or configuration change today, and payment and database dependencies report healthy status in both regions."}, {"speaker": "Incident coordinator", "text": "Runbook: classify using consecutive complete buckets; a value exactly on a boundary enters the higher severity band."}]}, "method": "c2d", "provenance": {"source_id": "diverse-134", "source_is_synthetic": true, "source_sha256": "ad1d4e1484cdb6a3a2e5104e486414ba4c97d04018b99270162dffeb80a0d8f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "observe"}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy, request scope, entity bindings, and time references. Each evidence array has two complete factual sentences, and the Billing change is coherent without contradiction or an embedded answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"case_note\":\"Trace T1 has finished processing the alert and contains a single slow-component record. The identifier in that record is the one used for the service-registry lookup. The registry snapshot is dated 2026-09-17 and contains a single service name for that identifier. During the same interval, the Payments dependency dashboard is available but shows no degradation confirmation. The alert concerns elevated ShopWave checkout latency, and no deployment occurred for Checkout or Payments during the preceding six hours. The routing review uses only the completed trace and the dashboard record, without inferring missing evidence.\",\"evidence\":[\"The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17.\",\"The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Payments.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "0"], "text": "The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17."}, {"path": ["evidence", "1"], "text": "The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Payments."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17.", "negative_left": "The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17.", "negative_right": "The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Billing.", "right": "The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Payments."}, "verifier_independent_model": false}, "family": "fast-41-diverse-136-002", "id": "fast-41-diverse-136-002-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"case_note": "Trace T1 has finished processing the alert and contains a single slow-component record. The identifier in that record is the one used for the service-registry lookup. The registry snapshot is dated 2026-09-17 and contains a single service name for that identifier. During the same interval, the Payments dependency dashboard is available but shows no degradation confirmation. The alert concerns elevated ShopWave checkout latency, and no deployment occurred for Checkout or Payments during the preceding six hours. The routing review uses only the completed trace and the dashboard record, without inferring missing evidence.", "context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17.", "The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Payments."]}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy, request scope, entity bindings, and time references. Each evidence array has two complete factual sentences, and the Billing change is coherent without contradiction or an embedded answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"case_note\":\"Trace T1 has finished processing the alert and contains a single slow-component record. The identifier in that record is the one used for the service-registry lookup. The registry snapshot is dated 2026-09-17 and contains a single service name for that identifier. During the same interval, the Payments dependency dashboard is available but shows no degradation confirmation. The alert concerns elevated ShopWave checkout latency, and no deployment occurred for Checkout or Payments during the preceding six hours. The routing review uses only the completed trace and the dashboard record, without inferring missing evidence.\",\"evidence\":[\"The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17.\",\"The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Payments.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "0"], "text": "The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17."}, {"path": ["evidence", "1"], "text": "The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Payments."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17.", "negative_left": "The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17.", "negative_right": "The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Billing.", "right": "The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Payments."}, "verifier_independent_model": false}, "family": "fast-41-diverse-136-002", "id": "fast-41-diverse-136-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"case_note": "Trace T1 has finished processing the alert and contains a single slow-component record. The identifier in that record is the one used for the service-registry lookup. The registry snapshot is dated 2026-09-17 and contains a single service name for that identifier. During the same interval, the Payments dependency dashboard is available but shows no degradation confirmation. The alert concerns elevated ShopWave checkout latency, and no deployment occurred for Checkout or Payments during the preceding six hours. The routing review uses only the completed trace and the dashboard record, without inferring missing evidence.", "context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["The completed Trace T1 for the ShopWave production alert records its sole slow component under service identifier SC-17.", "The 2026-09-17 service-registry entry for the identifier recorded in Trace T1 lists exactly one service name: Billing."]}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question policy, scope, entity, path, and time bindings; each has two complete factual evidence sentences, and the counterfactual coherently changes only the catalog classification from Payments to Fraud without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation. The assessment concerns the production alert during the interval beginning at 14:05 UTC. Trace T1 finished processing and reports exactly one slow component. Its component record identifies service SC-17. The dependency dashboard contains no confirming observation for Payments degradation during that interval. The registry and catalog entries below are the dated records used to interpret the component identifier.\",\"evidence\":[\"On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service.\",\"The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Payments service.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "0"], "text": "On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service."}, {"path": ["evidence", "1"], "text": "The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service.", "negative_left": "On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service.", "negative_right": "The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Fraud service.", "right": "The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Payments service."}, "verifier_independent_model": false}, "family": "fast-41-diverse-136-010", "id": "fast-41-diverse-136-010-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation. The assessment concerns the production alert during the interval beginning at 14:05 UTC. Trace T1 finished processing and reports exactly one slow component. Its component record identifies service SC-17. The dependency dashboard contains no confirming observation for Payments degradation during that interval. The registry and catalog entries below are the dated records used to interpret the component identifier.", "evidence": ["On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service.", "The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Payments service."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question policy, scope, entity, path, and time bindings; each has two complete factual evidence sentences, and the counterfactual coherently changes only the catalog classification from Payments to Fraud without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation. The assessment concerns the production alert during the interval beginning at 14:05 UTC. Trace T1 finished processing and reports exactly one slow component. Its component record identifies service SC-17. The dependency dashboard contains no confirming observation for Payments degradation during that interval. The registry and catalog entries below are the dated records used to interpret the component identifier.\",\"evidence\":[\"On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service.\",\"The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Payments service.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "0"], "text": "On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service."}, {"path": ["evidence", "1"], "text": "The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service.", "negative_left": "On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service.", "negative_right": "The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Fraud service.", "right": "The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Payments service."}, "verifier_independent_model": false}, "family": "fast-41-diverse-136-010", "id": "fast-41-diverse-136-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation. The assessment concerns the production alert during the interval beginning at 14:05 UTC. Trace T1 finished processing and reports exactly one slow component. Its component record identifies service SC-17. The dependency dashboard contains no confirming observation for Payments degradation during that interval. The registry and catalog entries below are the dated records used to interpret the component identifier.", "evidence": ["On 2026-09-17, the ShopWave service registry maps service identifier SC-17 to service code PAY-042, and PAY-042 is not assigned to any other service.", "The 2026-09-17 canonical ShopWave service catalog lists service code PAY-042 as the Fraud service."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Base context is coherent: \"The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.\" \"The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service.\" Counterfactual context coherently changes the code's identified service to Inventory without changing the policy, request bindings, or embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"Trace T1 for the ShopWave production alert is complete and identifies exactly one slow component.\",\"The component named by Trace T1 has service identifier SC-17.\",\"The dependency dashboard has no confirming data for Payments degradation during the relevant alert interval.\",\"The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.\",\"The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service.\",\"The alert remains under review while the routing record is reconciled with the trace and registry.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services."}, {"path": ["evidence", "4"], "text": "The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.", "negative_left": "The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.", "negative_right": "The same ShopWave service registry records that secondary code SVC-BLUE identifies the Inventory service.", "right": "The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service."}, "verifier_independent_model": false}, "family": "fast-41-diverse-136-016", "id": "fast-41-diverse-136-016-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for the ShopWave production alert is complete and identifies exactly one slow component.", "The component named by Trace T1 has service identifier SC-17.", "The dependency dashboard has no confirming data for Payments degradation during the relevant alert interval.", "The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.", "The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service.", "The alert remains under review while the routing record is reconciled with the trace and registry."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Base context is coherent: \"The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.\" \"The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service.\" Counterfactual context coherently changes the code's identified service to Inventory without changing the policy, request bindings, or embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"Trace T1 for the ShopWave production alert is complete and identifies exactly one slow component.\",\"The component named by Trace T1 has service identifier SC-17.\",\"The dependency dashboard has no confirming data for Payments degradation during the relevant alert interval.\",\"The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.\",\"The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service.\",\"The alert remains under review while the routing record is reconciled with the trace and registry.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services."}, {"path": ["evidence", "4"], "text": "The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.", "negative_left": "The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.", "negative_right": "The same ShopWave service registry records that secondary code SVC-BLUE identifies the Inventory service.", "right": "The same ShopWave service registry records that secondary code SVC-BLUE identifies the Payments service."}, "verifier_independent_model": false}, "family": "fast-41-diverse-136-016", "id": "fast-41-diverse-136-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for the ShopWave production alert is complete and identifies exactly one slow component.", "The component named by Trace T1 has service identifier SC-17.", "The dependency dashboard has no confirming data for Payments degradation during the relevant alert interval.", "The ShopWave service registry dated 2026-09-17 records that SC-17 is the identifier of the service carrying unique secondary code SVC-BLUE, and that the Payments service and Inventory service are different services.", "The same ShopWave service registry records that secondary code SVC-BLUE identifies the Inventory service.", "The alert remains under review while the routing record is reconciled with the trace and registry."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing scope, thresholds, escalation rule, and 10-minute window. Both contexts retain the required region, traffic, and time bindings. The evidence consists of two complete factual sentences: “During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window.” and “During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window.” The counterfactual’s less-than-5% assertion for other regions is consistent with eu-west reaching at least 5%, a 12% maximum, and no region reaching 25%. Neither context embeds an answer, label rationale, proposition identifier, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual regional-failure relationship; A2 is a permissible universally quantified relation and A3 is a single region-count relation. A3 is factual rather than a policy conclusion. The base and counter assignments are realizable while changing only A3: the base can have two regions between 5% and 25%, while the counter can have only eu-west in that range, with no region at 25% or above in either case. The policy evidence correctly preserves the relevant state-originating scope, exclusions, threshold inclusivity, and escalation binding; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 alone entails the level-3 condition that at least 5% fail in two or more regions for the full 10-minute window. No exclusion of lower outcomes is needed because the ordered rubric classifies this directly as major broad impact.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes at least one qualifying region at 5% or more for 10 minutes. Refuted A3 entails that fewer than two regions qualify, so eu-west is exactly one qualifying region. A2 excludes every region from reaching the competing 25% level-3 threshold. Together these conditions entail level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During 14:20–14:30 UTC, at least 5% of in-scope production customer requests failed in eu-west for the full 10-minute window."}, {"id": "A2", "statement": "During 14:20–14:30 UTC, no region had at least 25% of in-scope production customer requests fail for the full 10-minute window."}, {"id": "A3", "statement": "During 14:20–14:30 UTC, at least two distinct regions each had at least 5% of in-scope production customer requests fail for the full 10-minute window."}], "base_state_json": "[{\"speaker\":\"Incident record\",\"text\":\"The audit covers only in-scope production customer requests during 14:20–14:30 UTC. The highest regional failure rate observed in the window was 12%, and the eu-west deployment was active throughout.\"},{\"speaker\":\"Counterfactual comparison\",\"text\":\"During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window.\"},{\"speaker\":\"Observed incident\",\"text\":\"During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window.\"},{\"speaker\":\"Runbook\",\"text\":\"Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window."}, {"path": ["2", "text"], "text": "During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}], "rules": [{"justification": "At least two regions each meeting the inclusive 5% threshold for the full 10-minute window is sufficient for major broad impact.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Eu-west supplies at least one region meeting 5% for 10 minutes; refutation of A3 limits the count of such regions to one, and A2 excludes the competing level-3 condition of any region reaching 25%.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window.", "negative_left": "During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window.", "negative_right": "During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had less than 5% of in-scope production customer requests fail for the full 10-minute window.", "right": "During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window."}, "verifier_independent_model": false}, "family": "fast-41-diverse-137-014", "id": "fast-41-diverse-137-014-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified production impact: only excluded synthetic or beta traffic is affected, or in-scope production failures remain below 1% in every region during the window.", "1 — Minor verified impact: at least 1% but less than 5% of in-scope production requests fail in one region for 10 minutes, with no other region at 1% or higher.", "2 — Significant regional impact: at least 5% but less than 25% of in-scope production requests fail in exactly one region for at least 10 minutes; escalate to the owning service team.", "3 — Major broad impact: at least 25% of in-scope production requests fail in any region for at least 10 minutes, or at least 5% fail in two or more regions for that duration."], "instructions": "Assign the alert an impact level using the ordered rubric. Assess only the stated 10-minute window and apply the runbook’s scope and exceptions.", "type": "score"}}, "state": [{"speaker": "Incident record", "text": "The audit covers only in-scope production customer requests during 14:20–14:30 UTC. The highest regional failure rate observed in the window was 12%, and the eu-west deployment was active throughout."}, {"speaker": "Counterfactual comparison", "text": "During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window."}, {"speaker": "Observed incident", "text": "During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window."}, {"speaker": "Runbook", "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}]}, "method": "c2d", "provenance": {"source_id": "diverse-137", "source_is_synthetic": true, "source_sha256": "6ad980dc01b48352a342f771b5dfe60d2cbfd6e9d0710110e93512959974536f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing scope, thresholds, escalation rule, and 10-minute window. Both contexts retain the required region, traffic, and time bindings. The evidence consists of two complete factual sentences: “During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window.” and “During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window.” The counterfactual’s less-than-5% assertion for other regions is consistent with eu-west reaching at least 5%, a 12% maximum, and no region reaching 25%. Neither context embeds an answer, label rationale, proposition identifier, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual regional-failure relationship; A2 is a permissible universally quantified relation and A3 is a single region-count relation. A3 is factual rather than a policy conclusion. The base and counter assignments are realizable while changing only A3: the base can have two regions between 5% and 25%, while the counter can have only eu-west in that range, with no region at 25% or above in either case. The policy evidence correctly preserves the relevant state-originating scope, exclusions, threshold inclusivity, and escalation binding; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 alone entails the level-3 condition that at least 5% fail in two or more regions for the full 10-minute window. No exclusion of lower outcomes is needed because the ordered rubric classifies this directly as major broad impact.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes at least one qualifying region at 5% or more for 10 minutes. Refuted A3 entails that fewer than two regions qualify, so eu-west is exactly one qualifying region. A2 excludes every region from reaching the competing 25% level-3 threshold. Together these conditions entail level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During 14:20–14:30 UTC, at least 5% of in-scope production customer requests failed in eu-west for the full 10-minute window."}, {"id": "A2", "statement": "During 14:20–14:30 UTC, no region had at least 25% of in-scope production customer requests fail for the full 10-minute window."}, {"id": "A3", "statement": "During 14:20–14:30 UTC, at least two distinct regions each had at least 5% of in-scope production customer requests fail for the full 10-minute window."}], "base_state_json": "[{\"speaker\":\"Incident record\",\"text\":\"The audit covers only in-scope production customer requests during 14:20–14:30 UTC. The highest regional failure rate observed in the window was 12%, and the eu-west deployment was active throughout.\"},{\"speaker\":\"Counterfactual comparison\",\"text\":\"During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window.\"},{\"speaker\":\"Observed incident\",\"text\":\"During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window.\"},{\"speaker\":\"Runbook\",\"text\":\"Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window."}, {"path": ["2", "text"], "text": "During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window."}], "policy_evidence": [{"path": ["3", "text"], "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}], "rules": [{"justification": "At least two regions each meeting the inclusive 5% threshold for the full 10-minute window is sufficient for major broad impact.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Eu-west supplies at least one region meeting 5% for 10 minutes; refutation of A3 limits the count of such regions to one, and A2 excludes the competing level-3 condition of any region reaching 25%.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window.", "negative_left": "During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window.", "negative_right": "During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had less than 5% of in-scope production customer requests fail for the full 10-minute window.", "right": "During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had at least 5% of in-scope production customer requests fail for the full 10-minute window."}, "verifier_independent_model": false}, "family": "fast-41-diverse-137-014", "id": "fast-41-diverse-137-014-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified production impact: only excluded synthetic or beta traffic is affected, or in-scope production failures remain below 1% in every region during the window.", "1 — Minor verified impact: at least 1% but less than 5% of in-scope production requests fail in one region for 10 minutes, with no other region at 1% or higher.", "2 — Significant regional impact: at least 5% but less than 25% of in-scope production requests fail in exactly one region for at least 10 minutes; escalate to the owning service team.", "3 — Major broad impact: at least 25% of in-scope production requests fail in any region for at least 10 minutes, or at least 5% fail in two or more regions for that duration."], "instructions": "Assign the alert an impact level using the ordered rubric. Assess only the stated 10-minute window and apply the runbook’s scope and exceptions.", "type": "score"}}, "state": [{"speaker": "Incident record", "text": "The audit covers only in-scope production customer requests during 14:20–14:30 UTC. The highest regional failure rate observed in the window was 12%, and the eu-west deployment was active throughout."}, {"speaker": "Counterfactual comparison", "text": "During 14:20–14:30 UTC, eu-west had at least 5% of in-scope production customer requests fail for the full 10-minute window, and no region had at least 25% fail for the full 10-minute window."}, {"speaker": "Observed incident", "text": "During 14:20–14:30 UTC, every region other than eu-west, including ap-southeast, had less than 5% of in-scope production customer requests fail for the full 10-minute window."}, {"speaker": "Runbook", "text": "Runbook: use production customer traffic only. Synthetic and beta traffic never count. A threshold written “at least” is inclusive; regional incidents meeting it are escalated to the service owner."}]}, "method": "c2d", "provenance": {"source_id": "diverse-137", "source_is_synthetic": true, "source_sha256": "6ad980dc01b48352a342f771b5dfe60d2cbfd6e9d0710110e93512959974536f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the relevant Checkout API, purchase-completion, and 14:20 UTC bindings. The two evidence spans are complete factual sentences: \"During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint.\" and \"During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z.\" The counterfactual coherently relocates affected customers to disjoint Region Y, making Procedure Q unavailable without contradicting other measurements. Neither context contains an explicit answer, code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom concerns the factual availability of an alternative purchase procedure rather than policy. The base and counter assignments differ only on that focus and are both realizable: the service owner's determination can remain accurate while changing consistently with whether a workaround exists. Policy evidence preserves the only substantive state-originated runbook rule needed for interpretation; criteria and instructions in the questions object need not and must not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than 20% of requests fail and the ordinary purchase path is blocked, but an alternative completion procedure exists. Together with the accurate service-owner determination, this establishes a confirmed workaround. That satisfies level 2's degraded-purchases-with-workaround branch and excludes level 3, which requires no workaround.", "rule_index": 0, "sound": true}, {"reason": "More than 20% of customer-facing requests fail, the ordinary purchase path is blocked, and refutation of the existential workaround atom entails that no alternative purchase-completion procedure exists. These conditions are sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, more than 20% of customer-facing Checkout API requests failed."}, {"id": "a2", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Checkout API failures blocked the ordinary purchase-completion path for affected customers."}, {"id": "a3", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, at least one alternative procedure was available to affected customers for completing purchases despite the Checkout API failures."}, {"id": "a4", "statement": "The Checkout service owner's determination of whether any alternative purchase-completion procedure was available to affected customers during the 14-minute Checkout API alert interval beginning at 14:20 UTC was accurate."}], "base_state_json": "\"At 14:20 UTC, the Checkout API entered a 14-minute customer-facing alert interval. During that interval, 42% of customer-facing Checkout API requests failed, exceeding 20%, and the ordinary purchase-completion path was blocked for affected customers. The service owner reviewed the determination concerning purchase-completion procedures for that interval and recorded it accurately. During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint. During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z. No Checkout deployment occurred that day, and the last release was three days earlier. Dependency monitoring recorded PaymentGateway latency above 12 seconds while other dependencies remained healthy. The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint."}, {"path": [], "text": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z."}], "policy_evidence": [{"path": [], "text": "The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation."}], "rules": [{"justification": "The ordinary purchase path is blocked, but an alternative purchase-completion procedure is available. Because the service owner's workaround-availability determination is accurate, customers have a confirmed workaround, satisfying level 2 and excluding level 3's no-workaround requirement.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "More than 20% of customer-facing requests fail, the ordinary purchase-completion path is blocked, and refutation of the existential workaround atom establishes that no alternative purchase-completion procedure is available. These facts satisfy level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint.", "negative_left": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint.", "negative_right": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Y.", "right": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z."}, "verifier_independent_model": false}, "family": "fast-41-diverse-138-017", "id": "fast-41-diverse-138-017-base", "input": {"questions": {"decision": {"criteria": ["0 — No incident: the alert is false or expected, customer-facing error rates remain at baseline, and completed purchases are unaffected.", "1 — Minor operational issue: a real fault exists, but it is confined to internal processing or affects fewer than 5% of customer requests without blocking purchases.", "2 — Moderate user impact: 5% to 20% of customer requests fail, or purchases are degraded but customers have a confirmed workaround; service-team response is required without major-incident escalation.", "3 — Major user impact: more than 20% of customer-facing requests fail and purchases are blocked without a workaround; notify the incident coordinator and immediately escalate to the responsible service or dependency team."], "instructions": "Assess the alert's user-impact level using the evidence and runbook. Select exactly one ordered level, resolving references such as “the former” and “the latter” to the named services.", "type": "score"}}, "state": "At 14:20 UTC, the Checkout API entered a 14-minute customer-facing alert interval. During that interval, 42% of customer-facing Checkout API requests failed, exceeding 20%, and the ordinary purchase-completion path was blocked for affected customers. The service owner reviewed the determination concerning purchase-completion procedures for that interval and recorded it accurately. During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint. During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z. No Checkout deployment occurred that day, and the last release was three days earlier. Dependency monitoring recorded PaymentGateway latency above 12 seconds while other dependencies remained healthy. The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation."}, "method": "c2d", "provenance": {"source_id": "diverse-138", "source_is_synthetic": true, "source_sha256": "d4407ec5f95dddf32844f514d87988ba4151765246a49398f43049b05769c0c4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the relevant Checkout API, purchase-completion, and 14:20 UTC bindings. The two evidence spans are complete factual sentences: \"During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint.\" and \"During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z.\" The counterfactual coherently relocates affected customers to disjoint Region Y, making Procedure Q unavailable without contradicting other measurements. Neither context contains an explicit answer, code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom concerns the factual availability of an alternative purchase procedure rather than policy. The base and counter assignments differ only on that focus and are both realizable: the service owner's determination can remain accurate while changing consistently with whether a workaround exists. Policy evidence preserves the only substantive state-originated runbook rule needed for interpretation; criteria and instructions in the questions object need not and must not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than 20% of requests fail and the ordinary purchase path is blocked, but an alternative completion procedure exists. Together with the accurate service-owner determination, this establishes a confirmed workaround. That satisfies level 2's degraded-purchases-with-workaround branch and excludes level 3, which requires no workaround.", "rule_index": 0, "sound": true}, {"reason": "More than 20% of customer-facing requests fail, the ordinary purchase path is blocked, and refutation of the existential workaround atom entails that no alternative purchase-completion procedure exists. These conditions are sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, more than 20% of customer-facing Checkout API requests failed."}, {"id": "a2", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Checkout API failures blocked the ordinary purchase-completion path for affected customers."}, {"id": "a3", "statement": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, at least one alternative procedure was available to affected customers for completing purchases despite the Checkout API failures."}, {"id": "a4", "statement": "The Checkout service owner's determination of whether any alternative purchase-completion procedure was available to affected customers during the 14-minute Checkout API alert interval beginning at 14:20 UTC was accurate."}], "base_state_json": "\"At 14:20 UTC, the Checkout API entered a 14-minute customer-facing alert interval. During that interval, 42% of customer-facing Checkout API requests failed, exceeding 20%, and the ordinary purchase-completion path was blocked for affected customers. The service owner reviewed the determination concerning purchase-completion procedures for that interval and recorded it accurately. During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint. During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z. No Checkout deployment occurred that day, and the last release was three days earlier. Dependency monitoring recorded PaymentGateway latency above 12 seconds while other dependencies remained healthy. The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint."}, {"path": [], "text": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z."}], "policy_evidence": [{"path": [], "text": "The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation."}], "rules": [{"justification": "The ordinary purchase path is blocked, but an alternative purchase-completion procedure is available. Because the service owner's workaround-availability determination is accurate, customers have a confirmed workaround, satisfying level 2 and excluding level 3's no-workaround requirement.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "More than 20% of customer-facing requests fail, the ordinary purchase-completion path is blocked, and refutation of the existential workaround atom establishes that no alternative purchase-completion procedure is available. These facts satisfy level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint.", "negative_left": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint.", "negative_right": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Y.", "right": "During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Z."}, "verifier_independent_model": false}, "family": "fast-41-diverse-138-017", "id": "fast-41-diverse-138-017-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No incident: the alert is false or expected, customer-facing error rates remain at baseline, and completed purchases are unaffected.", "1 — Minor operational issue: a real fault exists, but it is confined to internal processing or affects fewer than 5% of customer requests without blocking purchases.", "2 — Moderate user impact: 5% to 20% of customer requests fail, or purchases are degraded but customers have a confirmed workaround; service-team response is required without major-incident escalation.", "3 — Major user impact: more than 20% of customer-facing requests fail and purchases are blocked without a workaround; notify the incident coordinator and immediately escalate to the responsible service or dependency team."], "instructions": "Assess the alert's user-impact level using the evidence and runbook. Select exactly one ordered level, resolving references such as “the former” and “the latter” to the named services.", "type": "score"}}, "state": "At 14:20 UTC, the Checkout API entered a 14-minute customer-facing alert interval. During that interval, 42% of customer-facing Checkout API requests failed, exceeding 20%, and the ordinary purchase-completion path was blocked for affected customers. The service owner reviewed the determination concerning purchase-completion procedures for that interval and recorded it accurately. During the 14-minute Checkout API alert interval beginning at 14:20 UTC, Procedure Q was the only alternative purchase-completion procedure, and it was available only to customers in Region Z; Region Z and Region Y were disjoint. During the 14-minute Checkout API alert interval beginning at 14:20 UTC, all customers affected by the Checkout API failures were located in Region Y. No Checkout deployment occurred that day, and the last release was three days earlier. Dependency monitoring recorded PaymentGateway latency above 12 seconds while other dependencies remained healthy. The runbook says a customer-facing failure rate above 20% with blocked purchases is a major incident requiring the incident coordinator and immediate dependency-team escalation."}, "method": "c2d", "provenance": {"source_id": "diverse-138", "source_is_synthetic": true, "source_sha256": "d4407ec5f95dddf32844f514d87988ba4151765246a49398f43049b05769c0c4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions and both contexts retain the governing rules. Question bindings remain batch C-184, the 14:20 report, and the readiness disposition. The evidence consists of two complete factual sentences: “The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2.” and “The immutable SHA-256 digest of batch C-184 is 6A91F3C2.” The counterfactual coherently changes the immutable digest while retaining the report’s recorded digest, creating a mismatch without duplicate measurements. Neither context embeds a gold answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Quality analyst\",\"text\":\"The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2.\"},{\"speaker\":\"Quality analyst\",\"text\":\"The immutable SHA-256 digest of batch C-184 is 6A91F3C2.\"},{\"speaker\":\"Records coordinator\",\"text\":\"The 14:20 report is the newest submission for the readiness disposition of batch C-184, and it is marked completed.\"},{\"speaker\":\"Validation lead\",\"text\":\"That report records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%.\"},{\"speaker\":\"Pipeline auditor\",\"text\":\"Every transformation used to generate the 14:20 report is supported; no unsupported coercion was found.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2."}, {"path": ["1", "text"], "text": "The immutable SHA-256 digest of batch C-184 is 6A91F3C2."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2.", "negative_left": "The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2.", "negative_right": "The immutable SHA-256 digest of batch C-184 is 8D47B1E9.", "right": "The immutable SHA-256 digest of batch C-184 is 6A91F3C2."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-001", "id": "fast-41-diverse-139-001-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Quality analyst", "text": "The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2."}, {"speaker": "Quality analyst", "text": "The immutable SHA-256 digest of batch C-184 is 6A91F3C2."}, {"speaker": "Records coordinator", "text": "The 14:20 report is the newest submission for the readiness disposition of batch C-184, and it is marked completed."}, {"speaker": "Validation lead", "text": "That report records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%."}, {"speaker": "Pipeline auditor", "text": "Every transformation used to generate the 14:20 report is supported; no unsupported coercion was found."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions and both contexts retain the governing rules. Question bindings remain batch C-184, the 14:20 report, and the readiness disposition. The evidence consists of two complete factual sentences: “The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2.” and “The immutable SHA-256 digest of batch C-184 is 6A91F3C2.” The counterfactual coherently changes the immutable digest while retaining the report’s recorded digest, creating a mismatch without duplicate measurements. Neither context embeds a gold answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Quality analyst\",\"text\":\"The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2.\"},{\"speaker\":\"Quality analyst\",\"text\":\"The immutable SHA-256 digest of batch C-184 is 6A91F3C2.\"},{\"speaker\":\"Records coordinator\",\"text\":\"The 14:20 report is the newest submission for the readiness disposition of batch C-184, and it is marked completed.\"},{\"speaker\":\"Validation lead\",\"text\":\"That report records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%.\"},{\"speaker\":\"Pipeline auditor\",\"text\":\"Every transformation used to generate the 14:20 report is supported; no unsupported coercion was found.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2."}, {"path": ["1", "text"], "text": "The immutable SHA-256 digest of batch C-184 is 6A91F3C2."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2.", "negative_left": "The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2.", "negative_right": "The immutable SHA-256 digest of batch C-184 is 8D47B1E9.", "right": "The immutable SHA-256 digest of batch C-184 is 6A91F3C2."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-001", "id": "fast-41-diverse-139-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Quality analyst", "text": "The newest report submitted for the readiness disposition of batch C-184 at 14:20 records the SHA-256 digest 6A91F3C2."}, {"speaker": "Quality analyst", "text": "The immutable SHA-256 digest of batch C-184 is 8D47B1E9."}, {"speaker": "Records coordinator", "text": "The 14:20 report is the newest submission for the readiness disposition of batch C-184, and it is marked completed."}, {"speaker": "Validation lead", "text": "That report records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%."}, {"speaker": "Pipeline auditor", "text": "Every transformation used to generate the 14:20 report is supported; no unsupported coercion was found."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy, bindings, complete evidence sentences, and coherent hash relationships without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Readiness analyst\",\"text\":\"The readiness review concerns batch C-184. The newest submission for its disposition is timestamped 14:20, and the report is completed.\"},{\"speaker\":\"Verification log\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614.\"},{\"speaker\":\"Verification log\",\"text\":\"The immutable SHA-256 hash of batch C-184 has the value 7f3a91c2e8b64d105ab7c9f0d3e2a614.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The 14:20 report records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%. The audit log marks every transformation used to generate it as supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}, {"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 has the value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614.", "negative_right": "The immutable SHA-256 hash of batch C-184 has the value 4c8e27b1d965fa3027e41c8b0a6d593f.", "right": "The immutable SHA-256 hash of batch C-184 has the value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-005", "id": "fast-41-diverse-139-005-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Readiness analyst", "text": "The readiness review concerns batch C-184. The newest submission for its disposition is timestamped 14:20, and the report is completed."}, {"speaker": "Verification log", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}, {"speaker": "Verification log", "text": "The immutable SHA-256 hash of batch C-184 has the value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}, {"speaker": "Quality reviewer", "text": "The 14:20 report records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%. The audit log marks every transformation used to generate it as supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy, bindings, complete evidence sentences, and coherent hash relationships without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Readiness analyst\",\"text\":\"The readiness review concerns batch C-184. The newest submission for its disposition is timestamped 14:20, and the report is completed.\"},{\"speaker\":\"Verification log\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614.\"},{\"speaker\":\"Verification log\",\"text\":\"The immutable SHA-256 hash of batch C-184 has the value 7f3a91c2e8b64d105ab7c9f0d3e2a614.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"The 14:20 report records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%. The audit log marks every transformation used to generate it as supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}, {"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 has the value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614.", "negative_right": "The immutable SHA-256 hash of batch C-184 has the value 4c8e27b1d965fa3027e41c8b0a6d593f.", "right": "The immutable SHA-256 hash of batch C-184 has the value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-005", "id": "fast-41-diverse-139-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Readiness analyst", "text": "The readiness review concerns batch C-184. The newest submission for its disposition is timestamped 14:20, and the report is completed."}, {"speaker": "Verification log", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 7f3a91c2e8b64d105ab7c9f0d3e2a614."}, {"speaker": "Verification log", "text": "The immutable SHA-256 hash of batch C-184 has the value 4c8e27b1d965fa3027e41c8b0a6d593f."}, {"speaker": "Quality reviewer", "text": "The 14:20 report records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%. The audit log marks every transformation used to generate it as supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy and scoring criteria. Both contexts retain the batch, report time, hash, and disposition bindings. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes the immutable hash and creates a report-to-batch mismatch without contradictory duplicate measurements. Neither context embeds an answer, label rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31.\"},{\"speaker\":\"Records clerk\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31.\"},{\"speaker\":\"Readiness analyst\",\"text\":\"The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184, and it is completed. Its required-field completeness is exactly 98.0%, its duplicate rate is 0.9%, and its schema-error rate is 0.0%. The audit records every transformation used to generate it as supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}, {"path": ["1", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31.", "negative_left": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 4b82e6d1a9037c5f29b8e047d6a1c93f.", "right": "The immutable SHA-256 hash of batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-007", "id": "fast-41-diverse-139-007-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Records clerk", "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}, {"speaker": "Records clerk", "text": "The immutable SHA-256 hash of batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}, {"speaker": "Readiness analyst", "text": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184, and it is completed. Its required-field completeness is exactly 98.0%, its duplicate rate is 0.9%, and its schema-error rate is 0.0%. The audit records every transformation used to generate it as supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy and scoring criteria. Both contexts retain the batch, report time, hash, and disposition bindings. The focus evidence contains exactly two complete factual sentences. The counterfactual coherently changes the immutable hash and creates a report-to-batch mismatch without contradictory duplicate measurements. Neither context embeds an answer, label rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31.\"},{\"speaker\":\"Records clerk\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31.\"},{\"speaker\":\"Readiness analyst\",\"text\":\"The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184, and it is completed. Its required-field completeness is exactly 98.0%, its duplicate rate is 0.9%, and its schema-error rate is 0.0%. The audit records every transformation used to generate it as supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}, {"path": ["1", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31.", "negative_left": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 4b82e6d1a9037c5f29b8e047d6a1c93f.", "right": "The immutable SHA-256 hash of batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-007", "id": "fast-41-diverse-139-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Records clerk", "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a91c2d8e640b5a17c9d04e6f28a31."}, {"speaker": "Records clerk", "text": "The immutable SHA-256 hash of batch C-184 is 4b82e6d1a9037c5f29b8e047d6a1c93f."}, {"speaker": "Readiness analyst", "text": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184, and it is completed. Its required-field completeness is exactly 98.0%, its duplicate rate is 0.9%, and its schema-error rate is 0.0%. The audit records every transformation used to generate it as supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, and both contexts retain the governing policy. The batch, report time, and evidence paths remain bound to C-184 and the 14:20 report. The two focus spans are complete factual sentences: \"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.\" \"The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.\" The counterfactual coherently represents a report hash that differs from the immutable batch hash, which the preserved policy explicitly handles. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.\"},{\"speaker\":\"Readiness reviewer\",\"text\":\"The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184, and it is completed.\"},{\"speaker\":\"Readiness reviewer\",\"text\":\"That report measures required-field completeness at exactly 98.0%, duplicate rate at 0.9%, and schema-error rate at 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"Every transformation used to generate the newest 14:20 report for batch C-184 is supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}, {"path": ["1", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 93b6e1a4d8c2f0b7e5a9c3d1f6b8e2a0c4d7f9b1e3a5c8d6f2b0a4e7c9d1f3.", "right": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-008", "id": "fast-41-diverse-139-008-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}, {"speaker": "Data quality analyst", "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}, {"speaker": "Readiness reviewer", "text": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184, and it is completed."}, {"speaker": "Readiness reviewer", "text": "That report measures required-field completeness at exactly 98.0%, duplicate rate at 0.9%, and schema-error rate at 0.0%."}, {"speaker": "Data engineer", "text": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, and both contexts retain the governing policy. The batch, report time, and evidence paths remain bound to C-184 and the 14:20 report. The two focus spans are complete factual sentences: \"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.\" \"The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.\" The counterfactual coherently represents a report hash that differs from the immutable batch hash, which the preserved policy explicitly handles. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.\"},{\"speaker\":\"Readiness reviewer\",\"text\":\"The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184, and it is completed.\"},{\"speaker\":\"Readiness reviewer\",\"text\":\"That report measures required-field completeness at exactly 98.0%, duplicate rate at 0.9%, and schema-error rate at 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"Every transformation used to generate the newest 14:20 report for batch C-184 is supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}, {"path": ["1", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 93b6e1a4d8c2f0b7e5a9c3d1f6b8e2a0c4d7f9b1e3a5c8d6f2b0a4e7c9d1f3.", "right": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-008", "id": "fast-41-diverse-139-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash 7f3a9c2d1e8b4a6f0c5d7e9b2a1f3c8d6e4b0a9c7d5f2e1a8b6c4d3e2f1a0b9."}, {"speaker": "Data quality analyst", "text": "The immutable SHA-256 hash of batch C-184 is 93b6e1a4d8c2f0b7e5a9c3d1f6b8e2a0c4d7f9b1e3a5c8d6f2b0a4e7c9d1f3."}, {"speaker": "Readiness reviewer", "text": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184, and it is completed."}, {"speaker": "Readiness reviewer", "text": "That report measures required-field completeness at exactly 98.0%, duplicate rate at 0.9%, and schema-error rate at 0.0%."}, {"speaker": "Data engineer", "text": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain batch C-184 and the 14:20 report binding. The two evidence spans are complete factual sentences: \"The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa.\" \"The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa.\" The counterfactual coherently changes the immutable hash and thereby represents a report-hash mismatch covered by the policy. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Readiness reviewer\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa.\"},{\"speaker\":\"Readiness reviewer\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa.\"},{\"speaker\":\"Quality analyst\",\"text\":\"The 14:20 report is the newest report submitted for the readiness disposition of batch C-184 and is completed. It records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%.\"},{\"speaker\":\"Pipeline auditor\",\"text\":\"Every transformation used to generate the newest 14:20 report for batch C-184 is supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa."}, {"path": ["1", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 9c26f4a8e173bd506d2190f3c8475ab1.", "right": "The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-010", "id": "fast-41-diverse-139-010-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Readiness reviewer", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa."}, {"speaker": "Readiness reviewer", "text": "The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa."}, {"speaker": "Quality analyst", "text": "The 14:20 report is the newest report submitted for the readiness disposition of batch C-184 and is completed. It records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%."}, {"speaker": "Pipeline auditor", "text": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain batch C-184 and the 14:20 report binding. The two evidence spans are complete factual sentences: \"The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa.\" \"The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa.\" The counterfactual coherently changes the immutable hash and thereby represents a report-hash mismatch covered by the policy. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Readiness reviewer\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa.\"},{\"speaker\":\"Readiness reviewer\",\"text\":\"The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa.\"},{\"speaker\":\"Quality analyst\",\"text\":\"The 14:20 report is the newest report submitted for the readiness disposition of batch C-184 and is completed. It records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%.\"},{\"speaker\":\"Pipeline auditor\",\"text\":\"Every transformation used to generate the newest 14:20 report for batch C-184 is supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa."}, {"path": ["1", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 9c26f4a8e173bd506d2190f3c8475ab1.", "right": "The immutable SHA-256 hash of batch C-184 is 4b7e3a91d2f6c805a1149e2b6d7380fa."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-010", "id": "fast-41-diverse-139-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Readiness reviewer", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 value 4b7e3a91d2f6c805a1149e2b6d7380fa."}, {"speaker": "Readiness reviewer", "text": "The immutable SHA-256 hash of batch C-184 is 9c26f4a8e173bd506d2190f3c8475ab1."}, {"speaker": "Quality analyst", "text": "The 14:20 report is the newest report submitted for the readiness disposition of batch C-184 and is completed. It records required-field completeness of exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%."}, {"speaker": "Pipeline auditor", "text": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full policy, bindings, precedence, and criteria; both contexts remain scoped to batch C-184 and its 14:20 readiness report, while the counterfactual coherently changes the registry hash to create a policy-relevant mismatch rather than a duplicate measurement. Required evidence quotes: “The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.” “The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.\"},{\"speaker\":\"Hash registry clerk\",\"text\":\"The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.\"},{\"speaker\":\"Records clerk\",\"text\":\"The 14:20 report is the newest submission for the readiness disposition of batch C-184 and is marked completed.\"},{\"speaker\":\"Quality analyst\",\"text\":\"The report measures required-field completeness at exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"The transformation audit confirms that every transformation used to generate the 14:20 report is supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}, {"path": ["1", "text"], "text": "The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.", "negative_left": "The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.", "negative_right": "The immutable SHA-256 value stored for batch C-184 is 2c8e74a1b6935d0f47e219ac83b760de14f5a9268c3b1d7e50a4f692b8c01357.", "right": "The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-012", "id": "fast-41-diverse-139-012-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Records clerk", "text": "The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}, {"speaker": "Hash registry clerk", "text": "The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}, {"speaker": "Records clerk", "text": "The 14:20 report is the newest submission for the readiness disposition of batch C-184 and is marked completed."}, {"speaker": "Quality analyst", "text": "The report measures required-field completeness at exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%."}, {"speaker": "Audit reviewer", "text": "The transformation audit confirms that every transformation used to generate the 14:20 report is supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full policy, bindings, precedence, and criteria; both contexts remain scoped to batch C-184 and its 14:20 readiness report, while the counterfactual coherently changes the registry hash to create a policy-relevant mismatch rather than a duplicate measurement. Required evidence quotes: “The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.” “The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.\"},{\"speaker\":\"Hash registry clerk\",\"text\":\"The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.\"},{\"speaker\":\"Records clerk\",\"text\":\"The 14:20 report is the newest submission for the readiness disposition of batch C-184 and is marked completed.\"},{\"speaker\":\"Quality analyst\",\"text\":\"The report measures required-field completeness at exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"The transformation audit confirms that every transformation used to generate the 14:20 report is supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}, {"path": ["1", "text"], "text": "The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.", "negative_left": "The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a.", "negative_right": "The immutable SHA-256 value stored for batch C-184 is 2c8e74a1b6935d0f47e219ac83b760de14f5a9268c3b1d7e50a4f692b8c01357.", "right": "The immutable SHA-256 value stored for batch C-184 is 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-012", "id": "fast-41-diverse-139-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Records clerk", "text": "The newest report submitted for the readiness disposition of batch C-184, timestamped 14:20, records SHA-256 value 7f3a91c0d5e8246b19aa40f7c2d68311e5b9a07c4f6d2e8a31b5c9d7e0f2468a."}, {"speaker": "Hash registry clerk", "text": "The immutable SHA-256 value stored for batch C-184 is 2c8e74a1b6935d0f47e219ac83b760de14f5a9268c3b1d7e50a4f692b8c01357."}, {"speaker": "Records clerk", "text": "The 14:20 report is the newest submission for the readiness disposition of batch C-184 and is marked completed."}, {"speaker": "Quality analyst", "text": "The report measures required-field completeness at exactly 98.0%, a duplicate rate of 0.9%, and a schema-error rate of 0.0%."}, {"speaker": "Audit reviewer", "text": "The transformation audit confirms that every transformation used to generate the 14:20 report is supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policy, including thresholds, precedence, and the identical-hash rule. Both preserve the batch C-184 and 14:20 report bindings. The two evidence spans are complete factual sentences: “The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.” and “The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.” The counterfactual coherently changes the immutable hash, creating a meaningful mismatch with the report hash without duplicate measurements. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"Batch C-184 is under readiness review. The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2. The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"The 14:20 report is the newest report submitted for the readiness disposition of batch C-184, and it is completed.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Its required-field completeness is exactly 98.0%, its duplicate rate is 0.9%, and its schema-error rate is 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The audit records every transformation used to generate the 14:20 report as supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2."}, {"path": ["0", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.", "negative_left": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 4c8e1a6d9f2b7e5c0a3d8f1b6e4c9a2d7f5b0e3c8a6d1f9b4e2c7a5d0f8b3.", "right": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-013", "id": "fast-41-diverse-139-013-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Case note", "text": "Batch C-184 is under readiness review. The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2. The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2."}, {"speaker": "Data quality analyst", "text": "The 14:20 report is the newest report submitted for the readiness disposition of batch C-184, and it is completed."}, {"speaker": "Data quality analyst", "text": "Its required-field completeness is exactly 98.0%, its duplicate rate is 0.9%, and its schema-error rate is 0.0%."}, {"speaker": "Data engineer", "text": "The audit records every transformation used to generate the 14:20 report as supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policy, including thresholds, precedence, and the identical-hash rule. Both preserve the batch C-184 and 14:20 report bindings. The two evidence spans are complete factual sentences: “The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.” and “The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.” The counterfactual coherently changes the immutable hash, creating a meaningful mismatch with the report hash without duplicate measurements. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"Batch C-184 is under readiness review. The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2. The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"The 14:20 report is the newest report submitted for the readiness disposition of batch C-184, and it is completed.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Its required-field completeness is exactly 98.0%, its duplicate rate is 0.9%, and its schema-error rate is 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The audit records every transformation used to generate the 14:20 report as supported.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2."}, {"path": ["0", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.", "negative_left": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 4c8e1a6d9f2b7e5c0a3d8f1b6e4c9a2d7f5b0e3c8a6d1f9b4e2c7a5d0f8b3.", "right": "The immutable SHA-256 hash of batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2."}, "verifier_independent_model": false}, "family": "fast-41-diverse-139-013", "id": "fast-41-diverse-139-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Case note", "text": "Batch C-184 is under readiness review. The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 7f3a9c2d8e1b4a6f0c5d7e9b2a4c6f8d1e3b5a7c9d0f2e4a6b8c1d3e5f7a9b2. The immutable SHA-256 hash of batch C-184 is 4c8e1a6d9f2b7e5c0a3d8f1b6e4c9a2d7f5b0e3c8a6d1f9b4e2c7a5d0f8b3."}, {"speaker": "Data quality analyst", "text": "The 14:20 report is the newest report submitted for the readiness disposition of batch C-184, and it is completed."}, {"speaker": "Data quality analyst", "text": "Its required-field completeness is exactly 98.0%, its duplicate rate is 0.9%, and its schema-error rate is 0.0%."}, {"speaker": "Data engineer", "text": "The audit records every transformation used to generate the 14:20 report as supported."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the acceptance requirements and documented history-record exception. The batch binding remains B-142, with no changed question entity or time/path binding. The focus evidence contains two complete factual sentences: \"In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history.\" and \"In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value.\" The counterfactual consistently changes the completeness count to 19,900 without contradictory duplicate measurements. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Quality auditor\",\"text\":\"The post-transformation audit for batch B-142 recorded zero duplicate customer_id values.\"},{\"speaker\":\"Scope reviewer\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Scope reviewer\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Quality auditor\",\"text\":\"In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history.\"},{\"speaker\":\"Quality auditor\",\"text\":\"In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history."}, {"path": ["4", "text"], "text": "In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history.", "negative_left": "In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history.", "negative_right": "In that same snapshot, 19,900 of those batch B-142 records contain a non-empty ship_country value.", "right": "In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value."}, "verifier_independent_model": false}, "family": "fast-41-diverse-142-005", "id": "fast-41-diverse-142-005-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Quality auditor", "text": "The post-transformation audit for batch B-142 recorded zero duplicate customer_id values."}, {"speaker": "Scope reviewer", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Scope reviewer", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Quality auditor", "text": "In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history."}, {"speaker": "Quality auditor", "text": "In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the acceptance requirements and documented history-record exception. The batch binding remains B-142, with no changed question entity or time/path binding. The focus evidence contains two complete factual sentences: \"In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history.\" and \"In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value.\" The counterfactual consistently changes the completeness count to 19,900 without contradictory duplicate measurements. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Quality auditor\",\"text\":\"The post-transformation audit for batch B-142 recorded zero duplicate customer_id values.\"},{\"speaker\":\"Scope reviewer\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Scope reviewer\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Quality auditor\",\"text\":\"In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history.\"},{\"speaker\":\"Quality auditor\",\"text\":\"In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history."}, {"path": ["4", "text"], "text": "In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history.", "negative_left": "In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history.", "negative_right": "In that same snapshot, 19,900 of those batch B-142 records contain a non-empty ship_country value.", "right": "In that same snapshot, 19,999 of those batch B-142 records contain a non-empty ship_country value."}, "verifier_independent_model": false}, "family": "fast-41-diverse-142-005", "id": "fast-41-diverse-142-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Quality auditor", "text": "The post-transformation audit for batch B-142 recorded zero duplicate customer_id values."}, {"speaker": "Scope reviewer", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Scope reviewer", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Quality auditor", "text": "In the 2026-09-16 quality snapshot, batch B-142 contains exactly 20,000 records whose import_mode is not history."}, {"speaker": "Quality auditor", "text": "In that same snapshot, 19,900 of those batch B-142 records contain a non-empty ship_country value."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete acceptance policy, and both contexts retain it without additions or omissions. Batch B-142 and the acceptance decision remain correctly bound. The evidence consists of exactly two complete factual sentences: \"In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records.\" and \"The audit found ship_country present in 9,997 of 10,000 included records in batch B-142.\" The counterfactual changes only the completeness measurement to 9,994 of 10,000, which is internally coherent and introduces no contradictory duplicate assertion. Neither context contains an explicit answer, answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"The audit found ship_country present in 9,997 of 10,000 included records in batch B-142.\"},{\"speaker\":\"Data engineer\",\"text\":\"The batch contains 10,000 customer records. The approved transformation retained the newest record for each repeated customer_id; its final report lists no duplicate customer_id values. The validation log shows no remaining schema errors, and the source export confirms that the archived records carry the history tag.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["1", "text"], "text": "In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records."}, {"path": ["2", "text"], "text": "The audit found ship_country present in 9,997 of 10,000 included records in batch B-142."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records.", "negative_left": "In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records.", "negative_right": "The audit found ship_country present in 9,994 of 10,000 included records in batch B-142.", "right": "The audit found ship_country present in 9,997 of 10,000 included records in batch B-142."}, "verifier_independent_model": false}, "family": "fast-41-diverse-142-009", "id": "fast-41-diverse-142-009-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Audit reviewer", "text": "In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records."}, {"speaker": "Audit reviewer", "text": "The audit found ship_country present in 9,997 of 10,000 included records in batch B-142."}, {"speaker": "Data engineer", "text": "The batch contains 10,000 customer records. The approved transformation retained the newest record for each repeated customer_id; its final report lists no duplicate customer_id values. The validation log shows no remaining schema errors, and the source export confirms that the archived records carry the history tag."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete acceptance policy, and both contexts retain it without additions or omissions. Batch B-142 and the acceptance decision remain correctly bound. The evidence consists of exactly two complete factual sentences: \"In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records.\" and \"The audit found ship_country present in 9,997 of 10,000 included records in batch B-142.\" The counterfactual changes only the completeness measurement to 9,994 of 10,000, which is internally coherent and introduces no contradictory duplicate assertion. Neither context contains an explicit answer, answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records.\"},{\"speaker\":\"Audit reviewer\",\"text\":\"The audit found ship_country present in 9,997 of 10,000 included records in batch B-142.\"},{\"speaker\":\"Data engineer\",\"text\":\"The batch contains 10,000 customer records. The approved transformation retained the newest record for each repeated customer_id; its final report lists no duplicate customer_id values. The validation log shows no remaining schema errors, and the source export confirms that the archived records carry the history tag.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["1", "text"], "text": "In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records."}, {"path": ["2", "text"], "text": "The audit found ship_country present in 9,997 of 10,000 included records in batch B-142."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records.", "negative_left": "In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records.", "negative_right": "The audit found ship_country present in 9,994 of 10,000 included records in batch B-142.", "right": "The audit found ship_country present in 9,997 of 10,000 included records in batch B-142."}, "verifier_independent_model": false}, "family": "fast-41-diverse-142-009", "id": "fast-41-diverse-142-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Audit reviewer", "text": "In batch B-142, the audit scope excludes every record tagged import_mode=history and includes all other records."}, {"speaker": "Audit reviewer", "text": "The audit found ship_country present in 9,994 of 10,000 included records in batch B-142."}, {"speaker": "Data engineer", "text": "The batch contains 10,000 customer records. The approved transformation retained the newest record for each repeated customer_id; its final report lists no duplicate customer_id values. The validation log shows no remaining schema errors, and the source export confirms that the archived records carry the history tag."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, bindings, criteria, and instructions; both contexts retain the governing scope rules and exact evidence, while the counterfactual changes only the completeness count consistently.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142 was processed with the approved deduplication step, which retains the newest updated_at record for each customer_id. The post-processing audit found no repeated customer_id values.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The source export marks every archived history record with import_mode=history, so those records are outside the product-search population. The batch review found no other validation or schema errors.\"},{\"speaker\":\"Audit log\",\"text\":\"At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history.\"},{\"speaker\":\"Audit log\",\"text\":\"At 2026-09-17 10:05 UTC, 1,999 of those records had a recorded ship_country value.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["4", "text"], "text": "At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history."}, {"path": ["5", "text"], "text": "At 2026-09-17 10:05 UTC, 1,999 of those records had a recorded ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history.", "negative_left": "At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history.", "negative_right": "At 2026-09-17 10:05 UTC, 1,998 of those records had a recorded ship_country value.", "right": "At 2026-09-17 10:05 UTC, 1,999 of those records had a recorded ship_country value."}, "verifier_independent_model": false}, "family": "fast-41-diverse-142-018", "id": "fast-41-diverse-142-018-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 was processed with the approved deduplication step, which retains the newest updated_at record for each customer_id. The post-processing audit found no repeated customer_id values."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "The source export marks every archived history record with import_mode=history, so those records are outside the product-search population. The batch review found no other validation or schema errors."}, {"speaker": "Audit log", "text": "At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history."}, {"speaker": "Audit log", "text": "At 2026-09-17 10:05 UTC, 1,999 of those records had a recorded ship_country value."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, bindings, criteria, and instructions; both contexts retain the governing scope rules and exact evidence, while the counterfactual changes only the completeness count consistently.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142 was processed with the approved deduplication step, which retains the newest updated_at record for each customer_id. The post-processing audit found no repeated customer_id values.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The source export marks every archived history record with import_mode=history, so those records are outside the product-search population. The batch review found no other validation or schema errors.\"},{\"speaker\":\"Audit log\",\"text\":\"At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history.\"},{\"speaker\":\"Audit log\",\"text\":\"At 2026-09-17 10:05 UTC, 1,999 of those records had a recorded ship_country value.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["4", "text"], "text": "At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history."}, {"path": ["5", "text"], "text": "At 2026-09-17 10:05 UTC, 1,999 of those records had a recorded ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history.", "negative_left": "At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history.", "negative_right": "At 2026-09-17 10:05 UTC, 1,998 of those records had a recorded ship_country value.", "right": "At 2026-09-17 10:05 UTC, 1,999 of those records had a recorded ship_country value."}, "verifier_independent_model": false}, "family": "fast-41-diverse-142-018", "id": "fast-41-diverse-142-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 was processed with the approved deduplication step, which retains the newest updated_at record for each customer_id. The post-processing audit found no repeated customer_id values."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "The source export marks every archived history record with import_mode=history, so those records are outside the product-search population. The batch review found no other validation or schema errors."}, {"speaker": "Audit log", "text": "At 2026-09-17 10:00 UTC, batch B-142 contained 2,000 records without the tag import_mode=history."}, {"speaker": "Audit log", "text": "At 2026-09-17 10:05 UTC, 1,998 of those records had a recorded ship_country value."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, scope, and bindings. Both contexts retain the transformation rule without stating a classification answer. The two evidence spans are complete factual sentences. The counterfactual coherently changes one postal-country record without creating contradictory duplicate assertions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A data-quality case note records the following. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63. Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-63. The validation summary records 12 blank required cells in total. Exactly six of those blanks are in region_code; the remaining six are distributed among customer_name and email. No other required-cell or transformation facts were reported.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63."}, {"path": [], "text": "Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-63."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63.", "negative_left": "Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63.", "negative_right": "Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-64.", "right": "Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-63."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-001", "id": "fast-41-diverse-143-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A data-quality case note records the following. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63. Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-63. The validation summary records 12 blank required cells in total. Exactly six of those blanks are in region_code; the remaining six are distributed among customer_name and email. No other required-cell or transformation facts were reported."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, scope, and bindings. Both contexts retain the transformation rule without stating a classification answer. The two evidence spans are complete factual sentences. The counterfactual coherently changes one postal-country record without creating contradictory duplicate assertions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A data-quality case note records the following. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63. Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-63. The validation summary records 12 blank required cells in total. Exactly six of those blanks are in region_code; the remaining six are distributed among customer_name and email. No other required-cell or transformation facts were reported.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63."}, {"path": [], "text": "Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-63."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63.", "negative_left": "Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63.", "negative_right": "Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-64.", "right": "Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-63."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-001", "id": "fast-41-diverse-143-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A data-quality case note records the following. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before transformation, the six blank required region_code cells have record identifiers C-17, C-24, C-31, C-42, C-58, and C-63. Before transformation, present postal_country values are attached to record identifiers C-17, C-24, C-31, C-42, C-58, and C-64. The validation summary records 12 blank required cells in total. Exactly six of those blanks are in region_code; the remaining six are distributed among customer_name and email. No other required-cell or transformation facts were reported."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing derivation rule and the completeness task remain unchanged. The customer-batch entity and post-transformation assessment scope remain bound to the question. The evidence contains exactly two complete factual sentences: \"Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6.\" and \"Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6.\" The counterfactual coherently changes only postal_country presence for R6, leaving counts and transformation consequences consistent. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A validation analyst reviews a customer batch with 300 required cells before transformation. The pre-transformation report identifies exactly 12 blank required cells, of which exactly six are region_code cells; the other six blanks are in different required fields. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6. Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6. The remaining blank cells receive no transformation, and no other validation adjustment is recorded. The analyst must assess completeness after applying the stated rule.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6."}, {"path": [], "text": "Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6.", "negative_left": "Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6.", "negative_right": "Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, and R5, but not R6.", "right": "Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-004", "id": "fast-41-diverse-143-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A validation analyst reviews a customer batch with 300 required cells before transformation. The pre-transformation report identifies exactly 12 blank required cells, of which exactly six are region_code cells; the other six blanks are in different required fields. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6. Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6. The remaining blank cells receive no transformation, and no other validation adjustment is recorded. The analyst must assess completeness after applying the stated rule."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing derivation rule and the completeness task remain unchanged. The customer-batch entity and post-transformation assessment scope remain bound to the question. The evidence contains exactly two complete factual sentences: \"Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6.\" and \"Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6.\" The counterfactual coherently changes only postal_country presence for R6, leaving counts and transformation consequences consistent. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A validation analyst reviews a customer batch with 300 required cells before transformation. The pre-transformation report identifies exactly 12 blank required cells, of which exactly six are region_code cells; the other six blanks are in different required fields. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6. Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6. The remaining blank cells receive no transformation, and no other validation adjustment is recorded. The analyst must assess completeness after applying the stated rule.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6."}, {"path": [], "text": "Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6.", "negative_left": "Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6.", "negative_right": "Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, and R5, but not R6.", "right": "Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, R5, and R6."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-004", "id": "fast-41-diverse-143-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A validation analyst reviews a customer batch with 300 required cells before transformation. The pre-transformation report identifies exactly 12 blank required cells, of which exactly six are region_code cells; the other six blanks are in different required fields. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before the stated transformation, the six blank required region_code cells are attached to record identifiers R1, R2, R3, R4, R5, and R6. Before the stated transformation, present postal_country values are attached to record identifiers R1, R2, R3, R4, and R5, but not R6. The remaining blank cells receive no transformation, and no other validation adjustment is recorded. The analyst must assess completeness after applying the stated rule."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, score criteria, scope, and bindings, while both contexts retain the stated derivation rule without adding exceptions. Evidence quotes: \"Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81.\" \"Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81.\" The counterfactual coherently changes one postal_country identifier, so only the intersection qualifies for derivation. Neither context contains an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Case note: A source-system operator delivered a customer batch for validation. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The validation summary records 12 blank required cells before transformation, including six blank region_code cells. The remaining six blanks are distributed among email and customer_name fields. Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81. Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81. The analyst is asked to assess completeness after applying the stated rule.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81."}, {"path": [], "text": "Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81.", "negative_left": "Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81.", "negative_right": "Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-82.", "right": "Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-006", "id": "fast-41-diverse-143-006-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Case note: A source-system operator delivered a customer batch for validation. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The validation summary records 12 blank required cells before transformation, including six blank region_code cells. The remaining six blanks are distributed among email and customer_name fields. Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81. Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81. The analyst is asked to assess completeness after applying the stated rule."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, score criteria, scope, and bindings, while both contexts retain the stated derivation rule without adding exceptions. Evidence quotes: \"Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81.\" \"Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81.\" The counterfactual coherently changes one postal_country identifier, so only the intersection qualifies for derivation. Neither context contains an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Case note: A source-system operator delivered a customer batch for validation. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The validation summary records 12 blank required cells before transformation, including six blank region_code cells. The remaining six blanks are distributed among email and customer_name fields. Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81. Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81. The analyst is asked to assess completeness after applying the stated rule.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81."}, {"path": [], "text": "Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81.", "negative_left": "Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81.", "negative_right": "Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-82.", "right": "Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-81."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-006", "id": "fast-41-diverse-143-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Case note: A source-system operator delivered a customer batch for validation. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The validation summary records 12 blank required cells before transformation, including six blank region_code cells. The remaining six blanks are distributed among email and customer_name fields. Before transformation, the six customer-batch records with blank required region_code cells have record identifiers C-14, C-27, C-39, C-52, C-68, and C-81. Before transformation, the complete set of records with present postal_country values has identifiers C-14, C-27, C-39, C-52, C-68, and C-82. The analyst is asked to assess completeness after applying the stated rule."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rule, scope, entity, timing, and scoring criteria. Both contexts preserve the schema, transformation condition, and post-transformation counting rule without adding exceptions. The evidence consists of two complete factual sentences: “Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79.” “Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91.” The counterfactual changes the postal_country observation consistently while retaining six affected records and the same policy. Neither context embeds an answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A quality-control case note covers a 100-record customer batch. The product schema requires customer_name, email, and region_code, giving 300 required cells. Before processing, the validation ledger records 12 blank required cells: six in region_code, three in email, and three in customer_name. The six affected region_code records and the records with available postal_country values are identified in the audit sentences below. The transformation is applied only where the stated condition is met, and a derived region_code is counted as populated afterward. No other population changes, duplicate-related adjustments, or schema changes are recorded. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79. Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79."}, {"path": [], "text": "Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79.", "negative_left": "Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79.", "negative_right": "Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, and C-91.", "right": "Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-012", "id": "fast-41-diverse-143-012-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A quality-control case note covers a 100-record customer batch. The product schema requires customer_name, email, and region_code, giving 300 required cells. Before processing, the validation ledger records 12 blank required cells: six in region_code, three in email, and three in customer_name. The six affected region_code records and the records with available postal_country values are identified in the audit sentences below. The transformation is applied only where the stated condition is met, and a derived region_code is counted as populated afterward. No other population changes, duplicate-related adjustments, or schema changes are recorded. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79. Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rule, scope, entity, timing, and scoring criteria. Both contexts preserve the schema, transformation condition, and post-transformation counting rule without adding exceptions. The evidence consists of two complete factual sentences: “Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79.” “Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91.” The counterfactual changes the postal_country observation consistently while retaining six affected records and the same policy. Neither context embeds an answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A quality-control case note covers a 100-record customer batch. The product schema requires customer_name, email, and region_code, giving 300 required cells. Before processing, the validation ledger records 12 blank required cells: six in region_code, three in email, and three in customer_name. The six affected region_code records and the records with available postal_country values are identified in the audit sentences below. The transformation is applied only where the stated condition is met, and a derived region_code is counted as populated afterward. No other population changes, duplicate-related adjustments, or schema changes are recorded. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79. Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79."}, {"path": [], "text": "Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79.", "negative_left": "Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79.", "negative_right": "Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, and C-91.", "right": "Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, C-79, and C-91."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-012", "id": "fast-41-diverse-143-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A quality-control case note covers a 100-record customer batch. The product schema requires customer_name, email, and region_code, giving 300 required cells. Before processing, the validation ledger records 12 blank required cells: six in region_code, three in email, and three in customer_name. The six affected region_code records and the records with available postal_country values are identified in the audit sentences below. The transformation is applied only where the stated condition is met, and a derived region_code is counted as populated afterward. No other population changes, duplicate-related adjustments, or schema changes are recorded. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before the stated transformation, the six record identifiers attached to the customer batch’s blank required region_code cells are C-11, C-24, C-37, C-52, C-68, and C-79. Before the stated transformation, the complete list of record identifiers attached to present postal_country values in the customer batch is C-11, C-24, C-37, C-52, C-68, and C-91."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because both contexts retain the derivation rule and the unchanged questions preserve all criteria and instructions. Question bindings remain fixed to customer batch Q7, required cells, the stated transformation, and the resulting completeness rating. Evidence spans are complete factual sentences: \"Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109.\" and \"Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109.\" The counterfactual coherently changes one observed postal_country identifier without contradicting its own counts or governing rule. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Case note: Customer batch Q7 contains 100 customer records. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before review, 12 required cells are blank: six region_code cells, three email cells, and three customer_name cells. The six region_code omissions belong to the records identified below. A validation export lists the postal_country status for the same batch, and the analyst is instructed to apply the stated derivation rule only where its condition is met. No duplicate records were reported, and no other transformation changes required-cell status. Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109. Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109. The administrator requests the resulting completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109."}, {"path": [], "text": "Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109.", "negative_left": "Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109.", "negative_right": "Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-110.", "right": "Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-013", "id": "fast-41-diverse-143-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Case note: Customer batch Q7 contains 100 customer records. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before review, 12 required cells are blank: six region_code cells, three email cells, and three customer_name cells. The six region_code omissions belong to the records identified below. A validation export lists the postal_country status for the same batch, and the analyst is instructed to apply the stated derivation rule only where its condition is met. No duplicate records were reported, and no other transformation changes required-cell status. Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109. Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109. The administrator requests the resulting completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because both contexts retain the derivation rule and the unchanged questions preserve all criteria and instructions. Question bindings remain fixed to customer batch Q7, required cells, the stated transformation, and the resulting completeness rating. Evidence spans are complete factual sentences: \"Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109.\" and \"Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109.\" The counterfactual coherently changes one observed postal_country identifier without contradicting its own counts or governing rule. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Case note: Customer batch Q7 contains 100 customer records. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before review, 12 required cells are blank: six region_code cells, three email cells, and three customer_name cells. The six region_code omissions belong to the records identified below. A validation export lists the postal_country status for the same batch, and the analyst is instructed to apply the stated derivation rule only where its condition is met. No duplicate records were reported, and no other transformation changes required-cell status. Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109. Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109. The administrator requests the resulting completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109."}, {"path": [], "text": "Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109.", "negative_left": "Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109.", "negative_right": "Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-110.", "right": "Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-109."}, "verifier_independent_model": false}, "family": "fast-41-diverse-143-013", "id": "fast-41-diverse-143-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Case note: Customer batch Q7 contains 100 customer records. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” Before review, 12 required cells are blank: six region_code cells, three email cells, and three customer_name cells. The six region_code omissions belong to the records identified below. A validation export lists the postal_country status for the same batch, and the analyst is instructed to apply the stated derivation rule only where its condition is met. No duplicate records were reported, and no other transformation changes required-cell status. Before the stated transformation, the six blank required region_code cells in customer batch Q7 have record identifiers C-104, C-105, C-106, C-107, C-108, and C-109. Before the stated transformation, the record identifiers attached to present postal_country values in customer batch Q7 are exactly C-104, C-105, C-106, C-107, C-108, and C-110. The administrator requests the resulting completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question and policy, retain the relevant bindings, provide two complete factual evidence sentences, change only the defect severity coherently, and contain no answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Case note: Northstar Notes 6.2 RC3 is under release review. The authoritative signed defect list was finalized on 14 September 2026. The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open. The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-1. The compatibility report records passing results for every required Windows 11 and macOS 15 check. The installation report records a successful required installation, and the signing report confirms that the required signing check passed. The rollback test does not pass: after restoration, the updater service remains disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging. The release dashboard's green summary is retained only as a non-authoritative observation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open."}, {"path": [], "text": "The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-1."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open.", "negative_right": "The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-3.", "right": "The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-1."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-003", "id": "fast-41-diverse-145-003-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Case note: Northstar Notes 6.2 RC3 is under release review. The authoritative signed defect list was finalized on 14 September 2026. The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open. The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-1. The compatibility report records passing results for every required Windows 11 and macOS 15 check. The installation report records a successful required installation, and the signing report confirms that the required signing check passed. The rollback test does not pass: after restoration, the updater service remains disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging. The release dashboard's green summary is retained only as a non-authoritative observation."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question and policy, retain the relevant bindings, provide two complete factual evidence sentences, change only the defect severity coherently, and contain no answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Case note: Northstar Notes 6.2 RC3 is under release review. The authoritative signed defect list was finalized on 14 September 2026. The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open. The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-1. The compatibility report records passing results for every required Windows 11 and macOS 15 check. The installation report records a successful required installation, and the signing report confirms that the required signing check passed. The rollback test does not pass: after restoration, the updater service remains disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging. The release dashboard's green summary is retained only as a non-authoritative observation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open."}, {"path": [], "text": "The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-1."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open.", "negative_right": "The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-3.", "right": "The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-1."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-003", "id": "fast-41-diverse-145-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Case note: Northstar Notes 6.2 RC3 is under release review. The authoritative signed defect list was finalized on 14 September 2026. The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 September 2026, contains exactly one defect entry, ticket NN-4821, and records that entry as open. The 14 September 2026 signed record for ticket NN-4821 rates the defect Sev-3. The compatibility report records passing results for every required Windows 11 and macOS 15 check. The installation report records a successful required installation, and the signing report confirms that the required signing check passed. The rollback test does not pass: after restoration, the updater service remains disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging. The release dashboard's green summary is retained only as a non-authoritative observation."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and Northstar Notes 6.2 RC3 scope; the changed defect ID and severity are observations. The two complete evidence quotes are “The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open.” and “A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1.” The counterfactual changes only Sev-1 to Sev-3 and remains coherent with the unchanged open defect and rollback failure. Neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Case note: Northstar Notes 6.2 RC3 was reviewed from signed reports rather than the release manager dashboard. The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open. A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1. The compatibility report records passing results for every required Windows 11 and macOS 15 check. The installation report records a successful required installation, and the signing report records a successful required signing check. The rollback report records that the updater service remains disabled after rollback, so the required rollback check does not pass. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open."}, {"path": [], "text": "A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open.", "negative_right": "A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-3.", "right": "A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-006", "id": "fast-41-diverse-145-006-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Case note: Northstar Notes 6.2 RC3 was reviewed from signed reports rather than the release manager dashboard. The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open. A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1. The compatibility report records passing results for every required Windows 11 and macOS 15 check. The installation report records a successful required installation, and the signing report records a successful required signing check. The rollback report records that the updater service remains disabled after rollback, so the required rollback check does not pass. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and Northstar Notes 6.2 RC3 scope; the changed defect ID and severity are observations. The two complete evidence quotes are “The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open.” and “A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1.” The counterfactual changes only Sev-1 to Sev-3 and remains coherent with the unchanged open defect and rollback failure. Neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Case note: Northstar Notes 6.2 RC3 was reviewed from signed reports rather than the release manager dashboard. The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open. A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1. The compatibility report records passing results for every required Windows 11 and macOS 15 check. The installation report records a successful required installation, and the signing report records a successful required signing check. The rollback report records that the updater service remains disabled after rollback, so the required rollback check does not pass. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open."}, {"path": [], "text": "A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open.", "negative_right": "A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-3.", "right": "A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-1."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-006", "id": "fast-41-diverse-145-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Case note: Northstar Notes 6.2 RC3 was reviewed from signed reports rather than the release manager dashboard. The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 14 August 2026, contains exactly one entry, defect NS-643, and marks that entry as open. A severity record dated 14 August 2026 rates defect NS-643 in Northstar Notes 6.2 RC3 as Sev-3. The compatibility report records passing results for every required Windows 11 and macOS 15 check. The installation report records a successful required installation, and the signing report records a successful required signing check. The rollback report records that the updater service remains disabled after rollback, so the required rollback check does not pass. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question, policy, Northstar Notes 6.2 RC3 binding, and time/scope details; the two evidence spans are complete factual sentences, the counterfactual coherently changes D-417 from open to closed without duplicate contradictions, and neither context embeds an answer or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Northstar Notes 6.2 RC3 is undergoing release review. The signed compatibility report records passing results for every required Windows 11 and macOS 15 check. The signed installation report records a passing installation gate, and the signing report records a passing signing gate. The rollback test fails because the updater service remains disabled afterward; this failed gate is assigned to packaging. The dashboard’s readiness summary is not controlling because detailed signed reports govern the record. For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417. The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked open and rated Sev-1. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417."}, {"path": [], "text": "The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked open and rated Sev-1."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417.", "negative_left": "For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417.", "negative_right": "The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked closed and rated Sev-1.", "right": "The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked open and rated Sev-1."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-007", "id": "fast-41-diverse-145-007-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Northstar Notes 6.2 RC3 is undergoing release review. The signed compatibility report records passing results for every required Windows 11 and macOS 15 check. The signed installation report records a passing installation gate, and the signing report records a passing signing gate. The rollback test fails because the updater service remains disabled afterward; this failed gate is assigned to packaging. The dashboard’s readiness summary is not controlling because detailed signed reports govern the record. For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417. The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked open and rated Sev-1. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question, policy, Northstar Notes 6.2 RC3 binding, and time/scope details; the two evidence spans are complete factual sentences, the counterfactual coherently changes D-417 from open to closed without duplicate contradictions, and neither context embeds an answer or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Northstar Notes 6.2 RC3 is undergoing release review. The signed compatibility report records passing results for every required Windows 11 and macOS 15 check. The signed installation report records a passing installation gate, and the signing report records a passing signing gate. The rollback test fails because the updater service remains disabled afterward; this failed gate is assigned to packaging. The dashboard’s readiness summary is not controlling because detailed signed reports govern the record. For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417. The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked open and rated Sev-1. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417."}, {"path": [], "text": "The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked open and rated Sev-1."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417.", "negative_left": "For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417.", "negative_right": "The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked closed and rated Sev-1.", "right": "The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked open and rated Sev-1."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-007", "id": "fast-41-diverse-145-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Northstar Notes 6.2 RC3 is undergoing release review. The signed compatibility report records passing results for every required Windows 11 and macOS 15 check. The signed installation report records a passing installation gate, and the signing report records a passing signing gate. The rollback test fails because the updater service remains disabled afterward; this failed gate is assigned to packaging. The dashboard’s readiness summary is not controlling because detailed signed reports govern the record. For Northstar Notes 6.2 RC3, the authoritative signed defect list contains exactly one defect record, labeled D-417. The sole record labeled D-417 on the Northstar Notes 6.2 RC3 defect list is marked closed and rated Sev-1. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the supplied policy, Northstar Notes 6.2 RC3 release scope, and classification request without leaking an answer; the evidence consists of the complete factual sentences “On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open.” and “On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2.”; the counterfactual coherently changes only the severities to Sev-3 and Sev-4 without contradictory duplicate assertions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Northstar Notes 6.2 RC3 is under review for the 14 September 2026 release window. On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open. On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2. The compatibility report records every required Windows 11 and macOS 15 check as passed. The installation report records the required installation gate as passed, and the signing report records the required signing gate as passed. The rollback report records the required rollback gate as failed. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging. The release manager has requested classification from these signed records rather than from the dashboard label.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open."}, {"path": [], "text": "On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open.", "negative_left": "On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open.", "negative_right": "On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-3 and NS-423 severity Sev-4.", "right": "On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-010", "id": "fast-41-diverse-145-010-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Northstar Notes 6.2 RC3 is under review for the 14 September 2026 release window. On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open. On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2. The compatibility report records every required Windows 11 and macOS 15 check as passed. The installation report records the required installation gate as passed, and the signing report records the required signing gate as passed. The rollback report records the required rollback gate as failed. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging. The release manager has requested classification from these signed records rather than from the dashboard label."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the supplied policy, Northstar Notes 6.2 RC3 release scope, and classification request without leaking an answer; the evidence consists of the complete factual sentences “On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open.” and “On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2.”; the counterfactual coherently changes only the severities to Sev-3 and Sev-4 without contradictory duplicate assertions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Northstar Notes 6.2 RC3 is under review for the 14 September 2026 release window. On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open. On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2. The compatibility report records every required Windows 11 and macOS 15 check as passed. The installation report records the required installation gate as passed, and the signing report records the required signing gate as passed. The rollback report records the required rollback gate as failed. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging. The release manager has requested classification from these signed records rather than from the dashboard label.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open."}, {"path": [], "text": "On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open.", "negative_left": "On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open.", "negative_right": "On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-3 and NS-423 severity Sev-4.", "right": "On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-1 and NS-423 severity Sev-2."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-010", "id": "fast-41-diverse-145-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Northstar Notes 6.2 RC3 is under review for the 14 September 2026 release window. On 14 September 2026, the authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly two entries, NS-417 and NS-423, and marks both entries open. On 14 September 2026, that authoritative signed defect list assigns NS-417 severity Sev-3 and NS-423 severity Sev-4. The compatibility report records every required Windows 11 and macOS 15 check as passed. The installation report records the required installation gate as passed, and the signing report records the required signing gate as passed. The rollback report records the required rollback gate as failed. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging. The release manager has requested classification from these signed records rather than from the dashboard label."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions and governing policy, with the same release, evidence paths, and 2026-09-17 binding. The two evidence spans are complete factual sentences: \"The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417.\" and \"Defect D-417 is open and rated Sev-1 as of 2026-09-17.\" The counterfactual coherently changes only D-417 from open to closed while retaining the failed rollback. Neither context embeds a gold answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Northstar Notes 6.2 RC3 was reviewed using signed release records rather than the release manager’s dashboard, which says “Ready: all gates green.” The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417. Defect D-417 is open and rated Sev-1 as of 2026-09-17. The compatibility tester’s signed report records every required Windows 11 and macOS 15 check as passed. The deployment specialist’s signed records show that the required installation gate passes and the required signing gate passes. Its rollback exercise does not pass because the updater service remains disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417."}, {"path": [], "text": "Defect D-417 is open and rated Sev-1 as of 2026-09-17."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417.", "negative_right": "Defect D-417 is closed and rated Sev-1 as of 2026-09-17.", "right": "Defect D-417 is open and rated Sev-1 as of 2026-09-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-011", "id": "fast-41-diverse-145-011-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Northstar Notes 6.2 RC3 was reviewed using signed release records rather than the release manager’s dashboard, which says “Ready: all gates green.” The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417. Defect D-417 is open and rated Sev-1 as of 2026-09-17. The compatibility tester’s signed report records every required Windows 11 and macOS 15 check as passed. The deployment specialist’s signed records show that the required installation gate passes and the required signing gate passes. Its rollback exercise does not pass because the updater service remains disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions and governing policy, with the same release, evidence paths, and 2026-09-17 binding. The two evidence spans are complete factual sentences: \"The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417.\" and \"Defect D-417 is open and rated Sev-1 as of 2026-09-17.\" The counterfactual coherently changes only D-417 from open to closed while retaining the failed rollback. Neither context embeds a gold answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Northstar Notes 6.2 RC3 was reviewed using signed release records rather than the release manager’s dashboard, which says “Ready: all gates green.” The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417. Defect D-417 is open and rated Sev-1 as of 2026-09-17. The compatibility tester’s signed report records every required Windows 11 and macOS 15 check as passed. The deployment specialist’s signed records show that the required installation gate passes and the required signing gate passes. Its rollback exercise does not pass because the updater service remains disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417."}, {"path": [], "text": "Defect D-417 is open and rated Sev-1 as of 2026-09-17."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417.", "negative_right": "Defect D-417 is closed and rated Sev-1 as of 2026-09-17.", "right": "Defect D-417 is open and rated Sev-1 as of 2026-09-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-011", "id": "fast-41-diverse-145-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Northstar Notes 6.2 RC3 was reviewed using signed release records rather than the release manager’s dashboard, which says “Ready: all gates green.” The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one listed defect, D-417. Defect D-417 is closed and rated Sev-1 as of 2026-09-17. The compatibility tester’s signed report records every required Windows 11 and macOS 15 check as passed. The deployment specialist’s signed records show that the required installation gate passes and the required signing gate passes. Its rollback exercise does not pass because the updater service remains disabled. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original release, policy, gate ownership, and evidence-authority rules. The release entity remains Northstar Notes 6.2 RC3, with no restricted question binding changed. The two evidence spans are complete factual sentences and contain no policy instructions. The counterfactual changes only the defect severity from Sev-1 to Sev-3 while remaining consistent with the exactly-one-defect assertion. Neither context states a classification, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Northstar Notes 6.2 RC3 was reviewed using signed reports rather than the release dashboard. The compatibility report records passing required checks for Windows 11 and macOS 15. The installation report records a successful required installation, and the signing report confirms the required signing gate passes. The rollback exercise did not pass: it left the updater service disabled. The authoritative defect record states: \\\"The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417.\\\" It also states: \\\"Defect NNS-417 is rated Sev-1 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17.\\\" Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417."}, {"path": [], "text": "Defect NNS-417 is rated Sev-1 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417.", "negative_right": "Defect NNS-417 is rated Sev-3 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17.", "right": "Defect NNS-417 is rated Sev-1 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-013", "id": "fast-41-diverse-145-013-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Northstar Notes 6.2 RC3 was reviewed using signed reports rather than the release dashboard. The compatibility report records passing required checks for Windows 11 and macOS 15. The installation report records a successful required installation, and the signing report confirms the required signing gate passes. The rollback exercise did not pass: it left the updater service disabled. The authoritative defect record states: \"The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417.\" It also states: \"Defect NNS-417 is rated Sev-1 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17.\" Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original release, policy, gate ownership, and evidence-authority rules. The release entity remains Northstar Notes 6.2 RC3, with no restricted question binding changed. The two evidence spans are complete factual sentences and contain no policy instructions. The counterfactual changes only the defect severity from Sev-1 to Sev-3 while remaining consistent with the exactly-one-defect assertion. Neither context states a classification, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Northstar Notes 6.2 RC3 was reviewed using signed reports rather than the release dashboard. The compatibility report records passing required checks for Windows 11 and macOS 15. The installation report records a successful required installation, and the signing report confirms the required signing gate passes. The rollback exercise did not pass: it left the updater service disabled. The authoritative defect record states: \\\"The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417.\\\" It also states: \\\"Defect NNS-417 is rated Sev-1 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17.\\\" Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417."}, {"path": [], "text": "Defect NNS-417 is rated Sev-1 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417.", "negative_right": "Defect NNS-417 is rated Sev-3 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17.", "right": "Defect NNS-417 is rated Sev-1 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-145-013", "id": "fast-41-diverse-145-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Northstar Notes 6.2 RC3 was reviewed using signed reports rather than the release dashboard. The compatibility report records passing required checks for Windows 11 and macOS 15. The installation report records a successful required installation, and the signing report confirms the required signing gate passes. The rollback exercise did not pass: it left the updater service disabled. The authoritative defect record states: \"The authoritative signed defect list for Northstar Notes 6.2 RC3, finalized on 2026-09-17, contains exactly one open defect, identified as NNS-417.\" It also states: \"Defect NNS-417 is rated Sev-3 on the authoritative signed defect list for Northstar Notes 6.2 RC3 finalized on 2026-09-17.\" Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, and both contexts retain the complete gate policy. The RC-17, approval-decision-now, and in-scope-platform bindings remain unchanged. The evidence consists of two complete factual sentences: \"At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms.\" and \"At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed.\" The counterfactual changes readiness consistently from both platforms passing to both platforms failing without duplicate contradictions. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Scope register\",\"text\":\"At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms.\"},{\"speaker\":\"Readiness auditor\",\"text\":\"At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed.\"},{\"speaker\":\"Defect coordinator\",\"text\":\"At the approval decision now, no in-scope Severity-1 or Severity-2 defect remains open for RC-17.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"At the approval decision now, every RC-17 installer is signed.\"},{\"speaker\":\"Release engineer\",\"text\":\"At the approval decision now, the rollback procedure for RC-17 has been rehearsed successfully.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms."}, {"path": ["2", "text"], "text": "At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms.", "negative_left": "At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms.", "negative_right": "At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both failed.", "right": "At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-003", "id": "fast-41-diverse-147-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Scope register", "text": "At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms."}, {"speaker": "Readiness auditor", "text": "At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed."}, {"speaker": "Defect coordinator", "text": "At the approval decision now, no in-scope Severity-1 or Severity-2 defect remains open for RC-17."}, {"speaker": "Deployment specialist", "text": "At the approval decision now, every RC-17 installer is signed."}, {"speaker": "Release engineer", "text": "At the approval decision now, the rollback procedure for RC-17 has been rehearsed successfully."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, and both contexts retain the complete gate policy. The RC-17, approval-decision-now, and in-scope-platform bindings remain unchanged. The evidence consists of two complete factual sentences: \"At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms.\" and \"At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed.\" The counterfactual changes readiness consistently from both platforms passing to both platforms failing without duplicate contradictions. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Scope register\",\"text\":\"At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms.\"},{\"speaker\":\"Readiness auditor\",\"text\":\"At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed.\"},{\"speaker\":\"Defect coordinator\",\"text\":\"At the approval decision now, no in-scope Severity-1 or Severity-2 defect remains open for RC-17.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"At the approval decision now, every RC-17 installer is signed.\"},{\"speaker\":\"Release engineer\",\"text\":\"At the approval decision now, the rollback procedure for RC-17 has been rehearsed successfully.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms."}, {"path": ["2", "text"], "text": "At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms.", "negative_left": "At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms.", "negative_right": "At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both failed.", "right": "At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both passed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-003", "id": "fast-41-diverse-147-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Scope register", "text": "At the approval decision now, release candidate RC-17 has Linux Atlas and macOS Birch as its in-scope platforms."}, {"speaker": "Readiness auditor", "text": "At the approval decision now, readiness testing records show that Linux Atlas and macOS Birch have both failed."}, {"speaker": "Defect coordinator", "text": "At the approval decision now, no in-scope Severity-1 or Severity-2 defect remains open for RC-17."}, {"speaker": "Deployment specialist", "text": "At the approval decision now, every RC-17 installer is signed."}, {"speaker": "Release engineer", "text": "At the approval decision now, the rollback procedure for RC-17 has been rehearsed successfully."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The gate rule and exception rule are preserved in both contexts. The release-candidate and current-decision bindings remain intact. The two evidence spans are complete factual sentences: \"At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3.\" and \"At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test.\" The counterfactual changes readiness results consistently while retaining all other facts. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Release coordinator\",\"text\":\"Approval review concerns the current release candidate. The scope record lists no additional platform, defect, installer, or rollback exception relevant to this decision.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3. At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test.\"},{\"speaker\":\"Quality lead\",\"text\":\"No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"Every installer for this release candidate is signed at the approval decision now. The rollback procedure for this release candidate has been rehearsed by the approval decision now.\"},{\"speaker\":\"Review clerk\",\"text\":\"For a consistency check, if the readiness records instead showed the two in-scope platforms failing, the remaining recorded gate facts would be unchanged.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["2", "text"], "text": "At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3."}, {"path": ["2", "text"], "text": "At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3.", "negative_left": "At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3.", "negative_right": "At the approval decision now, Orion-7 and Vega-3 each have a recorded failing readiness test.", "right": "At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-008", "id": "fast-41-diverse-147-008-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Release coordinator", "text": "Approval review concerns the current release candidate. The scope record lists no additional platform, defect, installer, or rollback exception relevant to this decision."}, {"speaker": "Compatibility tester", "text": "At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3. At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test."}, {"speaker": "Quality lead", "text": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"speaker": "Deployment specialist", "text": "Every installer for this release candidate is signed at the approval decision now. The rollback procedure for this release candidate has been rehearsed by the approval decision now."}, {"speaker": "Review clerk", "text": "For a consistency check, if the readiness records instead showed the two in-scope platforms failing, the remaining recorded gate facts would be unchanged."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The gate rule and exception rule are preserved in both contexts. The release-candidate and current-decision bindings remain intact. The two evidence spans are complete factual sentences: \"At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3.\" and \"At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test.\" The counterfactual changes readiness results consistently while retaining all other facts. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Release coordinator\",\"text\":\"Approval review concerns the current release candidate. The scope record lists no additional platform, defect, installer, or rollback exception relevant to this decision.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3. At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test.\"},{\"speaker\":\"Quality lead\",\"text\":\"No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"Every installer for this release candidate is signed at the approval decision now. The rollback procedure for this release candidate has been rehearsed by the approval decision now.\"},{\"speaker\":\"Review clerk\",\"text\":\"For a consistency check, if the readiness records instead showed the two in-scope platforms failing, the remaining recorded gate facts would be unchanged.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["2", "text"], "text": "At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3."}, {"path": ["2", "text"], "text": "At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3.", "negative_left": "At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3.", "negative_right": "At the approval decision now, Orion-7 and Vega-3 each have a recorded failing readiness test.", "right": "At the approval decision now, Orion-7 and Vega-3 each have a recorded passing readiness test."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-008", "id": "fast-41-diverse-147-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Release coordinator", "text": "Approval review concerns the current release candidate. The scope record lists no additional platform, defect, installer, or rollback exception relevant to this decision."}, {"speaker": "Compatibility tester", "text": "At the approval decision now, the in-scope platforms for this release candidate are exactly Orion-7 and Vega-3. At the approval decision now, Orion-7 and Vega-3 each have a recorded failing readiness test."}, {"speaker": "Quality lead", "text": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"speaker": "Deployment specialist", "text": "Every installer for this release candidate is signed at the approval decision now. The rollback procedure for this release candidate has been rehearsed by the approval decision now."}, {"speaker": "Review clerk", "text": "For a consistency check, if the readiness records instead showed the two in-scope platforms failing, the remaining recorded gate facts would be unchanged."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the gate policy and approval-now release-candidate scope. The evidence contains exactly two complete factual sentences: “At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega.” and “The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing.” The counterfactual coherently changes readiness from passed to failed without duplicate contradictions, and neither context states an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Release coordinator\",\"text\":\"At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega.\"},{\"speaker\":\"Readiness auditor\",\"text\":\"The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing.\"},{\"speaker\":\"Defect manager\",\"text\":\"At the approval decision now, no in-scope Severity-1 or Severity-2 defect remains open for RC-47.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"At the approval decision now, every installer for RC-47 is signed, and the rollback procedure has been rehearsed successfully.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega."}, {"path": ["2", "text"], "text": "The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega.", "negative_left": "At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega.", "negative_right": "The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each failed readiness testing.", "right": "The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-009", "id": "fast-41-diverse-147-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Release coordinator", "text": "At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega."}, {"speaker": "Readiness auditor", "text": "The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing."}, {"speaker": "Defect manager", "text": "At the approval decision now, no in-scope Severity-1 or Severity-2 defect remains open for RC-47."}, {"speaker": "Deployment specialist", "text": "At the approval decision now, every installer for RC-47 is signed, and the rollback procedure has been rehearsed successfully."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the gate policy and approval-now release-candidate scope. The evidence contains exactly two complete factual sentences: “At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega.” and “The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing.” The counterfactual coherently changes readiness from passed to failed without duplicate contradictions, and neither context states an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Release coordinator\",\"text\":\"At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega.\"},{\"speaker\":\"Readiness auditor\",\"text\":\"The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing.\"},{\"speaker\":\"Defect manager\",\"text\":\"At the approval decision now, no in-scope Severity-1 or Severity-2 defect remains open for RC-47.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"At the approval decision now, every installer for RC-47 is signed, and the rollback procedure has been rehearsed successfully.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega."}, {"path": ["2", "text"], "text": "The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega.", "negative_left": "At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega.", "negative_right": "The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each failed readiness testing.", "right": "The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each passed readiness testing."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-009", "id": "fast-41-diverse-147-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Release coordinator", "text": "At the approval decision now, the in-scope platforms for release candidate RC-47 are Linux Atlas and Solaris Vega."}, {"speaker": "Readiness auditor", "text": "The readiness records at the approval decision now show that Linux Atlas and Solaris Vega each failed readiness testing."}, {"speaker": "Defect manager", "text": "At the approval decision now, no in-scope Severity-1 or Severity-2 defect remains open for RC-47."}, {"speaker": "Deployment specialist", "text": "At the approval decision now, every installer for RC-47 is signed, and the rollback procedure has been rehearsed successfully."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing instructions and criteria, and both contexts retain the gate rule. The release candidate, decision date, and platform scope remain consistently bound in both contexts. Required evidence quote 1: \"At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04.\" Required evidence quote 2: \"At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing.\" The counterfactual changes only Ubuntu 24.04 readiness and introduces no contradictory duplicate assertion. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04.\"},{\"speaker\":\"Readiness lead\",\"text\":\"At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing.\"},{\"speaker\":\"Defect coordinator\",\"text\":\"At the approval decision, no in-scope Severity-1 or Severity-2 defect remains open for RC-47.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"At the approval decision, every installer for RC-47 is signed, and the rollback procedure has been rehearsed successfully.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04."}, {"path": ["2", "text"], "text": "At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04.", "negative_left": "At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04.", "negative_right": "At the approval decision on 17 September 2026, readiness records show that Windows 11 passed readiness testing and Ubuntu 24.04 failed readiness testing.", "right": "At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-012", "id": "fast-41-diverse-147-012-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Compatibility tester", "text": "At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04."}, {"speaker": "Readiness lead", "text": "At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing."}, {"speaker": "Defect coordinator", "text": "At the approval decision, no in-scope Severity-1 or Severity-2 defect remains open for RC-47."}, {"speaker": "Deployment specialist", "text": "At the approval decision, every installer for RC-47 is signed, and the rollback procedure has been rehearsed successfully."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing instructions and criteria, and both contexts retain the gate rule. The release candidate, decision date, and platform scope remain consistently bound in both contexts. Required evidence quote 1: \"At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04.\" Required evidence quote 2: \"At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing.\" The counterfactual changes only Ubuntu 24.04 readiness and introduces no contradictory duplicate assertion. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04.\"},{\"speaker\":\"Readiness lead\",\"text\":\"At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing.\"},{\"speaker\":\"Defect coordinator\",\"text\":\"At the approval decision, no in-scope Severity-1 or Severity-2 defect remains open for RC-47.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"At the approval decision, every installer for RC-47 is signed, and the rollback procedure has been rehearsed successfully.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04."}, {"path": ["2", "text"], "text": "At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04.", "negative_left": "At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04.", "negative_right": "At the approval decision on 17 September 2026, readiness records show that Windows 11 passed readiness testing and Ubuntu 24.04 failed readiness testing.", "right": "At the approval decision on 17 September 2026, readiness records show that Windows 11 and Ubuntu 24.04 passed readiness testing."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-012", "id": "fast-41-diverse-147-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Compatibility tester", "text": "At the approval decision on 17 September 2026, the complete set of platforms in scope for release candidate RC-47 is Windows 11 and Ubuntu 24.04."}, {"speaker": "Readiness lead", "text": "At the approval decision on 17 September 2026, readiness records show that Windows 11 passed readiness testing and Ubuntu 24.04 failed readiness testing."}, {"speaker": "Defect coordinator", "text": "At the approval decision, no in-scope Severity-1 or Severity-2 defect remains open for RC-47."}, {"speaker": "Deployment specialist", "text": "At the approval decision, every installer for RC-47 is signed, and the rollback procedure has been rehearsed successfully."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy, release scope, approval timing, and entity bindings are preserved in both contexts. The evidence consists of two complete factual sentences: \"At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1.\" and \"The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1.\" The counterfactual coherently changes the readiness results from passed to failed without contradictory duplicates. Neither context includes answer codes, rationale, rule tables, proposition IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Approval record\",\"text\":\"At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1.\"},{\"speaker\":\"Readiness coordinator\",\"text\":\"The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1.\"},{\"speaker\":\"Defect manager\",\"text\":\"The release-candidate review found no in-scope Severity-1 or Severity-2 defect still open at the approval decision.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"Every installer for the release candidate was signed at the approval decision, and the rollback procedure had been rehearsed successfully by then.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1."}, {"path": ["2", "text"], "text": "The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1.", "negative_left": "At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1.", "negative_right": "The readiness-test ledger timestamped 17 September 2026 records failed results for OrionOS 4.2 and VegaOS 9.1.", "right": "The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-013", "id": "fast-41-diverse-147-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Approval record", "text": "At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1."}, {"speaker": "Readiness coordinator", "text": "The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1."}, {"speaker": "Defect manager", "text": "The release-candidate review found no in-scope Severity-1 or Severity-2 defect still open at the approval decision."}, {"speaker": "Deployment specialist", "text": "Every installer for the release candidate was signed at the approval decision, and the rollback procedure had been rehearsed successfully by then."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy, release scope, approval timing, and entity bindings are preserved in both contexts. The evidence consists of two complete factual sentences: \"At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1.\" and \"The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1.\" The counterfactual coherently changes the readiness results from passed to failed without contradictory duplicates. Neither context includes answer codes, rationale, rule tables, proposition IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual gate condition rather than a policy classification. The focus A1 is factual. The cited state evidence preserves the complete gate and Windows 10 exception rule needed from the original state; instructions and decision criteria in the questions object need not be duplicated. The base and counter assignments are realizable with only A1 changing because a platform can fail readiness testing without leaving an open Severity-1/2 defect, while signing and rollback remain successful.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all four applicable release gates pass: all in-scope platforms passed, no in-scope Severity-1/2 defect remains, all installers are signed, and rollback was rehearsed. This is sufficient for approval.", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that at least one in-scope platform has not passed. That failure defeats an applicable release gate regardless of the other three gates passing, so non-approval is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Every platform in scope for this release candidate at the approval decision now has passed readiness testing."}, {"id": "A2", "statement": "No in-scope Severity-1 or Severity-2 defect remains open for this release candidate at the approval decision now."}, {"id": "A3", "statement": "Every installer for this release candidate is signed at the approval decision now."}, {"id": "A4", "statement": "The rollback procedure for this release candidate has been rehearsed by the approval decision now."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported.\"},{\"speaker\":\"Approval record\",\"text\":\"At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1.\"},{\"speaker\":\"Readiness coordinator\",\"text\":\"The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1.\"},{\"speaker\":\"Defect manager\",\"text\":\"The release-candidate review found no in-scope Severity-1 or Severity-2 defect still open at the approval decision.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"Every installer for the release candidate was signed at the approval decision, and the rollback procedure had been rehearsed successfully by then.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1."}, {"path": ["2", "text"], "text": "The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1."}], "policy_evidence": [{"path": ["0", "text"], "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}], "rules": [{"justification": "Every applicable release gate passes now.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "At least one currently in-scope platform fails, so the all-in-scope-platforms gate fails even though the other gates pass.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1.", "negative_left": "At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1.", "negative_right": "The readiness-test ledger timestamped 17 September 2026 records failed results for OrionOS 4.2 and VegaOS 9.1.", "right": "The readiness-test ledger timestamped 17 September 2026 records passed results for OrionOS 4.2 and VegaOS 9.1."}, "verifier_independent_model": false}, "family": "fast-41-diverse-147-013", "id": "fast-41-diverse-147-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve it now because at least one applicable release gate fails.", "true": "Yes — approve the release candidate now because every applicable release gate passes."}, "instructions": "Readiness verification: decide whether the release candidate passes the stated approval gate now. Apply the scope and exception rule exactly as written.", "type": "noul"}}, "state": [{"speaker": "Release manager", "text": "Gate rule: approve only if all in-scope platforms pass, no in-scope Severity-1/2 defects remain, installers are signed, and rollback is rehearsed. Windows 10 failures may be excepted only after release notes explicitly mark Windows 10 unsupported."}, {"speaker": "Approval record", "text": "At the approval decision on 17 September 2026, the in-scope platforms were OrionOS 4.2 and VegaOS 9.1."}, {"speaker": "Readiness coordinator", "text": "The readiness-test ledger timestamped 17 September 2026 records failed results for OrionOS 4.2 and VegaOS 9.1."}, {"speaker": "Defect manager", "text": "The release-candidate review found no in-scope Severity-1 or Severity-2 defect still open at the approval decision."}, {"speaker": "Deployment specialist", "text": "Every installer for the release candidate was signed at the approval decision, and the rollback procedure had been rehearsed successfully by then."}]}, "method": "c2d", "provenance": {"source_id": "diverse-147", "source_is_synthetic": true, "source_sha256": "4a0247d799192b68b120cd17c264036fdd05657d487744c13544b3b647a72155", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and bindings are preserved, including the product, date, blocker, and packaging-routing scope. The base context contains contradictory report assertions because D and S are separately named as sole defects in R, while the counterfactual places D in R and S in Q with disjoint lists. The evidence consists of complete factual sentences, including “In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R.”, “In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R.”, and “In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report Q.” The counterfactual is coherent because its report assignments do not contradict disjointness or sole-defect claims. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"The 2026-09-17 release case concerns Northstar Notes 4.2 RC3. The deployment specialist records one open blocker, and that blocker concerns exactly one defect, D. The Windows installer omits the required signature from its updater executable; this is defect S. The release inventory states that every defect involving installer construction or signing is identical to S. The application binaries pass runtime checks, compatibility checks pass on Windows 11, macOS 15, and Ubuntu 24.04, and rollback to version 4.1 succeeds. In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R. In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations. Any open blocker fails the release gate.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R."}, {"path": [], "text": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R.", "negative_left": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R.", "negative_right": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report Q.", "right": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R."}, "verifier_independent_model": false}, "family": "fast-41-diverse-148-010", "id": "fast-41-diverse-148-010-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "The 2026-09-17 release case concerns Northstar Notes 4.2 RC3. The deployment specialist records one open blocker, and that blocker concerns exactly one defect, D. The Windows installer omits the required signature from its updater executable; this is defect S. The release inventory states that every defect involving installer construction or signing is identical to S. The application binaries pass runtime checks, compatibility checks pass on Windows 11, macOS 15, and Ubuntu 24.04, and rollback to version 4.1 succeeds. In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R. In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations. Any open blocker fails the release gate."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and bindings are preserved, including the product, date, blocker, and packaging-routing scope. The base context contains contradictory report assertions because D and S are separately named as sole defects in R, while the counterfactual places D in R and S in Q with disjoint lists. The evidence consists of complete factual sentences, including “In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R.”, “In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R.”, and “In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report Q.” The counterfactual is coherent because its report assignments do not contradict disjointness or sole-defect claims. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"The 2026-09-17 release case concerns Northstar Notes 4.2 RC3. The deployment specialist records one open blocker, and that blocker concerns exactly one defect, D. The Windows installer omits the required signature from its updater executable; this is defect S. The release inventory states that every defect involving installer construction or signing is identical to S. The application binaries pass runtime checks, compatibility checks pass on Windows 11, macOS 15, and Ubuntu 24.04, and rollback to version 4.1 succeeds. In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R. In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations. Any open blocker fails the release gate.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R."}, {"path": [], "text": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R.", "negative_left": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R.", "negative_right": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report Q.", "right": "In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report R."}, "verifier_independent_model": false}, "family": "fast-41-diverse-148-010", "id": "fast-41-diverse-148-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "The 2026-09-17 release case concerns Northstar Notes 4.2 RC3. The deployment specialist records one open blocker, and that blocker concerns exactly one defect, D. The Windows installer omits the required signature from its updater executable; this is defect S. The release inventory states that every defect involving installer construction or signing is identical to S. The application binaries pass runtime checks, compatibility checks pass on Windows 11, macOS 15, and Ubuntu 24.04, and rollback to version 4.1 succeeds. In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, report R and report Q are distinct, their defect lists are disjoint, and defect D is the sole defect listed in report R. In the 2026-09-17 deployment records for Northstar Notes 4.2 RC3, defect S is the sole defect listed in report Q. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations. Any open blocker fails the release gate."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, use two complete factual evidence sentences, and the counterfactual coherently changes the specification to conflict with the observed state without embedding an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":\"Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes. The release manager reviewed Northstar Notes 6.2 RC3 on 2026-09-16. All supported-platform compatibility tests passed, and every installer rollback other than the Windows MSI rollback succeeded. No release-blocking defect is open. Every release package has a valid signature. The release record documents a keyboard-shortcuts-table typo; it is classified as nonblocking but still requires an approved release-note action. The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate. On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17. On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R17.\",\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["context"], "text": "On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17."}, {"path": ["context"], "text": "On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R17."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17.", "negative_left": "On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17.", "negative_right": "On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R18.", "right": "On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-149-002", "id": "fast-41-diverse-149-002-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes. The release manager reviewed Northstar Notes 6.2 RC3 on 2026-09-16. All supported-platform compatibility tests passed, and every installer rollback other than the Windows MSI rollback succeeded. No release-blocking defect is open. Every release package has a valid signature. The release record documents a keyboard-shortcuts-table typo; it is classified as nonblocking but still requires an approved release-note action. The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate. On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17. On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R17.", "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and bindings, use two complete factual evidence sentences, and the counterfactual coherently changes the specification to conflict with the observed state without embedding an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":\"Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes. The release manager reviewed Northstar Notes 6.2 RC3 on 2026-09-16. All supported-platform compatibility tests passed, and every installer rollback other than the Windows MSI rollback succeeded. No release-blocking defect is open. Every release package has a valid signature. The release record documents a keyboard-shortcuts-table typo; it is classified as nonblocking but still requires an approved release-note action. The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate. On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17. On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R17.\",\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["context"], "text": "On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17."}, {"path": ["context"], "text": "On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R17."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17.", "negative_left": "On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17.", "negative_right": "On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R18.", "right": "On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-149-002", "id": "fast-41-diverse-149-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes. The release manager reviewed Northstar Notes 6.2 RC3 on 2026-09-16. All supported-platform compatibility tests passed, and every installer rollback other than the Windows MSI rollback succeeded. No release-blocking defect is open. Every release package has a valid signature. The release record documents a keyboard-shortcuts-table typo; it is classified as nonblocking but still requires an approved release-note action. The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate. On 2026-09-16 at 14:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was state code R17. On 2026-09-16 at 14:20 UTC, the signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI specified registration state code R18.", "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete original policy, criteria, instructions, entity, scope, and request. Both contexts retain the governing release policy and the RC3 readiness and gate-failure bindings. All evidence entries are complete factual sentences rather than instructions or policy definitions. The base evidence is internally consistent: \"On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.\" The matching specification states: \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417.\" The counterfactual coherently changes the specification to code R-418, creating a failed rollback comparison without contradictory duplicate measurements. Neither context embeds a gold answer, answer code, label rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":{\"case_note\":\"Northstar Notes 6.2 RC3 release review, 2026-09-14. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\",\"evidence\":[\"The compatibility report records passing results for every supported platform, and all installer rollbacks other than the Windows MSI rollback completed successfully.\",\"On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.\",\"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417.\",\"The release record documents a keyboard-shortcuts-table typo; documentation review classifies it as nonblocking, but an approved release-note action is still required.\",\"No release-blocking defect for RC3 is open, and every RC3 release package has a valid signature.\",\"The Northstar Notes deployment team owns resolution of any Windows MSI rollback gate failure.\"],\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["context", "evidence", "1"], "text": "On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417."}, {"path": ["context", "evidence", "2"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.", "negative_left": "On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.", "negative_right": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-418.", "right": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-149-012", "id": "fast-41-diverse-149-012-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": {"case_note": "Northstar Notes 6.2 RC3 release review, 2026-09-14. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility report records passing results for every supported platform, and all installer rollbacks other than the Windows MSI rollback completed successfully.", "On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417.", "The release record documents a keyboard-shortcuts-table typo; documentation review classifies it as nonblocking, but an approved release-note action is still required.", "No release-blocking defect for RC3 is open, and every RC3 release package has a valid signature.", "The Northstar Notes deployment team owns resolution of any Windows MSI rollback gate failure."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete original policy, criteria, instructions, entity, scope, and request. Both contexts retain the governing release policy and the RC3 readiness and gate-failure bindings. All evidence entries are complete factual sentences rather than instructions or policy definitions. The base evidence is internally consistent: \"On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.\" The matching specification states: \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417.\" The counterfactual coherently changes the specification to code R-418, creating a failed rollback comparison without contradictory duplicate measurements. Neither context embeds a gold answer, answer code, label rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":{\"case_note\":\"Northstar Notes 6.2 RC3 release review, 2026-09-14. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\",\"evidence\":[\"The compatibility report records passing results for every supported platform, and all installer rollbacks other than the Windows MSI rollback completed successfully.\",\"On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.\",\"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417.\",\"The release record documents a keyboard-shortcuts-table typo; documentation review classifies it as nonblocking, but an approved release-note action is still required.\",\"No release-blocking defect for RC3 is open, and every RC3 release package has a valid signature.\",\"The Northstar Notes deployment team owns resolution of any Windows MSI rollback gate failure.\"],\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["context", "evidence", "1"], "text": "On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417."}, {"path": ["context", "evidence", "2"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.", "negative_left": "On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.", "negative_right": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-418.", "right": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-149-012", "id": "fast-41-diverse-149-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": {"case_note": "Northstar Notes 6.2 RC3 release review, 2026-09-14. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility report records passing results for every supported platform, and all installer rollbacks other than the Windows MSI rollback completed successfully.", "On 2026-09-14 at 16:20 UTC, the observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI was code R-417.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI lists registration state code R-418.", "The release record documents a keyboard-shortcuts-table typo; documentation review classifies it as nonblocking, but an approved release-note action is still required.", "No release-blocking defect for RC3 is open, and every RC3 release package has a valid signature.", "The Northstar Notes deployment team owns resolution of any Windows MSI rollback gate failure."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and bindings, the focus evidence is complete and exact—“For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build.” and “The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing.”—and the counterfactual coherently changes only the trial result without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom a6 is factual rather than a policy classification. The base and counter assignments are jointly realizable while changing only whether the cleanup-command trial passes; a7 merely excludes other tested workarounds and need not assert that the cleanup command succeeds. Empty policy_evidence is correct because all governing readiness criteria and interpretive instructions occur in the retained questions object, while the original state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The Ubuntu required-platform check fails without remediation, the cleanup command does not produce a passing result, and no other tested workaround exists. Therefore a required-platform check fails without a tested workaround, which is sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes no open blocker/high-severity defects, passing core and required-platform checks through the tested Ubuntu workaround, verified installer signing/integrity, validated rollback, and exactly one open medium packaging-or-compatibility defect whose documentation approval remains pending. This is sufficient for level 1 and excludes the stated level-0 conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of open blocker defects for RC 4 is zero."}, {"id": "a2", "statement": "The number of open high-severity defects for RC 4 is zero."}, {"id": "a3", "statement": "Each of the 186 required smoke and regression tests for RC 4 passed."}, {"id": "a4", "statement": "Every required-platform check for RC 4 other than the supported-Ubuntu obsolete-tray-helper check passed."}, {"id": "a5", "statement": "Without remediation, the obsolete-tray-helper defect causes its required-platform check on the supported Ubuntu build for RC 4 to fail."}, {"id": "a6", "statement": "The cleanup-command trial for the obsolete-tray-helper defect produced a passing result for its required-platform check on the supported Ubuntu build for RC 4."}, {"id": "a7", "statement": "No tested workaround other than the cleanup command exists for the obsolete-tray-helper defect in RC 4."}, {"id": "a8", "statement": "The obsolete-tray-helper defect in RC 4 is open."}, {"id": "a9", "statement": "The obsolete-tray-helper defect in RC 4 has medium severity."}, {"id": "a10", "statement": "The obsolete-tray-helper defect in RC 4 is a compatibility defect."}, {"id": "a11", "statement": "Every open packaging or compatibility defect in RC 4 other than the obsolete-tray-helper defect has severity below medium."}, {"id": "a12", "statement": "Every RC 4 installer has the expected publisher identity."}, {"id": "a13", "statement": "Every RC 4 installer digest matches its build record."}, {"id": "a14", "statement": "The RC 4 removal package restored the prior version 7.3."}, {"id": "a15", "statement": "The RC 4 removal package retained every user file during restoration of version 7.3."}, {"id": "a16", "statement": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed.\"},{\"speaker\":\"Platform lead\",\"text\":\"Every required-platform check for RC 4 passed except the supported-Ubuntu check affected by the obsolete-tray-helper defect. Without remediation, that check fails, and no tested workaround other than the cleanup command exists.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing.\"},{\"speaker\":\"Release manager\",\"text\":\"The obsolete-tray-helper defect is open, medium severity, and a compatibility defect. Every other open packaging or compatibility defect has severity below medium.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"Every RC 4 installer has the expected publisher identity, and every installer digest matches its build record. The removal package restored prior version 7.3 and retained every user file during restoration.\"},{\"speaker\":\"Documentation lead\",\"text\":\"Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build."}, {"path": ["3", "text"], "text": "The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing."}], "policy_evidence": [], "rules": [{"justification": "The required supported-Ubuntu platform check fails, the cleanup-command trial does not make it pass, and no other tested workaround exists; this satisfies the blocked criterion.", "target": "0", "when": [{"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "There are no blocker or high-severity defects; the core and required-platform gates pass, including the Ubuntu gate through a tested cleanup workaround; signing, integrity, and rollback are verified; and exactly one medium compatibility defect remains with documentation approval pending.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build.", "negative_left": "For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build.", "negative_right": "The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as failing.", "right": "The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing."}, "verifier_independent_model": false}, "family": "fast-41-diverse-150-014", "id": "fast-41-diverse-150-014-base", "input": {"questions": {"decision": {"criteria": ["0 — Blocked: At least one blocker/high-severity defect is open, a required platform or core test fails without a tested workaround, installer signing/integrity is unverified, or rollback has not been validated.", "1 — Conditionally ready: No blocker/high-severity defects remain; core tests, required-platform checks, installer signing/integrity, and rollback pass, but exactly one medium packaging or compatibility defect remains with a tested workaround that still needs final documentation or release-operations approval.", "2 — Fully ready: All core gates pass, rollback and installer integrity are validated, no medium-or-higher packaging or compatibility defects remain, and all release notes and operational approvals are final."], "instructions": "Assign the release-readiness level. Interpret evidence by meaning rather than requiring exact rubric phrases; for example, restoring the prior version without losing files constitutes validated rollback, and matching publisher identity/digests constitutes signing and integrity verification.", "type": "score"}}, "state": [{"speaker": "Release manager", "text": "RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed."}, {"speaker": "Platform lead", "text": "Every required-platform check for RC 4 passed except the supported-Ubuntu check affected by the obsolete-tray-helper defect. Without remediation, that check fails, and no tested workaround other than the cleanup command exists."}, {"speaker": "Compatibility tester", "text": "For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build."}, {"speaker": "Compatibility tester", "text": "The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing."}, {"speaker": "Release manager", "text": "The obsolete-tray-helper defect is open, medium severity, and a compatibility defect. Every other open packaging or compatibility defect has severity below medium."}, {"speaker": "Deployment specialist", "text": "Every RC 4 installer has the expected publisher identity, and every installer digest matches its build record. The removal package restored prior version 7.3 and retained every user file during restoration."}, {"speaker": "Documentation lead", "text": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}]}, "method": "c2d", "provenance": {"source_id": "diverse-150", "source_is_synthetic": true, "source_sha256": "debcb2db8f1984d1db79d80259af413a1a60045f28238c51df71d3deaccb8e61", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and bindings, the focus evidence is complete and exact—“For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build.” and “The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing.”—and the counterfactual coherently changes only the trial result without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom a6 is factual rather than a policy classification. The base and counter assignments are jointly realizable while changing only whether the cleanup-command trial passes; a7 merely excludes other tested workarounds and need not assert that the cleanup command succeeds. Empty policy_evidence is correct because all governing readiness criteria and interpretive instructions occur in the retained questions object, while the original state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The Ubuntu required-platform check fails without remediation, the cleanup command does not produce a passing result, and no other tested workaround exists. Therefore a required-platform check fails without a tested workaround, which is sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes no open blocker/high-severity defects, passing core and required-platform checks through the tested Ubuntu workaround, verified installer signing/integrity, validated rollback, and exactly one open medium packaging-or-compatibility defect whose documentation approval remains pending. This is sufficient for level 1 and excludes the stated level-0 conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of open blocker defects for RC 4 is zero."}, {"id": "a2", "statement": "The number of open high-severity defects for RC 4 is zero."}, {"id": "a3", "statement": "Each of the 186 required smoke and regression tests for RC 4 passed."}, {"id": "a4", "statement": "Every required-platform check for RC 4 other than the supported-Ubuntu obsolete-tray-helper check passed."}, {"id": "a5", "statement": "Without remediation, the obsolete-tray-helper defect causes its required-platform check on the supported Ubuntu build for RC 4 to fail."}, {"id": "a6", "statement": "The cleanup-command trial for the obsolete-tray-helper defect produced a passing result for its required-platform check on the supported Ubuntu build for RC 4."}, {"id": "a7", "statement": "No tested workaround other than the cleanup command exists for the obsolete-tray-helper defect in RC 4."}, {"id": "a8", "statement": "The obsolete-tray-helper defect in RC 4 is open."}, {"id": "a9", "statement": "The obsolete-tray-helper defect in RC 4 has medium severity."}, {"id": "a10", "statement": "The obsolete-tray-helper defect in RC 4 is a compatibility defect."}, {"id": "a11", "statement": "Every open packaging or compatibility defect in RC 4 other than the obsolete-tray-helper defect has severity below medium."}, {"id": "a12", "statement": "Every RC 4 installer has the expected publisher identity."}, {"id": "a13", "statement": "Every RC 4 installer digest matches its build record."}, {"id": "a14", "statement": "The RC 4 removal package restored the prior version 7.3."}, {"id": "a15", "statement": "The RC 4 removal package retained every user file during restoration of version 7.3."}, {"id": "a16", "statement": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed.\"},{\"speaker\":\"Platform lead\",\"text\":\"Every required-platform check for RC 4 passed except the supported-Ubuntu check affected by the obsolete-tray-helper defect. Without remediation, that check fails, and no tested workaround other than the cleanup command exists.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing.\"},{\"speaker\":\"Release manager\",\"text\":\"The obsolete-tray-helper defect is open, medium severity, and a compatibility defect. Every other open packaging or compatibility defect has severity below medium.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"Every RC 4 installer has the expected publisher identity, and every installer digest matches its build record. The removal package restored prior version 7.3 and retained every user file during restoration.\"},{\"speaker\":\"Documentation lead\",\"text\":\"Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build."}, {"path": ["3", "text"], "text": "The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing."}], "policy_evidence": [], "rules": [{"justification": "The required supported-Ubuntu platform check fails, the cleanup-command trial does not make it pass, and no other tested workaround exists; this satisfies the blocked criterion.", "target": "0", "when": [{"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "There are no blocker or high-severity defects; the core and required-platform gates pass, including the Ubuntu gate through a tested cleanup workaround; signing, integrity, and rollback are verified; and exactly one medium compatibility defect remains with documentation approval pending.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build.", "negative_left": "For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build.", "negative_right": "The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as failing.", "right": "The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as passing."}, "verifier_independent_model": false}, "family": "fast-41-diverse-150-014", "id": "fast-41-diverse-150-014-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Blocked: At least one blocker/high-severity defect is open, a required platform or core test fails without a tested workaround, installer signing/integrity is unverified, or rollback has not been validated.", "1 — Conditionally ready: No blocker/high-severity defects remain; core tests, required-platform checks, installer signing/integrity, and rollback pass, but exactly one medium packaging or compatibility defect remains with a tested workaround that still needs final documentation or release-operations approval.", "2 — Fully ready: All core gates pass, rollback and installer integrity are validated, no medium-or-higher packaging or compatibility defects remain, and all release notes and operational approvals are final."], "instructions": "Assign the release-readiness level. Interpret evidence by meaning rather than requiring exact rubric phrases; for example, restoring the prior version without losing files constitutes validated rollback, and matching publisher identity/digests constitutes signing and integrity verification.", "type": "score"}}, "state": [{"speaker": "Release manager", "text": "RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed."}, {"speaker": "Platform lead", "text": "Every required-platform check for RC 4 passed except the supported-Ubuntu check affected by the obsolete-tray-helper defect. Without remediation, that check fails, and no tested workaround other than the cleanup command exists."}, {"speaker": "Compatibility tester", "text": "For the obsolete-tray-helper defect in RC 4, the cleanup-command trial was the trial designated to remediate the required-platform check on the supported Ubuntu build."}, {"speaker": "Compatibility tester", "text": "The designated trial for the obsolete-tray-helper defect in RC 4 was recorded as failing."}, {"speaker": "Release manager", "text": "The obsolete-tray-helper defect is open, medium severity, and a compatibility defect. Every other open packaging or compatibility defect has severity below medium."}, {"speaker": "Deployment specialist", "text": "Every RC 4 installer has the expected publisher identity, and every installer digest matches its build record. The removal package restored prior version 7.3 and retained every user file during restoration."}, {"speaker": "Documentation lead", "text": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}]}, "method": "c2d", "provenance": {"source_id": "diverse-150", "source_is_synthetic": true, "source_sha256": "debcb2db8f1984d1db79d80259af413a1a60045f28238c51df71d3deaccb8e61", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, bindings remain about RC 4 release readiness, both evidence spans are complete factual sentences, the changed failed check is consistent with the surrounding facts, and neither context includes an encoded answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom a6 is factual rather than a policy classification. The base and counter assignments are jointly realizable while changing only whether the cleanup-command trial passes; a7 merely excludes other tested workarounds and need not assert that the cleanup command succeeds. Empty policy_evidence is correct because all governing readiness criteria and interpretive instructions occur in the retained questions object, while the original state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The Ubuntu required-platform check fails without remediation, the cleanup command does not produce a passing result, and no other tested workaround exists. Therefore a required-platform check fails without a tested workaround, which is sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes no open blocker/high-severity defects, passing core and required-platform checks through the tested Ubuntu workaround, verified installer signing/integrity, validated rollback, and exactly one open medium packaging-or-compatibility defect whose documentation approval remains pending. This is sufficient for level 1 and excludes the stated level-0 conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of open blocker defects for RC 4 is zero."}, {"id": "a2", "statement": "The number of open high-severity defects for RC 4 is zero."}, {"id": "a3", "statement": "Each of the 186 required smoke and regression tests for RC 4 passed."}, {"id": "a4", "statement": "Every required-platform check for RC 4 other than the supported-Ubuntu obsolete-tray-helper check passed."}, {"id": "a5", "statement": "Without remediation, the obsolete-tray-helper defect causes its required-platform check on the supported Ubuntu build for RC 4 to fail."}, {"id": "a6", "statement": "The cleanup-command trial for the obsolete-tray-helper defect produced a passing result for its required-platform check on the supported Ubuntu build for RC 4."}, {"id": "a7", "statement": "No tested workaround other than the cleanup command exists for the obsolete-tray-helper defect in RC 4."}, {"id": "a8", "statement": "The obsolete-tray-helper defect in RC 4 is open."}, {"id": "a9", "statement": "The obsolete-tray-helper defect in RC 4 has medium severity."}, {"id": "a10", "statement": "The obsolete-tray-helper defect in RC 4 is a compatibility defect."}, {"id": "a11", "statement": "Every open packaging or compatibility defect in RC 4 other than the obsolete-tray-helper defect has severity below medium."}, {"id": "a12", "statement": "Every RC 4 installer has the expected publisher identity."}, {"id": "a13", "statement": "Every RC 4 installer digest matches its build record."}, {"id": "a14", "statement": "The RC 4 removal package restored the prior version 7.3."}, {"id": "a15", "statement": "The RC 4 removal package retained every user file during restoration of version 7.3."}, {"id": "a16", "statement": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed. Every required-platform check passed except the supported-Ubuntu check for the obsolete-tray-helper defect.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"Without remediation, the obsolete-tray-helper defect causes that supported-Ubuntu check to fail. The defect is open, medium severity, and a compatibility defect; every other open packaging or compatibility defect is below medium.\"},{\"speaker\":\"Test lead\",\"text\":\"For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect.\"},{\"speaker\":\"Test lead\",\"text\":\"Test record UBT-19 marks the required-platform check as passed.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"No tested workaround other than the cleanup command exists. Every RC 4 installer has the expected publisher identity, and every digest matches its build record. The removal package restored version 7.3 while retaining every user file.\"},{\"speaker\":\"Release manager\",\"text\":\"Final approval of the release-note documentation for the obsolete-tray-helper defect remains pending.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect."}, {"path": ["3", "text"], "text": "Test record UBT-19 marks the required-platform check as passed."}], "policy_evidence": [], "rules": [{"justification": "The required supported-Ubuntu platform check fails, the cleanup-command trial does not make it pass, and no other tested workaround exists; this satisfies the blocked criterion.", "target": "0", "when": [{"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "There are no blocker or high-severity defects; the core and required-platform gates pass, including the Ubuntu gate through a tested cleanup workaround; signing, integrity, and rollback are verified; and exactly one medium compatibility defect remains with documentation approval pending.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect.", "negative_left": "For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect.", "negative_right": "Test record UBT-19 marks the required-platform check as failed.", "right": "Test record UBT-19 marks the required-platform check as passed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-150-015", "id": "fast-41-diverse-150-015-base", "input": {"questions": {"decision": {"criteria": ["0 — Blocked: At least one blocker/high-severity defect is open, a required platform or core test fails without a tested workaround, installer signing/integrity is unverified, or rollback has not been validated.", "1 — Conditionally ready: No blocker/high-severity defects remain; core tests, required-platform checks, installer signing/integrity, and rollback pass, but exactly one medium packaging or compatibility defect remains with a tested workaround that still needs final documentation or release-operations approval.", "2 — Fully ready: All core gates pass, rollback and installer integrity are validated, no medium-or-higher packaging or compatibility defects remain, and all release notes and operational approvals are final."], "instructions": "Assign the release-readiness level. Interpret evidence by meaning rather than requiring exact rubric phrases; for example, restoring the prior version without losing files constitutes validated rollback, and matching publisher identity/digests constitutes signing and integrity verification.", "type": "score"}}, "state": [{"speaker": "Release manager", "text": "RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed. Every required-platform check passed except the supported-Ubuntu check for the obsolete-tray-helper defect."}, {"speaker": "Compatibility tester", "text": "Without remediation, the obsolete-tray-helper defect causes that supported-Ubuntu check to fail. The defect is open, medium severity, and a compatibility defect; every other open packaging or compatibility defect is below medium."}, {"speaker": "Test lead", "text": "For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect."}, {"speaker": "Test lead", "text": "Test record UBT-19 marks the required-platform check as passed."}, {"speaker": "Deployment specialist", "text": "No tested workaround other than the cleanup command exists. Every RC 4 installer has the expected publisher identity, and every digest matches its build record. The removal package restored version 7.3 while retaining every user file."}, {"speaker": "Release manager", "text": "Final approval of the release-note documentation for the obsolete-tray-helper defect remains pending."}]}, "method": "c2d", "provenance": {"source_id": "diverse-150", "source_is_synthetic": true, "source_sha256": "debcb2db8f1984d1db79d80259af413a1a60045f28238c51df71d3deaccb8e61", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, bindings remain about RC 4 release readiness, both evidence spans are complete factual sentences, the changed failed check is consistent with the surrounding facts, and neither context includes an encoded answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom a6 is factual rather than a policy classification. The base and counter assignments are jointly realizable while changing only whether the cleanup-command trial passes; a7 merely excludes other tested workarounds and need not assert that the cleanup command succeeds. Empty policy_evidence is correct because all governing readiness criteria and interpretive instructions occur in the retained questions object, while the original state contains case observations rather than additional substantive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The Ubuntu required-platform check fails without remediation, the cleanup command does not produce a passing result, and no other tested workaround exists. Therefore a required-platform check fails without a tested workaround, which is sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes no open blocker/high-severity defects, passing core and required-platform checks through the tested Ubuntu workaround, verified installer signing/integrity, validated rollback, and exactly one open medium packaging-or-compatibility defect whose documentation approval remains pending. This is sufficient for level 1 and excludes the stated level-0 conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of open blocker defects for RC 4 is zero."}, {"id": "a2", "statement": "The number of open high-severity defects for RC 4 is zero."}, {"id": "a3", "statement": "Each of the 186 required smoke and regression tests for RC 4 passed."}, {"id": "a4", "statement": "Every required-platform check for RC 4 other than the supported-Ubuntu obsolete-tray-helper check passed."}, {"id": "a5", "statement": "Without remediation, the obsolete-tray-helper defect causes its required-platform check on the supported Ubuntu build for RC 4 to fail."}, {"id": "a6", "statement": "The cleanup-command trial for the obsolete-tray-helper defect produced a passing result for its required-platform check on the supported Ubuntu build for RC 4."}, {"id": "a7", "statement": "No tested workaround other than the cleanup command exists for the obsolete-tray-helper defect in RC 4."}, {"id": "a8", "statement": "The obsolete-tray-helper defect in RC 4 is open."}, {"id": "a9", "statement": "The obsolete-tray-helper defect in RC 4 has medium severity."}, {"id": "a10", "statement": "The obsolete-tray-helper defect in RC 4 is a compatibility defect."}, {"id": "a11", "statement": "Every open packaging or compatibility defect in RC 4 other than the obsolete-tray-helper defect has severity below medium."}, {"id": "a12", "statement": "Every RC 4 installer has the expected publisher identity."}, {"id": "a13", "statement": "Every RC 4 installer digest matches its build record."}, {"id": "a14", "statement": "The RC 4 removal package restored the prior version 7.3."}, {"id": "a15", "statement": "The RC 4 removal package retained every user file during restoration of version 7.3."}, {"id": "a16", "statement": "Final approval of the release-note documentation for the obsolete-tray-helper defect in RC 4 is pending."}], "base_state_json": "[{\"speaker\":\"Release manager\",\"text\":\"RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed. Every required-platform check passed except the supported-Ubuntu check for the obsolete-tray-helper defect.\"},{\"speaker\":\"Compatibility tester\",\"text\":\"Without remediation, the obsolete-tray-helper defect causes that supported-Ubuntu check to fail. The defect is open, medium severity, and a compatibility defect; every other open packaging or compatibility defect is below medium.\"},{\"speaker\":\"Test lead\",\"text\":\"For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect.\"},{\"speaker\":\"Test lead\",\"text\":\"Test record UBT-19 marks the required-platform check as passed.\"},{\"speaker\":\"Deployment specialist\",\"text\":\"No tested workaround other than the cleanup command exists. Every RC 4 installer has the expected publisher identity, and every digest matches its build record. The removal package restored version 7.3 while retaining every user file.\"},{\"speaker\":\"Release manager\",\"text\":\"Final approval of the release-note documentation for the obsolete-tray-helper defect remains pending.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect."}, {"path": ["3", "text"], "text": "Test record UBT-19 marks the required-platform check as passed."}], "policy_evidence": [], "rules": [{"justification": "The required supported-Ubuntu platform check fails, the cleanup-command trial does not make it pass, and no other tested workaround exists; this satisfies the blocked criterion.", "target": "0", "when": [{"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "There are no blocker or high-severity defects; the core and required-platform gates pass, including the Ubuntu gate through a tested cleanup workaround; signing, integrity, and rollback are verified; and exactly one medium compatibility defect remains with documentation approval pending.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect.", "negative_left": "For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect.", "negative_right": "Test record UBT-19 marks the required-platform check as failed.", "right": "Test record UBT-19 marks the required-platform check as passed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-150-015", "id": "fast-41-diverse-150-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Blocked: At least one blocker/high-severity defect is open, a required platform or core test fails without a tested workaround, installer signing/integrity is unverified, or rollback has not been validated.", "1 — Conditionally ready: No blocker/high-severity defects remain; core tests, required-platform checks, installer signing/integrity, and rollback pass, but exactly one medium packaging or compatibility defect remains with a tested workaround that still needs final documentation or release-operations approval.", "2 — Fully ready: All core gates pass, rollback and installer integrity are validated, no medium-or-higher packaging or compatibility defects remain, and all release notes and operational approvals are final."], "instructions": "Assign the release-readiness level. Interpret evidence by meaning rather than requiring exact rubric phrases; for example, restoring the prior version without losing files constitutes validated rollback, and matching publisher identity/digests constitutes signing and integrity verification.", "type": "score"}}, "state": [{"speaker": "Release manager", "text": "RC 4 has zero open blocker defects and zero open high-severity defects. All 186 required smoke and regression tests passed. Every required-platform check passed except the supported-Ubuntu check for the obsolete-tray-helper defect."}, {"speaker": "Compatibility tester", "text": "Without remediation, the obsolete-tray-helper defect causes that supported-Ubuntu check to fail. The defect is open, medium severity, and a compatibility defect; every other open packaging or compatibility defect is below medium."}, {"speaker": "Test lead", "text": "For RC 4, the cleanup-command trial identified in test record UBT-19 was conducted on the supported Ubuntu build for the obsolete-tray-helper defect."}, {"speaker": "Test lead", "text": "Test record UBT-19 marks the required-platform check as failed."}, {"speaker": "Deployment specialist", "text": "No tested workaround other than the cleanup command exists. Every RC 4 installer has the expected publisher identity, and every digest matches its build record. The removal package restored version 7.3 while retaining every user file."}, {"speaker": "Release manager", "text": "Final approval of the release-note documentation for the obsolete-tray-helper defect remains pending."}]}, "method": "c2d", "provenance": {"source_id": "diverse-150", "source_is_synthetic": true, "source_sha256": "debcb2db8f1984d1db79d80259af413a1a60045f28238c51df71d3deaccb8e61", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings; the exact evidence quotes are \"The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18.\" and \"Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m.\"; the counterfactual changes only P-18 and remains coherent without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The missed-pickup claim concerns 18 Larch Court in Zone R4, where recycling was scheduled for Tuesday. The report was received within the 48-hour deadline. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18. Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m. No obstruction, leaking waste, or visible pests are documented for the claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18."}, {"path": ["context"], "text": "Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18.", "negative_left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18.", "negative_right": "Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart curbside beside 18 Larch Court at 6:55 a.m.", "right": "Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-005", "id": "fast-41-diverse-151-005-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The missed-pickup claim concerns 18 Larch Court in Zone R4, where recycling was scheduled for Tuesday. The report was received within the 48-hour deadline. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18. Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m. No obstruction, leaking waste, or visible pests are documented for the claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings; the exact evidence quotes are \"The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18.\" and \"Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m.\"; the counterfactual changes only P-18 and remains coherent without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The missed-pickup claim concerns 18 Larch Court in Zone R4, where recycling was scheduled for Tuesday. The report was received within the 48-hour deadline. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18. Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m. No obstruction, leaking waste, or visible pests are documented for the claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18."}, {"path": ["context"], "text": "Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18.", "negative_left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18.", "negative_right": "Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart curbside beside 18 Larch Court at 6:55 a.m.", "right": "Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart inside the garage at 6:55 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-005", "id": "fast-41-diverse-151-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The missed-pickup claim concerns 18 Larch Court in Zone R4, where recycling was scheduled for Tuesday. The report was received within the 48-hour deadline. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs P-17 and P-18. Photograph P-17 shows the blue cart in an enclosed rear yard behind a locked fence at 7:05 a.m., and photograph P-18 shows the blue cart curbside beside 18 Larch Court at 6:55 a.m. No obstruction, leaking waste, or visible pests are documented for the claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions and governing policy. Entity, waste-stream, date, and routing-request bindings remain intact. The two focus spans are complete factual sentences. The counterfactual changes only the placement time from 7:05 a.m. to 6:55 a.m. and introduces no contradiction. Neither context contains an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The address is 18 Larch Court, which is in Zone R4, and recycling was scheduled there for Tuesday.\",\"The missed-pickup report for Tuesday recycling at 18 Larch Court arrived within the 48-hour reporting deadline.\",\"The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb.\",\"The photograph in that submission was taken at 7:05 a.m. on Tuesday.\",\"The Tuesday recycling claim documents no obstruction, leaking waste, or visible pests.\",\"The resident requests a routing decision for this missed-pickup claim.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb."}, {"path": ["evidence", "3"], "text": "The photograph in that submission was taken at 7:05 a.m. on Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb.", "negative_left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb.", "negative_right": "The photograph in that submission was taken at 6:55 a.m. on Tuesday.", "right": "The photograph in that submission was taken at 7:05 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-006", "id": "fast-41-diverse-151-006-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The address is 18 Larch Court, which is in Zone R4, and recycling was scheduled there for Tuesday.", "The missed-pickup report for Tuesday recycling at 18 Larch Court arrived within the 48-hour reporting deadline.", "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb.", "The photograph in that submission was taken at 7:05 a.m. on Tuesday.", "The Tuesday recycling claim documents no obstruction, leaking waste, or visible pests.", "The resident requests a routing decision for this missed-pickup claim."]}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions and governing policy. Entity, waste-stream, date, and routing-request bindings remain intact. The two focus spans are complete factual sentences. The counterfactual changes only the placement time from 7:05 a.m. to 6:55 a.m. and introduces no contradiction. Neither context contains an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The address is 18 Larch Court, which is in Zone R4, and recycling was scheduled there for Tuesday.\",\"The missed-pickup report for Tuesday recycling at 18 Larch Court arrived within the 48-hour reporting deadline.\",\"The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb.\",\"The photograph in that submission was taken at 7:05 a.m. on Tuesday.\",\"The Tuesday recycling claim documents no obstruction, leaking waste, or visible pests.\",\"The resident requests a routing decision for this missed-pickup claim.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb."}, {"path": ["evidence", "3"], "text": "The photograph in that submission was taken at 7:05 a.m. on Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb.", "negative_left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb.", "negative_right": "The photograph in that submission was taken at 6:55 a.m. on Tuesday.", "right": "The photograph in that submission was taken at 7:05 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-006", "id": "fast-41-diverse-151-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The address is 18 Larch Court, which is in Zone R4, and recycling was scheduled there for Tuesday.", "The missed-pickup report for Tuesday recycling at 18 Larch Court arrived within the 48-hour reporting deadline.", "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court contains one item, a photograph of the blue cart taken at the curb.", "The photograph in that submission was taken at 6:55 a.m. on Tuesday.", "The Tuesday recycling claim documents no obstruction, leaking waste, or visible pests.", "The resident requests a routing decision for this missed-pickup claim."]}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, and the original questions object preserves all question rules and criteria. Entity, waste-stream, deadline, curb-placement, and hazard bindings remain consistent. Each evidence item is a complete factual sentence. The counterfactual changes P12 to verified without contradicting its other facts. Neither context contains an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The address, 18 Larch Court, is recorded in Zone R4, where recycling was scheduled for Tuesday.\",\"The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline.\",\"The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14.\",\"The verification fields for P12, P13, and P14 are blank for the blue cart's curbside check at 7:00 a.m.\",\"The case file documents no obstruction, leaking waste, or visible pests for the Tuesday recycling claim at 18 Larch Court.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14."}, {"path": ["evidence", "3"], "text": "The verification fields for P12, P13, and P14 are blank for the blue cart's curbside check at 7:00 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14.", "negative_left": "The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14.", "negative_right": "The verification field for P12 is marked verified for the blue cart's curbside check at 7:00 a.m.", "right": "The verification fields for P12, P13, and P14 are blank for the blue cart's curbside check at 7:00 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-012", "id": "fast-41-diverse-151-012-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The address, 18 Larch Court, is recorded in Zone R4, where recycling was scheduled for Tuesday.", "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline.", "The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14.", "The verification fields for P12, P13, and P14 are blank for the blue cart's curbside check at 7:00 a.m.", "The case file documents no obstruction, leaking waste, or visible pests for the Tuesday recycling claim at 18 Larch Court."]}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, and the original questions object preserves all question rules and criteria. Entity, waste-stream, deadline, curb-placement, and hazard bindings remain consistent. Each evidence item is a complete factual sentence. The counterfactual changes P12 to verified without contradicting its other facts. Neither context contains an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"evidence\":[\"The address, 18 Larch Court, is recorded in Zone R4, where recycling was scheduled for Tuesday.\",\"The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline.\",\"The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14.\",\"The verification fields for P12, P13, and P14 are blank for the blue cart's curbside check at 7:00 a.m.\",\"The case file documents no obstruction, leaking waste, or visible pests for the Tuesday recycling claim at 18 Larch Court.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14."}, {"path": ["evidence", "3"], "text": "The verification fields for P12, P13, and P14 are blank for the blue cart's curbside check at 7:00 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14.", "negative_left": "The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14.", "negative_right": "The verification field for P12 is marked verified for the blue cart's curbside check at 7:00 a.m.", "right": "The verification fields for P12, P13, and P14 are blank for the blue cart's curbside check at 7:00 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-012", "id": "fast-41-diverse-151-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "evidence": ["The address, 18 Larch Court, is recorded in Zone R4, where recycling was scheduled for Tuesday.", "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline.", "The placement-evidence packet for the Tuesday recycling claim at 18 Larch Court contains exactly three submitted items, labeled P12, P13, and P14.", "The verification field for P12 is marked verified for the blue cart's curbside check at 7:00 a.m.", "The case file documents no obstruction, leaking waste, or visible pests for the Tuesday recycling claim at 18 Larch Court."]}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The verbatim questions preserve the governing policy, criteria, instructions, request scope, and bindings. Each context has two complete factual evidence sentences. Base evidence: \"The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.\" and \"Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m.\" Counterfactual evidence: \"The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.\" and \"At least one of the timestamped photograph and signed witness statement verifies that the blue cart was curbside by 7:00 a.m.\" The counterfactual changes only the placement finding and creates no contradiction. Neither context contains an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The case concerns 18 Larch Court, which is in Zone R4, where recycling was scheduled for Tuesday. The report was received within the 48-hour deadline. The file documents no obstruction, leaking waste, or visible pests. The placement materials are identified below.\",\"evidence\":[\"The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.\",\"Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "0"], "text": "The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement."}, {"path": ["evidence", "1"], "text": "Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.", "negative_left": "The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.", "negative_right": "At least one of the timestamped photograph and signed witness statement verifies that the blue cart was curbside by 7:00 a.m.", "right": "Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-020", "id": "fast-41-diverse-151-020-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The case concerns 18 Larch Court, which is in Zone R4, where recycling was scheduled for Tuesday. The report was received within the 48-hour deadline. The file documents no obstruction, leaking waste, or visible pests. The placement materials are identified below.", "evidence": ["The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.", "Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m."]}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The verbatim questions preserve the governing policy, criteria, instructions, request scope, and bindings. Each context has two complete factual evidence sentences. Base evidence: \"The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.\" and \"Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m.\" Counterfactual evidence: \"The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.\" and \"At least one of the timestamped photograph and signed witness statement verifies that the blue cart was curbside by 7:00 a.m.\" The counterfactual changes only the placement finding and creates no contradiction. Neither context contains an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The case concerns 18 Larch Court, which is in Zone R4, where recycling was scheduled for Tuesday. The report was received within the 48-hour deadline. The file documents no obstruction, leaking waste, or visible pests. The placement materials are identified below.\",\"evidence\":[\"The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.\",\"Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "0"], "text": "The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement."}, {"path": ["evidence", "1"], "text": "Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.", "negative_left": "The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.", "negative_right": "At least one of the timestamped photograph and signed witness statement verifies that the blue cart was curbside by 7:00 a.m.", "right": "Neither the timestamped photograph nor the signed witness statement verifies that the blue cart was curbside by 7:00 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-020", "id": "fast-41-diverse-151-020-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The case concerns 18 Larch Court, which is in Zone R4, where recycling was scheduled for Tuesday. The report was received within the 48-hour deadline. The file documents no obstruction, leaking waste, or visible pests. The placement materials are identified below.", "evidence": ["The complete set of submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of a timestamped photograph and a signed witness statement.", "At least one of the timestamped photograph and signed witness statement verifies that the blue cart was curbside by 7:00 a.m."]}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, and the original questions object is preserved verbatim. The claim remains tied to recycling at 18 Larch Court on Tuesday with the applicable timing and placement threshold. The two evidence spans are complete factual sentences. The counterfactual changes only the cart’s placement and remains internally consistent with the single photograph. Neither context states a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The missed-pickup claim concerns recycling at 18 Larch Court. Municipal records place the address in Zone R4, whose recycling route was scheduled for Tuesday. The report was logged within the applicable 48-hour deadline. The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417. Photograph P-417 depicts the blue cart inside the fenced yard, rather than at the curb, at 6:52 a.m. on Tuesday. The file was attached to the resident’s claim for review. The incident record documents no blocked travel lane or other obstruction, no leaking waste, and no visible pests. The dispatcher is reviewing the claim under the following governing policy: A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417."}, {"path": ["context"], "text": "Photograph P-417 depicts the blue cart inside the fenced yard, rather than at the curb, at 6:52 a.m. on Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417.", "negative_left": "The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417.", "negative_right": "Photograph P-417 depicts the blue cart at the curb at 6:52 a.m. on Tuesday.", "right": "Photograph P-417 depicts the blue cart inside the fenced yard, rather than at the curb, at 6:52 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-034", "id": "fast-41-diverse-151-034-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The missed-pickup claim concerns recycling at 18 Larch Court. Municipal records place the address in Zone R4, whose recycling route was scheduled for Tuesday. The report was logged within the applicable 48-hour deadline. The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417. Photograph P-417 depicts the blue cart inside the fenced yard, rather than at the curb, at 6:52 a.m. on Tuesday. The file was attached to the resident’s claim for review. The incident record documents no blocked travel lane or other obstruction, no leaking waste, and no visible pests. The dispatcher is reviewing the claim under the following governing policy: A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, and the original questions object is preserved verbatim. The claim remains tied to recycling at 18 Larch Court on Tuesday with the applicable timing and placement threshold. The two evidence spans are complete factual sentences. The counterfactual changes only the cart’s placement and remains internally consistent with the single photograph. Neither context states a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The missed-pickup claim concerns recycling at 18 Larch Court. Municipal records place the address in Zone R4, whose recycling route was scheduled for Tuesday. The report was logged within the applicable 48-hour deadline. The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417. Photograph P-417 depicts the blue cart inside the fenced yard, rather than at the curb, at 6:52 a.m. on Tuesday. The file was attached to the resident’s claim for review. The incident record documents no blocked travel lane or other obstruction, no leaking waste, and no visible pests. The dispatcher is reviewing the claim under the following governing policy: A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417."}, {"path": ["context"], "text": "Photograph P-417 depicts the blue cart inside the fenced yard, rather than at the curb, at 6:52 a.m. on Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417.", "negative_left": "The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417.", "negative_right": "Photograph P-417 depicts the blue cart at the curb at 6:52 a.m. on Tuesday.", "right": "Photograph P-417 depicts the blue cart inside the fenced yard, rather than at the curb, at 6:52 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-034", "id": "fast-41-diverse-151-034-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The missed-pickup claim concerns recycling at 18 Larch Court. Municipal records place the address in Zone R4, whose recycling route was scheduled for Tuesday. The report was logged within the applicable 48-hour deadline. The complete placement-evidence file submitted for the Tuesday recycling claim at 18 Larch Court contains exactly one item, photograph P-417. Photograph P-417 depicts the blue cart at the curb at 6:52 a.m. on Tuesday. The file was attached to the resident’s claim for review. The incident record documents no blocked travel lane or other obstruction, no leaking waste, and no visible pests. The dispatcher is reviewing the claim under the following governing policy: A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, the evidence contains the exact factual sentences \"For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record.\" and \"Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m.\", the counterfactual coherently changes only that observation, and neither context leaks an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"case_note\":\"A resident filed a missed-pickup claim for 18 Larch Court. The address is assigned to Zone R4, where recycling collection was scheduled for Tuesday. The report arrived within the applicable 48-hour deadline. For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record. Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m. The case file documents no blocked travel lane or other obstruction, no leaking waste, and no visible pests. The collection stream and address were checked against the route schedule, and the report timestamp was logged by intake. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The intake clerk attached the route lookup, deadline record, and hazard review to the claim for administrative handling.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["case_note"], "text": "For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record."}, {"path": ["case_note"], "text": "Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record.", "negative_left": "For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record.", "negative_right": "Record PE-472 contains a time-stamped image of the curb that verifies that the blue cart was curbside by 7:00 a.m.", "right": "Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-038", "id": "fast-41-diverse-151-038-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"case_note": "A resident filed a missed-pickup claim for 18 Larch Court. The address is assigned to Zone R4, where recycling collection was scheduled for Tuesday. The report arrived within the applicable 48-hour deadline. For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record. Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m. The case file documents no blocked travel lane or other obstruction, no leaking waste, and no visible pests. The collection stream and address were checked against the route schedule, and the report timestamp was logged by intake. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The intake clerk attached the route lookup, deadline record, and hazard review to the claim for administrative handling."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, the evidence contains the exact factual sentences \"For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record.\" and \"Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m.\", the counterfactual coherently changes only that observation, and neither context leaks an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"case_note\":\"A resident filed a missed-pickup claim for 18 Larch Court. The address is assigned to Zone R4, where recycling collection was scheduled for Tuesday. The report arrived within the applicable 48-hour deadline. For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record. Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m. The case file documents no blocked travel lane or other obstruction, no leaking waste, and no visible pests. The collection stream and address were checked against the route schedule, and the report timestamp was logged by intake. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The intake clerk attached the route lookup, deadline record, and hazard review to the claim for administrative handling.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["case_note"], "text": "For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record."}, {"path": ["case_note"], "text": "Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record.", "negative_left": "For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record.", "negative_right": "Record PE-472 contains a time-stamped image of the curb that verifies that the blue cart was curbside by 7:00 a.m.", "right": "Record PE-472 contains a time-stamped image of the curb, but the image does not verify that the blue cart was curbside by 7:00 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-038", "id": "fast-41-diverse-151-038-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"case_note": "A resident filed a missed-pickup claim for 18 Larch Court. The address is assigned to Zone R4, where recycling collection was scheduled for Tuesday. The report arrived within the applicable 48-hour deadline. For the Tuesday recycling claim at 18 Larch Court, PE-472 is the sole submitted placement-evidence record. Record PE-472 contains a time-stamped image of the curb that verifies that the blue cart was curbside by 7:00 a.m. The case file documents no blocked travel lane or other obstruction, no leaking waste, and no visible pests. The collection stream and address were checked against the route schedule, and the report timestamp was logged by intake. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The intake clerk attached the route lookup, deadline record, and hazard review to the claim for administrative handling."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and unchanged question bindings. The evidence spans are complete factual sentences: \"The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317.\" and \"Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside.\" The counterfactual changes only the placement observation and remains coherent. Neither context contains a gold answer, answer code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"case_note\":\"A missed-pickup claim was filed for 18 Larch Court, which is assigned to Zone R4. Recycling service in that zone was scheduled for Tuesday, and the report was received within the applicable 48-hour reporting deadline. A resident submitted a placement record for the claim. The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317. Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside. The file contains no documented obstruction, no leaking waste, and no visible pests associated with the claim. A routing decision is requested. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["case_note"], "text": "The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317."}, {"path": ["case_note"], "text": "Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317.", "negative_left": "The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317.", "negative_right": "Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was curbside.", "right": "Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-044", "id": "fast-41-diverse-151-044-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"case_note": "A missed-pickup claim was filed for 18 Larch Court, which is assigned to Zone R4. Recycling service in that zone was scheduled for Tuesday, and the report was received within the applicable 48-hour reporting deadline. A resident submitted a placement record for the claim. The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317. Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside. The file contains no documented obstruction, no leaking waste, and no visible pests associated with the claim. A routing decision is requested. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and unchanged question bindings. The evidence spans are complete factual sentences: \"The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317.\" and \"Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside.\" The counterfactual changes only the placement observation and remains coherent. Neither context contains a gold answer, answer code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"case_note\":\"A missed-pickup claim was filed for 18 Larch Court, which is assigned to Zone R4. Recycling service in that zone was scheduled for Tuesday, and the report was received within the applicable 48-hour reporting deadline. A resident submitted a placement record for the claim. The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317. Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside. The file contains no documented obstruction, no leaking waste, and no visible pests associated with the claim. A routing decision is requested. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["case_note"], "text": "The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317."}, {"path": ["case_note"], "text": "Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317.", "negative_left": "The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317.", "negative_right": "Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was curbside.", "right": "Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was inside the garage rather than curbside."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-044", "id": "fast-41-diverse-151-044-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"case_note": "A missed-pickup claim was filed for 18 Larch Court, which is assigned to Zone R4. Recycling service in that zone was scheduled for Tuesday, and the report was received within the applicable 48-hour reporting deadline. A resident submitted a placement record for the claim. The Tuesday recycling claim for 18 Larch Court has exactly one submitted placement-evidence item, identified as video record P-317. Video record P-317 shows continuously from 6:55 a.m. through 7:00 a.m. on Tuesday that the blue cart was curbside. The file contains no documented obstruction, no leaking waste, and no visible pests associated with the claim. A routing decision is requested. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and all question bindings; each evidence span is a complete factual sentence, and the counterfactual changes only the placement time consistently. Neither context includes answer labels, rationale, rule tables, proposition IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The sanitation file concerns 18 Larch Court, which is assigned to Zone R4, and the missed collection involved recycling scheduled for Tuesday. The report was received within the 48-hour reporting deadline. The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731. The intake log for item PE-731 records the blue cart as first observed curbside at 7:18 a.m. on Tuesday. The file records no obstruction, no leaking waste, and no visible pests for this claim. A review entry identifies the matter as recycling, not organics or trash. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731."}, {"path": ["context"], "text": "The intake log for item PE-731 records the blue cart as first observed curbside at 7:18 a.m. on Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731.", "negative_left": "The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731.", "negative_right": "The intake log for item PE-731 records the blue cart as first observed curbside at 6:52 a.m. on Tuesday.", "right": "The intake log for item PE-731 records the blue cart as first observed curbside at 7:18 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-056", "id": "fast-41-diverse-151-056-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The sanitation file concerns 18 Larch Court, which is assigned to Zone R4, and the missed collection involved recycling scheduled for Tuesday. The report was received within the 48-hour reporting deadline. The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731. The intake log for item PE-731 records the blue cart as first observed curbside at 7:18 a.m. on Tuesday. The file records no obstruction, no leaking waste, and no visible pests for this claim. A review entry identifies the matter as recycling, not organics or trash. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and all question bindings; each evidence span is a complete factual sentence, and the counterfactual changes only the placement time consistently. Neither context includes answer labels, rationale, rule tables, proposition IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The sanitation file concerns 18 Larch Court, which is assigned to Zone R4, and the missed collection involved recycling scheduled for Tuesday. The report was received within the 48-hour reporting deadline. The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731. The intake log for item PE-731 records the blue cart as first observed curbside at 7:18 a.m. on Tuesday. The file records no obstruction, no leaking waste, and no visible pests for this claim. A review entry identifies the matter as recycling, not organics or trash. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731."}, {"path": ["context"], "text": "The intake log for item PE-731 records the blue cart as first observed curbside at 7:18 a.m. on Tuesday."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731.", "negative_left": "The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731.", "negative_right": "The intake log for item PE-731 records the blue cart as first observed curbside at 6:52 a.m. on Tuesday.", "right": "The intake log for item PE-731 records the blue cart as first observed curbside at 7:18 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-056", "id": "fast-41-diverse-151-056-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The sanitation file concerns 18 Larch Court, which is assigned to Zone R4, and the missed collection involved recycling scheduled for Tuesday. The report was received within the 48-hour reporting deadline. The sole submitted placement-evidence item for the Tuesday recycling claim at 18 Larch Court is item PE-731. The intake log for item PE-731 records the blue cart as first observed curbside at 6:52 a.m. on Tuesday. The file records no obstruction, no leaking waste, and no visible pests for this claim. A review entry identifies the matter as recycling, not organics or trash. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and verbatim governing policy. The address, recycling stream, Tuesday schedule, reporting deadline, and evidence paths remain bound consistently. The evidence sentences are factual: \"The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.\" and \"Record P-417 lists the blue cart's curbside time as 7:12 a.m.\" The counterfactual coherently changes only the recorded time from 7:12 a.m. to 6:52 a.m. without duplicate contradictions. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"case_note\":{\"records\":[\"The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.\",\"Record P-417 lists the blue cart's curbside time as 7:12 a.m.\"],\"details\":\"A resident reporter submitted a missed-pickup claim for 18 Larch Court. The address is assigned to Zone R4, and recycling was scheduled there for Tuesday. The report was received within the 48-hour reporting deadline. The file identifies the reporter, address, waste stream, scheduled service day, and report timestamp. An inspection note for this claim records no obstruction, no leaking waste, and no visible pests. The dispatcher is reviewing the file for routing and response priority, while the service desk remains available to clarify any record needed before a field return is considered.\"},\"policy\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["case_note", "records", "0"], "text": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417."}, {"path": ["case_note", "records", "1"], "text": "Record P-417 lists the blue cart's curbside time as 7:12 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.", "negative_left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.", "negative_right": "Record P-417 lists the blue cart's curbside time as 6:52 a.m.", "right": "Record P-417 lists the blue cart's curbside time as 7:12 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-061", "id": "fast-41-diverse-151-061-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"case_note": {"details": "A resident reporter submitted a missed-pickup claim for 18 Larch Court. The address is assigned to Zone R4, and recycling was scheduled there for Tuesday. The report was received within the 48-hour reporting deadline. The file identifies the reporter, address, waste stream, scheduled service day, and report timestamp. An inspection note for this claim records no obstruction, no leaking waste, and no visible pests. The dispatcher is reviewing the file for routing and response priority, while the service desk remains available to clarify any record needed before a field return is considered.", "records": ["The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.", "Record P-417 lists the blue cart's curbside time as 7:12 a.m."]}, "policy": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and verbatim governing policy. The address, recycling stream, Tuesday schedule, reporting deadline, and evidence paths remain bound consistently. The evidence sentences are factual: \"The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.\" and \"Record P-417 lists the blue cart's curbside time as 7:12 a.m.\" The counterfactual coherently changes only the recorded time from 7:12 a.m. to 6:52 a.m. without duplicate contradictions. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"case_note\":{\"records\":[\"The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.\",\"Record P-417 lists the blue cart's curbside time as 7:12 a.m.\"],\"details\":\"A resident reporter submitted a missed-pickup claim for 18 Larch Court. The address is assigned to Zone R4, and recycling was scheduled there for Tuesday. The report was received within the 48-hour reporting deadline. The file identifies the reporter, address, waste stream, scheduled service day, and report timestamp. An inspection note for this claim records no obstruction, no leaking waste, and no visible pests. The dispatcher is reviewing the file for routing and response priority, while the service desk remains available to clarify any record needed before a field return is considered.\"},\"policy\":\"A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["case_note", "records", "0"], "text": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417."}, {"path": ["case_note", "records", "1"], "text": "Record P-417 lists the blue cart's curbside time as 7:12 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.", "negative_left": "The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.", "negative_right": "Record P-417 lists the blue cart's curbside time as 6:52 a.m.", "right": "Record P-417 lists the blue cart's curbside time as 7:12 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-061", "id": "fast-41-diverse-151-061-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"case_note": {"details": "A resident reporter submitted a missed-pickup claim for 18 Larch Court. The address is assigned to Zone R4, and recycling was scheduled there for Tuesday. The report was received within the 48-hour reporting deadline. The file identifies the reporter, address, waste stream, scheduled service day, and report timestamp. An inspection note for this claim records no obstruction, no leaking waste, and no visible pests. The dispatcher is reviewing the file for routing and response priority, while the service desk remains available to clarify any record needed before a field return is considered.", "records": ["The complete placement-evidence submission for the Tuesday recycling claim at 18 Larch Court consists solely of record P-417.", "Record P-417 lists the blue cart's curbside time as 6:52 a.m."]}, "policy": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy and bindings without answer leakage; the two required evidence quotes are complete factual sentences: “The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42.” and “File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m.” The counterfactual’s changed Q-41 observation and unchanged Q-42 observation are not contradictory.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The sanitation intake concerns a missed household-recycling collection at 18 Larch Court, assigned to Zone R4's Tuesday route. The report was logged at 10:14 a.m. on Wednesday, within the applicable 48-hour deadline. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42. File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m. The field inspection entry records no blocked travel lane, leaking waste, or visible pests connected with this claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The intake clerk attached the route sheet and timestamp receipt to the case before forwarding it for review.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42."}, {"path": ["context"], "text": "File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42.", "negative_left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42.", "negative_right": "File Q-41 shows the blue cart curbside beside 18 Larch Court at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m.", "right": "File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-074", "id": "fast-41-diverse-151-074-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The sanitation intake concerns a missed household-recycling collection at 18 Larch Court, assigned to Zone R4's Tuesday route. The report was logged at 10:14 a.m. on Wednesday, within the applicable 48-hour deadline. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42. File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m. The field inspection entry records no blocked travel lane, leaking waste, or visible pests connected with this claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The intake clerk attached the route sheet and timestamp receipt to the case before forwarding it for review."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy and bindings without answer leakage; the two required evidence quotes are complete factual sentences: “The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42.” and “File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m.” The counterfactual’s changed Q-41 observation and unchanged Q-42 observation are not contradictory.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The sanitation intake concerns a missed household-recycling collection at 18 Larch Court, assigned to Zone R4's Tuesday route. The report was logged at 10:14 a.m. on Wednesday, within the applicable 48-hour deadline. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42. File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m. The field inspection entry records no blocked travel lane, leaking waste, or visible pests connected with this claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The intake clerk attached the route sheet and timestamp receipt to the case before forwarding it for review.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42."}, {"path": ["context"], "text": "File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42.", "negative_left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42.", "negative_right": "File Q-41 shows the blue cart curbside beside 18 Larch Court at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m.", "right": "File Q-41 shows the blue cart inside a locked shed at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-074", "id": "fast-41-diverse-151-074-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The sanitation intake concerns a missed household-recycling collection at 18 Larch Court, assigned to Zone R4's Tuesday route. The report was logged at 10:14 a.m. on Wednesday, within the applicable 48-hour deadline. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of files Q-41 and Q-42. File Q-41 shows the blue cart curbside beside 18 Larch Court at 6:52 a.m., and file Q-42 shows it behind a closed gate at 6:58 a.m. The field inspection entry records no blocked travel lane, leaking waste, or visible pests connected with this claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine. The intake clerk attached the route sheet and timestamp receipt to the case before forwarding it for review."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing sanitation policy and the unchanged questions preserve all criteria and instructions. The address, zone, route, waste stream, deadline, evidence paths, and placement-time bindings remain unchanged. The evidence spans are complete factual sentences: \"The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43.\" and \"Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m.\" The counterfactual is coherent because the two photographs show different cart locations at different times without duplicate contradictory measurements. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The sanitation intake concerns 18 Larch Court, assigned to Zone R4, where the Tuesday collection route was the recycling stream. The electronic report was logged before the 48-hour deadline expired. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43. Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m. The inspection log records no blocked travel lane, leaking waste, or visible pests associated with this claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"request\":\"Choose the correct routing and response classification.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43."}, {"path": ["context"], "text": "Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43.", "negative_left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43.", "negative_right": "Photograph Q-42 shows the blue cart curbside at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m.", "right": "Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-083", "id": "fast-41-diverse-151-083-base", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The sanitation intake concerns 18 Larch Court, assigned to Zone R4, where the Tuesday collection route was the recycling stream. The electronic report was logged before the 48-hour deadline expired. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43. Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m. The inspection log records no blocked travel lane, leaking waste, or visible pests associated with this claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "request": "Choose the correct routing and response classification."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "service_desk_missing_placement_routine"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing sanitation policy and the unchanged questions preserve all criteria and instructions. The address, zone, route, waste stream, deadline, evidence paths, and placement-time bindings remain unchanged. The evidence spans are complete factual sentences: \"The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43.\" and \"Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m.\" The counterfactual is coherent because the two photographs show different cart locations at different times without duplicate contradictory measurements. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or evidentiary relationship; A4 is a permissible universally quantified evidentiary fact rather than a bundled classification. The focus is factual, and the base/counter assignments are realizable with only whether submitted evidence verifies timely curb placement changing. The cited state context preserves all state-originating eligibility, fallback-routing, hazard, and priority rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies the correct zone, scheduled recycling stream, and timely report, while A4 establishes that curb placement by 7:00 a.m. remains unverified. Refutation of A5–A7 excludes every listed urgent hazard. The policy therefore requires routine service-desk routing to resolve placement evidence.", "rule_index": 0, "sound": true}, {"reason": "Refuting A4 entails that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, recycling schedule, timely report, and refutation of all listed urgent hazards, this satisfies every condition for a routine recycling crew return.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Larch Court is in Zone R4."}, {"id": "A2", "statement": "Zone R4 had recycling scheduled for Tuesday."}, {"id": "A3", "statement": "The missed-pickup report for Tuesday recycling at 18 Larch Court was received within the 48-hour reporting deadline."}, {"id": "A4", "statement": "No item submitted as placement evidence for the Tuesday recycling claim at 18 Larch Court verifies that the blue cart was curbside by 7:00 a.m."}, {"id": "A5", "statement": "An obstruction is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A6", "statement": "Leaking waste is documented for the Tuesday recycling claim at 18 Larch Court."}, {"id": "A7", "statement": "Visible pests are documented for the Tuesday recycling claim at 18 Larch Court."}], "base_state_json": "{\"context\":\"Case note: The sanitation intake concerns 18 Larch Court, assigned to Zone R4, where the Tuesday collection route was the recycling stream. The electronic report was logged before the 48-hour deadline expired. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43. Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m. The inspection log records no blocked travel lane, leaking waste, or visible pests associated with this claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.\",\"request\":\"Choose the correct routing and response classification.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["context"], "text": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43."}, {"path": ["context"], "text": "Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m."}], "policy_evidence": [{"path": ["context"], "text": "A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine."}], "rules": [{"justification": "The zone, scheduled recycling stream, and timely report are verified, but the submission provides no verification of curb placement by 7:00 a.m.; with every listed urgent hazard explicitly absent, policy requires routine service-desk routing to resolve the missing placement evidence.", "target": "service_desk_missing_placement_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}, {"justification": "Refutation of A4 establishes that at least one submitted item verifies curb placement by 7:00 a.m. Together with the verified zone, Tuesday recycling schedule, timely report, and explicit absence of every urgent hazard, all requirements for a routine recycling return are satisfied.", "target": "recycling_return_routine", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}]}]}, "verified_pair": {"left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43.", "negative_left": "The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43.", "negative_right": "Photograph Q-42 shows the blue cart curbside at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m.", "right": "Photograph Q-42 shows the blue cart inside a locked shed at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-151-083", "id": "fast-41-diverse-151-083-counterfactual", "input": {"questions": {"decision": {"criteria": {"organics_return_urgent": "Dispatch the organics crew urgently, used only for an eligible organics miss with a documented obstruction, leak, or pest hazard.", "recycling_return_routine": "Dispatch the recycling crew routinely, used only when all return-pickup requirements are verified and no urgent hazard exists.", "service_desk_missing_placement_routine": "Route to the service desk at routine priority to obtain placement time or clear curbside evidence because return eligibility is incomplete and no hazard is documented.", "trash_return_urgent": "Dispatch the trash crew urgently, used only for an eligible trash miss with a documented obstruction, leak, or pest hazard."}, "instructions": "Apply the stated policy. Select exactly one option. A crew return requires every listed eligibility fact to be verified; missing curb-placement evidence must be resolved by the service desk.", "type": "choice"}}, "state": {"context": "Case note: The sanitation intake concerns 18 Larch Court, assigned to Zone R4, where the Tuesday collection route was the recycling stream. The electronic report was logged before the 48-hour deadline expired. The submitted placement evidence for the Tuesday recycling claim at 18 Larch Court consists exactly of photographs Q-42 and Q-43. Photograph Q-42 shows the blue cart curbside at 6:58 a.m., and photograph Q-43 shows the blue cart inside a garage at 7:04 a.m. The inspection log records no blocked travel lane, leaking waste, or visible pests associated with this claim. A resident reporter submits a missed-pickup claim. City policy allows a sanitation dispatcher to order a return only when the correct zone, scheduled waste stream, timely report, and curb placement by 7:00 a.m. are all verified. Otherwise, the service desk must obtain missing information. A collection crew supervisor rates verified hazards urgent only for blocked travel lanes, leaking waste, or visible pests; all other cases are routine.", "request": "Choose the correct routing and response classification."}}, "method": "c2d", "provenance": {"source_id": "diverse-151", "source_is_synthetic": true, "source_sha256": "3655127c5209926929eb121826539ea874a6cfc4e5a6bd58ba43c50b4ac3b910", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "recycling_return_routine"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, contain two complete factual evidence sentences, remain internally coherent, and include no gold answer or classifier instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relations or closely defined factual alternatives, and the focus atom is the factual timeliness relation. The base and counter assignments differ only on timeliness and are jointly realizable: an otherwise identical unemptied organics cart can have been timely or late. Policy evidence correctly preserves the substantive return-scope and preparation-exception rule originating in the original state; the remaining governing criteria and priority rules are already retained in the questions object and need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes organics rather than recycling, an unemptied timely and accessible cart, a lid gap over 5 cm, and the explicit absence of every listed urgent-hazard category. The preparation exception therefore applies and uniquely supports routine organics-desk handling with no return pickup.", "rule_index": 0, "sound": true}, {"reason": "Refuting the no-later-than-deadline atom entails a late set-out, placing the cart outside the return-pickup scope. Under the instruction to apply scope before the preparation exception, the over-5-cm gap does not independently establish outcome D. Recycling and all urgent hazards are excluded, so none of A–D applies and E is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported cart at 18 Lark Street was an organics cart for the Tuesday service event."}, {"id": "a2", "statement": "Organics collection was scheduled at 18 Lark Street for the Tuesday service event."}, {"id": "a3", "statement": "The evidence identified recycling as the scheduled or reported material at 18 Lark Street for the Tuesday service event."}, {"id": "a4", "statement": "The reported cart at 18 Lark Street remained unemptied after the Tuesday collection pass."}, {"id": "a5", "statement": "The reported cart's Tuesday curb set-out time was no later than the applicable 7:00 a.m. set-out deadline."}, {"id": "a6", "statement": "The reported cart at 18 Lark Street was accessible to the collection crew during the Tuesday collection pass."}, {"id": "a7", "statement": "The reported cart's lid gap during the Tuesday collection pass was greater than 5 cm."}, {"id": "a8", "statement": "Waste from the reported cart was spilled during the relevant Tuesday service period."}, {"id": "a9", "statement": "The reported cart or its waste blocked travel during the relevant Tuesday service period."}, {"id": "a10", "statement": "The reported cart leaked during the relevant Tuesday service period."}, {"id": "a11", "statement": "Pests were present at the reported cart during the relevant Tuesday service period."}, {"id": "a12", "statement": "The reported cart presented another stated immediate sanitation hazard during the relevant Tuesday service period."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024. The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 7:00 a.m.\"},{\"speaker\":\"Dispatcher\",\"text\":\"The address was scheduled for organics collection during Tuesday's service event, and the reported organics cart remained unemptied after the collection pass. The cart was accessible to the collection crew.\"},{\"speaker\":\"Crew supervisor\",\"text\":\"A branch held the cart lid approximately 8 cm open during the pass. No recycling material was identified as the scheduled or reported material. The service record reports no spilled waste, blocked travel, leakage, pests, or other stated immediate sanitation hazard during the relevant period. Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024."}, {"path": ["0", "text"], "text": "The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 7:00 a.m."}], "policy_evidence": [{"path": ["3", "text"], "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}], "rules": [{"justification": "A timely, accessible, apparently missed organics cart is within return-pickup scope, but its lid gap over 5 cm triggers the preparation exception. With every stated urgent-hazard category explicitly absent and recycling explicitly excluded, the organics service desk handles it at routine priority with no return pickup.", "target": "D_organics_desk_routine_no_return", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "An explicitly late cart is outside the timely return-pickup scope, so neither organics return-pickup rubric applies and the scope-first instruction prevents use of the preparation exception as an independent routing basis. Recycling and every urgent-hazard category are explicitly excluded, leaving none of the four substantive rubrics uniquely satisfied.", "target": "E_none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024.", "negative_left": "The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024.", "negative_right": "The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 6:30 a.m.", "right": "The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 7:00 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-152-015", "id": "fast-41-diverse-152-015-base", "input": {"questions": {"decision": {"criteria": {"A_organics_return_routine": "Route to the organics collection crew for a routine return pickup because the cart was timely, accessible, properly prepared within the 5 cm lid limit, and apparently missed.", "B_organics_return_urgent": "Route to the organics collection crew for an urgent return pickup because an eligible missed cart also presents a stated spill, obstruction, leakage, pest, or immediate sanitation hazard.", "C_recycling_desk_routine": "Route to the recycling service desk at routine priority because the evidence identifies recycling, rather than organics, as the scheduled or reported material.", "D_organics_desk_routine_no_return": "Route to the organics service desk at routine priority with no return pickup because a protruding branch or lid gap over 5 cm triggers the preparation exception and no urgent hazard is present.", "E_none_of_above": "Use only if the evidence does not uniquely satisfy any of the four routing, return, and priority rubrics above."}, "instructions": "Choose the routing, return-pickup decision, and priority that match the supplied policy. Apply the return-pickup scope first, then its preparation exception. Urgent priority requires spilled waste, blocked travel, leakage, pests, or another stated immediate sanitation hazard; otherwise use routine priority.", "type": "choice"}}, "state": [{"speaker": "Case note", "text": "The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024. The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 7:00 a.m."}, {"speaker": "Dispatcher", "text": "The address was scheduled for organics collection during Tuesday's service event, and the reported organics cart remained unemptied after the collection pass. The cart was accessible to the collection crew."}, {"speaker": "Crew supervisor", "text": "A branch held the cart lid approximately 8 cm open during the pass. No recycling material was identified as the scheduled or reported material. The service record reports no spilled waste, blocked travel, leakage, pests, or other stated immediate sanitation hazard during the relevant period. Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-152", "source_is_synthetic": true, "source_sha256": "1dfc7c00860982758ea45c9c5b47d8291aac858825f6143cf430090701b0ce5f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_organics_desk_routine_no_return"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, contain two complete factual evidence sentences, remain internally coherent, and include no gold answer or classifier instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relations or closely defined factual alternatives, and the focus atom is the factual timeliness relation. The base and counter assignments differ only on timeliness and are jointly realizable: an otherwise identical unemptied organics cart can have been timely or late. Policy evidence correctly preserves the substantive return-scope and preparation-exception rule originating in the original state; the remaining governing criteria and priority rules are already retained in the questions object and need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes organics rather than recycling, an unemptied timely and accessible cart, a lid gap over 5 cm, and the explicit absence of every listed urgent-hazard category. The preparation exception therefore applies and uniquely supports routine organics-desk handling with no return pickup.", "rule_index": 0, "sound": true}, {"reason": "Refuting the no-later-than-deadline atom entails a late set-out, placing the cart outside the return-pickup scope. Under the instruction to apply scope before the preparation exception, the over-5-cm gap does not independently establish outcome D. Recycling and all urgent hazards are excluded, so none of A–D applies and E is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported cart at 18 Lark Street was an organics cart for the Tuesday service event."}, {"id": "a2", "statement": "Organics collection was scheduled at 18 Lark Street for the Tuesday service event."}, {"id": "a3", "statement": "The evidence identified recycling as the scheduled or reported material at 18 Lark Street for the Tuesday service event."}, {"id": "a4", "statement": "The reported cart at 18 Lark Street remained unemptied after the Tuesday collection pass."}, {"id": "a5", "statement": "The reported cart's Tuesday curb set-out time was no later than the applicable 7:00 a.m. set-out deadline."}, {"id": "a6", "statement": "The reported cart at 18 Lark Street was accessible to the collection crew during the Tuesday collection pass."}, {"id": "a7", "statement": "The reported cart's lid gap during the Tuesday collection pass was greater than 5 cm."}, {"id": "a8", "statement": "Waste from the reported cart was spilled during the relevant Tuesday service period."}, {"id": "a9", "statement": "The reported cart or its waste blocked travel during the relevant Tuesday service period."}, {"id": "a10", "statement": "The reported cart leaked during the relevant Tuesday service period."}, {"id": "a11", "statement": "Pests were present at the reported cart during the relevant Tuesday service period."}, {"id": "a12", "statement": "The reported cart presented another stated immediate sanitation hazard during the relevant Tuesday service period."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024. The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 7:00 a.m.\"},{\"speaker\":\"Dispatcher\",\"text\":\"The address was scheduled for organics collection during Tuesday's service event, and the reported organics cart remained unemptied after the collection pass. The cart was accessible to the collection crew.\"},{\"speaker\":\"Crew supervisor\",\"text\":\"A branch held the cart lid approximately 8 cm open during the pass. No recycling material was identified as the scheduled or reported material. The service record reports no spilled waste, blocked travel, leakage, pests, or other stated immediate sanitation hazard during the relevant period. Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["0", "text"], "text": "The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024."}, {"path": ["0", "text"], "text": "The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 7:00 a.m."}], "policy_evidence": [{"path": ["3", "text"], "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}], "rules": [{"justification": "A timely, accessible, apparently missed organics cart is within return-pickup scope, but its lid gap over 5 cm triggers the preparation exception. With every stated urgent-hazard category explicitly absent and recycling explicitly excluded, the organics service desk handles it at routine priority with no return pickup.", "target": "D_organics_desk_routine_no_return", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "An explicitly late cart is outside the timely return-pickup scope, so neither organics return-pickup rubric applies and the scope-first instruction prevents use of the preparation exception as an independent routing basis. Recycling and every urgent-hazard category are explicitly excluded, leaving none of the four substantive rubrics uniquely satisfied.", "target": "E_none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024.", "negative_left": "The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024.", "negative_right": "The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 6:30 a.m.", "right": "The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 7:00 a.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-152-015", "id": "fast-41-diverse-152-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_organics_return_routine": "Route to the organics collection crew for a routine return pickup because the cart was timely, accessible, properly prepared within the 5 cm lid limit, and apparently missed.", "B_organics_return_urgent": "Route to the organics collection crew for an urgent return pickup because an eligible missed cart also presents a stated spill, obstruction, leakage, pest, or immediate sanitation hazard.", "C_recycling_desk_routine": "Route to the recycling service desk at routine priority because the evidence identifies recycling, rather than organics, as the scheduled or reported material.", "D_organics_desk_routine_no_return": "Route to the organics service desk at routine priority with no return pickup because a protruding branch or lid gap over 5 cm triggers the preparation exception and no urgent hazard is present.", "E_none_of_above": "Use only if the evidence does not uniquely satisfy any of the four routing, return, and priority rubrics above."}, "instructions": "Choose the routing, return-pickup decision, and priority that match the supplied policy. Apply the return-pickup scope first, then its preparation exception. Urgent priority requires spilled waste, blocked travel, leakage, pests, or another stated immediate sanitation hazard; otherwise use routine priority.", "type": "choice"}}, "state": [{"speaker": "Case note", "text": "The reported cart at 18 Lark Street was recorded at the curb at 6:42 a.m. on Tuesday, 14 May 2024. The applicable set-out deadline for the Tuesday, 14 May 2024 service at 18 Lark Street was 6:30 a.m."}, {"speaker": "Dispatcher", "text": "The address was scheduled for organics collection during Tuesday's service event, and the reported organics cart remained unemptied after the collection pass. The cart was accessible to the collection crew."}, {"speaker": "Crew supervisor", "text": "A branch held the cart lid approximately 8 cm open during the pass. No recycling material was identified as the scheduled or reported material. The service record reports no spilled waste, blocked travel, leakage, pests, or other stated immediate sanitation hazard during the relevant period. Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-152", "source_is_synthetic": true, "source_sha256": "1dfc7c00860982758ea45c9c5b47d8291aac858825f6143cf430090701b0ce5f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "E_none_of_above"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged routing request, governing policy, and Mara Chen’s Tuesday report at 18 Birch Lane. The two evidence spans are complete factual sentences and appear verbatim in both contexts. The counterfactual changes only the photographed asset from CA-417 to CB-209, which coheres with the stated cart properties and priorities. Neither context states an answer, label, rule table, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual identity of the reported cart. The policy evidence preserves the state-originating routing and urgency rules; question-originating criteria need not be repeated. The base and counter assignments can describe the same two carts and circumstances while changing only which cart is the reported cart, without violating the policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the reported cart as Cart A, establish a timely and properly placed missed collection, give organics as its scheduled material, and explicitly set the controlling priority to urgent. The leaking-waste condition also satisfies the policy's necessary urgent-condition restriction.", "rule_index": 0, "sound": true}, {"reason": "Membership in {Cart A, Cart B} together with refutation of Cart A identifies the reported cart as Cart B. The remaining conditions establish a timely and properly placed missed collection, scheduled recycling, routine priority, and absence of both policy-listed urgent conditions, which is sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane belongs to the candidate asset set {Cart A, Cart B}."}, {"id": "a2", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is Cart A."}, {"id": "a3", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was properly placed for collection."}, {"id": "a4", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was curbside no later than the 6:30 a.m. Tuesday cutoff."}, {"id": "a5", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane remained uncollected after the scheduled Tuesday collection."}, {"id": "a6", "statement": "The scheduled Tuesday material for Cart A at 18 Birch Lane is organics."}, {"id": "a7", "statement": "The priority field in the controlling dispatch record for Cart A’s Tuesday missed collection at 18 Birch Lane is urgent."}, {"id": "a8", "statement": "Waste in Cart A was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a9", "statement": "The scheduled Tuesday material for Cart B at 18 Birch Lane is recycling."}, {"id": "a10", "statement": "The priority field in the controlling dispatch record for Cart B’s Tuesday missed collection at 18 Birch Lane is routine."}, {"id": "a11", "statement": "Waste in Cart B was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a12", "statement": "Cart B blocked the roadway at 18 Birch Lane during the Tuesday collection period."}], "base_state_json": "\"Mara Chen’s Tuesday missed-collection report concerns 18 Birch Lane. The cart was properly placed at the curb by the 6:30 a.m. cutoff and remained uncollected after Tuesday’s scheduled service. The Tuesday schedule assigns organics to Cart A and recycling to Cart B. The controlling dispatch record lists urgent priority for Cart A and routine priority for Cart B. Cart A’s waste was leaking during the collection period. Cart B’s waste was not leaking and did not block the roadway. The intake record states: The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209. The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CA-417. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209."}, {"path": [], "text": "The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CA-417."}], "policy_evidence": [{"path": [], "text": "Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}], "rules": [{"justification": "The report concerns Cart A, which was properly placed, timely, and missed; its scheduled material is organics and its controlling priority is urgent. Its leaking waste also satisfies the policy’s necessary condition for urgent status.", "target": "true", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The reported cart is one of Cart A and Cart B but is not Cart A, so it is Cart B. It was properly placed, timely, and missed; its scheduled material is recycling and its controlling priority is routine, while both policy-listed urgent conditions are explicitly absent.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209.", "negative_left": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209.", "negative_right": "The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CB-209.", "right": "The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CA-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-153-007", "id": "fast-41-diverse-153-007-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route it to the organics crew as urgent; route it to the recycling return-pickup crew with routine priority.", "true": "Route the report to the organics return-pickup crew with urgent priority."}, "instructions": "Decide whether this report should be routed to the organics return-pickup crew as an urgent case.", "type": "noul"}}, "state": "Mara Chen’s Tuesday missed-collection report concerns 18 Birch Lane. The cart was properly placed at the curb by the 6:30 a.m. cutoff and remained uncollected after Tuesday’s scheduled service. The Tuesday schedule assigns organics to Cart A and recycling to Cart B. The controlling dispatch record lists urgent priority for Cart A and routine priority for Cart B. Cart A’s waste was leaking during the collection period. Cart B’s waste was not leaking and did not block the roadway. The intake record states: The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209. The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CA-417. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}, "method": "c2d", "provenance": {"source_id": "diverse-153", "source_is_synthetic": true, "source_sha256": "f8cc49d0970071d68e752b9fc8ad734adab047d8cd33245838da80f07b42cdfc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged routing request, governing policy, and Mara Chen’s Tuesday report at 18 Birch Lane. The two evidence spans are complete factual sentences and appear verbatim in both contexts. The counterfactual changes only the photographed asset from CA-417 to CB-209, which coheres with the stated cart properties and priorities. Neither context states an answer, label, rule table, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a12": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual identity of the reported cart. The policy evidence preserves the state-originating routing and urgency rules; question-originating criteria need not be repeated. The base and counter assignments can describe the same two carts and circumstances while changing only which cart is the reported cart, without violating the policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the reported cart as Cart A, establish a timely and properly placed missed collection, give organics as its scheduled material, and explicitly set the controlling priority to urgent. The leaking-waste condition also satisfies the policy's necessary urgent-condition restriction.", "rule_index": 0, "sound": true}, {"reason": "Membership in {Cart A, Cart B} together with refutation of Cart A identifies the reported cart as Cart B. The remaining conditions establish a timely and properly placed missed collection, scheduled recycling, routine priority, and absence of both policy-listed urgent conditions, which is sufficient for the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane belongs to the candidate asset set {Cart A, Cart B}."}, {"id": "a2", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is Cart A."}, {"id": "a3", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was properly placed for collection."}, {"id": "a4", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane was curbside no later than the 6:30 a.m. Tuesday cutoff."}, {"id": "a5", "statement": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane remained uncollected after the scheduled Tuesday collection."}, {"id": "a6", "statement": "The scheduled Tuesday material for Cart A at 18 Birch Lane is organics."}, {"id": "a7", "statement": "The priority field in the controlling dispatch record for Cart A’s Tuesday missed collection at 18 Birch Lane is urgent."}, {"id": "a8", "statement": "Waste in Cart A was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a9", "statement": "The scheduled Tuesday material for Cart B at 18 Birch Lane is recycling."}, {"id": "a10", "statement": "The priority field in the controlling dispatch record for Cart B’s Tuesday missed collection at 18 Birch Lane is routine."}, {"id": "a11", "statement": "Waste in Cart B was leaking during the Tuesday collection period at 18 Birch Lane."}, {"id": "a12", "statement": "Cart B blocked the roadway at 18 Birch Lane during the Tuesday collection period."}], "base_state_json": "\"Mara Chen’s Tuesday missed-collection report concerns 18 Birch Lane. The cart was properly placed at the curb by the 6:30 a.m. cutoff and remained uncollected after Tuesday’s scheduled service. The Tuesday schedule assigns organics to Cart A and recycling to Cart B. The controlling dispatch record lists urgent priority for Cart A and routine priority for Cart B. Cart A’s waste was leaking during the collection period. Cart B’s waste was not leaking and did not block the roadway. The intake record states: The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209. The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CA-417. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209."}, {"path": [], "text": "The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CA-417."}], "policy_evidence": [{"path": [], "text": "Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}], "rules": [{"justification": "The report concerns Cart A, which was properly placed, timely, and missed; its scheduled material is organics and its controlling priority is urgent. Its leaking waste also satisfies the policy’s necessary condition for urgent status.", "target": "true", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The reported cart is one of Cart A and Cart B but is not Cart A, so it is Cart B. It was properly placed, timely, and missed; its scheduled material is recycling and its controlling priority is routine, while both policy-listed urgent conditions are explicitly absent.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209.", "negative_left": "The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209.", "negative_right": "The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CB-209.", "right": "The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CA-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-153-007", "id": "fast-41-diverse-153-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route it to the organics crew as urgent; route it to the recycling return-pickup crew with routine priority.", "true": "Route the report to the organics return-pickup crew with urgent priority."}, "instructions": "Decide whether this report should be routed to the organics return-pickup crew as an urgent case.", "type": "noul"}}, "state": "Mara Chen’s Tuesday missed-collection report concerns 18 Birch Lane. The cart was properly placed at the curb by the 6:30 a.m. cutoff and remained uncollected after Tuesday’s scheduled service. The Tuesday schedule assigns organics to Cart A and recycling to Cart B. The controlling dispatch record lists urgent priority for Cart A and routine priority for Cart B. Cart A’s waste was leaking during the collection period. Cart B’s waste was not leaking and did not block the roadway. The intake record states: The cart in Mara Chen’s Tuesday missed-collection report at 18 Birch Lane is either Cart A, whose asset tag is CA-417, or Cart B, whose asset tag is CB-209. The intake photograph for Mara Chen’s Tuesday missed-collection report at 18 Birch Lane records asset tag CB-209. Policy routes a properly placed, timely missed cart to the crew matching its scheduled material; urgent status applies only to leaking waste or a blocked roadway, neither of which appears in the photo."}, "method": "c2d", "provenance": {"source_id": "diverse-153", "source_is_synthetic": true, "source_sha256": "f8cc49d0970071d68e752b9fc8ad734adab047d8cd33245838da80f07b42cdfc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions object and do not alter governing policy. Both preserve Mara Chen, 18 Alder Court, Tuesday, Zone C, and the return-pickup scope. Each context has two complete factual evidence sentences: base evidence \"The cart depicted in both submitted images for Mara Chen’s report is designated as organics.\" and \"Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report.\"; counterfactual evidence \"The cart depicted in both submitted images for Mara Chen’s report is designated as residuals.\" and \"Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report.\" The counterfactual coherently changes the cart designation while retaining the scheduled material and all unchanged observations. Neither generated context embeds a gold answer, answer code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court and requested a return collection. Address records place the property in Zone C. The placement image was captured at 5:58 a.m.; it shows a cart at the curb in compliant condition. The noncollection image was captured at 7:14 p.m.; it shows that cart still uncollected. Both submitted images depict the same cart at 18 Alder Court when captured. The cart shown in the placement image is the cart shown in the later image. The report concerns the Tuesday collection scheduled for this address.\",\"evidence\":[\"The cart depicted in both submitted images for Mara Chen’s report is designated as organics.\",\"Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report.\"],\"request\":\"Is this report ready for direct routing to the Zone C crew for a return pickup?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "The cart depicted in both submitted images for Mara Chen’s report is designated as organics."}, {"path": ["evidence", "1"], "text": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The cart depicted in both submitted images for Mara Chen’s report is designated as organics.", "negative_left": "The cart depicted in both submitted images for Mara Chen’s report is designated as residuals.", "negative_right": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report.", "right": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report."}, "verifier_independent_model": false}, "family": "fast-41-diverse-154-002", "id": "fast-41-diverse-154-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court and requested a return collection. Address records place the property in Zone C. The placement image was captured at 5:58 a.m.; it shows a cart at the curb in compliant condition. The noncollection image was captured at 7:14 p.m.; it shows that cart still uncollected. Both submitted images depict the same cart at 18 Alder Court when captured. The cart shown in the placement image is the cart shown in the later image. The report concerns the Tuesday collection scheduled for this address.", "evidence": ["The cart depicted in both submitted images for Mara Chen’s report is designated as organics.", "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report."], "request": "Is this report ready for direct routing to the Zone C crew for a return pickup?"}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions object and do not alter governing policy. Both preserve Mara Chen, 18 Alder Court, Tuesday, Zone C, and the return-pickup scope. Each context has two complete factual evidence sentences: base evidence \"The cart depicted in both submitted images for Mara Chen’s report is designated as organics.\" and \"Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report.\"; counterfactual evidence \"The cart depicted in both submitted images for Mara Chen’s report is designated as residuals.\" and \"Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report.\" The counterfactual coherently changes the cart designation while retaining the scheduled material and all unchanged observations. Neither generated context embeds a gold answer, answer code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court and requested a return collection. Address records place the property in Zone C. The placement image was captured at 5:58 a.m.; it shows a cart at the curb in compliant condition. The noncollection image was captured at 7:14 p.m.; it shows that cart still uncollected. Both submitted images depict the same cart at 18 Alder Court when captured. The cart shown in the placement image is the cart shown in the later image. The report concerns the Tuesday collection scheduled for this address.\",\"evidence\":[\"The cart depicted in both submitted images for Mara Chen’s report is designated as organics.\",\"Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report.\"],\"request\":\"Is this report ready for direct routing to the Zone C crew for a return pickup?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "The cart depicted in both submitted images for Mara Chen’s report is designated as organics."}, {"path": ["evidence", "1"], "text": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The cart depicted in both submitted images for Mara Chen’s report is designated as organics.", "negative_left": "The cart depicted in both submitted images for Mara Chen’s report is designated as residuals.", "negative_right": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report.", "right": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report."}, "verifier_independent_model": false}, "family": "fast-41-diverse-154-002", "id": "fast-41-diverse-154-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court and requested a return collection. Address records place the property in Zone C. The placement image was captured at 5:58 a.m.; it shows a cart at the curb in compliant condition. The noncollection image was captured at 7:14 p.m.; it shows that cart still uncollected. Both submitted images depict the same cart at 18 Alder Court when captured. The cart shown in the placement image is the cart shown in the later image. The report concerns the Tuesday collection scheduled for this address.", "evidence": ["The cart depicted in both submitted images for Mara Chen’s report is designated as residuals.", "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen’s report."], "request": "Is this report ready for direct routing to the Zone C crew for a return pickup?"}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and neither context adds exceptions or altered rules. The original bindings for Mara Chen, 18 Alder Court, Tuesday, the requested return collection, and the material and zone path remain unchanged. The base evidence retains “The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics.” and “The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics.” as complete factual sentences. The counterfactual instead states that the cart designation is recyclables while the scheduled material is organics, which is a coherent changed observation rather than a duplicate contradiction. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or classifier instruction. Both contexts preserve the complete address, timing, placement, identity, and noncollection evidence needed by the unchanged question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Case note: Mara Chen submitted a Tuesday missed-pickup report for 18 Alder Court and requested a return collection. Address records place the property in Zone C. The scheduled material there that day is identified below. The placement image was captured by 6:00 a.m.; it shows a cart at the property in a compliant position. The noncollection image was captured after 7:00 p.m. and shows that cart uncollected. The two images depict the same cart, and every submitted cart image for this report depicts a cart at 18 Alder Court when captured. The cart's material designation and the day's scheduled material are recorded in the evidence list.\",\"evidence\":[\"The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics.\",\"The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics.\",\"18 Alder Court is in Zone C.\",\"The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report.\",\"The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time.\",\"The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report.\",\"The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report.\",\"The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time.\",\"Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time.\"],\"request\":\"Is this report ready for the appropriate return-pickup handling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics."}, {"path": ["evidence", "1"], "text": "The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics.", "negative_left": "The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is recyclables.", "negative_right": "The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics.", "right": "The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics."}, "verifier_independent_model": false}, "family": "fast-41-diverse-154-006", "id": "fast-41-diverse-154-006-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Case note: Mara Chen submitted a Tuesday missed-pickup report for 18 Alder Court and requested a return collection. Address records place the property in Zone C. The scheduled material there that day is identified below. The placement image was captured by 6:00 a.m.; it shows a cart at the property in a compliant position. The noncollection image was captured after 7:00 p.m. and shows that cart uncollected. The two images depict the same cart, and every submitted cart image for this report depicts a cart at 18 Alder Court when captured. The cart's material designation and the day's scheduled material are recorded in the evidence list.", "evidence": ["The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics.", "The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics.", "18 Alder Court is in Zone C.", "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report.", "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time.", "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report.", "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report.", "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time.", "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."], "request": "Is this report ready for the appropriate return-pickup handling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and neither context adds exceptions or altered rules. The original bindings for Mara Chen, 18 Alder Court, Tuesday, the requested return collection, and the material and zone path remain unchanged. The base evidence retains “The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics.” and “The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics.” as complete factual sentences. The counterfactual instead states that the cart designation is recyclables while the scheduled material is organics, which is a coherent changed observation rather than a duplicate contradiction. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or classifier instruction. Both contexts preserve the complete address, timing, placement, identity, and noncollection evidence needed by the unchanged question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Case note: Mara Chen submitted a Tuesday missed-pickup report for 18 Alder Court and requested a return collection. Address records place the property in Zone C. The scheduled material there that day is identified below. The placement image was captured by 6:00 a.m.; it shows a cart at the property in a compliant position. The noncollection image was captured after 7:00 p.m. and shows that cart uncollected. The two images depict the same cart, and every submitted cart image for this report depicts a cart at 18 Alder Court when captured. The cart's material designation and the day's scheduled material are recorded in the evidence list.\",\"evidence\":[\"The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics.\",\"The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics.\",\"18 Alder Court is in Zone C.\",\"The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report.\",\"The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time.\",\"The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report.\",\"The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report.\",\"The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time.\",\"Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time.\"],\"request\":\"Is this report ready for the appropriate return-pickup handling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics."}, {"path": ["evidence", "1"], "text": "The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is organics.", "negative_left": "The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is recyclables.", "negative_right": "The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics.", "right": "The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics."}, "verifier_independent_model": false}, "family": "fast-41-diverse-154-006", "id": "fast-41-diverse-154-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Case note: Mara Chen submitted a Tuesday missed-pickup report for 18 Alder Court and requested a return collection. Address records place the property in Zone C. The scheduled material there that day is identified below. The placement image was captured by 6:00 a.m.; it shows a cart at the property in a compliant position. The noncollection image was captured after 7:00 p.m. and shows that cart uncollected. The two images depict the same cart, and every submitted cart image for this report depicts a cart at 18 Alder Court when captured. The cart's material designation and the day's scheduled material are recorded in the evidence list.", "evidence": ["The material designation recorded for the cart depicted in the submitted images for Mara Chen's report is recyclables.", "The material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report is organics.", "18 Alder Court is in Zone C.", "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report.", "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time.", "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report.", "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report.", "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time.", "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."], "request": "Is this report ready for the appropriate return-pickup handling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, and instructions. Both contexts retain the address, zone, collection day, cart type, and service event bindings. The two focus evidence spans are complete factual sentences. The counterfactual changes only the report time from 5:00 p.m. to 7:00 p.m. on 15 May 2025, without creating contradictory duplicate assertions. Neither context states a gold label, answer code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Service record\",\"text\":\"18 Alder Mews is assigned to Zone C, which collects trash on Wednesday using gray carts. The cart presented for that Wednesday collection was gray and had been curbside by 7:00 a.m.; route records show the scheduled crew did not empty it.\"},{\"speaker\":\"Verified timing record\",\"text\":\"The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025. The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on 15 May 2025.\"},{\"speaker\":\"Site inspection\",\"text\":\"The remaining waste is fully contained. Inspection found no exposed waste, leakage, odor, animal access, pests, loose material, obstruction, or acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on 15 May 2025."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025.", "negative_left": "The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 7:00 p.m. on 15 May 2025.", "right": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on 15 May 2025."}, "verifier_independent_model": false}, "family": "fast-41-diverse-155-005", "id": "fast-41-diverse-155-005-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Service record", "text": "18 Alder Mews is assigned to Zone C, which collects trash on Wednesday using gray carts. The cart presented for that Wednesday collection was gray and had been curbside by 7:00 a.m.; route records show the scheduled crew did not empty it."}, {"speaker": "Verified timing record", "text": "The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025. The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on 15 May 2025."}, {"speaker": "Site inspection", "text": "The remaining waste is fully contained. Inspection found no exposed waste, leakage, odor, animal access, pests, loose material, obstruction, or acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, and instructions. Both contexts retain the address, zone, collection day, cart type, and service event bindings. The two focus evidence spans are complete factual sentences. The counterfactual changes only the report time from 5:00 p.m. to 7:00 p.m. on 15 May 2025, without creating contradictory duplicate assertions. Neither context states a gold label, answer code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Service record\",\"text\":\"18 Alder Mews is assigned to Zone C, which collects trash on Wednesday using gray carts. The cart presented for that Wednesday collection was gray and had been curbside by 7:00 a.m.; route records show the scheduled crew did not empty it.\"},{\"speaker\":\"Verified timing record\",\"text\":\"The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025. The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on 15 May 2025.\"},{\"speaker\":\"Site inspection\",\"text\":\"The remaining waste is fully contained. Inspection found no exposed waste, leakage, odor, animal access, pests, loose material, obstruction, or acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["1", "text"], "text": "The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on 15 May 2025."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025.", "negative_left": "The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 7:00 p.m. on 15 May 2025.", "right": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on 15 May 2025."}, "verifier_independent_model": false}, "family": "fast-41-diverse-155-005", "id": "fast-41-diverse-155-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Service record", "text": "18 Alder Mews is assigned to Zone C, which collects trash on Wednesday using gray carts. The cart presented for that Wednesday collection was gray and had been curbside by 7:00 a.m.; route records show the scheduled crew did not empty it."}, {"speaker": "Verified timing record", "text": "The relevant Zone C Wednesday trash route closed at 6:00 p.m. on 14 May 2025. The missed-collection report for 18 Alder Mews was submitted at 7:00 p.m. on 15 May 2025."}, {"speaker": "Site inspection", "text": "The remaining waste is fully contained. Inspection found no exposed waste, leakage, odor, animal access, pests, loose material, obstruction, or acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object remains verbatim, and neither context adds, removes, or changes governing policy. Entity and path bindings remain 18 Alder Mews, Zone C, Wednesday trash, and gray-cart service. The evidence consists of exactly two complete factual sentences: “The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC.” and “The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC.” The counterfactual coherently changes only the report date to Friday, 12 June 2026, making it later than the unchanged 24-hour policy window. Neither context contains a gold answer, answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route log\",\"text\":\"The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC.\"},{\"speaker\":\"Resident report\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC.\"},{\"speaker\":\"Address and collection records\",\"text\":\"18 Alder Mews is assigned to Zone C. Zone C collects trash on Wednesday, using gray carts, and the cart presented at this address was gray. It was curbside by 07:00 Wednesday. Supplied route evidence shows the scheduled crew did not empty it.\"},{\"speaker\":\"Site inspection\",\"text\":\"The remaining waste is fully contained. It is not exposed, leaking, odorous, disturbed by animals, or affected by pests. No loose waste is outside the cart, and neither the waste nor cart obstructs access or creates an acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC.", "negative_left": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted on Friday, 12 June 2026, at 13:00 UTC.", "right": "The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-155-006", "id": "fast-41-diverse-155-006-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route log", "text": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC."}, {"speaker": "Resident report", "text": "The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC."}, {"speaker": "Address and collection records", "text": "18 Alder Mews is assigned to Zone C. Zone C collects trash on Wednesday, using gray carts, and the cart presented at this address was gray. It was curbside by 07:00 Wednesday. Supplied route evidence shows the scheduled crew did not empty it."}, {"speaker": "Site inspection", "text": "The remaining waste is fully contained. It is not exposed, leaking, odorous, disturbed by animals, or affected by pests. No loose waste is outside the cart, and neither the waste nor cart obstructs access or creates an acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object remains verbatim, and neither context adds, removes, or changes governing policy. Entity and path bindings remain 18 Alder Mews, Zone C, Wednesday trash, and gray-cart service. The evidence consists of exactly two complete factual sentences: “The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC.” and “The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC.” The counterfactual coherently changes only the report date to Friday, 12 June 2026, making it later than the unchanged 24-hour policy window. Neither context contains a gold answer, answer code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route log\",\"text\":\"The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC.\"},{\"speaker\":\"Resident report\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC.\"},{\"speaker\":\"Address and collection records\",\"text\":\"18 Alder Mews is assigned to Zone C. Zone C collects trash on Wednesday, using gray carts, and the cart presented at this address was gray. It was curbside by 07:00 Wednesday. Supplied route evidence shows the scheduled crew did not empty it.\"},{\"speaker\":\"Site inspection\",\"text\":\"The remaining waste is fully contained. It is not exposed, leaking, odorous, disturbed by animals, or affected by pests. No loose waste is outside the cart, and neither the waste nor cart obstructs access or creates an acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC.", "negative_left": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted on Friday, 12 June 2026, at 13:00 UTC.", "right": "The missed-collection report for 18 Alder Mews was submitted on Thursday, 11 June 2026, at 13:00 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-155-006", "id": "fast-41-diverse-155-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route log", "text": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed on Wednesday, 10 June 2026, at 14:00 UTC."}, {"speaker": "Resident report", "text": "The missed-collection report for 18 Alder Mews was submitted on Friday, 12 June 2026, at 13:00 UTC."}, {"speaker": "Address and collection records", "text": "18 Alder Mews is assigned to Zone C. Zone C collects trash on Wednesday, using gray carts, and the cart presented at this address was gray. It was curbside by 07:00 Wednesday. Supplied route evidence shows the scheduled crew did not empty it."}, {"speaker": "Site inspection", "text": "The remaining waste is fully contained. It is not exposed, leaking, odorous, disturbed by animals, or affected by pests. No loose waste is outside the cart, and neither the waste nor cart obstructs access or creates an acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions, bindings, complete evidence quotes, and coherent observations without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route record\",\"text\":\"The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025.\"},{\"speaker\":\"Report record\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on Thursday, 15 May 2025.\"},{\"speaker\":\"Address file\",\"text\":\"18 Alder Mews is assigned to Zone C, which has scheduled trash collection on Wednesday using gray carts.\"},{\"speaker\":\"Collection log\",\"text\":\"The submitted container was a gray cart, and it was curbside by 7:00 a.m. Wednesday. Supplied route evidence shows that the scheduled crew did not empty it.\"},{\"speaker\":\"Site inspection\",\"text\":\"The waste remaining at 18 Alder Mews is fully contained. It is not exposed, leaking, odorous, accessible to animals, or affected by pests.\"},{\"speaker\":\"Site inspection\",\"text\":\"No loose waste is outside the cart, and the waste or cart creates neither an obstruction nor an acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on Thursday, 15 May 2025."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025.", "negative_left": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 7:00 p.m. on Thursday, 15 May 2025.", "right": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on Thursday, 15 May 2025."}, "verifier_independent_model": false}, "family": "fast-41-diverse-155-010", "id": "fast-41-diverse-155-010-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route record", "text": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025."}, {"speaker": "Report record", "text": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on Thursday, 15 May 2025."}, {"speaker": "Address file", "text": "18 Alder Mews is assigned to Zone C, which has scheduled trash collection on Wednesday using gray carts."}, {"speaker": "Collection log", "text": "The submitted container was a gray cart, and it was curbside by 7:00 a.m. Wednesday. Supplied route evidence shows that the scheduled crew did not empty it."}, {"speaker": "Site inspection", "text": "The waste remaining at 18 Alder Mews is fully contained. It is not exposed, leaking, odorous, accessible to animals, or affected by pests."}, {"speaker": "Site inspection", "text": "No loose waste is outside the cart, and the waste or cart creates neither an obstruction nor an acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged questions, bindings, complete evidence quotes, and coherent observations without embedding an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Route record\",\"text\":\"The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025.\"},{\"speaker\":\"Report record\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on Thursday, 15 May 2025.\"},{\"speaker\":\"Address file\",\"text\":\"18 Alder Mews is assigned to Zone C, which has scheduled trash collection on Wednesday using gray carts.\"},{\"speaker\":\"Collection log\",\"text\":\"The submitted container was a gray cart, and it was curbside by 7:00 a.m. Wednesday. Supplied route evidence shows that the scheduled crew did not empty it.\"},{\"speaker\":\"Site inspection\",\"text\":\"The waste remaining at 18 Alder Mews is fully contained. It is not exposed, leaking, odorous, accessible to animals, or affected by pests.\"},{\"speaker\":\"Site inspection\",\"text\":\"No loose waste is outside the cart, and the waste or cart creates neither an obstruction nor an acute dangerous spill.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["0", "text"], "text": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025."}, {"path": ["1", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on Thursday, 15 May 2025."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025.", "negative_left": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 7:00 p.m. on Thursday, 15 May 2025.", "right": "The missed-collection report for 18 Alder Mews was submitted at 5:00 p.m. on Thursday, 15 May 2025."}, "verifier_independent_model": false}, "family": "fast-41-diverse-155-010", "id": "fast-41-diverse-155-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Route record", "text": "The relevant Zone C Wednesday trash route for 18 Alder Mews closed at 6:00 p.m. on Wednesday, 14 May 2025."}, {"speaker": "Report record", "text": "The missed-collection report for 18 Alder Mews was submitted at 7:00 p.m. on Thursday, 15 May 2025."}, {"speaker": "Address file", "text": "18 Alder Mews is assigned to Zone C, which has scheduled trash collection on Wednesday using gray carts."}, {"speaker": "Collection log", "text": "The submitted container was a gray cart, and it was curbside by 7:00 a.m. Wednesday. Supplied route evidence shows that the scheduled crew did not empty it."}, {"speaker": "Site inspection", "text": "The waste remaining at 18 Alder Mews is fully contained. It is not exposed, leaking, odorous, accessible to animals, or affected by pests."}, {"speaker": "Site inspection", "text": "No loose waste is outside the cart, and the waste or cart creates neither an obstruction nor an acute dangerous spill."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, criteria, and instructions. Entity, address, service, zone, schedule, report time, photo, and crew-completion bindings remain intact. The evidence spans are complete factual sentences: “The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report.” “The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday.” The counterfactual's 4:18 p.m. photo and 4:43 p.m. completion time are temporally coherent. Neither context contains an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a 2:10 p.m. Tuesday report concerning trash service at 18 Alder Lane. The address was in municipal collection service, and trash collection was scheduled there that day. Her report requested trash pickup and identified the gray cart shown in the attached photo. The cart was accessible to the assigned Zone C crew when it completed service at the address. The report's only evidence about whether the cart was emptied was that photo, which depicted the cart as unemptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday. The report contains no documentation of an immediate sanitation or obstruction hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report."}, {"path": [], "text": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report.", "negative_left": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report.", "negative_right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 4:43 p.m. on that Tuesday.", "right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-001", "id": "fast-41-diverse-156-001-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a 2:10 p.m. Tuesday report concerning trash service at 18 Alder Lane. The address was in municipal collection service, and trash collection was scheduled there that day. Her report requested trash pickup and identified the gray cart shown in the attached photo. The cart was accessible to the assigned Zone C crew when it completed service at the address. The report's only evidence about whether the cart was emptied was that photo, which depicted the cart as unemptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday. The report contains no documentation of an immediate sanitation or obstruction hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, criteria, and instructions. Entity, address, service, zone, schedule, report time, photo, and crew-completion bindings remain intact. The evidence spans are complete factual sentences: “The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report.” “The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday.” The counterfactual's 4:18 p.m. photo and 4:43 p.m. completion time are temporally coherent. Neither context contains an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a 2:10 p.m. Tuesday report concerning trash service at 18 Alder Lane. The address was in municipal collection service, and trash collection was scheduled there that day. Her report requested trash pickup and identified the gray cart shown in the attached photo. The cart was accessible to the assigned Zone C crew when it completed service at the address. The report's only evidence about whether the cart was emptied was that photo, which depicted the cart as unemptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday. The report contains no documentation of an immediate sanitation or obstruction hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report."}, {"path": [], "text": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report.", "negative_left": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report.", "negative_right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 4:43 p.m. on that Tuesday.", "right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:52 p.m. on that Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-001", "id": "fast-41-diverse-156-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a 2:10 p.m. Tuesday report concerning trash service at 18 Alder Lane. The address was in municipal collection service, and trash collection was scheduled there that day. Her report requested trash pickup and identified the gray cart shown in the attached photo. The cart was accessible to the assigned Zone C crew when it completed service at the address. The report's only evidence about whether the cart was emptied was that photo, which depicted the cart as unemptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the Tuesday when she filed the 2:10 p.m. report. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 4:43 p.m. on that Tuesday. The report contains no documentation of an immediate sanitation or obstruction hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, and neither context adds exceptions or altered priorities. The address, reporter, report time, Tuesday service day, trash type, Zone C, cart, and collection-status bindings remain consistent. The evidence spans are exactly two complete factual sentences: “Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday.” and “The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday.” The counterfactual is coherent because its 4:00 p.m. completion occurs after the 3:20 p.m. photograph. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a 2:10 p.m. Tuesday report about an uncollected trash cart at 18 Alder Lane. The address receives municipal collection, and trash service was scheduled there that day. Her report requested trash pickup and identified the gray cart shown in the attached photograph. The Zone C trash crew was assigned to the address and completed its route there. The cart was accessible to that crew at completion. Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday. The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday. The photograph, which is the report's only evidence about collection status, depicts the identified gray cart as unemptied. The report records no immediate sanitation or obstruction hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday."}, {"path": [], "text": "The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday.", "negative_left": "Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday.", "negative_right": "The assigned Zone C trash crew recorded completion at 4:00 p.m. at 18 Alder Lane on the report Tuesday.", "right": "The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-002", "id": "fast-41-diverse-156-002-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a 2:10 p.m. Tuesday report about an uncollected trash cart at 18 Alder Lane. The address receives municipal collection, and trash service was scheduled there that day. Her report requested trash pickup and identified the gray cart shown in the attached photograph. The Zone C trash crew was assigned to the address and completed its route there. The cart was accessible to that crew at completion. Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday. The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday. The photograph, which is the report's only evidence about collection status, depicts the identified gray cart as unemptied. The report records no immediate sanitation or obstruction hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, and neither context adds exceptions or altered priorities. The address, reporter, report time, Tuesday service day, trash type, Zone C, cart, and collection-status bindings remain consistent. The evidence spans are exactly two complete factual sentences: “Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday.” and “The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday.” The counterfactual is coherent because its 4:00 p.m. completion occurs after the 3:20 p.m. photograph. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a 2:10 p.m. Tuesday report about an uncollected trash cart at 18 Alder Lane. The address receives municipal collection, and trash service was scheduled there that day. Her report requested trash pickup and identified the gray cart shown in the attached photograph. The Zone C trash crew was assigned to the address and completed its route there. The cart was accessible to that crew at completion. Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday. The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday. The photograph, which is the report's only evidence about collection status, depicts the identified gray cart as unemptied. The report records no immediate sanitation or obstruction hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday."}, {"path": [], "text": "The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday.", "negative_left": "Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday.", "negative_right": "The assigned Zone C trash crew recorded completion at 4:00 p.m. at 18 Alder Lane on the report Tuesday.", "right": "The assigned Zone C trash crew recorded completion at 3:00 p.m. at 18 Alder Lane on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-002", "id": "fast-41-diverse-156-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a 2:10 p.m. Tuesday report about an uncollected trash cart at 18 Alder Lane. The address receives municipal collection, and trash service was scheduled there that day. Her report requested trash pickup and identified the gray cart shown in the attached photograph. The Zone C trash crew was assigned to the address and completed its route there. The cart was accessible to that crew at completion. Maya Chen's attached photo was captured at 3:20 p.m. on the report Tuesday. The assigned Zone C trash crew recorded completion at 4:00 p.m. at 18 Alder Lane on the report Tuesday. The photograph, which is the report's only evidence about collection status, depicts the identified gray cart as unemptied. The report records no immediate sanitation or obstruction hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, scope, criteria, and thresholds for both contexts. Both contexts retain Maya Chen, 18 Alder Lane, Tuesday, trash collection, the gray cart, the assigned Zone C crew, and the relevant completion-time binding. The focus evidence consists of two complete factual sentences: \"The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday.\" and \"The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday.\" The counterfactual's 3:18 p.m. photo precedes the 3:42 p.m. completion and does not contradict any assertion about the cart at completion. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a 2:10 p.m. Tuesday report for 18 Alder Lane, which was within municipal collection service and scheduled for trash pickup. Her report requested trash collection and identified the gray cart shown in its attached photo. The Zone C trash crew was assigned to the address and completed its service there at the recorded completion time. The cart was accessible to that crew when service was completed, while the photo depicts it as unemptied. The photo is the report’s only evidence about whether the cart was emptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday. The report records no immediate sanitation-or-obstruction hazard caused by the collection issue.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday."}, {"path": [], "text": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday.", "negative_left": "The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on the report Tuesday.", "negative_right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday.", "right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-004", "id": "fast-41-diverse-156-004-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a 2:10 p.m. Tuesday report for 18 Alder Lane, which was within municipal collection service and scheduled for trash pickup. Her report requested trash collection and identified the gray cart shown in its attached photo. The Zone C trash crew was assigned to the address and completed its service there at the recorded completion time. The cart was accessible to that crew when service was completed, while the photo depicts it as unemptied. The photo is the report’s only evidence about whether the cart was emptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday. The report records no immediate sanitation-or-obstruction hazard caused by the collection issue."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, scope, criteria, and thresholds for both contexts. Both contexts retain Maya Chen, 18 Alder Lane, Tuesday, trash collection, the gray cart, the assigned Zone C crew, and the relevant completion-time binding. The focus evidence consists of two complete factual sentences: \"The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday.\" and \"The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday.\" The counterfactual's 3:18 p.m. photo precedes the 3:42 p.m. completion and does not contradict any assertion about the cart at completion. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a 2:10 p.m. Tuesday report for 18 Alder Lane, which was within municipal collection service and scheduled for trash pickup. Her report requested trash collection and identified the gray cart shown in its attached photo. The Zone C trash crew was assigned to the address and completed its service there at the recorded completion time. The cart was accessible to that crew when service was completed, while the photo depicts it as unemptied. The photo is the report’s only evidence about whether the cart was emptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday. The report records no immediate sanitation-or-obstruction hazard caused by the collection issue.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday."}, {"path": [], "text": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday.", "negative_left": "The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on the report Tuesday.", "negative_right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday.", "right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-004", "id": "fast-41-diverse-156-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a 2:10 p.m. Tuesday report for 18 Alder Lane, which was within municipal collection service and scheduled for trash pickup. Her report requested trash collection and identified the gray cart shown in its attached photo. The Zone C trash crew was assigned to the address and completed its service there at the recorded completion time. The cart was accessible to that crew when service was completed, while the photo depicts it as unemptied. The photo is the report’s only evidence about whether the cart was emptied. The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on the report Tuesday. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday. The report records no immediate sanitation-or-obstruction hazard caused by the collection issue."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions object and preserve the report, address, trash, Tuesday, and crew bindings. The evidence spans are complete factual sentences: \"Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday.\" and \"The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday.\" The counterfactual changes only the completion time to 4:56 p.m., which is coherent with the 4:18 p.m. photo. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen submitted a 2:10 p.m. Tuesday report about trash at 18 Alder Lane. The address was within municipal collection service, and trash pickup was scheduled there that day. Her report requested collection of trash and identified the gray cart shown in the attached photo. The Zone C trash crew was assigned to the address and completed its scheduled service there. The cart was accessible to that crew when service was completed. The photo depicts the reported gray cart as unemptied and is the report’s only evidence about whether it was emptied. Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday. The report contains no documented immediate sanitation or obstruction hazard caused by the collection issue. No other evidence establishes a different collection status or hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday."}, {"path": [], "text": "The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday.", "negative_left": "Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday.", "negative_right": "The assigned Zone C trash crew recorded completion at 4:56 p.m. at 18 Alder Lane on the report Tuesday.", "right": "The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-006", "id": "fast-41-diverse-156-006-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen submitted a 2:10 p.m. Tuesday report about trash at 18 Alder Lane. The address was within municipal collection service, and trash pickup was scheduled there that day. Her report requested collection of trash and identified the gray cart shown in the attached photo. The Zone C trash crew was assigned to the address and completed its scheduled service there. The cart was accessible to that crew when service was completed. The photo depicts the reported gray cart as unemptied and is the report’s only evidence about whether it was emptied. Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday. The report contains no documented immediate sanitation or obstruction hazard caused by the collection issue. No other evidence establishes a different collection status or hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions object and preserve the report, address, trash, Tuesday, and crew bindings. The evidence spans are complete factual sentences: \"Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday.\" and \"The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday.\" The counterfactual changes only the completion time to 4:56 p.m., which is coherent with the 4:18 p.m. photo. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen submitted a 2:10 p.m. Tuesday report about trash at 18 Alder Lane. The address was within municipal collection service, and trash pickup was scheduled there that day. Her report requested collection of trash and identified the gray cart shown in the attached photo. The Zone C trash crew was assigned to the address and completed its scheduled service there. The cart was accessible to that crew when service was completed. The photo depicts the reported gray cart as unemptied and is the report’s only evidence about whether it was emptied. Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday. The report contains no documented immediate sanitation or obstruction hazard caused by the collection issue. No other evidence establishes a different collection status or hazard.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday."}, {"path": [], "text": "The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday.", "negative_left": "Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday.", "negative_right": "The assigned Zone C trash crew recorded completion at 4:56 p.m. at 18 Alder Lane on the report Tuesday.", "right": "The assigned Zone C trash crew recorded completion at 3:42 p.m. at 18 Alder Lane on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-006", "id": "fast-41-diverse-156-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen submitted a 2:10 p.m. Tuesday report about trash at 18 Alder Lane. The address was within municipal collection service, and trash pickup was scheduled there that day. Her report requested collection of trash and identified the gray cart shown in the attached photo. The Zone C trash crew was assigned to the address and completed its scheduled service there. The cart was accessible to that crew when service was completed. The photo depicts the reported gray cart as unemptied and is the report’s only evidence about whether it was emptied. Maya Chen's attached photo was captured at 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew recorded completion at 4:56 p.m. at 18 Alder Lane on the report Tuesday. The report contains no documented immediate sanitation or obstruction hazard caused by the collection issue. No other evidence establishes a different collection status or hazard."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the bound report details and policy-relevant facts, and the evidence spans are complete factual sentences: \"The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday.\" and \"The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday.\" The counterfactual's 3:18 p.m. photo time is coherent because it precedes the unchanged 3:42 p.m. completion time, with no duplicate contradiction or embedded answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a missed-trash report for 18 Alder Lane at 2:10 p.m. Tuesday. The address was within municipal collection service, and municipal trash collection was scheduled there that day. Her report requested trash collection and identified the gray cart shown in her attached photo as the reported container. The Zone C trash crew was assigned to the address and completed service there at its recorded completion time. The cart was accessible to the assigned crew when service was completed. The photo depicts the cart as unemptied and is the report’s only evidence about whether it was emptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday. The report documents no immediate sanitation-or-obstruction hazard caused by the missed collection.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday."}, {"path": [], "text": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday.", "negative_left": "The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on the report Tuesday.", "negative_right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday.", "right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-009", "id": "fast-41-diverse-156-009-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a missed-trash report for 18 Alder Lane at 2:10 p.m. Tuesday. The address was within municipal collection service, and municipal trash collection was scheduled there that day. Her report requested trash collection and identified the gray cart shown in her attached photo as the reported container. The Zone C trash crew was assigned to the address and completed service there at its recorded completion time. The cart was accessible to the assigned crew when service was completed. The photo depicts the cart as unemptied and is the report’s only evidence about whether it was emptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday. The report documents no immediate sanitation-or-obstruction hazard caused by the missed collection."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the bound report details and policy-relevant facts, and the evidence spans are complete factual sentences: \"The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday.\" and \"The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday.\" The counterfactual's 3:18 p.m. photo time is coherent because it precedes the unchanged 3:42 p.m. completion time, with no duplicate contradiction or embedded answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen filed a missed-trash report for 18 Alder Lane at 2:10 p.m. Tuesday. The address was within municipal collection service, and municipal trash collection was scheduled there that day. Her report requested trash collection and identified the gray cart shown in her attached photo as the reported container. The Zone C trash crew was assigned to the address and completed service there at its recorded completion time. The cart was accessible to the assigned crew when service was completed. The photo depicts the cart as unemptied and is the report’s only evidence about whether it was emptied. The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday. The report documents no immediate sanitation-or-obstruction hazard caused by the missed collection.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday."}, {"path": [], "text": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The recorded capture time of Maya Chen's attached photo was 4:18 p.m. on the report Tuesday.", "negative_left": "The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on the report Tuesday.", "negative_right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday.", "right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-009", "id": "fast-41-diverse-156-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen filed a missed-trash report for 18 Alder Lane at 2:10 p.m. Tuesday. The address was within municipal collection service, and municipal trash collection was scheduled there that day. Her report requested trash collection and identified the gray cart shown in her attached photo as the reported container. The Zone C trash crew was assigned to the address and completed service there at its recorded completion time. The cart was accessible to the assigned crew when service was completed. The photo depicts the cart as unemptied and is the report’s only evidence about whether it was emptied. The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on the report Tuesday. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane was 3:42 p.m. on the report Tuesday. The report documents no immediate sanitation-or-obstruction hazard caused by the missed collection."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing rubric and instructions, and neither context adds or alters policy. Both contexts retain Maya Chen, 18 Alder Lane, Tuesday trash service, Zone C, the reported cart, and the relevant crew-completion event. The two evidence spans are complete factual sentences: \"The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025.\" and \"The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m.\" The counterfactual's 2:18 p.m. photo precedes the unchanged 2:47 p.m. completion without contradicting another assertion. Neither context contains a gold answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen reported at 2:10 p.m. Tuesday that municipal trash collection had been missed at 18 Alder Lane, an address within municipal service. Tuesday trash service was scheduled there, and her report identified the requested material as trash. The photographed gray cart was the reported container, and it was accessible to the assigned Zone C trash crew when that crew completed service at the address. The crew was assigned to the location and its service record confirms completion. The attached photo depicts the gray cart as unemptied, and it is the report’s only evidence about whether the cart was emptied. The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m. The report documents no immediate sanitation-or-obstruction hazard caused by the collection issue.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025."}, {"path": [], "text": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025.", "negative_left": "The recorded capture time of Maya Chen's attached photo was 2:18 p.m. on Tuesday, May 6, 2025.", "negative_right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m.", "right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-011", "id": "fast-41-diverse-156-011-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen reported at 2:10 p.m. Tuesday that municipal trash collection had been missed at 18 Alder Lane, an address within municipal service. Tuesday trash service was scheduled there, and her report identified the requested material as trash. The photographed gray cart was the reported container, and it was accessible to the assigned Zone C trash crew when that crew completed service at the address. The crew was assigned to the location and its service record confirms completion. The attached photo depicts the gray cart as unemptied, and it is the report’s only evidence about whether the cart was emptied. The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m. The report documents no immediate sanitation-or-obstruction hazard caused by the collection issue."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing rubric and instructions, and neither context adds or alters policy. Both contexts retain Maya Chen, 18 Alder Lane, Tuesday trash service, Zone C, the reported cart, and the relevant crew-completion event. The two evidence spans are complete factual sentences: \"The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025.\" and \"The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m.\" The counterfactual's 2:18 p.m. photo precedes the unchanged 2:47 p.m. completion without contradicting another assertion. Neither context contains a gold answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Maya Chen reported at 2:10 p.m. Tuesday that municipal trash collection had been missed at 18 Alder Lane, an address within municipal service. Tuesday trash service was scheduled there, and her report identified the requested material as trash. The photographed gray cart was the reported container, and it was accessible to the assigned Zone C trash crew when that crew completed service at the address. The crew was assigned to the location and its service record confirms completion. The attached photo depicts the gray cart as unemptied, and it is the report’s only evidence about whether the cart was emptied. The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m. The report documents no immediate sanitation-or-obstruction hazard caused by the collection issue.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025."}, {"path": [], "text": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The recorded capture time of Maya Chen's attached photo was 3:18 p.m. on Tuesday, May 6, 2025.", "negative_left": "The recorded capture time of Maya Chen's attached photo was 2:18 p.m. on Tuesday, May 6, 2025.", "negative_right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m.", "right": "The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m."}, "verifier_independent_model": false}, "family": "fast-41-diverse-156-011", "id": "fast-41-diverse-156-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Maya Chen reported at 2:10 p.m. Tuesday that municipal trash collection had been missed at 18 Alder Lane, an address within municipal service. Tuesday trash service was scheduled there, and her report identified the requested material as trash. The photographed gray cart was the reported container, and it was accessible to the assigned Zone C trash crew when that crew completed service at the address. The crew was assigned to the location and its service record confirms completion. The attached photo depicts the gray cart as unemptied, and it is the report’s only evidence about whether the cart was emptied. The recorded capture time of Maya Chen's attached photo was 2:18 p.m. on Tuesday, May 6, 2025. The assigned Zone C trash crew's recorded completion time at 18 Alder Lane on Tuesday, May 6, 2025, was 2:47 p.m. The report documents no immediate sanitation-or-obstruction hazard caused by the collection issue."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions and governing policy; bindings remain the activity-room request for September 29, 2026, from 6–8 p.m.; the two evidence quotes are complete factual sentences: “The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18.” and “Urns U-17 and U-18 are standard equipment.”; the counterfactual changes only the urn classification without contradicting counts or other assertions; neither context leaks an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Event organizer\",\"text\":\"The activity-room request is for September 29, 2026, from 6–8 p.m., for 38 attendees. It requests 40 chairs, the built-in projector, and two hot-water urns for packaged drinks. The room request form is attached, and no kitchen use is requested beyond that service.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"},{\"speaker\":\"Inventory attendant\",\"text\":\"The requested equipment consists only of chairs, a built-in projector, and exactly two hot-water urns. The chairs are standard furniture, the projector is built-in audiovisual equipment, and both urns belong to the community center and will serve packaged drinks.\"},{\"speaker\":\"Event organizer\",\"text\":\"The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18.\"},{\"speaker\":\"Inventory attendant\",\"text\":\"Urns U-17 and U-18 are standard equipment.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The activity room is a standard room, and the calendar shows the requested slot available.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["3", "text"], "text": "The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18."}, {"path": ["4", "text"], "text": "Urns U-17 and U-18 are standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18.", "negative_left": "The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18.", "negative_right": "Urns U-17 and U-18 are nonstandard equipment.", "right": "Urns U-17 and U-18 are standard equipment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-157-009", "id": "fast-41-diverse-157-009-base", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Event organizer", "text": "The activity-room request is for September 29, 2026, from 6–8 p.m., for 38 attendees. It requests 40 chairs, the built-in projector, and two hot-water urns for packaged drinks. The room request form is attached, and no kitchen use is requested beyond that service."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}, {"speaker": "Inventory attendant", "text": "The requested equipment consists only of chairs, a built-in projector, and exactly two hot-water urns. The chairs are standard furniture, the projector is built-in audiovisual equipment, and both urns belong to the community center and will serve packaged drinks."}, {"speaker": "Event organizer", "text": "The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18."}, {"speaker": "Inventory attendant", "text": "Urns U-17 and U-18 are standard equipment."}, {"speaker": "Facility booking coordinator", "text": "The activity room is a standard room, and the calendar shows the requested slot available."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_low"}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions and governing policy; bindings remain the activity-room request for September 29, 2026, from 6–8 p.m.; the two evidence quotes are complete factual sentences: “The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18.” and “Urns U-17 and U-18 are standard equipment.”; the counterfactual changes only the urn classification without contradicting counts or other assertions; neither context leaks an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Event organizer\",\"text\":\"The activity-room request is for September 29, 2026, from 6–8 p.m., for 38 attendees. It requests 40 chairs, the built-in projector, and two hot-water urns for packaged drinks. The room request form is attached, and no kitchen use is requested beyond that service.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"},{\"speaker\":\"Inventory attendant\",\"text\":\"The requested equipment consists only of chairs, a built-in projector, and exactly two hot-water urns. The chairs are standard furniture, the projector is built-in audiovisual equipment, and both urns belong to the community center and will serve packaged drinks.\"},{\"speaker\":\"Event organizer\",\"text\":\"The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18.\"},{\"speaker\":\"Inventory attendant\",\"text\":\"Urns U-17 and U-18 are standard equipment.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The activity room is a standard room, and the calendar shows the requested slot available.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["3", "text"], "text": "The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18."}, {"path": ["4", "text"], "text": "Urns U-17 and U-18 are standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18.", "negative_left": "The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18.", "negative_right": "Urns U-17 and U-18 are nonstandard equipment.", "right": "Urns U-17 and U-18 are standard equipment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-157-009", "id": "fast-41-diverse-157-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Event organizer", "text": "The activity-room request is for September 29, 2026, from 6–8 p.m., for 38 attendees. It requests 40 chairs, the built-in projector, and two hot-water urns for packaged drinks. The room request form is attached, and no kitchen use is requested beyond that service."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}, {"speaker": "Inventory attendant", "text": "The requested equipment consists only of chairs, a built-in projector, and exactly two hot-water urns. The chairs are standard furniture, the projector is built-in audiovisual equipment, and both urns belong to the community center and will serve packaged drinks."}, {"speaker": "Event organizer", "text": "The September 29, 2026 activity-room request lists exactly two hot-water urns, identified as U-17 and U-18."}, {"speaker": "Inventory attendant", "text": "Urns U-17 and U-18 are nonstandard equipment."}, {"speaker": "Facility booking coordinator", "text": "The activity room is a standard room, and the calendar shows the requested slot available."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_medium"}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, and the original questions object is unchanged. Dates, time, attendance, activity-room scope, and equipment bindings are preserved. The evidence consists of exactly two complete factual sentences. The counterfactual consistently changes U-417 to nonstandard while leaving U-982 standard. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Event organizer\",\"text\":\"The activity-room request is for September 29, 2026, from 6:00 to 8:00 p.m., for a 38-person lecture. The submission provides the date, time, attendance, requested equipment, and attached room request form.\"},{\"speaker\":\"Event organizer\",\"text\":\"The requested equipment consists of chairs, the built-in projector, and hot-water urns; no other equipment is requested. Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request.\"},{\"speaker\":\"Event organizer\",\"text\":\"U-417 and U-982 are standard equipment.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The chairs are standard furniture, and the projector is built-in audiovisual equipment. The urns are owned by the community center and will be used for packaged drinks.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The requested activity room is a standard room, and the request includes no kitchen use other than the hot-water-urn service.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["1", "text"], "text": "Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request."}, {"path": ["2", "text"], "text": "U-417 and U-982 are standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request.", "negative_left": "Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request.", "negative_right": "U-417 is nonstandard equipment, and U-982 is standard equipment.", "right": "U-417 and U-982 are standard equipment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-157-011", "id": "fast-41-diverse-157-011-base", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Event organizer", "text": "The activity-room request is for September 29, 2026, from 6:00 to 8:00 p.m., for a 38-person lecture. The submission provides the date, time, attendance, requested equipment, and attached room request form."}, {"speaker": "Event organizer", "text": "The requested equipment consists of chairs, the built-in projector, and hot-water urns; no other equipment is requested. Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request."}, {"speaker": "Event organizer", "text": "U-417 and U-982 are standard equipment."}, {"speaker": "Facility booking coordinator", "text": "The chairs are standard furniture, and the projector is built-in audiovisual equipment. The urns are owned by the community center and will be used for packaged drinks."}, {"speaker": "Facility booking coordinator", "text": "The requested activity room is a standard room, and the request includes no kitchen use other than the hot-water-urn service."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_low"}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy, and the original questions object is unchanged. Dates, time, attendance, activity-room scope, and equipment bindings are preserved. The evidence consists of exactly two complete factual sentences. The counterfactual consistently changes U-417 to nonstandard while leaving U-982 standard. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "refuted", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a13": "unknown"}, "remove_right": {"a13": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a13": "unknown"}, "negative_pair": {"a13": "refuted"}, "negative_sentence": {"a13": "unknown"}, "positive_pair": {"a13": "supported"}, "right": {"a13": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including the universally quantified equipment claims. The focus atom concerns whether all requested urns are standard equipment, not a policy conclusion. The base and counter assignments can differ only on that fact: two center-owned urns used for packaged drinks can all be standard in one scenario and include a nonstandard urn in the other without altering the exemption or other facts. Policy evidence correctly preserves the substantive routing, readiness, exception, and complexity rules originating in the original state; calendar availability and staffing observations are not needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every readiness field, establishes attendance at no more than 40, limits equipment to standard furniture, built-in AV, and standard exempt urns, and excludes other kitchen use. The center-owned packaged-drink urn exception therefore supports standard-room routing, booking readiness, and low complexity.", "rule_index": 0, "sound": true}, {"reason": "Refuting the universal claim that every requested urn is standard, together with the existence of exactly two requested urns, entails at least one nonstandard urn. All readiness fields remain supplied; attendance is no more than 40; and the center-owned packaged-drink exception plus the exclusion of other kitchen use prevents kitchen-use routing or high complexity. Nonstandard equipment supports medium complexity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested event date is provided for the September 29, 2026 activity-room request."}, {"id": "a2", "statement": "The requested event time is provided for the September 29, 2026 activity-room request."}, {"id": "a3", "statement": "The requested attendance is provided for the September 29, 2026 activity-room request."}, {"id": "a4", "statement": "The requested attendance for the September 29, 2026 activity-room request is no more than 40 people."}, {"id": "a5", "statement": "The requested equipment is provided for the September 29, 2026 activity-room request."}, {"id": "a6", "statement": "A room request form is attached to the September 29, 2026 activity-room request."}, {"id": "a7", "statement": "Every equipment item requested for the September 29, 2026 activity-room request is a chair, projector, or hot-water urn."}, {"id": "a8", "statement": "Every chair requested for the September 29, 2026 activity-room request is standard furniture."}, {"id": "a9", "statement": "Every projector requested for the September 29, 2026 activity-room request is built-in audiovisual equipment."}, {"id": "a10", "statement": "Exactly two hot-water urns are requested for the September 29, 2026 activity-room request."}, {"id": "a11", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is owned by the community center."}, {"id": "a12", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is used for packaged drinks."}, {"id": "a13", "statement": "Every hot-water urn requested for the September 29, 2026 activity-room request is standard equipment."}, {"id": "a14", "statement": "The activity room requested for September 29, 2026 is a standard room."}, {"id": "a15", "statement": "The September 29, 2026 activity-room request includes no kitchen use other than the requested hot-water-urn service."}], "base_state_json": "[{\"speaker\":\"Event organizer\",\"text\":\"The activity-room request is for September 29, 2026, from 6:00 to 8:00 p.m., for a 38-person lecture. The submission provides the date, time, attendance, requested equipment, and attached room request form.\"},{\"speaker\":\"Event organizer\",\"text\":\"The requested equipment consists of chairs, the built-in projector, and hot-water urns; no other equipment is requested. Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request.\"},{\"speaker\":\"Event organizer\",\"text\":\"U-417 and U-982 are standard equipment.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The chairs are standard furniture, and the projector is built-in audiovisual equipment. The urns are owned by the community center and will be used for packaged drinks.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The requested activity room is a standard room, and the request includes no kitchen use other than the hot-water-urn service.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a13", "focus_evidence": [{"path": ["1", "text"], "text": "Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request."}, {"path": ["2", "text"], "text": "U-417 and U-982 are standard equipment."}], "policy_evidence": [{"path": ["1", "text"], "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}], "rules": [{"justification": "The standard-room request supplies every required readiness field. Its attendance is within 40, all requested furniture is standard, all requested AV is built in, and the two standard center-owned urns are used for packaged drinks. The urn service is exempt from kitchen use, and no other kitchen use is included, so the request remains in the standard-room queue and has low complexity.", "target": "standard_ready_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Exactly two urns are requested, and refuting that every requested urn is standard equipment entails that at least one requested urn is nonstandard equipment. The center-owned packaged-drink exception still prevents the urn service from being kitchen use, and no other kitchen use is included. Every readiness field is supplied, while the nonstandard equipment triggers medium complexity.", "target": "standard_ready_medium", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "refuted"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}]}, "verified_pair": {"left": "Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request.", "negative_left": "Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request.", "negative_right": "U-417 is nonstandard equipment, and U-982 is standard equipment.", "right": "U-417 and U-982 are standard equipment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-157-011", "id": "fast-41-diverse-157-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"kitchen_not_ready_high": "Route to the kitchen-use queue, mark not ready pending a food-handling form, and rate high complexity under the kitchen-use rule.", "kitchen_ready_high": "Route to the facility booking coordinator’s kitchen-use queue, mark booking-ready, and rate high complexity because any heated beverage service counts as kitchen use.", "standard_not_ready_low": "Keep in the standard-room queue and rate low complexity, but mark not ready because a food-handling form is required for the urns.", "standard_ready_low": "Keep in the facility booking coordinator’s standard-room queue, mark booking-ready, and rate low complexity because the urn exception applies and attendance is within 40.", "standard_ready_medium": "Keep in the standard-room queue and mark booking-ready, but rate medium complexity because the urns are nonstandard equipment."}, "instructions": "Choose the single option that correctly determines the request's routing, booking readiness, and operational complexity under the stated scope, exception, and thresholds.", "type": "choice"}}, "state": [{"speaker": "Event organizer", "text": "The activity-room request is for September 29, 2026, from 6:00 to 8:00 p.m., for a 38-person lecture. The submission provides the date, time, attendance, requested equipment, and attached room request form."}, {"speaker": "Event organizer", "text": "The requested equipment consists of chairs, the built-in projector, and hot-water urns; no other equipment is requested. Exactly two hot-water urns, identified as U-417 and U-982, are requested for the September 29, 2026 activity-room request."}, {"speaker": "Event organizer", "text": "U-417 is nonstandard equipment, and U-982 is standard equipment."}, {"speaker": "Facility booking coordinator", "text": "The chairs are standard furniture, and the projector is built-in audiovisual equipment. The urns are owned by the community center and will be used for packaged drinks."}, {"speaker": "Facility booking coordinator", "text": "The requested activity room is a standard room, and the request includes no kitchen use other than the hot-water-urn service."}, {"speaker": "Facility booking coordinator", "text": "Standard-room requests stay in my standard-room queue. Kitchen use goes to the kitchen-use queue and requires a food-handling form. Exception: center-owned urns for packaged drinks are not kitchen use. A request is ready when date, time, attendance, equipment, and room form are provided. Low complexity covers up to 40 people with standard furniture, built-in AV, and exempt urns; medium covers 41–80 or nonstandard equipment; high covers more than 80 or kitchen use."}]}, "method": "c2d", "provenance": {"source_id": "diverse-157", "source_is_synthetic": true, "source_sha256": "87324ef79bc272df0d81481c3cbb62d0a5ab75d5ffb7f556bd9239c932c81997", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "standard_ready_medium"}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking policy and coordinator-confirmation scope. The request remains reservation R-731 for October 14 and the same confirmation decision. The two focus spans are complete factual sentences: \"The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.\" and \"Document F-204 is titled “Food Use Form.”\" The counterfactual coherently changes the form title to “Equipment Use Form,” leaving the required Food Use Form absent without contradictory duplicate facts. Neither context states a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship; a6 is a single negative relationship concerning the set of applicable forms rather than a bundled final classification. The focus a7 concerns whether a form is included, not the decision policy itself. The base and counter assignments can differ only in whether the Food Use Form is included while all other facts remain fixed. Policy evidence correctly cites the original state’s substantive readiness and kitchen-form rules; instructions and criteria in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the selected kitchen, all required reservation details, that no form other than the Food Use Form applies, and inclusion of that form. Therefore every stated readiness requirement is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the selected space is a kitchen while the required Food Use Form is absent. A missing required form is sufficient for a false decision regardless of the other satisfied details.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The selected space for the reservation request is the teaching kitchen."}, {"id": "a2", "statement": "The reservation request includes the event date October 14."}, {"id": "a3", "statement": "The reservation request includes the event time interval 6–9 p.m."}, {"id": "a4", "statement": "The reservation request includes an attendance estimate of 22 attendees."}, {"id": "a5", "statement": "The reservation request includes an equipment-needs specification."}, {"id": "a6", "statement": "No form other than the Food Use Form is required for this selected teaching-kitchen reservation."}, {"id": "a7", "statement": "The reservation request includes a Food Use Form."}], "base_state_json": "{\"context\":\"The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold.\",\"evidence\":[\"The event organizer selected the teaching kitchen for October 14, from 6–9 p.m., expecting 22 attendees.\",\"The reservation lists its equipment needs, including two induction burners and the center’s coffee urn.\",\"The facilities log records no additional form requirement for this teaching-kitchen reservation.\",\"The kitchen capacity is 24, the listed equipment is available, and a closer is scheduled for that evening.\",\"The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.\",\"Document F-204 is titled “Food Use Form.”\",\"Reservation R-731 is under review before coordinator confirmation.\"],\"request\":\"Is reservation R-731 ready for confirmation by the facility booking coordinator?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "4"], "text": "The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204."}, {"path": ["evidence", "5"], "text": "Document F-204 is titled “Food Use Form.”"}], "policy_evidence": [{"path": ["context"], "text": "The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold."}], "rules": [{"justification": "The request supplies every required reservation detail, and the required Food Use Form is included; no other form applies to the selected teaching kitchen.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The selected teaching kitchen requires a Food Use Form, and the request does not include that required form, so it is not booking-ready despite containing the other details.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.", "negative_left": "The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.", "negative_right": "Document F-204 is titled “Equipment Use Form.”", "right": "Document F-204 is titled “Food Use Form.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-159-010", "id": "fast-41-diverse-159-010-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks at least one required detail or applicable form, so it is not ready for coordinator confirmation.", "true": "The request contains all required reservation details and every form required for the selected space, so it is ready for coordinator confirmation."}, "instructions": "Answer yes only if all booking-readiness requirements in the context are satisfied. Answer no if any required detail or form is missing, even when capacity, equipment, and staffing are adequate.", "type": "noul"}}, "state": {"context": "The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold.", "evidence": ["The event organizer selected the teaching kitchen for October 14, from 6–9 p.m., expecting 22 attendees.", "The reservation lists its equipment needs, including two induction burners and the center’s coffee urn.", "The facilities log records no additional form requirement for this teaching-kitchen reservation.", "The kitchen capacity is 24, the listed equipment is available, and a closer is scheduled for that evening.", "The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.", "Document F-204 is titled “Food Use Form.”", "Reservation R-731 is under review before coordinator confirmation."], "request": "Is reservation R-731 ready for confirmation by the facility booking coordinator?"}}, "method": "c2d", "provenance": {"source_id": "diverse-159", "source_is_synthetic": true, "source_sha256": "bef4997ffa59657f3e9d777596ff439ee1a02aa66ae3676bb18f51233df96d1a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing booking policy and coordinator-confirmation scope. The request remains reservation R-731 for October 14 and the same confirmation decision. The two focus spans are complete factual sentences: \"The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.\" and \"Document F-204 is titled “Food Use Form.”\" The counterfactual coherently changes the form title to “Equipment Use Form,” leaving the required Food Use Form absent without contradictory duplicate facts. Neither context states a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship; a6 is a single negative relationship concerning the set of applicable forms rather than a bundled final classification. The focus a7 concerns whether a form is included, not the decision policy itself. The base and counter assignments can differ only in whether the Food Use Form is included while all other facts remain fixed. Policy evidence correctly cites the original state’s substantive readiness and kitchen-form rules; instructions and criteria in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the selected kitchen, all required reservation details, that no form other than the Food Use Form applies, and inclusion of that form. Therefore every stated readiness requirement is satisfied.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the selected space is a kitchen while the required Food Use Form is absent. A missing required form is sufficient for a false decision regardless of the other satisfied details.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The selected space for the reservation request is the teaching kitchen."}, {"id": "a2", "statement": "The reservation request includes the event date October 14."}, {"id": "a3", "statement": "The reservation request includes the event time interval 6–9 p.m."}, {"id": "a4", "statement": "The reservation request includes an attendance estimate of 22 attendees."}, {"id": "a5", "statement": "The reservation request includes an equipment-needs specification."}, {"id": "a6", "statement": "No form other than the Food Use Form is required for this selected teaching-kitchen reservation."}, {"id": "a7", "statement": "The reservation request includes a Food Use Form."}], "base_state_json": "{\"context\":\"The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold.\",\"evidence\":[\"The event organizer selected the teaching kitchen for October 14, from 6–9 p.m., expecting 22 attendees.\",\"The reservation lists its equipment needs, including two induction burners and the center’s coffee urn.\",\"The facilities log records no additional form requirement for this teaching-kitchen reservation.\",\"The kitchen capacity is 24, the listed equipment is available, and a closer is scheduled for that evening.\",\"The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.\",\"Document F-204 is titled “Food Use Form.”\",\"Reservation R-731 is under review before coordinator confirmation.\"],\"request\":\"Is reservation R-731 ready for confirmation by the facility booking coordinator?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "4"], "text": "The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204."}, {"path": ["evidence", "5"], "text": "Document F-204 is titled “Food Use Form.”"}], "policy_evidence": [{"path": ["context"], "text": "The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold."}], "rules": [{"justification": "The request supplies every required reservation detail, and the required Food Use Form is included; no other form applies to the selected teaching kitchen.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The selected teaching kitchen requires a Food Use Form, and the request does not include that required form, so it is not booking-ready despite containing the other details.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.", "negative_left": "The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.", "negative_right": "Document F-204 is titled “Equipment Use Form.”", "right": "Document F-204 is titled “Food Use Form.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-159-010", "id": "fast-41-diverse-159-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks at least one required detail or applicable form, so it is not ready for coordinator confirmation.", "true": "The request contains all required reservation details and every form required for the selected space, so it is ready for coordinator confirmation."}, "instructions": "Answer yes only if all booking-readiness requirements in the context are satisfied. Answer no if any required detail or form is missing, even when capacity, equipment, and staffing are adequate.", "type": "noul"}}, "state": {"context": "The Eastgate Community Center screens reservation requests before the facility booking coordinator confirms them. A request is booking-ready only if it includes the date and time, attendance estimate, equipment needs, and every form required for the selected space. Kitchen reservations always require a Food Use Form, regardless of whether food is sold.", "evidence": ["The event organizer selected the teaching kitchen for October 14, from 6–9 p.m., expecting 22 attendees.", "The reservation lists its equipment needs, including two induction burners and the center’s coffee urn.", "The facilities log records no additional form requirement for this teaching-kitchen reservation.", "The kitchen capacity is 24, the listed equipment is available, and a closer is scheduled for that evening.", "The sole form included in Eastgate Community Center reservation request R-731 for October 14, 2026, is document F-204.", "Document F-204 is titled “Equipment Use Form.”", "Reservation R-731 is under review before coordinator confirmation."], "request": "Is reservation R-731 ready for confirmation by the facility booking coordinator?"}}, "method": "c2d", "provenance": {"source_id": "diverse-159", "source_is_synthetic": true, "source_sha256": "bef4997ffa59657f3e9d777596ff439ee1a02aa66ae3676bb18f51233df96d1a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts add no altered rules. The date, time, teaching kitchen, workshop, and equipment bindings remain consistent. The evidence contains two complete factual sentences: \"The planned attendance for the October 24, 2026, dumpling workshop is 18 people.\" and \"The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people.\" The counterfactual changes capacity to 12 while retaining attendance of 18 without contradictory duplicate assertions. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Booking coordinator\",\"text\":\"The teaching-kitchen reservation request is for October 24, 2026, from 5:30 p.m. through 8:30 p.m.; both the date and the full time interval are available.\"},{\"speaker\":\"Workshop file note\",\"text\":\"The planned attendance for the October 24, 2026, dumpling workshop is 18 people.\"},{\"speaker\":\"Facility record\",\"text\":\"The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people.\"},{\"speaker\":\"Organizer\",\"text\":\"The dumpling workshop requires all four induction burners and no other equipment; the equipment request is complete.\"},{\"speaker\":\"Kitchen technician\",\"text\":\"Every induction burner requested for the October 24, 2026, dumpling workshop is operational.\"},{\"speaker\":\"Booking coordinator\",\"text\":\"The organizer uploaded a signed sheet. It is the required Food Preparation Form for preparing food onsite, identified by its name and purpose, and the form has been accepted in the reservation file.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The planned attendance for the October 24, 2026, dumpling workshop is 18 people."}, {"path": ["2", "text"], "text": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The planned attendance for the October 24, 2026, dumpling workshop is 18 people.", "negative_left": "The planned attendance for the October 24, 2026, dumpling workshop is 18 people.", "negative_right": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 12 people.", "right": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people."}, "verifier_independent_model": false}, "family": "fast-41-diverse-160-007", "id": "fast-41-diverse-160-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Booking coordinator", "text": "The teaching-kitchen reservation request is for October 24, 2026, from 5:30 p.m. through 8:30 p.m.; both the date and the full time interval are available."}, {"speaker": "Workshop file note", "text": "The planned attendance for the October 24, 2026, dumpling workshop is 18 people."}, {"speaker": "Facility record", "text": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people."}, {"speaker": "Organizer", "text": "The dumpling workshop requires all four induction burners and no other equipment; the equipment request is complete."}, {"speaker": "Kitchen technician", "text": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"speaker": "Booking coordinator", "text": "The organizer uploaded a signed sheet. It is the required Food Preparation Form for preparing food onsite, identified by its name and purpose, and the form has been accepted in the reservation file."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts add no altered rules. The date, time, teaching kitchen, workshop, and equipment bindings remain consistent. The evidence contains two complete factual sentences: \"The planned attendance for the October 24, 2026, dumpling workshop is 18 people.\" and \"The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people.\" The counterfactual changes capacity to 12 while retaining attendance of 18 without contradictory duplicate assertions. Neither context embeds an answer, code, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Booking coordinator\",\"text\":\"The teaching-kitchen reservation request is for October 24, 2026, from 5:30 p.m. through 8:30 p.m.; both the date and the full time interval are available.\"},{\"speaker\":\"Workshop file note\",\"text\":\"The planned attendance for the October 24, 2026, dumpling workshop is 18 people.\"},{\"speaker\":\"Facility record\",\"text\":\"The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people.\"},{\"speaker\":\"Organizer\",\"text\":\"The dumpling workshop requires all four induction burners and no other equipment; the equipment request is complete.\"},{\"speaker\":\"Kitchen technician\",\"text\":\"Every induction burner requested for the October 24, 2026, dumpling workshop is operational.\"},{\"speaker\":\"Booking coordinator\",\"text\":\"The organizer uploaded a signed sheet. It is the required Food Preparation Form for preparing food onsite, identified by its name and purpose, and the form has been accepted in the reservation file.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The planned attendance for the October 24, 2026, dumpling workshop is 18 people."}, {"path": ["2", "text"], "text": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The planned attendance for the October 24, 2026, dumpling workshop is 18 people.", "negative_left": "The planned attendance for the October 24, 2026, dumpling workshop is 18 people.", "negative_right": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 12 people.", "right": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 24 people."}, "verifier_independent_model": false}, "family": "fast-41-diverse-160-007", "id": "fast-41-diverse-160-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Booking coordinator", "text": "The teaching-kitchen reservation request is for October 24, 2026, from 5:30 p.m. through 8:30 p.m.; both the date and the full time interval are available."}, {"speaker": "Workshop file note", "text": "The planned attendance for the October 24, 2026, dumpling workshop is 18 people."}, {"speaker": "Facility record", "text": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 12 people."}, {"speaker": "Organizer", "text": "The dumpling workshop requires all four induction burners and no other equipment; the equipment request is complete."}, {"speaker": "Kitchen technician", "text": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"speaker": "Booking coordinator", "text": "The organizer uploaded a signed sheet. It is the required Food Preparation Form for preparing food onsite, identified by its name and purpose, and the form has been accepted in the reservation file."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the teaching-kitchen, date, time, equipment, and form bindings. The evidence contains exactly two complete factual sentences: \"The planned attendance for the October 24, 2026, dumpling workshop is 42 people.\" and \"The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people.\" The counterfactual coherently changes capacity to 39 without duplicating or contradicting another measurement. Neither context includes a gold answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Facility booking coordinator\",\"text\":\"The planned attendance for the October 24, 2026, dumpling workshop is 42 people.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people.\"},{\"speaker\":\"Booking record\",\"text\":\"The teaching-kitchen reservation is requested for October 24, 2026, from 5:30 p.m. through 8:30 p.m., and both the date and interval are open.\"},{\"speaker\":\"Event organizer\",\"text\":\"The workshop's equipment needs are fully listed: all four induction burners, with no other equipment requested.\"},{\"speaker\":\"Kitchen attendant\",\"text\":\"Every requested induction burner is operational.\"},{\"speaker\":\"Booking coordinator\",\"text\":\"The organizer uploaded a signed sheet.\"},{\"speaker\":\"Booking coordinator\",\"text\":\"That sheet is the required Food Preparation Form for preparing food onsite, identified by its name and purpose.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The planned attendance for the October 24, 2026, dumpling workshop is 42 people."}, {"path": ["1", "text"], "text": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The planned attendance for the October 24, 2026, dumpling workshop is 42 people.", "negative_left": "The planned attendance for the October 24, 2026, dumpling workshop is 42 people.", "negative_right": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 39 people.", "right": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people."}, "verifier_independent_model": false}, "family": "fast-41-diverse-160-009", "id": "fast-41-diverse-160-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Facility booking coordinator", "text": "The planned attendance for the October 24, 2026, dumpling workshop is 42 people."}, {"speaker": "Facility booking coordinator", "text": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people."}, {"speaker": "Booking record", "text": "The teaching-kitchen reservation is requested for October 24, 2026, from 5:30 p.m. through 8:30 p.m., and both the date and interval are open."}, {"speaker": "Event organizer", "text": "The workshop's equipment needs are fully listed: all four induction burners, with no other equipment requested."}, {"speaker": "Kitchen attendant", "text": "Every requested induction burner is operational."}, {"speaker": "Booking coordinator", "text": "The organizer uploaded a signed sheet."}, {"speaker": "Booking coordinator", "text": "That sheet is the required Food Preparation Form for preparing food onsite, identified by its name and purpose."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the teaching-kitchen, date, time, equipment, and form bindings. The evidence contains exactly two complete factual sentences: \"The planned attendance for the October 24, 2026, dumpling workshop is 42 people.\" and \"The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people.\" The counterfactual coherently changes capacity to 39 without duplicating or contradicting another measurement. Neither context includes a gold answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Facility booking coordinator\",\"text\":\"The planned attendance for the October 24, 2026, dumpling workshop is 42 people.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people.\"},{\"speaker\":\"Booking record\",\"text\":\"The teaching-kitchen reservation is requested for October 24, 2026, from 5:30 p.m. through 8:30 p.m., and both the date and interval are open.\"},{\"speaker\":\"Event organizer\",\"text\":\"The workshop's equipment needs are fully listed: all four induction burners, with no other equipment requested.\"},{\"speaker\":\"Kitchen attendant\",\"text\":\"Every requested induction burner is operational.\"},{\"speaker\":\"Booking coordinator\",\"text\":\"The organizer uploaded a signed sheet.\"},{\"speaker\":\"Booking coordinator\",\"text\":\"That sheet is the required Food Preparation Form for preparing food onsite, identified by its name and purpose.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The planned attendance for the October 24, 2026, dumpling workshop is 42 people."}, {"path": ["1", "text"], "text": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The planned attendance for the October 24, 2026, dumpling workshop is 42 people.", "negative_left": "The planned attendance for the October 24, 2026, dumpling workshop is 42 people.", "negative_right": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 39 people.", "right": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 48 people."}, "verifier_independent_model": false}, "family": "fast-41-diverse-160-009", "id": "fast-41-diverse-160-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Facility booking coordinator", "text": "The planned attendance for the October 24, 2026, dumpling workshop is 42 people."}, {"speaker": "Facility booking coordinator", "text": "The teaching kitchen's maximum attendance capacity for the October 24, 2026, dumpling workshop is 39 people."}, {"speaker": "Booking record", "text": "The teaching-kitchen reservation is requested for October 24, 2026, from 5:30 p.m. through 8:30 p.m., and both the date and interval are open."}, {"speaker": "Event organizer", "text": "The workshop's equipment needs are fully listed: all four induction burners, with no other equipment requested."}, {"speaker": "Kitchen attendant", "text": "Every requested induction burner is operational."}, {"speaker": "Booking coordinator", "text": "The organizer uploaded a signed sheet."}, {"speaker": "Booking coordinator", "text": "That sheet is the required Food Preparation Form for preparing food onsite, identified by its name and purpose."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original_input includes the questions object verbatim, preserving all governing criteria and instructions. The requester, event, room, date, and time bindings remain unchanged in both contexts. The evidence spans are complete factual sentences: \"Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers.\" and \"Roster NS-417 names Bright Fork Catering as an outside service provider.\" The counterfactual consistently changes the roster to association staff while retaining the same event and operational facts. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Facility coordinator Luis recorded Priya Shah’s neighborhood association meeting in Meeting Room B on October 22, 2026, from 6:00–8:00 p.m. The reservation is for 38 attendees, uses only Meeting Room B, and requests neither kitchen access nor room reconfiguration. Standard projector and microphone use are listed. The community center attendant will cover the event during the regular shift, and the custodial scheduler confirmed that no added staff hours are needed. The date and room are available. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers. Roster NS-417 names Bright Fork Catering as an outside service provider. In a review-only alternate record, the provider roster is replaced with one containing association staff only; all scheduling, attendance, space, equipment, and staffing facts remain unchanged.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers."}, {"path": [], "text": "Roster NS-417 names Bright Fork Catering as an outside service provider."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers.", "negative_left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers.", "negative_right": "Roster NS-417 names only association staff as service providers.", "right": "Roster NS-417 names Bright Fork Catering as an outside service provider."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-001", "id": "fast-41-diverse-161-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Facility coordinator Luis recorded Priya Shah’s neighborhood association meeting in Meeting Room B on October 22, 2026, from 6:00–8:00 p.m. The reservation is for 38 attendees, uses only Meeting Room B, and requests neither kitchen access nor room reconfiguration. Standard projector and microphone use are listed. The community center attendant will cover the event during the regular shift, and the custodial scheduler confirmed that no added staff hours are needed. The date and room are available. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers. Roster NS-417 names Bright Fork Catering as an outside service provider. In a review-only alternate record, the provider roster is replaced with one containing association staff only; all scheduling, attendance, space, equipment, and staffing facts remain unchanged."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original_input includes the questions object verbatim, preserving all governing criteria and instructions. The requester, event, room, date, and time bindings remain unchanged in both contexts. The evidence spans are complete factual sentences: \"Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers.\" and \"Roster NS-417 names Bright Fork Catering as an outside service provider.\" The counterfactual consistently changes the roster to association staff while retaining the same event and operational facts. Neither context contains a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Facility coordinator Luis recorded Priya Shah’s neighborhood association meeting in Meeting Room B on October 22, 2026, from 6:00–8:00 p.m. The reservation is for 38 attendees, uses only Meeting Room B, and requests neither kitchen access nor room reconfiguration. Standard projector and microphone use are listed. The community center attendant will cover the event during the regular shift, and the custodial scheduler confirmed that no added staff hours are needed. The date and room are available. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers. Roster NS-417 names Bright Fork Catering as an outside service provider. In a review-only alternate record, the provider roster is replaced with one containing association staff only; all scheduling, attendance, space, equipment, and staffing facts remain unchanged.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers."}, {"path": [], "text": "Roster NS-417 names Bright Fork Catering as an outside service provider."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers.", "negative_left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers.", "negative_right": "Roster NS-417 names only association staff as service providers.", "right": "Roster NS-417 names Bright Fork Catering as an outside service provider."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-001", "id": "fast-41-diverse-161-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Facility coordinator Luis recorded Priya Shah’s neighborhood association meeting in Meeting Room B on October 22, 2026, from 6:00–8:00 p.m. The reservation is for 38 attendees, uses only Meeting Room B, and requests neither kitchen access nor room reconfiguration. Standard projector and microphone use are listed. The community center attendant will cover the event during the regular shift, and the custodial scheduler confirmed that no added staff hours are needed. The date and room are available. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. identifies roster NS-417 as its complete list of service providers. Roster NS-417 names only association staff as service providers. In a review-only alternate record, the provider roster is replaced with one containing association staff only; all scheduling, attendance, space, equipment, and staffing facts remain unchanged."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing criteria and instructions. Priya Shah, the meeting, date, time, and Meeting Room B remain bound correctly. Both evidence spans are complete factual sentences. The counterfactual changes the provider’s organizational status without creating a contradiction. Neither context states a label, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in Meeting Room B. The reservation lists 38 attendees, so attendance is above 25 but not above 60. The final request does not include kitchen use or room reconfiguration. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services. Harbor Table Services is an independent commercial company not operated by the neighborhood association. The room is the only facility space reserved. The attendant’s work remains within the regular shift, and the custodial scheduler reports zero additional staff hours. The signed form is complete, and the booking coordinator confirmed that the date and room are available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services."}, {"path": [], "text": "Harbor Table Services is an independent commercial company not operated by the neighborhood association."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services.", "negative_left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services.", "negative_right": "Harbor Table Services is an employee of the neighborhood association.", "right": "Harbor Table Services is an independent commercial company not operated by the neighborhood association."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-002", "id": "fast-41-diverse-161-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in Meeting Room B. The reservation lists 38 attendees, so attendance is above 25 but not above 60. The final request does not include kitchen use or room reconfiguration. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services. Harbor Table Services is an independent commercial company not operated by the neighborhood association. The room is the only facility space reserved. The attendant’s work remains within the regular shift, and the custodial scheduler reports zero additional staff hours. The signed form is complete, and the booking coordinator confirmed that the date and room are available."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing criteria and instructions. Priya Shah, the meeting, date, time, and Meeting Room B remain bound correctly. Both evidence spans are complete factual sentences. The counterfactual changes the provider’s organizational status without creating a contradiction. Neither context states a label, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in Meeting Room B. The reservation lists 38 attendees, so attendance is above 25 but not above 60. The final request does not include kitchen use or room reconfiguration. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services. Harbor Table Services is an independent commercial company not operated by the neighborhood association. The room is the only facility space reserved. The attendant’s work remains within the regular shift, and the custodial scheduler reports zero additional staff hours. The signed form is complete, and the booking coordinator confirmed that the date and room are available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services."}, {"path": [], "text": "Harbor Table Services is an independent commercial company not operated by the neighborhood association."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services.", "negative_left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services.", "negative_right": "Harbor Table Services is an employee of the neighborhood association.", "right": "Harbor Table Services is an independent commercial company not operated by the neighborhood association."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-002", "id": "fast-41-diverse-161-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in Meeting Room B. The reservation lists 38 attendees, so attendance is above 25 but not above 60. The final request does not include kitchen use or room reconfiguration. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists exactly one service provider, Harbor Table Services. Harbor Table Services is an employee of the neighborhood association. The room is the only facility space reserved. The attendant’s work remains within the regular shift, and the custodial scheduler reports zero additional staff hours. The signed form is complete, and the booking coordinator confirmed that the date and room are available."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy and scoring criteria. Both contexts retain Priya Shah, the meeting, Meeting Room B, and the specified date and time. The required evidence quotes are \"Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting.\" and \"Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association.\" Both evidence spans are complete factual sentences. The counterfactual changes only the provider’s affiliation, and its in-house status is coherent with the surrounding facts. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in Meeting Room B. The reservation record lists 38 attendees, one facility room, and no kitchen access or room reconfiguration. The requested setup uses the room’s standard projector and one microphone. The attendant will provide that equipment during the regular shift, and the custodial scheduler reports that ordinary post-meeting cleaning is sufficient, with no added staff hours. Priya submitted the final reservation form after confirming the date and room availability with facility booking coordinator Luis. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting. Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association. The coordinator recorded no other special-space request or operational change.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting."}, {"path": [], "text": "Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting.", "negative_left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting.", "negative_right": "Harbor Spoon Catering LLC is an in-house service operated by Priya Shah’s neighborhood association.", "right": "Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-005", "id": "fast-41-diverse-161-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in Meeting Room B. The reservation record lists 38 attendees, one facility room, and no kitchen access or room reconfiguration. The requested setup uses the room’s standard projector and one microphone. The attendant will provide that equipment during the regular shift, and the custodial scheduler reports that ordinary post-meeting cleaning is sufficient, with no added staff hours. Priya submitted the final reservation form after confirming the date and room availability with facility booking coordinator Luis. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting. Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association. The coordinator recorded no other special-space request or operational change."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy and scoring criteria. Both contexts retain Priya Shah, the meeting, Meeting Room B, and the specified date and time. The required evidence quotes are \"Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting.\" and \"Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association.\" Both evidence spans are complete factual sentences. The counterfactual changes only the provider’s affiliation, and its in-house status is coherent with the surrounding facts. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in Meeting Room B. The reservation record lists 38 attendees, one facility room, and no kitchen access or room reconfiguration. The requested setup uses the room’s standard projector and one microphone. The attendant will provide that equipment during the regular shift, and the custodial scheduler reports that ordinary post-meeting cleaning is sufficient, with no added staff hours. Priya submitted the final reservation form after confirming the date and room availability with facility booking coordinator Luis. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting. Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association. The coordinator recorded no other special-space request or operational change.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting."}, {"path": [], "text": "Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting.", "negative_left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting.", "negative_right": "Harbor Spoon Catering LLC is an in-house service operated by Priya Shah’s neighborhood association.", "right": "Harbor Spoon Catering LLC is a company unaffiliated with Priya Shah’s neighborhood association."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-005", "id": "fast-41-diverse-161-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in Meeting Room B. The reservation record lists 38 attendees, one facility room, and no kitchen access or room reconfiguration. The requested setup uses the room’s standard projector and one microphone. The attendant will provide that equipment during the regular shift, and the custodial scheduler reports that ordinary post-meeting cleaning is sufficient, with no added staff hours. Priya submitted the final reservation form after confirming the date and room availability with facility booking coordinator Luis. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists Harbor Spoon Catering LLC as the sole service provider for the meeting. Harbor Spoon Catering LLC is an in-house service operated by Priya Shah’s neighborhood association. The coordinator recorded no other special-space request or operational change."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain Priya Shah, Meeting Room B, October 22, 2026, and 6:00–8:00 p.m. The evidence is factual and complete: \"The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services.\" \"The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting.\" The counterfactual coherently changes only Maple Catering LLC’s classification from outside vendor to association employee, without contradictory duplicate measurements or assertions. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"On October 22, 2026, Priya Shah’s neighborhood association meeting is scheduled for 6:00–8:00 p.m. in Meeting Room B. The signed attendance estimate is 38 people, so it exceeds 25 but does not exceed 60. The final request does not include kitchen use or room reconfiguration and uses only Meeting Room B. Standard equipment will be provided during the attendant’s regular shift; the custodial scheduler reports that ordinary cleaning is sufficient, with no added staff hours. The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services. The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting. Luis, the facility booking coordinator, confirmed that the room and date are available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services."}, {"path": [], "text": "The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services.", "negative_left": "The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services.", "negative_right": "The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an association employee for Priya Shah’s October 22, 2026, meeting.", "right": "The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-008", "id": "fast-41-diverse-161-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "On October 22, 2026, Priya Shah’s neighborhood association meeting is scheduled for 6:00–8:00 p.m. in Meeting Room B. The signed attendance estimate is 38 people, so it exceeds 25 but does not exceed 60. The final request does not include kitchen use or room reconfiguration and uses only Meeting Room B. Standard equipment will be provided during the attendant’s regular shift; the custodial scheduler reports that ordinary cleaning is sufficient, with no added staff hours. The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services. The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting. Luis, the facility booking coordinator, confirmed that the room and date are available."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain Priya Shah, Meeting Room B, October 22, 2026, and 6:00–8:00 p.m. The evidence is factual and complete: \"The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services.\" \"The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting.\" The counterfactual coherently changes only Maple Catering LLC’s classification from outside vendor to association employee, without contradictory duplicate measurements or assertions. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"On October 22, 2026, Priya Shah’s neighborhood association meeting is scheduled for 6:00–8:00 p.m. in Meeting Room B. The signed attendance estimate is 38 people, so it exceeds 25 but does not exceed 60. The final request does not include kitchen use or room reconfiguration and uses only Meeting Room B. Standard equipment will be provided during the attendant’s regular shift; the custodial scheduler reports that ordinary cleaning is sufficient, with no added staff hours. The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services. The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting. Luis, the facility booking coordinator, confirmed that the room and date are available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services."}, {"path": [], "text": "The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services.", "negative_left": "The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services.", "negative_right": "The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an association employee for Priya Shah’s October 22, 2026, meeting.", "right": "The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an outside vendor for Priya Shah’s October 22, 2026, meeting."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-008", "id": "fast-41-diverse-161-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "On October 22, 2026, Priya Shah’s neighborhood association meeting is scheduled for 6:00–8:00 p.m. in Meeting Room B. The signed attendance estimate is 38 people, so it exceeds 25 but does not exceed 60. The final request does not include kitchen use or room reconfiguration and uses only Meeting Room B. Standard equipment will be provided during the attendant’s regular shift; the custodial scheduler reports that ordinary cleaning is sufficient, with no added staff hours. The signed final request for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. names Maple Catering LLC as the only provider for all requested services. The association’s event classification record dated October 23, 2026, classifies Maple Catering LLC as an association employee for Priya Shah’s October 22, 2026, meeting. Luis, the facility booking coordinator, confirmed that the room and date are available."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria. Both contexts retain Priya Shah, the neighborhood meeting, the October 22, 2026 date, the 6:00–8:00 p.m. time, and the facility-use scope. The evidence contains two complete factual sentences: \"Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting.\" and \"The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor.\" The counterfactual coherently changes the vendor count from one to zero without creating duplicate contradictions. Neither context states a score, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in one facility space. The attendance estimate is 38 people, so it is above 25 but not above 60. The final request does not include kitchen use or room reconfiguration. The request requires no added staff time; routine opening, equipment support, and cleanup remain within the regular shift. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting. The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor. The booking coordinator confirmed the room and date, and the signed reservation form is on file.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting."}, {"path": [], "text": "The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting.", "negative_left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting.", "negative_right": "The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists zero outside vendors.", "right": "The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-009", "id": "fast-41-diverse-161-009-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in one facility space. The attendance estimate is 38 people, so it is above 25 but not above 60. The final request does not include kitchen use or room reconfiguration. The request requires no added staff time; routine opening, equipment support, and cleanup remain within the regular shift. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting. The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor. The booking coordinator confirmed the room and date, and the signed reservation form is on file."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria. Both contexts retain Priya Shah, the neighborhood meeting, the October 22, 2026 date, the 6:00–8:00 p.m. time, and the facility-use scope. The evidence contains two complete factual sentences: \"Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting.\" and \"The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor.\" The counterfactual coherently changes the vendor count from one to zero without creating duplicate contradictions. Neither context states a score, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual presence of an outside vendor rather than a policy classification. The base and counter assignments differ only on that focus and are jointly realizable because the policy does not require an outside vendor to entail added staff time or any other listed condition. Empty policy_evidence is correct: all governing scoring rules are already preserved in the questions object, while the original state contains only case-specific observations that need not be retained for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Outside-vendor use triggers level 3. Refuting attendance above 60 excludes all higher attendance thresholds; refuting added staff time above zero excludes the more-than-six-hours trigger; and refuting use of more than one facility space excludes the remaining level-4 trigger. Thus no level-4 condition can supersede level 3.", "rule_index": 0, "sound": true}, {"reason": "Attendance above 25 and not above 60 triggers level 1. The conjunction excludes every higher-level trigger: attendance above 60, kitchen use, room reconfiguration, outside vendors, any added staff time, and multiple-space use are all refuted. Therefore level 1 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 25 people."}, {"id": "a2", "statement": "The requested attendance for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than 60 people."}, {"id": "a3", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes kitchen use."}, {"id": "a4", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes room reconfiguration."}, {"id": "a5", "statement": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. includes at least one outside vendor."}, {"id": "a6", "statement": "The total added staff time required for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. is greater than zero hours."}, {"id": "a7", "statement": "Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. uses more than one facility space."}], "base_state_json": "\"Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in one facility space. The attendance estimate is 38 people, so it is above 25 but not above 60. The final request does not include kitchen use or room reconfiguration. The request requires no added staff time; routine opening, equipment support, and cleanup remain within the regular shift. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting. The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor. The booking coordinator confirmed the room and date, and the signed reservation form is on file.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting."}, {"path": [], "text": "The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor."}], "policy_evidence": [], "rules": [{"justification": "An outside vendor establishes High complexity. Attendance not exceeding 60, zero added staff hours, and use of no more than one facility space exclude every Very high trigger.", "target": "3", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Attendance greater than 25 and not greater than 60 establishes Low complexity. The remaining conditions exclude kitchen use, room reconfiguration, outside vendors, added staff hours, and multiple facility spaces, while the attendance and staffing bounds also exclude all higher attendance and staffing triggers.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting.", "negative_left": "Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting.", "negative_right": "The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists zero outside vendors.", "right": "The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists one outside vendor."}, "verifier_independent_model": false}, "family": "fast-41-diverse-161-009", "id": "fast-41-diverse-161-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: 25 or fewer attendees, no equipment or special space use, and no staffing beyond normal room opening and closing.", "1 — Low: 26–60 attendees or use of standard in-room audiovisual equipment, with no kitchen use, room reconfiguration, outside vendor, or added staff hours.", "2 — Moderate: 61–100 attendees, kitchen use, substantial room reconfiguration, or up to two added staff hours.", "3 — High: 101–175 attendees, an outside vendor, or more than two but no more than six added attendant or custodial hours.", "4 — Very high: More than 175 attendees, use of multiple facility spaces, or more than six added staff hours."], "instructions": "Rate the request’s operational complexity and staffing demand using the five levels below. Use only the explicit evidence in the request and assign the highest level whose stated condition applies.", "type": "score"}}, "state": "Case note: Priya Shah’s neighborhood association meeting is scheduled for October 22, 2026, from 6:00–8:00 p.m. in one facility space. The attendance estimate is 38 people, so it is above 25 but not above 60. The final request does not include kitchen use or room reconfiguration. The request requires no added staff time; routine opening, equipment support, and cleanup remain within the regular shift. Priya Shah’s final request for the neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. contains the only itemized logistics attachment submitted for that meeting. The only itemized logistics attachment submitted for Priya Shah’s neighborhood association meeting on October 22, 2026, from 6:00–8:00 p.m. lists zero outside vendors. The booking coordinator confirmed the room and date, and the signed reservation form is on file."}, "method": "c2d", "provenance": {"source_id": "diverse-161", "source_is_synthetic": true, "source_sha256": "a03554c974d3dcfe27827a9aca647db041abbc23be133e566d636ad60be32ef2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy and scoring criteria. Both contexts retain the requested notice, facility, department, time, and email bindings. The two evidence spans are complete factual sentences. The blank format in the counterfactual does not contradict any other assertion. Neither context contains a gold answer, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled final classification. A2 is a factual count relation and is therefore an appropriate focus. The base and counter assignments differ only in A2 and are realizable: an otherwise identical identifiable, routed request for a physical archived record can either specify a format or omit it. Empty policy_evidence is correct because all governing classification rules, thresholds, priorities, scope, and exceptions originate in the retained questions object; no substantive rule from the original state is needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 establishes identifiability under the two-matching-reference rule; A2 and A3 establish supplied format and delivery; A4 establishes correct routing; and A5 establishes required manual handling of a physical record by Archive Unit staff. Together these are sufficient for complete_manual_archive_retrieval and exclude the incomplete, non-identifiable, and immediate-digital outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes identifiability, while refutation of A2 entails that no copy format is specified. A missing format is sufficient for identifiable_but_incomplete even though delivery, routing, and manual physical retrieval are established. The missing format excludes both complete retrieval outcomes, and A1 excludes not_identifiable.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The number of matching reference details among date, issuing department, facility, and subject between the Records requester's request for the March 2017 Maple Street Gym shutdown notice and the archived notice sought by that request is at least two."}, {"id": "A2", "statement": "The number of copy formats specified in the Records requester's request for the March 2017 Maple Street Gym shutdown notice is at least one."}, {"id": "A3", "statement": "The number of delivery methods specified in the Records requester's request for the March 2017 Maple Street Gym shutdown notice is at least one."}, {"id": "A4", "statement": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice is correctly routed to the Archive Unit."}, {"id": "A5", "statement": "Archive Unit staff must retrieve or scan a physical record manually to fulfill the Records requester's request for the March 2017 Maple Street Gym shutdown notice."}], "base_state_json": "[{\"speaker\":\"Records requester\",\"text\":\"I am seeking the March 2017 shutdown notice for Maple Street Gym, issued by Parks, and request that the result be sent by email to dana@example.test.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"The request has no notice number, but its date, issuing department, facility, and shutdown subject provide matching reference details. The request is correctly routed to the Archive Unit.\"},{\"speaker\":\"Department records coordinator\",\"text\":\"The index confirms one corresponding Parks notice dated March 8, 2017, concerning Maple Street Recreation Center, and identifies no other 2017 notice for that facility.\"},{\"speaker\":\"Records system note\",\"text\":\"The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”\"},{\"speaker\":\"Records system note\",\"text\":\"The value recorded in that field is “PDF.”\"},{\"speaker\":\"Archive technician\",\"text\":\"The entry appears only in a box-level finding aid. Staff must retrieve and scan the paper notice manually; no indexed electronic copy exists.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["3", "text"], "text": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”"}, {"path": ["4", "text"], "text": "The value recorded in that field is “PDF.”"}], "policy_evidence": [], "rules": [{"justification": "At least two matching references make the material identifiable; specified copy format and delivery method make the request complete; routing is correct; and required physical-record retrieval or scanning makes processing manual.", "target": "complete_manual_archive_retrieval", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "At least two matching references make the material identifiable, while refutation of the at-least-one copy-format count entails that no copy format is specified. The request is therefore incomplete even though delivery, routing, and the physical retrieval requirement are known.", "target": "identifiable_but_incomplete", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}]}, "verified_pair": {"left": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”", "negative_left": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”", "negative_right": "The value recorded in that field is blank.", "right": "The value recorded in that field is “PDF.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-175-005", "id": "fast-41-diverse-175-005-base", "input": {"questions": {"decision": {"criteria": {"complete_immediate_digital_retrieval": "The request is complete and correctly routed, and an existing indexed electronic copy can be retrieved immediately without physical archive work.", "complete_manual_archive_retrieval": "The archived material is uniquely identifiable, format and delivery are supplied, routing is correct, and staff must retrieve or scan a physical record manually.", "identifiable_but_incomplete": "The material is identifiable, but either the requested copy format or delivery method is missing.", "not_identifiable": "The material lacks a notice number and fewer than two reference details match, so clarification is required before routing or retrieval."}, "instructions": "Assign the request’s processing status. Apply these ordered rubrics: a paraphrased subject counts as matching reference information only when the index confirms a unique semantic match. A request is identifiable with either a notice number or two matching references among date, issuing department, facility, and subject. It is complete only if it also specifies copy format and delivery method. Archived paper requiring box retrieval is manual retrieval; an existing indexed electronic file is immediate retrieval.", "type": "choice"}}, "state": [{"speaker": "Records requester", "text": "I am seeking the March 2017 shutdown notice for Maple Street Gym, issued by Parks, and request that the result be sent by email to dana@example.test."}, {"speaker": "Municipal information clerk", "text": "The request has no notice number, but its date, issuing department, facility, and shutdown subject provide matching reference details. The request is correctly routed to the Archive Unit."}, {"speaker": "Department records coordinator", "text": "The index confirms one corresponding Parks notice dated March 8, 2017, concerning Maple Street Recreation Center, and identifies no other 2017 notice for that facility."}, {"speaker": "Records system note", "text": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”"}, {"speaker": "Records system note", "text": "The value recorded in that field is “PDF.”"}, {"speaker": "Archive technician", "text": "The entry appears only in a box-level finding aid. Staff must retrieve and scan the paper notice manually; no indexed electronic copy exists."}]}, "method": "c2d", "provenance": {"source_id": "diverse-175", "source_is_synthetic": true, "source_sha256": "2f72697e9f6816680b99ab45bf6d7fe59788e84cf12827594c07bcea161e107a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_manual_archive_retrieval"}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy and scoring criteria. Both contexts retain the requested notice, facility, department, time, and email bindings. The two evidence spans are complete factual sentences. The blank format in the counterfactual does not contradict any other assertion. Neither context contains a gold answer, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled final classification. A2 is a factual count relation and is therefore an appropriate focus. The base and counter assignments differ only in A2 and are realizable: an otherwise identical identifiable, routed request for a physical archived record can either specify a format or omit it. Empty policy_evidence is correct because all governing classification rules, thresholds, priorities, scope, and exceptions originate in the retained questions object; no substantive rule from the original state is needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 establishes identifiability under the two-matching-reference rule; A2 and A3 establish supplied format and delivery; A4 establishes correct routing; and A5 establishes required manual handling of a physical record by Archive Unit staff. Together these are sufficient for complete_manual_archive_retrieval and exclude the incomplete, non-identifiable, and immediate-digital outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes identifiability, while refutation of A2 entails that no copy format is specified. A missing format is sufficient for identifiable_but_incomplete even though delivery, routing, and manual physical retrieval are established. The missing format excludes both complete retrieval outcomes, and A1 excludes not_identifiable.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The number of matching reference details among date, issuing department, facility, and subject between the Records requester's request for the March 2017 Maple Street Gym shutdown notice and the archived notice sought by that request is at least two."}, {"id": "A2", "statement": "The number of copy formats specified in the Records requester's request for the March 2017 Maple Street Gym shutdown notice is at least one."}, {"id": "A3", "statement": "The number of delivery methods specified in the Records requester's request for the March 2017 Maple Street Gym shutdown notice is at least one."}, {"id": "A4", "statement": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice is correctly routed to the Archive Unit."}, {"id": "A5", "statement": "Archive Unit staff must retrieve or scan a physical record manually to fulfill the Records requester's request for the March 2017 Maple Street Gym shutdown notice."}], "base_state_json": "[{\"speaker\":\"Records requester\",\"text\":\"I am seeking the March 2017 shutdown notice for Maple Street Gym, issued by Parks, and request that the result be sent by email to dana@example.test.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"The request has no notice number, but its date, issuing department, facility, and shutdown subject provide matching reference details. The request is correctly routed to the Archive Unit.\"},{\"speaker\":\"Department records coordinator\",\"text\":\"The index confirms one corresponding Parks notice dated March 8, 2017, concerning Maple Street Recreation Center, and identifies no other 2017 notice for that facility.\"},{\"speaker\":\"Records system note\",\"text\":\"The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”\"},{\"speaker\":\"Records system note\",\"text\":\"The value recorded in that field is “PDF.”\"},{\"speaker\":\"Archive technician\",\"text\":\"The entry appears only in a box-level finding aid. Staff must retrieve and scan the paper notice manually; no indexed electronic copy exists.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["3", "text"], "text": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”"}, {"path": ["4", "text"], "text": "The value recorded in that field is “PDF.”"}], "policy_evidence": [], "rules": [{"justification": "At least two matching references make the material identifiable; specified copy format and delivery method make the request complete; routing is correct; and required physical-record retrieval or scanning makes processing manual.", "target": "complete_manual_archive_retrieval", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "At least two matching references make the material identifiable, while refutation of the at-least-one copy-format count entails that no copy format is specified. The request is therefore incomplete even though delivery, routing, and the physical retrieval requirement are known.", "target": "identifiable_but_incomplete", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}]}, "verified_pair": {"left": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”", "negative_left": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”", "negative_right": "The value recorded in that field is blank.", "right": "The value recorded in that field is “PDF.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-175-005", "id": "fast-41-diverse-175-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_immediate_digital_retrieval": "The request is complete and correctly routed, and an existing indexed electronic copy can be retrieved immediately without physical archive work.", "complete_manual_archive_retrieval": "The archived material is uniquely identifiable, format and delivery are supplied, routing is correct, and staff must retrieve or scan a physical record manually.", "identifiable_but_incomplete": "The material is identifiable, but either the requested copy format or delivery method is missing.", "not_identifiable": "The material lacks a notice number and fewer than two reference details match, so clarification is required before routing or retrieval."}, "instructions": "Assign the request’s processing status. Apply these ordered rubrics: a paraphrased subject counts as matching reference information only when the index confirms a unique semantic match. A request is identifiable with either a notice number or two matching references among date, issuing department, facility, and subject. It is complete only if it also specifies copy format and delivery method. Archived paper requiring box retrieval is manual retrieval; an existing indexed electronic file is immediate retrieval.", "type": "choice"}}, "state": [{"speaker": "Records requester", "text": "I am seeking the March 2017 shutdown notice for Maple Street Gym, issued by Parks, and request that the result be sent by email to dana@example.test."}, {"speaker": "Municipal information clerk", "text": "The request has no notice number, but its date, issuing department, facility, and shutdown subject provide matching reference details. The request is correctly routed to the Archive Unit."}, {"speaker": "Department records coordinator", "text": "The index confirms one corresponding Parks notice dated March 8, 2017, concerning Maple Street Recreation Center, and identifies no other 2017 notice for that facility."}, {"speaker": "Records system note", "text": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice contains a field labeled “copy format.”"}, {"speaker": "Records system note", "text": "The value recorded in that field is blank."}, {"speaker": "Archive technician", "text": "The entry appears only in a box-level finding aid. Staff must retrieve and scan the paper notice manually; no indexed electronic copy exists."}]}, "method": "c2d", "provenance": {"source_id": "diverse-175", "source_is_synthetic": true, "source_sha256": "2f72697e9f6816680b99ab45bf6d7fe59788e84cf12827594c07bcea161e107a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identifiable_but_incomplete"}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubrics, while both contexts retain the same request scope, facility, department, date, routing, and delivery binding. The required evidence quotes are complete factual sentences: “In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor.” and “For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format.” The counterfactual changes only the catalog classification and remains internally coherent without contradictory duplicate assertions. Neither context states a status label, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled final classification. A2 is a factual count relation and is therefore an appropriate focus. The base and counter assignments differ only in A2 and are realizable: an otherwise identical identifiable, routed request for a physical archived record can either specify a format or omit it. Empty policy_evidence is correct because all governing classification rules, thresholds, priorities, scope, and exceptions originate in the retained questions object; no substantive rule from the original state is needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 establishes identifiability under the two-matching-reference rule; A2 and A3 establish supplied format and delivery; A4 establishes correct routing; and A5 establishes required manual handling of a physical record by Archive Unit staff. Together these are sufficient for complete_manual_archive_retrieval and exclude the incomplete, non-identifiable, and immediate-digital outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes identifiability, while refutation of A2 entails that no copy format is specified. A missing format is sufficient for identifiable_but_incomplete even though delivery, routing, and manual physical retrieval are established. The missing format excludes both complete retrieval outcomes, and A1 excludes not_identifiable.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The number of matching reference details among date, issuing department, facility, and subject between the Records requester's request for the March 2017 Maple Street Gym shutdown notice and the archived notice sought by that request is at least two."}, {"id": "A2", "statement": "The number of copy formats specified in the Records requester's request for the March 2017 Maple Street Gym shutdown notice is at least one."}, {"id": "A3", "statement": "The number of delivery methods specified in the Records requester's request for the March 2017 Maple Street Gym shutdown notice is at least one."}, {"id": "A4", "statement": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice is correctly routed to the Archive Unit."}, {"id": "A5", "statement": "Archive Unit staff must retrieve or scan a physical record manually to fulfill the Records requester's request for the March 2017 Maple Street Gym shutdown notice."}], "base_state_json": "[{\"speaker\":\"Records requester\",\"text\":\"The request seeks the March 2017 shutdown notice for Maple Street Gym, associated with Parks and the March 8 date.\"},{\"speaker\":\"Department records coordinator\",\"text\":\"The index identifies one relevant Parks notice dated March 8, 2017, concerning Maple Street Recreation Center, with no other 2017 notice for that facility; the subject match is unique.\"},{\"speaker\":\"Records requester\",\"text\":\"The requested delivery method is email to dana@example.test.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"The request is correctly routed to the Archive Unit.\"},{\"speaker\":\"Archive technician\",\"text\":\"The record appears only in a box-level finding aid, so staff must retrieve and scan the physical notice manually; no indexed electronic copy is available.\"},{\"speaker\":\"Records clerk\",\"text\":\"In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor.\"},{\"speaker\":\"Archive catalog note\",\"text\":\"For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["5", "text"], "text": "In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor."}, {"path": ["6", "text"], "text": "For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format."}], "policy_evidence": [], "rules": [{"justification": "At least two matching references make the material identifiable; specified copy format and delivery method make the request complete; routing is correct; and required physical-record retrieval or scanning makes processing manual.", "target": "complete_manual_archive_retrieval", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "At least two matching references make the material identifiable, while refutation of the at-least-one copy-format count entails that no copy format is specified. The request is therefore incomplete even though delivery, routing, and the physical retrieval requirement are known.", "target": "identifiable_but_incomplete", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}]}, "verified_pair": {"left": "In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor.", "negative_left": "In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor.", "negative_right": "For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as not a copy format.", "right": "For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format."}, "verifier_independent_model": false}, "family": "fast-41-diverse-175-018", "id": "fast-41-diverse-175-018-base", "input": {"questions": {"decision": {"criteria": {"complete_immediate_digital_retrieval": "The request is complete and correctly routed, and an existing indexed electronic copy can be retrieved immediately without physical archive work.", "complete_manual_archive_retrieval": "The archived material is uniquely identifiable, format and delivery are supplied, routing is correct, and staff must retrieve or scan a physical record manually.", "identifiable_but_incomplete": "The material is identifiable, but either the requested copy format or delivery method is missing.", "not_identifiable": "The material lacks a notice number and fewer than two reference details match, so clarification is required before routing or retrieval."}, "instructions": "Assign the request’s processing status. Apply these ordered rubrics: a paraphrased subject counts as matching reference information only when the index confirms a unique semantic match. A request is identifiable with either a notice number or two matching references among date, issuing department, facility, and subject. It is complete only if it also specifies copy format and delivery method. Archived paper requiring box retrieval is manual retrieval; an existing indexed electronic file is immediate retrieval.", "type": "choice"}}, "state": [{"speaker": "Records requester", "text": "The request seeks the March 2017 shutdown notice for Maple Street Gym, associated with Parks and the March 8 date."}, {"speaker": "Department records coordinator", "text": "The index identifies one relevant Parks notice dated March 8, 2017, concerning Maple Street Recreation Center, with no other 2017 notice for that facility; the subject match is unique."}, {"speaker": "Records requester", "text": "The requested delivery method is email to dana@example.test."}, {"speaker": "Municipal information clerk", "text": "The request is correctly routed to the Archive Unit."}, {"speaker": "Archive technician", "text": "The record appears only in a box-level finding aid, so staff must retrieve and scan the physical notice manually; no indexed electronic copy is available."}, {"speaker": "Records clerk", "text": "In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor."}, {"speaker": "Archive catalog note", "text": "For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format."}]}, "method": "c2d", "provenance": {"source_id": "diverse-175", "source_is_synthetic": true, "source_sha256": "2f72697e9f6816680b99ab45bf6d7fe59788e84cf12827594c07bcea161e107a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_manual_archive_retrieval"}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubrics, while both contexts retain the same request scope, facility, department, date, routing, and delivery binding. The required evidence quotes are complete factual sentences: “In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor.” and “For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format.” The counterfactual changes only the catalog classification and remains internally coherent without contradictory duplicate assertions. Neither context states a status label, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled final classification. A2 is a factual count relation and is therefore an appropriate focus. The base and counter assignments differ only in A2 and are realizable: an otherwise identical identifiable, routed request for a physical archived record can either specify a format or omit it. Empty policy_evidence is correct because all governing classification rules, thresholds, priorities, scope, and exceptions originate in the retained questions object; no substantive rule from the original state is needed to interpret the task.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 establishes identifiability under the two-matching-reference rule; A2 and A3 establish supplied format and delivery; A4 establishes correct routing; and A5 establishes required manual handling of a physical record by Archive Unit staff. Together these are sufficient for complete_manual_archive_retrieval and exclude the incomplete, non-identifiable, and immediate-digital outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1 establishes identifiability, while refutation of A2 entails that no copy format is specified. A missing format is sufficient for identifiable_but_incomplete even though delivery, routing, and manual physical retrieval are established. The missing format excludes both complete retrieval outcomes, and A1 excludes not_identifiable.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The number of matching reference details among date, issuing department, facility, and subject between the Records requester's request for the March 2017 Maple Street Gym shutdown notice and the archived notice sought by that request is at least two."}, {"id": "A2", "statement": "The number of copy formats specified in the Records requester's request for the March 2017 Maple Street Gym shutdown notice is at least one."}, {"id": "A3", "statement": "The number of delivery methods specified in the Records requester's request for the March 2017 Maple Street Gym shutdown notice is at least one."}, {"id": "A4", "statement": "The Records requester's request for the March 2017 Maple Street Gym shutdown notice is correctly routed to the Archive Unit."}, {"id": "A5", "statement": "Archive Unit staff must retrieve or scan a physical record manually to fulfill the Records requester's request for the March 2017 Maple Street Gym shutdown notice."}], "base_state_json": "[{\"speaker\":\"Records requester\",\"text\":\"The request seeks the March 2017 shutdown notice for Maple Street Gym, associated with Parks and the March 8 date.\"},{\"speaker\":\"Department records coordinator\",\"text\":\"The index identifies one relevant Parks notice dated March 8, 2017, concerning Maple Street Recreation Center, with no other 2017 notice for that facility; the subject match is unique.\"},{\"speaker\":\"Records requester\",\"text\":\"The requested delivery method is email to dana@example.test.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"The request is correctly routed to the Archive Unit.\"},{\"speaker\":\"Archive technician\",\"text\":\"The record appears only in a box-level finding aid, so staff must retrieve and scan the physical notice manually; no indexed electronic copy is available.\"},{\"speaker\":\"Records clerk\",\"text\":\"In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor.\"},{\"speaker\":\"Archive catalog note\",\"text\":\"For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["5", "text"], "text": "In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor."}, {"path": ["6", "text"], "text": "For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format."}], "policy_evidence": [], "rules": [{"justification": "At least two matching references make the material identifiable; specified copy format and delivery method make the request complete; routing is correct; and required physical-record retrieval or scanning makes processing manual.", "target": "complete_manual_archive_retrieval", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "At least two matching references make the material identifiable, while refutation of the at-least-one copy-format count entails that no copy format is specified. The request is therefore incomplete even though delivery, routing, and the physical retrieval requirement are known.", "target": "identifiable_but_incomplete", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}]}, "verified_pair": {"left": "In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor.", "negative_left": "In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor.", "negative_right": "For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as not a copy format.", "right": "For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as a copy format."}, "verifier_independent_model": false}, "family": "fast-41-diverse-175-018", "id": "fast-41-diverse-175-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_immediate_digital_retrieval": "The request is complete and correctly routed, and an existing indexed electronic copy can be retrieved immediately without physical archive work.", "complete_manual_archive_retrieval": "The archived material is uniquely identifiable, format and delivery are supplied, routing is correct, and staff must retrieve or scan a physical record manually.", "identifiable_but_incomplete": "The material is identifiable, but either the requested copy format or delivery method is missing.", "not_identifiable": "The material lacks a notice number and fewer than two reference details match, so clarification is required before routing or retrieval."}, "instructions": "Assign the request’s processing status. Apply these ordered rubrics: a paraphrased subject counts as matching reference information only when the index confirms a unique semantic match. A request is identifiable with either a notice number or two matching references among date, issuing department, facility, and subject. It is complete only if it also specifies copy format and delivery method. Archived paper requiring box retrieval is manual retrieval; an existing indexed electronic file is immediate retrieval.", "type": "choice"}}, "state": [{"speaker": "Records requester", "text": "The request seeks the March 2017 shutdown notice for Maple Street Gym, associated with Parks and the March 8 date."}, {"speaker": "Department records coordinator", "text": "The index identifies one relevant Parks notice dated March 8, 2017, concerning Maple Street Recreation Center, with no other 2017 notice for that facility; the subject match is unique."}, {"speaker": "Records requester", "text": "The requested delivery method is email to dana@example.test."}, {"speaker": "Municipal information clerk", "text": "The request is correctly routed to the Archive Unit."}, {"speaker": "Archive technician", "text": "The record appears only in a box-level finding aid, so staff must retrieve and scan the physical notice manually; no indexed electronic copy is available."}, {"speaker": "Records clerk", "text": "In the Records requester's request for the March 2017 Maple Street Gym shutdown notice, the output-descriptor field contains exactly D-17 and no other output descriptor."}, {"speaker": "Archive catalog note", "text": "For the March 2017 Maple Street Gym shutdown notice, the Archive Unit's 2026-09-17 catalog classifies output descriptor D-17 as not a copy format."}]}, "method": "c2d", "provenance": {"source_id": "diverse-175", "source_is_synthetic": true, "source_sha256": "2f72697e9f6816680b99ab45bf6d7fe59788e84cf12827594c07bcea161e107a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identifiable_but_incomplete"}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and both contexts preserve the governing routing, readiness, fallback, retrieval-order, and difficulty rules. Both contexts retain the Parks Advisory Board, date, reference, requested packet, and fallback path. The two evidence spans are complete factual sentences. The counterfactual changes only availability from unconfirmed to confirmed, without contradiction. Neither context states a gold answer, answer code, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The intake record identifies the Parks Advisory Board, the May 14, 2018 meeting date, and reference PB-2018-05-14. It requests the complete meeting packet and permits the archived agenda and available attachments only if the complete packet cannot be found. The clerk marked the request ready for processing.\"},{\"speaker\":\"Archive Unit routing note\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive Unit audit record\",\"text\":\"In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code U.\"},{\"speaker\":\"Archive Unit audit legend\",\"text\":\"In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED.\"},{\"speaker\":\"Retrieval guidance\",\"text\":\"The complete packet must be checked before the permitted alternative is supplied.\"},{\"speaker\":\"Difficulty guidance\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["2", "text"], "text": "In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code U."}, {"path": ["3", "text"], "text": "In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code U.", "negative_left": "In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code C.", "negative_right": "In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED.", "right": "In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-002", "id": "fast-41-diverse-178-002-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "The intake record identifies the Parks Advisory Board, the May 14, 2018 meeting date, and reference PB-2018-05-14. It requests the complete meeting packet and permits the archived agenda and available attachments only if the complete packet cannot be found. The clerk marked the request ready for processing."}, {"speaker": "Archive Unit routing note", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive Unit audit record", "text": "In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code U."}, {"speaker": "Archive Unit audit legend", "text": "In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED."}, {"speaker": "Retrieval guidance", "text": "The complete packet must be checked before the permitted alternative is supplied."}, {"speaker": "Difficulty guidance", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and both contexts preserve the governing routing, readiness, fallback, retrieval-order, and difficulty rules. Both contexts retain the Parks Advisory Board, date, reference, requested packet, and fallback path. The two evidence spans are complete factual sentences. The counterfactual changes only availability from unconfirmed to confirmed, without contradiction. Neither context states a gold answer, answer code, label rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The intake record identifies the Parks Advisory Board, the May 14, 2018 meeting date, and reference PB-2018-05-14. It requests the complete meeting packet and permits the archived agenda and available attachments only if the complete packet cannot be found. The clerk marked the request ready for processing.\"},{\"speaker\":\"Archive Unit routing note\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive Unit audit record\",\"text\":\"In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code U.\"},{\"speaker\":\"Archive Unit audit legend\",\"text\":\"In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED.\"},{\"speaker\":\"Retrieval guidance\",\"text\":\"The complete packet must be checked before the permitted alternative is supplied.\"},{\"speaker\":\"Difficulty guidance\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["2", "text"], "text": "In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code U."}, {"path": ["3", "text"], "text": "In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code U.", "negative_left": "In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code C.", "negative_right": "In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED.", "right": "In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-002", "id": "fast-41-diverse-178-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "The intake record identifies the Parks Advisory Board, the May 14, 2018 meeting date, and reference PB-2018-05-14. It requests the complete meeting packet and permits the archived agenda and available attachments only if the complete packet cannot be found. The clerk marked the request ready for processing."}, {"speaker": "Archive Unit routing note", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive Unit audit record", "text": "In the Archive Unit's 2026-09-17 audit record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14, the binary exact-item availability field contains the code C."}, {"speaker": "Archive Unit audit legend", "text": "In that same audit record, the field legend defines U as UNCONFIRMED and C as CONFIRMED."}, {"speaker": "Retrieval guidance", "text": "The complete packet must be checked before the permitted alternative is supplied."}, {"speaker": "Difficulty guidance", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original request, routing rule, fallback condition, and difficulty scale, while the unchanged questions object preserves all criteria and instructions. Entity, date, reference, requested packet, and Archive Unit bindings remain unchanged. The required evidence consists of two complete factual sentences: \"In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471.\" and \"Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED.\" The counterfactual coherently changes only the audit status to CONFIRMED, which implies Routine under the preserved scale without asserting an answer. Neither context includes a complete gold answer, answer code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"Request PB-2018-05-14 seeks the complete Parks Advisory Board meeting packet for May 14, 2018. The request identifies the board and date, and the requester authorizes the archived agenda and available attachments only if the complete packet cannot be found.\"},{\"speaker\":\"Archive routing note\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Processing record\",\"text\":\"The submission was marked process-ready after the identifying details and conditional fallback were checked. The reference identifies the requested board meeting, and the requested item is the complete packet.\"},{\"speaker\":\"Archive inventory\",\"text\":\"In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471.\"},{\"speaker\":\"Audit entry\",\"text\":\"Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED.\"},{\"speaker\":\"Retrieval guidance\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["3", "text"], "text": "In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471."}, {"path": ["4", "text"], "text": "Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471.", "negative_left": "In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471.", "negative_right": "Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as CONFIRMED.", "right": "Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-003", "id": "fast-41-diverse-178-003-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "Request PB-2018-05-14 seeks the complete Parks Advisory Board meeting packet for May 14, 2018. The request identifies the board and date, and the requester authorizes the archived agenda and available attachments only if the complete packet cannot be found."}, {"speaker": "Archive routing note", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Processing record", "text": "The submission was marked process-ready after the identifying details and conditional fallback were checked. The reference identifies the requested board meeting, and the requested item is the complete packet."}, {"speaker": "Archive inventory", "text": "In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471."}, {"speaker": "Audit entry", "text": "Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED."}, {"speaker": "Retrieval guidance", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original request, routing rule, fallback condition, and difficulty scale, while the unchanged questions object preserves all criteria and instructions. Entity, date, reference, requested packet, and Archive Unit bindings remain unchanged. The required evidence consists of two complete factual sentences: \"In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471.\" and \"Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED.\" The counterfactual coherently changes only the audit status to CONFIRMED, which implies Routine under the preserved scale without asserting an answer. Neither context includes a complete gold answer, answer code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"Request PB-2018-05-14 seeks the complete Parks Advisory Board meeting packet for May 14, 2018. The request identifies the board and date, and the requester authorizes the archived agenda and available attachments only if the complete packet cannot be found.\"},{\"speaker\":\"Archive routing note\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Processing record\",\"text\":\"The submission was marked process-ready after the identifying details and conditional fallback were checked. The reference identifies the requested board meeting, and the requested item is the complete packet.\"},{\"speaker\":\"Archive inventory\",\"text\":\"In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471.\"},{\"speaker\":\"Audit entry\",\"text\":\"Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED.\"},{\"speaker\":\"Retrieval guidance\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["3", "text"], "text": "In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471."}, {"path": ["4", "text"], "text": "Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471.", "negative_left": "In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471.", "negative_right": "Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as CONFIRMED.", "right": "Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as UNCONFIRMED."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-003", "id": "fast-41-diverse-178-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "Request PB-2018-05-14 seeks the complete Parks Advisory Board meeting packet for May 14, 2018. The request identifies the board and date, and the requester authorizes the archived agenda and available attachments only if the complete packet cannot be found."}, {"speaker": "Archive routing note", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Processing record", "text": "The submission was marked process-ready after the identifying details and conditional fallback were checked. The reference identifies the requested board meeting, and the requested item is the complete packet."}, {"speaker": "Archive inventory", "text": "In the Archive Unit's 2026-09-17 inventory, the binary exact-item availability field for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 is assigned the status recorded in audit entry AU-4471."}, {"speaker": "Audit entry", "text": "Audit entry AU-4471, created by the Archive Unit on 2026-09-17, records that status as CONFIRMED."}, {"speaker": "Retrieval guidance", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the original policy and all request entity, path, and time bindings. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the status-code interpretation, so it remains internally coherent and introduces no duplicate contradiction. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The intake log identifies reference PB-2018-05-14 as a request for the complete Parks Advisory Board meeting packet from the May 14, 2018 meeting. The submission names the board and date, and the intake review marks it process-ready.\"},{\"speaker\":\"Requester\",\"text\":\"If the complete packet cannot be found, the archived agenda and available attachments may be supplied instead; they are not substitutes when the complete packet exists.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive technician\",\"text\":\"The retrieval team will check the referenced packet before using the permitted fallback.\"},{\"speaker\":\"Archive auditor\",\"text\":\"The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17.\"},{\"speaker\":\"Archive technician\",\"text\":\"The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is unconfirmed.\"},{\"speaker\":\"Records clerk\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["4", "text"], "text": "The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17."}, {"path": ["5", "text"], "text": "The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is unconfirmed."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17.", "negative_left": "The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17.", "negative_right": "The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is confirmed.", "right": "The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is unconfirmed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-007", "id": "fast-41-diverse-178-007-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "The intake log identifies reference PB-2018-05-14 as a request for the complete Parks Advisory Board meeting packet from the May 14, 2018 meeting. The submission names the board and date, and the intake review marks it process-ready."}, {"speaker": "Requester", "text": "If the complete packet cannot be found, the archived agenda and available attachments may be supplied instead; they are not substitutes when the complete packet exists."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive technician", "text": "The retrieval team will check the referenced packet before using the permitted fallback."}, {"speaker": "Archive auditor", "text": "The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17."}, {"speaker": "Archive technician", "text": "The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is unconfirmed."}, {"speaker": "Records clerk", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the original policy and all request entity, path, and time bindings. The two focus-evidence spans are complete factual sentences. The counterfactual changes only the status-code interpretation, so it remains internally coherent and introduces no duplicate contradiction. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The intake log identifies reference PB-2018-05-14 as a request for the complete Parks Advisory Board meeting packet from the May 14, 2018 meeting. The submission names the board and date, and the intake review marks it process-ready.\"},{\"speaker\":\"Requester\",\"text\":\"If the complete packet cannot be found, the archived agenda and available attachments may be supplied instead; they are not substitutes when the complete packet exists.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive technician\",\"text\":\"The retrieval team will check the referenced packet before using the permitted fallback.\"},{\"speaker\":\"Archive auditor\",\"text\":\"The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17.\"},{\"speaker\":\"Archive technician\",\"text\":\"The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is unconfirmed.\"},{\"speaker\":\"Records clerk\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["4", "text"], "text": "The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17."}, {"path": ["5", "text"], "text": "The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is unconfirmed."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17.", "negative_left": "The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17.", "negative_right": "The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is confirmed.", "right": "The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is unconfirmed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-007", "id": "fast-41-diverse-178-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "The intake log identifies reference PB-2018-05-14 as a request for the complete Parks Advisory Board meeting packet from the May 14, 2018 meeting. The submission names the board and date, and the intake review marks it process-ready."}, {"speaker": "Requester", "text": "If the complete packet cannot be found, the archived agenda and available attachments may be supplied instead; they are not substitutes when the complete packet exists."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive technician", "text": "The retrieval team will check the referenced packet before using the permitted fallback."}, {"speaker": "Archive auditor", "text": "The Archive Unit's audit entry dated September 16, 2026, for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 lists status code U-17."}, {"speaker": "Archive technician", "text": "The Archive Unit's codebook dated September 16, 2026, records status code U-17 as meaning that exact-item availability is confirmed."}, {"speaker": "Records clerk", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings without answer leakage; the complete factual evidence sentences are “The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14.” and “The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field.”; the counterfactual coherently changes the mapping to CONFIRMED without contradictory duplicates.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records requester\",\"text\":\"Please process reference PB-2018-05-14 for the complete Parks Advisory Board meeting packet from May 14, 2018. If that complete packet cannot be found, the archived agenda and available attachments may be supplied instead; do not substitute them if the complete packet exists.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Records intake note\",\"text\":\"The reference identifies the Parks Advisory Board and the May 14, 2018 meeting. The intake record contains the body, meeting date, requested complete packet, and conditional fallback, so the request is process-ready.\"},{\"speaker\":\"Archive Unit export\",\"text\":\"The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14.\"},{\"speaker\":\"Archive Unit data dictionary\",\"text\":\"The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["3", "text"], "text": "The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14."}, {"path": ["4", "text"], "text": "The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14.", "negative_left": "The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14.", "negative_right": "The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to CONFIRMED and code 1 to UNCONFIRMED for that binary field.", "right": "The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-008", "id": "fast-41-diverse-178-008-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records requester", "text": "Please process reference PB-2018-05-14 for the complete Parks Advisory Board meeting packet from May 14, 2018. If that complete packet cannot be found, the archived agenda and available attachments may be supplied instead; do not substitute them if the complete packet exists."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Records intake note", "text": "The reference identifies the Parks Advisory Board and the May 14, 2018 meeting. The intake record contains the body, meeting date, requested complete packet, and conditional fallback, so the request is process-ready."}, {"speaker": "Archive Unit export", "text": "The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14."}, {"speaker": "Archive Unit data dictionary", "text": "The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings without answer leakage; the complete factual evidence sentences are “The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14.” and “The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field.”; the counterfactual coherently changes the mapping to CONFIRMED without contradictory duplicates.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records requester\",\"text\":\"Please process reference PB-2018-05-14 for the complete Parks Advisory Board meeting packet from May 14, 2018. If that complete packet cannot be found, the archived agenda and available attachments may be supplied instead; do not substitute them if the complete packet exists.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Records intake note\",\"text\":\"The reference identifies the Parks Advisory Board and the May 14, 2018 meeting. The intake record contains the body, meeting date, requested complete packet, and conditional fallback, so the request is process-ready.\"},{\"speaker\":\"Archive Unit export\",\"text\":\"The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14.\"},{\"speaker\":\"Archive Unit data dictionary\",\"text\":\"The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["3", "text"], "text": "The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14."}, {"path": ["4", "text"], "text": "The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14.", "negative_left": "The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14.", "negative_right": "The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to CONFIRMED and code 1 to UNCONFIRMED for that binary field.", "right": "The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to UNCONFIRMED and code 1 to CONFIRMED for that binary field."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-008", "id": "fast-41-diverse-178-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records requester", "text": "Please process reference PB-2018-05-14 for the complete Parks Advisory Board meeting packet from May 14, 2018. If that complete packet cannot be found, the archived agenda and available attachments may be supplied instead; do not substitute them if the complete packet exists."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Records intake note", "text": "The reference identifies the Parks Advisory Board and the May 14, 2018 meeting. The intake record contains the body, meeting date, requested complete packet, and conditional fallback, so the request is process-ready."}, {"speaker": "Archive Unit export", "text": "The Archive Unit's 2026-09-17 export records binary exact-item availability code 0 for the complete packet requested under reference PB-2018-05-14."}, {"speaker": "Archive Unit data dictionary", "text": "The Archive Unit's signed 2026-09-17 data dictionary maps code 0 to CONFIRMED and code 1 to UNCONFIRMED for that binary field."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question and policy, retain the required bindings, include the exact evidence quotes “The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17.” and “The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED.”, and present a coherent observation change without leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"Case note: Reference PB-2018-05-14 identifies the complete Parks Advisory Board meeting packet for the board's May 14, 2018 meeting. The meeting date is before January 1, 2022, so the matter is handled as a legacy archive request.\"},{\"speaker\":\"Records requester\",\"text\":\"If the complete packet cannot be found, please provide the archived agenda and any available attachments. That fallback applies only when the requested complete packet is not found.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"},{\"speaker\":\"Archive technician\",\"text\":\"The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17.\"},{\"speaker\":\"Archive technician\",\"text\":\"The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["4", "text"], "text": "The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17."}, {"path": ["5", "text"], "text": "The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17.", "negative_left": "The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17.", "negative_right": "The Archive Unit's 2026-09-17 code legend states that U-17 means CONFIRMED rather than UNCONFIRMED.", "right": "The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-009", "id": "fast-41-diverse-178-009-base", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "Case note: Reference PB-2018-05-14 identifies the complete Parks Advisory Board meeting packet for the board's May 14, 2018 meeting. The meeting date is before January 1, 2022, so the matter is handled as a legacy archive request."}, {"speaker": "Records requester", "text": "If the complete packet cannot be found, please provide the archived agenda and any available attachments. That fallback applies only when the requested complete packet is not found."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}, {"speaker": "Archive technician", "text": "The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17."}, {"speaker": "Archive technician", "text": "The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question and policy, retain the required bindings, include the exact evidence quotes “The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17.” and “The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED.”, and present a coherent observation change without leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual status, identity, date relation, permission, or field value. The focus A5 is a factual availability-field relation rather than a policy rule. The cited state evidence preserves the substantive routing/readiness and difficulty rules needed to interpret the unchanged question; case-specific observations need not be preserved. Both assignments are realizable with only availability confirmation changing: the request can remain process-ready, correctly routed, and equipped with the same fallback whether exact availability is unconfirmed or confirmed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes pre-2022 Archive Unit routing, process readiness, an exact packet request, an acceptable conditional fallback, and unconfirmed exact-item availability. Under the preserved difficulty scale, checking the exact item before using that fallback is Moderate, so the requested combined outcome follows.", "rule_index": 0, "sound": true}, {"reason": "A5 explicitly describes a binary UNCONFIRMED-versus-CONFIRMED field. Refuting A5 therefore entails CONFIRMED rather than merely leaving availability unknown. Confirmed exact availability is Routine, not Moderate, so the combined Archive Unit/process-ready/Moderate outcome is false even though routing and readiness remain satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The meeting date of the Parks Advisory Board record requested under reference PB-2018-05-14 is earlier than January 1, 2022."}, {"id": "A2", "statement": "The records request is process-ready."}, {"id": "A3", "statement": "The exact item requested under reference PB-2018-05-14 is the complete Parks Advisory Board meeting packet."}, {"id": "A4", "statement": "The requester permits the archived agenda and available attachments to be supplied if the complete packet requested under reference PB-2018-05-14 cannot be found."}, {"id": "A5", "statement": "The Archive Unit's binary exact-item availability field for the complete packet requested under reference PB-2018-05-14 has the value UNCONFIRMED rather than CONFIRMED."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"Case note: Reference PB-2018-05-14 identifies the complete Parks Advisory Board meeting packet for the board's May 14, 2018 meeting. The meeting date is before January 1, 2022, so the matter is handled as a legacy archive request.\"},{\"speaker\":\"Records requester\",\"text\":\"If the complete packet cannot be found, please provide the archived agenda and any available attachments. That fallback applies only when the requested complete packet is not found.\"},{\"speaker\":\"Municipal information clerk\",\"text\":\"Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback.\"},{\"speaker\":\"Archive technician\",\"text\":\"Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required.\"},{\"speaker\":\"Archive technician\",\"text\":\"The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17.\"},{\"speaker\":\"Archive technician\",\"text\":\"The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A5", "focus_evidence": [{"path": ["4", "text"], "text": "The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17."}, {"path": ["5", "text"], "text": "The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED."}], "policy_evidence": [{"path": ["1", "text"], "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"path": ["3", "text"], "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}], "rules": [{"justification": "The pre-2022 date routes the process-ready request to the Archive Unit. Because the exact packet's availability is unconfirmed and the requester supplied a fallback that applies only if that packet cannot be found, the exact packet must be checked before the fallback is applied; the stated scale classifies that retrieval as Moderate.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}]}, {"justification": "The pre-2022 date still routes the process-ready request to the Archive Unit. Because the binary availability field is not UNCONFIRMED, it is CONFIRMED, and the stated scale classifies confirmed exact availability as Routine rather than Moderate. The requested combined Archive Unit, process-ready, and Moderate outcome is therefore false.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17.", "negative_left": "The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17.", "negative_right": "The Archive Unit's 2026-09-17 code legend states that U-17 means CONFIRMED rather than UNCONFIRMED.", "right": "The Archive Unit's 2026-09-17 code legend states that U-17 means UNCONFIRMED rather than CONFIRMED."}, "verifier_independent_model": false}, "family": "fast-41-diverse-178-009", "id": "fast-41-diverse-178-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The request lacks required details or fallback permission, routes elsewhere, or does not meet the Moderate category.", "true": "The request is complete enough to process, routes to the Archive Unit, and receives a Moderate difficulty rating."}, "instructions": "Answer yes or no: Should this request be routed to the Archive Unit and classified as process-ready with Moderate retrieval difficulty under the stated rules?", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "Case note: Reference PB-2018-05-14 identifies the complete Parks Advisory Board meeting packet for the board's May 14, 2018 meeting. The meeting date is before January 1, 2022, so the matter is handled as a legacy archive request."}, {"speaker": "Records requester", "text": "If the complete packet cannot be found, please provide the archived agenda and any available attachments. That fallback applies only when the requested complete packet is not found."}, {"speaker": "Municipal information clerk", "text": "Requests dated before 2022 go to the Archive Unit. A legacy request is process-ready only if it identifies the body and date and, when exact-item availability is uncertain, gives an explicit acceptable fallback."}, {"speaker": "Archive technician", "text": "Our difficulty scale is Routine when exact availability is confirmed, Moderate when we must check for an exact item before applying a supplied fallback, and Complex when clarification is required."}, {"speaker": "Archive technician", "text": "The Archive Unit's 2026-09-17 availability record for the complete Parks Advisory Board meeting packet requested under reference PB-2018-05-14 displays code U-17."}, {"speaker": "Archive technician", "text": "The Archive Unit's 2026-09-17 code legend states that U-17 means CONFIRMED rather than UNCONFIRMED."}]}, "method": "c2d", "provenance": {"source_id": "diverse-178", "source_is_synthetic": true, "source_sha256": "c50b28b311005f594a9965926b8bddf4019e3843c320fbdb34e306e1a41493bb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing policy, question scope, washer entity, and 9:10 p.m. binding are preserved in both contexts. The evidence consists of two complete factual sentences: \"The two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.\" and \"The 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier.\" The counterfactual consistently changes only the spill-to-washer identifier relation, without contradictory duplicate assertions. Neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"At 9:10 p.m., a floor-spill sample and two possible sources were logged: the uncapped detergent bottle on the shelf and the washer. The bottle is outside the washer, and its liquid is a detergent product. Each source-profile identifier used in this inquiry uniquely identifies its associated possible origin within the investigated set, and the spill has exactly one origin. No active flooding or electrical signs are present at the washer. Washer logs show no recurring fault, while the cycle notes do not confirm either an unbalanced load or an overloaded load.\\n\\nThe two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.\\nThe 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier.\\n\\nPolicy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency.\\nAppliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set."}, {"path": [], "text": "The 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.", "negative_left": "The two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.", "negative_right": "The 9:10 p.m. floor-spill identifier is identical to the washer-moisture identifier.", "right": "The 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier."}, "verifier_independent_model": false}, "family": "fast-41-diverse-181-003", "id": "fast-41-diverse-181-003-base", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "At 9:10 p.m., a floor-spill sample and two possible sources were logged: the uncapped detergent bottle on the shelf and the washer. The bottle is outside the washer, and its liquid is a detergent product. Each source-profile identifier used in this inquiry uniquely identifies its associated possible origin within the investigated set, and the spill has exactly one origin. No active flooding or electrical signs are present at the washer. Washer logs show no recurring fault, while the cycle notes do not confirm either an unbalanced load or an overloaded load.\n\nThe two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.\nThe 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier.\n\nPolicy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency.\nAppliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing policy, question scope, washer entity, and 9:10 p.m. binding are preserved in both contexts. The evidence consists of two complete factual sentences: \"The two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.\" and \"The 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier.\" The counterfactual consistently changes only the spill-to-washer identifier relation, without contradictory duplicate assertions. Neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"At 9:10 p.m., a floor-spill sample and two possible sources were logged: the uncapped detergent bottle on the shelf and the washer. The bottle is outside the washer, and its liquid is a detergent product. Each source-profile identifier used in this inquiry uniquely identifies its associated possible origin within the investigated set, and the spill has exactly one origin. No active flooding or electrical signs are present at the washer. Washer logs show no recurring fault, while the cycle notes do not confirm either an unbalanced load or an overloaded load.\\n\\nThe two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.\\nThe 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier.\\n\\nPolicy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency.\\nAppliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set."}, {"path": [], "text": "The 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.", "negative_left": "The two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.", "negative_right": "The 9:10 p.m. floor-spill identifier is identical to the washer-moisture identifier.", "right": "The 9:10 p.m. floor-spill identifier differs from the washer-moisture identifier."}, "verifier_independent_model": false}, "family": "fast-41-diverse-181-003", "id": "fast-41-diverse-181-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "At 9:10 p.m., a floor-spill sample and two possible sources were logged: the uncapped detergent bottle on the shelf and the washer. The bottle is outside the washer, and its liquid is a detergent product. Each source-profile identifier used in this inquiry uniquely identifies its associated possible origin within the investigated set, and the spill has exactly one origin. No active flooding or electrical signs are present at the washer. Washer logs show no recurring fault, while the cycle notes do not confirm either an unbalanced load or an overloaded load.\n\nThe two candidate source-profile identifiers are exactly the bottle-liquid identifier and the washer-moisture identifier, those identifiers differ, and the 9:10 p.m. floor-spill identifier belongs to that two-member set.\nThe 9:10 p.m. floor-spill identifier is identical to the washer-moisture identifier.\n\nPolicy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency.\nAppliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "basic_inspection_hold_high"}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and all question entity, path, and time bindings. The two evidence spans are complete factual sentences, and the counterfactual consistently changes the bottle identifier while allowing the washer identifier to be DP-417. Neither context embeds an answer, label code, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"At 9:10 p.m., investigators collected a floor-spill sample and liquid from an uncapped detergent bottle on a shelf. The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417. The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was DP-417. The floor-spill identifier is among the bottle-liquid identifier and the washer-moisture identifier. The bottle-liquid identifier differs from the washer-moisture identifier. Each source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set. That set consists exactly of the uncapped detergent bottle on the shelf and the washer, and the spill has exactly one origin. The bottle is external to the washer, and its liquid is a detergent product. No active flooding or electrical signs are present at the washer. The washer logs show no recurring fault. The cycle notes do not confirm an unbalanced load or an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417."}, {"path": [], "text": "The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was DP-417."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417.", "negative_left": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417.", "negative_right": "The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was WM-892.", "right": "The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was DP-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-181-005", "id": "fast-41-diverse-181-005-base", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "At 9:10 p.m., investigators collected a floor-spill sample and liquid from an uncapped detergent bottle on a shelf. The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417. The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was DP-417. The floor-spill identifier is among the bottle-liquid identifier and the washer-moisture identifier. The bottle-liquid identifier differs from the washer-moisture identifier. Each source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set. That set consists exactly of the uncapped detergent bottle on the shelf and the washer, and the spill has exactly one origin. The bottle is external to the washer, and its liquid is a detergent product. No active flooding or electrical signs are present at the washer. The washer logs show no recurring fault. The cycle notes do not confirm an unbalanced load or an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and all question entity, path, and time bindings. The two evidence spans are complete factual sentences, and the counterfactual consistently changes the bottle identifier while allowing the washer identifier to be DP-417. Neither context embeds an answer, label code, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"At 9:10 p.m., investigators collected a floor-spill sample and liquid from an uncapped detergent bottle on a shelf. The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417. The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was DP-417. The floor-spill identifier is among the bottle-liquid identifier and the washer-moisture identifier. The bottle-liquid identifier differs from the washer-moisture identifier. Each source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set. That set consists exactly of the uncapped detergent bottle on the shelf and the washer, and the spill has exactly one origin. The bottle is external to the washer, and its liquid is a detergent product. No active flooding or electrical signs are present at the washer. The washer logs show no recurring fault. The cycle notes do not confirm an unbalanced load or an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417."}, {"path": [], "text": "The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was DP-417."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417.", "negative_left": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417.", "negative_right": "The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was WM-892.", "right": "The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was DP-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-181-005", "id": "fast-41-diverse-181-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "At 9:10 p.m., investigators collected a floor-spill sample and liquid from an uncapped detergent bottle on a shelf. The source-profile identifier measured for the 9:10 p.m. floor-spill sample was DP-417. The source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf was WM-892. The floor-spill identifier is among the bottle-liquid identifier and the washer-moisture identifier. The bottle-liquid identifier differs from the washer-moisture identifier. Each source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set. That set consists exactly of the uncapped detergent bottle on the shelf and the washer, and the spill has exactly one origin. The bottle is external to the washer, and its liquid is a detergent product. No active flooding or electrical signs are present at the washer. The washer logs show no recurring fault. The cycle notes do not confirm an unbalanced load or an overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "basic_inspection_hold_high"}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the washer, reuse condition, test-run identities, and time relation without conflicting duplicate facts. The two evidence quotes are complete factual sentences, and changing S from two to four coherently changes R from zero to two. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "refuted", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A15": "unknown"}, "remove_right": {"A15": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A15": "unknown"}, "negative_pair": {"A15": "refuted"}, "negative_sentence": {"A15": "unknown"}, "positive_pair": {"A15": "supported"}, "right": {"A15": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a factual relationship; interval-bounded and uniqueness propositions remain atomic despite quantifying over times or tests. A15 is a factual observation about the banging count, not a policy conclusion. The base and counter assignments differ only on A15 and are realizable: a spin may be balanced yet still produce banging. Policy evidence correctly preserves the resident’s state-originating, case-specific reuse condition, while the general governing readiness rule remains automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that inspector P completed the resident’s specified test in the relevant interval: an empty, balanced spin with zero banging. This successfully fulfills the express reuse condition, so readiness follows. The additional incident and uniqueness conditions are unnecessary but do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "A15 being refuted entails that test R had a nonzero banging count, so it did not satisfy the required no-banging test. A16 excludes any other completed inspector spin test in the relevant interval, preventing a competing successful test. Therefore the resident’s express condition remains unfulfilled and the false target follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During the period covered by the evidence through readiness decision D, incident E was the washer's only banging incident."}, {"id": "A2", "statement": "The washer contained exactly one item during incident E."}, {"id": "A3", "statement": "The item in the washer was off-center during incident E."}, {"id": "A4", "statement": "No water leaked from the washer during or after incident E through readiness decision D."}, {"id": "A5", "statement": "No odor was detected from the washer during or after incident E through readiness decision D."}, {"id": "A6", "statement": "No scraping was detected from the washer during or after incident E through readiness decision D."}, {"id": "A7", "statement": "Inspection of the washer after incident E detected no looseness."}, {"id": "A8", "statement": "Inspection of the washer after incident E detected no damage."}, {"id": "A9", "statement": "Person P was an inspector at the time of test run R."}, {"id": "A10", "statement": "Person P performed test run R on the washer."}, {"id": "A11", "statement": "Test run R was completed."}, {"id": "A12", "statement": "Test run R occurred after the resident stated the reuse condition and before readiness decision D."}, {"id": "A13", "statement": "The washer drum contained zero items during test run R."}, {"id": "A14", "statement": "The washer's spin was balanced during test run R."}, {"id": "A15", "statement": "The observed banging-event count during test run R was zero."}, {"id": "A16", "statement": "Test run R was the only washer spin test completed by any inspector after the resident stated the reuse condition and before readiness decision D."}], "base_state_json": "[{\"speaker\":\"Resident reporting the incident\",\"text\":\"The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging.\"},{\"speaker\":\"Maintenance case note\",\"text\":\"Through readiness decision D, incident E was the washer's only banging incident. It involved exactly one off-center item. No water leaked, odor appeared, or scraping occurred during or after the incident, and inspection found no looseness or damage.\"},{\"speaker\":\"Maintenance case note\",\"text\":\"Person P was the inspector who performed test run R. R was completed after the resident stated the reuse condition and before D, with zero items in the drum and balanced spin. R was the only washer spin test completed by any inspector in that interval.\"},{\"speaker\":\"Verified observation\",\"text\":\"The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S.\"},{\"speaker\":\"Verified observation\",\"text\":\"The observed banging-event count during test run S was two.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "supported"}], "focus_atom": "A15", "focus_evidence": [{"path": ["3", "text"], "text": "The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S."}, {"path": ["4", "text"], "text": "The observed banging-event count during test run S was two."}], "policy_evidence": [{"path": ["0", "text"], "text": "The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging."}], "rules": [{"justification": "The incident satisfies every stated routine load-correction qualification, and inspector P completed an empty, balanced spin without banging after the resident imposed the condition. The resident's specified test was therefore completed successfully.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}]}, {"justification": "Although the incident satisfies every stated routine load-correction qualification, the only inspector-completed spin test in the relevant interval had a nonzero banging count. Thus no competing inspector-completed test could have fulfilled the resident's requirement for a balanced empty spin without banging.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "supported"}]}]}, "verified_pair": {"left": "The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S.", "negative_left": "The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S.", "negative_right": "The observed banging-event count during test run S was four.", "right": "The observed banging-event count during test run S was two."}, "verifier_independent_model": false}, "family": "fast-41-diverse-183-022", "id": "fast-41-diverse-183-022-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one stated condition for the resident’s intended reuse remains unfulfilled, even if the incident otherwise appears routine.", "true": "Yes — the evidence establishes readiness for the resident’s intended reuse, including fulfillment of every stated condition."}, "instructions": "Answer whether the washer is ready for the resident’s intended reuse. Routing options are cleanup, load correction, basic inspection, and deferred maintenance. Urgency is ordered: 1 = routine, 2 = prompt, 3 = urgent, 4 = immediate shutdown. Under the household policy, an isolated banging event with an off-center single item, no leak, odor, scraping, looseness, or damage routes to load correction at urgency 1. However, when the resident expressly conditions reuse on a specified test, readiness is established only after that test is completed successfully. An inspection plan or other favorable evidence does not satisfy the condition.", "type": "noul"}}, "state": [{"speaker": "Resident reporting the incident", "text": "The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging."}, {"speaker": "Maintenance case note", "text": "Through readiness decision D, incident E was the washer's only banging incident. It involved exactly one off-center item. No water leaked, odor appeared, or scraping occurred during or after the incident, and inspection found no looseness or damage."}, {"speaker": "Maintenance case note", "text": "Person P was the inspector who performed test run R. R was completed after the resident stated the reuse condition and before D, with zero items in the drum and balanced spin. R was the only washer spin test completed by any inspector in that interval."}, {"speaker": "Verified observation", "text": "The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S."}, {"speaker": "Verified observation", "text": "The observed banging-event count during test run S was two."}]}, "method": "c2d", "provenance": {"source_id": "diverse-183", "source_is_synthetic": true, "source_sha256": "e3c0610d70e581b628bc4f9c3a5638fa1feb83c55d1a41d1e45a1f8fe0e3e6ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the washer, reuse condition, test-run identities, and time relation without conflicting duplicate facts. The two evidence quotes are complete factual sentences, and changing S from two to four coherently changes R from zero to two. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "refuted", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "supported", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A11": "supported", "A12": "supported", "A13": "supported", "A14": "supported", "A15": "refuted", "A16": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A15": "unknown"}, "remove_right": {"A15": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A15": "unknown"}, "negative_pair": {"A15": "refuted"}, "negative_sentence": {"A15": "unknown"}, "positive_pair": {"A15": "supported"}, "right": {"A15": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a factual relationship; interval-bounded and uniqueness propositions remain atomic despite quantifying over times or tests. A15 is a factual observation about the banging count, not a policy conclusion. The base and counter assignments differ only on A15 and are realizable: a spin may be balanced yet still produce banging. Policy evidence correctly preserves the resident’s state-originating, case-specific reuse condition, while the general governing readiness rule remains automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that inspector P completed the resident’s specified test in the relevant interval: an empty, balanced spin with zero banging. This successfully fulfills the express reuse condition, so readiness follows. The additional incident and uniqueness conditions are unnecessary but do not undermine sufficiency.", "rule_index": 0, "sound": true}, {"reason": "A15 being refuted entails that test R had a nonzero banging count, so it did not satisfy the required no-banging test. A16 excludes any other completed inspector spin test in the relevant interval, preventing a competing successful test. Therefore the resident’s express condition remains unfulfilled and the false target follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "During the period covered by the evidence through readiness decision D, incident E was the washer's only banging incident."}, {"id": "A2", "statement": "The washer contained exactly one item during incident E."}, {"id": "A3", "statement": "The item in the washer was off-center during incident E."}, {"id": "A4", "statement": "No water leaked from the washer during or after incident E through readiness decision D."}, {"id": "A5", "statement": "No odor was detected from the washer during or after incident E through readiness decision D."}, {"id": "A6", "statement": "No scraping was detected from the washer during or after incident E through readiness decision D."}, {"id": "A7", "statement": "Inspection of the washer after incident E detected no looseness."}, {"id": "A8", "statement": "Inspection of the washer after incident E detected no damage."}, {"id": "A9", "statement": "Person P was an inspector at the time of test run R."}, {"id": "A10", "statement": "Person P performed test run R on the washer."}, {"id": "A11", "statement": "Test run R was completed."}, {"id": "A12", "statement": "Test run R occurred after the resident stated the reuse condition and before readiness decision D."}, {"id": "A13", "statement": "The washer drum contained zero items during test run R."}, {"id": "A14", "statement": "The washer's spin was balanced during test run R."}, {"id": "A15", "statement": "The observed banging-event count during test run R was zero."}, {"id": "A16", "statement": "Test run R was the only washer spin test completed by any inspector after the resident stated the reuse condition and before readiness decision D."}], "base_state_json": "[{\"speaker\":\"Resident reporting the incident\",\"text\":\"The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging.\"},{\"speaker\":\"Maintenance case note\",\"text\":\"Through readiness decision D, incident E was the washer's only banging incident. It involved exactly one off-center item. No water leaked, odor appeared, or scraping occurred during or after the incident, and inspection found no looseness or damage.\"},{\"speaker\":\"Maintenance case note\",\"text\":\"Person P was the inspector who performed test run R. R was completed after the resident stated the reuse condition and before D, with zero items in the drum and balanced spin. R was the only washer spin test completed by any inspector in that interval.\"},{\"speaker\":\"Verified observation\",\"text\":\"The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S.\"},{\"speaker\":\"Verified observation\",\"text\":\"The observed banging-event count during test run S was two.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "supported"}], "focus_atom": "A15", "focus_evidence": [{"path": ["3", "text"], "text": "The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S."}, {"path": ["4", "text"], "text": "The observed banging-event count during test run S was two."}], "policy_evidence": [{"path": ["0", "text"], "text": "The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging."}], "rules": [{"justification": "The incident satisfies every stated routine load-correction qualification, and inspector P completed an empty, balanced spin without banging after the resident imposed the condition. The resident's specified test was therefore completed successfully.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "supported"}, {"atom_id": "A16", "state": "supported"}]}, {"justification": "Although the incident satisfies every stated routine load-correction qualification, the only inspector-completed spin test in the relevant interval had a nonzero banging count. Thus no competing inspector-completed test could have fulfilled the resident's requirement for a balanced empty spin without banging.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "supported"}, {"atom_id": "A12", "state": "supported"}, {"atom_id": "A13", "state": "supported"}, {"atom_id": "A14", "state": "supported"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "supported"}]}]}, "verified_pair": {"left": "The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S.", "negative_left": "The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S.", "negative_right": "The observed banging-event count during test run S was four.", "right": "The observed banging-event count during test run S was two."}, "verifier_independent_model": false}, "family": "fast-41-diverse-183-022", "id": "fast-41-diverse-183-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one stated condition for the resident’s intended reuse remains unfulfilled, even if the incident otherwise appears routine.", "true": "Yes — the evidence establishes readiness for the resident’s intended reuse, including fulfillment of every stated condition."}, "instructions": "Answer whether the washer is ready for the resident’s intended reuse. Routing options are cleanup, load correction, basic inspection, and deferred maintenance. Urgency is ordered: 1 = routine, 2 = prompt, 3 = urgent, 4 = immediate shutdown. Under the household policy, an isolated banging event with an off-center single item, no leak, odor, scraping, looseness, or damage routes to load correction at urgency 1. However, when the resident expressly conditions reuse on a specified test, readiness is established only after that test is completed successfully. An inspection plan or other favorable evidence does not satisfy the condition.", "type": "noul"}}, "state": [{"speaker": "Resident reporting the incident", "text": "The washer banged twice during spin with one wet bath mat inside. I stopped it. I will use it again only if an inspector completes a balanced empty spin without banging."}, {"speaker": "Maintenance case note", "text": "Through readiness decision D, incident E was the washer's only banging incident. It involved exactly one off-center item. No water leaked, odor appeared, or scraping occurred during or after the incident, and inspection found no looseness or damage."}, {"speaker": "Maintenance case note", "text": "Person P was the inspector who performed test run R. R was completed after the resident stated the reuse condition and before D, with zero items in the drum and balanced spin. R was the only washer spin test completed by any inspector in that interval."}, {"speaker": "Verified observation", "text": "The observed banging-event count during test run R was two fewer than the observed banging-event count during test run S."}, {"speaker": "Verified observation", "text": "The observed banging-event count during test run S was four."}]}, "method": "c2d", "provenance": {"source_id": "diverse-183", "source_is_synthetic": true, "source_sha256": "e3c0610d70e581b628bc4f9c3a5638fa1feb83c55d1a41d1e45a1f8fe0e3e6ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing criteria and instructions, and both contexts retain the same request and bindings. The base evidence has two complete factual sentences: “During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.” and “During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged.” The counterfactual evidence likewise has two complete factual sentences: “During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.” and “During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 250 milliliters discharged.” The changed measurement is compatible with discharge through the monitored hose and does not assert an appliance leak. Neither context contains an answer, code, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Mara reported a small puddle beside her washer after the spin cycle. The puddle was traced to detergent from an uncapped bottle, while Inez's inspection did not confirm an appliance-origin leak. The washer completed the reported cycle and the follow-up rinse with zero error codes. No electrical symptom, burning symptom, or uncontrolled heat was observed during any of the three relevant events. Inez found no damaged washer component, and no washer symptom other than a possible leak repeated across those events. The laundry nook was monitored during the follow-up rinse, with the drain route instrumented and the surrounding area checked afterward. Cleanup of the detergent spill and ordinary load correction are available before deciding whether the machine may return to normal use.\\n\\nPolicy: Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\",\"evidence\":[\"During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.\",\"During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "0"], "text": "During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose."}, {"path": ["evidence", "1"], "text": "During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.", "negative_left": "During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.", "negative_right": "During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 250 milliliters discharged.", "right": "During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged."}, "verifier_independent_model": false}, "family": "fast-41-diverse-185-016", "id": "fast-41-diverse-185-016-base", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Mara reported a small puddle beside her washer after the spin cycle. The puddle was traced to detergent from an uncapped bottle, while Inez's inspection did not confirm an appliance-origin leak. The washer completed the reported cycle and the follow-up rinse with zero error codes. No electrical symptom, burning symptom, or uncontrolled heat was observed during any of the three relevant events. Inez found no damaged washer component, and no washer symptom other than a possible leak repeated across those events. The laundry nook was monitored during the follow-up rinse, with the drain route instrumented and the surrounding area checked afterward. Cleanup of the detergent spill and ordinary load correction are available before deciding whether the machine may return to normal use.\n\nPolicy: Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.", "evidence": ["During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.", "During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged."]}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing criteria and instructions, and both contexts retain the same request and bindings. The base evidence has two complete factual sentences: “During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.” and “During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged.” The counterfactual evidence likewise has two complete factual sentences: “During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.” and “During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 250 milliliters discharged.” The changed measurement is compatible with discharge through the monitored hose and does not assert an appliance leak. Neither context contains an answer, code, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Mara reported a small puddle beside her washer after the spin cycle. The puddle was traced to detergent from an uncapped bottle, while Inez's inspection did not confirm an appliance-origin leak. The washer completed the reported cycle and the follow-up rinse with zero error codes. No electrical symptom, burning symptom, or uncontrolled heat was observed during any of the three relevant events. Inez found no damaged washer component, and no washer symptom other than a possible leak repeated across those events. The laundry nook was monitored during the follow-up rinse, with the drain route instrumented and the surrounding area checked afterward. Cleanup of the detergent spill and ordinary load correction are available before deciding whether the machine may return to normal use.\\n\\nPolicy: Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\",\"evidence\":[\"During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.\",\"During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "0"], "text": "During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose."}, {"path": ["evidence", "1"], "text": "During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.", "negative_left": "During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.", "negative_right": "During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 250 milliliters discharged.", "right": "During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 0 milliliters discharged."}, "verifier_independent_model": false}, "family": "fast-41-diverse-185-016", "id": "fast-41-diverse-185-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Mara reported a small puddle beside her washer after the spin cycle. The puddle was traced to detergent from an uncapped bottle, while Inez's inspection did not confirm an appliance-origin leak. The washer completed the reported cycle and the follow-up rinse with zero error codes. No electrical symptom, burning symptom, or uncontrolled heat was observed during any of the three relevant events. Inez found no damaged washer component, and no washer symptom other than a possible leak repeated across those events. The laundry nook was monitored during the follow-up rinse, with the drain route instrumented and the surrounding area checked afterward. Cleanup of the detergent spill and ordinary load correction are available before deciding whether the machine may return to normal use.\n\nPolicy: Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.", "evidence": ["During Inez's post-inspection short rinse of Mara's washer, its sole water-discharge path into the laundry nook was the monitored drain hose.", "During that short rinse, the calibrated sensor on Mara's washer's drain hose recorded 250 milliliters discharged."]}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the verbatim policy, request, and bindings; the evidence is factual and complete, including “For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop.” and “During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation.” The counterfactual coherently changes only the sensor observation and does not reveal an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Case note: Mara reported a puddle after the washer's spin cycle. The puddle was beside the machine, and Inez found an uncapped detergent bottle above it; the bottle's detergent was identified as the source. Inez's inspection did not confirm an appliance-origin leak. The cycle log and the post-inspection short rinse together showed zero error codes. No electrical symptom, burning symptom, uncontrolled heat, or damaged washer component was identified during any of those events. No washer symptom other than a possible leak during the short rinse repeated across the reported cycle, the inspection, or the rinse. For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop. During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation.\",\"request\":\"Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["context"], "text": "For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop."}, {"path": ["context"], "text": "During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop.", "negative_left": "For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop.", "negative_right": "During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded one activation.", "right": "During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation."}, "verifier_independent_model": false}, "family": "fast-41-diverse-185-021", "id": "fast-41-diverse-185-021-base", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Case note: Mara reported a puddle after the washer's spin cycle. The puddle was beside the machine, and Inez found an uncapped detergent bottle above it; the bottle's detergent was identified as the source. Inez's inspection did not confirm an appliance-origin leak. The cycle log and the post-inspection short rinse together showed zero error codes. No electrical symptom, burning symptom, uncontrolled heat, or damaged washer component was identified during any of those events. No washer symptom other than a possible leak during the short rinse repeated across the reported cycle, the inspection, or the rinse. For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop. During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation.", "request": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts preserve the verbatim policy, request, and bindings; the evidence is factual and complete, including “For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop.” and “During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation.” The counterfactual coherently changes only the sensor observation and does not reveal an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; explicitly scoped zero-error and repetition claims remain atomic. The focus a3 is factual. The base and counter assignments differ only on a3 and are jointly realizable: an external detergent spill can explain the original puddle while a later rinse either does or does not release washer-origin water. Policy evidence properly cites only the original state; the retained questions already preserve the disposition criteria and priorities.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a3 entails that washer-origin water escaped during the repeat rinse. That repeat operation therefore confirms an appliance leak, directly satisfying level 2.", "rule_index": 0, "sound": true}, {"reason": "The conjunction identifies the puddle as an external detergent spill, excludes a confirmed appliance leak and recurrence on the repeat rinse, records no errors, and excludes the specified repeated symptoms and high-urgency hazards. These conditions are sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The small puddle beside Mara's washer after the reported spin cycle was caused by detergent from the uncapped bottle."}, {"id": "a2", "statement": "Inez's inspection after the reported cycle confirmed an appliance-origin leak from Mara's washer."}, {"id": "a3", "statement": "During Inez's post-inspection short rinse of Mara's washer, no water originating from the washer escaped into the laundry nook."}, {"id": "a4", "statement": "Across the reported cycle and Inez's post-inspection short rinse, Mara's washer recorded zero error codes."}, {"id": "a5", "statement": "An electrical symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a6", "statement": "A burning symptom occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a7", "statement": "Uncontrolled heat occurred during the reported cycle, Inez's inspection, or the post-inspection short rinse."}, {"id": "a8", "statement": "Inez's inspection identified a damaged component in Mara's washer."}, {"id": "a9", "statement": "A washer symptom other than a possible leak during the post-inspection short rinse repeated during the reported cycle, Inez's inspection, or the post-inspection short rinse."}], "base_state_json": "{\"context\":\"Case note: Mara reported a puddle after the washer's spin cycle. The puddle was beside the machine, and Inez found an uncapped detergent bottle above it; the bottle's detergent was identified as the source. Inez's inspection did not confirm an appliance-origin leak. The cycle log and the post-inspection short rinse together showed zero error codes. No electrical symptom, burning symptom, uncontrolled heat, or damaged washer component was identified during any of those events. No washer symptom other than a possible leak during the short rinse repeated across the reported cycle, the inspection, or the rinse. For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop. During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation.\",\"request\":\"Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["context"], "text": "For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop."}, {"path": ["context"], "text": "During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation."}], "policy_evidence": [{"path": ["request"], "text": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}], "rules": [{"justification": "If water originating from the washer escaped during Inez's repeat operation, that repeat operation confirms an appliance leak and independently satisfies the high-urgency criterion.", "target": "2", "when": [{"atom_id": "a3", "state": "refuted"}]}, {"justification": "The puddle is identified as an external detergent spill, inspection does not confirm an appliance leak, the repeat rinse has no washer-origin leak, no errors or other symptoms repeat, and no high-urgency hazard is present; therefore cleanup followed by normal use satisfies the low-urgency criterion.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop.", "negative_left": "For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop.", "negative_right": "During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded one activation.", "right": "During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded no activation."}, "verifier_independent_model": false}, "family": "fast-41-diverse-185-021", "id": "fast-41-diverse-185-021-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low urgency: Route to cleanup or load correction when evidence identifies an external spill or loading issue and appliance checks show no recurring fault. Clean or correct the load, then normal use is allowed.", "1 — Moderate urgency: Route to basic inspection when the source remains ambiguous, a symptom repeats, or logs and physical checks conflict without confirming a hazardous defect. Pause use until hoses, seals, venting, leveling, or load balance are checked.", "2 — High urgency: Route to deferred maintenance and remove the appliance from use when inspection or repeat operation confirms an appliance leak, electrical or burning symptom, uncontrolled heat, damaged component, or repeated error. Arrange repair before reuse."], "instructions": "Choose exactly one ordered disposition level. Apply the higher level only when its stated conditions are met; an unsupported resident claim alone does not override direct inspection and a successful repeat check.", "type": "score"}}, "state": {"context": "Case note: Mara reported a puddle after the washer's spin cycle. The puddle was beside the machine, and Inez found an uncapped detergent bottle above it; the bottle's detergent was identified as the source. Inez's inspection did not confirm an appliance-origin leak. The cycle log and the post-inspection short rinse together showed zero error codes. No electrical symptom, burning symptom, uncontrolled heat, or damaged washer component was identified during any of those events. No washer symptom other than a possible leak during the short rinse repeated across the reported cycle, the inspection, or the rinse. For Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor would activate for every drop of water originating from Mara's washer that entered the laundry nook, and it would activate only for such a drop. During Inez's post-inspection short rinse of Mara's washer, the calibrated threshold sensor at the laundry nook's entrance recorded one activation.", "request": "Select the disposition level using corroborated photos, logs, and inspection findings; explain how the conflicting evidence affects reuse."}}, "method": "c2d", "provenance": {"source_id": "diverse-185", "source_is_synthetic": true, "source_sha256": "9dd3afa8cfbbdcf9ca20cade2645de2579a46dcdb657502583d2291603ae01d3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, bindings, and criteria; both contexts are coherent, and the two quoted evidence spans are complete factual sentences without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or incident-assessment relationship rather than bundling unrelated requirements or encoding the final response level. A4 is a factual recurrence proposition, not a policy conclusion. The base and counter assignments can differ only in recurrence: two completed leakage occasions can occur without active leakage at decision time, while the counter can describe only one occasion; the remaining facts can stay fixed. Empty policy_evidence is correct because the governing rubric, evidence requirements, priorities, and exceptions all originate in the retained questions object, while the state contains only case observations that need not be preserved for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A4 establishes leakage on at least two distinct occasions, which is recurring leakage. Recurring leakage independently requires level 2, regardless of whether leakage is active at decision time.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a minor, contained incident; missing required rear-connection photographic evidence; and explicit absence of every enumerated level-2 indicator. Missing required evidence excludes level 0 and directs the case to level 1.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The reported washer incident was minor."}, {"id": "A2", "statement": "The reported washer incident was contained."}, {"id": "A3", "statement": "At the response-level decision time for the reported washer incident, the required incident-area photos lacked a view of the washer's rear connections."}, {"id": "A4", "statement": "From the start of the reported washer incident through the response-level decision time, the washer produced water leakage on at least two distinct occasions."}, {"id": "A5", "statement": "At the response-level decision time for the reported incident, the washer was actively leaking."}, {"id": "A6", "statement": "During the reported washer incident, water was near the electrical outlet serving the washer."}, {"id": "A7", "statement": "During the reported washer incident, the washer emitted smoke."}, {"id": "A8", "statement": "During the reported washer incident, the washer produced sparks."}, {"id": "A9", "statement": "During the reported washer incident, the washer emitted an electrical odor."}, {"id": "A10", "statement": "The washer's logs contained repeated fault entries associated with the reported incident."}, {"id": "A11", "statement": "The washer failed the basic inspection required before reuse after the reported incident."}], "base_state_json": "[{\"speaker\":\"Maintenance record\",\"text\":\"On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations. The maintenance log records leakage observations at 08:25 and 09:10 on 17 September 2026.\"},{\"speaker\":\"Response coordinator\",\"text\":\"The incident was minor and contained. At the decision time, the washer was not actively leaking, and no water had been near the electrical outlet. There was no smoke, sparking, or electrical odor during the incident. The required incident-area photos lacked a view of the washer's rear connections. The cycle log showed an unbalanced-load warning after one heavy blanket, but there were no repeated fault entries.\"},{\"speaker\":\"Inspector\",\"text\":\"The washer passed the basic inspection required before reuse. The available records describe a small puddle that was cleaned up, with no continuing hazard observed while the appliance was off. The cycle log and symptom notes were available, but the rear-connection view was not.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["0", "text"], "text": "On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations."}, {"path": ["0", "text"], "text": "The maintenance log records leakage observations at 08:25 and 09:10 on 17 September 2026."}], "policy_evidence": [], "rules": [{"justification": "At least two distinct leakage occasions during the incident-to-decision interval establish recurring leakage, which independently requires the high-response route.", "target": "2", "when": [{"atom_id": "A4", "state": "supported"}]}, {"justification": "The incident is minor and contained, required rear-connection photographic evidence is missing, and every enumerated level-2 indicator is explicitly absent, so the moderate-response route applies.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations.", "negative_left": "On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations.", "negative_right": "The maintenance log records one leakage observation at 08:25 on 17 September 2026.", "right": "The maintenance log records leakage observations at 08:25 and 09:10 on 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-186-013", "id": "fast-41-diverse-186-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Low response: Cleanup or load correction; reuse is allowed now only when all required evidence is present, shows no ongoing leak or danger indicator, and supports a one-time spill or load-balance cause.", "1 — Moderate response: Basic inspection before reuse; apply when the incident is minor and contained with no level-2 indicator, but required evidence is missing or a simple load-related cause remains unconfirmed. Complete the inspection the same day before another cycle.", "2 — High response: Stop use and arrange urgent maintenance; apply for active or recurring leakage, water near an outlet, smoke, sparks, electrical odor, repeated fault logs, or failure of the basic inspection."], "instructions": "Assign one response-intensity level using the ordered rubric. Required evidence for immediate reuse is: incident-area photos including rear connections, the cycle log, and symptom notes covering recurrence after unloading. Missing required evidence prevents level 0. Determine the route, whether the washer may be used again, and the urgency.", "type": "score"}}, "state": [{"speaker": "Maintenance record", "text": "On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations. The maintenance log records leakage observations at 08:25 and 09:10 on 17 September 2026."}, {"speaker": "Response coordinator", "text": "The incident was minor and contained. At the decision time, the washer was not actively leaking, and no water had been near the electrical outlet. There was no smoke, sparking, or electrical odor during the incident. The required incident-area photos lacked a view of the washer's rear connections. The cycle log showed an unbalanced-load warning after one heavy blanket, but there were no repeated fault entries."}, {"speaker": "Inspector", "text": "The washer passed the basic inspection required before reuse. The available records describe a small puddle that was cleaned up, with no continuing hazard observed while the appliance was off. The cycle log and symptom notes were available, but the rear-connection view was not."}]}, "method": "c2d", "provenance": {"source_id": "diverse-186", "source_is_synthetic": true, "source_sha256": "bedc0c805625693d340a70785d59687aa7dc865190604e3fc159604b114d855a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, bindings, and criteria; both contexts are coherent, and the two quoted evidence spans are complete factual sentences without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or incident-assessment relationship rather than bundling unrelated requirements or encoding the final response level. A4 is a factual recurrence proposition, not a policy conclusion. The base and counter assignments can differ only in recurrence: two completed leakage occasions can occur without active leakage at decision time, while the counter can describe only one occasion; the remaining facts can stay fixed. Empty policy_evidence is correct because the governing rubric, evidence requirements, priorities, and exceptions all originate in the retained questions object, while the state contains only case observations that need not be preserved for synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A4 establishes leakage on at least two distinct occasions, which is recurring leakage. Recurring leakage independently requires level 2, regardless of whether leakage is active at decision time.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a minor, contained incident; missing required rear-connection photographic evidence; and explicit absence of every enumerated level-2 indicator. Missing required evidence excludes level 0 and directs the case to level 1.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The reported washer incident was minor."}, {"id": "A2", "statement": "The reported washer incident was contained."}, {"id": "A3", "statement": "At the response-level decision time for the reported washer incident, the required incident-area photos lacked a view of the washer's rear connections."}, {"id": "A4", "statement": "From the start of the reported washer incident through the response-level decision time, the washer produced water leakage on at least two distinct occasions."}, {"id": "A5", "statement": "At the response-level decision time for the reported incident, the washer was actively leaking."}, {"id": "A6", "statement": "During the reported washer incident, water was near the electrical outlet serving the washer."}, {"id": "A7", "statement": "During the reported washer incident, the washer emitted smoke."}, {"id": "A8", "statement": "During the reported washer incident, the washer produced sparks."}, {"id": "A9", "statement": "During the reported washer incident, the washer emitted an electrical odor."}, {"id": "A10", "statement": "The washer's logs contained repeated fault entries associated with the reported incident."}, {"id": "A11", "statement": "The washer failed the basic inspection required before reuse after the reported incident."}], "base_state_json": "[{\"speaker\":\"Maintenance record\",\"text\":\"On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations. The maintenance log records leakage observations at 08:25 and 09:10 on 17 September 2026.\"},{\"speaker\":\"Response coordinator\",\"text\":\"The incident was minor and contained. At the decision time, the washer was not actively leaking, and no water had been near the electrical outlet. There was no smoke, sparking, or electrical odor during the incident. The required incident-area photos lacked a view of the washer's rear connections. The cycle log showed an unbalanced-load warning after one heavy blanket, but there were no repeated fault entries.\"},{\"speaker\":\"Inspector\",\"text\":\"The washer passed the basic inspection required before reuse. The available records describe a small puddle that was cleaned up, with no continuing hazard observed while the appliance was off. The cycle log and symptom notes were available, but the rear-connection view was not.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["0", "text"], "text": "On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations."}, {"path": ["0", "text"], "text": "The maintenance log records leakage observations at 08:25 and 09:10 on 17 September 2026."}], "policy_evidence": [], "rules": [{"justification": "At least two distinct leakage occasions during the incident-to-decision interval establish recurring leakage, which independently requires the high-response route.", "target": "2", "when": [{"atom_id": "A4", "state": "supported"}]}, {"justification": "The incident is minor and contained, required rear-connection photographic evidence is missing, and every enumerated level-2 indicator is explicitly absent, so the moderate-response route applies.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations.", "negative_left": "On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations.", "negative_right": "The maintenance log records one leakage observation at 08:25 on 17 September 2026.", "right": "The maintenance log records leakage observations at 08:25 and 09:10 on 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-186-013", "id": "fast-41-diverse-186-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Low response: Cleanup or load correction; reuse is allowed now only when all required evidence is present, shows no ongoing leak or danger indicator, and supports a one-time spill or load-balance cause.", "1 — Moderate response: Basic inspection before reuse; apply when the incident is minor and contained with no level-2 indicator, but required evidence is missing or a simple load-related cause remains unconfirmed. Complete the inspection the same day before another cycle.", "2 — High response: Stop use and arrange urgent maintenance; apply for active or recurring leakage, water near an outlet, smoke, sparks, electrical odor, repeated fault logs, or failure of the basic inspection."], "instructions": "Assign one response-intensity level using the ordered rubric. Required evidence for immediate reuse is: incident-area photos including rear connections, the cycle log, and symptom notes covering recurrence after unloading. Missing required evidence prevents level 0. Determine the route, whether the washer may be used again, and the urgency.", "type": "score"}}, "state": [{"speaker": "Maintenance record", "text": "On 17 September 2026, the reported washer incident ran from 08:00 until the 10:00 response-level decision, and its maintenance log is exhaustive for leakage observations. The maintenance log records one leakage observation at 08:25 on 17 September 2026."}, {"speaker": "Response coordinator", "text": "The incident was minor and contained. At the decision time, the washer was not actively leaking, and no water had been near the electrical outlet. There was no smoke, sparking, or electrical odor during the incident. The required incident-area photos lacked a view of the washer's rear connections. The cycle log showed an unbalanced-load warning after one heavy blanket, but there were no repeated fault entries."}, {"speaker": "Inspector", "text": "The washer passed the basic inspection required before reuse. The available records describe a small puddle that was cleaned up, with no continuing hazard observed while the appliance was off. The cycle log and symptom notes were available, but the rear-connection view was not."}]}, "method": "c2d", "provenance": {"source_id": "diverse-186", "source_is_synthetic": true, "source_sha256": "bedc0c805625693d340a70785d59687aa7dc865190604e3fc159604b114d855a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings, use two complete factual evidence sentences, and the counterfactual consistently changes only the serving time to make the timed step incomplete.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\\n\\nA prep checklist confirms that the rice is fully cooked and the tofu is browned on every side listed by the recipe. The ingredient inventory is complete, all food in the batch is uncontaminated, and the four bowls are portioned. Each portion designated for refrigerated storage is labeled. A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time. The same log records the serving time for the four teriyaki tofu bowls as 18:44:25 on 2026-09-17.\",\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["context"], "text": "A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time."}, {"path": ["context"], "text": "The same log records the serving time for the four teriyaki tofu bowls as 18:44:25 on 2026-09-17."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time.", "negative_left": "A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time.", "negative_right": "The same log records the serving time for the four teriyaki tofu bowls as 18:43:25 on 2026-09-17.", "right": "The same log records the serving time for the four teriyaki tofu bowls as 18:44:25 on 2026-09-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-187-001", "id": "fast-41-diverse-187-001-base", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\n\nA prep checklist confirms that the rice is fully cooked and the tofu is browned on every side listed by the recipe. The ingredient inventory is complete, all food in the batch is uncontaminated, and the four bowls are portioned. Each portion designated for refrigerated storage is labeled. A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time. The same log records the serving time for the four teriyaki tofu bowls as 18:44:25 on 2026-09-17.", "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_5_ready"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and bindings, use two complete factual evidence sentences, and the counterfactual consistently changes only the serving time to make the timed step incomplete.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\\n\\nA prep checklist confirms that the rice is fully cooked and the tofu is browned on every side listed by the recipe. The ingredient inventory is complete, all food in the batch is uncontaminated, and the four bowls are portioned. Each portion designated for refrigerated storage is labeled. A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time. The same log records the serving time for the four teriyaki tofu bowls as 18:44:25 on 2026-09-17.\",\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["context"], "text": "A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time."}, {"path": ["context"], "text": "The same log records the serving time for the four teriyaki tofu bowls as 18:44:25 on 2026-09-17."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time.", "negative_left": "A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time.", "negative_right": "The same log records the serving time for the four teriyaki tofu bowls as 18:43:25 on 2026-09-17.", "right": "The same log records the serving time for the four teriyaki tofu bowls as 18:44:25 on 2026-09-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-187-001", "id": "fast-41-diverse-187-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\n\nA prep checklist confirms that the rice is fully cooked and the tofu is browned on every side listed by the recipe. The ingredient inventory is complete, all food in the batch is uncontaminated, and the four bowls are portioned. Each portion designated for refrigerated storage is labeled. A calibrated kitchen log records that the glaze for the four teriyaki tofu bowls began bubbling continuously at 18:42:10 on 2026-09-17 and remained continuously bubbling until the logged serving time. The same log records the serving time for the four teriyaki tofu bowls as 18:43:25 on 2026-09-17.", "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_2_major_correction"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and the unchanged questions object preserves all criteria and instructions. The recipe, four-bowl entity, request, and time bindings remain unchanged. The evidence consists of two complete factual sentences: \"For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026.\" and \"For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00.\" The counterfactual’s changed ending time is coherent with its unchanged beginning time and introduces no duplicate contradiction. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions. The batch record confirms that the rice is fully cooked and the tofu is browned on every recipe-listed side. All required ingredients are present, and every food item in the batch is uncontaminated. The four bowls are portioned, and every portion designated for refrigerated storage is labeled. For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026. For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00.\",\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["context"], "text": "For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026."}, {"path": ["context"], "text": "For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026.", "negative_left": "For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026.", "negative_right": "For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:43:25 on 12 June 2026, before serving at 18:45:00.", "right": "For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00."}, "verifier_independent_model": false}, "family": "fast-41-diverse-187-006", "id": "fast-41-diverse-187-006-base", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions. The batch record confirms that the rice is fully cooked and the tofu is browned on every recipe-listed side. All required ingredients are present, and every food item in the batch is uncontaminated. The four bowls are portioned, and every portion designated for refrigerated storage is labeled. For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026. For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00.", "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_5_ready"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and the unchanged questions object preserves all criteria and instructions. The recipe, four-bowl entity, request, and time bindings remain unchanged. The evidence consists of two complete factual sentences: \"For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026.\" and \"For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00.\" The counterfactual’s changed ending time is coherent with its unchanged beginning time and introduces no duplicate contradiction. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions. The batch record confirms that the rice is fully cooked and the tofu is browned on every recipe-listed side. All required ingredients are present, and every food item in the batch is uncontaminated. The four bowls are portioned, and every portion designated for refrigerated storage is labeled. For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026. For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00.\",\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["context"], "text": "For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026."}, {"path": ["context"], "text": "For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026.", "negative_left": "For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026.", "negative_right": "For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:43:25 on 12 June 2026, before serving at 18:45:00.", "right": "For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:44:25 on 12 June 2026, before serving at 18:45:00."}, "verifier_independent_model": false}, "family": "fast-41-diverse-187-006", "id": "fast-41-diverse-187-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions. The batch record confirms that the rice is fully cooked and the tofu is browned on every recipe-listed side. All required ingredients are present, and every food item in the batch is uncontaminated. The four bowls are portioned, and every portion designated for refrigerated storage is labeled. For the four teriyaki tofu bowls, continuous glaze bubbling began at 18:42:10 on 12 June 2026. For the four teriyaki tofu bowls, continuous glaze bubbling ended at 18:43:25 on 12 June 2026, before serving at 18:45:00.", "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_2_major_correction"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy, request, recipe, roles, and bindings. The focus evidence contains two complete factual sentences: \"Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.\" and \"The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds.\" in the base context, with only the recorded duration changed to 75 seconds in the counterfactual. The counterfactual is coherent because the timer measurement changes without duplicating or contradicting any unchanged assertion. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.\",\"The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds.\",\"The rice for all four bowls is fully cooked, and the tofu is browned on every recipe-listed side.\",\"All ingredients required for the batch are present, and every food item is uncontaminated.\",\"The four bowls are portioned, and every portion designated for refrigerated storage is labeled.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "0"], "text": "Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17."}, {"path": ["evidence", "1"], "text": "The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.", "negative_left": "Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.", "negative_right": "The synchronized kitchen timer assigned to observation B-17 recorded 75 seconds.", "right": "The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds."}, "verifier_independent_model": false}, "family": "fast-41-diverse-187-011", "id": "fast-41-diverse-187-011-base", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.", "The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds.", "The rice for all four bowls is fully cooked, and the tofu is browned on every recipe-listed side.", "All ingredients required for the batch are present, and every food item is uncontaminated.", "The four bowls are portioned, and every portion designated for refrigerated storage is labeled."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_5_ready"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy, request, recipe, roles, and bindings. The focus evidence contains two complete factual sentences: \"Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.\" and \"The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds.\" in the base context, with only the recorded duration changed to 75 seconds in the counterfactual. The counterfactual is coherent because the timer measurement changes without duplicating or contradicting any unchanged assertion. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.\",\"The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds.\",\"The rice for all four bowls is fully cooked, and the tofu is browned on every recipe-listed side.\",\"All ingredients required for the batch are present, and every food item is uncontaminated.\",\"The four bowls are portioned, and every portion designated for refrigerated storage is labeled.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "0"], "text": "Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17."}, {"path": ["evidence", "1"], "text": "The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.", "negative_left": "Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.", "negative_right": "The synchronized kitchen timer assigned to observation B-17 recorded 75 seconds.", "right": "The synchronized kitchen timer assigned to observation B-17 recorded 135 seconds."}, "verifier_independent_model": false}, "family": "fast-41-diverse-187-011", "id": "fast-41-diverse-187-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Observation B-17 was the glaze's continuous bubbling interval before serving on 2026-09-17.", "The synchronized kitchen timer assigned to observation B-17 recorded 75 seconds.", "The rice for all four bowls is fully cooked, and the tofu is browned on every recipe-listed side.", "All ingredients required for the batch are present, and every food item is uncontaminated.", "The four bowls are portioned, and every portion designated for refrigerated storage is labeled."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_2_major_correction"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, bindings, criteria, and thresholds. The two retained evidence quotes are complete factual sentences. The counterfactual changes only C3’s refrigerator status and introduces no contradictory duplicate assertion. Neither context contains a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Recipe card\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Kitchen log\",\"text\":\"At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C. Exactly two dinner portions of rice had been prepared, and each weighed exactly 350 g.\"},{\"speaker\":\"Kitchen log\",\"text\":\"The leftover rice had a temperature of 60°C when its cooling began and reached 21°C or lower within two hours.\"},{\"speaker\":\"Kitchen log\",\"text\":\"At 8:00 p.m., the leftover rice had a temperature of 21°C or lower.\"},{\"speaker\":\"Kitchen log\",\"text\":\"All of the leftover rice had been transferred from the cooling bowl into containers. Every container holding it was shallow and labeled.\"},{\"speaker\":\"Container log\",\"text\":\"By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice.\"},{\"speaker\":\"Container log\",\"text\":\"At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R.\"},{\"speaker\":\"Supervisor\",\"text\":\"The temperature readings, portion measurements, transfer, and container descriptions were recorded at the workflow evaluation time.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["6", "text"], "text": "By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice."}, {"path": ["7", "text"], "text": "At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice.", "negative_left": "By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice.", "negative_right": "At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R except C3.", "right": "At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R."}, "verifier_independent_model": false}, "family": "fast-41-diverse-188-004", "id": "fast-41-diverse-188-004-base", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Recipe card", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Kitchen log", "text": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C. Exactly two dinner portions of rice had been prepared, and each weighed exactly 350 g."}, {"speaker": "Kitchen log", "text": "The leftover rice had a temperature of 60°C when its cooling began and reached 21°C or lower within two hours."}, {"speaker": "Kitchen log", "text": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"speaker": "Kitchen log", "text": "All of the leftover rice had been transferred from the cooling bowl into containers. Every container holding it was shallow and labeled."}, {"speaker": "Container log", "text": "By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice."}, {"speaker": "Container log", "text": "At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R."}, {"speaker": "Supervisor", "text": "The temperature readings, portion measurements, transfer, and container descriptions were recorded at the workflow evaluation time."}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_3_COMPLETE"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, bindings, criteria, and thresholds. The two retained evidence quotes are complete factual sentences. The counterfactual changes only C3’s refrigerator status and introduces no contradictory duplicate assertion. Neither context contains a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Recipe card\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Kitchen log\",\"text\":\"At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C. Exactly two dinner portions of rice had been prepared, and each weighed exactly 350 g.\"},{\"speaker\":\"Kitchen log\",\"text\":\"The leftover rice had a temperature of 60°C when its cooling began and reached 21°C or lower within two hours.\"},{\"speaker\":\"Kitchen log\",\"text\":\"At 8:00 p.m., the leftover rice had a temperature of 21°C or lower.\"},{\"speaker\":\"Kitchen log\",\"text\":\"All of the leftover rice had been transferred from the cooling bowl into containers. Every container holding it was shallow and labeled.\"},{\"speaker\":\"Container log\",\"text\":\"By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice.\"},{\"speaker\":\"Container log\",\"text\":\"At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R.\"},{\"speaker\":\"Supervisor\",\"text\":\"The temperature readings, portion measurements, transfer, and container descriptions were recorded at the workflow evaluation time.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["6", "text"], "text": "By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice."}, {"path": ["7", "text"], "text": "At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice.", "negative_left": "By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice.", "negative_right": "At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R except C3.", "right": "At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R."}, "verifier_independent_model": false}, "family": "fast-41-diverse-188-004", "id": "fast-41-diverse-188-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Recipe card", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Kitchen log", "text": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C. Exactly two dinner portions of rice had been prepared, and each weighed exactly 350 g."}, {"speaker": "Kitchen log", "text": "The leftover rice had a temperature of 60°C when its cooling began and reached 21°C or lower within two hours."}, {"speaker": "Kitchen log", "text": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"speaker": "Kitchen log", "text": "All of the leftover rice had been transferred from the cooling bowl into containers. Every container holding it was shallow and labeled."}, {"speaker": "Container log", "text": "By the workflow evaluation time, exactly three containers—C1, C2, and C3—held all the leftover rice."}, {"speaker": "Container log", "text": "At the workflow evaluation time, containers C1, C2, and C3 were in refrigerator R except C3."}, {"speaker": "Supervisor", "text": "The temperature readings, portion measurements, transfer, and container descriptions were recorded at the workflow evaluation time."}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_2_READY_CORRECTION_DUE"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions, and both contexts retain the same policy. Entity, path, and time bindings remain unchanged. The two evidence spans are complete factual sentences: “At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice.” and “At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator.” The counterfactual coherently changes refrigeration status to outside without contradictory duplicate assertions. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Home cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Prep log\",\"text\":\"At the workflow evaluation time, the chicken was 75°C. Exactly two dinner portions of rice had been prepared, each weighing 350 g. The leftover rice began cooling at 60°C, reached 20°C within two hours, and measured 20°C at 8:00 p.m.\"},{\"speaker\":\"Storage log\",\"text\":\"By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Every container holding the leftover rice was shallow and labeled.\"},{\"speaker\":\"Container audit\",\"text\":\"At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice.\"},{\"speaker\":\"Refrigeration audit\",\"text\":\"At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["4", "text"], "text": "At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice."}, {"path": ["5", "text"], "text": "At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice.", "negative_left": "At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice.", "negative_right": "At the workflow evaluation time, containers C-17 and C-23 were outside the refrigerator.", "right": "At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-188-005", "id": "fast-41-diverse-188-005-base", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Meal planner", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Home cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Prep log", "text": "At the workflow evaluation time, the chicken was 75°C. Exactly two dinner portions of rice had been prepared, each weighing 350 g. The leftover rice began cooling at 60°C, reached 20°C within two hours, and measured 20°C at 8:00 p.m."}, {"speaker": "Storage log", "text": "By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Every container holding the leftover rice was shallow and labeled."}, {"speaker": "Container audit", "text": "At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice."}, {"speaker": "Refrigeration audit", "text": "At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator."}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_3_COMPLETE"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions, and both contexts retain the same policy. Entity, path, and time bindings remain unchanged. The two evidence spans are complete factual sentences: “At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice.” and “At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator.” The counterfactual coherently changes refrigeration status to outside without contradictory duplicate assertions. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified container atoms concern one property over an explicit set rather than bundling unrelated requirements. The focus atom is the factual refrigeration status of the leftover-rice containers. The base and counter assignments are jointly realizable with only that status changing: all rice can be in shallow labeled containers, with every container refrigerated in the base and at least one not refrigerated in the counter. Policy evidence correctly cites the substantive recipe, cooling, threshold, and conditional-storage rules originating in the original state; criteria and ordering from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliant chicken temperature, exactly two correctly weighted portions, compliant cooling from 60°C to at most 21°C within two hours, satisfaction of the inclusive 8:00 condition, and completion of transfer, shallow-container, labeling, and refrigeration requirements. It excludes the stated discard and pending-action outcomes, so Level 3 follows.", "rule_index": 0, "sound": true}, {"reason": "The cooking, portioning, cooling, and inclusive 8:00 trigger conditions are satisfied, excluding a safety-limit breach or discard condition. Transfer into shallow labeled containers is complete, while refutation of the universal refrigeration atom entails that at least one container holding leftover rice is not refrigerated. Thus a triggered storage action remains pending and Level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the workflow evaluation time, the chicken for the bowl dinner had a temperature of at least 74°C."}, {"id": "a2", "statement": "At the workflow evaluation time, exactly two dinner portions of rice had been prepared."}, {"id": "a3", "statement": "At the workflow evaluation time, each of the two dinner portions of rice weighed exactly 350 g."}, {"id": "a4", "statement": "The leftover rice had a temperature of 60°C when its cooling began."}, {"id": "a5", "statement": "The leftover rice reached a temperature of 21°C or lower within two hours after its cooling began."}, {"id": "a6", "statement": "At 8:00 p.m., the leftover rice had a temperature of 21°C or lower."}, {"id": "a7", "statement": "By the workflow evaluation time, all of the leftover rice had been transferred from the cooling bowl into containers."}, {"id": "a8", "statement": "Every container holding the leftover rice at the workflow evaluation time was shallow."}, {"id": "a9", "statement": "Every container holding the leftover rice at the workflow evaluation time was labeled."}, {"id": "a10", "statement": "Every container holding the leftover rice at the workflow evaluation time was in the refrigerator."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive.\"},{\"speaker\":\"Home cook\",\"text\":\"Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it.\"},{\"speaker\":\"Prep log\",\"text\":\"At the workflow evaluation time, the chicken was 75°C. Exactly two dinner portions of rice had been prepared, each weighing 350 g. The leftover rice began cooling at 60°C, reached 20°C within two hours, and measured 20°C at 8:00 p.m.\"},{\"speaker\":\"Storage log\",\"text\":\"By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Every container holding the leftover rice was shallow and labeled.\"},{\"speaker\":\"Container audit\",\"text\":\"At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice.\"},{\"speaker\":\"Refrigeration audit\",\"text\":\"At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}], "focus_atom": "a10", "focus_evidence": [{"path": ["4", "text"], "text": "At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice."}, {"path": ["5", "text"], "text": "At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator."}], "policy_evidence": [{"path": ["0", "text"], "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"path": ["1", "text"], "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}], "rules": [{"justification": "The inclusive cooking and cooling limits, the required count and weight of dinner portions, and every triggered transfer, container, labeling, and refrigeration step have been completed.", "target": "LEVEL_3_COMPLETE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Cooking, portioning, and cooling comply, and the inclusive 8:00 p.m. condition triggers refrigerated storage. The rice is already in shallow labeled containers, but at least one container holding it is not refrigerated, so a required storage action remains pending without a safety-limit breach.", "target": "LEVEL_2_READY_CORRECTION_DUE", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}]}]}, "verified_pair": {"left": "At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice.", "negative_left": "At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice.", "negative_right": "At the workflow evaluation time, containers C-17 and C-23 were outside the refrigerator.", "right": "At the workflow evaluation time, containers C-17 and C-23 were inside the refrigerator."}, "verifier_independent_model": false}, "family": "fast-41-diverse-188-005", "id": "fast-41-diverse-188-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"LEVEL_1_FAILED_DISCARD": "A cooking or cooling limit was missed, or the stated condition requires discard; do not serve or store the affected food.", "LEVEL_2_READY_CORRECTION_DUE": "Cooking and portioning comply, and no safety limit was breached, but a triggered corrective or storage action is still pending; transfer the rice now into shallow labeled containers and refrigerate.", "LEVEL_3_COMPLETE": "All recipe, portioning, and triggered storage steps have already been completed; serve the dinner portions and leave the leftovers refrigerated."}, "instructions": "Choose the single overall workflow level. The levels are ordered from Level 3 (complete) to Level 1 (failed). Apply the cook’s conditional storage intent at the inclusive temperature boundary.", "type": "choice"}}, "state": [{"speaker": "Meal planner", "text": "The bowl recipe requires chicken at 74°C or higher, rice divided into two 350 g dinner portions, and leftover rice cooled from 60°C to 21°C or lower within two hours. Thresholds are inclusive."}, {"speaker": "Home cook", "text": "Cooling began at 6:00 p.m. If the leftover rice is 21°C or lower at 8:00, I intend it to go immediately into shallow labeled containers in the refrigerator; if it is above 21°C, discard it."}, {"speaker": "Prep log", "text": "At the workflow evaluation time, the chicken was 75°C. Exactly two dinner portions of rice had been prepared, each weighing 350 g. The leftover rice began cooling at 60°C, reached 20°C within two hours, and measured 20°C at 8:00 p.m."}, {"speaker": "Storage log", "text": "By the workflow evaluation time, all leftover rice had been transferred from the cooling bowl into containers. Every container holding the leftover rice was shallow and labeled."}, {"speaker": "Container audit", "text": "At the workflow evaluation time, all leftover rice was in containers C-17 and C-23, and no other container held leftover rice."}, {"speaker": "Refrigeration audit", "text": "At the workflow evaluation time, containers C-17 and C-23 were outside the refrigerator."}]}, "method": "c2d", "provenance": {"source_id": "diverse-188", "source_is_synthetic": true, "source_sha256": "42ed89bfd97cc8f6703794160c0ef61303e58af68d3688a9812ca619765bea8b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "LEVEL_2_READY_CORRECTION_DUE"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, scope, and criteria for both contexts. Both contexts retain Mara, the meal, the four portions, the two serving portions, and the two storage portions. The required base evidence quotes are preserved exactly: “For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.” and “For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time.” All evidence entries are complete factual sentences. The counterfactual's 20:25 cooling time changes an observation without creating contradictory duplicate assertions. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"Home cook Mara prepared paprika chicken and rice. She finished simmering the sauce before starting the chicken. A probe used for two chicken readings passed its ice-water calibration that day and produced readings of 75°C and 74°C. Sol observed slight pinkness in one cut piece, while Inez noted that the kitchen rule governs conflicting visual and thermometer evidence. Mara divided the cooked meal into exactly four equal portions. Two portions were designated for serving now and prepared for serving; the other two were designated for storage and began compliant cooling in labeled shallow containers. The following logged entries describe the storage portions:\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.\",\"For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time."}, {"path": ["evidence", "3"], "text": "For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.", "negative_left": "For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.", "negative_right": "For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 20:25 local time.", "right": "For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-021", "id": "fast-41-diverse-190-021-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara prepared paprika chicken and rice. She finished simmering the sauce before starting the chicken. A probe used for two chicken readings passed its ice-water calibration that day and produced readings of 75°C and 74°C. Sol observed slight pinkness in one cut piece, while Inez noted that the kitchen rule governs conflicting visual and thermometer evidence. Mara divided the cooked meal into exactly four equal portions. Two portions were designated for serving now and prepared for serving; the other two were designated for storage and began compliant cooling in labeled shallow containers. The following logged entries describe the storage portions:", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.", "For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time."]}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, scope, and criteria for both contexts. Both contexts retain Mara, the meal, the four portions, the two serving portions, and the two storage portions. The required base evidence quotes are preserved exactly: “For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.” and “For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time.” All evidence entries are complete factual sentences. The counterfactual's 20:25 cooling time changes an observation without creating contradictory duplicate assertions. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"Home cook Mara prepared paprika chicken and rice. She finished simmering the sauce before starting the chicken. A probe used for two chicken readings passed its ice-water calibration that day and produced readings of 75°C and 74°C. Sol observed slight pinkness in one cut piece, while Inez noted that the kitchen rule governs conflicting visual and thermometer evidence. Mara divided the cooked meal into exactly four equal portions. Two portions were designated for serving now and prepared for serving; the other two were designated for storage and began compliant cooling in labeled shallow containers. The following logged entries describe the storage portions:\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.\",\"For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time."}, {"path": ["evidence", "3"], "text": "For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.", "negative_left": "For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.", "negative_right": "For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 20:25 local time.", "right": "For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 19:40 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-021", "id": "fast-41-diverse-190-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara prepared paprika chicken and rice. She finished simmering the sauce before starting the chicken. A probe used for two chicken readings passed its ice-water calibration that day and produced readings of 75°C and 74°C. Sol observed slight pinkness in one cut piece, while Inez noted that the kitchen rule governs conflicting visual and thermometer evidence. Mara divided the cooked meal into exactly four equal portions. Two portions were designated for serving now and prepared for serving; the other two were designated for storage and began compliant cooling in labeled shallow containers. The following logged entries describe the storage portions:", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "For Mara's two designated storage portions, cooking ended on 14 June 2026 at 18:10 local time.", "For Mara's two designated storage portions, compliant cooling began on 14 June 2026 at 20:25 local time."]}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric and criteria. The meal, portions, serving scope, and cooling path remain bound to the original request. The focus evidence contains exactly two complete factual sentences: \"Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.\" and \"The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026.\" The counterfactual changes only the cooling-start time to 21:05 and introduces no contradictory duplicate measurement or assertion. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"At a community kitchen on 22 September 2026, Mara prepared paprika chicken and rice. She finished simmering the sauce before starting the chicken. The calibrated probe used for the chicken passed its ice-water check that day and produced two readings, 76°C and 74°C. One cut surface looked faintly pink, but the thermometer readings governed the inspection. Mara divided the completed meal into exactly four equal portions. Two portions were designated for serving immediately and were plated for service. The other two were designated for storage and placed in shallow containers, where compliant cooling began. Inez logged the relevant preparation steps, while Dev handled the storage containers. The meal is being reviewed against the kitchen’s readiness standard.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.\",\"The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026.\"],\"request\":\"Is this meal at Level 3 readiness, allowing the two prepared portions to be served now while the other two continue cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026."}, {"path": ["evidence", "3"], "text": "The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.", "negative_left": "Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.", "negative_right": "The same meal record shows that compliant cooling began for the two designated storage portions at 21:05 local time on 22 September 2026.", "right": "The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-030", "id": "fast-41-diverse-190-030-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "At a community kitchen on 22 September 2026, Mara prepared paprika chicken and rice. She finished simmering the sauce before starting the chicken. The calibrated probe used for the chicken passed its ice-water check that day and produced two readings, 76°C and 74°C. One cut surface looked faintly pink, but the thermometer readings governed the inspection. Mara divided the completed meal into exactly four equal portions. Two portions were designated for serving immediately and were plated for service. The other two were designated for storage and placed in shallow containers, where compliant cooling began. Inez logged the relevant preparation steps, while Dev handled the storage containers. The meal is being reviewed against the kitchen’s readiness standard.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.", "The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026."], "request": "Is this meal at Level 3 readiness, allowing the two prepared portions to be served now while the other two continue cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric and criteria. The meal, portions, serving scope, and cooling path remain bound to the original request. The focus evidence contains exactly two complete factual sentences: \"Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.\" and \"The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026.\" The counterfactual changes only the cooling-start time to 21:05 and introduces no contradictory duplicate measurement or assertion. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"At a community kitchen on 22 September 2026, Mara prepared paprika chicken and rice. She finished simmering the sauce before starting the chicken. The calibrated probe used for the chicken passed its ice-water check that day and produced two readings, 76°C and 74°C. One cut surface looked faintly pink, but the thermometer readings governed the inspection. Mara divided the completed meal into exactly four equal portions. Two portions were designated for serving immediately and were plated for service. The other two were designated for storage and placed in shallow containers, where compliant cooling began. Inez logged the relevant preparation steps, while Dev handled the storage containers. The meal is being reviewed against the kitchen’s readiness standard.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.\",\"The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026.\"],\"request\":\"Is this meal at Level 3 readiness, allowing the two prepared portions to be served now while the other two continue cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026."}, {"path": ["evidence", "3"], "text": "The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.", "negative_left": "Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.", "negative_right": "The same meal record shows that compliant cooling began for the two designated storage portions at 21:05 local time on 22 September 2026.", "right": "The same meal record shows that compliant cooling began for the two designated storage portions at 20:05 local time on 22 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-030", "id": "fast-41-diverse-190-030-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "At a community kitchen on 22 September 2026, Mara prepared paprika chicken and rice. She finished simmering the sauce before starting the chicken. The calibrated probe used for the chicken passed its ice-water check that day and produced two readings, 76°C and 74°C. One cut surface looked faintly pink, but the thermometer readings governed the inspection. Mara divided the completed meal into exactly four equal portions. Two portions were designated for serving immediately and were plated for service. The other two were designated for storage and placed in shallow containers, where compliant cooling began. Inez logged the relevant preparation steps, while Dev handled the storage containers. The meal is being reviewed against the kitchen’s readiness standard.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara's meal record shows that cooking ended for the two designated storage portions at 18:20 local time on 22 September 2026.", "The same meal record shows that compliant cooling began for the two designated storage portions at 21:05 local time on 22 September 2026."], "request": "Is this meal at Level 3 readiness, allowing the two prepared portions to be served now while the other two continue cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object remains verbatim and the recipe and kitchen rule remain intact. Question bindings are preserved because both requests concern Mara’s meal, Level 3 readiness, serving portions, storage portions, and the same timing scope. Evidence passes because “Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.” and “The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026.” are complete factual sentences. The counterfactual is coherent because 20:05 is a changed cooling-start observation with unchanged meal identity, counts, portions, and cooking measurements. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"Home cook Mara prepared paprika chicken and rice while Sol assisted with ingredient preparation and Inez reviewed the kitchen records. Mara finished simmering the sauce before starting the chicken. The calibrated probe used for the chicken produced readings of 76°C and 74°C, and its ice-water check was recorded that day. Mara divided the meal into four equal portions: two were designated for immediate serving and prepared for serving, while two were designated for storage and placed in shallow containers for compliant cooling. The preparation log and portion labels identify the same meal throughout.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.\",\"The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026.\",\"The kitchen log records two thermometer readings, both at least 74°C, and identifies exactly four equal portions. Two portions are marked for serving now and prepared for serving; the other two are marked for storage and began compliant cooling.\"] ,\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026."}, {"path": ["evidence", "3"], "text": "The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.", "negative_left": "Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.", "negative_right": "The compliant-cooling record for Mara's two designated storage portions begins at 20:05 local time on 12 September 2026.", "right": "The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-035", "id": "fast-41-diverse-190-035-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara prepared paprika chicken and rice while Sol assisted with ingredient preparation and Inez reviewed the kitchen records. Mara finished simmering the sauce before starting the chicken. The calibrated probe used for the chicken produced readings of 76°C and 74°C, and its ice-water check was recorded that day. Mara divided the meal into four equal portions: two were designated for immediate serving and prepared for serving, while two were designated for storage and placed in shallow containers for compliant cooling. The preparation log and portion labels identify the same meal throughout.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.", "The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026.", "The kitchen log records two thermometer readings, both at least 74°C, and identifies exactly four equal portions. Two portions are marked for serving now and prepared for serving; the other two are marked for storage and began compliant cooling."], "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object remains verbatim and the recipe and kitchen rule remain intact. Question bindings are preserved because both requests concern Mara’s meal, Level 3 readiness, serving portions, storage portions, and the same timing scope. Evidence passes because “Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.” and “The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026.” are complete factual sentences. The counterfactual is coherent because 20:05 is a changed cooling-start observation with unchanged meal identity, counts, portions, and cooking measurements. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"Home cook Mara prepared paprika chicken and rice while Sol assisted with ingredient preparation and Inez reviewed the kitchen records. Mara finished simmering the sauce before starting the chicken. The calibrated probe used for the chicken produced readings of 76°C and 74°C, and its ice-water check was recorded that day. Mara divided the meal into four equal portions: two were designated for immediate serving and prepared for serving, while two were designated for storage and placed in shallow containers for compliant cooling. The preparation log and portion labels identify the same meal throughout.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.\",\"The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026.\",\"The kitchen log records two thermometer readings, both at least 74°C, and identifies exactly four equal portions. Two portions are marked for serving now and prepared for serving; the other two are marked for storage and began compliant cooling.\"] ,\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026."}, {"path": ["evidence", "3"], "text": "The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.", "negative_left": "Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.", "negative_right": "The compliant-cooling record for Mara's two designated storage portions begins at 20:05 local time on 12 September 2026.", "right": "The compliant-cooling record for Mara's two designated storage portions begins at 18:45 local time on 12 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-035", "id": "fast-41-diverse-190-035-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara prepared paprika chicken and rice while Sol assisted with ingredient preparation and Inez reviewed the kitchen records. Mara finished simmering the sauce before starting the chicken. The calibrated probe used for the chicken produced readings of 76°C and 74°C, and its ice-water check was recorded that day. Mara divided the meal into four equal portions: two were designated for immediate serving and prepared for serving, while two were designated for storage and placed in shallow containers for compliant cooling. The preparation log and portion labels identify the same meal throughout.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara's two designated storage portions finished cooking at 17:20 local time on 12 September 2026.", "The compliant-cooling record for Mara's two designated storage portions begins at 20:05 local time on 12 September 2026.", "The kitchen log records two thermometer readings, both at least 74°C, and identifies exactly four equal portions. Two portions are marked for serving now and prepared for serving; the other two are marked for storage and began compliant cooling."], "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is included verbatim and neither context adds or changes governing rules. Question bindings remain Mara, the designated serving portions, the storage portions, and the same readiness request. Base evidence includes “The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.” and “The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time.” Counterfactual evidence includes “The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.” and “The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 20:10 local time.” The evidence spans are complete factual sentences, and the changed time creates no duplicate or internal contradiction. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"At a neighborhood cooking class, Mara prepared paprika chicken and rice while Sol measured the listed ingredients, Inez checked the preparation, and Dev handled the meal containers. Mara finished simmering the sauce before she started cooking the chicken. A calibrated probe produced two readings from the chicken, both at least 74°C, and the probe had passed its ice-water check that day. One cut surface looked faintly pink, but the kitchen record identifies the probe readings as controlling for that observation. Mara divided the finished meal into exactly four equal portions. Two portions were marked for serving immediately and were prepared for serving. The remaining two portions were marked for storage and placed in shallow containers to begin compliant cooling. The kitchen log contains the following dated entries for those storage portions.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.\",\"The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time.\"],\"request\":\"Is Mara's meal ready for serving the designated portions while the other portions continue through storage handling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time."}, {"path": ["evidence", "3"], "text": "The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.", "negative_left": "The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.", "negative_right": "The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 20:10 local time.", "right": "The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-044", "id": "fast-41-diverse-190-044-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "At a neighborhood cooking class, Mara prepared paprika chicken and rice while Sol measured the listed ingredients, Inez checked the preparation, and Dev handled the meal containers. Mara finished simmering the sauce before she started cooking the chicken. A calibrated probe produced two readings from the chicken, both at least 74°C, and the probe had passed its ice-water check that day. One cut surface looked faintly pink, but the kitchen record identifies the probe readings as controlling for that observation. Mara divided the finished meal into exactly four equal portions. Two portions were marked for serving immediately and were prepared for serving. The remaining two portions were marked for storage and placed in shallow containers to begin compliant cooling. The kitchen log contains the following dated entries for those storage portions.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.", "The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time."], "request": "Is Mara's meal ready for serving the designated portions while the other portions continue through storage handling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is included verbatim and neither context adds or changes governing rules. Question bindings remain Mara, the designated serving portions, the storage portions, and the same readiness request. Base evidence includes “The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.” and “The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time.” Counterfactual evidence includes “The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.” and “The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 20:10 local time.” The evidence spans are complete factual sentences, and the changed time creates no duplicate or internal contradiction. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"At a neighborhood cooking class, Mara prepared paprika chicken and rice while Sol measured the listed ingredients, Inez checked the preparation, and Dev handled the meal containers. Mara finished simmering the sauce before she started cooking the chicken. A calibrated probe produced two readings from the chicken, both at least 74°C, and the probe had passed its ice-water check that day. One cut surface looked faintly pink, but the kitchen record identifies the probe readings as controlling for that observation. Mara divided the finished meal into exactly four equal portions. Two portions were marked for serving immediately and were prepared for serving. The remaining two portions were marked for storage and placed in shallow containers to begin compliant cooling. The kitchen log contains the following dated entries for those storage portions.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.\",\"The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time.\"],\"request\":\"Is Mara's meal ready for serving the designated portions while the other portions continue through storage handling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time."}, {"path": ["evidence", "3"], "text": "The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.", "negative_left": "The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.", "negative_right": "The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 20:10 local time.", "right": "The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 19:05 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-044", "id": "fast-41-diverse-190-044-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "At a neighborhood cooking class, Mara prepared paprika chicken and rice while Sol measured the listed ingredients, Inez checked the preparation, and Dev handled the meal containers. Mara finished simmering the sauce before she started cooking the chicken. A calibrated probe produced two readings from the chicken, both at least 74°C, and the probe had passed its ice-water check that day. One cut surface looked faintly pink, but the kitchen record identifies the probe readings as controlling for that observation. Mara divided the finished meal into exactly four equal portions. Two portions were marked for serving immediately and were prepared for serving. The remaining two portions were marked for storage and placed in shallow containers to begin compliant cooling. The kitchen log contains the following dated entries for those storage portions.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "The kitchen log for Mara's two designated storage portions records that cooking ended on 21 June 2026 at 17:35 local time.", "The same kitchen log records that compliant cooling began for those portions on 21 June 2026 at 20:10 local time."], "request": "Is Mara's meal ready for serving the designated portions while the other portions continue through storage handling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric, scope, and bindings. The focus evidence contains two complete factual sentences: \"The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.\" and \"Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time.\" The counterfactual changes only the cooling-start observation to 19:10 while retaining coherent meal, portion, record, and time references. Neither context includes an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"Home cook Mara prepared paprika chicken and rice while Sol assisted with ingredient preparation, Inez reviewed the kitchen record, and Dev handled the portions. Mara finished simmering the sauce before she began cooking the chicken. The probe used for the chicken checks passed its ice-water calibration that day and produced two readings, 75°C and 74°C, from the thickest pieces. One cut piece looked slightly pink, but the kitchen rule governs that visual conflict. Mara divided the finished meal into exactly four equal portions. Two portions were designated for serving now and were prepared for serving; the other two were designated for storage and began compliant cooling in shallow containers. The relevant storage record is reproduced below.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.\",\"Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time.\"],\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time."}, {"path": ["evidence", "3"], "text": "Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.", "negative_left": "The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.", "negative_right": "Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 19:10 local time.", "right": "Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-049", "id": "fast-41-diverse-190-049-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara prepared paprika chicken and rice while Sol assisted with ingredient preparation, Inez reviewed the kitchen record, and Dev handled the portions. Mara finished simmering the sauce before she began cooking the chicken. The probe used for the chicken checks passed its ice-water calibration that day and produced two readings, 75°C and 74°C, from the thickest pieces. One cut piece looked slightly pink, but the kitchen rule governs that visual conflict. Mara divided the finished meal into exactly four equal portions. Two portions were designated for serving now and were prepared for serving; the other two were designated for storage and began compliant cooling in shallow containers. The relevant storage record is reproduced below.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.", "Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time."], "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric, scope, and bindings. The focus evidence contains two complete factual sentences: \"The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.\" and \"Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time.\" The counterfactual changes only the cooling-start observation to 19:10 while retaining coherent meal, portion, record, and time references. Neither context includes an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"Home cook Mara prepared paprika chicken and rice while Sol assisted with ingredient preparation, Inez reviewed the kitchen record, and Dev handled the portions. Mara finished simmering the sauce before she began cooking the chicken. The probe used for the chicken checks passed its ice-water calibration that day and produced two readings, 75°C and 74°C, from the thickest pieces. One cut piece looked slightly pink, but the kitchen rule governs that visual conflict. Mara divided the finished meal into exactly four equal portions. Two portions were designated for serving now and were prepared for serving; the other two were designated for storage and began compliant cooling in shallow containers. The relevant storage record is reproduced below.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.\",\"Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time.\"],\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "2"], "text": "The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time."}, {"path": ["evidence", "3"], "text": "Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.", "negative_left": "The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.", "negative_right": "Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 19:10 local time.", "right": "Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 18:05 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-190-049", "id": "fast-41-diverse-190-049-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara prepared paprika chicken and rice while Sol assisted with ingredient preparation, Inez reviewed the kitchen record, and Dev handled the portions. Mara finished simmering the sauce before she began cooking the chicken. The probe used for the chicken checks passed its ice-water calibration that day and produced two readings, 75°C and 74°C, from the thickest pieces. One cut piece looked slightly pink, but the kitchen rule governs that visual conflict. Mara divided the finished meal into exactly four equal portions. Two portions were designated for serving now and were prepared for serving; the other two were designated for storage and began compliant cooling in shallow containers. The relevant storage record is reproduced below.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "The kitchen log identifies record KS-417 as Mara’s two designated storage portions, with cooking ending on 22 August 2026 at 16:35 local time.", "Record KS-417 records compliant cooling for Mara’s two designated storage portions beginning on 22 August 2026 at 19:10 local time."], "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim, preserve the recipe and handling requirements, maintain the lemon-chicken and event bindings, and contain the two quoted factual evidence sentences. The counterfactual coherently moves portioning event B before measurement M without duplicating or contradicting measurements, and neither context embeds an answer, code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Recipe card\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"Before either logged event, I recognized that no internal-temperature reading for the lemon chicken had been documented or announced. The selected route schedules exactly one thermometer event, M, and requires further cooking if M records below 74°C.\"},{\"speaker\":\"Thermometer log\",\"text\":\"On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time.\"},{\"speaker\":\"Handling log\",\"text\":\"On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:09 local time.\"},{\"speaker\":\"Home cook\",\"text\":\"Event M measured the thickest piece and recorded an internal temperature of at least 74°C. Event B was the first serving or portioning of the lemon chicken into the four dinners.\"},{\"speaker\":\"Storage note\",\"text\":\"The selected route refrigerates every unserved portion within the required two-hour window.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time."}, {"path": ["3", "text"], "text": "On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:09 local time."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time.", "negative_left": "On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time.", "negative_right": "On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:02 local time.", "right": "On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:09 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-191-008", "id": "fast-41-diverse-191-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Recipe card", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "Before either logged event, I recognized that no internal-temperature reading for the lemon chicken had been documented or announced. The selected route schedules exactly one thermometer event, M, and requires further cooking if M records below 74°C."}, {"speaker": "Thermometer log", "text": "On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time."}, {"speaker": "Handling log", "text": "On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:09 local time."}, {"speaker": "Home cook", "text": "Event M measured the thickest piece and recorded an internal temperature of at least 74°C. Event B was the first serving or portioning of the lemon chicken into the four dinners."}, {"speaker": "Storage note", "text": "The selected route refrigerates every unserved portion within the required two-hour window."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim, preserve the recipe and handling requirements, maintain the lemon-chicken and event bindings, and contain the two quoted factual evidence sentences. The counterfactual coherently moves portioning event B before measurement M without duplicating or contradicting measurements, and neither context embeds an answer, code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Recipe card\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"Before either logged event, I recognized that no internal-temperature reading for the lemon chicken had been documented or announced. The selected route schedules exactly one thermometer event, M, and requires further cooking if M records below 74°C.\"},{\"speaker\":\"Thermometer log\",\"text\":\"On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time.\"},{\"speaker\":\"Handling log\",\"text\":\"On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:09 local time.\"},{\"speaker\":\"Home cook\",\"text\":\"Event M measured the thickest piece and recorded an internal temperature of at least 74°C. Event B was the first serving or portioning of the lemon chicken into the four dinners.\"},{\"speaker\":\"Storage note\",\"text\":\"The selected route refrigerates every unserved portion within the required two-hour window.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time."}, {"path": ["3", "text"], "text": "On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:09 local time."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time.", "negative_left": "On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time.", "negative_right": "On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:02 local time.", "right": "On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:09 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-191-008", "id": "fast-41-diverse-191-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Recipe card", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "Before either logged event, I recognized that no internal-temperature reading for the lemon chicken had been documented or announced. The selected route schedules exactly one thermometer event, M, and requires further cooking if M records below 74°C."}, {"speaker": "Thermometer log", "text": "On 17 September 2026, thermometer event M for the lemon chicken was recorded at 18:04 local time."}, {"speaker": "Handling log", "text": "On 17 September 2026, chicken-handling event B for the lemon chicken was recorded at 18:02 local time."}, {"speaker": "Home cook", "text": "Event M measured the thickest piece and recorded an internal temperature of at least 74°C. Event B was the first serving or portioning of the lemon chicken into the four dinners."}, {"speaker": "Storage note", "text": "The selected route refrigerates every unserved portion within the required two-hour window."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy and bindings, the two evidence spans are complete factual sentences, the time change is coherent without duplicate contradictions, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"Before either scheduled event, I recognize that no internal-temperature reading for the lemon chicken has been documented or announced. Event M will measure the thickest piece, and the selected route schedules exactly one thermometer event for this chicken.\"},{\"speaker\":\"Kitchen log\",\"text\":\"On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12.\"},{\"speaker\":\"Kitchen log\",\"text\":\"On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:27.\"},{\"speaker\":\"Meal planner\",\"text\":\"Event B is the first serving or portioning of the lemon chicken into the four dinners. Event M records an internal temperature of at least 74°C; the route requires further cooking if its result is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The selected route refrigerates every unserved portion within that required window.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12."}, {"path": ["3", "text"], "text": "On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:27."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12.", "negative_left": "On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12.", "negative_right": "On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:03.", "right": "On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:27."}, "verifier_independent_model": false}, "family": "fast-41-diverse-191-022", "id": "fast-41-diverse-191-022-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "Before either scheduled event, I recognize that no internal-temperature reading for the lemon chicken has been documented or announced. Event M will measure the thickest piece, and the selected route schedules exactly one thermometer event for this chicken."}, {"speaker": "Kitchen log", "text": "On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12."}, {"speaker": "Kitchen log", "text": "On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:27."}, {"speaker": "Meal planner", "text": "Event B is the first serving or portioning of the lemon chicken into the four dinners. Event M records an internal temperature of at least 74°C; the route requires further cooking if its result is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The selected route refrigerates every unserved portion within that required window."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy and bindings, the two evidence spans are complete factual sentences, the time change is coherent without duplicate contradictions, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"Before either scheduled event, I recognize that no internal-temperature reading for the lemon chicken has been documented or announced. Event M will measure the thickest piece, and the selected route schedules exactly one thermometer event for this chicken.\"},{\"speaker\":\"Kitchen log\",\"text\":\"On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12.\"},{\"speaker\":\"Kitchen log\",\"text\":\"On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:27.\"},{\"speaker\":\"Meal planner\",\"text\":\"Event B is the first serving or portioning of the lemon chicken into the four dinners. Event M records an internal temperature of at least 74°C; the route requires further cooking if its result is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The selected route refrigerates every unserved portion within that required window.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12."}, {"path": ["3", "text"], "text": "On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:27."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12.", "negative_left": "On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12.", "negative_right": "On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:03.", "right": "On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:27."}, "verifier_independent_model": false}, "family": "fast-41-diverse-191-022", "id": "fast-41-diverse-191-022-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "Before either scheduled event, I recognize that no internal-temperature reading for the lemon chicken has been documented or announced. Event M will measure the thickest piece, and the selected route schedules exactly one thermometer event for this chicken."}, {"speaker": "Kitchen log", "text": "On 17 September 2026, thermometer event M for the lemon chicken has a recorded time of 18:12."}, {"speaker": "Kitchen log", "text": "On 17 September 2026, chicken-handling event B for the lemon chicken has a recorded time of 18:03."}, {"speaker": "Meal planner", "text": "Event B is the first serving or portioning of the lemon chicken into the four dinners. Event M records an internal temperature of at least 74°C; the route requires further cooking if its result is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The selected route refrigerates every unserved portion within that required window."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria, while both contexts retain the same recipe, threshold, chicken, events, date, and routing scope. The question bindings remain stable because M, B, lemon chicken, the four dinners, and the requested readiness decision are unchanged. The evidence consists of exactly two complete factual sentences: “On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time.” and “On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time.” The counterfactual is coherent because B occurs at 17:18 before M at 17:26 without duplicate or contradictory measurements. Neither context embeds a gold answer, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Recipe card\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"Before either logged action, I recognized that no internal-temperature reading for the lemon chicken had been documented or announced. The chosen route schedules exactly one thermometer event, M, and requires further cooking if that check is below 74°C.\"},{\"speaker\":\"Thermometer log\",\"text\":\"On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time.\"},{\"speaker\":\"Handling log\",\"text\":\"On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time.\"},{\"speaker\":\"Thermometer log\",\"text\":\"Event M measured the thickest piece of the lemon chicken and recorded an internal temperature of at least 74°C.\"},{\"speaker\":\"Handling log\",\"text\":\"Event B was the first serving or portioning of the lemon chicken into the four dinners.\"},{\"speaker\":\"Storage note\",\"text\":\"The chosen route refrigerates every unserved portion within the required two-hour window.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time."}, {"path": ["3", "text"], "text": "On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time.", "negative_left": "On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time.", "negative_right": "On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:18 local time.", "right": "On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-191-029", "id": "fast-41-diverse-191-029-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Recipe card", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "Before either logged action, I recognized that no internal-temperature reading for the lemon chicken had been documented or announced. The chosen route schedules exactly one thermometer event, M, and requires further cooking if that check is below 74°C."}, {"speaker": "Thermometer log", "text": "On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time."}, {"speaker": "Handling log", "text": "On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time."}, {"speaker": "Thermometer log", "text": "Event M measured the thickest piece of the lemon chicken and recorded an internal temperature of at least 74°C."}, {"speaker": "Handling log", "text": "Event B was the first serving or portioning of the lemon chicken into the four dinners."}, {"speaker": "Storage note", "text": "The chosen route refrigerates every unserved portion within the required two-hour window."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and scoring criteria, while both contexts retain the same recipe, threshold, chicken, events, date, and routing scope. The question bindings remain stable because M, B, lemon chicken, the four dinners, and the requested readiness decision are unchanged. The evidence consists of exactly two complete factual sentences: “On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time.” and “On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time.” The counterfactual is coherent because B occurs at 17:18 before M at 17:26 without duplicate or contradictory measurements. Neither context embeds a gold answer, answer code, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Recipe card\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"Before either logged action, I recognized that no internal-temperature reading for the lemon chicken had been documented or announced. The chosen route schedules exactly one thermometer event, M, and requires further cooking if that check is below 74°C.\"},{\"speaker\":\"Thermometer log\",\"text\":\"On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time.\"},{\"speaker\":\"Handling log\",\"text\":\"On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time.\"},{\"speaker\":\"Thermometer log\",\"text\":\"Event M measured the thickest piece of the lemon chicken and recorded an internal temperature of at least 74°C.\"},{\"speaker\":\"Handling log\",\"text\":\"Event B was the first serving or portioning of the lemon chicken into the four dinners.\"},{\"speaker\":\"Storage note\",\"text\":\"The chosen route refrigerates every unserved portion within the required two-hour window.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time."}, {"path": ["3", "text"], "text": "On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time.", "negative_left": "On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time.", "negative_right": "On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:18 local time.", "right": "On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:41 local time."}, "verifier_independent_model": false}, "family": "fast-41-diverse-191-029", "id": "fast-41-diverse-191-029-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Recipe card", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "Before either logged action, I recognized that no internal-temperature reading for the lemon chicken had been documented or announced. The chosen route schedules exactly one thermometer event, M, and requires further cooking if that check is below 74°C."}, {"speaker": "Thermometer log", "text": "On 12 November 2026, thermometer event M for the lemon chicken was recorded at 17:26 local time."}, {"speaker": "Handling log", "text": "On 12 November 2026, chicken-handling event B for the lemon chicken was recorded at 17:18 local time."}, {"speaker": "Thermometer log", "text": "Event M measured the thickest piece of the lemon chicken and recorded an internal temperature of at least 74°C."}, {"speaker": "Handling log", "text": "Event B was the first serving or portioning of the lemon chicken into the four dinners."}, {"speaker": "Storage note", "text": "The chosen route refrigerates every unserved portion within the required two-hour window."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and neither context adds or removes a policy rule. The shelf, divider-use condition, orientation, measurements, fasteners, load, and specified 20 kg test remain correctly bound. The two evidence quotes are complete factual sentences: \"The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room.\" and \"During inspection, the installed back-panel's white face was oriented toward the room.\" The counterfactual changes only the observed orientation, creating a coherent noncompliant condition without duplicate contradictions. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"connectors_flush": "supported", "connectors_tight": "supported", "diagonal_within_tolerance": "supported", "joints_stable": "supported", "orientation_match": "supported", "specified_load_no_sway": "supported"}, "full_context_fact_states": {"base": {"connectors_flush": "supported", "connectors_tight": "supported", "diagonal_within_tolerance": "supported", "joints_stable": "supported", "orientation_match": "supported", "specified_load_no_sway": "supported"}, "counterfactual": {"connectors_flush": "supported", "connectors_tight": "supported", "diagonal_within_tolerance": "supported", "joints_stable": "supported", "orientation_match": "refuted", "specified_load_no_sway": "supported"}, "remove_left": {"orientation_match": "unknown"}, "remove_right": {"orientation_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"orientation_match": "unknown"}, "negative_pair": {"orientation_match": "refuted"}, "negative_sentence": {"orientation_match": "unknown"}, "positive_pair": {"orientation_match": "supported"}, "right": {"orientation_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, including permissible universal claims over the connector or joint set. The focus is the factual match between installed orientation and required final-use orientation, not a policy outcome. Base and counter assignments can differ only in panel orientation while retaining the same connector, diagonal, and load-test facts, so both are realizable. Policy evidence preserves the state-originating conditional-use binding, orientation mapping, and specified-load scope needed to interpret the unchanged question; the governing decision rubric itself is already retained in original_input.questions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every condition required for final placement/pass: matching orientation, flush and non-loose connectors, immobile joints, diagonal difference at most 3 mm, and no sway under the specified load. No competing defect remains possible under these literals.", "rule_index": 0, "sound": true}, {"reason": "A refuted orientation_match means the installed orientation conflicts with the conditionally determined final use, which is independently sufficient for step correction. The remaining literals also exclude connector defects, excessive diagonal difference, and sway, though those exclusions are not needed for sufficiency.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "orientation_match", "statement": "The installed back-panel face orientation matches the orientation required for the shelf's final use as determined after the specified 20 kg distributed-load test."}, {"id": "connectors_flush", "statement": "Every installed cam and bolt is flush with the surrounding surface."}, {"id": "connectors_tight", "statement": "Every installed cam and bolt is snug rather than loose."}, {"id": "joints_stable", "statement": "Every assembled joint remains immobile under hand pressure."}, {"id": "diagonal_within_tolerance", "statement": "The absolute difference between the shelf's two measured diagonals is 3 mm or less."}, {"id": "specified_load_no_sway", "statement": "The shelf shows no sway under the specified 20 kg distributed load."}], "base_state_json": "[{\"speaker\":\"Inspection record\",\"text\":\"The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room.\"},{\"speaker\":\"Inspection record\",\"text\":\"During inspection, the installed back-panel's white face was oriented toward the room.\"},{\"speaker\":\"Parts checker\",\"text\":\"I intend to use the shelf as a freestanding divider if it passes the load test; otherwise I will place it against the wall.\"},{\"speaker\":\"Parts checker\",\"text\":\"The instructions require the white face toward the room for divider use, but toward the wall for wall placement.\"},{\"speaker\":\"Measurement log\",\"text\":\"The two diagonal measurements were 1200 mm and 1203 mm, giving an absolute difference of 3 mm. Every installed cam and bolt was snug and flush with the surrounding surface.\"},{\"speaker\":\"Assembly check\",\"text\":\"Hand pressure produced no movement at any assembled joint. With the specified 20 kg distributed load, the shelf remained level and showed no sway.\"},{\"speaker\":\"Test record\",\"text\":\"the specified 20 kg distributed load\"}]", "base_states": [{"atom_id": "orientation_match", "state": "supported"}, {"atom_id": "connectors_flush", "state": "supported"}, {"atom_id": "connectors_tight", "state": "supported"}, {"atom_id": "joints_stable", "state": "supported"}, {"atom_id": "diagonal_within_tolerance", "state": "supported"}, {"atom_id": "specified_load_no_sway", "state": "supported"}], "counter_states": [{"atom_id": "orientation_match", "state": "refuted"}, {"atom_id": "connectors_flush", "state": "supported"}, {"atom_id": "connectors_tight", "state": "supported"}, {"atom_id": "joints_stable", "state": "supported"}, {"atom_id": "diagonal_within_tolerance", "state": "supported"}, {"atom_id": "specified_load_no_sway", "state": "supported"}], "focus_atom": "orientation_match", "focus_evidence": [{"path": ["0", "text"], "text": "The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room."}, {"path": ["1", "text"], "text": "During inspection, the installed back-panel's white face was oriented toward the room."}], "policy_evidence": [{"path": ["0", "text"], "text": "I intend to use the shelf as a freestanding divider if it passes the load test; otherwise I will place it against the wall."}, {"path": ["1", "text"], "text": "The instructions require the white face toward the room for divider use, but toward the wall for wall placement."}, {"path": ["3", "text"], "text": "the specified 20 kg distributed load"}], "rules": [{"justification": "The panel orientation matches the final use determined after the specified load test; every connector is flush and snug; no joint moves; the diagonal difference is within the inclusive 3 mm limit; and the shelf has no sway under the specified 20 kg distributed load.", "target": "final_placement_pass", "when": [{"atom_id": "orientation_match", "state": "supported"}, {"atom_id": "connectors_flush", "state": "supported"}, {"atom_id": "connectors_tight", "state": "supported"}, {"atom_id": "joints_stable", "state": "supported"}, {"atom_id": "diagonal_within_tolerance", "state": "supported"}, {"atom_id": "specified_load_no_sway", "state": "supported"}]}, {"justification": "The installed panel orientation conflicts with the final use determined after the specified load test, which independently requires step correction; the favorable connector, diagonal, and load conditions exclude the other stated defects.", "target": "step_correction", "when": [{"atom_id": "orientation_match", "state": "refuted"}, {"atom_id": "connectors_flush", "state": "supported"}, {"atom_id": "connectors_tight", "state": "supported"}, {"atom_id": "joints_stable", "state": "supported"}, {"atom_id": "diagonal_within_tolerance", "state": "supported"}, {"atom_id": "specified_load_no_sway", "state": "supported"}]}]}, "verified_pair": {"left": "The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room.", "negative_left": "The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room.", "negative_right": "During inspection, the installed back-panel's white face was oriented toward the wall.", "right": "During inspection, the installed back-panel's white face was oriented toward the room."}, "verifier_independent_model": false}, "family": "fast-41-diverse-193-015", "id": "fast-41-diverse-193-015-base", "input": {"questions": {"decision": {"criteria": {"final_placement_pass": "Pass inspection and move to final divider placement because orientation, fasteners, diagonal tolerance, and load stability all satisfy the rubric.", "step_correction": "Disassemble or repeat an assembly step because orientation is wrong for the conditionally determined use, or because diagonal difference is greater than 3 mm.", "tightening": "Keep the current assembly sequence but tighten connectors because at least one is proud, loose, or allows joint movement."}, "instructions": "Select the inspection route under this rubric. Choose step correction if panel orientation conflicts with the conditionally determined final use or if diagonal difference exceeds 3 mm. Choose tightening if orientation is correct but any connector is proud, loose, or permits joint movement. Choose final placement/pass only if orientation matches final use, connectors are flush and stable, diagonal difference is 3 mm or less, and the load test shows no sway.", "type": "choice"}}, "state": [{"speaker": "Inspection record", "text": "The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room."}, {"speaker": "Inspection record", "text": "During inspection, the installed back-panel's white face was oriented toward the room."}, {"speaker": "Parts checker", "text": "I intend to use the shelf as a freestanding divider if it passes the load test; otherwise I will place it against the wall."}, {"speaker": "Parts checker", "text": "The instructions require the white face toward the room for divider use, but toward the wall for wall placement."}, {"speaker": "Measurement log", "text": "The two diagonal measurements were 1200 mm and 1203 mm, giving an absolute difference of 3 mm. Every installed cam and bolt was snug and flush with the surrounding surface."}, {"speaker": "Assembly check", "text": "Hand pressure produced no movement at any assembled joint. With the specified 20 kg distributed load, the shelf remained level and showed no sway."}, {"speaker": "Test record", "text": "the specified 20 kg distributed load"}]}, "method": "c2d", "provenance": {"source_id": "diverse-193", "source_is_synthetic": true, "source_sha256": "3ec98e0e6419e3e1c5f2c6a06ebc4044f7ad14392d4468bb25beb6ad6ccfb17f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "final_placement_pass"}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and neither context adds or removes a policy rule. The shelf, divider-use condition, orientation, measurements, fasteners, load, and specified 20 kg test remain correctly bound. The two evidence quotes are complete factual sentences: \"The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room.\" and \"During inspection, the installed back-panel's white face was oriented toward the room.\" The counterfactual changes only the observed orientation, creating a coherent noncompliant condition without duplicate contradictions. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"connectors_flush": "supported", "connectors_tight": "supported", "diagonal_within_tolerance": "supported", "joints_stable": "supported", "orientation_match": "refuted", "specified_load_no_sway": "supported"}, "full_context_fact_states": {"base": {"connectors_flush": "supported", "connectors_tight": "supported", "diagonal_within_tolerance": "supported", "joints_stable": "supported", "orientation_match": "supported", "specified_load_no_sway": "supported"}, "counterfactual": {"connectors_flush": "supported", "connectors_tight": "supported", "diagonal_within_tolerance": "supported", "joints_stable": "supported", "orientation_match": "refuted", "specified_load_no_sway": "supported"}, "remove_left": {"orientation_match": "unknown"}, "remove_right": {"orientation_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"orientation_match": "unknown"}, "negative_pair": {"orientation_match": "refuted"}, "negative_sentence": {"orientation_match": "unknown"}, "positive_pair": {"orientation_match": "supported"}, "right": {"orientation_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, including permissible universal claims over the connector or joint set. The focus is the factual match between installed orientation and required final-use orientation, not a policy outcome. Base and counter assignments can differ only in panel orientation while retaining the same connector, diagonal, and load-test facts, so both are realizable. Policy evidence preserves the state-originating conditional-use binding, orientation mapping, and specified-load scope needed to interpret the unchanged question; the governing decision rubric itself is already retained in original_input.questions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every condition required for final placement/pass: matching orientation, flush and non-loose connectors, immobile joints, diagonal difference at most 3 mm, and no sway under the specified load. No competing defect remains possible under these literals.", "rule_index": 0, "sound": true}, {"reason": "A refuted orientation_match means the installed orientation conflicts with the conditionally determined final use, which is independently sufficient for step correction. The remaining literals also exclude connector defects, excessive diagonal difference, and sway, though those exclusions are not needed for sufficiency.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "orientation_match", "statement": "The installed back-panel face orientation matches the orientation required for the shelf's final use as determined after the specified 20 kg distributed-load test."}, {"id": "connectors_flush", "statement": "Every installed cam and bolt is flush with the surrounding surface."}, {"id": "connectors_tight", "statement": "Every installed cam and bolt is snug rather than loose."}, {"id": "joints_stable", "statement": "Every assembled joint remains immobile under hand pressure."}, {"id": "diagonal_within_tolerance", "statement": "The absolute difference between the shelf's two measured diagonals is 3 mm or less."}, {"id": "specified_load_no_sway", "statement": "The shelf shows no sway under the specified 20 kg distributed load."}], "base_state_json": "[{\"speaker\":\"Inspection record\",\"text\":\"The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room.\"},{\"speaker\":\"Inspection record\",\"text\":\"During inspection, the installed back-panel's white face was oriented toward the room.\"},{\"speaker\":\"Parts checker\",\"text\":\"I intend to use the shelf as a freestanding divider if it passes the load test; otherwise I will place it against the wall.\"},{\"speaker\":\"Parts checker\",\"text\":\"The instructions require the white face toward the room for divider use, but toward the wall for wall placement.\"},{\"speaker\":\"Measurement log\",\"text\":\"The two diagonal measurements were 1200 mm and 1203 mm, giving an absolute difference of 3 mm. Every installed cam and bolt was snug and flush with the surrounding surface.\"},{\"speaker\":\"Assembly check\",\"text\":\"Hand pressure produced no movement at any assembled joint. With the specified 20 kg distributed load, the shelf remained level and showed no sway.\"},{\"speaker\":\"Test record\",\"text\":\"the specified 20 kg distributed load\"}]", "base_states": [{"atom_id": "orientation_match", "state": "supported"}, {"atom_id": "connectors_flush", "state": "supported"}, {"atom_id": "connectors_tight", "state": "supported"}, {"atom_id": "joints_stable", "state": "supported"}, {"atom_id": "diagonal_within_tolerance", "state": "supported"}, {"atom_id": "specified_load_no_sway", "state": "supported"}], "counter_states": [{"atom_id": "orientation_match", "state": "refuted"}, {"atom_id": "connectors_flush", "state": "supported"}, {"atom_id": "connectors_tight", "state": "supported"}, {"atom_id": "joints_stable", "state": "supported"}, {"atom_id": "diagonal_within_tolerance", "state": "supported"}, {"atom_id": "specified_load_no_sway", "state": "supported"}], "focus_atom": "orientation_match", "focus_evidence": [{"path": ["0", "text"], "text": "The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room."}, {"path": ["1", "text"], "text": "During inspection, the installed back-panel's white face was oriented toward the room."}], "policy_evidence": [{"path": ["0", "text"], "text": "I intend to use the shelf as a freestanding divider if it passes the load test; otherwise I will place it against the wall."}, {"path": ["1", "text"], "text": "The instructions require the white face toward the room for divider use, but toward the wall for wall placement."}, {"path": ["3", "text"], "text": "the specified 20 kg distributed load"}], "rules": [{"justification": "The panel orientation matches the final use determined after the specified load test; every connector is flush and snug; no joint moves; the diagonal difference is within the inclusive 3 mm limit; and the shelf has no sway under the specified 20 kg distributed load.", "target": "final_placement_pass", "when": [{"atom_id": "orientation_match", "state": "supported"}, {"atom_id": "connectors_flush", "state": "supported"}, {"atom_id": "connectors_tight", "state": "supported"}, {"atom_id": "joints_stable", "state": "supported"}, {"atom_id": "diagonal_within_tolerance", "state": "supported"}, {"atom_id": "specified_load_no_sway", "state": "supported"}]}, {"justification": "The installed panel orientation conflicts with the final use determined after the specified load test, which independently requires step correction; the favorable connector, diagonal, and load conditions exclude the other stated defects.", "target": "step_correction", "when": [{"atom_id": "orientation_match", "state": "refuted"}, {"atom_id": "connectors_flush", "state": "supported"}, {"atom_id": "connectors_tight", "state": "supported"}, {"atom_id": "joints_stable", "state": "supported"}, {"atom_id": "diagonal_within_tolerance", "state": "supported"}, {"atom_id": "specified_load_no_sway", "state": "supported"}]}]}, "verified_pair": {"left": "The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room.", "negative_left": "The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room.", "negative_right": "During inspection, the installed back-panel's white face was oriented toward the wall.", "right": "During inspection, the installed back-panel's white face was oriented toward the room."}, "verifier_independent_model": false}, "family": "fast-41-diverse-193-015", "id": "fast-41-diverse-193-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"final_placement_pass": "Pass inspection and move to final divider placement because orientation, fasteners, diagonal tolerance, and load stability all satisfy the rubric.", "step_correction": "Disassemble or repeat an assembly step because orientation is wrong for the conditionally determined use, or because diagonal difference is greater than 3 mm.", "tightening": "Keep the current assembly sequence but tighten connectors because at least one is proud, loose, or allows joint movement."}, "instructions": "Select the inspection route under this rubric. Choose step correction if panel orientation conflicts with the conditionally determined final use or if diagonal difference exceeds 3 mm. Choose tightening if orientation is correct but any connector is proud, loose, or permits joint movement. Choose final placement/pass only if orientation matches final use, connectors are flush and stable, diagonal difference is 3 mm or less, and the load test shows no sway.", "type": "choice"}}, "state": [{"speaker": "Inspection record", "text": "The specified 20 kg distributed-load test determined that the shelf's final use would be a freestanding divider, for which the instructions require the white back-panel face toward the room."}, {"speaker": "Inspection record", "text": "During inspection, the installed back-panel's white face was oriented toward the wall."}, {"speaker": "Parts checker", "text": "I intend to use the shelf as a freestanding divider if it passes the load test; otherwise I will place it against the wall."}, {"speaker": "Parts checker", "text": "The instructions require the white face toward the room for divider use, but toward the wall for wall placement."}, {"speaker": "Measurement log", "text": "The two diagonal measurements were 1200 mm and 1203 mm, giving an absolute difference of 3 mm. Every installed cam and bolt was snug and flush with the surrounding surface."}, {"speaker": "Assembly check", "text": "Hand pressure produced no movement at any assembled joint. With the specified 20 kg distributed load, the shelf remained level and showed no sway."}, {"speaker": "Test record", "text": "the specified 20 kg distributed load"}]}, "method": "c2d", "provenance": {"source_id": "diverse-193", "source_is_synthetic": true, "source_sha256": "3ec98e0e6419e3e1c5f2c6a06ebc4044f7ad14392d4468bb25beb6ad6ccfb17f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "step_correction"}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions object retains all governing instructions and criteria. Question bindings remain unchanged for the Alderline unit, the decision immediately after 14:25, and test AL-47. The evidence contains two complete factual sentences: \"For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test.\" \"Test AL-47 records a wobble amplitude of 0 millimetres.\" The counterfactual coherently changes only the wobble measurement while retaining correct parts and assembly configuration. Neither context includes a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"At the current-routing decision immediately after 14:25, the evidence log identifies the relevant test and records the latest inspection findings for the fictional Alderline three-shelf unit. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test. Test AL-47 records a wobble amplitude of 0 millimetres. The maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres. Every component required by the instructions is present, every present component is undamaged, and every present component is required by the instructions. Every installed component has the identity specified by the instructions. Every observed assembly-configuration feature matches the corresponding instruction. The log is restricted to evidence available at this decision point; no earlier corrected condition is carried forward.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test."}, {"path": [], "text": "Test AL-47 records a wobble amplitude of 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test.", "negative_left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test.", "negative_right": "Test AL-47 records a wobble amplitude of 3.5 millimetres.", "right": "Test AL-47 records a wobble amplitude of 0 millimetres."}, "verifier_independent_model": false}, "family": "fast-41-diverse-194-001", "id": "fast-41-diverse-194-001-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "At the current-routing decision immediately after 14:25, the evidence log identifies the relevant test and records the latest inspection findings for the fictional Alderline three-shelf unit. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test. Test AL-47 records a wobble amplitude of 0 millimetres. The maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres. Every component required by the instructions is present, every present component is undamaged, and every present component is required by the instructions. Every installed component has the identity specified by the instructions. Every observed assembly-configuration feature matches the corresponding instruction. The log is restricted to evidence available at this decision point; no earlier corrected condition is carried forward."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions object retains all governing instructions and criteria. Question bindings remain unchanged for the Alderline unit, the decision immediately after 14:25, and test AL-47. The evidence contains two complete factual sentences: \"For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test.\" \"Test AL-47 records a wobble amplitude of 0 millimetres.\" The counterfactual coherently changes only the wobble measurement while retaining correct parts and assembly configuration. Neither context includes a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"At the current-routing decision immediately after 14:25, the evidence log identifies the relevant test and records the latest inspection findings for the fictional Alderline three-shelf unit. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test. Test AL-47 records a wobble amplitude of 0 millimetres. The maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres. Every component required by the instructions is present, every present component is undamaged, and every present component is required by the instructions. Every installed component has the identity specified by the instructions. Every observed assembly-configuration feature matches the corresponding instruction. The log is restricted to evidence available at this decision point; no earlier corrected condition is carried forward.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test."}, {"path": [], "text": "Test AL-47 records a wobble amplitude of 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test.", "negative_left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test.", "negative_right": "Test AL-47 records a wobble amplitude of 3.5 millimetres.", "right": "Test AL-47 records a wobble amplitude of 0 millimetres."}, "verifier_independent_model": false}, "family": "fast-41-diverse-194-001", "id": "fast-41-diverse-194-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "At the current-routing decision immediately after 14:25, the evidence log identifies the relevant test and records the latest inspection findings for the fictional Alderline three-shelf unit. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, load-and-push test AL-47 is the latest applicable test. Test AL-47 records a wobble amplitude of 3.5 millimetres. The maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres. Every component required by the instructions is present, every present component is undamaged, and every present component is required by the instructions. Every installed component has the identity specified by the instructions. Every observed assembly-configuration feature matches the corresponding instruction. The log is restricted to evidence available at this decision point; no earlier corrected condition is carried forward."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "tightening"}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings without answer leakage; the retained evidence quotes are \"At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417.\" and \"At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres.\", and the counterfactual coherently changes the wobble to 1 millimetre.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"Case note: The Alderline three-shelf unit remained under review at the current-routing decision immediately after 14:25. At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417. At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres. The latest applicable evidence records a nonnegative real-number wobble measurement in millimetres and a maximum fastening-play measurement of 0 millimetres. It confirms that every component required by the instructions is present, every present component is undamaged, and every present component is required by the instructions. The same evidence confirms that every installed component has the identity specified by the instructions and that every observed assembly-configuration feature matches its corresponding instruction. These findings concern the current evidence only; earlier observations are not carried forward as current defects. The unit is awaiting routing.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417."}, {"path": [], "text": "At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417.", "negative_left": "At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly one millimetre above the value assigned to marker M-417.", "negative_right": "At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres.", "right": "At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres."}, "verifier_independent_model": false}, "family": "fast-41-diverse-194-003", "id": "fast-41-diverse-194-003-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "Case note: The Alderline three-shelf unit remained under review at the current-routing decision immediately after 14:25. At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417. At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres. The latest applicable evidence records a nonnegative real-number wobble measurement in millimetres and a maximum fastening-play measurement of 0 millimetres. It confirms that every component required by the instructions is present, every present component is undamaged, and every present component is required by the instructions. The same evidence confirms that every installed component has the identity specified by the instructions and that every observed assembly-configuration feature matches its corresponding instruction. These findings concern the current evidence only; earlier observations are not carried forward as current defects. The unit is awaiting routing."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings without answer leakage; the retained evidence quotes are \"At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417.\" and \"At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres.\", and the counterfactual coherently changes the wobble to 1 millimetre.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"Case note: The Alderline three-shelf unit remained under review at the current-routing decision immediately after 14:25. At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417. At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres. The latest applicable evidence records a nonnegative real-number wobble measurement in millimetres and a maximum fastening-play measurement of 0 millimetres. It confirms that every component required by the instructions is present, every present component is undamaged, and every present component is required by the instructions. The same evidence confirms that every installed component has the identity specified by the instructions and that every observed assembly-configuration feature matches its corresponding instruction. These findings concern the current evidence only; earlier observations are not carried forward as current defects. The unit is awaiting routing.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417."}, {"path": [], "text": "At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly the value assigned to marker M-417.", "negative_left": "At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly one millimetre above the value assigned to marker M-417.", "negative_right": "At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres.", "right": "At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres."}, "verifier_independent_model": false}, "family": "fast-41-diverse-194-003", "id": "fast-41-diverse-194-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "Case note: The Alderline three-shelf unit remained under review at the current-routing decision immediately after 14:25. At the current-routing decision immediately after 14:25, the latest applicable load-and-push test records the fictional Alderline three-shelf unit's wobble amplitude as exactly one millimetre above the value assigned to marker M-417. At the current-routing decision immediately after 14:25, marker M-417 has the assigned value 0 millimetres. The latest applicable evidence records a nonnegative real-number wobble measurement in millimetres and a maximum fastening-play measurement of 0 millimetres. It confirms that every component required by the instructions is present, every present component is undamaged, and every present component is required by the instructions. The same evidence confirms that every installed component has the identity specified by the instructions and that every observed assembly-configuration feature matches its corresponding instruction. These findings concern the current evidence only; earlier observations are not carried forward as current defects. The unit is awaiting routing."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "tightening"}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy. Both contexts preserve the Alderline unit, current-routing timing, latest-evidence scope, and intervention bindings. Each focus-evidence set contains two complete factual sentences. The counterfactual coherently changes R-17 from 0 to 3 millimetres without duplicate measurements or contradiction. Neither context states a routing answer, answer code, rule table, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"At the current-routing decision immediately after 14:25, the latest applicable evidence for the fictional Alderline three-shelf unit includes the following two independently verified observations: “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17.” “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 0 millimetres.” The recorded wobble measurement is a nonnegative real-number measurement in millimetres. The same latest applicable test recorded maximum fastening play of 0 millimetres. Evidence confirms that every component required by the instructions is present, every present component is undamaged and required, every installed component has the identity specified by the instructions, and every observed assembly-configuration feature matches the corresponding instruction. The evidence was collected after the specified fastening intervention, and no earlier defect remains part of the current record.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17."}, {"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17.", "negative_left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17.", "negative_right": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 3 millimetres.", "right": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 0 millimetres."}, "verifier_independent_model": false}, "family": "fast-41-diverse-194-007", "id": "fast-41-diverse-194-007-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "At the current-routing decision immediately after 14:25, the latest applicable evidence for the fictional Alderline three-shelf unit includes the following two independently verified observations: “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17.” “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 0 millimetres.” The recorded wobble measurement is a nonnegative real-number measurement in millimetres. The same latest applicable test recorded maximum fastening play of 0 millimetres. Evidence confirms that every component required by the instructions is present, every present component is undamaged and required, every installed component has the identity specified by the instructions, and every observed assembly-configuration feature matches the corresponding instruction. The evidence was collected after the specified fastening intervention, and no earlier defect remains part of the current record."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy. Both contexts preserve the Alderline unit, current-routing timing, latest-evidence scope, and intervention bindings. Each focus-evidence set contains two complete factual sentences. The counterfactual coherently changes R-17 from 0 to 3 millimetres without duplicate measurements or contradiction. Neither context states a routing answer, answer code, rule table, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"At the current-routing decision immediately after 14:25, the latest applicable evidence for the fictional Alderline three-shelf unit includes the following two independently verified observations: “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17.” “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 0 millimetres.” The recorded wobble measurement is a nonnegative real-number measurement in millimetres. The same latest applicable test recorded maximum fastening play of 0 millimetres. Evidence confirms that every component required by the instructions is present, every present component is undamaged and required, every installed component has the identity specified by the instructions, and every observed assembly-configuration feature matches the corresponding instruction. The evidence was collected after the specified fastening intervention, and no earlier defect remains part of the current record.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17."}, {"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 0 millimetres."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17.", "negative_left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17.", "negative_right": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 3 millimetres.", "right": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 0 millimetres."}, "verifier_independent_model": false}, "family": "fast-41-diverse-194-007", "id": "fast-41-diverse-194-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "At the current-routing decision immediately after 14:25, the latest applicable evidence for the fictional Alderline three-shelf unit includes the following two independently verified observations: “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is identical to the calibrated reference reading labeled R-17.” “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, calibrated reference reading R-17 equals 3 millimetres.” The recorded wobble measurement is a nonnegative real-number measurement in millimetres. The same latest applicable test recorded maximum fastening play of 0 millimetres. Evidence confirms that every component required by the instructions is present, every present component is undamaged and required, every installed component has the identity specified by the instructions, and every observed assembly-configuration feature matches the corresponding instruction. The evidence was collected after the specified fastening intervention, and no earlier defect remains part of the current record."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "tightening"}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and bindings. The evidence spans are complete factual sentences: “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units.” and “For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude.” The base context is coherent with a zero-millimetre calibrated measurement. The counterfactual coherently changes that calibration to 1 millimetre without contradictory duplicate measurements. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"At the Alderline three-shelf unit’s routing review immediately after 14:25, the latest applicable evidence consisted of the completed build, its repeat test, and the calibration record. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units. For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude. The instrument’s millimetre measurements are nonnegative real numbers. The latest test also recorded maximum fastening play of 0 millimetres. Every component required by the instructions was present; every present component was undamaged and required; every installed component had the identity specified by the instructions; and every observed assembly-configuration feature matched its corresponding instruction. The review relied only on this latest evidence, while earlier defects had been corrected before the repeat test.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units."}, {"path": [], "text": "For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units.", "negative_left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units.", "negative_right": "For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 1 millimetre of wobble amplitude.", "right": "For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude."}, "verifier_independent_model": false}, "family": "fast-41-diverse-194-013", "id": "fast-41-diverse-194-013-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "At the Alderline three-shelf unit’s routing review immediately after 14:25, the latest applicable evidence consisted of the completed build, its repeat test, and the calibration record. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units. For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude. The instrument’s millimetre measurements are nonnegative real numbers. The latest test also recorded maximum fastening play of 0 millimetres. Every component required by the instructions was present; every present component was undamaged and required; every installed component had the identity specified by the instructions; and every observed assembly-configuration feature matched its corresponding instruction. The review relied only on this latest evidence, while earlier defects had been corrected before the repeat test."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and bindings. The evidence spans are complete factual sentences: “For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units.” and “For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude.” The base context is coherent with a zero-millimetre calibrated measurement. The counterfactual coherently changes that calibration to 1 millimetre without contradictory duplicate measurements. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified component and configuration relations. The focus is the factual equality of the latest wobble measurement to zero, not a routing policy. The base and counter assignments can differ only in that measurement: with a nonnegative measurement, the base permits zero and the counter entails a positive value, while all other facts remain fixed. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state contributes only case observations that need not be preserved for new synthetic contexts.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes all three issue routes: zero wobble and zero fastening play exclude tightening; complete, undamaged, non-extra, correctly identified components exclude parts sorting; and matching observed assembly configuration excludes step correction. Therefore none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refuting an exact-zero wobble measurement while supporting that it is a nonnegative real measurement entails positive wobble. The remaining conditions exclude component defects and assembly-configuration mismatches, so the policy's tightening route applies.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test equals 0 millimetres."}, {"id": "a2", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the wobble amplitude recorded by the latest applicable load-and-push test is a nonnegative real-number measurement in millimetres."}, {"id": "a3", "statement": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the maximum fastening-play measurement recorded by the latest applicable test equals 0 millimetres."}, {"id": "a4", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every component required by the instructions is present."}, {"id": "a5", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is undamaged."}, {"id": "a6", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every present component is required by the instructions."}, {"id": "a7", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every installed component has the identity specified by the instructions."}, {"id": "a8", "statement": "In the latest applicable evidence for the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, every observed assembly-configuration feature matches the corresponding instruction."}], "base_state_json": "\"At the Alderline three-shelf unit’s routing review immediately after 14:25, the latest applicable evidence consisted of the completed build, its repeat test, and the calibration record. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units. For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude. The instrument’s millimetre measurements are nonnegative real numbers. The latest test also recorded maximum fastening play of 0 millimetres. Every component required by the instructions was present; every present component was undamaged and required; every installed component had the identity specified by the instructions; and every observed assembly-configuration feature matched its corresponding instruction. The review relied only on this latest evidence, while earlier defects had been corrected before the repeat test.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units."}, {"path": [], "text": "For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude."}], "policy_evidence": [], "rules": [{"justification": "The latest applicable test has zero measured wobble and zero measured fastening play. The latest applicable evidence also establishes that all required parts are present, no present part is damaged or extra, every installed part has the specified identity, and every observed assembly-configuration feature matches the instructions. Thus none of the three issue routes fits.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The latest applicable test has an existing nonnegative real-number wobble measurement that is not zero, so its measured wobble amplitude is positive. All component and assembly-configuration correctness conditions are established, excluding parts sorting and step correction; the current positive wobble therefore requires tightening.", "target": "tightening", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units.", "negative_left": "For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units.", "negative_right": "For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 1 millimetre of wobble amplitude.", "right": "For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 0 millimetres of wobble amplitude."}, "verifier_independent_model": false}, "family": "fast-41-diverse-194-013", "id": "fast-41-diverse-194-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when none of the three issue routes fits the latest evidence, including when the build now passes inspection and is ready for final placement.", "parts_sorting": "Route to parts sorting only when the latest evidence shows a missing, damaged, extra, or substituted component; do not use for assembly or fastening issues.", "step_correction": "Route to step correction only when all required parts are present but the latest evidence shows an incorrect orientation, sequence, hole choice, or other mismatch with the instructions; do not use for a fastening-only defect.", "tightening": "Route to tightening only when the latest test still shows looseness or wobble while the parts and assembly configuration are otherwise correct."}, "instructions": "Using only the latest applicable evidence, choose the correct current routing. Earlier defects that were successfully corrected do not count as current issues.", "type": "choice"}}, "state": "At the Alderline three-shelf unit’s routing review immediately after 14:25, the latest applicable evidence consisted of the completed build, its repeat test, and the calibration record. For the fictional Alderline three-shelf unit at the current-routing decision immediately after 14:25, the latest applicable load-and-push test recorded the wobble amplitude as 7.00 calibrated wobble units. For that test and unit, the calibration record states that 7.00 calibrated wobble units corresponds to 1 millimetre of wobble amplitude. The instrument’s millimetre measurements are nonnegative real numbers. The latest test also recorded maximum fastening play of 0 millimetres. Every component required by the instructions was present; every present component was undamaged and required; every installed component had the identity specified by the instructions; and every observed assembly-configuration feature matched its corresponding instruction. The review relied only on this latest evidence, while earlier defects had been corrected before the repeat test."}, "method": "c2d", "provenance": {"source_id": "diverse-194", "source_is_synthetic": true, "source_sha256": "d2603e3ee1d996eedf2ed52bb3af37b0f669167318abd381e7c6af83a82127ef", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "tightening"}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both policy arrays preserve the governing rule and conflict policy. The cabinet, back-panel, inspection, and condition bindings remain unchanged. The base evidence retains “At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.” and “At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face.” The counterfactual retains the first sentence and changes the second to “At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its rough face.” The counterfactual is coherent because the rough face can point inward without duplicating measurements. Neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"case_note\":\"The Alderline C-4 cabinet was assembled for a living-room workshop and presented for final inspection at 14:00 UTC on 2026-09-17. The inspection record identifies two required anti-tip straps; the first is secured to its mounting point, and the second is secured to its mounting point. During final-use testing, the cabinet remains steady without rocking. The inspector records the following paired observations:\",\"evidence\":[\"At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.\",\"At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face.\"],\"policy\":[\"The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing.\",\"When notes conflict with timestamped images, the later image controls.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle."}, {"path": ["evidence", "1"], "text": "At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.", "negative_left": "At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.", "negative_right": "At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its rough face.", "right": "At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face."}, "verifier_independent_model": false}, "family": "fast-41-diverse-195-010", "id": "fast-41-diverse-195-010-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"case_note": "The Alderline C-4 cabinet was assembled for a living-room workshop and presented for final inspection at 14:00 UTC on 2026-09-17. The inspection record identifies two required anti-tip straps; the first is secured to its mounting point, and the second is secured to its mounting point. During final-use testing, the cabinet remains steady without rocking. The inspector records the following paired observations:", "evidence": ["At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.", "At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face."], "policy": ["The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing.", "When notes conflict with timestamped images, the later image controls."]}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both policy arrays preserve the governing rule and conflict policy. The cabinet, back-panel, inspection, and condition bindings remain unchanged. The base evidence retains “At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.” and “At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face.” The counterfactual retains the first sentence and changes the second to “At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its rough face.” The counterfactual is coherent because the rough face can point inward without duplicating measurements. Neither context embeds an answer, code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"case_note\":\"The Alderline C-4 cabinet was assembled for a living-room workshop and presented for final inspection at 14:00 UTC on 2026-09-17. The inspection record identifies two required anti-tip straps; the first is secured to its mounting point, and the second is secured to its mounting point. During final-use testing, the cabinet remains steady without rocking. The inspector records the following paired observations:\",\"evidence\":[\"At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.\",\"At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face.\"],\"policy\":[\"The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing.\",\"When notes conflict with timestamped images, the later image controls.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle."}, {"path": ["evidence", "1"], "text": "At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.", "negative_left": "At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.", "negative_right": "At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its rough face.", "right": "At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its smooth face."}, "verifier_independent_model": false}, "family": "fast-41-diverse-195-010", "id": "fast-41-diverse-195-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"case_note": "The Alderline C-4 cabinet was assembled for a living-room workshop and presented for final inspection at 14:00 UTC on 2026-09-17. The inspection record identifies two required anti-tip straps; the first is secured to its mounting point, and the second is secured to its mounting point. During final-use testing, the cabinet remains steady without rocking. The inspector records the following paired observations:", "evidence": ["At final inspection at 14:00 UTC on 2026-09-17, the side of the back panel installed in the fictional Alderline C-4 cabinet that faces the cabinet interior bears a blue triangle.", "At final inspection at 14:00 UTC on 2026-09-17, the blue triangle on the fictional Alderline C-4 cabinet’s back panel is painted on its rough face."], "policy": ["The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing.", "When notes conflict with timestamped images, the later image controls."]}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the supplied pass rule, conflict-resolution policy, and unchanged question scope. Both requests remain bound to whether the fictional Alderline C-4 cabinet passes final build inspection. Each context has two complete factual evidence sentences. The counterfactual changes only image 7’s orientation finding, without creating contradictory duplicate measurements or assertions. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction. Retained evidence quotes are: \"At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.\" \"Inspection image 7 records surface S pointing inward.\" \"Inspection image 7 records surface S pointing outward.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"A final inspection was conducted on the fictional Alderline C-4 cabinet after assembly in the workshop. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The inspection record identifies image 7 as the later close-up and links it to the cabinet’s final condition. The tester logged the cabinet as stable during the final-use test, and the installation checklist marks each of the two required strap positions complete. Surface S is the surface identified in image 7’s annotation. The inspection question concerns whether this cabinet passes final build inspection.\",\"evidence\":[\"At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.\",\"Inspection image 7 records surface S pointing inward.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7."}, {"path": ["evidence", "1"], "text": "Inspection image 7 records surface S pointing inward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.", "negative_left": "At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.", "negative_right": "Inspection image 7 records surface S pointing outward.", "right": "Inspection image 7 records surface S pointing inward."}, "verifier_independent_model": false}, "family": "fast-41-diverse-195-018", "id": "fast-41-diverse-195-018-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "A final inspection was conducted on the fictional Alderline C-4 cabinet after assembly in the workshop. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The inspection record identifies image 7 as the later close-up and links it to the cabinet’s final condition. The tester logged the cabinet as stable during the final-use test, and the installation checklist marks each of the two required strap positions complete. Surface S is the surface identified in image 7’s annotation. The inspection question concerns whether this cabinet passes final build inspection.", "evidence": ["At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.", "Inspection image 7 records surface S pointing inward."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the supplied pass rule, conflict-resolution policy, and unchanged question scope. Both requests remain bound to whether the fictional Alderline C-4 cabinet passes final build inspection. Each context has two complete factual evidence sentences. The counterfactual changes only image 7’s orientation finding, without creating contradictory duplicate measurements or assertions. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction. Retained evidence quotes are: \"At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.\" \"Inspection image 7 records surface S pointing inward.\" \"Inspection image 7 records surface S pointing outward.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"A final inspection was conducted on the fictional Alderline C-4 cabinet after assembly in the workshop. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The inspection record identifies image 7 as the later close-up and links it to the cabinet’s final condition. The tester logged the cabinet as stable during the final-use test, and the installation checklist marks each of the two required strap positions complete. Surface S is the surface identified in image 7’s annotation. The inspection question concerns whether this cabinet passes final build inspection.\",\"evidence\":[\"At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.\",\"Inspection image 7 records surface S pointing inward.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7."}, {"path": ["evidence", "1"], "text": "Inspection image 7 records surface S pointing inward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.", "negative_left": "At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.", "negative_right": "Inspection image 7 records surface S pointing outward.", "right": "Inspection image 7 records surface S pointing inward."}, "verifier_independent_model": false}, "family": "fast-41-diverse-195-018", "id": "fast-41-diverse-195-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "A final inspection was conducted on the fictional Alderline C-4 cabinet after assembly in the workshop. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The inspection record identifies image 7 as the later close-up and links it to the cabinet’s final condition. The tester logged the cabinet as stable during the final-use test, and the installation checklist marks each of the two required strap positions complete. Surface S is the surface identified in image 7’s annotation. The inspection question concerns whether this cabinet passes final build inspection.", "evidence": ["At final inspection, both required anti-tip straps are installed on the fictional Alderline C-4 cabinet, the cabinet does not rock during final-use testing, and the smooth face of its back panel is designated surface S in inspection image 7.", "Inspection image 7 records surface S pointing outward."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions retain the governing rule and decision instructions. Question bindings are preserved for the fictional Alderline C-4 cabinet and final build inspection. Both evidence items are complete factual sentences. The counterfactual coherently changes the blue seal’s direction without creating duplicate contradictory measurements. Neither context embeds a gold answer, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. A final checklist records that the first required anti-tip strap is secured, and a separate fastening photo confirms the second required anti-tip strap is secured. The cabinet remains steady while drawers are opened, loaded, and closed during final-use testing. The inspection occurred after assembly was complete, with no further installation work recorded.\",\"evidence\":[\"At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal.\",\"At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces inward.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal."}, {"path": ["evidence", "1"], "text": "At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces inward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal.", "negative_left": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal.", "negative_right": "At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces outward.", "right": "At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces inward."}, "verifier_independent_model": false}, "family": "fast-41-diverse-195-021", "id": "fast-41-diverse-195-021-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. A final checklist records that the first required anti-tip strap is secured, and a separate fastening photo confirms the second required anti-tip strap is secured. The cabinet remains steady while drawers are opened, loaded, and closed during final-use testing. The inspection occurred after assembly was complete, with no further installation work recorded.", "evidence": ["At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal.", "At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces inward."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions retain the governing rule and decision instructions. Question bindings are preserved for the fictional Alderline C-4 cabinet and final build inspection. Both evidence items are complete factual sentences. The counterfactual coherently changes the blue seal’s direction without creating duplicate contradictory measurements. Neither context embeds a gold answer, answer code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. A final checklist records that the first required anti-tip strap is secured, and a separate fastening photo confirms the second required anti-tip strap is secured. The cabinet remains steady while drawers are opened, loaded, and closed during final-use testing. The inspection occurred after assembly was complete, with no further installation work recorded.\",\"evidence\":[\"At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal.\",\"At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces inward.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal."}, {"path": ["evidence", "1"], "text": "At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces inward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal.", "negative_left": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal.", "negative_right": "At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces outward.", "right": "At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces inward."}, "verifier_independent_model": false}, "family": "fast-41-diverse-195-021", "id": "fast-41-diverse-195-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. A final checklist records that the first required anti-tip strap is secured, and a separate fastening photo confirms the second required anti-tip strap is secured. The cabinet remains steady while drawers are opened, loaded, and closed during final-use testing. The inspection occurred after assembly was complete, with no further installation work recorded.", "evidence": ["At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet is the face bearing the blue seal.", "At final inspection, the blue seal on the back panel installed in the fictional Alderline C-4 cabinet faces outward."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and cabinet-inspection binding; the two evidence sentences are complete factual claims—“At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel.” and “The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward.”—and the counterfactual changes only that image’s direction without contradiction or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"Case note for the fictional Alderline C-4 cabinet: At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel. The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward. At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed, and the second required anti-tip strap is installed. During final-use testing, the cabinet does not rock. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The inspection question concerns whether this completed cabinet passes final build inspection.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["context"], "text": "At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel."}, {"path": ["context"], "text": "The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel.", "negative_left": "At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel.", "negative_right": "The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing outward.", "right": "The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward."}, "verifier_independent_model": false}, "family": "fast-41-diverse-195-026", "id": "fast-41-diverse-195-026-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "Case note for the fictional Alderline C-4 cabinet: At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel. The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward. At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed, and the second required anti-tip strap is installed. During final-use testing, the cabinet does not rock. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The inspection question concerns whether this completed cabinet passes final build inspection."}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the policy and cabinet-inspection binding; the two evidence sentences are complete factual claims—“At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel.” and “The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward.”—and the counterfactual changes only that image’s direction without contradiction or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"Case note for the fictional Alderline C-4 cabinet: At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel. The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward. At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed, and the second required anti-tip strap is installed. During final-use testing, the cabinet does not rock. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The inspection question concerns whether this completed cabinet passes final build inspection.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["context"], "text": "At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel."}, {"path": ["context"], "text": "The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel.", "negative_left": "At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel.", "negative_right": "The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing outward.", "right": "The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing inward."}, "verifier_independent_model": false}, "family": "fast-41-diverse-195-026", "id": "fast-41-diverse-195-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "Case note for the fictional Alderline C-4 cabinet: At 18:00 on 14 September 2026, final inspection of the fictional Alderline C-4 cabinet used a timestamped image of its installed back panel. The timestamped 18:00 image from 14 September 2026 records the smooth face of the back panel pointing outward. At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed, and the second required anti-tip strap is installed. During final-use testing, the cabinet does not rock. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The inspection question concerns whether this completed cabinet passes final build inspection."}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and Larkspur, step-6, placement-date, and reopening bindings, the two evidence spans are complete factual sentences—“During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit.” and “The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821.”—and the counterfactual coherently changes only the submitted unit’s serial record to L-7394 without embedding an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Review note\",\"text\":\"For the 2026-09-17 placement decision, every required safety test for the physical Larkspur unit had a passing result, and every concealed-orientation checkpoint other than step 6 had qualifying evidence. The step-6 back-panel checkpoint had zero required progress images. During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward and documented that observation in the step-6 verification file used for the decision. Apart from session S, there were zero current direct visual confirmations available for the submitted unit’s step-6 back-panel orientation. The review materials include the following serial-number records.\"},{\"speaker\":\"Inspection record\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit.\"},{\"speaker\":\"Placement record\",\"text\":\"The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821.\"},{\"speaker\":\"Governing policy\",\"text\":\"Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit."}, {"path": ["2", "text"], "text": "The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit.", "negative_left": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit.", "negative_right": "The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-7394.", "right": "The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821."}, "verifier_independent_model": false}, "family": "fast-41-diverse-196-014", "id": "fast-41-diverse-196-014-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Review note", "text": "For the 2026-09-17 placement decision, every required safety test for the physical Larkspur unit had a passing result, and every concealed-orientation checkpoint other than step 6 had qualifying evidence. The step-6 back-panel checkpoint had zero required progress images. During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward and documented that observation in the step-6 verification file used for the decision. Apart from session S, there were zero current direct visual confirmations available for the submitted unit’s step-6 back-panel orientation. The review materials include the following serial-number records."}, {"speaker": "Inspection record", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit."}, {"speaker": "Placement record", "text": "The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821."}, {"speaker": "Governing policy", "text": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and Larkspur, step-6, placement-date, and reopening bindings, the two evidence spans are complete factual sentences—“During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit.” and “The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821.”—and the counterfactual coherently changes only the submitted unit’s serial record to L-7394 without embedding an answer or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Review note\",\"text\":\"For the 2026-09-17 placement decision, every required safety test for the physical Larkspur unit had a passing result, and every concealed-orientation checkpoint other than step 6 had qualifying evidence. The step-6 back-panel checkpoint had zero required progress images. During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward and documented that observation in the step-6 verification file used for the decision. Apart from session S, there were zero current direct visual confirmations available for the submitted unit’s step-6 back-panel orientation. The review materials include the following serial-number records.\"},{\"speaker\":\"Inspection record\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit.\"},{\"speaker\":\"Placement record\",\"text\":\"The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821.\"},{\"speaker\":\"Governing policy\",\"text\":\"Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit."}, {"path": ["2", "text"], "text": "The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit.", "negative_left": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit.", "negative_right": "The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-7394.", "right": "The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-4821."}, "verifier_independent_model": false}, "family": "fast-41-diverse-196-014", "id": "fast-41-diverse-196-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Review note", "text": "For the 2026-09-17 placement decision, every required safety test for the physical Larkspur unit had a passing result, and every concealed-orientation checkpoint other than step 6 had qualifying evidence. The step-6 back-panel checkpoint had zero required progress images. During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward and documented that observation in the step-6 verification file used for the decision. Apart from session S, there were zero current direct visual confirmations available for the submitted unit’s step-6 back-panel orientation. The review materials include the following serial-number records."}, {"speaker": "Inspection record", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector recorded the uniquely identifying serial number L-4821 on the inspected physical unit."}, {"speaker": "Placement record", "text": "The placement record for the Larkspur unit submitted for final placement on 2026-09-17 lists its uniquely identifying serial number as L-7394."}, {"speaker": "Governing policy", "text": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question and governing policy. The focus evidence contains two complete factual sentences. The counterfactual consistently identifies the inspected and submitted units with different serial numbers, without contradictory duplicate claims or embedded answer instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Placement record\",\"text\":\"For the 2026-09-17 decision, every required safety test for the physical Larkspur unit passed. Every concealed-orientation checkpoint other than step 6 had qualifying evidence. No required progress image was available for the step-6 back-panel checkpoint.\"},{\"speaker\":\"Inspection record\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward. The observation was entered in the step-6 verification file used for the Larkspur placement decision.\"},{\"speaker\":\"Identity record\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482.\"},{\"speaker\":\"Identity record\",\"text\":\"The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-482.\"},{\"speaker\":\"Evidence review\",\"text\":\"Excluding pre-placement visual-inspection session S, zero current direct visual confirmations were available for the submitted Larkspur unit’s step-6 back-panel orientation. The rear trim covered the panel edges, and removing it was required for another direct check. The assembler’s note and successful use test did not expose the concealed panel.\"},{\"speaker\":\"Governing policy\",\"text\":\"Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482."}, {"path": ["3", "text"], "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-482."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482.", "negative_left": "During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482.", "negative_right": "The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-913.", "right": "The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-482."}, "verifier_independent_model": false}, "family": "fast-41-diverse-196-021", "id": "fast-41-diverse-196-021-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Placement record", "text": "For the 2026-09-17 decision, every required safety test for the physical Larkspur unit passed. Every concealed-orientation checkpoint other than step 6 had qualifying evidence. No required progress image was available for the step-6 back-panel checkpoint."}, {"speaker": "Inspection record", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward. The observation was entered in the step-6 verification file used for the Larkspur placement decision."}, {"speaker": "Identity record", "text": "During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482."}, {"speaker": "Identity record", "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-482."}, {"speaker": "Evidence review", "text": "Excluding pre-placement visual-inspection session S, zero current direct visual confirmations were available for the submitted Larkspur unit’s step-6 back-panel orientation. The rear trim covered the panel edges, and removing it was required for another direct check. The assembler’s note and successful use test did not expose the concealed panel."}, {"speaker": "Governing policy", "text": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question and governing policy. The focus evidence contains two complete factual sentences. The counterfactual consistently identifies the inspected and submitted units with different serial numbers, without contradictory duplicate claims or embedded answer instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Placement record\",\"text\":\"For the 2026-09-17 decision, every required safety test for the physical Larkspur unit passed. Every concealed-orientation checkpoint other than step 6 had qualifying evidence. No required progress image was available for the step-6 back-panel checkpoint.\"},{\"speaker\":\"Inspection record\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward. The observation was entered in the step-6 verification file used for the Larkspur placement decision.\"},{\"speaker\":\"Identity record\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482.\"},{\"speaker\":\"Identity record\",\"text\":\"The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-482.\"},{\"speaker\":\"Evidence review\",\"text\":\"Excluding pre-placement visual-inspection session S, zero current direct visual confirmations were available for the submitted Larkspur unit’s step-6 back-panel orientation. The rear trim covered the panel edges, and removing it was required for another direct check. The assembler’s note and successful use test did not expose the concealed panel.\"},{\"speaker\":\"Governing policy\",\"text\":\"Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["2", "text"], "text": "During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482."}, {"path": ["3", "text"], "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-482."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482.", "negative_left": "During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482.", "negative_right": "The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-913.", "right": "The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-482."}, "verifier_independent_model": false}, "family": "fast-41-diverse-196-021", "id": "fast-41-diverse-196-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Placement record", "text": "For the 2026-09-17 decision, every required safety test for the physical Larkspur unit passed. Every concealed-orientation checkpoint other than step 6 had qualifying evidence. No required progress image was available for the step-6 back-panel checkpoint."}, {"speaker": "Inspection record", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward. The observation was entered in the step-6 verification file used for the Larkspur placement decision."}, {"speaker": "Identity record", "text": "During pre-placement visual-inspection session S on 2026-09-16, the physical unit inspected was marked with uniquely identifying serial number LS-482."}, {"speaker": "Identity record", "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 was marked with uniquely identifying serial number LS-913."}, {"speaker": "Evidence review", "text": "Excluding pre-placement visual-inspection session S, zero current direct visual confirmations were available for the submitted Larkspur unit’s step-6 back-panel orientation. The rear trim covered the panel edges, and removing it was required for another direct check. The assembler’s note and successful use test did not expose the concealed panel."}, {"speaker": "Governing policy", "text": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and identical policy preserve all governing requirements. The Larkspur, step-6, and dated-session bindings remain intact despite the counterfactual serial change. The evidence consists of two complete factual sentences: \"The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417.\" and \"The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417.\" The counterfactual consistently distinguishes inspected unit L-417 from submitted unit L-418 without contradictory duplicate assertions. Neither generated context embeds a gold answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417.\"},{\"speaker\":\"Records clerk\",\"text\":\"The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417.\"},{\"speaker\":\"Inspector\",\"text\":\"During session S on 2026-09-16, the inspected unit's back-panel grooved face was oriented inward, and the observation was entered in the step-6 verification file used for the placement decision.\"},{\"speaker\":\"Placement coordinator\",\"text\":\"Every required safety test for the submitted unit passed. All concealed-orientation checkpoints other than step 6 had qualifying evidence, while zero required progress images were available for step 6.\"},{\"speaker\":\"Placement coordinator\",\"text\":\"Apart from session S, zero current direct visual confirmations were available for the submitted unit's step-6 back-panel orientation. The rear trim remained in place at the decision time.\"},{\"speaker\":\"Policy\",\"text\":\"a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["0", "text"], "text": "The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417."}, {"path": ["1", "text"], "text": "The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417.", "negative_left": "The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417.", "negative_right": "The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-418.", "right": "The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-196-024", "id": "fast-41-diverse-196-024-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417."}, {"speaker": "Records clerk", "text": "The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417."}, {"speaker": "Inspector", "text": "During session S on 2026-09-16, the inspected unit's back-panel grooved face was oriented inward, and the observation was entered in the step-6 verification file used for the placement decision."}, {"speaker": "Placement coordinator", "text": "Every required safety test for the submitted unit passed. All concealed-orientation checkpoints other than step 6 had qualifying evidence, while zero required progress images were available for step 6."}, {"speaker": "Placement coordinator", "text": "Apart from session S, zero current direct visual confirmations were available for the submitted unit's step-6 back-panel orientation. The rear trim remained in place at the decision time."}, {"speaker": "Policy", "text": "a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and identical policy preserve all governing requirements. The Larkspur, step-6, and dated-session bindings remain intact despite the counterfactual serial change. The evidence consists of two complete factual sentences: \"The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417.\" and \"The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417.\" The counterfactual consistently distinguishes inspected unit L-417 from submitted unit L-418 without contradictory duplicate assertions. Neither generated context embeds a gold answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Records clerk\",\"text\":\"The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417.\"},{\"speaker\":\"Records clerk\",\"text\":\"The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417.\"},{\"speaker\":\"Inspector\",\"text\":\"During session S on 2026-09-16, the inspected unit's back-panel grooved face was oriented inward, and the observation was entered in the step-6 verification file used for the placement decision.\"},{\"speaker\":\"Placement coordinator\",\"text\":\"Every required safety test for the submitted unit passed. All concealed-orientation checkpoints other than step 6 had qualifying evidence, while zero required progress images were available for step 6.\"},{\"speaker\":\"Placement coordinator\",\"text\":\"Apart from session S, zero current direct visual confirmations were available for the submitted unit's step-6 back-panel orientation. The rear trim remained in place at the decision time.\"},{\"speaker\":\"Policy\",\"text\":\"a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["0", "text"], "text": "The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417."}, {"path": ["1", "text"], "text": "The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417.", "negative_left": "The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417.", "negative_right": "The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-418.", "right": "The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-196-024", "id": "fast-41-diverse-196-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Records clerk", "text": "The inspection record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number L-417."}, {"speaker": "Records clerk", "text": "The final-placement record for 2026-09-17 identifies the submitted physical Larkspur unit by serial number L-418."}, {"speaker": "Inspector", "text": "During session S on 2026-09-16, the inspected unit's back-panel grooved face was oriented inward, and the observation was entered in the step-6 verification file used for the placement decision."}, {"speaker": "Placement coordinator", "text": "Every required safety test for the submitted unit passed. All concealed-orientation checkpoints other than step 6 had qualifying evidence, while zero required progress images were available for step 6."}, {"speaker": "Placement coordinator", "text": "Apart from session S, zero current direct visual confirmations were available for the submitted unit's step-6 back-panel orientation. The rear trim remained in place at the decision time."}, {"speaker": "Policy", "text": "a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and the same governing policy. Both preserve the Larkspur unit, step 6, and the stated dates; the counterfactual serial-number change is a coherent observation change, not a question-binding change. The evidence spans are complete factual sentences: “The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827.” and “The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827.” The counterfactual consistently distinguishes the inspected unit from the submitted unit without contradictory duplicate assertions. Neither context adds a gold answer, answer code, rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Placement record\",\"text\":\"On 2026-09-17, the Larkspur unit was submitted for final placement. Every required safety test had a passing result, and every concealed-orientation checkpoint other than step 6 had qualifying evidence. No required progress image was available for the step-6 back-panel checkpoint.\"},{\"speaker\":\"Inspection record\",\"text\":\"The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827. During that session, an inspector directly observed the back-panel grooved face oriented inward, and the observation was documented in the step-6 verification file used for the placement decision.\"},{\"speaker\":\"Submission record\",\"text\":\"The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827. Excluding session S, no current direct visual confirmation was available for the submitted unit’s step-6 back-panel orientation.\"},{\"speaker\":\"Governing policy\",\"text\":\"Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827."}, {"path": ["2", "text"], "text": "The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827.", "negative_left": "The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827.", "negative_right": "The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-9136.", "right": "The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827."}, "verifier_independent_model": false}, "family": "fast-41-diverse-196-035", "id": "fast-41-diverse-196-035-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Placement record", "text": "On 2026-09-17, the Larkspur unit was submitted for final placement. Every required safety test had a passing result, and every concealed-orientation checkpoint other than step 6 had qualifying evidence. No required progress image was available for the step-6 back-panel checkpoint."}, {"speaker": "Inspection record", "text": "The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827. During that session, an inspector directly observed the back-panel grooved face oriented inward, and the observation was documented in the step-6 verification file used for the placement decision."}, {"speaker": "Submission record", "text": "The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827. Excluding session S, no current direct visual confirmation was available for the submitted unit’s step-6 back-panel orientation."}, {"speaker": "Governing policy", "text": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and the same governing policy. Both preserve the Larkspur unit, step 6, and the stated dates; the counterfactual serial-number change is a coherent observation change, not a question-binding change. The evidence spans are complete factual sentences: “The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827.” and “The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827.” The counterfactual consistently distinguishes the inspected unit from the submitted unit without contradictory duplicate assertions. Neither context adds a gold answer, answer code, rationale, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Placement record\",\"text\":\"On 2026-09-17, the Larkspur unit was submitted for final placement. Every required safety test had a passing result, and every concealed-orientation checkpoint other than step 6 had qualifying evidence. No required progress image was available for the step-6 back-panel checkpoint.\"},{\"speaker\":\"Inspection record\",\"text\":\"The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827. During that session, an inspector directly observed the back-panel grooved face oriented inward, and the observation was documented in the step-6 verification file used for the placement decision.\"},{\"speaker\":\"Submission record\",\"text\":\"The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827. Excluding session S, no current direct visual confirmation was available for the submitted unit’s step-6 back-panel orientation.\"},{\"speaker\":\"Governing policy\",\"text\":\"Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["1", "text"], "text": "The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827."}, {"path": ["2", "text"], "text": "The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827.", "negative_left": "The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827.", "negative_right": "The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-9136.", "right": "The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-4827."}, "verifier_independent_model": false}, "family": "fast-41-diverse-196-035", "id": "fast-41-diverse-196-035-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Placement record", "text": "On 2026-09-17, the Larkspur unit was submitted for final placement. Every required safety test had a passing result, and every concealed-orientation checkpoint other than step 6 had qualifying evidence. No required progress image was available for the step-6 back-panel checkpoint."}, {"speaker": "Inspection record", "text": "The record for pre-placement visual-inspection session S on 2026-09-16 identifies the inspected physical unit by serial number LK-4827. During that session, an inspector directly observed the back-panel grooved face oriented inward, and the observation was documented in the step-6 verification file used for the placement decision."}, {"speaker": "Submission record", "text": "The final-placement submission record for the Larkspur unit on 2026-09-17 identifies the submitted physical unit by serial number LK-9136. Excluding session S, no current direct visual confirmation was available for the submitted unit’s step-6 back-panel orientation."}, {"speaker": "Governing policy", "text": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the scope and 4 mm exception without inventing rules. Entity, test, measurement, and date bindings remain consistent with the original question. The evidence spans are complete factual sentences: “On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001.” and “The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17.” The counterfactual coherently changes the measurement from 6 mm to 3 mm without creating contradictory duplicate assertions within either context. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"A 2026-09-17 inspection covered Alderline five-shelf unit serial AL-001. The inventory contained every required in-scope part, and installed parts matched their specified identifiers and orientations. The assembler completed instruction steps 1 through 12, with each step conforming to its corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error posing an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was identified. On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001. The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17. A rear-only cosmetic mark was noted, and wall anchoring was not included in this inspection.\\n\\nThe manual permits at most 4 mm.\\n\\nWall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001."}, {"path": [], "text": "The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001.", "negative_left": "On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001.", "negative_right": "The calibrated displacement gauge recorded 3 mm for that test on 2026-09-17.", "right": "The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-007", "id": "fast-41-diverse-197-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "A 2026-09-17 inspection covered Alderline five-shelf unit serial AL-001. The inventory contained every required in-scope part, and installed parts matched their specified identifiers and orientations. The assembler completed instruction steps 1 through 12, with each step conforming to its corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error posing an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was identified. On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001. The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17. A rear-only cosmetic mark was noted, and wall anchoring was not included in this inspection.\n\nThe manual permits at most 4 mm.\n\nWall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the scope and 4 mm exception without inventing rules. Entity, test, measurement, and date bindings remain consistent with the original question. The evidence spans are complete factual sentences: “On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001.” and “The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17.” The counterfactual coherently changes the measurement from 6 mm to 3 mm without creating contradictory duplicate assertions within either context. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"A 2026-09-17 inspection covered Alderline five-shelf unit serial AL-001. The inventory contained every required in-scope part, and installed parts matched their specified identifiers and orientations. The assembler completed instruction steps 1 through 12, with each step conforming to its corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error posing an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was identified. On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001. The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17. A rear-only cosmetic mark was noted, and wall anchoring was not included in this inspection.\\n\\nThe manual permits at most 4 mm.\\n\\nWall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001."}, {"path": [], "text": "The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001.", "negative_left": "On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001.", "negative_right": "The calibrated displacement gauge recorded 3 mm for that test on 2026-09-17.", "right": "The calibrated displacement gauge recorded 6 mm for that test on 2026-09-17."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-007", "id": "fast-41-diverse-197-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "A 2026-09-17 inspection covered Alderline five-shelf unit serial AL-001. The inventory contained every required in-scope part, and installed parts matched their specified identifiers and orientations. The assembler completed instruction steps 1 through 12, with each step conforming to its corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error posing an immediate collapse or injury risk. Every accessible fastener was secure, and no inaccessible fastening error was identified. On 2026-09-17, unloaded stability test UT-01 was the designated test for measuring top-corner displacement of Alderline five-shelf unit serial AL-001. The calibrated displacement gauge recorded 3 mm for that test on 2026-09-17. A rear-only cosmetic mark was noted, and wall anchoring was not included in this inspection.\n\nThe manual permits at most 4 mm.\n\nWall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing scope and exception, retain the Alderline and unloaded-test bindings, and contain no answer code or classifier instruction; the evidence consists of the exact factual sentences “During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17.” and “The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm.” The counterfactual coherently changes the calibration meaning to 4 mm or less without contradicting its other observations.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Case note: The inspected inventory for Alderline five-shelf unit serial AL-001 contains every required in-scope part, and installed parts match their specified identifiers. Structural parts are oriented as specified. Assembly records show that instruction steps 1 through 12 were all performed and followed. The unit shows no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. All accessible fasteners are secure, and inspection found no inaccessible fastening error. During unloaded stability test UT-01, nonzero top-corner movement was observed. During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17. The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17."}, {"path": [], "text": "The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17.", "negative_left": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17.", "negative_right": "The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement of 4 mm or less.", "right": "The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-008", "id": "fast-41-diverse-197-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Case note: The inspected inventory for Alderline five-shelf unit serial AL-001 contains every required in-scope part, and installed parts match their specified identifiers. Structural parts are oriented as specified. Assembly records show that instruction steps 1 through 12 were all performed and followed. The unit shows no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. All accessible fasteners are secure, and inspection found no inaccessible fastening error. During unloaded stability test UT-01, nonzero top-corner movement was observed. During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17. The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing scope and exception, retain the Alderline and unloaded-test bindings, and contain no answer code or classifier instruction; the evidence consists of the exact factual sentences “During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17.” and “The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm.” The counterfactual coherently changes the calibration meaning to 4 mm or less without contradicting its other observations.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Case note: The inspected inventory for Alderline five-shelf unit serial AL-001 contains every required in-scope part, and installed parts match their specified identifiers. Structural parts are oriented as specified. Assembly records show that instruction steps 1 through 12 were all performed and followed. The unit shows no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. All accessible fasteners are secure, and inspection found no inaccessible fastening error. During unloaded stability test UT-01, nonzero top-corner movement was observed. During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17. The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17."}, {"path": [], "text": "The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17.", "negative_left": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17.", "negative_right": "The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement of 4 mm or less.", "right": "The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement exceeding 4 mm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-008", "id": "fast-41-diverse-197-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Case note: The inspected inventory for Alderline five-shelf unit serial AL-001 contains every required in-scope part, and installed parts match their specified identifiers. Structural parts are oriented as specified. Assembly records show that instruction steps 1 through 12 were all performed and followed. The unit shows no structural damage, severe instability, or assembly error creating an immediate collapse or injury risk. All accessible fasteners are secure, and inspection found no inaccessible fastening error. During unloaded stability test UT-01, nonzero top-corner movement was observed. During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001, the calibrated instrument display showed reading code B-17. The calibration record for unloaded stability test UT-01 states that reading code B-17 represents a top-corner displacement of 4 mm or less. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain its relevant scope and exception without inventing priorities or defaults. Question bindings for the Alderline unit, unloaded stability test, 4 mm threshold, inspection scope, and placement decision remain intact. The evidence consists of two complete factual sentences: “At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm.” and “During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm.” The counterfactual is coherent because its displacement is 3.2 mm, and no contradictory duplicate measurement or assertion appears. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm. During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm. The UT-01 record notes nonzero top-corner displacement. Inventory review found every in-scope part present; installed parts matched their identifiers, and structural parts had their specified orientations. The assembly log records completion of instruction steps 1 through 12, each conforming to its corresponding instruction. Inspection found no structural damage, severe instability, immediate collapse or injury risk from assembly, or inaccessible fastening error. Every accessible fastener was secure. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm."}, {"path": [], "text": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm.", "negative_left": "At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm.", "negative_right": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value minus 3.0 mm.", "right": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-018", "id": "fast-41-diverse-197-018-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm. During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm. The UT-01 record notes nonzero top-corner displacement. Inventory review found every in-scope part present; installed parts matched their identifiers, and structural parts had their specified orientations. The assembly log records completion of instruction steps 1 through 12, each conforming to its corresponding instruction. Inspection found no structural damage, severe instability, immediate collapse or injury risk from assembly, or inaccessible fastening error. Every accessible fastener was secure. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain its relevant scope and exception without inventing priorities or defaults. Question bindings for the Alderline unit, unloaded stability test, 4 mm threshold, inspection scope, and placement decision remain intact. The evidence consists of two complete factual sentences: “At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm.” and “During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm.” The counterfactual is coherent because its displacement is 3.2 mm, and no contradictory duplicate measurement or assertion appears. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm. During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm. The UT-01 record notes nonzero top-corner displacement. Inventory review found every in-scope part present; installed parts matched their identifiers, and structural parts had their specified orientations. The assembly log records completion of instruction steps 1 through 12, each conforming to its corresponding instruction. Inspection found no structural damage, severe instability, immediate collapse or injury risk from assembly, or inaccessible fastening error. Every accessible fastener was secure. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm."}, {"path": [], "text": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm.", "negative_left": "At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm.", "negative_right": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value minus 3.0 mm.", "right": "During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value plus 1.0 mm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-018", "id": "fast-41-diverse-197-018-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "At 10:00 UTC on 2026-09-17, calibrated gauge G-18 displayed a reading of 6.2 mm. During unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 at 10:00 UTC on 2026-09-17, the measured top-corner displacement equaled gauge G-18's displayed value minus 3.0 mm. The UT-01 record notes nonzero top-corner displacement. Inventory review found every in-scope part present; installed parts matched their identifiers, and structural parts had their specified orientations. The assembly log records completion of instruction steps 1 through 12, each conforming to its corresponding instruction. Inspection found no structural damage, severe instability, immediate collapse or injury risk from assembly, or inaccessible fastening error. Every accessible fastener was secure. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain its scope and exception without inventing rules. The Alderline unit, inspection record, unloaded-stability path, and 4 mm threshold remain bound correctly. The evidence consists of two complete factual sentences: \"Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001.\" and \"Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm.\" The counterfactual coherently changes the displacement to 3 mm, which is consistent with the unchanged 4 mm limit and other observations. Neither context states a level, route, answer code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"The workshop inspection concerned Alderline five-shelf unit serial AL-001. The inventory sheet lists every in-scope component, and identification checks found each installed part matched its specification. Structural parts were oriented correctly. The assembly log records completion of instruction steps 1 through 12, with each step checked against the corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error posing an immediate collapse or injury risk. All accessible fasteners were secure, and no inaccessible fastening error was found. Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001. Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement. A rear-panel surface mark was noted, while the unit otherwise remained intact during review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001."}, {"path": [], "text": "Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001.", "negative_left": "Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001.", "negative_right": "Inspection record UT-01-R17 reports a measured top-corner displacement of 3 mm.", "right": "Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-034", "id": "fast-41-diverse-197-034-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "The workshop inspection concerned Alderline five-shelf unit serial AL-001. The inventory sheet lists every in-scope component, and identification checks found each installed part matched its specification. Structural parts were oriented correctly. The assembly log records completion of instruction steps 1 through 12, with each step checked against the corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error posing an immediate collapse or injury risk. All accessible fasteners were secure, and no inaccessible fastening error was found. Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001. Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement. A rear-panel surface mark was noted, while the unit otherwise remained intact during review."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain its scope and exception without inventing rules. The Alderline unit, inspection record, unloaded-stability path, and 4 mm threshold remain bound correctly. The evidence consists of two complete factual sentences: \"Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001.\" and \"Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm.\" The counterfactual coherently changes the displacement to 3 mm, which is consistent with the unchanged 4 mm limit and other observations. Neither context states a level, route, answer code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"The workshop inspection concerned Alderline five-shelf unit serial AL-001. The inventory sheet lists every in-scope component, and identification checks found each installed part matched its specification. Structural parts were oriented correctly. The assembly log records completion of instruction steps 1 through 12, with each step checked against the corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error posing an immediate collapse or injury risk. All accessible fasteners were secure, and no inaccessible fastening error was found. Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001. Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement. A rear-panel surface mark was noted, while the unit otherwise remained intact during review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001."}, {"path": [], "text": "Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001.", "negative_left": "Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001.", "negative_right": "Inspection record UT-01-R17 reports a measured top-corner displacement of 3 mm.", "right": "Inspection record UT-01-R17 reports a measured top-corner displacement of 6 mm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-034", "id": "fast-41-diverse-197-034-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "The workshop inspection concerned Alderline five-shelf unit serial AL-001. The inventory sheet lists every in-scope component, and identification checks found each installed part matched its specification. Structural parts were oriented correctly. The assembly log records completion of instruction steps 1 through 12, with each step checked against the corresponding instruction. Inspectors found no structural damage, severe instability, or assembly error posing an immediate collapse or injury risk. All accessible fasteners were secure, and no inaccessible fastening error was found. Inspection record UT-01-R17 is the unloaded stability record for Alderline five-shelf unit serial AL-001. Inspection record UT-01-R17 reports a measured top-corner displacement of 3 mm. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement. A rear-panel surface mark was noted, while the unit otherwise remained intact during review."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy and scoring criteria for both contexts. Both contexts retain the Alderline unit, five-shelf build, unloaded stability test, 4 mm threshold, and inspection scope. The evidence consists of two complete factual sentences: \"Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001.\" and \"Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test.\" The counterfactual coherently changes the displacement from 6 mm to 3 mm without retaining a contradictory measurement. Neither context states a score, answer code, label rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Case note for Alderline five-shelf unit serial AL-001: The parts inventory lists every in-scope panel, connector, and fastener required for the five-shelf build. A component audit found each installed part matched its specified identifier, and structural members were oriented as shown in the assembly record. The technician completed instruction steps 1 through 12 and recorded conformity at each step. The unit showed no structural damage, severe instability, or immediate collapse or injury risk. All accessible fasteners were secure, with no inaccessible fastening error identified. Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001. Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test. The test log confirms that top-corner movement was observed. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement. A rear-panel surface mark was noted but excluded from this inspection.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001."}, {"path": [], "text": "Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001.", "negative_left": "Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001.", "negative_right": "Inspection record ST-UT-41 records a top-corner displacement of 3 mm during the unloaded test.", "right": "Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-041", "id": "fast-41-diverse-197-041-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Case note for Alderline five-shelf unit serial AL-001: The parts inventory lists every in-scope panel, connector, and fastener required for the five-shelf build. A component audit found each installed part matched its specified identifier, and structural members were oriented as shown in the assembly record. The technician completed instruction steps 1 through 12 and recorded conformity at each step. The unit showed no structural damage, severe instability, or immediate collapse or injury risk. All accessible fasteners were secure, with no inaccessible fastening error identified. Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001. Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test. The test log confirms that top-corner movement was observed. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement. A rear-panel surface mark was noted but excluded from this inspection."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing policy and scoring criteria for both contexts. Both contexts retain the Alderline unit, five-shelf build, unloaded stability test, 4 mm threshold, and inspection scope. The evidence consists of two complete factual sentences: \"Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001.\" and \"Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test.\" The counterfactual coherently changes the displacement from 6 mm to 3 mm without retaining a contradictory measurement. Neither context states a score, answer code, label rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified inventory, orientation, step, and fastener relations. The focus atom is the factual threshold relation that unloaded displacement exceeds 4 mm. Base and counter assignments differ only on that focus and are realizable: secure fasteners can coexist with displacement above 4 mm, and nonzero displacement can be 4 mm or less. Policy evidence cites the original state and preserves the threshold/scope statements; no additional state-origin rule is needed because the unchanged questions already retain the full governing criteria and exceptions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes level-0 hazards and level-1 defects, establishes complete and correct parts and steps, and supports unloaded wobble above 4 mm. Under the explicit criterion, wobble above 4 mm prevents a pass and is sufficient for level 2 even though accessible fasteners are secure.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete and correct parts and steps, secure accessible fasteners, no inaccessible fastening error, no level-0 hazard, and displacement not exceeding 4 mm. Observed nonzero displacement excludes level 4's no-observed-wobble requirement, while the stated facts satisfy level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every in-scope part required for Alderline five-shelf unit serial AL-001 is present in the unit's inspected parts inventory."}, {"id": "a2", "statement": "Every installed in-scope part on Alderline five-shelf unit serial AL-001 matches its specified part identifier."}, {"id": "a3", "statement": "Every installed structural part on Alderline five-shelf unit serial AL-001 has its specified orientation."}, {"id": "a4", "statement": "Every instruction step from step 1 through step 12 was performed for Alderline five-shelf unit serial AL-001."}, {"id": "a5", "statement": "Every performed instruction step from step 1 through step 12 for Alderline five-shelf unit serial AL-001 conforms to the corresponding instruction."}, {"id": "a6", "statement": "Alderline five-shelf unit serial AL-001 has structural damage."}, {"id": "a7", "statement": "Alderline five-shelf unit serial AL-001 has severe instability."}, {"id": "a8", "statement": "An assembly error in Alderline five-shelf unit serial AL-001 creates an immediate collapse or injury risk."}, {"id": "a9", "statement": "Every accessible fastener on Alderline five-shelf unit serial AL-001 is secure."}, {"id": "a10", "statement": "Alderline five-shelf unit serial AL-001 has an inaccessible fastening error."}, {"id": "a11", "statement": "The measured top-corner displacement during unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 exceeds 4 mm."}, {"id": "a12", "statement": "Unloaded stability test UT-01 of Alderline five-shelf unit serial AL-001 observed nonzero top-corner displacement."}], "base_state_json": "\"Case note for Alderline five-shelf unit serial AL-001: The parts inventory lists every in-scope panel, connector, and fastener required for the five-shelf build. A component audit found each installed part matched its specified identifier, and structural members were oriented as shown in the assembly record. The technician completed instruction steps 1 through 12 and recorded conformity at each step. The unit showed no structural damage, severe instability, or immediate collapse or injury risk. All accessible fasteners were secure, with no inaccessible fastening error identified. Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001. Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test. The test log confirms that top-corner movement was observed. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement. A rear-panel surface mark was noted but excluded from this inspection.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a11", "focus_evidence": [{"path": [], "text": "Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001."}, {"path": [], "text": "Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test."}], "policy_evidence": [{"path": [], "text": "The manual permits at most 4 mm."}, {"path": [], "text": "Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement."}], "rules": [{"justification": "All in-scope parts and steps are complete and correct, all fastening errors are excluded, and every level-0 hazard is explicitly refuted. The unloaded displacement nevertheless exceeds 4 mm, so the explicit exception prevents a passing level and requires level 2 tightening and retesting.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "All in-scope parts, steps, and fasteners satisfy the standard-pass requirements; level-0 and level-1 defects are excluded; and explicit refutation of displacement above 4 mm establishes displacement of 4 mm or less. Nonzero displacement was observed, excluding the level-4 requirement of no observed wobble, so level 3 is sufficient.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001.", "negative_left": "Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001.", "negative_right": "Inspection record ST-UT-41 records a top-corner displacement of 3 mm during the unloaded test.", "right": "Inspection record ST-UT-41 records a top-corner displacement of 6 mm during the unloaded test."}, "verifier_independent_model": false}, "family": "fast-41-diverse-197-041", "id": "fast-41-diverse-197-041-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or unusable: structural damage, severe instability, or an assembly error creates an immediate collapse or injury risk; stop use and disassemble or replace affected parts.", "1 — Major correction required: missing or wrongly oriented structural parts, skipped steps, or multiple inaccessible fastening errors require parts sorting or substantial step correction before retesting.", "2 — Minor correction required before placement: the build is otherwise complete and correctly assembled, but one or a few accessible fasteners need tightening, or unloaded wobble exceeds 4 mm; route to tightening and retest.", "3 — Standard pass: all in-scope parts and steps are correct, accessible fasteners are secure, unloaded wobble is 4 mm or less, and only excluded matters such as rear-only cosmetic marks or pending wall anchoring remain.", "4 — Excellent pass: meets level 3 and also has clean visible surfaces, consistently flush joints, even shelf alignment, and no observed wobble during the unloaded test."], "instructions": "Rate assembly readiness using the five ordered levels. Consider parts, step accuracy, fastening, and unloaded stability. Ignore wall anchoring and rear-only cosmetic marks under the stated scope. Apply the explicit exception that unloaded wobble above 4 mm prevents a passing level even when wall anchoring is not yet required. Select one level and identify the appropriate route implied by that level.", "type": "score"}}, "state": "Case note for Alderline five-shelf unit serial AL-001: The parts inventory lists every in-scope panel, connector, and fastener required for the five-shelf build. A component audit found each installed part matched its specified identifier, and structural members were oriented as shown in the assembly record. The technician completed instruction steps 1 through 12 and recorded conformity at each step. The unit showed no structural damage, severe instability, or immediate collapse or injury risk. All accessible fasteners were secure, with no inaccessible fastening error identified. Inspection record ST-UT-41 is the unloaded stability test UT-01 record for Alderline five-shelf unit serial AL-001. Inspection record ST-UT-41 records a top-corner displacement of 3 mm during the unloaded test. The test log confirms that top-corner movement was observed. The manual permits at most 4 mm. Wall anchoring and rear-only cosmetic marks are outside this inspection’s scope, except unloaded wobble above 4 mm must still be corrected before placement. A rear-panel surface mark was noted but excluded from this inspection."}, "method": "c2d", "provenance": {"source_id": "diverse-197", "source_is_synthetic": true, "source_sha256": "719cc9dcf23a914a3933b90d83a84b33e439f3b69f8f92cfafde9793ca501292", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged original questions object supplies the criteria and instructions, while both contexts retain the same cabinet, request, and date bindings. The two focus spans are complete factual sentences: “The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.” and “The same log records installation of the upper shelf at 09:25 local time on 15 September 2026.” The base context is factually incoherent because its timestamps show the required order while its tester assertion says wobbling persisted until that sequence was corrected. The counterfactual is coherent because the upper shelf at 08:55 precedes the rear crossbar at 09:10 and aligns with the stated sequence problem. Neither context embeds a gold answer, answer code, label rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relationship, including the universally quantified conformity and corrective-action relations. The focus atom is the factual installation order. The unchanged question already preserves the scoring criteria and highest-applicable instruction; the state-derived policy evidence additionally preserves the cabinet-specific required sequence, so policy completeness is satisfied. The base and counter assignments are realizable with only the order relation changing: both installation events can remain distinct while their order reverses, and the remaining conditions can stay fixed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the specified order, conformity of every other assembly feature, stability, tight hardware, safe positioning, level and clearance, and no other need for corrective action. It also excludes missing, wrong, damaged, or mixed-up required components. Atom a13 only states what a violation would require; with a1 supported, no such violation exists. Thus no level 1–4 condition applies and level 0 is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Given that both installations occurred at distinct times, refuting that the rear crossbar preceded the upper shelf entails the reverse order, violating the supplied sequence. Atom a13 establishes that this violation requires undoing and rebuilding, satisfying level 3. The component atoms exclude level 4, while level 3 outranks any lower condition.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "In the inspected Larkspur wall cabinet, installation of the rear crossbar occurred before installation of the upper shelf."}, {"id": "a2", "statement": "In the inspected Larkspur wall cabinet, the rear crossbar and upper shelf were installed at distinct times."}, {"id": "a3", "statement": "Every component required for the inspected Larkspur wall cabinet is available."}, {"id": "a4", "statement": "Every component selected for the inspected Larkspur wall cabinet is the correct specified component."}, {"id": "a5", "statement": "Every component required for the inspected Larkspur wall cabinet is undamaged."}, {"id": "a6", "statement": "No components for the inspected Larkspur wall cabinet are mixed up."}, {"id": "a7", "statement": "Every assembly feature of the inspected Larkspur wall cabinet other than the relative installation order of the rear crossbar and upper shelf matches the supplied instructions."}, {"id": "a8", "statement": "The inspected Larkspur wall cabinet is stable."}, {"id": "a9", "statement": "Every item of installed hardware in the inspected Larkspur wall cabinet is tight."}, {"id": "a10", "statement": "The inspected Larkspur wall cabinet is safely positioned."}, {"id": "a11", "statement": "The inspected Larkspur wall cabinet is level."}, {"id": "a12", "statement": "The inspected Larkspur wall cabinet has the required clearance."}, {"id": "a13", "statement": "A violation of the specified relative installation order of the rear crossbar and upper shelf in the inspected Larkspur wall cabinet requires that portion of the assembly to be undone and rebuilt."}, {"id": "a14", "statement": "No feature of the inspected Larkspur wall cabinet other than the relative installation order of the rear crossbar and upper shelf requires corrective action."}], "base_state_json": "{\"context\":\"A household assembler built a fictional Larkspur wall cabinet in the living room. Inspection routes range from a clean pass to parts problems. When several conditions appear, use the highest applicable route.\",\"evidence\":[\"The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.\",\"The same log records installation of the upper shelf at 09:25 local time on 15 September 2026.\",\"The instructions place the rear crossbar before the upper shelf. The final-use tester reports wobbling until that sequence is corrected.\",\"The parts inventory confirms that every required component is present, specified correctly, undamaged, and not mixed up.\",\"All other assembly features match the supplied instructions. The inspected cabinet is stable, safely positioned, level, and provided with the required clearance; every installed hardware item is tight.\",\"The inspection note says that any violation of the specified installation order requires that portion of the assembly to be undone and rebuilt, and that no other feature requires corrective action.\"],\"request\":\"Choose the inspection route and severity level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time."}, {"path": ["evidence", "1"], "text": "The same log records installation of the upper shelf at 09:25 local time on 15 September 2026."}], "policy_evidence": [{"path": ["context"], "text": "A household assembler built a fictional Larkspur wall cabinet in the living room. Inspection routes range from a clean pass to parts problems. When several conditions appear, use the highest applicable route."}, {"path": ["evidence", "2"], "text": "The instructions place the rear crossbar before the upper shelf. The final-use tester reports wobbling until that sequence is corrected."}], "rules": [{"justification": "The rear crossbar precedes the upper shelf as instructed; every other assembly feature matches the instructions; the cabinet is stable, properly placed, level, adequately cleared, and tight; and no corrective action is required. The exclusions for missing, wrong, damaged, or mixed-up parts also rule out the higher parts-sorting route.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "Because the two installations occurred at distinct times, refuting that the rear crossbar preceded the upper shelf entails that the upper shelf preceded the rear crossbar, contrary to the supplied sequence. That violation requires undoing and rebuilding. All required parts are available, correct, undamaged, and not mixed up, so level 4 is excluded; level 3 outranks any lower route.", "target": "3", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.", "negative_left": "The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.", "negative_right": "The same log records installation of the upper shelf at 08:55 local time on 15 September 2026.", "right": "The same log records installation of the upper shelf at 09:25 local time on 15 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-198-015", "id": "fast-41-diverse-198-015-base", "input": {"questions": {"decision": {"criteria": ["0 — Pass: Assembly matches the instructions, is stable, and needs no corrective action.", "1 — Final placement: Assembly is correct and stable, but safe positioning, leveling, or clearance adjustment remains.", "2 — Tightening: Correct parts and sequence were used, but loose installed hardware must be tightened.", "3 — Step correction: All required parts are available, but an incorrect sequence, orientation, or connection must be undone and rebuilt.", "4 — Parts sorting: A required component is missing, wrong, damaged, or mixed up and must be identified or replaced before rebuilding."], "instructions": "Resolve references in the notes, compare the evidence with the supplied assembly sequence, and select exactly one indexed level. Apply the highest applicable level.", "type": "score"}}, "state": {"context": "A household assembler built a fictional Larkspur wall cabinet in the living room. Inspection routes range from a clean pass to parts problems. When several conditions appear, use the highest applicable route.", "evidence": ["The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.", "The same log records installation of the upper shelf at 09:25 local time on 15 September 2026.", "The instructions place the rear crossbar before the upper shelf. The final-use tester reports wobbling until that sequence is corrected.", "The parts inventory confirms that every required component is present, specified correctly, undamaged, and not mixed up.", "All other assembly features match the supplied instructions. The inspected cabinet is stable, safely positioned, level, and provided with the required clearance; every installed hardware item is tight.", "The inspection note says that any violation of the specified installation order requires that portion of the assembly to be undone and rebuilt, and that no other feature requires corrective action."], "request": "Choose the inspection route and severity level."}}, "method": "c2d", "provenance": {"source_id": "diverse-198", "source_is_synthetic": true, "source_sha256": "7420a111a2090f69057ec276d78f13ce0713c0c572301b328ac2ae2d6f80beab", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged original questions object supplies the criteria and instructions, while both contexts retain the same cabinet, request, and date bindings. The two focus spans are complete factual sentences: “The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.” and “The same log records installation of the upper shelf at 09:25 local time on 15 September 2026.” The base context is factually incoherent because its timestamps show the required order while its tester assertion says wobbling persisted until that sequence was corrected. The counterfactual is coherent because the upper shelf at 08:55 precedes the rear crossbar at 09:10 and aligns with the stated sequence problem. Neither context embeds a gold answer, answer code, label rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relationship, including the universally quantified conformity and corrective-action relations. The focus atom is the factual installation order. The unchanged question already preserves the scoring criteria and highest-applicable instruction; the state-derived policy evidence additionally preserves the cabinet-specific required sequence, so policy completeness is satisfied. The base and counter assignments are realizable with only the order relation changing: both installation events can remain distinct while their order reverses, and the remaining conditions can stay fixed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the specified order, conformity of every other assembly feature, stability, tight hardware, safe positioning, level and clearance, and no other need for corrective action. It also excludes missing, wrong, damaged, or mixed-up required components. Atom a13 only states what a violation would require; with a1 supported, no such violation exists. Thus no level 1–4 condition applies and level 0 is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Given that both installations occurred at distinct times, refuting that the rear crossbar preceded the upper shelf entails the reverse order, violating the supplied sequence. Atom a13 establishes that this violation requires undoing and rebuilding, satisfying level 3. The component atoms exclude level 4, while level 3 outranks any lower condition.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "In the inspected Larkspur wall cabinet, installation of the rear crossbar occurred before installation of the upper shelf."}, {"id": "a2", "statement": "In the inspected Larkspur wall cabinet, the rear crossbar and upper shelf were installed at distinct times."}, {"id": "a3", "statement": "Every component required for the inspected Larkspur wall cabinet is available."}, {"id": "a4", "statement": "Every component selected for the inspected Larkspur wall cabinet is the correct specified component."}, {"id": "a5", "statement": "Every component required for the inspected Larkspur wall cabinet is undamaged."}, {"id": "a6", "statement": "No components for the inspected Larkspur wall cabinet are mixed up."}, {"id": "a7", "statement": "Every assembly feature of the inspected Larkspur wall cabinet other than the relative installation order of the rear crossbar and upper shelf matches the supplied instructions."}, {"id": "a8", "statement": "The inspected Larkspur wall cabinet is stable."}, {"id": "a9", "statement": "Every item of installed hardware in the inspected Larkspur wall cabinet is tight."}, {"id": "a10", "statement": "The inspected Larkspur wall cabinet is safely positioned."}, {"id": "a11", "statement": "The inspected Larkspur wall cabinet is level."}, {"id": "a12", "statement": "The inspected Larkspur wall cabinet has the required clearance."}, {"id": "a13", "statement": "A violation of the specified relative installation order of the rear crossbar and upper shelf in the inspected Larkspur wall cabinet requires that portion of the assembly to be undone and rebuilt."}, {"id": "a14", "statement": "No feature of the inspected Larkspur wall cabinet other than the relative installation order of the rear crossbar and upper shelf requires corrective action."}], "base_state_json": "{\"context\":\"A household assembler built a fictional Larkspur wall cabinet in the living room. Inspection routes range from a clean pass to parts problems. When several conditions appear, use the highest applicable route.\",\"evidence\":[\"The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.\",\"The same log records installation of the upper shelf at 09:25 local time on 15 September 2026.\",\"The instructions place the rear crossbar before the upper shelf. The final-use tester reports wobbling until that sequence is corrected.\",\"The parts inventory confirms that every required component is present, specified correctly, undamaged, and not mixed up.\",\"All other assembly features match the supplied instructions. The inspected cabinet is stable, safely positioned, level, and provided with the required clearance; every installed hardware item is tight.\",\"The inspection note says that any violation of the specified installation order requires that portion of the assembly to be undone and rebuilt, and that no other feature requires corrective action.\"],\"request\":\"Choose the inspection route and severity level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time."}, {"path": ["evidence", "1"], "text": "The same log records installation of the upper shelf at 09:25 local time on 15 September 2026."}], "policy_evidence": [{"path": ["context"], "text": "A household assembler built a fictional Larkspur wall cabinet in the living room. Inspection routes range from a clean pass to parts problems. When several conditions appear, use the highest applicable route."}, {"path": ["evidence", "2"], "text": "The instructions place the rear crossbar before the upper shelf. The final-use tester reports wobbling until that sequence is corrected."}], "rules": [{"justification": "The rear crossbar precedes the upper shelf as instructed; every other assembly feature matches the instructions; the cabinet is stable, properly placed, level, adequately cleared, and tight; and no corrective action is required. The exclusions for missing, wrong, damaged, or mixed-up parts also rule out the higher parts-sorting route.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}, {"justification": "Because the two installations occurred at distinct times, refuting that the rear crossbar preceded the upper shelf entails that the upper shelf preceded the rear crossbar, contrary to the supplied sequence. That violation requires undoing and rebuilding. All required parts are available, correct, undamaged, and not mixed up, so level 4 is excluded; level 3 outranks any lower route.", "target": "3", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}]}]}, "verified_pair": {"left": "The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.", "negative_left": "The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.", "negative_right": "The same log records installation of the upper shelf at 08:55 local time on 15 September 2026.", "right": "The same log records installation of the upper shelf at 09:25 local time on 15 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-198-015", "id": "fast-41-diverse-198-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Pass: Assembly matches the instructions, is stable, and needs no corrective action.", "1 — Final placement: Assembly is correct and stable, but safe positioning, leveling, or clearance adjustment remains.", "2 — Tightening: Correct parts and sequence were used, but loose installed hardware must be tightened.", "3 — Step correction: All required parts are available, but an incorrect sequence, orientation, or connection must be undone and rebuilt.", "4 — Parts sorting: A required component is missing, wrong, damaged, or mixed up and must be identified or replaced before rebuilding."], "instructions": "Resolve references in the notes, compare the evidence with the supplied assembly sequence, and select exactly one indexed level. Apply the highest applicable level.", "type": "score"}}, "state": {"context": "A household assembler built a fictional Larkspur wall cabinet in the living room. Inspection routes range from a clean pass to parts problems. When several conditions appear, use the highest applicable route.", "evidence": ["The 15 September 2026 assembly log for the inspected Larkspur wall cabinet records installation of the rear crossbar at 09:10 local time.", "The same log records installation of the upper shelf at 08:55 local time on 15 September 2026.", "The instructions place the rear crossbar before the upper shelf. The final-use tester reports wobbling until that sequence is corrected.", "The parts inventory confirms that every required component is present, specified correctly, undamaged, and not mixed up.", "All other assembly features match the supplied instructions. The inspected cabinet is stable, safely positioned, level, and provided with the required clearance; every installed hardware item is tight.", "The inspection note says that any violation of the specified installation order requires that portion of the assembly to be undone and rebuilt, and that no other feature requires corrective action."], "request": "Choose the inspection route and severity level."}}, "method": "c2d", "provenance": {"source_id": "diverse-198", "source_is_synthetic": true, "source_sha256": "7420a111a2090f69057ec276d78f13ce0713c0c572301b328ac2ae2d6f80beab", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, and both contexts preserve the 3 mm tolerance without adding policy. Priya’s 45 cm cushion cover and the relevant seam and timestamp bindings remain intact while observations change. The evidence contains exactly two complete factual sentences: “At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge.” and “At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge.” The counterfactual is coherent because its 42 mm and 40 mm measurements differ by 2 mm without duplicating contradictory assertions. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"At the latest review of Priya’s 45 cm cushion cover, the measurement log confirms that every required dimension was verified, and the cutting checklist confirms that every required cut piece is present. The construction checklist confirms that every required seam is present and that all other construction requirements passed. The project sheet allows no more than 3 mm seam deviation. At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge. At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge. Later checks recorded trimming passed, pressing passed, and final inspection passed. No other issues were recorded for the cover.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge."}, {"path": [], "text": "At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge.", "negative_left": "At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge.", "negative_right": "At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 40 mm from the cover’s left edge.", "right": "At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge."}, "verifier_independent_model": false}, "family": "fast-41-diverse-199-001", "id": "fast-41-diverse-199-001-base", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "At the latest review of Priya’s 45 cm cushion cover, the measurement log confirms that every required dimension was verified, and the cutting checklist confirms that every required cut piece is present. The construction checklist confirms that every required seam is present and that all other construction requirements passed. The project sheet allows no more than 3 mm seam deviation. At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge. At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge. Later checks recorded trimming passed, pressing passed, and final inspection passed. No other issues were recorded for the cover."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, and both contexts preserve the 3 mm tolerance without adding policy. Priya’s 45 cm cushion cover and the relevant seam and timestamp bindings remain intact while observations change. The evidence contains exactly two complete factual sentences: “At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge.” and “At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge.” The counterfactual is coherent because its 42 mm and 40 mm measurements differ by 2 mm without duplicating contradictory assertions. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"At the latest review of Priya’s 45 cm cushion cover, the measurement log confirms that every required dimension was verified, and the cutting checklist confirms that every required cut piece is present. The construction checklist confirms that every required seam is present and that all other construction requirements passed. The project sheet allows no more than 3 mm seam deviation. At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge. At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge. Later checks recorded trimming passed, pressing passed, and final inspection passed. No other issues were recorded for the cover.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge."}, {"path": [], "text": "At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge.", "negative_left": "At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge.", "negative_right": "At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 40 mm from the cover’s left edge.", "right": "At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 37 mm from the cover’s left edge."}, "verifier_independent_model": false}, "family": "fast-41-diverse-199-001", "id": "fast-41-diverse-199-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "At the latest review of Priya’s 45 cm cushion cover, the measurement log confirms that every required dimension was verified, and the cutting checklist confirms that every required cut piece is present. The construction checklist confirms that every required seam is present and that all other construction requirements passed. The project sheet allows no more than 3 mm seam deviation. At 2026-09-17T10:00Z, the measured location of the existing right-edge seam on Priya’s 45 cm cushion cover was 42 mm from the cover’s left edge. At 2026-09-17T10:05Z, the recorded location of the marked line for that right-edge seam on Priya’s 45 cm cushion cover was 40 mm from the cover’s left edge. Later checks recorded trimming passed, pressing passed, and final inspection passed. No other issues were recorded for the cover."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_clean_4"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, both contexts retain the Priya 45 cm cover and latest-timestamp bindings, the evidence consists of the two complete factual sentences “The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge.” and “The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge.”, and the counterfactual’s 125 mm value is coherent because it differs from 128 mm by exactly the allowed 3 mm without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"Priya’s 45 cm cushion cover was assessed from the latest timestamped records. All required dimensions were verified, every required cut piece was present, and every required construction seam was present. All construction requirements other than the specified seam-location check passed. Trimming, pressing, and final inspection also passed. The project sheet allows no more than 3 mm seam deviation. The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge. The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge. These records concern the existing seam and its corresponding marked line on the same cover.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge."}, {"path": [], "text": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge.", "negative_left": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge.", "negative_right": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 125 mm from the left edge.", "right": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge."}, "verifier_independent_model": false}, "family": "fast-41-diverse-199-009", "id": "fast-41-diverse-199-009-base", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "Priya’s 45 cm cushion cover was assessed from the latest timestamped records. All required dimensions were verified, every required cut piece was present, and every required construction seam was present. All construction requirements other than the specified seam-location check passed. Trimming, pressing, and final inspection also passed. The project sheet allows no more than 3 mm seam deviation. The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge. The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge. These records concern the existing seam and its corresponding marked line on the same cover."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, both contexts retain the Priya 45 cm cover and latest-timestamp bindings, the evidence consists of the two complete factual sentences “The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge.” and “The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge.”, and the counterfactual’s 125 mm value is coherent because it differs from 128 mm by exactly the allowed 3 mm without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"Priya’s 45 cm cushion cover was assessed from the latest timestamped records. All required dimensions were verified, every required cut piece was present, and every required construction seam was present. All construction requirements other than the specified seam-location check passed. Trimming, pressing, and final inspection also passed. The project sheet allows no more than 3 mm seam deviation. The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge. The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge. These records concern the existing seam and its corresponding marked line on the same cover.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge."}, {"path": [], "text": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge.", "negative_left": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge.", "negative_right": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 125 mm from the left edge.", "right": "The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 124 mm from the left edge."}, "verifier_independent_model": false}, "family": "fast-41-diverse-199-009", "id": "fast-41-diverse-199-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "Priya’s 45 cm cushion cover was assessed from the latest timestamped records. All required dimensions were verified, every required cut piece was present, and every required construction seam was present. All construction requirements other than the specified seam-location check passed. Trimming, pressing, and final inspection also passed. The project sheet allows no more than 3 mm seam deviation. The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, measured the existing right-edge seam at 128 mm from the left edge. The latest timestamped inspection of Priya’s 45 cm cushion cover, recorded on 2026-09-10 at 14:00 UTC, recorded the marked line for that right-edge seam at 125 mm from the left edge. These records concern the existing seam and its corresponding marked line on the same cover."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_clean_4"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and neither context adds exceptions or defaults. Entity and latest-evidence bindings remain Priya’s 45 cm cushion cover and its right-edge seam. The evidence consists of two complete factual sentences: “The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge.” and “The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge.” The counterfactual coherently changes the marked-line measurement from 123 mm to 126 mm without creating contradictory duplicate measurements. Neither context states an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"At the latest review, Priya’s 45 cm cushion cover had verified dimensions, all required cut pieces, and every required construction seam. The remaining construction checks, excluding seam presence and the specified location check, had passed. Trimming, pressing, and final inspection were also recorded as passed. The project sheet allows no more than 3 mm seam deviation. The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge. The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge. All records concern the same cover and the latest available evidence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge."}, {"path": [], "text": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge.", "negative_left": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge.", "negative_right": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 126 mm from the cover’s left edge.", "right": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge."}, "verifier_independent_model": false}, "family": "fast-41-diverse-199-011", "id": "fast-41-diverse-199-011-base", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "At the latest review, Priya’s 45 cm cushion cover had verified dimensions, all required cut pieces, and every required construction seam. The remaining construction checks, excluding seam presence and the specified location check, had passed. Trimming, pressing, and final inspection were also recorded as passed. The project sheet allows no more than 3 mm seam deviation. The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge. The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge. All records concern the same cover and the latest available evidence."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and neither context adds exceptions or defaults. Entity and latest-evidence bindings remain Priya’s 45 cm cushion cover and its right-edge seam. The evidence consists of two complete factual sentences: “The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge.” and “The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge.” The counterfactual coherently changes the marked-line measurement from 123 mm to 126 mm without creating contradictory duplicate measurements. Neither context states an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"At the latest review, Priya’s 45 cm cushion cover had verified dimensions, all required cut pieces, and every required construction seam. The remaining construction checks, excluding seam presence and the specified location check, had passed. Trimming, pressing, and final inspection were also recorded as passed. The project sheet allows no more than 3 mm seam deviation. The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge. The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge. All records concern the same cover and the latest available evidence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge."}, {"path": [], "text": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge.", "negative_left": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge.", "negative_right": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 126 mm from the cover’s left edge.", "right": "The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 123 mm from the cover’s left edge."}, "verifier_independent_model": false}, "family": "fast-41-diverse-199-011", "id": "fast-41-diverse-199-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "At the latest review, Priya’s 45 cm cushion cover had verified dimensions, all required cut pieces, and every required construction seam. The remaining construction checks, excluding seam presence and the specified location check, had passed. Trimming, pressing, and final inspection were also recorded as passed. The project sheet allows no more than 3 mm seam deviation. The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the measured location of the existing right-edge seam as 128 mm from the cover’s left edge. The latest timestamped inspection record for Priya’s 45 cm cushion cover, dated 2026-09-17T10:00:00Z, lists the recorded location of the right-edge seam’s marked line as 126 mm from the cover’s left edge. All records concern the same cover and the latest available evidence."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_clean_4"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and the unchanged questions preserve the same requirements and request. Mara, Lena, the cushion cover, and all question bindings remain unchanged. The evidence quotes are “The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2.” and “The registry records identifier C-417 for Mara's submitted zippered cushion cover.” Changing the submitted cover's identifier to C-982 remains coherent because C-417 remains exclusive to IMG-2's distinct object. Neither context states a gold answer, label code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At a household craft station, Mara submitted a zippered cushion cover for Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion. Direct inspection measured the finished width and height within 0.5 cm of 45.0 cm. Inspectors found every corner square, every thread trimmed, and the cover pressed. The cosmetic-defect count on the submitted cover is zero. Inspection records state that every inspected portion of its zipper seam outside the object shown in IMG-2 is even. IMG-2 shows a puckered zipper seam. The construction-defect count on the submitted cover outside any zipper-seam portion shown in IMG-2 is zero. The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2. The registry records identifier C-417 for Mara's submitted zippered cushion cover.\",\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\",\"policy\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2."}, {"path": ["context"], "text": "The registry records identifier C-417 for Mara's submitted zippered cushion cover."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2.", "negative_left": "The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2.", "negative_right": "The registry records identifier C-982 for Mara's submitted zippered cushion cover.", "right": "The registry records identifier C-417 for Mara's submitted zippered cushion cover."}, "verifier_independent_model": false}, "family": "fast-41-diverse-200-001", "id": "fast-41-diverse-200-001-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At a household craft station, Mara submitted a zippered cushion cover for Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion. Direct inspection measured the finished width and height within 0.5 cm of 45.0 cm. Inspectors found every corner square, every thread trimmed, and the cover pressed. The cosmetic-defect count on the submitted cover is zero. Inspection records state that every inspected portion of its zipper seam outside the object shown in IMG-2 is even. IMG-2 shows a puckered zipper seam. The construction-defect count on the submitted cover outside any zipper-seam portion shown in IMG-2 is zero. The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2. The registry records identifier C-417 for Mara's submitted zippered cushion cover.", "policy": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and the unchanged questions preserve the same requirements and request. Mara, Lena, the cushion cover, and all question bindings remain unchanged. The evidence quotes are “The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2.” and “The registry records identifier C-417 for Mara's submitted zippered cushion cover.” Changing the submitted cover's identifier to C-982 remains coherent because C-417 remains exclusive to IMG-2's distinct object. Neither context states a gold answer, label code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At a household craft station, Mara submitted a zippered cushion cover for Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion. Direct inspection measured the finished width and height within 0.5 cm of 45.0 cm. Inspectors found every corner square, every thread trimmed, and the cover pressed. The cosmetic-defect count on the submitted cover is zero. Inspection records state that every inspected portion of its zipper seam outside the object shown in IMG-2 is even. IMG-2 shows a puckered zipper seam. The construction-defect count on the submitted cover outside any zipper-seam portion shown in IMG-2 is zero. The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2. The registry records identifier C-417 for Mara's submitted zippered cushion cover.\",\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\",\"policy\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2."}, {"path": ["context"], "text": "The registry records identifier C-417 for Mara's submitted zippered cushion cover."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2.", "negative_left": "The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2.", "negative_right": "The registry records identifier C-982 for Mara's submitted zippered cushion cover.", "right": "The registry records identifier C-417 for Mara's submitted zippered cushion cover."}, "verifier_independent_model": false}, "family": "fast-41-diverse-200-001", "id": "fast-41-diverse-200-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At a household craft station, Mara submitted a zippered cushion cover for Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion. Direct inspection measured the finished width and height within 0.5 cm of 45.0 cm. Inspectors found every corner square, every thread trimmed, and the cover pressed. The cosmetic-defect count on the submitted cover is zero. Inspection records state that every inspected portion of its zipper seam outside the object shown in IMG-2 is even. IMG-2 shows a puckered zipper seam. The construction-defect count on the submitted cover outside any zipper-seam portion shown in IMG-2 is zero. The registry assigns different identifiers to different physical objects and assigns identifier C-417 exclusively to the physical object shown in inspection image IMG-2. The registry records identifier C-982 for Mara's submitted zippered cushion cover.", "policy": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and request, and both contexts repeat the same policy. Entity, request, and temporal bindings remain consistent. The two evidence spans are complete factual sentences: \"At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417.\" and \"At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417.\" The counterfactual coherently distinguishes IMG-2 from Mara's submitted cover by changing only the latter identifier to MRC-902. Neither context contains an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"Case note, 2026-09-17. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion. Direct inspection measured Mara's submitted zippered cushion cover at 45.2 × 44.8 cm. Every corner was square, every thread was trimmed, and the cover had been pressed. Direct inspection found zero cosmetic defects. Every inspected portion of the submitted cover's zipper seam not part of the object shown in IMG-2 was even. IMG-2 showed a puckered zipper seam. The number of construction defects on Mara's submitted cover outside any zipper-seam portion shown in IMG-2 was zero. At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417. At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417. Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\",\"policy\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417."}, {"path": ["context"], "text": "At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417.", "negative_left": "At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417.", "negative_right": "At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-902.", "right": "At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-200-015", "id": "fast-41-diverse-200-015-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "Case note, 2026-09-17. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion. Direct inspection measured Mara's submitted zippered cushion cover at 45.2 × 44.8 cm. Every corner was square, every thread was trimmed, and the cover had been pressed. Direct inspection found zero cosmetic defects. Every inspected portion of the submitted cover's zipper seam not part of the object shown in IMG-2 was even. IMG-2 showed a puckered zipper seam. The number of construction defects on Mara's submitted cover outside any zipper-seam portion shown in IMG-2 was zero. At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417. At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417. Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.", "policy": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and request, and both contexts repeat the same policy. Entity, request, and temporal bindings remain consistent. The two evidence spans are complete factual sentences: \"At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417.\" and \"At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417.\" The counterfactual coherently distinguishes IMG-2 from Mara's submitted cover by changing only the latter identifier to MRC-902. Neither context contains an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"Case note, 2026-09-17. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion. Direct inspection measured Mara's submitted zippered cushion cover at 45.2 × 44.8 cm. Every corner was square, every thread was trimmed, and the cover had been pressed. Direct inspection found zero cosmetic defects. Every inspected portion of the submitted cover's zipper seam not part of the object shown in IMG-2 was even. IMG-2 showed a puckered zipper seam. The number of construction defects on Mara's submitted cover outside any zipper-seam portion shown in IMG-2 was zero. At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417. At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417. Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\",\"policy\":\"The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417."}, {"path": ["context"], "text": "At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417.", "negative_left": "At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417.", "negative_right": "At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-902.", "right": "At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-417."}, "verifier_independent_model": false}, "family": "fast-41-diverse-200-015", "id": "fast-41-diverse-200-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "Case note, 2026-09-17. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion. Direct inspection measured Mara's submitted zippered cushion cover at 45.2 × 44.8 cm. Every corner was square, every thread was trimmed, and the cover had been pressed. Direct inspection found zero cosmetic defects. Every inspected portion of the submitted cover's zipper seam not part of the object shown in IMG-2 was even. IMG-2 showed a puckered zipper seam. The number of construction defects on Mara's submitted cover outside any zipper-seam portion shown in IMG-2 was zero. At the 2026-09-17 inspection, the unique object identifier recorded for the object shown in IMG-2 was MRC-417. At the 2026-09-17 inspection, the unique object identifier recorded for Mara's submitted zippered cushion cover was MRC-902. Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.", "policy": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both full inputs because the unchanged original questions and both contexts retain the governing scope and quality rules. Question bindings remain unchanged for the cushion-cover project, its two covers, and the 40 cm measurement target. The evidence consists of two complete factual sentences: “Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm.” “The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively.” The counterfactual is coherent because only CC-17’s measurement changes to 40.8 cm, with no contradictory duplicate assertion. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm. The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Inez tested the envelope closures on CC-17 and CC-18, and both opened and closed normally. Her inspection log marks every loose thread as trimmed and every visible seam as pressed. The photographs show one visible topstitch wobble across the pair, and the gauge records it at 3 mm. No other visible topstitch wobble was recorded, and the only unfinished edges are the approved hidden inner raw edges.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm."}, {"path": [], "text": "The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm.", "negative_left": "Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm.", "negative_right": "The recorded measurements for CC-17 and CC-18 are 40.8 cm and 39.8 cm, respectively.", "right": "The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively."}, "verifier_independent_model": false}, "family": "fast-41-diverse-202-010", "id": "fast-41-diverse-202-010-base", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm. The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Inez tested the envelope closures on CC-17 and CC-18, and both opened and closed normally. Her inspection log marks every loose thread as trimmed and every visible seam as pressed. The photographs show one visible topstitch wobble across the pair, and the gauge records it at 3 mm. No other visible topstitch wobble was recorded, and the only unfinished edges are the approved hidden inner raw edges."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both full inputs because the unchanged original questions and both contexts retain the governing scope and quality rules. Question bindings remain unchanged for the cushion-cover project, its two covers, and the 40 cm measurement target. The evidence consists of two complete factual sentences: “Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm.” “The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively.” The counterfactual is coherent because only CC-17’s measurement changes to 40.8 cm, with no contradictory duplicate assertion. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm. The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Inez tested the envelope closures on CC-17 and CC-18, and both opened and closed normally. Her inspection log marks every loose thread as trimmed and every visible seam as pressed. The photographs show one visible topstitch wobble across the pair, and the gauge records it at 3 mm. No other visible topstitch wobble was recorded, and the only unfinished edges are the approved hidden inner raw edges.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm."}, {"path": [], "text": "The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm.", "negative_left": "Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm.", "negative_right": "The recorded measurements for CC-17 and CC-18 are 40.8 cm and 39.8 cm, respectively.", "right": "The recorded measurements for CC-17 and CC-18 are 40.3 cm and 39.8 cm, respectively."}, "verifier_independent_model": false}, "family": "fast-41-diverse-202-010", "id": "fast-41-diverse-202-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "Mara’s project scope contains exactly two cushion covers, labeled CC-17 and CC-18, and their target measurement is 40 cm. The recorded measurements for CC-17 and CC-18 are 40.8 cm and 39.8 cm, respectively. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Inez tested the envelope closures on CC-17 and CC-18, and both opened and closed normally. Her inspection log marks every loose thread as trimmed and every visible seam as pressed. The photographs show one visible topstitch wobble across the pair, and the gauge records it at 3 mm. No other visible topstitch wobble was recorded, and the only unfinished edges are the approved hidden inner raw edges."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question, governing scope and quality policy, relevant bindings, complete factual evidence sentences, and coherent measurements without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"At Mara’s household craft station, recipient Leo’s project record concerns exactly two cushion covers, MC-17 and MC-18. On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm. On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.6 cm. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Inez confirms both envelope closures work, every loose thread is trimmed, and every visible seam is pressed. Her inspection records one visible topstitch wobble measuring 3 mm across the pair, while the hidden inner raw edges remain unfinished under the approved exception.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm."}, {"path": [], "text": "On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.6 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm.", "negative_left": "On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm.", "negative_right": "On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.2 cm.", "right": "On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.6 cm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-202-017", "id": "fast-41-diverse-202-017-base", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "At Mara’s household craft station, recipient Leo’s project record concerns exactly two cushion covers, MC-17 and MC-18. On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm. On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.6 cm. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Inez confirms both envelope closures work, every loose thread is trimmed, and every visible seam is pressed. Her inspection records one visible topstitch wobble measuring 3 mm across the pair, while the hidden inner raw edges remain unfinished under the approved exception."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged question, governing scope and quality policy, relevant bindings, complete factual evidence sentences, and coherent measurements without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"At Mara’s household craft station, recipient Leo’s project record concerns exactly two cushion covers, MC-17 and MC-18. On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm. On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.6 cm. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Inez confirms both envelope closures work, every loose thread is trimmed, and every visible seam is pressed. Her inspection records one visible topstitch wobble measuring 3 mm across the pair, while the hidden inner raw edges remain unfinished under the approved exception.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm."}, {"path": [], "text": "On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.6 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm.", "negative_left": "On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm.", "negative_right": "On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.2 cm.", "right": "On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.6 cm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-202-017", "id": "fast-41-diverse-202-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "At Mara’s household craft station, recipient Leo’s project record concerns exactly two cushion covers, MC-17 and MC-18. On 2026-09-17, Mara’s stated project scope contains exactly two cushion covers, MC-17 and MC-18, and the recorded measurement of MC-17 is 40.2 cm. On 2026-09-17, the recorded measurement of cushion cover MC-18 in Mara’s stated project scope is 39.2 cm. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Inez confirms both envelope closures work, every loose thread is trimmed, and every visible seam is pressed. Her inspection records one visible topstitch wobble measuring 3 mm across the pair, while the hidden inner raw edges remain unfinished under the approved exception."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric and route rules. Both contexts address the blue tote and its project sheet without changing the request scope or entity bindings. Each context has two complete factual evidence sentences, including “Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.” and “The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote.” in the base context. The counterfactual instead coherently states that the physical seam is S-17 while the required seam is S-18, without duplicating contradictory measurements or assertions. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal claims over an explicit class remain atomic. The focus atom concerns whether a particular physical seam is project-sheet-required, not a policy classification. The base and counter assignments are jointly realizable while changing only a1: the opened seam can be required in the base and an additional, nonrequired seam in the counter. The policy evidence correctly preserves the state-originating project-sheet requirements needed to interpret the unchanged rubric; observations and the generic request need not be copied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 and a2 together entail that a project-sheet-required seam has a positive-length opening. Under the ordered rubric, that is sufficient for score 0 and the seam-correction or stitching route, regardless of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "With a1 refuted, the opened inspected seam is not required. Because a3 then covers every required seam, and a4 establishes that every required non-seam component is secure, score 0 is excluded. Refuted a6 establishes unfinished pressing, while the other finish and measurement conditions hold, making score 1 and the finishing route sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The physical seam identified in blue-tote inspection record IR-1 at 09:00 on 17 September 2026 is one of the blue tote’s seams required by its project sheet."}, {"id": "a2", "statement": "The physical seam identified in blue-tote inspection record IR-1 at 09:00 on 17 September 2026 has an opening longer than 0 cm."}, {"id": "a3", "statement": "At 09:00 on 17 September 2026, every project-sheet-required seam of the blue tote other than the physical seam identified in inspection record IR-1 is secure."}, {"id": "a4", "statement": "At 09:00 on 17 September 2026, every project-sheet-required non-seam component of the blue tote is secure."}, {"id": "a5", "statement": "At 09:00 on 17 September 2026, the measured dimensions of every blue-tote panel match the dimensions specified for that panel by the blue tote’s project sheet."}, {"id": "a6", "statement": "At 09:00 on 17 September 2026, pressing of the blue tote is finished."}, {"id": "a7", "statement": "At 09:00 on 17 September 2026, topstitching of the blue tote is finished."}, {"id": "a8", "statement": "At 09:00 on 17 September 2026, thread cleanup on the blue tote is finished."}], "base_state_json": "{\"context\":\"Case note: At 09:00 on 17 September 2026, IR-1 records a blue tote inspection. The blue tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching. The inspected opening measures 4 cm. All other project-sheet-required seams are secure, every required non-seam component is secure, and both panel measurements match the project sheet. Topstitching and thread cleanup are finished, but pressing is not finished. The tote is being held for the next work route.\\n\\nThe tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching.\",\"evidence\":[\"Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.\",\"The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17."}, {"path": ["evidence", "1"], "text": "The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote."}], "policy_evidence": [{"path": ["context"], "text": "The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching."}], "rules": [{"justification": "The inspected physical seam is a project-sheet-required seam and has a positive-length opening. A required seam is therefore open, which is sufficient for level 0 and routing to seam correction or stitching.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "The seam having the opening is not a project-sheet-required seam, while every required seam is secure and every required non-seam component is secure. Panel measurements match, topstitching and thread cleanup are finished, but pressing remains unfinished. The tote is therefore complete but has a finishing issue, which is sufficient for level 1 and routing to finishing.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.", "negative_left": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.", "negative_right": "The blue tote’s project sheet lists seam ID S-18, and no other seam ID, among the seams required for the blue tote.", "right": "The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote."}, "verifier_independent_model": false}, "family": "fast-41-diverse-203-001", "id": "fast-41-diverse-203-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not complete: A required seam is open, hardware or handles are insecure, or a required component is missing. Route to seam correction or stitching before finishing.", "1 — Complete but needs finishing: All required seams and components are secure, but minor issues such as unpressed fabric, uneven topstitching, or loose thread ends remain. Route to finishing.", "2 — Complete with good finish: Measurements match the project sheet, all required seams and components are secure, and pressing, topstitching, and thread cleanup are finished. Release to the project recipient."], "instructions": "Use the ordered rubric. Resolve pronouns and possessives from context, verify the project-sheet requirements against the evidence, and choose one level. The route specified by the chosen level is the next station.", "type": "score"}}, "state": {"context": "Case note: At 09:00 on 17 September 2026, IR-1 records a blue tote inspection. The blue tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching. The inspected opening measures 4 cm. All other project-sheet-required seams are secure, every required non-seam component is secure, and both panel measurements match the project sheet. Topstitching and thread cleanup are finished, but pressing is not finished. The tote is being held for the next work route.\n\nThe tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching.", "evidence": ["Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.", "The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote."]}}, "method": "c2d", "provenance": {"source_id": "diverse-203", "source_is_synthetic": true, "source_sha256": "62b02fcb4a0e58796e32491572ec31e832a7aba2391247dcb435399a4ede249f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the complete governing rubric and route rules. Both contexts address the blue tote and its project sheet without changing the request scope or entity bindings. Each context has two complete factual evidence sentences, including “Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.” and “The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote.” in the base context. The counterfactual instead coherently states that the physical seam is S-17 while the required seam is S-18, without duplicating contradictory measurements or assertions. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal claims over an explicit class remain atomic. The focus atom concerns whether a particular physical seam is project-sheet-required, not a policy classification. The base and counter assignments are jointly realizable while changing only a1: the opened seam can be required in the base and an additional, nonrequired seam in the counter. The policy evidence correctly preserves the state-originating project-sheet requirements needed to interpret the unchanged rubric; observations and the generic request need not be copied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 and a2 together entail that a project-sheet-required seam has a positive-length opening. Under the ordered rubric, that is sufficient for score 0 and the seam-correction or stitching route, regardless of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "With a1 refuted, the opened inspected seam is not required. Because a3 then covers every required seam, and a4 establishes that every required non-seam component is secure, score 0 is excluded. Refuted a6 establishes unfinished pressing, while the other finish and measurement conditions hold, making score 1 and the finishing route sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The physical seam identified in blue-tote inspection record IR-1 at 09:00 on 17 September 2026 is one of the blue tote’s seams required by its project sheet."}, {"id": "a2", "statement": "The physical seam identified in blue-tote inspection record IR-1 at 09:00 on 17 September 2026 has an opening longer than 0 cm."}, {"id": "a3", "statement": "At 09:00 on 17 September 2026, every project-sheet-required seam of the blue tote other than the physical seam identified in inspection record IR-1 is secure."}, {"id": "a4", "statement": "At 09:00 on 17 September 2026, every project-sheet-required non-seam component of the blue tote is secure."}, {"id": "a5", "statement": "At 09:00 on 17 September 2026, the measured dimensions of every blue-tote panel match the dimensions specified for that panel by the blue tote’s project sheet."}, {"id": "a6", "statement": "At 09:00 on 17 September 2026, pressing of the blue tote is finished."}, {"id": "a7", "statement": "At 09:00 on 17 September 2026, topstitching of the blue tote is finished."}, {"id": "a8", "statement": "At 09:00 on 17 September 2026, thread cleanup on the blue tote is finished."}], "base_state_json": "{\"context\":\"Case note: At 09:00 on 17 September 2026, IR-1 records a blue tote inspection. The blue tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching. The inspected opening measures 4 cm. All other project-sheet-required seams are secure, every required non-seam component is secure, and both panel measurements match the project sheet. Topstitching and thread cleanup are finished, but pressing is not finished. The tote is being held for the next work route.\\n\\nThe tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching.\",\"evidence\":[\"Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.\",\"The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17."}, {"path": ["evidence", "1"], "text": "The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote."}], "policy_evidence": [{"path": ["context"], "text": "The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching."}], "rules": [{"justification": "The inspected physical seam is a project-sheet-required seam and has a positive-length opening. A required seam is therefore open, which is sufficient for level 0 and routing to seam correction or stitching.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "The seam having the opening is not a project-sheet-required seam, while every required seam is secure and every required non-seam component is secure. Panel measurements match, topstitching and thread cleanup are finished, but pressing remains unfinished. The tote is therefore complete but has a finishing issue, which is sufficient for level 1 and routing to finishing.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.", "negative_left": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.", "negative_right": "The blue tote’s project sheet lists seam ID S-18, and no other seam ID, among the seams required for the blue tote.", "right": "The blue tote’s project sheet lists seam ID S-17, and no other seam ID, among the seams required for the blue tote."}, "verifier_independent_model": false}, "family": "fast-41-diverse-203-001", "id": "fast-41-diverse-203-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not complete: A required seam is open, hardware or handles are insecure, or a required component is missing. Route to seam correction or stitching before finishing.", "1 — Complete but needs finishing: All required seams and components are secure, but minor issues such as unpressed fabric, uneven topstitching, or loose thread ends remain. Route to finishing.", "2 — Complete with good finish: Measurements match the project sheet, all required seams and components are secure, and pressing, topstitching, and thread cleanup are finished. Release to the project recipient."], "instructions": "Use the ordered rubric. Resolve pronouns and possessives from context, verify the project-sheet requirements against the evidence, and choose one level. The route specified by the chosen level is the next station.", "type": "score"}}, "state": {"context": "Case note: At 09:00 on 17 September 2026, IR-1 records a blue tote inspection. The blue tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching. The inspected opening measures 4 cm. All other project-sheet-required seams are secure, every required non-seam component is secure, and both panel measurements match the project sheet. Topstitching and thread cleanup are finished, but pressing is not finished. The tote is being held for the next work route.\n\nThe tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching.", "evidence": ["Inspection record IR-1 at 09:00 on 17 September 2026 identifies the blue tote’s physical seam as seam ID S-17.", "The blue tote’s project sheet lists seam ID S-18, and no other seam ID, among the seams required for the blue tote."]}}, "method": "c2d", "provenance": {"source_id": "diverse-203", "source_is_synthetic": true, "source_sha256": "62b02fcb4a0e58796e32491572ec31e832a7aba2391247dcb435399a4ede249f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged policy and bindings, use two complete factual evidence sentences, and present a coherent seam-identity change without contradictions or embedded answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal claims over an explicit class remain atomic. The focus atom concerns whether a particular physical seam is project-sheet-required, not a policy classification. The base and counter assignments are jointly realizable while changing only a1: the opened seam can be required in the base and an additional, nonrequired seam in the counter. The policy evidence correctly preserves the state-originating project-sheet requirements needed to interpret the unchanged rubric; observations and the generic request need not be copied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 and a2 together entail that a project-sheet-required seam has a positive-length opening. Under the ordered rubric, that is sufficient for score 0 and the seam-correction or stitching route, regardless of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "With a1 refuted, the opened inspected seam is not required. Because a3 then covers every required seam, and a4 establishes that every required non-seam component is secure, score 0 is excluded. Refuted a6 establishes unfinished pressing, while the other finish and measurement conditions hold, making score 1 and the finishing route sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The physical seam identified in blue-tote inspection record IR-1 at 09:00 on 17 September 2026 is one of the blue tote’s seams required by its project sheet."}, {"id": "a2", "statement": "The physical seam identified in blue-tote inspection record IR-1 at 09:00 on 17 September 2026 has an opening longer than 0 cm."}, {"id": "a3", "statement": "At 09:00 on 17 September 2026, every project-sheet-required seam of the blue tote other than the physical seam identified in inspection record IR-1 is secure."}, {"id": "a4", "statement": "At 09:00 on 17 September 2026, every project-sheet-required non-seam component of the blue tote is secure."}, {"id": "a5", "statement": "At 09:00 on 17 September 2026, the measured dimensions of every blue-tote panel match the dimensions specified for that panel by the blue tote’s project sheet."}, {"id": "a6", "statement": "At 09:00 on 17 September 2026, pressing of the blue tote is finished."}, {"id": "a7", "statement": "At 09:00 on 17 September 2026, topstitching of the blue tote is finished."}, {"id": "a8", "statement": "At 09:00 on 17 September 2026, thread cleanup on the blue tote is finished."}], "base_state_json": "{\"context\":\"Case note — blue tote inspection at 09:00 on 17 September 2026. The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching. Inspection measured an opening longer than 0 cm at the physical seam identified in IR-1. All other project-sheet-required seams are secure, every required non-seam component is secure, and both panel dimensions match the project sheet. Topstitching and thread cleanup are finished, but pressing is unfinished. The tote is held for the appropriate next station. The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching.\",\"evidence\":[\"Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-17.\",\"The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-17."}, {"path": ["evidence", "1"], "text": "The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18."}], "policy_evidence": [{"path": ["context"], "text": "The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching."}], "rules": [{"justification": "The inspected physical seam is a project-sheet-required seam and has a positive-length opening. A required seam is therefore open, which is sufficient for level 0 and routing to seam correction or stitching.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "The seam having the opening is not a project-sheet-required seam, while every required seam is secure and every required non-seam component is secure. Panel measurements match, topstitching and thread cleanup are finished, but pressing remains unfinished. The tote is therefore complete but has a finishing issue, which is sufficient for level 1 and routing to finishing.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-17.", "negative_left": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-18.", "negative_right": "The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18.", "right": "The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18."}, "verifier_independent_model": false}, "family": "fast-41-diverse-203-009", "id": "fast-41-diverse-203-009-base", "input": {"questions": {"decision": {"criteria": ["0 — Not complete: A required seam is open, hardware or handles are insecure, or a required component is missing. Route to seam correction or stitching before finishing.", "1 — Complete but needs finishing: All required seams and components are secure, but minor issues such as unpressed fabric, uneven topstitching, or loose thread ends remain. Route to finishing.", "2 — Complete with good finish: Measurements match the project sheet, all required seams and components are secure, and pressing, topstitching, and thread cleanup are finished. Release to the project recipient."], "instructions": "Use the ordered rubric. Resolve pronouns and possessives from context, verify the project-sheet requirements against the evidence, and choose one level. The route specified by the chosen level is the next station.", "type": "score"}}, "state": {"context": "Case note — blue tote inspection at 09:00 on 17 September 2026. The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching. Inspection measured an opening longer than 0 cm at the physical seam identified in IR-1. All other project-sheet-required seams are secure, every required non-seam component is secure, and both panel dimensions match the project sheet. Topstitching and thread cleanup are finished, but pressing is unfinished. The tote is held for the appropriate next station. The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching.", "evidence": ["Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-17.", "The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18."]}}, "method": "c2d", "provenance": {"source_id": "diverse-203", "source_is_synthetic": true, "source_sha256": "62b02fcb4a0e58796e32491572ec31e832a7aba2391247dcb435399a4ede249f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged policy and bindings, use two complete factual evidence sentences, and present a coherent seam-identity change without contradictions or embedded answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal claims over an explicit class remain atomic. The focus atom concerns whether a particular physical seam is project-sheet-required, not a policy classification. The base and counter assignments are jointly realizable while changing only a1: the opened seam can be required in the base and an additional, nonrequired seam in the counter. The policy evidence correctly preserves the state-originating project-sheet requirements needed to interpret the unchanged rubric; observations and the generic request need not be copied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "a1 and a2 together entail that a project-sheet-required seam has a positive-length opening. Under the ordered rubric, that is sufficient for score 0 and the seam-correction or stitching route, regardless of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "With a1 refuted, the opened inspected seam is not required. Because a3 then covers every required seam, and a4 establishes that every required non-seam component is secure, score 0 is excluded. Refuted a6 establishes unfinished pressing, while the other finish and measurement conditions hold, making score 1 and the finishing route sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The physical seam identified in blue-tote inspection record IR-1 at 09:00 on 17 September 2026 is one of the blue tote’s seams required by its project sheet."}, {"id": "a2", "statement": "The physical seam identified in blue-tote inspection record IR-1 at 09:00 on 17 September 2026 has an opening longer than 0 cm."}, {"id": "a3", "statement": "At 09:00 on 17 September 2026, every project-sheet-required seam of the blue tote other than the physical seam identified in inspection record IR-1 is secure."}, {"id": "a4", "statement": "At 09:00 on 17 September 2026, every project-sheet-required non-seam component of the blue tote is secure."}, {"id": "a5", "statement": "At 09:00 on 17 September 2026, the measured dimensions of every blue-tote panel match the dimensions specified for that panel by the blue tote’s project sheet."}, {"id": "a6", "statement": "At 09:00 on 17 September 2026, pressing of the blue tote is finished."}, {"id": "a7", "statement": "At 09:00 on 17 September 2026, topstitching of the blue tote is finished."}, {"id": "a8", "statement": "At 09:00 on 17 September 2026, thread cleanup on the blue tote is finished."}], "base_state_json": "{\"context\":\"Case note — blue tote inspection at 09:00 on 17 September 2026. The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching. Inspection measured an opening longer than 0 cm at the physical seam identified in IR-1. All other project-sheet-required seams are secure, every required non-seam component is secure, and both panel dimensions match the project sheet. Topstitching and thread cleanup are finished, but pressing is unfinished. The tote is held for the appropriate next station. The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching.\",\"evidence\":[\"Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-17.\",\"The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-17."}, {"path": ["evidence", "1"], "text": "The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18."}], "policy_evidence": [{"path": ["context"], "text": "The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching."}], "rules": [{"justification": "The inspected physical seam is a project-sheet-required seam and has a positive-length opening. A required seam is therefore open, which is sufficient for level 0 and routing to seam correction or stitching.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}]}, {"justification": "The seam having the opening is not a project-sheet-required seam, while every required seam is secure and every required non-seam component is secure. Panel measurements match, topstitching and thread cleanup are finished, but pressing remains unfinished. The tote is therefore complete but has a finishing issue, which is sufficient for level 1 and routing to finishing.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-17.", "negative_left": "Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-18.", "negative_right": "The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18.", "right": "The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18."}, "verifier_independent_model": false}, "family": "fast-41-diverse-203-009", "id": "fast-41-diverse-203-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not complete: A required seam is open, hardware or handles are insecure, or a required component is missing. Route to seam correction or stitching before finishing.", "1 — Complete but needs finishing: All required seams and components are secure, but minor issues such as unpressed fabric, uneven topstitching, or loose thread ends remain. Route to finishing.", "2 — Complete with good finish: Measurements match the project sheet, all required seams and components are secure, and pressing, topstitching, and thread cleanup are finished. Release to the project recipient."], "instructions": "Use the ordered rubric. Resolve pronouns and possessives from context, verify the project-sheet requirements against the evidence, and choose one level. The route specified by the chosen level is the next station.", "type": "score"}}, "state": {"context": "Case note — blue tote inspection at 09:00 on 17 September 2026. The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching. Inspection measured an opening longer than 0 cm at the physical seam identified in IR-1. All other project-sheet-required seams are secure, every required non-seam component is secure, and both panel dimensions match the project sheet. Topstitching and thread cleanup are finished, but pressing is unfinished. The tote is held for the appropriate next station. The tote’s project sheet requires two 36 × 40 cm panels, closed side and bottom seams, two secure handles, and pressed topstitching.", "evidence": ["Inspection record IR-1 at 09:00 on 17 September 2026 identifies the physical seam as seam S-18.", "The blue tote’s project sheet lists seam S-17 as a required seam and does not list seam S-18."]}}, "method": "c2d", "provenance": {"source_id": "diverse-203", "source_is_synthetic": true, "source_sha256": "62b02fcb4a0e58796e32491572ec31e832a7aba2391247dcb435399a4ede249f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and question bindings remain preserved through the verbatim questions and cushion-cover scope. The two evidence spans are complete factual sentences: \"Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction.\" and \"At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle.\" The counterfactual changes only visibility of the completed correction and introduces no contradictory duplicate measurement or assertion. Neither context embeds an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Assessment note\",\"text\":\"The cushion cover was examined at the final assessment after all required operations were finished. The zipper opened and closed through its full travel; every seam was secured, and every raw edge was enclosed.\"},{\"speaker\":\"Measurement record\",\"text\":\"The horizontal finished dimension was 40.1 cm and the vertical finished dimension was 39.9 cm, both within 39.8–40.2 cm inclusive. Maximum visible stitch drift was 1 mm, and no puckering was visible.\"},{\"speaker\":\"Correction log\",\"text\":\"Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction.\"},{\"speaker\":\"Visual inspection\",\"text\":\"At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle.\"},{\"speaker\":\"Scope note\",\"text\":\"The assessment concerned only this cushion cover; a separate tote project, color preferences, and optional decorations were outside the review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction."}, {"path": ["3", "text"], "text": "At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction.", "negative_left": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction.", "negative_right": "At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could be seen from any viewing angle.", "right": "At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle."}, "verifier_independent_model": false}, "family": "fast-41-diverse-204-015", "id": "fast-41-diverse-204-015-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Assessment note", "text": "The cushion cover was examined at the final assessment after all required operations were finished. The zipper opened and closed through its full travel; every seam was secured, and every raw edge was enclosed."}, {"speaker": "Measurement record", "text": "The horizontal finished dimension was 40.1 cm and the vertical finished dimension was 39.9 cm, both within 39.8–40.2 cm inclusive. Maximum visible stitch drift was 1 mm, and no puckering was visible."}, {"speaker": "Correction log", "text": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction."}, {"speaker": "Visual inspection", "text": "At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle."}, {"speaker": "Scope note", "text": "The assessment concerned only this cushion cover; a separate tote project, color preferences, and optional decorations were outside the review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and question bindings remain preserved through the verbatim questions and cushion-cover scope. The two evidence spans are complete factual sentences: \"Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction.\" and \"At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle.\" The counterfactual changes only visibility of the completed correction and introduces no contradictory duplicate measurement or assertion. Neither context embeds an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Assessment note\",\"text\":\"The cushion cover was examined at the final assessment after all required operations were finished. The zipper opened and closed through its full travel; every seam was secured, and every raw edge was enclosed.\"},{\"speaker\":\"Measurement record\",\"text\":\"The horizontal finished dimension was 40.1 cm and the vertical finished dimension was 39.9 cm, both within 39.8–40.2 cm inclusive. Maximum visible stitch drift was 1 mm, and no puckering was visible.\"},{\"speaker\":\"Correction log\",\"text\":\"Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction.\"},{\"speaker\":\"Visual inspection\",\"text\":\"At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle.\"},{\"speaker\":\"Scope note\",\"text\":\"The assessment concerned only this cushion cover; a separate tote project, color preferences, and optional decorations were outside the review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction."}, {"path": ["3", "text"], "text": "At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction.", "negative_left": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction.", "negative_right": "At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could be seen from any viewing angle.", "right": "At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could not be seen from any viewing angle."}, "verifier_independent_model": false}, "family": "fast-41-diverse-204-015", "id": "fast-41-diverse-204-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Assessment note", "text": "The cushion cover was examined at the final assessment after all required operations were finished. The zipper opened and closed through its full travel; every seam was secured, and every raw edge was enclosed."}, {"speaker": "Measurement record", "text": "The horizontal finished dimension was 40.1 cm and the vertical finished dimension was 39.9 cm, both within 39.8–40.2 cm inclusive. Maximum visible stitch drift was 1 mm, and no puckering was visible."}, {"speaker": "Correction log", "text": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated, and that sole correction was a completed seam correction."}, {"speaker": "Visual inspection", "text": "At the final assessment used for question decision, the completed seam correction's thread and fabric alteration could be seen from any viewing angle."}, {"speaker": "Scope note", "text": "The assessment concerned only this cushion cover; a separate tote project, color preferences, and optional decorations were outside the review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing rubric, scope, and cushion-cover bindings. Both contexts retain all policy-relevant facts without inventing exceptions or defaults. The evidence consists of two complete factual sentences: \"The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction.\" and \"At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric.\" The counterfactual changes only the visibility observation and is consistent with the unchanged measurements, operations, and correction record. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Inspection lead\",\"text\":\"The cushion cover's required operations are finished. At final assessment, its zipper opens and closes through full travel, every seam is secured, every raw edge is enclosed, and its horizontal and vertical finished dimensions are 40.1 cm and 39.9 cm, respectively. Maximum visible stitch drift is 0.8 mm, and no puckering is visible.\"},{\"speaker\":\"Sewing record\",\"text\":\"Before final assessment, exactly one correction was made to the rated cushion cover. It was a seam correction, and that correction is completed. The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction.\"},{\"speaker\":\"Final assessor\",\"text\":\"At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["1", "text"], "text": "The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction."}, {"path": ["2", "text"], "text": "At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction.", "negative_left": "The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction.", "negative_right": "At the final assessment used for question decision, seam sections S-14 and S-15 are both visible on the cover's outer fabric.", "right": "At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric."}, "verifier_independent_model": false}, "family": "fast-41-diverse-204-018", "id": "fast-41-diverse-204-018-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Inspection lead", "text": "The cushion cover's required operations are finished. At final assessment, its zipper opens and closes through full travel, every seam is secured, every raw edge is enclosed, and its horizontal and vertical finished dimensions are 40.1 cm and 39.9 cm, respectively. Maximum visible stitch drift is 0.8 mm, and no puckering is visible."}, {"speaker": "Sewing record", "text": "Before final assessment, exactly one correction was made to the rated cushion cover. It was a seam correction, and that correction is completed. The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction."}, {"speaker": "Final assessor", "text": "At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing rubric, scope, and cushion-cover bindings. Both contexts retain all policy-relevant facts without inventing exceptions or defaults. The evidence consists of two complete factual sentences: \"The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction.\" and \"At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric.\" The counterfactual changes only the visibility observation and is consistent with the unchanged measurements, operations, and correction record. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Inspection lead\",\"text\":\"The cushion cover's required operations are finished. At final assessment, its zipper opens and closes through full travel, every seam is secured, every raw edge is enclosed, and its horizontal and vertical finished dimensions are 40.1 cm and 39.9 cm, respectively. Maximum visible stitch drift is 0.8 mm, and no puckering is visible.\"},{\"speaker\":\"Sewing record\",\"text\":\"Before final assessment, exactly one correction was made to the rated cushion cover. It was a seam correction, and that correction is completed. The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction.\"},{\"speaker\":\"Final assessor\",\"text\":\"At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["1", "text"], "text": "The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction."}, {"path": ["2", "text"], "text": "At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction.", "negative_left": "The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction.", "negative_right": "At the final assessment used for question decision, seam sections S-14 and S-15 are both visible on the cover's outer fabric.", "right": "At the final assessment used for question decision, seam sections S-14 and S-15 are both hidden beneath the cover's outer fabric."}, "verifier_independent_model": false}, "family": "fast-41-diverse-204-018", "id": "fast-41-diverse-204-018-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Inspection lead", "text": "The cushion cover's required operations are finished. At final assessment, its zipper opens and closes through full travel, every seam is secured, every raw edge is enclosed, and its horizontal and vertical finished dimensions are 40.1 cm and 39.9 cm, respectively. Maximum visible stitch drift is 0.8 mm, and no puckering is visible."}, {"speaker": "Sewing record", "text": "Before final assessment, exactly one correction was made to the rated cushion cover. It was a seam correction, and that correction is completed. The sole correction made to the cushion cover being rated consists exactly of seam sections S-14 and S-15, and those sections are the entire correction."}, {"speaker": "Final assessor", "text": "At the final assessment used for question decision, seam sections S-14 and S-15 are both visible on the cover's outer fabric."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric and instructions, and both contexts retain the cushion-cover scope. The entity, assessment path, correction identity, and timing remain bound consistently. The two evidence entries are complete factual sentences: \"The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work.\" and \"Record SC-17 reports no visible trace of the completed seam work at the final assessment.\" The counterfactual changes only the visibility observation, without creating contradictory duplicate measurements or assertions. Neither context states a score, answer code, rationale, rule table, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Final inspection note\",\"text\":\"The cushion cover underwent its decision assessment only after every required construction operation was finished. The zipper opened and closed across its full travel, all seams were secured, and every raw edge was enclosed. This record concerns the rated cushion cover rather than any other item.\"},{\"speaker\":\"Measurement and finish log\",\"text\":\"The horizontal finished dimension was 40.1 cm and the vertical finished dimension was 39.9 cm, each within 39.8–40.2 cm inclusive. The largest visible stitch drift measured 1 mm, and the inspector saw no puckering. No separate repairs or decorative changes were recorded for this cover.\"},{\"speaker\":\"Correction register\",\"text\":\"Before the decision assessment, exactly one correction had been made to the cushion cover. The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work.\"},{\"speaker\":\"Visual review\",\"text\":\"Record SC-17 reports no visible trace of the completed seam work at the final assessment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work."}, {"path": ["3", "text"], "text": "Record SC-17 reports no visible trace of the completed seam work at the final assessment."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work.", "negative_left": "The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work.", "negative_right": "Record SC-17 reports a visible trace of the completed seam work at the final assessment.", "right": "Record SC-17 reports no visible trace of the completed seam work at the final assessment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-204-029", "id": "fast-41-diverse-204-029-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Final inspection note", "text": "The cushion cover underwent its decision assessment only after every required construction operation was finished. The zipper opened and closed across its full travel, all seams were secured, and every raw edge was enclosed. This record concerns the rated cushion cover rather than any other item."}, {"speaker": "Measurement and finish log", "text": "The horizontal finished dimension was 40.1 cm and the vertical finished dimension was 39.9 cm, each within 39.8–40.2 cm inclusive. The largest visible stitch drift measured 1 mm, and the inspector saw no puckering. No separate repairs or decorative changes were recorded for this cover."}, {"speaker": "Correction register", "text": "Before the decision assessment, exactly one correction had been made to the cushion cover. The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work."}, {"speaker": "Visual review", "text": "Record SC-17 reports no visible trace of the completed seam work at the final assessment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric and instructions, and both contexts retain the cushion-cover scope. The entity, assessment path, correction identity, and timing remain bound consistently. The two evidence entries are complete factual sentences: \"The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work.\" and \"Record SC-17 reports no visible trace of the completed seam work at the final assessment.\" The counterfactual changes only the visibility observation, without creating contradictory duplicate measurements or assertions. Neither context states a score, answer code, rationale, rule table, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual property or quantified factual relationship; none is a bundled final classification. The focus, visibility of the sole correction, is factual rather than policy-based. The base and counter assignments can differ only in whether the completed correction remains visible, with all other properties unchanged. Empty policy_evidence is correct because the complete scoring rubric, scope, tolerances, and exceptions are already retained in the questions object; no additional substantive rule from the original state is needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all Level 1 requirements and all stricter Level 2 requirements. The narrow dimension bounds imply the Level 1 bounds, the drift is at most 1 mm, there is no puckering, and—because exactly one correction exists and no part of it is visible—there are no visible corrections. Target 2 is therefore required.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all Level 1 requirements. Refuting a12 entails that some part of the sole completed seam correction is visible, which violates Level 2's requirement of no visible corrections while remaining permitted as a completed seam correction at Level 1. Thus target 1 is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At the final assessment used for question decision, every required operation for the cushion cover being rated is finished."}, {"id": "a2", "statement": "At the final assessment used for question decision, the zipper on the cushion cover being rated opens and closes through its full travel."}, {"id": "a3", "statement": "At the final assessment used for question decision, every seam on the cushion cover being rated is secured."}, {"id": "a4", "statement": "At the final assessment used for question decision, every raw edge on the cushion cover being rated is enclosed."}, {"id": "a5", "statement": "At the final assessment used for question decision, the horizontal finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a6", "statement": "At the final assessment used for question decision, the vertical finished dimension of the cushion cover being rated is within 39.8–40.2 cm inclusive."}, {"id": "a7", "statement": "At the final assessment used for question decision, the maximum visible stitch drift on the cushion cover being rated is no more than 1 mm."}, {"id": "a8", "statement": "At the final assessment used for question decision, no puckering is visible on the cushion cover being rated."}, {"id": "a9", "statement": "Before the final assessment used for question decision, exactly one correction had been made to the cushion cover being rated."}, {"id": "a10", "statement": "The sole correction made to the cushion cover being rated is a seam correction."}, {"id": "a11", "statement": "The sole seam correction made to the cushion cover being rated is completed."}, {"id": "a12", "statement": "At the final assessment used for question decision, no part of the sole correction on the cushion cover being rated is visible."}], "base_state_json": "[{\"speaker\":\"Final inspection note\",\"text\":\"The cushion cover underwent its decision assessment only after every required construction operation was finished. The zipper opened and closed across its full travel, all seams were secured, and every raw edge was enclosed. This record concerns the rated cushion cover rather than any other item.\"},{\"speaker\":\"Measurement and finish log\",\"text\":\"The horizontal finished dimension was 40.1 cm and the vertical finished dimension was 39.9 cm, each within 39.8–40.2 cm inclusive. The largest visible stitch drift measured 1 mm, and the inspector saw no puckering. No separate repairs or decorative changes were recorded for this cover.\"},{\"speaker\":\"Correction register\",\"text\":\"Before the decision assessment, exactly one correction had been made to the cushion cover. The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work.\"},{\"speaker\":\"Visual review\",\"text\":\"Record SC-17 reports no visible trace of the completed seam work at the final assessment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work."}, {"path": ["3", "text"], "text": "Record SC-17 reports no visible trace of the completed seam work at the final assessment."}], "policy_evidence": [], "rules": [{"justification": "All Level 1 conditions are met; both dimensions meet the narrower inclusive Level 2 tolerance, visible stitch drift is no more than 1 mm, no puckering is visible, and the completed seam correction is not visible. Therefore the high-quality-finish score is required.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The narrower dimension bounds and 1 mm drift bound entail the corresponding Level 1 limits, and all other Level 1 requirements are met. Refutation of a12 establishes that the sole completed seam correction is visible, which Level 1 allows but Level 2 excludes, so score 1 is required.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work.", "negative_left": "The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work.", "negative_right": "Record SC-17 reports a visible trace of the completed seam work at the final assessment.", "right": "Record SC-17 reports no visible trace of the completed seam work at the final assessment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-204-029", "id": "fast-41-diverse-204-029-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: Any required operation is unfinished, the zipper fails, a seam is unsecured, a raw edge is exposed, or either finished dimension is outside 39.5–40.5 cm.", "1 — Complete with acceptable finish: All required operations are finished, the zipper works, seams are secured, raw edges are enclosed, and both dimensions are within 39.5–40.5 cm inclusive; minor visible stitch drift up to 3 mm or a completed seam correction is allowed.", "2 — Complete with high-quality finish: Level 1 conditions are met, both dimensions are within 39.8–40.2 cm inclusive, visible stitch drift is no more than 1 mm, and there are no visible corrections or puckering."], "instructions": "Rate the cushion cover’s completion and finish quality using the rubric. Treat measurement tolerances as inclusive. Consider only requirements and evidence for this cushion cover; unrelated future projects, color preferences, and optional decorations are distractors.", "type": "score"}}, "state": [{"speaker": "Final inspection note", "text": "The cushion cover underwent its decision assessment only after every required construction operation was finished. The zipper opened and closed across its full travel, all seams were secured, and every raw edge was enclosed. This record concerns the rated cushion cover rather than any other item."}, {"speaker": "Measurement and finish log", "text": "The horizontal finished dimension was 40.1 cm and the vertical finished dimension was 39.9 cm, each within 39.8–40.2 cm inclusive. The largest visible stitch drift measured 1 mm, and the inspector saw no puckering. No separate repairs or decorative changes were recorded for this cover."}, {"speaker": "Correction register", "text": "Before the decision assessment, exactly one correction had been made to the cushion cover. The final-assessment record labeled SC-17 belongs to the cushion cover being rated and records its sole correction as completed seam work."}, {"speaker": "Visual review", "text": "Record SC-17 reports a visible trace of the completed seam work at the final assessment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-204", "source_is_synthetic": true, "source_sha256": "6c29ec7eade372e7b9edd77092cad14c97a32f6da36bf9ce89cea271a61774f6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and the original question’s museum-livestream, sponsor-clip, channel, and duration bindings. The two evidence spans are complete factual sentences and match the base context. The counterfactual coherently changes the locally retained channel without contradictory duplicate measurements. Neither context states an answer, label, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "supported", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "full_context_fact_states": {"base": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "supported", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "counterfactual": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "refuted", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "remove_left": {"a_local_left_retained": "unknown"}, "remove_right": {"a_local_left_retained": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_local_left_retained": "unknown"}, "negative_pair": {"a_local_left_retained": "refuted"}, "negative_sentence": {"a_local_left_retained": "unknown"}, "positive_pair": {"a_local_left_retained": "supported"}, "right": {"a_local_left_retained": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a bundled policy conclusion. The focus atom is the factual condition of whether the local recording retained the left channel. Base and counter assignments differ only on that fact and are both realizable: one describes a platform-only loss and the other a locally present loss. The policy evidence correctly preserves the substantive routing, impact, exception, and intervention rules originating in the original state; question-level criteria need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a platform-output left-channel defect with a clean local recording, a stereo specification excluding intentional mono, and a timely cue excluding missed-cue ownership. It also establishes an 18-second duration with intelligible speech, satisfying Minor. Under the stated criteria, this is sufficient for the streaming-technician, Minor, intervention bundle.", "rule_index": 0, "sound": true}, {"reason": "Refuting local retention while supporting the same left-channel loss on the stream establishes that the non-exempt stereo audio defect is present locally. The state policy therefore assigns it to the audio engineer, which makes both listed substantive bundles incorrect and is sufficient for none_of_above. The duration and intelligibility also establish Minor, though the differing route alone already distinguishes the result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_stream_left_loss", "statement": "During the scheduled museum livestream, the streaming platform output did not carry the sponsor clip's left-channel program audio."}, {"id": "a_local_left_retained", "statement": "During the sponsor clip, the local recording retained the clip's left-channel program audio."}, {"id": "a_stereo_spec", "statement": "The rundown specifies stereo presentation for the sponsor clip."}, {"id": "a_duration", "statement": "The sponsor clip's left-channel loss on the streaming platform lasted 18 seconds."}, {"id": "a_intelligible", "statement": "Speech remained intelligible during the sponsor clip's left-channel loss."}, {"id": "a_cue_timing", "statement": "The stage manager fired the sponsor clip cue at its scheduled time."}], "base_state_json": "\"At the scheduled museum livestream, the sponsor clip’s streaming-platform output failed to carry its left-channel program audio for 18 seconds. The rundown specified stereo presentation for the clip, and speech remained intelligible throughout the interruption. The stage manager fired the sponsor clip cue at its scheduled time, so the event timing matched the planned rundown. During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel. That retained channel carried the sponsor clip's left-channel program audio. The incident was documented while the clip was active, rather than during a transition, and the local and platform records were reviewed independently. Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention.\"", "base_states": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "supported"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}], "counter_states": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "refuted"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}], "focus_atom": "a_local_left_retained", "focus_evidence": [{"path": [], "text": "During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel."}, {"path": [], "text": "That retained channel carried the sponsor clip's left-channel program audio."}], "policy_evidence": [{"path": [], "text": "Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention."}], "rules": [{"justification": "The left-channel defect is present in the streaming-platform output but absent from the local recording, the stereo specification excludes intentional mono, and the timely cue excludes a missed-cue route. The 18-second duration and intelligible speech satisfy Minor.", "target": "streaming_technician_minor_intervention", "when": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "supported"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}]}, {"justification": "The stereo clip's left-channel program audio is absent both from the platform output and from the local recording, so the defect is present locally and belongs to the audio engineer rather than either listed route. Although the 18-second intelligible incident is Minor, its correct route differs from both substantive bundles.", "target": "none_of_above", "when": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "refuted"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}]}]}, "verified_pair": {"left": "During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel.", "negative_left": "During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel.", "negative_right": "That retained channel carried the sponsor clip's right-channel program audio.", "right": "That retained channel carried the sponsor clip's left-channel program audio."}, "verifier_independent_model": false}, "family": "fast-41-diverse-217-013", "id": "fast-41-diverse-217-013-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when the correct route, intervention decision, or impact rating differs from both listed substantive decision bundles.", "stage_manager_minor_no_intervention": "Route to the stage manager, rate Minor, and take no intervention because the issue was a cueing problem that has ended.", "streaming_technician_minor_intervention": "Route to the streaming technician, rate Minor, and intervene because the defect was confined to the streaming platform while the local recording remained clean."}, "instructions": "Select the single option that correctly routes the incident, determines whether intervention is required, and rates its broadcast impact under the stated scope and exceptions.", "type": "choice"}}, "state": "At the scheduled museum livestream, the sponsor clip’s streaming-platform output failed to carry its left-channel program audio for 18 seconds. The rundown specified stereo presentation for the clip, and speech remained intelligible throughout the interruption. The stage manager fired the sponsor clip cue at its scheduled time, so the event timing matched the planned rundown. During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel. That retained channel carried the sponsor clip's left-channel program audio. The incident was documented while the clip was active, rather than during a transition, and the local and platform records were reviewed independently. Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention."}, "method": "c2d", "provenance": {"source_id": "diverse-217", "source_is_synthetic": true, "source_sha256": "56c54e57171feb278bcede4915e9a42909b58c37f80d15ac73b42a2c27244a6b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "streaming_technician_minor_intervention"}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and the original question’s museum-livestream, sponsor-clip, channel, and duration bindings. The two evidence spans are complete factual sentences and match the base context. The counterfactual coherently changes the locally retained channel without contradictory duplicate measurements. Neither context states an answer, label, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "refuted", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "full_context_fact_states": {"base": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "supported", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "counterfactual": {"a_cue_timing": "supported", "a_duration": "supported", "a_intelligible": "supported", "a_local_left_retained": "refuted", "a_stereo_spec": "supported", "a_stream_left_loss": "supported"}, "remove_left": {"a_local_left_retained": "unknown"}, "remove_right": {"a_local_left_retained": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_local_left_retained": "unknown"}, "negative_pair": {"a_local_left_retained": "refuted"}, "negative_sentence": {"a_local_left_retained": "unknown"}, "positive_pair": {"a_local_left_retained": "supported"}, "right": {"a_local_left_retained": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a bundled policy conclusion. The focus atom is the factual condition of whether the local recording retained the left channel. Base and counter assignments differ only on that fact and are both realizable: one describes a platform-only loss and the other a locally present loss. The policy evidence correctly preserves the substantive routing, impact, exception, and intervention rules originating in the original state; question-level criteria need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a platform-output left-channel defect with a clean local recording, a stereo specification excluding intentional mono, and a timely cue excluding missed-cue ownership. It also establishes an 18-second duration with intelligible speech, satisfying Minor. Under the stated criteria, this is sufficient for the streaming-technician, Minor, intervention bundle.", "rule_index": 0, "sound": true}, {"reason": "Refuting local retention while supporting the same left-channel loss on the stream establishes that the non-exempt stereo audio defect is present locally. The state policy therefore assigns it to the audio engineer, which makes both listed substantive bundles incorrect and is sufficient for none_of_above. The duration and intelligibility also establish Minor, though the differing route alone already distinguishes the result.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_stream_left_loss", "statement": "During the scheduled museum livestream, the streaming platform output did not carry the sponsor clip's left-channel program audio."}, {"id": "a_local_left_retained", "statement": "During the sponsor clip, the local recording retained the clip's left-channel program audio."}, {"id": "a_stereo_spec", "statement": "The rundown specifies stereo presentation for the sponsor clip."}, {"id": "a_duration", "statement": "The sponsor clip's left-channel loss on the streaming platform lasted 18 seconds."}, {"id": "a_intelligible", "statement": "Speech remained intelligible during the sponsor clip's left-channel loss."}, {"id": "a_cue_timing", "statement": "The stage manager fired the sponsor clip cue at its scheduled time."}], "base_state_json": "\"At the scheduled museum livestream, the sponsor clip’s streaming-platform output failed to carry its left-channel program audio for 18 seconds. The rundown specified stereo presentation for the clip, and speech remained intelligible throughout the interruption. The stage manager fired the sponsor clip cue at its scheduled time, so the event timing matched the planned rundown. During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel. That retained channel carried the sponsor clip's left-channel program audio. The incident was documented while the clip was active, rather than during a transition, and the local and platform records were reviewed independently. Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention.\"", "base_states": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "supported"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}], "counter_states": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "refuted"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}], "focus_atom": "a_local_left_retained", "focus_evidence": [{"path": [], "text": "During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel."}, {"path": [], "text": "That retained channel carried the sponsor clip's left-channel program audio."}], "policy_evidence": [{"path": [], "text": "Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention."}], "rules": [{"justification": "The left-channel defect is present in the streaming-platform output but absent from the local recording, the stereo specification excludes intentional mono, and the timely cue excludes a missed-cue route. The 18-second duration and intelligible speech satisfy Minor.", "target": "streaming_technician_minor_intervention", "when": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "supported"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}]}, {"justification": "The stereo clip's left-channel program audio is absent both from the platform output and from the local recording, so the defect is present locally and belongs to the audio engineer rather than either listed route. Although the 18-second intelligible incident is Minor, its correct route differs from both substantive bundles.", "target": "none_of_above", "when": [{"atom_id": "a_stream_left_loss", "state": "supported"}, {"atom_id": "a_local_left_retained", "state": "refuted"}, {"atom_id": "a_stereo_spec", "state": "supported"}, {"atom_id": "a_duration", "state": "supported"}, {"atom_id": "a_intelligible", "state": "supported"}, {"atom_id": "a_cue_timing", "state": "supported"}]}]}, "verified_pair": {"left": "During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel.", "negative_left": "During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel.", "negative_right": "That retained channel carried the sponsor clip's right-channel program audio.", "right": "That retained channel carried the sponsor clip's left-channel program audio."}, "verifier_independent_model": false}, "family": "fast-41-diverse-217-013", "id": "fast-41-diverse-217-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Choose when the correct route, intervention decision, or impact rating differs from both listed substantive decision bundles.", "stage_manager_minor_no_intervention": "Route to the stage manager, rate Minor, and take no intervention because the issue was a cueing problem that has ended.", "streaming_technician_minor_intervention": "Route to the streaming technician, rate Minor, and intervene because the defect was confined to the streaming platform while the local recording remained clean."}, "instructions": "Select the single option that correctly routes the incident, determines whether intervention is required, and rates its broadcast impact under the stated scope and exceptions.", "type": "choice"}}, "state": "At the scheduled museum livestream, the sponsor clip’s streaming-platform output failed to carry its left-channel program audio for 18 seconds. The rundown specified stereo presentation for the clip, and speech remained intelligible throughout the interruption. The stage manager fired the sponsor clip cue at its scheduled time, so the event timing matched the planned rundown. During the sponsor clip in the scheduled museum livestream, the local recording retained exactly one program-audio channel. That retained channel carried the sponsor clip's right-channel program audio. The incident was documented while the clip was active, rather than during a transition, and the local and platform records were reviewed independently. Policy: the stage manager owns missed cues; the streaming technician owns platform-only faults when the local recording is clean; the audio engineer owns defects present locally. Minor means under 30 seconds with intelligible programming. Active, non-exempt local audio defects require intervention."}, "method": "c2d", "provenance": {"source_id": "diverse-217", "source_is_synthetic": true, "source_sha256": "56c54e57171feb278bcede4915e9a42909b58c37f80d15ac73b42a2c27244a6b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged request, routing policy, and impact scale. The two evidence spans are complete factual sentences: \"During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.\" \"During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9.\" The counterfactual consistently changes A-310 from distinct from U-9 to identical to U-9 without contradictory duplicates. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"context\":\"Case note: During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9. The unit's malfunction interrupted its output for eight seconds. A technician reset U-9 at 19:42, and it returned to normal operation immediately afterward. No fault other than U-9's malfunction occurred, and no editorial decision occurred. L-204's sole function was providing the local confidence display. A-310's output supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"path": ["context"], "text": "During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_right": "During the scheduled awards livestream, public-program audio processor A-310 was the same device as malfunctioning unit U-9.", "right": "During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9."}, "verifier_independent_model": false}, "family": "fast-41-diverse-218-013", "id": "fast-41-diverse-218-013-base", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Case note: During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9. The unit's malfunction interrupted its output for eight seconds. A technician reset U-9 at 19:42, and it returned to normal operation immediately afterward. No fault other than U-9's malfunction occurred, and no editorial decision occurred. L-204's sole function was providing the local confidence display. A-310's output supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "streaming_level_1"}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged request, routing policy, and impact scale. The two evidence spans are complete factual sentences: \"During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.\" \"During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9.\" The counterfactual consistently changes A-310 from distinct from U-9 to identical to U-9 without contradictory duplicates. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"context\":\"Case note: During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9. The unit's malfunction interrupted its output for eight seconds. A technician reset U-9 at 19:42, and it returned to normal operation immediately afterward. No fault other than U-9's malfunction occurred, and no editorial decision occurred. L-204's sole function was providing the local confidence display. A-310's output supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"path": ["context"], "text": "During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_right": "During the scheduled awards livestream, public-program audio processor A-310 was the same device as malfunctioning unit U-9.", "right": "During the scheduled awards livestream, public-program audio processor A-310 was not the same device as malfunctioning unit U-9."}, "verifier_independent_model": false}, "family": "fast-41-diverse-218-013", "id": "fast-41-diverse-218-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Case note: During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, public-program audio processor A-310 was the same device as malfunctioning unit U-9. The unit's malfunction interrupted its output for eight seconds. A technician reset U-9 at 19:42, and it returned to normal operation immediately afterward. No fault other than U-9's malfunction occurred, and no editorial decision occurred. L-204's sole function was providing the local confidence display. A-310's output supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "audio_level_2"}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and both contexts preserve the governing routing policy and impact scale. The request, entities, livestream setting, and 19:42 time binding remain aligned. The evidence contains exactly two complete factual sentences: “During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310.” and “During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310.” The base context and counterfactual are each internally coherent, with the counterfactual consistently changing U-9’s identity to A-310. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or classifier output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"context\":\"During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310. The malfunction interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, and it returned to normal immediately. No other fault and no editorial decision occurred. L-204's sole function was providing the local confidence display, while A-310 supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what \\u201cit\\u201d refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_left": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_right": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to public-program audio processor A-310.", "right": "During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310."}, "verifier_independent_model": false}, "family": "fast-41-diverse-218-014", "id": "fast-41-diverse-218-014-base", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310. The malfunction interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, and it returned to normal immediately. No other fault and no editorial decision occurred. L-204's sole function was providing the local confidence display, while A-310 supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "streaming_level_1"}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and both contexts preserve the governing routing policy and impact scale. The request, entities, livestream setting, and 19:42 time binding remain aligned. The evidence contains exactly two complete factual sentences: “During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310.” and “During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310.” The base context and counterfactual are each internally coherent, with the counterfactual consistently changing U-9’s identity to A-310. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or classifier output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"context\":\"During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310. The malfunction interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, and it returned to normal immediately. No other fault and no editorial decision occurred. L-204's sole function was providing the local confidence display, while A-310 supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what \\u201cit\\u201d refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_left": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_right": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to public-program audio processor A-310.", "right": "During the scheduled awards livestream, malfunctioning unit U-9 was not identical to public-program audio processor A-310."}, "verifier_independent_model": false}, "family": "fast-41-diverse-218-014", "id": "fast-41-diverse-218-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "During the scheduled awards livestream, malfunctioning unit U-9 was identical to exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, malfunctioning unit U-9 was identical to public-program audio processor A-310. The malfunction interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, and it returned to normal immediately. No other fault and no editorial decision occurred. L-204's sole function was providing the local confidence display, while A-310 supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "audio_level_2"}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question, policy, scope, and bindings; the two exact evidence quotes are complete factual sentences. The counterfactual consistently changes the faulty unit to A-310 without contradictory duplicates or embedded answer instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"context\":\"Case note: During the scheduled awards livestream, the equipment log recorded a fault involving U-9. During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, malfunctioning unit U-9 was not public-program audio processor A-310. The fault interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, after which it immediately returned to normal operation. No other fault and no editorial decision occurred. L-204's sole function was providing the local confidence display, while A-310 supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was not public-program audio processor A-310."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_right": "During the scheduled awards livestream, malfunctioning unit U-9 was public-program audio processor A-310.", "right": "During the scheduled awards livestream, malfunctioning unit U-9 was not public-program audio processor A-310."}, "verifier_independent_model": false}, "family": "fast-41-diverse-218-024", "id": "fast-41-diverse-218-024-base", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Case note: During the scheduled awards livestream, the equipment log recorded a fault involving U-9. During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, malfunctioning unit U-9 was not public-program audio processor A-310. The fault interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, after which it immediately returned to normal operation. No other fault and no editorial decision occurred. L-204's sole function was providing the local confidence display, while A-310 supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "streaming_level_1"}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question, policy, scope, and bindings; the two exact evidence quotes are complete factual sentences. The counterfactual consistently changes the faulty unit to A-310 without contradictory duplicates or embedded answer instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; exclusive identity, sole-function, and no-other-fault propositions remain atomic despite their logical strength or quantified scope. The focus is the factual identity of U-9. Both assignments are realizable: the base makes U-9 the display, while the counter makes it the audio processor through the unchanged exclusive alternative, with all other listed atom states consistent. Policy evidence preserves the state-originating routing and impact rules needed alongside the automatically retained question. Including the generic request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that U-9 was the local confidence display, that its output alone was disrupted, that the audience-facing broadcast remained usable, and that the reset restored normal operation. The state-supplied routing policy sends display faults to Streaming, while the unchanged question specifies Level 1 and no further intervention for a successfully reset local confidence-display disruption.", "rule_index": 0, "sound": true}, {"reason": "Given that U-9 was exactly one of L-204 and A-310 and was not L-204, it was A-310. Because A-310 supplied public-program sound and its output was interrupted for eight seconds, the public program briefly lost or degraded sound. Continued broadcast usability excludes Level 3, the reset constitutes intervention during the broadcast, and the routing policy sends sound faults to Audio.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was the same device as local confidence display L-204."}, {"id": "a2", "statement": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"id": "a3", "statement": "During the scheduled awards livestream, the malfunction of unit U-9 interrupted U-9's output."}, {"id": "a4", "statement": "The interruption of unit U-9's output during the scheduled awards livestream lasted eight seconds."}, {"id": "a5", "statement": "A technician reset unit U-9 at 19:42 during the scheduled awards livestream."}, {"id": "a6", "statement": "Unit U-9 returned to normal operation immediately after its 19:42 reset during the scheduled awards livestream."}, {"id": "a7", "statement": "No fault other than unit U-9's malfunction occurred during the scheduled awards livestream."}, {"id": "a8", "statement": "No editorial decision occurred during the scheduled awards livestream."}, {"id": "a9", "statement": "During the scheduled awards livestream, the sole function of display L-204 was providing the local confidence display."}, {"id": "a10", "statement": "During the scheduled awards livestream, the output of audio processor A-310 supplied the public program sound."}, {"id": "a11", "statement": "The audience-facing broadcast remained usable throughout the interruption of unit U-9's output during the scheduled awards livestream."}], "base_state_json": "{\"context\":\"Case note: During the scheduled awards livestream, the equipment log recorded a fault involving U-9. During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, malfunctioning unit U-9 was not public-program audio processor A-310. The fault interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, after which it immediately returned to normal operation. No other fault and no editorial decision occurred. L-204's sole function was providing the local confidence display, while A-310 supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310."}, {"path": ["context"], "text": "During the scheduled awards livestream, malfunctioning unit U-9 was not public-program audio processor A-310."}], "policy_evidence": [{"path": ["context"], "text": "Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer."}, {"path": ["evidence", "2"], "text": "Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped."}, {"path": ["request"], "text": "Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}], "rules": [{"justification": "U-9 was L-204, whose sole function was the local confidence display. Its output was interrupted, but the audience-facing broadcast remained usable, no other fault or editorial decision occurred, and the reset restored normal operation. Thus only a local confidence display was disrupted, the display fault routes to Streaming, and no further intervention is required after the successful reset.", "target": "streaming_level_1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Because U-9 was exactly one of L-204 and A-310, and U-9 was not L-204, U-9 was A-310. A-310 supplied public program sound, so its eight-second output interruption briefly degraded or removed that sound. The reset was an intervention during the broadcast, the broadcast remained usable, and no competing fault or editorial decision occurred. The sound fault therefore routes to Audio and has Level 2 impact.", "target": "audio_level_2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_left": "During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310.", "negative_right": "During the scheduled awards livestream, malfunctioning unit U-9 was public-program audio processor A-310.", "right": "During the scheduled awards livestream, malfunctioning unit U-9 was not public-program audio processor A-310."}, "verifier_independent_model": false}, "family": "fast-41-diverse-218-024", "id": "fast-41-diverse-218-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"audio_level_2": "Route to the Audio engineer, intervene during the broadcast, and rate Level 2 because the public program briefly lost or degraded sound.", "producer_level_0": "Route to the Event producer, take no intervention, and rate Level 0 because no technical issue was verified anywhere.", "stage_level_3": "Route to the Stage manager, stop or delay the show, and rate Level 3 because the audience-facing broadcast became unusable.", "streaming_level_1": "Route to the Streaming technician, require no further intervention after the successful reset, and rate Level 1 because only a local confidence display was disrupted."}, "instructions": "Choose the single option whose routing, intervention decision, and impact rating all match the report, logs, and stated policy.", "type": "choice"}}, "state": {"context": "Case note: During the scheduled awards livestream, the equipment log recorded a fault involving U-9. During the scheduled awards livestream, malfunctioning unit U-9 was exactly one of local confidence display L-204 and public-program audio processor A-310. During the scheduled awards livestream, malfunctioning unit U-9 was public-program audio processor A-310. The fault interrupted U-9's output for eight seconds. A technician reset U-9 at 19:42, after which it immediately returned to normal operation. No other fault and no editorial decision occurred. L-204's sole function was providing the local confidence display, while A-310 supplied the public program sound. The audience-facing broadcast remained usable throughout the interruption. Policy routes display or encoder faults to Streaming; stage-cue errors to Stage; sound faults to Audio; editorial decisions to Producer. Impact scale: 0 = no verified issue; 1 = local workflow disruption only; 2 = brief broadcast degradation; 3 = broadcast unusable or stopped. Resolve what “it” refers to, route the incident, decide whether further intervention is required, and assign the impact level."}}, "method": "c2d", "provenance": {"source_id": "diverse-218", "source_is_synthetic": true, "source_sha256": "3e92ccc9ce8892fd9ed2b9ee353156c15fae779af4323cf25add5af3a5e54e4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "audio_level_2"}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policy, bind to the same public video incident, use two complete factual evidence sentences, and change only the consistent end timestamp without embedding an answer or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Operations note\",\"text\":\"During the public preshow countdown, the encoder log records video-freeze incident F, and the monitoring team confirms that the same incident appears in the public encoded output.\"},{\"speaker\":\"Encoder record\",\"text\":\"The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC.\"},{\"speaker\":\"Encoder record\",\"text\":\"The encoder log for public video-freeze incident F records its end timestamp as 14:07:04.250 UTC.\"},{\"speaker\":\"Production review\",\"text\":\"The review finds that incident F did not cause program content to be lost. The encoder record is the controlling source for the incident timing, rather than approximate observations from personnel. The affected item was video, and the streaming technician is the designated production function for video issues.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC."}, {"path": ["2", "text"], "text": "The encoder log for public video-freeze incident F records its end timestamp as 14:07:04.250 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC.", "negative_left": "The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC.", "negative_right": "The encoder log for public video-freeze incident F records its end timestamp as 14:07:03.750 UTC.", "right": "The encoder log for public video-freeze incident F records its end timestamp as 14:07:04.250 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-003", "id": "fast-41-diverse-219-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Operations note", "text": "During the public preshow countdown, the encoder log records video-freeze incident F, and the monitoring team confirms that the same incident appears in the public encoded output."}, {"speaker": "Encoder record", "text": "The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC."}, {"speaker": "Encoder record", "text": "The encoder log for public video-freeze incident F records its end timestamp as 14:07:04.250 UTC."}, {"speaker": "Production review", "text": "The review finds that incident F did not cause program content to be lost. The encoder record is the controlling source for the incident timing, rather than approximate observations from personnel. The affected item was video, and the streaming technician is the designated production function for video issues."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and governing policy, bind to the same public video incident, use two complete factual evidence sentences, and change only the consistent end timestamp without embedding an answer or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Operations note\",\"text\":\"During the public preshow countdown, the encoder log records video-freeze incident F, and the monitoring team confirms that the same incident appears in the public encoded output.\"},{\"speaker\":\"Encoder record\",\"text\":\"The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC.\"},{\"speaker\":\"Encoder record\",\"text\":\"The encoder log for public video-freeze incident F records its end timestamp as 14:07:04.250 UTC.\"},{\"speaker\":\"Production review\",\"text\":\"The review finds that incident F did not cause program content to be lost. The encoder record is the controlling source for the incident timing, rather than approximate observations from personnel. The affected item was video, and the streaming technician is the designated production function for video issues.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC."}, {"path": ["2", "text"], "text": "The encoder log for public video-freeze incident F records its end timestamp as 14:07:04.250 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC.", "negative_left": "The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC.", "negative_right": "The encoder log for public video-freeze incident F records its end timestamp as 14:07:03.750 UTC.", "right": "The encoder log for public video-freeze incident F records its end timestamp as 14:07:04.250 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-003", "id": "fast-41-diverse-219-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Operations note", "text": "During the public preshow countdown, the encoder log records video-freeze incident F, and the monitoring team confirms that the same incident appears in the public encoded output."}, {"speaker": "Encoder record", "text": "The encoder log for public video-freeze incident F records its start timestamp as 14:07:02.000 UTC."}, {"speaker": "Encoder record", "text": "The encoder log for public video-freeze incident F records its end timestamp as 14:07:03.750 UTC."}, {"speaker": "Production review", "text": "The review finds that incident F did not cause program content to be lost. The encoder record is the controlling source for the incident timing, rather than approximate observations from personnel. The affected item was video, and the streaming technician is the designated production function for video issues."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all policy and bindings, both evidence spans are complete factual sentences, the counterfactual consistently changes only the freeze end time, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"During the public preshow countdown, staff logged video-freeze incident F for review. The event occurred in the public stream, not merely on a backstage monitor. No program material was lost while the incident was being checked, and the stream returned to normal afterward. A separate microphone-pop report was confined to the backstage intercom and did not appear in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"For video-freeze incident F, the encoder log records an end timestamp of 14:03:13.100 UTC.\"},{\"speaker\":\"Event producer\",\"text\":\"The scheduled opening remains ahead, and the production team is monitoring the public feed. The review concerns only incident F; the backstage audio report is not part of it.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC."}, {"path": ["2", "text"], "text": "For video-freeze incident F, the encoder log records an end timestamp of 14:03:13.100 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC.", "negative_left": "For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC.", "negative_right": "For video-freeze incident F, the encoder log records an end timestamp of 14:03:12.000 UTC.", "right": "For video-freeze incident F, the encoder log records an end timestamp of 14:03:13.100 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-004", "id": "fast-41-diverse-219-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "During the public preshow countdown, staff logged video-freeze incident F for review. The event occurred in the public stream, not merely on a backstage monitor. No program material was lost while the incident was being checked, and the stream returned to normal afterward. A separate microphone-pop report was confined to the backstage intercom and did not appear in the public encoded output."}, {"speaker": "Streaming technician", "text": "For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC."}, {"speaker": "Streaming technician", "text": "For video-freeze incident F, the encoder log records an end timestamp of 14:03:13.100 UTC."}, {"speaker": "Event producer", "text": "The scheduled opening remains ahead, and the production team is monitoring the public feed. The review concerns only incident F; the backstage audio report is not part of it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all policy and bindings, both evidence spans are complete factual sentences, the counterfactual consistently changes only the freeze end time, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"During the public preshow countdown, staff logged video-freeze incident F for review. The event occurred in the public stream, not merely on a backstage monitor. No program material was lost while the incident was being checked, and the stream returned to normal afterward. A separate microphone-pop report was confined to the backstage intercom and did not appear in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"For video-freeze incident F, the encoder log records an end timestamp of 14:03:13.100 UTC.\"},{\"speaker\":\"Event producer\",\"text\":\"The scheduled opening remains ahead, and the production team is monitoring the public feed. The review concerns only incident F; the backstage audio report is not part of it.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC."}, {"path": ["2", "text"], "text": "For video-freeze incident F, the encoder log records an end timestamp of 14:03:13.100 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC.", "negative_left": "For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC.", "negative_right": "For video-freeze incident F, the encoder log records an end timestamp of 14:03:12.000 UTC.", "right": "For video-freeze incident F, the encoder log records an end timestamp of 14:03:13.100 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-004", "id": "fast-41-diverse-219-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "During the public preshow countdown, staff logged video-freeze incident F for review. The event occurred in the public stream, not merely on a backstage monitor. No program material was lost while the incident was being checked, and the stream returned to normal afterward. A separate microphone-pop report was confined to the backstage intercom and did not appear in the public encoded output."}, {"speaker": "Streaming technician", "text": "For video-freeze incident F, the encoder log records a start timestamp of 14:03:10.250 UTC."}, {"speaker": "Streaming technician", "text": "For video-freeze incident F, the encoder log records an end timestamp of 14:03:12.000 UTC."}, {"speaker": "Event producer", "text": "The scheduled opening remains ahead, and the production team is monitoring the public feed. The review concerns only incident F; the backstage audio report is not part of it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision scope. Both contexts bind to the verified public video incident and encoder timing. Base evidence: \"The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.\" \"The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC.\" Counterfactual evidence: \"The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.\" \"The encoder log records the end timestamp of video-freeze incident F as 20:14:04.900 UTC.\" The counterfactual changes only the end time and introduces no contradictory duplicate measurement. Neither context states a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Control-room log\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.\"},{\"speaker\":\"Control-room log\",\"text\":\"The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The freeze is visible in the public encoded stream, and the encoder record is the authoritative source for its timing rather than approximate observations.\"},{\"speaker\":\"Stage manager\",\"text\":\"The affected material was a preshow slate; the scheduled program continued without any program content being lost.\"},{\"speaker\":\"Event producer\",\"text\":\"The public stream remains under review before the opening, while the streaming technician prepares the incident handoff.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC."}, {"path": ["1", "text"], "text": "The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.", "negative_left": "The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.", "negative_right": "The encoder log records the end timestamp of video-freeze incident F as 20:14:04.900 UTC.", "right": "The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-009", "id": "fast-41-diverse-219-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Control-room log", "text": "The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC."}, {"speaker": "Control-room log", "text": "The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC."}, {"speaker": "Streaming technician", "text": "The freeze is visible in the public encoded stream, and the encoder record is the authoritative source for its timing rather than approximate observations."}, {"speaker": "Stage manager", "text": "The affected material was a preshow slate; the scheduled program continued without any program content being lost."}, {"speaker": "Event producer", "text": "The public stream remains under review before the opening, while the streaming technician prepares the incident handoff."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision scope. Both contexts bind to the verified public video incident and encoder timing. Base evidence: \"The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.\" \"The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC.\" Counterfactual evidence: \"The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.\" \"The encoder log records the end timestamp of video-freeze incident F as 20:14:04.900 UTC.\" The counterfactual changes only the end time and introduces no contradictory duplicate measurement. Neither context states a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Control-room log\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.\"},{\"speaker\":\"Control-room log\",\"text\":\"The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The freeze is visible in the public encoded stream, and the encoder record is the authoritative source for its timing rather than approximate observations.\"},{\"speaker\":\"Stage manager\",\"text\":\"The affected material was a preshow slate; the scheduled program continued without any program content being lost.\"},{\"speaker\":\"Event producer\",\"text\":\"The public stream remains under review before the opening, while the streaming technician prepares the incident handoff.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC."}, {"path": ["1", "text"], "text": "The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.", "negative_left": "The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC.", "negative_right": "The encoder log records the end timestamp of video-freeze incident F as 20:14:04.900 UTC.", "right": "The encoder log records the end timestamp of video-freeze incident F as 20:14:05.410 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-009", "id": "fast-41-diverse-219-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Control-room log", "text": "The encoder log records video-freeze incident F during the public preshow countdown, confirms incident F in the public encoded output, and records its start timestamp as 20:14:03.125 UTC."}, {"speaker": "Control-room log", "text": "The encoder log records the end timestamp of video-freeze incident F as 20:14:04.900 UTC."}, {"speaker": "Streaming technician", "text": "The freeze is visible in the public encoded stream, and the encoder record is the authoritative source for its timing rather than approximate observations."}, {"speaker": "Stage manager", "text": "The affected material was a preshow slate; the scheduled program continued without any program content being lost."}, {"speaker": "Event producer", "text": "The public stream remains under review before the opening, while the streaming technician prepares the incident handoff."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and criteria. Both contexts retain the same public video-freeze incident and relevant encoder-log bindings. The two evidence spans are complete factual sentences and match the base context. The counterfactual end time is after the start time and creates no duplicate contradiction. Neither context states a gold answer, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"During the public preshow countdown, the encoder audit identified video-freeze incident F. The incident is present in the public encoded output, and the encoder log is the controlling record for its timing. For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC.\"},{\"speaker\":\"Encoder auditor\",\"text\":\"For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:13.250 UTC.\"},{\"speaker\":\"Production reviewer\",\"text\":\"Review of the program track found that incident F did not cause any program content to be lost. The stream is otherwise stable, the opening is approaching, and no separate public audio defect was verified. Video-freeze incidents are assigned to the Streaming technician.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC."}, {"path": ["1", "text"], "text": "For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:13.250 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC.", "negative_left": "For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC.", "negative_right": "For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:12.000 UTC.", "right": "For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:13.250 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-010", "id": "fast-41-diverse-219-010-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "During the public preshow countdown, the encoder audit identified video-freeze incident F. The incident is present in the public encoded output, and the encoder log is the controlling record for its timing. For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC."}, {"speaker": "Encoder auditor", "text": "For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:13.250 UTC."}, {"speaker": "Production reviewer", "text": "Review of the program track found that incident F did not cause any program content to be lost. The stream is otherwise stable, the opening is approaching, and no separate public audio defect was verified. Video-freeze incidents are assigned to the Streaming technician."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and criteria. Both contexts retain the same public video-freeze incident and relevant encoder-log bindings. The two evidence spans are complete factual sentences and match the base context. The counterfactual end time is after the start time and creates no duplicate contradiction. Neither context states a gold answer, label rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"During the public preshow countdown, the encoder audit identified video-freeze incident F. The incident is present in the public encoded output, and the encoder log is the controlling record for its timing. For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC.\"},{\"speaker\":\"Encoder auditor\",\"text\":\"For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:13.250 UTC.\"},{\"speaker\":\"Production reviewer\",\"text\":\"Review of the program track found that incident F did not cause any program content to be lost. The stream is otherwise stable, the opening is approaching, and no separate public audio defect was verified. Video-freeze incidents are assigned to the Streaming technician.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC."}, {"path": ["1", "text"], "text": "For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:13.250 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC.", "negative_left": "For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC.", "negative_right": "For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:12.000 UTC.", "right": "For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:13.250 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-010", "id": "fast-41-diverse-219-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "During the public preshow countdown, the encoder audit identified video-freeze incident F. The incident is present in the public encoded output, and the encoder log is the controlling record for its timing. For video-freeze incident F, the encoder log records the start timestamp as 2026-09-17 14:00:10.125 UTC."}, {"speaker": "Encoder auditor", "text": "For video-freeze incident F, the encoder log records the end timestamp as 2026-09-17 14:00:12.000 UTC."}, {"speaker": "Production reviewer", "text": "Review of the program track found that incident F did not cause any program content to be lost. The stream is otherwise stable, the opening is approaching, and no separate public audio defect was verified. Video-freeze incidents are assigned to the Streaming technician."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, criteria, scope, and bindings; both contexts describe the same public video-freeze incident without answer leakage. The two focus spans are complete factual sentences, and the counterfactual changes only the end time consistently, yielding a 1.500-second incident without contradictory duplicate claims.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown, and the public encoded-output review confirms that F is present there.\"},{\"speaker\":\"Audit note\",\"text\":\"The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC.\"},{\"speaker\":\"Audit note\",\"text\":\"The encoder log for video-freeze incident F records its end timestamp as 14:03:12.500 UTC.\"},{\"speaker\":\"Continuity reviewer\",\"text\":\"The continuity record shows that video-freeze incident F did not cause program content to be lost.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The retained encoder export and public-output capture are available for review, and no separate public defect was identified.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC."}, {"path": ["2", "text"], "text": "The encoder log for video-freeze incident F records its end timestamp as 14:03:12.500 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC.", "negative_left": "The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC.", "negative_right": "The encoder log for video-freeze incident F records its end timestamp as 14:03:11.500 UTC.", "right": "The encoder log for video-freeze incident F records its end timestamp as 14:03:12.500 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-012", "id": "fast-41-diverse-219-012-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F during the public preshow countdown, and the public encoded-output review confirms that F is present there."}, {"speaker": "Audit note", "text": "The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC."}, {"speaker": "Audit note", "text": "The encoder log for video-freeze incident F records its end timestamp as 14:03:12.500 UTC."}, {"speaker": "Continuity reviewer", "text": "The continuity record shows that video-freeze incident F did not cause program content to be lost."}, {"speaker": "Streaming technician", "text": "The retained encoder export and public-output capture are available for review, and no separate public defect was identified."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, criteria, scope, and bindings; both contexts describe the same public video-freeze incident without answer leakage. The two focus spans are complete factual sentences, and the counterfactual changes only the end time consistently, yielding a 1.500-second incident without contradictory duplicate claims.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown, and the public encoded-output review confirms that F is present there.\"},{\"speaker\":\"Audit note\",\"text\":\"The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC.\"},{\"speaker\":\"Audit note\",\"text\":\"The encoder log for video-freeze incident F records its end timestamp as 14:03:12.500 UTC.\"},{\"speaker\":\"Continuity reviewer\",\"text\":\"The continuity record shows that video-freeze incident F did not cause program content to be lost.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The retained encoder export and public-output capture are available for review, and no separate public defect was identified.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC."}, {"path": ["2", "text"], "text": "The encoder log for video-freeze incident F records its end timestamp as 14:03:12.500 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC.", "negative_left": "The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC.", "negative_right": "The encoder log for video-freeze incident F records its end timestamp as 14:03:11.500 UTC.", "right": "The encoder log for video-freeze incident F records its end timestamp as 14:03:12.500 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-012", "id": "fast-41-diverse-219-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F during the public preshow countdown, and the public encoded-output review confirms that F is present there."}, {"speaker": "Audit note", "text": "The encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC."}, {"speaker": "Audit note", "text": "The encoder log for video-freeze incident F records its end timestamp as 14:03:11.500 UTC."}, {"speaker": "Continuity reviewer", "text": "The continuity record shows that video-freeze incident F did not cause program content to be lost."}, {"speaker": "Streaming technician", "text": "The retained encoder export and public-output capture are available for review, and no separate public defect was identified."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved by the unchanged questions object and neither context adds exceptions or defaults. The incident remains video-freeze F in the public preshow stream, preserving the relevant entity and path bindings. Evidence consists of two complete factual sentences: “The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown.” and “The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown.” The counterfactual changes only the end time to 14:03:11.742 UTC, which is coherent with the unchanged start time and other observations. Neither context states a gold answer, label, rule table, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoded-output review confirms that incident F is visible in the public stream, not merely in an internal preview or backstage monitor.\"},{\"speaker\":\"Production editor\",\"text\":\"Frame review found no removed, skipped, or obscured program content during incident F.\"},{\"speaker\":\"Event producer\",\"text\":\"The scheduled opening had not begun, the stream was stable afterward, and the incident was logged for the post-show production record.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown."}, {"path": ["1", "text"], "text": "The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown.", "negative_left": "The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown.", "negative_right": "The same encoder log records that video-freeze incident F ended at 14:03:11.742 UTC during the public preshow countdown.", "right": "The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-015", "id": "fast-41-diverse-219-015-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoded-output review confirms that incident F is visible in the public stream, not merely in an internal preview or backstage monitor."}, {"speaker": "Production editor", "text": "Frame review found no removed, skipped, or obscured program content during incident F."}, {"speaker": "Event producer", "text": "The scheduled opening had not begun, the stream was stable afterward, and the incident was logged for the post-show production record."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved by the unchanged questions object and neither context adds exceptions or defaults. The incident remains video-freeze F in the public preshow stream, preserving the relevant entity and path bindings. Evidence consists of two complete factual sentences: “The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown.” and “The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown.” The counterfactual changes only the end time to 14:03:11.742 UTC, which is coherent with the unchanged start time and other observations. Neither context states a gold answer, label, rule table, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoded-output review confirms that incident F is visible in the public stream, not merely in an internal preview or backstage monitor.\"},{\"speaker\":\"Production editor\",\"text\":\"Frame review found no removed, skipped, or obscured program content during incident F.\"},{\"speaker\":\"Event producer\",\"text\":\"The scheduled opening had not begun, the stream was stable afterward, and the incident was logged for the post-show production record.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown."}, {"path": ["1", "text"], "text": "The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown.", "negative_left": "The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown.", "negative_right": "The same encoder log records that video-freeze incident F ended at 14:03:11.742 UTC during the public preshow countdown.", "right": "The same encoder log records that video-freeze incident F ended at 14:03:12.347 UTC during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-015", "id": "fast-41-diverse-219-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The encoder log records that video-freeze incident F started at 14:03:10.000 UTC during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The same encoder log records that video-freeze incident F ended at 14:03:11.742 UTC during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoded-output review confirms that incident F is visible in the public stream, not merely in an internal preview or backstage monitor."}, {"speaker": "Production editor", "text": "Frame review found no removed, skipped, or obscured program content during incident F."}, {"speaker": "Event producer", "text": "The scheduled opening had not begun, the stream was stable afterward, and the incident was logged for the post-show production record."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and criteria. Both contexts describe the same public video-freeze incident and evidence path. The focus evidence contains two complete factual sentences. The counterfactual changes only the end time and remains internally consistent. Neither context states an answer, label, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Operations lead\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the end timestamp for video-freeze incident F as 14:22:12.487 UTC.\"},{\"speaker\":\"Production log\",\"text\":\"The review confirms that no program content was lost during incident F.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC."}, {"path": ["1", "text"], "text": "The encoder log records the end timestamp for video-freeze incident F as 14:22:12.487 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC.", "negative_left": "The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC.", "negative_right": "The encoder log records the end timestamp for video-freeze incident F as 14:22:11.987 UTC.", "right": "The encoder log records the end timestamp for video-freeze incident F as 14:22:12.487 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-016", "id": "fast-41-diverse-219-016-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Operations lead", "text": "The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC."}, {"speaker": "Streaming technician", "text": "The encoder log records the end timestamp for video-freeze incident F as 14:22:12.487 UTC."}, {"speaker": "Production log", "text": "The review confirms that no program content was lost during incident F."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and criteria. Both contexts describe the same public video-freeze incident and evidence path. The focus evidence contains two complete factual sentences. The counterfactual changes only the end time and remains internally consistent. Neither context states an answer, label, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Operations lead\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the end timestamp for video-freeze incident F as 14:22:12.487 UTC.\"},{\"speaker\":\"Production log\",\"text\":\"The review confirms that no program content was lost during incident F.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC."}, {"path": ["1", "text"], "text": "The encoder log records the end timestamp for video-freeze incident F as 14:22:12.487 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC.", "negative_left": "The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC.", "negative_right": "The encoder log records the end timestamp for video-freeze incident F as 14:22:11.987 UTC.", "right": "The encoder log records the end timestamp for video-freeze incident F as 14:22:12.487 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-016", "id": "fast-41-diverse-219-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Operations lead", "text": "The encoder log records video-freeze incident F during the public preshow countdown, shows incident F in the public encoded output, and records its start timestamp as 14:22:10.125 UTC."}, {"speaker": "Streaming technician", "text": "The encoder log records the end timestamp for video-freeze incident F as 14:22:11.987 UTC."}, {"speaker": "Production log", "text": "The review confirms that no program content was lost during incident F."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the public video-freeze scope, encoder-log authority, and relevant routing facts; the unchanged questions preserve all policy criteria. The two evidence quotes are complete factual sentences. The counterfactual consistently changes only F’s end time, yielding a shorter coherent incident without duplicate contradictions or embedded answer guidance.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"During the public preshow countdown, the encoder incident register identifies F as a video-freeze, and the public encoded-output review confirms that F is present in the stream. For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown. For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.701 UTC during the same public preshow countdown.\"},{\"speaker\":\"Production continuity lead\",\"text\":\"The continuity report records that F caused no program content to be lost. The public encoded-output review found no separate audio defect, and the stream remained stable after the countdown.\"},{\"speaker\":\"Stage manager\",\"text\":\"The scheduled opening follows shortly, and the Streaming technician is assigned to handle video-freeze incidents.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown."}, {"path": ["0", "text"], "text": "For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.701 UTC during the same public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown.", "negative_left": "For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown.", "negative_right": "For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.100 UTC during the same public preshow countdown.", "right": "For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.701 UTC during the same public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-017", "id": "fast-41-diverse-219-017-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "During the public preshow countdown, the encoder incident register identifies F as a video-freeze, and the public encoded-output review confirms that F is present in the stream. For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown. For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.701 UTC during the same public preshow countdown."}, {"speaker": "Production continuity lead", "text": "The continuity report records that F caused no program content to be lost. The public encoded-output review found no separate audio defect, and the stream remained stable after the countdown."}, {"speaker": "Stage manager", "text": "The scheduled opening follows shortly, and the Streaming technician is assigned to handle video-freeze incidents."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the public video-freeze scope, encoder-log authority, and relevant routing facts; the unchanged questions preserve all policy criteria. The two evidence quotes are complete factual sentences. The counterfactual consistently changes only F’s end time, yielding a shorter coherent incident without duplicate contradictions or embedded answer guidance.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"During the public preshow countdown, the encoder incident register identifies F as a video-freeze, and the public encoded-output review confirms that F is present in the stream. For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown. For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.701 UTC during the same public preshow countdown.\"},{\"speaker\":\"Production continuity lead\",\"text\":\"The continuity report records that F caused no program content to be lost. The public encoded-output review found no separate audio defect, and the stream remained stable after the countdown.\"},{\"speaker\":\"Stage manager\",\"text\":\"The scheduled opening follows shortly, and the Streaming technician is assigned to handle video-freeze incidents.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown."}, {"path": ["0", "text"], "text": "For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.701 UTC during the same public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown.", "negative_left": "For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown.", "negative_right": "For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.100 UTC during the same public preshow countdown.", "right": "For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.701 UTC during the same public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-017", "id": "fast-41-diverse-219-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "During the public preshow countdown, the encoder incident register identifies F as a video-freeze, and the public encoded-output review confirms that F is present in the stream. For video-freeze incident F, the encoder log records a start timestamp of 14:22:10.500 UTC during the public preshow countdown. For video-freeze incident F, the encoder log records an end timestamp of 14:22:12.100 UTC during the same public preshow countdown."}, {"speaker": "Production continuity lead", "text": "The continuity report records that F caused no program content to be lost. The public encoded-output review found no separate audio defect, and the stream remained stable after the countdown."}, {"speaker": "Stage manager", "text": "The scheduled opening follows shortly, and the Streaming technician is assigned to handle video-freeze incidents."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy for both contexts. The incident, public-output, encoder-record, and timing bindings remain consistent. Each focus-evidence item is a complete factual sentence. The counterfactual changes only the end timestamp, yielding a coherent two-second duration without duplicate contradictions. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC.\"},{\"speaker\":\"Case note\",\"text\":\"The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.500 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder archive is the controlling record for this review, rather than approximate backstage observations. The public encoded recording confirms that the affected material was a video freeze.\"},{\"speaker\":\"Program auditor\",\"text\":\"The program-content audit found no lost program content associated with incident F. The production log reports that the stream remained stable afterward and that the scheduled opening was still pending.\"},{\"speaker\":\"Production desk\",\"text\":\"The streaming technician is the assigned production contact for a video issue. An intercom-only microphone sound was excluded because no audio defect was detected in the public encoded output.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC."}, {"path": ["1", "text"], "text": "The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.500 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC.", "negative_left": "The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC.", "negative_right": "The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.000 UTC.", "right": "The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.500 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-020", "id": "fast-41-diverse-219-020-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC."}, {"speaker": "Case note", "text": "The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.500 UTC."}, {"speaker": "Streaming technician", "text": "The encoder archive is the controlling record for this review, rather than approximate backstage observations. The public encoded recording confirms that the affected material was a video freeze."}, {"speaker": "Program auditor", "text": "The program-content audit found no lost program content associated with incident F. The production log reports that the stream remained stable afterward and that the scheduled opening was still pending."}, {"speaker": "Production desk", "text": "The streaming technician is the assigned production contact for a video issue. An intercom-only microphone sound was excluded because no audio defect was detected in the public encoded output."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy for both contexts. The incident, public-output, encoder-record, and timing bindings remain consistent. Each focus-evidence item is a complete factual sentence. The counterfactual changes only the end timestamp, yielding a coherent two-second duration without duplicate contradictions. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC.\"},{\"speaker\":\"Case note\",\"text\":\"The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.500 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder archive is the controlling record for this review, rather than approximate backstage observations. The public encoded recording confirms that the affected material was a video freeze.\"},{\"speaker\":\"Program auditor\",\"text\":\"The program-content audit found no lost program content associated with incident F. The production log reports that the stream remained stable afterward and that the scheduled opening was still pending.\"},{\"speaker\":\"Production desk\",\"text\":\"The streaming technician is the assigned production contact for a video issue. An intercom-only microphone sound was excluded because no audio defect was detected in the public encoded output.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC."}, {"path": ["1", "text"], "text": "The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.500 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC.", "negative_left": "The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC.", "negative_right": "The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.000 UTC.", "right": "The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.500 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-020", "id": "fast-41-diverse-219-020-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The encoder log records video-freeze incident F during the public preshow countdown and identifies incident F in the public encoded output; the encoder-recorded start timestamp for incident F is 14:03:10.000 UTC."}, {"speaker": "Case note", "text": "The encoder-recorded end timestamp for video-freeze incident F is 14:03:12.000 UTC."}, {"speaker": "Streaming technician", "text": "The encoder archive is the controlling record for this review, rather than approximate backstage observations. The public encoded recording confirms that the affected material was a video freeze."}, {"speaker": "Program auditor", "text": "The program-content audit found no lost program content associated with incident F. The production log reports that the stream remained stable afterward and that the scheduled opening was still pending."}, {"speaker": "Production desk", "text": "The streaming technician is the assigned production contact for a video issue. An intercom-only microphone sound was excluded because no audio defect was detected in the public encoded output."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, decision scope, and criteria. The incident remains video-freeze incident F in the public preshow countdown on 2026-09-17. The evidence spans are complete factual sentences: \"On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown.\" and \"On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown.\" The counterfactual changes only the end time to 14:03:11.500 UTC and introduces no contradictory duplicate assertion. Neither context contains an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log identifies incident F as a video freeze, and review of the public encoded output confirms that the same incident is present there.\"},{\"speaker\":\"Production coordinator\",\"text\":\"The review found no loss of program content associated with incident F; the program feed continued without missing material after the freeze.\"},{\"speaker\":\"Stage manager\",\"text\":\"A backstage observer mentioned a microphone pop, but the public encoded audio contained no corresponding defect, so that observation is unrelated to incident F.\"},{\"speaker\":\"Event producer\",\"text\":\"The streaming technician retained the encoder record and public-output capture for the production file, while the event remained in the scheduled preshow period.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown."}, {"path": ["1", "text"], "text": "On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown.", "negative_left": "On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown.", "negative_right": "On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:11.500 UTC during the public preshow countdown.", "right": "On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-021", "id": "fast-41-diverse-219-021-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log identifies incident F as a video freeze, and review of the public encoded output confirms that the same incident is present there."}, {"speaker": "Production coordinator", "text": "The review found no loss of program content associated with incident F; the program feed continued without missing material after the freeze."}, {"speaker": "Stage manager", "text": "A backstage observer mentioned a microphone pop, but the public encoded audio contained no corresponding defect, so that observation is unrelated to incident F."}, {"speaker": "Event producer", "text": "The streaming technician retained the encoder record and public-output capture for the production file, while the event remained in the scheduled preshow period."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, decision scope, and criteria. The incident remains video-freeze incident F in the public preshow countdown on 2026-09-17. The evidence spans are complete factual sentences: \"On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown.\" and \"On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown.\" The counterfactual changes only the end time to 14:03:11.500 UTC and introduces no contradictory duplicate assertion. Neither context contains an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log identifies incident F as a video freeze, and review of the public encoded output confirms that the same incident is present there.\"},{\"speaker\":\"Production coordinator\",\"text\":\"The review found no loss of program content associated with incident F; the program feed continued without missing material after the freeze.\"},{\"speaker\":\"Stage manager\",\"text\":\"A backstage observer mentioned a microphone pop, but the public encoded audio contained no corresponding defect, so that observation is unrelated to incident F.\"},{\"speaker\":\"Event producer\",\"text\":\"The streaming technician retained the encoder record and public-output capture for the production file, while the event remained in the scheduled preshow period.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown."}, {"path": ["1", "text"], "text": "On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown.", "negative_left": "On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown.", "negative_right": "On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:11.500 UTC during the public preshow countdown.", "right": "On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:12.347 UTC during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-021", "id": "fast-41-diverse-219-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "On 2026-09-17, the encoder log for video-freeze incident F records its start timestamp as 14:03:10.000 UTC during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "On 2026-09-17, the encoder log for video-freeze incident F records its end timestamp as 14:03:11.500 UTC during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log identifies incident F as a video freeze, and review of the public encoded output confirms that the same incident is present there."}, {"speaker": "Production coordinator", "text": "The review found no loss of program content associated with incident F; the program feed continued without missing material after the freeze."}, {"speaker": "Stage manager", "text": "A backstage observer mentioned a microphone pop, but the public encoded audio contained no corresponding defect, so that observation is unrelated to incident F."}, {"speaker": "Event producer", "text": "The streaming technician retained the encoder record and public-output capture for the production file, while the event remained in the scheduled preshow period."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is unchanged and both contexts retain the encoder-log, public-output, impact, and routing scope. Entity, output-path, and incident-context bindings are preserved, while the timestamp change is an allowed observation change. The required evidence sentences are “The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC.” and “The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC.” The counterfactual coherently replaces the end time with “2026-09-17 20:14:10.400 UTC,” producing one consistent 1.900-second duration without duplicate measurements. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The overnight control-room review concerns video-freeze incident F during the public preshow countdown. The incident is recorded in the encoder log and confirmed in the public encoded output. The audit relies on encoder records rather than approximate observations. No program content was lost, and the scheduled material remained available afterward.\"},{\"speaker\":\"Audit record\",\"text\":\"The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC.\"},{\"speaker\":\"Audit record\",\"text\":\"The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC.\"},{\"speaker\":\"Case note\",\"text\":\"A separate microphone irregularity appeared only on the backstage intercom and is unrelated to incident F. The stream is stable, and any confirmed video issue is assigned to the Streaming technician.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC."}, {"path": ["2", "text"], "text": "The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC.", "negative_left": "The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC.", "negative_right": "The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.400 UTC.", "right": "The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-022", "id": "fast-41-diverse-219-022-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The overnight control-room review concerns video-freeze incident F during the public preshow countdown. The incident is recorded in the encoder log and confirmed in the public encoded output. The audit relies on encoder records rather than approximate observations. No program content was lost, and the scheduled material remained available afterward."}, {"speaker": "Audit record", "text": "The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC."}, {"speaker": "Audit record", "text": "The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC."}, {"speaker": "Case note", "text": "A separate microphone irregularity appeared only on the backstage intercom and is unrelated to incident F. The stream is stable, and any confirmed video issue is assigned to the Streaming technician."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is unchanged and both contexts retain the encoder-log, public-output, impact, and routing scope. Entity, output-path, and incident-context bindings are preserved, while the timestamp change is an allowed observation change. The required evidence sentences are “The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC.” and “The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC.” The counterfactual coherently replaces the end time with “2026-09-17 20:14:10.400 UTC,” producing one consistent 1.900-second duration without duplicate measurements. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"The overnight control-room review concerns video-freeze incident F during the public preshow countdown. The incident is recorded in the encoder log and confirmed in the public encoded output. The audit relies on encoder records rather than approximate observations. No program content was lost, and the scheduled material remained available afterward.\"},{\"speaker\":\"Audit record\",\"text\":\"The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC.\"},{\"speaker\":\"Audit record\",\"text\":\"The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC.\"},{\"speaker\":\"Case note\",\"text\":\"A separate microphone irregularity appeared only on the backstage intercom and is unrelated to incident F. The stream is stable, and any confirmed video issue is assigned to the Streaming technician.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC."}, {"path": ["2", "text"], "text": "The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC.", "negative_left": "The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC.", "negative_right": "The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.400 UTC.", "right": "The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.700 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-022", "id": "fast-41-diverse-219-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Case note", "text": "The overnight control-room review concerns video-freeze incident F during the public preshow countdown. The incident is recorded in the encoder log and confirmed in the public encoded output. The audit relies on encoder records rather than approximate observations. No program content was lost, and the scheduled material remained available afterward."}, {"speaker": "Audit record", "text": "The encoder-recorded start timestamp for video-freeze incident F is 2026-09-17 20:14:08.500 UTC."}, {"speaker": "Audit record", "text": "The encoder-recorded end timestamp for video-freeze incident F is 2026-09-17 20:14:10.400 UTC."}, {"speaker": "Case note", "text": "A separate microphone irregularity appeared only on the backstage intercom and is unrelated to incident F. The stream is stable, and any confirmed video issue is assigned to the Streaming technician."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, criteria, and scope. Both contexts retain the public preshow countdown, incident F, encoder ER-39, and public-output bindings. The required evidence quotes are “Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown.” and “In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp.” The counterfactual changes only the duration to 1.417 seconds without creating contradictory duplicate assertions. Neither context embeds an answer, code, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Operations coordinator\",\"text\":\"The review covers the public preshow countdown and relies on the encoder record rather than approximate witness impressions. Encoder record ER-39 is the record for the observed video-freeze incident during that countdown. Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown. The encoded-output review confirms that the freeze appears in the public stream viewed by audiences. In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp. The picture interruption did not remove or skip any program material; the program sequence continued normally. The event remained under active monitoring after the review, and the scheduled opening was still approaching. A backstage microphone artifact was confined to the intercom recording and was not present in the public encoded audio. The streaming technician retained responsibility for the picture issue, while no public-audio escalation was opened. No viewer complaints were recorded during the observation period.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown."}, {"path": ["0", "text"], "text": "In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown.", "negative_left": "Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown.", "negative_right": "In encoder record ER-39, the end timestamp is 1.417 seconds after the start timestamp.", "right": "In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-039", "id": "fast-41-diverse-219-039-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Operations coordinator", "text": "The review covers the public preshow countdown and relies on the encoder record rather than approximate witness impressions. Encoder record ER-39 is the record for the observed video-freeze incident during that countdown. Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown. The encoded-output review confirms that the freeze appears in the public stream viewed by audiences. In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp. The picture interruption did not remove or skip any program material; the program sequence continued normally. The event remained under active monitoring after the review, and the scheduled opening was still approaching. A backstage microphone artifact was confined to the intercom recording and was not present in the public encoded audio. The streaming technician retained responsibility for the picture issue, while no public-audio escalation was opened. No viewer complaints were recorded during the observation period."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, criteria, and scope. Both contexts retain the public preshow countdown, incident F, encoder ER-39, and public-output bindings. The required evidence quotes are “Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown.” and “In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp.” The counterfactual changes only the duration to 1.417 seconds without creating contradictory duplicate assertions. Neither context embeds an answer, code, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Operations coordinator\",\"text\":\"The review covers the public preshow countdown and relies on the encoder record rather than approximate witness impressions. Encoder record ER-39 is the record for the observed video-freeze incident during that countdown. Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown. The encoded-output review confirms that the freeze appears in the public stream viewed by audiences. In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp. The picture interruption did not remove or skip any program material; the program sequence continued normally. The event remained under active monitoring after the review, and the scheduled opening was still approaching. A backstage microphone artifact was confined to the intercom recording and was not present in the public encoded audio. The streaming technician retained responsibility for the picture issue, while no public-audio escalation was opened. No viewer complaints were recorded during the observation period.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown."}, {"path": ["0", "text"], "text": "In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown.", "negative_left": "Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown.", "negative_right": "In encoder record ER-39, the end timestamp is 1.417 seconds after the start timestamp.", "right": "In encoder record ER-39, the end timestamp is 2.417 seconds after the start timestamp."}, "verifier_independent_model": false}, "family": "fast-41-diverse-219-039", "id": "fast-41-diverse-219-039-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Operations coordinator", "text": "The review covers the public preshow countdown and relies on the encoder record rather than approximate witness impressions. Encoder record ER-39 is the record for the observed video-freeze incident during that countdown. Encoder record ER-39 is the record for video-freeze incident F during the public preshow countdown. The encoded-output review confirms that the freeze appears in the public stream viewed by audiences. In encoder record ER-39, the end timestamp is 1.417 seconds after the start timestamp. The picture interruption did not remove or skip any program material; the program sequence continued normally. The event remained under active monitoring after the review, and the scheduled opening was still approaching. A backstage microphone artifact was confined to the intercom recording and was not present in the public encoded audio. The streaming technician retained responsibility for the picture issue, while no public-audio escalation was opened. No viewer complaints were recorded during the observation period."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions and all governing policy, while changing only an allowed case observation. All Harbor Arts, incident, report, log, triage, technician, impact, and intervention bindings remain aligned. The evidence consists of two complete factual sentences: “The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault.” “At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault.” The counterfactual coherently changes the log evidence without duplicating contradictory measurements or assertions within itself. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is an allowed quantified proposition over the explicit set of submitted log evidence. The focus a2 is factual rather than policy-based. The base and counter assignments differ only on a2 and are jointly realizable: the same confidence-monitor cue disruption can occur either with or without equivalent submitted log evidence. The policy evidence accurately preserves the substantive rules originating in the original state, while the unchanged questions automatically preserve the decision criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for every required action: a2 supplies equivalent log evidence and therefore verification; a1 triggers streaming-technician routing; verified a2 plus cue disruption a3 requires intervention; and a4 plus a5 directly satisfies the L1 condition. No audience effect also excludes the stated L2 audience-degradation condition and is incompatible with broadcast loss.", "rule_index": 0, "sound": true}, {"reason": "Refuted a2 entails that no submitted log records a fault equivalent to the reported fault. Therefore the report cannot be treated as verified by equivalent log evidence, making at least one required action incorrect and satisfying the false criterion. This does not treat unknown as false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The Harbor Arts livestream incident is a confidence-monitor video timing fault."}, {"id": "a2", "statement": "Among all log evidence submitted for the Harbor Arts livestream incident, at least one log records a fault equivalent to the fault in the stage manager's report."}, {"id": "a3", "statement": "The Harbor Arts livestream incident disrupts the presenter's visual cues."}, {"id": "a4", "statement": "The Harbor Arts livestream incident causes internal disruption."}, {"id": "a5", "statement": "The Harbor Arts livestream incident has no audience effect."}], "base_state_json": "\"The Harbor Arts livestream incident involved a presenter who missed visual cues when the confidence display timing failed. The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault. At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault. The disruption remained within internal production operations, and the audience experienced no degradation. The event producer requested triage for the incident. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault."}, {"path": [], "text": "At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault."}], "policy_evidence": [{"path": [], "text": "Policy routes confidence-monitor video timing faults to the streaming technician."}, {"path": [], "text": "A report is verified when logs restate the same fault in equivalent words."}, {"path": [], "text": "Verified cue disruptions require intervention."}, {"path": [], "text": "Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}], "rules": [{"justification": "Equivalent log evidence verifies the report; the confidence-monitor video timing fault routes to the streaming technician; the verified cue disruption requires intervention; and internal disruption with no audience effect is L1.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "With no submitted log recording a fault equivalent to the reported fault, treating the report as verified is incorrect, so at least one required action is incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault.", "negative_left": "The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault.", "negative_right": "No submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault.", "right": "At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault."}, "verifier_independent_model": false}, "family": "fast-41-diverse-220-002", "id": "fast-41-diverse-220-002-base", "input": {"questions": {"decision": {"criteria": {"false": "At least one of these actions is incorrect: verification, streaming-technician routing, L1 impact, or intervention.", "true": "The report is verified by equivalent log evidence; route it to the streaming technician, rate it L1, and intervene."}, "instructions": "Decide whether the incident should be treated as verified, routed to the streaming technician, assigned L1 impact, and marked for intervention.", "type": "noul"}}, "state": "The Harbor Arts livestream incident involved a presenter who missed visual cues when the confidence display timing failed. The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault. At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault. The disruption remained within internal production operations, and the audience experienced no degradation. The event producer requested triage for the incident. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}, "method": "c2d", "provenance": {"source_id": "diverse-220", "source_is_synthetic": true, "source_sha256": "ba35b2ca55f4bdeb927287f576149f873d20519ea2ac4ac62e5acbb8a52c3730", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim questions and all governing policy, while changing only an allowed case observation. All Harbor Arts, incident, report, log, triage, technician, impact, and intervention bindings remain aligned. The evidence consists of two complete factual sentences: “The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault.” “At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault.” The counterfactual coherently changes the log evidence without duplicating contradictory measurements or assertions within itself. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is an allowed quantified proposition over the explicit set of submitted log evidence. The focus a2 is factual rather than policy-based. The base and counter assignments differ only on a2 and are jointly realizable: the same confidence-monitor cue disruption can occur either with or without equivalent submitted log evidence. The policy evidence accurately preserves the substantive rules originating in the original state, while the unchanged questions automatically preserve the decision criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for every required action: a2 supplies equivalent log evidence and therefore verification; a1 triggers streaming-technician routing; verified a2 plus cue disruption a3 requires intervention; and a4 plus a5 directly satisfies the L1 condition. No audience effect also excludes the stated L2 audience-degradation condition and is incompatible with broadcast loss.", "rule_index": 0, "sound": true}, {"reason": "Refuted a2 entails that no submitted log records a fault equivalent to the reported fault. Therefore the report cannot be treated as verified by equivalent log evidence, making at least one required action incorrect and satisfying the false criterion. This does not treat unknown as false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The Harbor Arts livestream incident is a confidence-monitor video timing fault."}, {"id": "a2", "statement": "Among all log evidence submitted for the Harbor Arts livestream incident, at least one log records a fault equivalent to the fault in the stage manager's report."}, {"id": "a3", "statement": "The Harbor Arts livestream incident disrupts the presenter's visual cues."}, {"id": "a4", "statement": "The Harbor Arts livestream incident causes internal disruption."}, {"id": "a5", "statement": "The Harbor Arts livestream incident has no audience effect."}], "base_state_json": "\"The Harbor Arts livestream incident involved a presenter who missed visual cues when the confidence display timing failed. The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault. At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault. The disruption remained within internal production operations, and the audience experienced no degradation. The event producer requested triage for the incident. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault."}, {"path": [], "text": "At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault."}], "policy_evidence": [{"path": [], "text": "Policy routes confidence-monitor video timing faults to the streaming technician."}, {"path": [], "text": "A report is verified when logs restate the same fault in equivalent words."}, {"path": [], "text": "Verified cue disruptions require intervention."}, {"path": [], "text": "Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}], "rules": [{"justification": "Equivalent log evidence verifies the report; the confidence-monitor video timing fault routes to the streaming technician; the verified cue disruption requires intervention; and internal disruption with no audience effect is L1.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "With no submitted log recording a fault equivalent to the reported fault, treating the report as verified is incorrect, so at least one required action is incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault.", "negative_left": "The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault.", "negative_right": "No submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault.", "right": "At least one submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault."}, "verifier_independent_model": false}, "family": "fast-41-diverse-220-002", "id": "fast-41-diverse-220-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "At least one of these actions is incorrect: verification, streaming-technician routing, L1 impact, or intervention.", "true": "The report is verified by equivalent log evidence; route it to the streaming technician, rate it L1, and intervene."}, "instructions": "Decide whether the incident should be treated as verified, routed to the streaming technician, assigned L1 impact, and marked for intervention.", "type": "noul"}}, "state": "The Harbor Arts livestream incident involved a presenter who missed visual cues when the confidence display timing failed. The stage manager's report for the Harbor Arts livestream incident identifies a confidence-monitor video timing fault. No submitted log for the Harbor Arts livestream incident records a confidence-monitor video timing fault. The disruption remained within internal production operations, and the audience experienced no degradation. The event producer requested triage for the incident. Policy routes confidence-monitor video timing faults to the streaming technician. A report is verified when logs restate the same fault in equivalent words. Verified cue disruptions require intervention. Impact is L1 for internal disruption with no audience effect, L2 for noticeable audience degradation, and L3 for broadcast loss."}, "method": "c2d", "provenance": {"source_id": "diverse-220", "source_is_synthetic": true, "source_sha256": "ba35b2ca55f4bdeb927287f576149f873d20519ea2ac4ac62e5acbb8a52c3730", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "All checks pass: the unchanged questions preserve the governing policy and bindings, the two evidence spans are complete factual sentences—“During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels.” and “During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels.”—and both contexts remain coherent without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels. During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels. Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05, and no audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold. The livestream remained usable throughout the event. No program-feed or stream-delivery impairment occurred other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05. Exactly one viewer reported audio loss. The encoder log recorded no stream-transport fault from 14:20 through 14:24, and the recorded source program-audio meter remained at normal speech level throughout that interval. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels."}, {"path": ["context"], "text": "During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels.", "negative_left": "During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels.", "negative_right": "During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 45 decibels.", "right": "During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels."}, "verifier_independent_model": false}, "family": "fast-41-diverse-221-011", "id": "fast-41-diverse-221-011-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels. During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels. Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05, and no audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold. The livestream remained usable throughout the event. No program-feed or stream-delivery impairment occurred other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05. Exactly one viewer reported audio loss. The encoder log recorded no stream-transport fault from 14:20 through 14:24, and the recorded source program-audio meter remained at normal speech level throughout that interval. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "All checks pass: the unchanged questions preserve the governing policy and bindings, the two evidence spans are complete factual sentences—“During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels.” and “During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels.”—and both contexts remain coherent without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels. During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels. Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05, and no audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold. The livestream remained usable throughout the event. No program-feed or stream-delivery impairment occurred other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05. Exactly one viewer reported audio loss. The encoder log recorded no stream-transport fault from 14:20 through 14:24, and the recorded source program-audio meter remained at normal speech level throughout that interval. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels."}, {"path": ["context"], "text": "During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels.", "negative_left": "During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels.", "negative_right": "During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 45 decibels.", "right": "During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 40 decibels."}, "verifier_independent_model": false}, "family": "fast-41-diverse-221-011", "id": "fast-41-diverse-221-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "During the scheduled museum-panel livestream, the calibrated audience-endpoint monitor recorded P-17's minimum delivered-program audio level as 42 decibels. During the scheduled museum-panel livestream, endpoint P-17's configured audibility threshold was 45 decibels. Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05, and no audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold. The livestream remained usable throughout the event. No program-feed or stream-delivery impairment occurred other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05. Exactly one viewer reported audio loss. The encoder log recorded no stream-transport fault from 14:20 through 14:24, and the recorded source program-audio meter remained at normal speech level throughout that interval. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy, question bindings, and scope; retain complete evidence sentences—“During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint.”, “Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB.”, “During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 68 dB at that endpoint.”, and “Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB.”—and the counterfactual is coherent without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"Case note: The scheduled museum-panel livestream remained usable throughout the event. Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05, and no audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold. No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05. Exactly one viewer reported audio loss. The encoder log recorded no stream-transport fault from 14:20 through 14:24, while the recorded source program-audio meter remained at normal speech level throughout that interval. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint.\",\"Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint."}, {"path": ["evidence", "1"], "text": "Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint.", "negative_left": "During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 68 dB at that endpoint.", "negative_right": "Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB.", "right": "Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB."}, "verifier_independent_model": false}, "family": "fast-41-diverse-221-013", "id": "fast-41-diverse-221-013-base", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "Case note: The scheduled museum-panel livestream remained usable throughout the event. Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05, and no audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold. No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05. Exactly one viewer reported audio loss. The encoder log recorded no stream-transport fault from 14:20 through 14:24, while the recorded source program-audio meter remained at normal speech level throughout that interval. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint.", "Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB."]}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy, question bindings, and scope; retain complete evidence sentences—“During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint.”, “Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB.”, “During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 68 dB at that endpoint.”, and “Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB.”—and the counterfactual is coherent without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; quantified or exception-bounded propositions such as a2, a3, and a5 remain atomic rather than bundling unrelated requirements. The focus a1 is a measurable factual relation, not a policy conclusion. The base and counter assignments are jointly realizable with only a1 changing: when a1 is supported, a2 may hold vacuously; when a1 is refuted, a2 confines the resulting below-threshold measurement while a3–a5 remain compatible. Policy evidence correctly preserves the substantive state-originating role bindings and verification/no-intervention policy. Rules contained in the unchanged questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that P-17 had no below-threshold delivered audio and that no other program-feed or delivery impairment occurred. Together with exactly one viewer report and corroborating normal encoder and source-audio records, this is sufficient for level 0 under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 entails at least one below-threshold delivered-audio measurement at P-17. The remaining conditions confine every such measurement to a one-second interval and one endpoint, establish continued usability, and exclude every other impairment. This is sufficient for a brief, localized verified issue at level 1 and excludes sustained, widespread, or content-blocking level-2 impact.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The minimum delivered-program audio level measured at audience endpoint P-17 during the scheduled museum-panel livestream was at least endpoint P-17's configured audibility threshold."}, {"id": "a2", "statement": "Every delivered-program audio measurement below endpoint P-17's configured audibility threshold during the scheduled museum-panel livestream was timestamped from 14:22:04 through 14:22:05."}, {"id": "a3", "statement": "No audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold during the scheduled museum-panel livestream."}, {"id": "a4", "statement": "The scheduled museum-panel livestream remained usable throughout the event."}, {"id": "a5", "statement": "No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05."}, {"id": "a6", "statement": "Exactly one viewer reported audio loss during the scheduled museum-panel livestream."}, {"id": "a7", "statement": "The encoder log recorded no stream-transport fault from 14:20 through 14:24 during the scheduled museum-panel livestream."}, {"id": "a8", "statement": "The recorded source program-audio meter remained at normal speech level from 14:20 through 14:24 during the scheduled museum-panel livestream."}], "base_state_json": "{\"context\":\"Case note: The scheduled museum-panel livestream remained usable throughout the event. Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05, and no audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold. No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05. Exactly one viewer reported audio loss. The encoder log recorded no stream-transport fault from 14:20 through 14:24, while the recorded source program-audio meter remained at normal speech level throughout that interval. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.\",\"evidence\":[\"During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint.\",\"Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint."}, {"path": ["evidence", "1"], "text": "Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB."}], "policy_evidence": [{"path": ["context"], "text": "The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes."}, {"path": ["context"], "text": "Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention."}], "rules": [{"justification": "Delivered-output monitoring shows no below-threshold audio, the encoder and source-audio records show no corroborating fault, and the competing-impact exclusion rules out any other program-feed or delivery impairment. The single report is therefore an isolated unverified complaint contradicted by monitoring and logs.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The refuted threshold comparison establishes delivered audio below P-17's configured audibility threshold. All such measurements are confined to a single one-second interval and to P-17, the stream remains usable, and no competing impairment occurred. This establishes a brief, localized issue while excluding sustained, widespread, and content-blocking impact.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 74 dB at that endpoint.", "negative_left": "During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 68 dB at that endpoint.", "negative_right": "Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB.", "right": "Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB."}, "verifier_independent_model": false}, "family": "fast-41-diverse-221-013", "id": "fast-41-diverse-221-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No verified broadcast impact: monitoring and logs contradict or fail to corroborate an isolated report. Route to the streaming technician for documentation and continued observation; make no live change.", "1 — Limited verified impact: evidence confirms a brief or localized delivery/program issue, but the stream remains usable. Route to the responsible technician or audio engineer for corrective action; producer notification is optional unless recurrence occurs.", "2 — Major verified impact: evidence confirms sustained, widespread, or content-blocking loss of program audio or stream delivery. Escalate immediately to the responsible engineer and event producer, with the stage manager coordinating an on-air pause or format change."], "instructions": "Apply the supplied policy and explicit evidence. Select exactly one indexed level; treat an unverified isolated complaint as insufficient to establish broadcast impact when logs and observations contradict it.", "type": "score"}}, "state": {"context": "Case note: The scheduled museum-panel livestream remained usable throughout the event. Every delivered-program audio measurement below endpoint P-17's configured audibility threshold was timestamped from 14:22:04 through 14:22:05, and no audience endpoint other than P-17 had a delivered-program audio measurement below its configured audibility threshold. No program-feed or stream-delivery impairment occurred during the scheduled museum-panel livestream other than any delivered-program audio below endpoint P-17's configured audibility threshold from 14:22:04 through 14:22:05. Exactly one viewer reported audio loss. The encoder log recorded no stream-transport fault from 14:20 through 14:24, while the recorded source program-audio meter remained at normal speech level throughout that interval. The streaming technician must verify delivery, the audio engineer owns program-audio faults, and the event producer authorizes on-air format changes. Production policy rates only verified program-feed or delivery impact; isolated reports contradicted by monitoring require documentation but no live intervention.", "evidence": ["During the scheduled museum-panel livestream, the calibrated audience-monitoring record for endpoint P-17 showed a minimum delivered-program audio level of 68 dB at that endpoint.", "Endpoint P-17's configured audibility threshold for the scheduled museum-panel livestream was 72 dB."]}}, "method": "c2d", "provenance": {"source_id": "diverse-221", "source_is_synthetic": true, "source_sha256": "f4945e994c173c991e9b9d019c354fbe087ea05aa516644450d47ae5fa8c5da5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and bindings are preserved, and neither context embeds an answer or output instruction. The relevant evidence quotes are \"The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.\" and \"The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21.\" The counterfactual changes HP-3’s checksum without asserting that it is the requested asset, so it remains coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies. The manifest identifies HP-3 as a project file that remains editable and links source media. It is neither raw captured sound nor a published final export. The requested asset belongs to the Harbor Echo transfer. Every transfer asset other than HP-3 is a published final export, is not raw captured sound, and is not editable.\",\"evidence\":[\"The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.\",\"The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21.\",\"The HP-3 record describes an editable project file linking source media; its record does not classify it as raw captured sound or a published final export.\",\"The transfer manifest places the requested file in the Harbor Echo transfer.\",\"For every Harbor Echo asset other than HP-3, the manifest records a published final export, not raw captured sound, and not an editable file.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21."}, {"path": ["evidence", "1"], "text": "The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.", "negative_left": "The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.", "negative_right": "The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 9C-48.", "right": "The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21."}, "verifier_independent_model": false}, "family": "fast-41-diverse-223-002", "id": "fast-41-diverse-223-002-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies. The manifest identifies HP-3 as a project file that remains editable and links source media. It is neither raw captured sound nor a published final export. The requested asset belongs to the Harbor Echo transfer. Every transfer asset other than HP-3 is a published final export, is not raw captured sound, and is not editable.", "evidence": ["The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.", "The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21.", "The HP-3 record describes an editable project file linking source media; its record does not classify it as raw captured sound or a published final export.", "The transfer manifest places the requested file in the Harbor Echo transfer.", "For every Harbor Echo asset other than HP-3, the manifest records a published final export, not raw captured sound, and not an editable file."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy and bindings are preserved, and neither context embeds an answer or output instruction. The relevant evidence quotes are \"The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.\" and \"The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21.\" The counterfactual changes HP-3’s checksum without asserting that it is the requested asset, so it remains coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies. The manifest identifies HP-3 as a project file that remains editable and links source media. It is neither raw captured sound nor a published final export. The requested asset belongs to the Harbor Echo transfer. Every transfer asset other than HP-3 is a published final export, is not raw captured sound, and is not editable.\",\"evidence\":[\"The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.\",\"The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21.\",\"The HP-3 record describes an editable project file linking source media; its record does not classify it as raw captured sound or a published final export.\",\"The transfer manifest places the requested file in the Harbor Echo transfer.\",\"For every Harbor Echo asset other than HP-3, the manifest records a published final export, not raw captured sound, and not an editable file.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21."}, {"path": ["evidence", "1"], "text": "The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.", "negative_left": "The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.", "negative_right": "The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 9C-48.", "right": "The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 7F-21."}, "verifier_independent_model": false}, "family": "fast-41-diverse-223-002", "id": "fast-41-diverse-223-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies. The manifest identifies HP-3 as a project file that remains editable and links source media. It is neither raw captured sound nor a published final export. The requested asset belongs to the Harbor Echo transfer. Every transfer asset other than HP-3 is a published final export, is not raw captured sound, and is not editable.", "evidence": ["The Harbor Echo manifest records harbor_mix.prproj as the requested asset with unique checksum 7F-21.", "The Harbor Echo manifest records the asset assigned transfer identifier HP-3 with unique checksum 9C-48.", "The HP-3 record describes an editable project file linking source media; its record does not classify it as raw captured sound or a published final export.", "The transfer manifest places the requested file in the Harbor Echo transfer.", "For every Harbor Echo asset other than HP-3, the manifest records a published final export, not raw captured sound, and not an editable file."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric and both contexts retain it. The request remains exactly about routing harbor_mix.prproj. Each context has two complete factual evidence sentences. Base evidence includes “The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.” and “The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3.” Counterfactual evidence includes “The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.” and “The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-205, and no other manifest record has transfer identifier HP-3.” The counterfactual uses distinct records without contradictory duplicate measurements or assertions. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies. The transfer manifest lists HP-3 as a project file that remains editable and links source media. HP-3 is not raw captured sound and is not a published final export. The requested file belongs to the Harbor Echo transfer. Every other asset in that transfer is a published final export, is not raw captured sound, and is not editable. A manifest cross-check recorded the following two identifying observations.\",\"evidence\":[\"The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.\",\"The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204."}, {"path": ["evidence", "1"], "text": "The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.", "negative_left": "The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.", "negative_right": "The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-205, and no other manifest record has transfer identifier HP-3.", "right": "The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-41-diverse-223-015", "id": "fast-41-diverse-223-015-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies. The transfer manifest lists HP-3 as a project file that remains editable and links source media. HP-3 is not raw captured sound and is not a published final export. The requested file belongs to the Harbor Echo transfer. Every other asset in that transfer is a published final export, is not raw captured sound, and is not editable. A manifest cross-check recorded the following two identifying observations.", "evidence": ["The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.", "The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full governing rubric and both contexts retain it. The request remains exactly about routing harbor_mix.prproj. Each context has two complete factual evidence sentences. Base evidence includes “The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.” and “The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3.” Counterfactual evidence includes “The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.” and “The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-205, and no other manifest record has transfer identifier HP-3.” The counterfactual uses distinct records without contradictory duplicate measurements or assertions. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies. The transfer manifest lists HP-3 as a project file that remains editable and links source media. HP-3 is not raw captured sound and is not a published final export. The requested file belongs to the Harbor Echo transfer. Every other asset in that transfer is a published final export, is not raw captured sound, and is not editable. A manifest cross-check recorded the following two identifying observations.\",\"evidence\":[\"The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.\",\"The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204."}, {"path": ["evidence", "1"], "text": "The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.", "negative_left": "The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.", "negative_right": "The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-205, and no other manifest record has transfer identifier HP-3.", "right": "The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-204, and no other manifest record has transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-41-diverse-223-015", "id": "fast-41-diverse-223-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies. The transfer manifest lists HP-3 as a project file that remains editable and links source media. HP-3 is not raw captured sound and is not a published final export. The requested file belongs to the Harbor Echo transfer. Every other asset in that transfer is a published final export, is not raw captured sound, and is not editable. A manifest cross-check recorded the following two identifying observations.", "evidence": ["The manifest identifies the requested asset harbor_mix.prproj as manifest record HX-204, and no other manifest record has code HX-204.", "The Harbor Echo manifest assigns transfer identifier HP-3 to manifest record HX-205, and no other manifest record has transfer identifier HP-3."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing scoring policy, and both contexts retain the same mandatory requirements and routing rule. Package Aurora_Interview and its ingestion scope remain bound consistently. The two evidence sentences are factual and exact: “The rights declaration governing the interview in Package Aurora_Interview bears status code R-17.” “In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration.” The counterfactual coherently changes only R-17’s recorded interpretation to absent clearance. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Intake clerk\",\"text\":\"Package Aurora_Interview contains WAV, MP4, TIFF, and project files; each file opens, and all submitted checksums match the manifest.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"The manifest supplies title, creator, date, collection code ORAL-7, formats, durations or dimensions, and matching filenames. The project file references all three media assets, the package identifier is stable, and the archive system marks the package usable.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"},{\"speaker\":\"Media librarian\",\"text\":\"No subject-keyword field, transcript, or equivalent description accompanied the submission.\"},{\"speaker\":\"Registry clerk\",\"text\":\"The rights declaration governing the interview in Package Aurora_Interview bears status code R-17.\"},{\"speaker\":\"Registry clerk\",\"text\":\"In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration.\"},{\"speaker\":\"Intake clerk\",\"text\":\"If the registry's interpretation of R-17 changed, that same recorded code would remain the item for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["4", "text"], "text": "The rights declaration governing the interview in Package Aurora_Interview bears status code R-17."}, {"path": ["5", "text"], "text": "In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The rights declaration governing the interview in Package Aurora_Interview bears status code R-17.", "negative_left": "The rights declaration governing the interview in Package Aurora_Interview bears status code R-17.", "negative_right": "In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes absent clearance for archive ingestion of the interview identified by its declaration.", "right": "In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration."}, "verifier_independent_model": false}, "family": "fast-41-diverse-227-010", "id": "fast-41-diverse-227-010-base", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Intake clerk", "text": "Package Aurora_Interview contains WAV, MP4, TIFF, and project files; each file opens, and all submitted checksums match the manifest."}, {"speaker": "Metadata specialist", "text": "The manifest supplies title, creator, date, collection code ORAL-7, formats, durations or dimensions, and matching filenames. The project file references all three media assets, the package identifier is stable, and the archive system marks the package usable."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}, {"speaker": "Media librarian", "text": "No subject-keyword field, transcript, or equivalent description accompanied the submission."}, {"speaker": "Registry clerk", "text": "The rights declaration governing the interview in Package Aurora_Interview bears status code R-17."}, {"speaker": "Registry clerk", "text": "In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration."}, {"speaker": "Intake clerk", "text": "If the registry's interpretation of R-17 changed, that same recorded code would remain the item for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing scoring policy, and both contexts retain the same mandatory requirements and routing rule. Package Aurora_Interview and its ingestion scope remain bound consistently. The two evidence sentences are factual and exact: “The rights declaration governing the interview in Package Aurora_Interview bears status code R-17.” “In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration.” The counterfactual coherently changes only R-17’s recorded interpretation to absent clearance. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Intake clerk\",\"text\":\"Package Aurora_Interview contains WAV, MP4, TIFF, and project files; each file opens, and all submitted checksums match the manifest.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"The manifest supplies title, creator, date, collection code ORAL-7, formats, durations or dimensions, and matching filenames. The project file references all three media assets, the package identifier is stable, and the archive system marks the package usable.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"},{\"speaker\":\"Media librarian\",\"text\":\"No subject-keyword field, transcript, or equivalent description accompanied the submission.\"},{\"speaker\":\"Registry clerk\",\"text\":\"The rights declaration governing the interview in Package Aurora_Interview bears status code R-17.\"},{\"speaker\":\"Registry clerk\",\"text\":\"In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration.\"},{\"speaker\":\"Intake clerk\",\"text\":\"If the registry's interpretation of R-17 changed, that same recorded code would remain the item for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["4", "text"], "text": "The rights declaration governing the interview in Package Aurora_Interview bears status code R-17."}, {"path": ["5", "text"], "text": "In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The rights declaration governing the interview in Package Aurora_Interview bears status code R-17.", "negative_left": "The rights declaration governing the interview in Package Aurora_Interview bears status code R-17.", "negative_right": "In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes absent clearance for archive ingestion of the interview identified by its declaration.", "right": "In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes affirmative clearance for archive ingestion of the interview identified by its declaration."}, "verifier_independent_model": false}, "family": "fast-41-diverse-227-010", "id": "fast-41-diverse-227-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Intake clerk", "text": "Package Aurora_Interview contains WAV, MP4, TIFF, and project files; each file opens, and all submitted checksums match the manifest."}, {"speaker": "Metadata specialist", "text": "The manifest supplies title, creator, date, collection code ORAL-7, formats, durations or dimensions, and matching filenames. The project file references all three media assets, the package identifier is stable, and the archive system marks the package usable."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}, {"speaker": "Media librarian", "text": "No subject-keyword field, transcript, or equivalent description accompanied the submission."}, {"speaker": "Registry clerk", "text": "The rights declaration governing the interview in Package Aurora_Interview bears status code R-17."}, {"speaker": "Registry clerk", "text": "In the archive-ingestion registry dated 17 September 2026, status code R-17 denotes absent clearance for archive ingestion of the interview identified by its declaration."}, {"speaker": "Intake clerk", "text": "If the registry's interpretation of R-17 changed, that same recorded code would remain the item for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy and bindings, the evidence is exactly “The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview.” and “The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration.”, and the counterfactual coherently changes authorization to denial without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Intake registrar\",\"text\":\"Package Aurora_Interview is identifiable under collection code ORAL-7, and its catalog record names the interview and supplying institution.\"},{\"speaker\":\"Digital archivist\",\"text\":\"The WAV, MP4, TIFF, and project file all open successfully. Recorded checksums match the manifest, and the project links to each submitted media asset.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"The manifest supplies title, creator, date, collection code, formats, durations or dimensions, and matching filenames. The package is usable for review.\"},{\"speaker\":\"Rights officer\",\"text\":\"The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview.\"},{\"speaker\":\"Rights officer\",\"text\":\"The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"No subject keywords accompany the package, and no transcript or equivalent description is supplied.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["3", "text"], "text": "The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview."}, {"path": ["4", "text"], "text": "The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview.", "negative_left": "The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview.", "negative_right": "The signed rights declaration dated 14 March 2026 records a denial of archive ingestion for the interview named in the declaration.", "right": "The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration."}, "verifier_independent_model": false}, "family": "fast-41-diverse-227-023", "id": "fast-41-diverse-227-023-base", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Intake registrar", "text": "Package Aurora_Interview is identifiable under collection code ORAL-7, and its catalog record names the interview and supplying institution."}, {"speaker": "Digital archivist", "text": "The WAV, MP4, TIFF, and project file all open successfully. Recorded checksums match the manifest, and the project links to each submitted media asset."}, {"speaker": "Metadata specialist", "text": "The manifest supplies title, creator, date, collection code, formats, durations or dimensions, and matching filenames. The package is usable for review."}, {"speaker": "Rights officer", "text": "The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview."}, {"speaker": "Rights officer", "text": "The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration."}, {"speaker": "Metadata specialist", "text": "No subject keywords accompany the package, and no transcript or equivalent description is supplied."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy and bindings, the evidence is exactly “The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview.” and “The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration.”, and the counterfactual coherently changes authorization to denial without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified file and link claims. The focus atom concerns the factual presence of affirmative rights clearance, not a policy classification. The base and counter assignments can differ only in the declaration’s substantive clearance while keeping all other facts fixed. The policy evidence correctly preserves the state-origin authority rule, mandatory-requirement list, and rights-review routing; the scoring criteria remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all five mandatory requirements pass and refutes both specified optional discovery aids. This entails level 2 and excludes level 3.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that exactly one mandatory requirement—affirmative rights clearance—fails, while the package remains identifiable and usable. This entails level 1; the preserved state policy assigns that blocker to rights review.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every file submitted in Package Aurora_Interview is readable."}, {"id": "a2", "statement": "Every checksum supplied for Package Aurora_Interview is valid."}, {"id": "a3", "statement": "Package Aurora_Interview has complete core metadata."}, {"id": "a4", "statement": "Every project-asset link in Package Aurora_Interview resolves to its corresponding submitted media asset."}, {"id": "a5", "statement": "The rights declaration governing the interview in Package Aurora_Interview affirmatively clears that interview for archive ingestion."}, {"id": "a6", "statement": "Package Aurora_Interview is identifiable."}, {"id": "a7", "statement": "Package Aurora_Interview is usable."}, {"id": "a8", "statement": "Package Aurora_Interview includes subject keywords."}, {"id": "a9", "statement": "Package Aurora_Interview includes a transcript or equivalent description."}], "base_state_json": "[{\"speaker\":\"Intake registrar\",\"text\":\"Package Aurora_Interview is identifiable under collection code ORAL-7, and its catalog record names the interview and supplying institution.\"},{\"speaker\":\"Digital archivist\",\"text\":\"The WAV, MP4, TIFF, and project file all open successfully. Recorded checksums match the manifest, and the project links to each submitted media asset.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"The manifest supplies title, creator, date, collection code, formats, durations or dimensions, and matching filenames. The package is usable for review.\"},{\"speaker\":\"Rights officer\",\"text\":\"The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview.\"},{\"speaker\":\"Rights officer\",\"text\":\"The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"No subject keywords accompany the package, and no transcript or equivalent description is supplied.\"},{\"speaker\":\"Collection manager\",\"text\":\"Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["3", "text"], "text": "The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview."}, {"path": ["4", "text"], "text": "The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}], "rules": [{"justification": "All five mandatory requirements pass, while both specified optional discovery aids are absent; this is Minimum-ready and excludes Fully enriched.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The package is identifiable and usable, and affirmative rights clearance is the only failed mandatory requirement; this is One-blocker review, with the governing policy assigning that blocker to rights review.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview.", "negative_left": "The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview.", "negative_right": "The signed rights declaration dated 14 March 2026 records a denial of archive ingestion for the interview named in the declaration.", "right": "The signed rights declaration dated 14 March 2026 records an affirmative authorization for archive ingestion of the interview named in the declaration."}, "verifier_independent_model": false}, "family": "fast-41-diverse-227-023", "id": "fast-41-diverse-227-023-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable or seriously deficient: files are unreadable or checksums fail, core identity is missing, or at least two mandatory requirements fail; route to package remediation.", "1 — One-blocker review: the package is identifiable and usable but fails exactly one mandatory requirement; do not ingest and route to the reviewer responsible for that requirement.", "2 — Minimum-ready: every mandatory requirement passes, though optional enhancements such as keywords or transcripts are missing; approve standard ingestion.", "3 — Fully enriched: every mandatory requirement passes and optional discovery aids, including subject keywords and a transcript or equivalent description, are present; approve priority ingestion."], "instructions": "Score ingestion readiness using the ordered levels below. Apply the stated evidence-authority policy, then select the single best level and route.", "type": "score"}}, "state": [{"speaker": "Intake registrar", "text": "Package Aurora_Interview is identifiable under collection code ORAL-7, and its catalog record names the interview and supplying institution."}, {"speaker": "Digital archivist", "text": "The WAV, MP4, TIFF, and project file all open successfully. Recorded checksums match the manifest, and the project links to each submitted media asset."}, {"speaker": "Metadata specialist", "text": "The manifest supplies title, creator, date, collection code, formats, durations or dimensions, and matching filenames. The package is usable for review."}, {"speaker": "Rights officer", "text": "The signed rights declaration dated 14 March 2026 is the governing rights document for the interview in Package Aurora_Interview."}, {"speaker": "Rights officer", "text": "The signed rights declaration dated 14 March 2026 records a denial of archive ingestion for the interview named in the declaration."}, {"speaker": "Metadata specialist", "text": "No subject keywords accompany the package, and no transcript or equivalent description is supplied."}, {"speaker": "Collection manager", "text": "Policy treats file contents as authoritative over filenames. Mandatory requirements are readable files, valid checksums, complete core metadata, linked project assets, and affirmative rights clearance. A package failing exactly one mandatory requirement goes to rights review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-227", "source_is_synthetic": true, "source_sha256": "650a4c3fc5ede6d3e71128e75bdc09f187189c5b40e8e3d9cb74ad747f6ce806", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions object is present verbatim and both contexts retain the donor’s conditional access policy. Question bindings are preserved because Harbor Day, the package, the video, and timestamp 00:04:12 remain identified consistently. Evidence consists of the complete factual sentences “The Harbor Day event occurred on 14 June 2024.” and “The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006.” Counterfactual observations are coherent because a person born on 14 June 2006 was 18 on the 14 June 2024 event date, and the changed derivative availability does not contradict the changed observations. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"Case note: The Harbor Day event occurred on 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The person is identifiable and unblurred, while every other person appearing in the video is blurred. The archivist verified the frame time against the preservation video and checked the event register against the birth record. The Harbor Day package contains an MP4 video, WAV audio, JPG poster, and PRPROJ project file; every required file is playable and uncorrupted, has an attached checksum, and matches its manifest checksum. Each file is reliably identified as part of the package. The authoritative metadata record supplies title, creator, event date, rights holder, and each file format, and the authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and a usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event occurred on 14 June 2024."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event occurred on 14 June 2024.", "negative_left": "The Harbor Day event occurred on 14 June 2024.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-001", "id": "fast-41-diverse-228-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "Case note: The Harbor Day event occurred on 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The person is identifiable and unblurred, while every other person appearing in the video is blurred. The archivist verified the frame time against the preservation video and checked the event register against the birth record. The Harbor Day package contains an MP4 video, WAV audio, JPG poster, and PRPROJ project file; every required file is playable and uncorrupted, has an attached checksum, and matches its manifest checksum. Each file is reliably identified as part of the package. The authoritative metadata record supplies title, creator, event date, rights holder, and each file format, and the authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and a usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions object is present verbatim and both contexts retain the donor’s conditional access policy. Question bindings are preserved because Harbor Day, the package, the video, and timestamp 00:04:12 remain identified consistently. Evidence consists of the complete factual sentences “The Harbor Day event occurred on 14 June 2024.” and “The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006.” Counterfactual observations are coherent because a person born on 14 June 2006 was 18 on the 14 June 2024 event date, and the changed derivative availability does not contradict the changed observations. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"Case note: The Harbor Day event occurred on 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The person is identifiable and unblurred, while every other person appearing in the video is blurred. The archivist verified the frame time against the preservation video and checked the event register against the birth record. The Harbor Day package contains an MP4 video, WAV audio, JPG poster, and PRPROJ project file; every required file is playable and uncorrupted, has an attached checksum, and matches its manifest checksum. Each file is reliably identified as part of the package. The authoritative metadata record supplies title, creator, event date, rights holder, and each file format, and the authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and a usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event occurred on 14 June 2024."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event occurred on 14 June 2024.", "negative_left": "The Harbor Day event occurred on 14 June 2024.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-001", "id": "fast-41-diverse-228-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "Case note: The Harbor Day event occurred on 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006. The person is identifiable and unblurred, while every other person appearing in the video is blurred. The archivist verified the frame time against the preservation video and checked the event register against the birth record. The Harbor Day package contains an MP4 video, WAV audio, JPG poster, and PRPROJ project file; every required file is playable and uncorrupted, has an attached checksum, and matches its manifest checksum. Each file is reliably identified as part of the package. The authoritative metadata record supplies title, creator, event date, rights holder, and each file format, and the authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and a usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the unchanged questions and governing policy. The Harbor Day, timestamp, subject, and event-date bindings are preserved. The two evidence spans are complete factual sentences: “The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed.” “The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007.” The counterfactual changes only the birth date to 14 June 2006, which is coherent with the other unchanged facts. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed. The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007. The subject is identifiable and unblurred, while every other person appearing in the video is blurred. All required preservation files—the MP4 video, WAV audio, JPG poster, and PRPROJ editing file—are playable and uncorrupted; each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, every required file is reliably identified as belonging to the Harbor Day package, and a usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed.", "negative_left": "The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-007", "id": "fast-41-diverse-228-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed. The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007. The subject is identifiable and unblurred, while every other person appearing in the video is blurred. All required preservation files—the MP4 video, WAV audio, JPG poster, and PRPROJ editing file—are playable and uncorrupted; each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, every required file is reliably identified as belonging to the Harbor Day package, and a usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the unchanged questions and governing policy. The Harbor Day, timestamp, subject, and event-date bindings are preserved. The two evidence spans are complete factual sentences: “The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed.” “The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007.” The counterfactual changes only the birth date to 14 June 2006, which is coherent with the other unchanged facts. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed. The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007. The subject is identifiable and unblurred, while every other person appearing in the video is blurred. All required preservation files—the MP4 video, WAV audio, JPG poster, and PRPROJ editing file—are playable and uncorrupted; each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, every required file is reliably identified as belonging to the Harbor Day package, and a usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed.", "negative_left": "The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 16 June 2007."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-007", "id": "fast-41-diverse-228-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day event date is 15 June 2024, and the person shown at 00:04:12 in the Harbor Day preservation video is the subject whose age is assessed. The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006. The subject is identifiable and unblurred, while every other person appearing in the video is blurred. All required preservation files—the MP4 video, WAV audio, JPG poster, and PRPROJ editing file—are playable and uncorrupted; each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, every required file is reliably identified as belonging to the Harbor Day package, and a usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the donor instruction and relevant requirements. The question bindings remain Harbor Day, 00:04:12, the preservation video, the person shown, and the event date. The evidence consists of exactly two complete factual sentences: “The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15.” and “The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03.” The counterfactual birth date makes the person an adult in 2024 and creates no contradictory duplicate assertion. Neither context states a gold answer, answer code, rationale, proposition identifier, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains a preservation video, audio, poster, and project file. Every required preservation file is playable and uncorrupted; each has an attached checksum, and every checksum matches the manifest. The authoritative metadata record supplies title, creator, event date, and rights-holder values, and records the format of every file. The authoritative donor record supplies the applicable access instruction. The rights permit archival preservation, and every required preservation file is reliably identified as belonging to the Harbor Day package. The package has a usable public-access derivative. At 00:04:12 in the preservation video, one person is identifiable and unblurred; every other person appearing in the video is blurred. The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15. The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15."}, {"path": [], "text": "The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15.", "negative_left": "The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15.", "negative_right": "The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2000-09-03.", "right": "The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-010", "id": "fast-41-diverse-228-010-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains a preservation video, audio, poster, and project file. Every required preservation file is playable and uncorrupted; each has an attached checksum, and every checksum matches the manifest. The authoritative metadata record supplies title, creator, event date, and rights-holder values, and records the format of every file. The authoritative donor record supplies the applicable access instruction. The rights permit archival preservation, and every required preservation file is reliably identified as belonging to the Harbor Day package. The package has a usable public-access derivative. At 00:04:12 in the preservation video, one person is identifiable and unblurred; every other person appearing in the video is blurred. The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15. The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the donor instruction and relevant requirements. The question bindings remain Harbor Day, 00:04:12, the preservation video, the person shown, and the event date. The evidence consists of exactly two complete factual sentences: “The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15.” and “The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03.” The counterfactual birth date makes the person an adult in 2024 and creates no contradictory duplicate assertion. Neither context states a gold answer, answer code, rationale, proposition identifier, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains a preservation video, audio, poster, and project file. Every required preservation file is playable and uncorrupted; each has an attached checksum, and every checksum matches the manifest. The authoritative metadata record supplies title, creator, event date, and rights-holder values, and records the format of every file. The authoritative donor record supplies the applicable access instruction. The rights permit archival preservation, and every required preservation file is reliably identified as belonging to the Harbor Day package. The package has a usable public-access derivative. At 00:04:12 in the preservation video, one person is identifiable and unblurred; every other person appearing in the video is blurred. The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15. The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15."}, {"path": [], "text": "The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15.", "negative_left": "The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15.", "negative_right": "The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2000-09-03.", "right": "The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2010-09-03."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-010", "id": "fast-41-diverse-228-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains a preservation video, audio, poster, and project file. Every required preservation file is playable and uncorrupted; each has an attached checksum, and every checksum matches the manifest. The authoritative metadata record supplies title, creator, event date, and rights-holder values, and records the format of every file. The authoritative donor record supplies the applicable access instruction. The rights permit archival preservation, and every required preservation file is reliably identified as belonging to the Harbor Day package. The package has a usable public-access derivative. At 00:04:12 in the preservation video, one person is identifiable and unblurred; every other person appearing in the video is blurred. The Harbor Day event date for the person shown at 00:04:12 in the Harbor Day preservation video is 2024-06-15. The birth date recorded for the person shown at 00:04:12 in the Harbor Day preservation video is 2000-09-03. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the donor policy and all governing criteria in the unchanged questions object. The question bindings remain Harbor Day, the preservation video, 00:04:12, and the event date. The two evidence quotes are complete factual sentences: \"The Harbor Day event date was 14 June 2024.\" and \"The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006.\" The counterfactual coherently changes the birth date to 14 June 2006, making the person eighteen on the event date without duplicating or contradicting measurements. Neither context states a label, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ project file. The Harbor Day event date was 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The person shown at 00:04:12 is identifiable and unblurred, while every other person appearing in the video is blurred. Every required preservation file is playable and uncorrupted, has an attached checksum, and has a checksum matching the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as belonging to the package. A usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date was 14 June 2024."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date was 14 June 2024.", "negative_left": "The Harbor Day event date was 14 June 2024.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-014", "id": "fast-41-diverse-228-014-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ project file. The Harbor Day event date was 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The person shown at 00:04:12 is identifiable and unblurred, while every other person appearing in the video is blurred. Every required preservation file is playable and uncorrupted, has an attached checksum, and has a checksum matching the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as belonging to the package. A usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the donor policy and all governing criteria in the unchanged questions object. The question bindings remain Harbor Day, the preservation video, 00:04:12, and the event date. The two evidence quotes are complete factual sentences: \"The Harbor Day event date was 14 June 2024.\" and \"The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006.\" The counterfactual coherently changes the birth date to 14 June 2006, making the person eighteen on the event date without duplicating or contradicting measurements. Neither context states a label, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ project file. The Harbor Day event date was 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The person shown at 00:04:12 is identifiable and unblurred, while every other person appearing in the video is blurred. Every required preservation file is playable and uncorrupted, has an attached checksum, and has a checksum matching the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as belonging to the package. A usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date was 14 June 2024."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date was 14 June 2024.", "negative_left": "The Harbor Day event date was 14 June 2024.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-014", "id": "fast-41-diverse-228-014-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ project file. The Harbor Day event date was 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 14 June 2006. The person shown at 00:04:12 is identifiable and unblurred, while every other person appearing in the video is blurred. Every required preservation file is playable and uncorrupted, has an attached checksum, and has a checksum matching the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as belonging to the package. A usable public-access derivative is available. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing criteria without adding exceptions or missing-evidence defaults. Both preserve the Harbor Day package, the 00:04:12 video path, the person, and the 10 February 2025 event binding. The evidence spans are complete factual sentences: “The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008.” and “The Harbor Day event date recorded for that preservation video is 10 February 2025.” The counterfactual’s 15 June 2006 birth date is coherent with the unchanged event date and introduces no contradictory duplicate assertion. Neither context states an answer level, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. Every required preservation file is playable and uncorrupted, has an attached checksum matching the manifest, and is reliably identified as belonging to the package. The authoritative metadata record supplies the title, creator, event date, rights holder, and each file’s format. The authoritative donor record supplies the applicable access instruction, and the rights permit archival preservation. The video frame at 00:04:12 shows an identifiable, unblurred person; every other person appearing in the video is blurred. The package has a usable public-access derivative. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008. The Harbor Day event date recorded for that preservation video is 10 February 2025. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008."}, {"path": [], "text": "The Harbor Day event date recorded for that preservation video is 10 February 2025."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008.", "negative_left": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006.", "negative_right": "The Harbor Day event date recorded for that preservation video is 10 February 2025.", "right": "The Harbor Day event date recorded for that preservation video is 10 February 2025."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-018", "id": "fast-41-diverse-228-018-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. Every required preservation file is playable and uncorrupted, has an attached checksum matching the manifest, and is reliably identified as belonging to the package. The authoritative metadata record supplies the title, creator, event date, rights holder, and each file’s format. The authoritative donor record supplies the applicable access instruction, and the rights permit archival preservation. The video frame at 00:04:12 shows an identifiable, unblurred person; every other person appearing in the video is blurred. The package has a usable public-access derivative. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008. The Harbor Day event date recorded for that preservation video is 10 February 2025. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing criteria without adding exceptions or missing-evidence defaults. Both preserve the Harbor Day package, the 00:04:12 video path, the person, and the 10 February 2025 event binding. The evidence spans are complete factual sentences: “The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008.” and “The Harbor Day event date recorded for that preservation video is 10 February 2025.” The counterfactual’s 15 June 2006 birth date is coherent with the unchanged event date and introduces no contradictory duplicate assertion. Neither context states an answer level, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. Every required preservation file is playable and uncorrupted, has an attached checksum matching the manifest, and is reliably identified as belonging to the package. The authoritative metadata record supplies the title, creator, event date, rights holder, and each file’s format. The authoritative donor record supplies the applicable access instruction, and the rights permit archival preservation. The video frame at 00:04:12 shows an identifiable, unblurred person; every other person appearing in the video is blurred. The package has a usable public-access derivative. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008. The Harbor Day event date recorded for that preservation video is 10 February 2025. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008."}, {"path": [], "text": "The Harbor Day event date recorded for that preservation video is 10 February 2025."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2008.", "negative_left": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006.", "negative_right": "The Harbor Day event date recorded for that preservation video is 10 February 2025.", "right": "The Harbor Day event date recorded for that preservation video is 10 February 2025."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-018", "id": "fast-41-diverse-228-018-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. Every required preservation file is playable and uncorrupted, has an attached checksum matching the manifest, and is reliably identified as belonging to the package. The authoritative metadata record supplies the title, creator, event date, rights holder, and each file’s format. The authoritative donor record supplies the applicable access instruction, and the rights permit archival preservation. The video frame at 00:04:12 shows an identifiable, unblurred person; every other person appearing in the video is blurred. The package has a usable public-access derivative. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The Harbor Day event date recorded for that preservation video is 10 February 2025. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, use the exact factual evidence “The Harbor Day event date is 14 June 2024.” and “The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008.”, and the counterfactual coherently changes only that birth-date observation without embedding an answer or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ project file. The Harbor Day event date is 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008. At that timestamp, the person is identifiable and unblurred; every other person appearing in the video is blurred. Every required preservation file is playable and uncorrupted, has an attached checksum, and has a checksum matching the manifest. Each required file is reliably identified as belonging to the package. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. The governing rights permit archival preservation. A usable public-access derivative is included. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date is 14 June 2024."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date is 14 June 2024.", "negative_left": "The Harbor Day event date is 14 June 2024.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 1990.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-020", "id": "fast-41-diverse-228-020-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ project file. The Harbor Day event date is 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008. At that timestamp, the person is identifiable and unblurred; every other person appearing in the video is blurred. Every required preservation file is playable and uncorrupted, has an attached checksum, and has a checksum matching the manifest. Each required file is reliably identified as belonging to the package. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. The governing rights permit archival preservation. A usable public-access derivative is included. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and bindings, use the exact factual evidence “The Harbor Day event date is 14 June 2024.” and “The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008.”, and the counterfactual coherently changes only that birth-date observation without embedding an answer or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ project file. The Harbor Day event date is 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008. At that timestamp, the person is identifiable and unblurred; every other person appearing in the video is blurred. Every required preservation file is playable and uncorrupted, has an attached checksum, and has a checksum matching the manifest. Each required file is reliably identified as belonging to the package. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. The governing rights permit archival preservation. A usable public-access derivative is included. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date is 14 June 2024."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date is 14 June 2024.", "negative_left": "The Harbor Day event date is 14 June 2024.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 1990.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 2008."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-020", "id": "fast-41-diverse-228-020-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ project file. The Harbor Day event date is 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 3 September 1990. At that timestamp, the person is identifiable and unblurred; every other person appearing in the video is blurred. Every required preservation file is playable and uncorrupted, has an attached checksum, and has a checksum matching the manifest. Each required file is reliably identified as belonging to the package. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. The governing rights permit archival preservation. A usable public-access derivative is included. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions object retains all governing criteria and neither context adds exceptions. Question bindings remain Harbor Day, the 14 June 2024 event date, and the person at 00:04:12. The evidence spans are complete factual sentences: “The Harbor Day event date was 14 June 2024.” and “The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006.” The counterfactual is coherent because only the birth date changes, making the person an adult on the unchanged event date. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, WAV audio, JPG poster, and PRPROJ project file. The Harbor Day event date was 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The frame permits identification of that person and shows that person without blur; every other person appearing in the video is blurred. A review log marks every required preservation file playable and uncorrupted. Each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as part of the Harbor Day package. A usable public-access derivative is present. The donor’s access instruction is recorded exactly as follows: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” The person’s birth date is recorded in the same authoritative case materials, while the event date is the date used for age review.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date was 14 June 2024."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date was 14 June 2024.", "negative_left": "The Harbor Day event date was 14 June 2024.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2005.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-022", "id": "fast-41-diverse-228-022-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, WAV audio, JPG poster, and PRPROJ project file. The Harbor Day event date was 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The frame permits identification of that person and shows that person without blur; every other person appearing in the video is blurred. A review log marks every required preservation file playable and uncorrupted. Each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as part of the Harbor Day package. A usable public-access derivative is present. The donor’s access instruction is recorded exactly as follows: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” The person’s birth date is recorded in the same authoritative case materials, while the event date is the date used for age review."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions object retains all governing criteria and neither context adds exceptions. Question bindings remain Harbor Day, the 14 June 2024 event date, and the person at 00:04:12. The evidence spans are complete factual sentences: “The Harbor Day event date was 14 June 2024.” and “The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006.” The counterfactual is coherent because only the birth date changes, making the person an adult on the unchanged event date. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, WAV audio, JPG poster, and PRPROJ project file. The Harbor Day event date was 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006. The frame permits identification of that person and shows that person without blur; every other person appearing in the video is blurred. A review log marks every required preservation file playable and uncorrupted. Each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as part of the Harbor Day package. A usable public-access derivative is present. The donor’s access instruction is recorded exactly as follows: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” The person’s birth date is recorded in the same authoritative case materials, while the event date is the date used for age review.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event date was 14 June 2024."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event date was 14 June 2024.", "negative_left": "The Harbor Day event date was 14 June 2024.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2005.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2006."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-022", "id": "fast-41-diverse-228-022-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, WAV audio, JPG poster, and PRPROJ project file. The Harbor Day event date was 14 June 2024. The person shown at 00:04:12 in the Harbor Day preservation video was born on 15 June 2005. The frame permits identification of that person and shows that person without blur; every other person appearing in the video is blurred. A review log marks every required preservation file playable and uncorrupted. Each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as part of the Harbor Day package. A usable public-access derivative is present. The donor’s access instruction is recorded exactly as follows: The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” The person’s birth date is recorded in the same authoritative case materials, while the event date is the date used for age review."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions object retains all criteria and both contexts retain the donor instruction and required case facts. Question bindings are preserved because both contexts retain the Harbor Day package, HD-417, the preservation video, and timestamp 00:04:12. Evidence consists of two complete factual sentences: “The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417.” and “Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026.” The counterfactual is coherent because changing the birth date makes the identified person an adult by the unchanged event date without creating contradictory duplicate claims. Neither context embeds an answer, label, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The digital archivist’s case note covers the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417. Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026. The frame shows that person clearly and without blur; every other person appearing in the preservation video is blurred. All required preservation files play correctly, are uncorrupted, and have attached checksums matching the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, each required file is reliably tied to this Harbor Day package, and a usable public-access derivative is present.\\n\\nThe donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417."}, {"path": [], "text": "Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417.", "negative_left": "The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417.", "negative_right": "Participant record HD-417 lists the person's birth date as 14 May 2007 and the Harbor Day event date as 22 June 2026.", "right": "Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-025", "id": "fast-41-diverse-228-025-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The digital archivist’s case note covers the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417. Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026. The frame shows that person clearly and without blur; every other person appearing in the preservation video is blurred. All required preservation files play correctly, are uncorrupted, and have attached checksums matching the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, each required file is reliably tied to this Harbor Day package, and a usable public-access derivative is present.\n\nThe donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged questions object retains all criteria and both contexts retain the donor instruction and required case facts. Question bindings are preserved because both contexts retain the Harbor Day package, HD-417, the preservation video, and timestamp 00:04:12. Evidence consists of two complete factual sentences: “The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417.” and “Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026.” The counterfactual is coherent because changing the birth date makes the identified person an adult by the unchanged event date without creating contradictory duplicate claims. Neither context embeds an answer, label, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The digital archivist’s case note covers the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417. Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026. The frame shows that person clearly and without blur; every other person appearing in the preservation video is blurred. All required preservation files play correctly, are uncorrupted, and have attached checksums matching the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, each required file is reliably tied to this Harbor Day package, and a usable public-access derivative is present.\\n\\nThe donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417."}, {"path": [], "text": "Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417.", "negative_left": "The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417.", "negative_right": "Participant record HD-417 lists the person's birth date as 14 May 2007 and the Harbor Day event date as 22 June 2026.", "right": "Participant record HD-417 lists the person's birth date as 14 May 2011 and the Harbor Day event date as 22 June 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-025", "id": "fast-41-diverse-228-025-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The digital archivist’s case note covers the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The Harbor Day event register identifies the person shown at 00:04:12 in the Harbor Day preservation video as participant record HD-417. Participant record HD-417 lists the person's birth date as 14 May 2007 and the Harbor Day event date as 22 June 2026. The frame shows that person clearly and without blur; every other person appearing in the preservation video is blurred. All required preservation files play correctly, are uncorrupted, and have attached checksums matching the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The authoritative donor record supplies the applicable access instruction. Rights permit archival preservation, each required file is reliably tied to this Harbor Day package, and a usable public-access derivative is present.\n\nThe donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and preserve the donor’s access instruction without inventing exceptions. Both contexts preserve the Harbor Day video, subject H-27, and timestamp 00:04:12 bindings. The evidence consists of two complete factual sentences: \"The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27.\" and \"The age field for subject H-27 on the Harbor Day event date reads 17 years.\" The counterfactual coherently changes H-27’s age to 19 while retaining the same timestamp and unblurred-person facts. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ editing file. The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27. The age field for subject H-27 on the Harbor Day event date reads 17 years. At that timestamp, H-27 is clearly visible without blur, while every other person appearing in the video is blurred. All required preservation files play correctly and pass integrity review; each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as part of the package. A usable public-access derivative is included. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27."}, {"path": [], "text": "The age field for subject H-27 on the Harbor Day event date reads 17 years."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27.", "negative_left": "The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27.", "negative_right": "The age field for subject H-27 on the Harbor Day event date reads 19 years.", "right": "The age field for subject H-27 on the Harbor Day event date reads 17 years."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-027", "id": "fast-41-diverse-228-027-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ editing file. The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27. The age field for subject H-27 on the Harbor Day event date reads 17 years. At that timestamp, H-27 is clearly visible without blur, while every other person appearing in the video is blurred. All required preservation files play correctly and pass integrity review; each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as part of the package. A usable public-access derivative is included. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and preserve the donor’s access instruction without inventing exceptions. Both contexts preserve the Harbor Day video, subject H-27, and timestamp 00:04:12 bindings. The evidence consists of two complete factual sentences: \"The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27.\" and \"The age field for subject H-27 on the Harbor Day event date reads 17 years.\" The counterfactual coherently changes H-27’s age to 19 while retaining the same timestamp and unblurred-person facts. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ editing file. The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27. The age field for subject H-27 on the Harbor Day event date reads 17 years. At that timestamp, H-27 is clearly visible without blur, while every other person appearing in the video is blurred. All required preservation files play correctly and pass integrity review; each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as part of the package. A usable public-access derivative is included. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27."}, {"path": [], "text": "The age field for subject H-27 on the Harbor Day event date reads 17 years."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27.", "negative_left": "The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27.", "negative_right": "The age field for subject H-27 on the Harbor Day event date reads 19 years.", "right": "The age field for subject H-27 on the Harbor Day event date reads 17 years."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-027", "id": "fast-41-diverse-228-027-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ editing file. The Harbor Day event record identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject H-27. The age field for subject H-27 on the Harbor Day event date reads 19 years. At that timestamp, H-27 is clearly visible without blur, while every other person appearing in the video is blurred. All required preservation files play correctly and pass integrity review; each has an attached checksum, and each checksum matches the manifest. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The donor record supplies the applicable access instruction. Rights permit archival preservation, and every required file is reliably identified as part of the package. A usable public-access derivative is included. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full policy, criteria, and instructions, and both contexts retain the donor’s conditional access rule. The timestamp, video, and subject bindings are preserved. Evidence contains exactly two complete factual sentences: “The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417.” and “The age field for subject HD-417 on the Harbor Day event date records 17 years.” The counterfactual coherently changes the subject’s recorded age to 18 while retaining the other observations. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ editing file. The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417. The age field for subject HD-417 on the Harbor Day event date records 17 years. The frame shows subject HD-417 without blur, while every other person appearing in the video is blurred. The package’s required preservation files open successfully and have no corruption indicators. Each required file has an attached checksum, and each checksum matches the corresponding manifest entry. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The donor record supplies the applicable access instruction. Rights documentation permits archival preservation, and every required file is reliably linked to the Harbor Day package. A usable public-access derivative is present. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417."}, {"path": [], "text": "The age field for subject HD-417 on the Harbor Day event date records 17 years."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417.", "negative_left": "The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417.", "negative_right": "The age field for subject HD-417 on the Harbor Day event date records 18 years.", "right": "The age field for subject HD-417 on the Harbor Day event date records 17 years."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-039", "id": "fast-41-diverse-228-039-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ editing file. The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417. The age field for subject HD-417 on the Harbor Day event date records 17 years. The frame shows subject HD-417 without blur, while every other person appearing in the video is blurred. The package’s required preservation files open successfully and have no corruption indicators. Each required file has an attached checksum, and each checksum matches the corresponding manifest entry. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The donor record supplies the applicable access instruction. Rights documentation permits archival preservation, and every required file is reliably linked to the Harbor Day package. A usable public-access derivative is present. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the full policy, criteria, and instructions, and both contexts retain the donor’s conditional access rule. The timestamp, video, and subject bindings are preserved. Evidence contains exactly two complete factual sentences: “The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417.” and “The age field for subject HD-417 on the Harbor Day event date records 17 years.” The counterfactual coherently changes the subject’s recorded age to 18 while retaining the other observations. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ editing file. The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417. The age field for subject HD-417 on the Harbor Day event date records 17 years. The frame shows subject HD-417 without blur, while every other person appearing in the video is blurred. The package’s required preservation files open successfully and have no corruption indicators. Each required file has an attached checksum, and each checksum matches the corresponding manifest entry. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The donor record supplies the applicable access instruction. Rights documentation permits archival preservation, and every required file is reliably linked to the Harbor Day package. A usable public-access derivative is present. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417."}, {"path": [], "text": "The age field for subject HD-417 on the Harbor Day event date records 17 years."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417.", "negative_left": "The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417.", "negative_right": "The age field for subject HD-417 on the Harbor Day event date records 18 years.", "right": "The age field for subject HD-417 on the Harbor Day event date records 17 years."}, "verifier_independent_model": false}, "family": "fast-41-diverse-228-039", "id": "fast-41-diverse-228-039-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "The Harbor Day package contains an MP4 preservation video, a WAV audio file, a JPG poster, and a PRPROJ editing file. The Harbor Day event registration identifies the person shown at 00:04:12 in the Harbor Day preservation video as subject HD-417. The age field for subject HD-417 on the Harbor Day event date records 18 years. The frame shows subject HD-417 without blur, while every other person appearing in the video is blurred. The package’s required preservation files open successfully and have no corruption indicators. Each required file has an attached checksum, and each checksum matches the corresponding manifest entry. The authoritative metadata record supplies the title, creator, event date, rights holder, and format for every file. The donor record supplies the applicable access instruction. Rights documentation permits archival preservation, and every required file is reliably linked to the Harbor Day package. A usable public-access derivative is present. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts restate compatible policy facts without adding exceptions. The Paper Moon revision, register, candidate records, and 2026-09-17 timing remain consistently bound. The evidence consists of two complete factual sentences: \"In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417.\" and \"In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592.\" The counterfactual coherently changes the sole register issue from PV-417 to PV-592 while retaining consistent candidate identifiers and other facts. Neither context contains a gold answer, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"Case note dated 2026-09-17: The proposed revision of “Paper Moon” has exactly one entry in its remaining-issue register. In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417. In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592. R-V lists visual continuity as its sole issue category and Storyboard artist as its sole responsible role. R-D lists dialogue as its sole issue category and Writer as its sole responsible role. The remaining issue does not conflict with the signed brief, and the sequence is ready for animatic. It does not block animatic advancement. The issue is significant, with urgency value 3. the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417."}, {"path": [], "text": "In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417.", "negative_left": "In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-592.", "negative_right": "In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592.", "right": "In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592."}, "verifier_independent_model": false}, "family": "fast-41-diverse-230-017", "id": "fast-41-diverse-230-017-base", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "Case note dated 2026-09-17: The proposed revision of “Paper Moon” has exactly one entry in its remaining-issue register. In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417. In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592. R-V lists visual continuity as its sole issue category and Storyboard artist as its sole responsible role. R-D lists dialogue as its sole issue category and Writer as its sole responsible role. The remaining issue does not conflict with the signed brief, and the sequence is ready for animatic. It does not block animatic advancement. The issue is significant, with urgency value 3. the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "storyboard_artist_ready_3"}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts restate compatible policy facts without adding exceptions. The Paper Moon revision, register, candidate records, and 2026-09-17 timing remain consistently bound. The evidence consists of two complete factual sentences: \"In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417.\" and \"In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592.\" The counterfactual coherently changes the sole register issue from PV-417 to PV-592 while retaining consistent candidate identifiers and other facts. Neither context contains a gold answer, answer code, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship; set cardinality and record-value statements do not improperly bundle final classifications. The focus atom is the factual identity relation between the remaining issue and R-V, not a policy proposition. The policy evidence correctly preserves the substantive state-originated brief, ownership, readiness, and urgency rules; case observations such as the cup and shouted line need not be preserved for synthetic contexts. The base and counter assignments differ only on the focus and are realizable as alternative sole remaining issues—R-V in the base and R-D in the counter—while sharing the stated nonblocking, ready, significant, urgency-3 properties.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies the sole remaining issue with R-V, whose category and responsible role are visual continuity and Storyboard artist. It also establishes readiness, nonblocking status, significance, and urgency 3. These facts are sufficient for storyboard_artist_ready_3 and exclude the blocked urgency-4 alternative.", "rule_index": 0, "sound": true}, {"reason": "Because the remaining issue matches exactly one of R-V and R-D and is refuted as matching R-V, it must match R-D. R-D supplies dialogue as the category and Writer as the responsible role, while the remaining conditions establish no brief conflict, animatic readiness, nonblocking status, and urgency 3. This is sufficient for writer_ready_3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Exactly one issue remains in the proposed revision of “Paper Moon.”"}, {"id": "a2", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to exactly one of the issue identifiers in candidate records R-V and R-D."}, {"id": "a3", "statement": "The issue identifier in the remaining-issue register for the proposed revision of “Paper Moon” is identical to the issue identifier in candidate record R-V."}, {"id": "a4", "statement": "The sole issue-category value in candidate record R-V is visual continuity."}, {"id": "a5", "statement": "The sole responsible-role value in candidate record R-V is Storyboard artist."}, {"id": "a6", "statement": "The sole issue-category value in candidate record R-D is dialogue."}, {"id": "a7", "statement": "The sole responsible-role value in candidate record R-D is Writer."}, {"id": "a8", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not conflict with the signed brief."}, {"id": "a9", "statement": "The sequence in the proposed revision of “Paper Moon” is ready for animatic."}, {"id": "a10", "statement": "The remaining issue in the proposed revision of “Paper Moon” does not block animatic advancement."}, {"id": "a11", "statement": "The remaining issue in the proposed revision of “Paper Moon” is significant."}, {"id": "a12", "statement": "The urgency value for the remaining issue in the proposed revision of “Paper Moon” is 3."}], "base_state_json": "\"Case note dated 2026-09-17: The proposed revision of “Paper Moon” has exactly one entry in its remaining-issue register. In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417. In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592. R-V lists visual continuity as its sole issue category and Storyboard artist as its sole responsible role. R-D lists dialogue as its sole issue category and Writer as its sole responsible role. The remaining issue does not conflict with the signed brief, and the sequence is ready for animatic. It does not block animatic advancement. The issue is significant, with urgency value 3. the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417."}, {"path": [], "text": "In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592."}], "policy_evidence": [{"path": [], "text": "the brief requires Mina’s realization to remain wordless."}, {"path": [], "text": "Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues."}, {"path": [], "text": "A sequence advances to animatic only with no brief-blocking conflict."}, {"path": [], "text": "Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}], "rules": [{"justification": "The remaining-issue identifier matches R-V, whose sole category is visual continuity and whose sole responsible role is Storyboard artist. The sole remaining issue is significant, nonblocking, ready for animatic, and assigned urgency 3, satisfying this option and excluding the blocked urgency-4 alternative.", "target": "storyboard_artist_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}, {"justification": "The remaining-issue identifier matches exactly one of R-V and R-D and is explicitly not identical to R-V’s identifier, so it matches R-D. R-D’s sole category is dialogue and its sole responsible role is Writer. The issue does not conflict with the brief or block advancement, and the sequence is ready for animatic at urgency 3.", "target": "writer_ready_3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-417.", "negative_left": "In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-592.", "negative_right": "In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592.", "right": "In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592."}, "verifier_independent_model": false}, "family": "fast-41-diverse-230-017", "id": "fast-41-diverse-230-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"animation_producer_blocked_5": "Select only if a safety or legal production issue belongs to the Animation producer, blocks animatic advancement, and has urgency 5.", "director_ready_2": "Select only if the remaining note is a minor performance-intent adjustment for the Director, the sequence is ready for animatic, and urgency is 2.", "none_of_above": "Select if the evidence supports a role, readiness status, and urgency combination not represented by any substantive option.", "storyboard_artist_blocked_4": "Select only if an unresolved visual-continuity error belongs to the Storyboard artist, blocks animatic advancement, and has urgency 4.", "storyboard_artist_ready_3": "Select only if a significant but nonblocking visual-continuity issue remains for the Storyboard artist, the sequence is ready for animatic, and urgency is 3.", "writer_ready_3": "Select only if a dialogue issue remains for the Writer but does not conflict with the brief, so the sequence is ready for animatic at urgency 3."}, "instructions": "Choose the single option that correctly paraphrases the remaining issue, routes it to the responsible role, determines animatic readiness, and applies the urgency rubric.", "type": "choice"}}, "state": "Case note dated 2026-09-17: The proposed revision of “Paper Moon” has exactly one entry in its remaining-issue register. In the remaining-issue register for the proposed revision of “Paper Moon” dated 2026-09-17, the issue identifier is PV-592. In the 2026-09-17 candidate records for the proposed revision of “Paper Moon,” R-V has issue identifier PV-417 and R-D has issue identifier PV-592. R-V lists visual continuity as its sole issue category and Storyboard artist as its sole responsible role. R-D lists dialogue as its sole issue category and Writer as its sole responsible role. The remaining issue does not conflict with the signed brief, and the sequence is ready for animatic. It does not block animatic advancement. The issue is significant, with urgency value 3. the brief requires Mina’s realization to remain wordless. Writer owns dialogue conflicts; storyboard artist owns visual continuity; director owns performance intent; animation producer owns schedule or resource issues. A sequence advances to animatic only with no brief-blocking conflict. Revision urgency is ordered: 1 cosmetic, 2 minor, 3 significant but nonblocking, 4 blocking brief violation, 5 safety or legal stop."}, "method": "c2d", "provenance": {"source_id": "diverse-230", "source_is_synthetic": true, "source_sha256": "f5e0aaa87d2b79ee5813c648cafdc623c032303fd64f189f82a055ded7fb75c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "writer_ready_3"}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The two exact evidence quotes are “The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026.” and “The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026.” in the base context, while the counterfactual changes only the seeing time to “14:06:12 UTC” and coherently represents a policy violation without duplicate measurements or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Storyboard artist\",\"text\":\"The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026. The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026. Review notes confirm that each event has an exact logged instant, and Mara's glances before the opening target only the scraping sound. The revision has exactly one minor issue: a missing storyboard annotation. That note is non-blocking, and no remaining note requires routing to a role other than the Storyboard artist.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role. The proposed revision is not subject to a hold.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026."}, {"path": ["1", "text"], "text": "The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026.", "negative_left": "The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026.", "negative_right": "The proposed revision's event log records Mara first seeing the brass compass at 14:06:12 UTC on 6 May 2026.", "right": "The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-232-011", "id": "fast-41-diverse-232-011-base", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Storyboard artist", "text": "The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026. The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026. Review notes confirm that each event has an exact logged instant, and Mara's glances before the opening target only the scraping sound. The revision has exactly one minor issue: a missing storyboard annotation. That note is non-blocking, and no remaining note requires routing to a role other than the Storyboard artist."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role. The proposed revision is not subject to a hold."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The two exact evidence quotes are “The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026.” and “The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026.” in the base context, while the counterfactual changes only the seeing time to “14:06:12 UTC” and coherently represents a policy violation without duplicate measurements or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Storyboard artist\",\"text\":\"The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026. The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026. Review notes confirm that each event has an exact logged instant, and Mara's glances before the opening target only the scraping sound. The revision has exactly one minor issue: a missing storyboard annotation. That note is non-blocking, and no remaining note requires routing to a role other than the Storyboard artist.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role. The proposed revision is not subject to a hold.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026."}, {"path": ["1", "text"], "text": "The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026.", "negative_left": "The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026.", "negative_right": "The proposed revision's event log records Mara first seeing the brass compass at 14:06:12 UTC on 6 May 2026.", "right": "The proposed revision's event log records Mara first seeing the brass compass at 14:07:12 UTC on 6 May 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-232-011", "id": "fast-41-diverse-232-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Storyboard artist", "text": "The proposed revision's event log records the attic hatch first opening at 14:07:05 UTC on 6 May 2026. The proposed revision's event log records Mara first seeing the brass compass at 14:06:12 UTC on 6 May 2026. Review notes confirm that each event has an exact logged instant, and Mara's glances before the opening target only the scraping sound. The revision has exactly one minor issue: a missing storyboard annotation. That note is non-blocking, and no remaining note requires routing to a role other than the Storyboard artist."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role. The proposed revision is not subject to a hold."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and policy, preserve all bindings, use the exact two factual evidence sentences, and contain no embedded answer; the counterfactual coherently changes only Mara’s viewing time, creating a policy violation without duplicate or contradictory measurements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Storyboard artist\",\"text\":\"In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026.\"},{\"speaker\":\"Storyboard artist\",\"text\":\"In the proposed revision's event log, Mara first sees the brass compass at 14:00:01 UTC on 3 May 2026.\"},{\"speaker\":\"Storyboard artist\",\"text\":\"Before the hatch first opens, every glance by Mara targets only the scraping sound. The proposed revision has exactly one minor issue, and the remaining note concerns a missing storyboard annotation.\"},{\"speaker\":\"Writer\",\"text\":\"The remaining note is non-blocking, and no remaining revision note requires routing to a role other than the Storyboard artist. The proposed revision is not subject to a hold.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026."}, {"path": ["2", "text"], "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:00:01 UTC on 3 May 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026.", "negative_left": "In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026.", "negative_right": "In the proposed revision's event log, Mara first sees the brass compass at 13:59:59 UTC on 3 May 2026.", "right": "In the proposed revision's event log, Mara first sees the brass compass at 14:00:01 UTC on 3 May 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-232-021", "id": "fast-41-diverse-232-021-base", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Storyboard artist", "text": "In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026."}, {"speaker": "Storyboard artist", "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:00:01 UTC on 3 May 2026."}, {"speaker": "Storyboard artist", "text": "Before the hatch first opens, every glance by Mara targets only the scraping sound. The proposed revision has exactly one minor issue, and the remaining note concerns a missing storyboard annotation."}, {"speaker": "Writer", "text": "The remaining note is non-blocking, and no remaining revision note requires routing to a role other than the Storyboard artist. The proposed revision is not subject to a hold."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and policy, preserve all bindings, use the exact two factual evidence sentences, and contain no embedded answer; the counterfactual coherently changes only Mara’s viewing time, creating a policy violation without duplicate or contradictory measurements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a0": "supported", "a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a0": "supported", "a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship or status, including the universally scoped pre-opening-glance and routing-exclusion atoms. The focus is a factual temporal relation. The base and counter assignments can differ only on the focus: the counter is realizable by making first sight simultaneous with hatch opening, which violates the strict-after requirement without contradicting the condition about glances strictly before opening. Policy evidence correctly preserves the state-originating brief and producer advancement condition; all general scale, routing, readiness, and decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes compliance with both brief constraints, exactly one minor issue, Storyboard routing for the sole annotation note, no competing-role requirement, non-blocking status, slot availability, and no hold. Under the stated scale and advancement conditions, this is sufficient for the three-part true outcome.", "rule_index": 0, "sound": true}, {"reason": "With exact instants established, refuting that first compass sight is later than first hatch opening places the sighting at or before opening. That contradicts the explicit brief, which forces urgency level 4 and therefore refutes the required level-2 conjunction.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a0", "statement": "The event log for the proposed revision records an exact instant when the attic hatch first opens."}, {"id": "a1", "statement": "The event log for the proposed revision records an exact instant when Mara first sees the brass compass."}, {"id": "a2", "statement": "The logged instant when Mara first sees the brass compass is later than the logged instant when the attic hatch first opens in the proposed revision."}, {"id": "a3", "statement": "Every glance by Mara before the attic hatch first opens in the proposed revision targets only the scraping sound."}, {"id": "a4", "statement": "The proposed revision has exactly one minor issue."}, {"id": "a5", "statement": "The remaining note for the proposed revision concerns a missing storyboard annotation."}, {"id": "a6", "statement": "The remaining note for the proposed revision is non-blocking."}, {"id": "a7", "statement": "An animatic slot is available for the proposed revision."}, {"id": "a8", "statement": "The proposed revision is not subject to a hold."}, {"id": "a9", "statement": "No remaining revision note requires routing to a role other than the Storyboard artist."}], "base_state_json": "[{\"speaker\":\"Director\",\"text\":\"Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound.\"},{\"speaker\":\"Storyboard artist\",\"text\":\"In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026.\"},{\"speaker\":\"Storyboard artist\",\"text\":\"In the proposed revision's event log, Mara first sees the brass compass at 14:00:01 UTC on 3 May 2026.\"},{\"speaker\":\"Storyboard artist\",\"text\":\"Before the hatch first opens, every glance by Mara targets only the scraping sound. The proposed revision has exactly one minor issue, and the remaining note concerns a missing storyboard annotation.\"},{\"speaker\":\"Writer\",\"text\":\"The remaining note is non-blocking, and no remaining revision note requires routing to a role other than the Storyboard artist. The proposed revision is not subject to a hold.\"},{\"speaker\":\"Animation producer\",\"text\":\"The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role.\"}]", "base_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["1", "text"], "text": "In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026."}, {"path": ["2", "text"], "text": "In the proposed revision's event log, Mara first sees the brass compass at 14:00:01 UTC on 3 May 2026."}], "policy_evidence": [{"path": ["0", "text"], "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"path": ["3", "text"], "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}], "rules": [{"justification": "The exact event times place Mara's first sight of the compass after the hatch first opens, and every pre-opening glance targets only the scraping sound, so the stated brief is satisfied. Exactly one minor issue with no brief contradiction gives urgency level 2. The sole remaining note concerns a missing storyboard annotation and requires no other role, so it routes to the Storyboard artist. The note is non-blocking, an animatic slot is available, and no hold applies, so the ready level-2 sequence advances to animatic.", "target": "true", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Because both exact event instants are recorded, refuting that the first compass sighting is later than the first hatch opening entails that the sighting occurs at or before the opening. This contradicts the brief's requirement that Mara not see the compass until after the hatch opens. Any brief contradiction gives urgency level 4, so the required three-part true outcome fails despite the unchanged note, routing, availability, and hold facts.", "target": "false", "when": [{"atom_id": "a0", "state": "supported"}, {"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026.", "negative_left": "In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026.", "negative_right": "In the proposed revision's event log, Mara first sees the brass compass at 13:59:59 UTC on 3 May 2026.", "right": "In the proposed revision's event log, Mara first sees the brass compass at 14:00:01 UTC on 3 May 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-232-021", "id": "fast-41-diverse-232-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The sequence does not meet all three conditions—for example, it has a brief contradiction, has more or fewer than one minor issue, requires another role, or must be held.", "true": "The sequence is urgency level 2, the annotation note goes to the Storyboard artist, and the sequence advances to animatic."}, "instructions": "Decide whether the sequence should receive revision-urgency level 2, have its remaining note routed to the Storyboard artist, and proceed to animatic. Scale: 1 = no issues; 2 = exactly one minor issue and no brief contradiction; 3 = at least two minor issues; 4 = any brief contradiction. Levels 1-2 are ready. A missing storyboard annotation is minor. Route visual framing or annotation notes to the Storyboard artist.", "type": "noul"}}, "state": [{"speaker": "Director", "text": "Brief: Mara must not see the brass compass until after the attic hatch opens; her pre-opening glance may target only the scraping sound."}, {"speaker": "Storyboard artist", "text": "In the proposed revision's event log, the attic hatch first opens at 14:00:00 UTC on 3 May 2026."}, {"speaker": "Storyboard artist", "text": "In the proposed revision's event log, Mara first sees the brass compass at 13:59:59 UTC on 3 May 2026."}, {"speaker": "Storyboard artist", "text": "Before the hatch first opens, every glance by Mara targets only the scraping sound. The proposed revision has exactly one minor issue, and the remaining note concerns a missing storyboard annotation."}, {"speaker": "Writer", "text": "The remaining note is non-blocking, and no remaining revision note requires routing to a role other than the Storyboard artist. The proposed revision is not subject to a hold."}, {"speaker": "Animation producer", "text": "The animatic slot is available. I can advance the sequence if the remaining note is non-blocking and sent to the correct role."}]}, "method": "c2d", "provenance": {"source_id": "diverse-232", "source_is_synthetic": true, "source_sha256": "273fbfbfb9cd9c13614f2bbaa604524feb72ec0b95b59d4dfa6dfb443f9469b6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions. Both contexts retain the same project, sequence, version, scope, and request bindings. The evidence consists of two complete factual sentences: \"Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.\" and \"The panel labeled X in Version 3 has an attached camera note.\" The counterfactual coherently changes only the camera-note status of panel X, with panel 17 excluded from the otherwise complete-note assertion. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "supported", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "full_context_fact_states": {"base": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "supported", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "counterfactual": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "refuted", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "remove_left": {"a_panel17_camera_note": "unknown"}, "remove_right": {"a_panel17_camera_note": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_panel17_camera_note": "unknown"}, "negative_pair": {"a_panel17_camera_note": "refuted"}, "negative_sentence": {"a_panel17_camera_note": "unknown"}, "positive_pair": {"a_panel17_camera_note": "supported"}, "right": {"a_panel17_camera_note": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universal camera-note and confirmation atoms remain single relations over their stated sets. The focus—whether panel 17 has a camera note—is factual rather than policy-based. The base and counter assignments differ only on that focus and are realizable: confirmations can report no unresolved revisions even if an objective panel-note defect exists, and the correction can remain non-redesigning in either assignment. Policy evidence preserves the substantive state-originating brief requirements. Supersession, scoring criteria, priorities, and task instructions are already retained in the unchanged questions object and correctly need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the latest assessable version satisfies every state-originating brief requirement: pursuit direction, rescue choice, exactly 24 numbered panels, and camera notes on all panels. They also establish that all relevant role confirmations report no unresolved revisions. This is sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all core story, continuity, and panel-count requirements are met and that panel 17 is the sole camera-note defect. They further establish that correcting it requires no sequence redesign, making it a single localized handoff defect sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_current", "statement": "Version 3 is the latest applicable version of the rooftop escape sequence."}, {"id": "a_storyboard", "statement": "The current storyboard for Version 3 is available for review."}, {"id": "a_brief", "statement": "The essential brief for the rooftop escape sequence is available for review."}, {"id": "a_direction", "statement": "Version 3 maintains readable left-to-right pursuit."}, {"id": "a_rescue", "statement": "Version 3 visibly depicts Lina choosing to rescue the trapped bird before escaping."}, {"id": "a_panel_count", "statement": "Version 3 contains exactly 24 numbered panels."}, {"id": "a_other_camera_notes", "statement": "Every numbered panel in Version 3 other than panel 17 has an attached camera note."}, {"id": "a_panel17_camera_note", "statement": "Numbered panel 17 in Version 3 has an attached camera note."}, {"id": "a_local_correction", "statement": "Attaching a camera note to panel 17 requires no redesign of the sequence."}, {"id": "a_understandable", "statement": "The sequence in Version 3 is understandable."}, {"id": "a_confirmations", "statement": "Every relevant role confirmation for Version 3 reports no unresolved revisions."}], "base_state_json": "{\"context\":\"The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff.\",\"request\":\"Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required.\",\"observations\":[\"Version 3 is the latest applicable version, and its current storyboard and the essential brief are available for review.\",\"The current storyboard for Version 3 is available for review.\",\"The essential brief for the rooftop escape sequence is available for review.\",\"Version 3 maintains readable left-to-right pursuit.\",\"Version 3 visibly depicts Lina choosing to rescue the trapped bird before escaping.\",\"Version 3 contains exactly 24 numbered panels.\",\"Every numbered panel in Version 3 other than panel 17 has an attached camera note.\",\"Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.\",\"The panel labeled X in Version 3 has an attached camera note.\",\"Attaching a camera note to panel 17 requires no redesign of the sequence.\",\"The sequence in Version 3 is understandable.\",\"Every relevant role confirmation for Version 3 reports no unresolved revisions.\",\"The latest submission superseded earlier storyboard notes, and review is based on that submission.\"]}", "base_states": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "supported"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}], "counter_states": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "refuted"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}], "focus_atom": "a_panel17_camera_note", "focus_evidence": [{"path": ["observations", "7"], "text": "Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3."}, {"path": ["observations", "8"], "text": "The panel labeled X in Version 3 has an attached camera note."}], "policy_evidence": [{"path": ["context"], "text": "The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff."}, {"path": ["request"], "text": "Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required."}], "rules": [{"justification": "The latest applicable version is assessable and meets every stated story, continuity, panel-count, and camera-note requirement, while all relevant confirmations report no unresolved revisions.", "target": "4", "when": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}]}, {"justification": "All core story, continuity, and panel-count requirements are met, but panel 17 is the sole panel without a camera note; attaching that note is a localized correction that requires no sequence redesign.", "target": "3", "when": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "refuted"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}]}]}, "verified_pair": {"left": "Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.", "negative_left": "Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.", "negative_right": "The panel labeled X in Version 3 has no attached camera note.", "right": "The panel labeled X in Version 3 has an attached camera note."}, "verifier_independent_model": false}, "family": "fast-41-diverse-234-018", "id": "fast-41-diverse-234-018-base", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable: the current storyboard or essential brief is absent, so readiness cannot be checked and the package must not advance.", "1 — Major rework required: the current version contradicts a core story or continuity requirement, or is too incomplete for meaningful handoff; route substantive revisions to the responsible creative roles.", "2 — Significant revision required: the sequence is understandable but still has at least one unresolved brief requirement or multiple handoff defects; keep it out of layout and route specific notes.", "3 — Conditionally ready: all core story and continuity requirements are met, but one minor, localized handoff defect remains that can be corrected without redesigning the sequence; advance only after that correction is verified.", "4 — Fully ready: the latest version meets every stated story, continuity, panel-count, and camera-note requirement, and relevant role confirmations show no unresolved revisions; approve for layout with no further note routing."], "instructions": "Select one readiness level. Later evidence supersedes earlier evidence when it explicitly replaces a prior version. Evaluate the current sequence against every stated brief and handoff requirement; older resolved notes must not lower the score.", "type": "score"}}, "state": {"context": "The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff.", "observations": ["Version 3 is the latest applicable version, and its current storyboard and the essential brief are available for review.", "The current storyboard for Version 3 is available for review.", "The essential brief for the rooftop escape sequence is available for review.", "Version 3 maintains readable left-to-right pursuit.", "Version 3 visibly depicts Lina choosing to rescue the trapped bird before escaping.", "Version 3 contains exactly 24 numbered panels.", "Every numbered panel in Version 3 other than panel 17 has an attached camera note.", "Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.", "The panel labeled X in Version 3 has an attached camera note.", "Attaching a camera note to panel 17 requires no redesign of the sequence.", "The sequence in Version 3 is understandable.", "Every relevant role confirmation for Version 3 reports no unresolved revisions.", "The latest submission superseded earlier storyboard notes, and review is based on that submission."], "request": "Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required."}}, "method": "c2d", "provenance": {"source_id": "diverse-234", "source_is_synthetic": true, "source_sha256": "e4655a6f708b74648e4a2cae39e134d37b2fbee776787cd553fa4dad14a35a6e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "media-04", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions. Both contexts retain the same project, sequence, version, scope, and request bindings. The evidence consists of two complete factual sentences: \"Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.\" and \"The panel labeled X in Version 3 has an attached camera note.\" The counterfactual coherently changes only the camera-note status of panel X, with panel 17 excluded from the otherwise complete-note assertion. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "refuted", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "full_context_fact_states": {"base": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "supported", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "counterfactual": {"a_brief": "supported", "a_confirmations": "supported", "a_current": "supported", "a_direction": "supported", "a_local_correction": "supported", "a_other_camera_notes": "supported", "a_panel17_camera_note": "refuted", "a_panel_count": "supported", "a_rescue": "supported", "a_storyboard": "supported", "a_understandable": "supported"}, "remove_left": {"a_panel17_camera_note": "unknown"}, "remove_right": {"a_panel17_camera_note": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_panel17_camera_note": "unknown"}, "negative_pair": {"a_panel17_camera_note": "refuted"}, "negative_sentence": {"a_panel17_camera_note": "unknown"}, "positive_pair": {"a_panel17_camera_note": "supported"}, "right": {"a_panel17_camera_note": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universal camera-note and confirmation atoms remain single relations over their stated sets. The focus—whether panel 17 has a camera note—is factual rather than policy-based. The base and counter assignments differ only on that focus and are realizable: confirmations can report no unresolved revisions even if an objective panel-note defect exists, and the correction can remain non-redesigning in either assignment. Policy evidence preserves the substantive state-originating brief requirements. Supersession, scoring criteria, priorities, and task instructions are already retained in the unchanged questions object and correctly need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the latest assessable version satisfies every state-originating brief requirement: pursuit direction, rescue choice, exactly 24 numbered panels, and camera notes on all panels. They also establish that all relevant role confirmations report no unresolved revisions. This is sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all core story, continuity, and panel-count requirements are met and that panel 17 is the sole camera-note defect. They further establish that correcting it requires no sequence redesign, making it a single localized handoff defect sufficient for level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_current", "statement": "Version 3 is the latest applicable version of the rooftop escape sequence."}, {"id": "a_storyboard", "statement": "The current storyboard for Version 3 is available for review."}, {"id": "a_brief", "statement": "The essential brief for the rooftop escape sequence is available for review."}, {"id": "a_direction", "statement": "Version 3 maintains readable left-to-right pursuit."}, {"id": "a_rescue", "statement": "Version 3 visibly depicts Lina choosing to rescue the trapped bird before escaping."}, {"id": "a_panel_count", "statement": "Version 3 contains exactly 24 numbered panels."}, {"id": "a_other_camera_notes", "statement": "Every numbered panel in Version 3 other than panel 17 has an attached camera note."}, {"id": "a_panel17_camera_note", "statement": "Numbered panel 17 in Version 3 has an attached camera note."}, {"id": "a_local_correction", "statement": "Attaching a camera note to panel 17 requires no redesign of the sequence."}, {"id": "a_understandable", "statement": "The sequence in Version 3 is understandable."}, {"id": "a_confirmations", "statement": "Every relevant role confirmation for Version 3 reports no unresolved revisions."}], "base_state_json": "{\"context\":\"The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff.\",\"request\":\"Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required.\",\"observations\":[\"Version 3 is the latest applicable version, and its current storyboard and the essential brief are available for review.\",\"The current storyboard for Version 3 is available for review.\",\"The essential brief for the rooftop escape sequence is available for review.\",\"Version 3 maintains readable left-to-right pursuit.\",\"Version 3 visibly depicts Lina choosing to rescue the trapped bird before escaping.\",\"Version 3 contains exactly 24 numbered panels.\",\"Every numbered panel in Version 3 other than panel 17 has an attached camera note.\",\"Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.\",\"The panel labeled X in Version 3 has an attached camera note.\",\"Attaching a camera note to panel 17 requires no redesign of the sequence.\",\"The sequence in Version 3 is understandable.\",\"Every relevant role confirmation for Version 3 reports no unresolved revisions.\",\"The latest submission superseded earlier storyboard notes, and review is based on that submission.\"]}", "base_states": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "supported"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}], "counter_states": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "refuted"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}], "focus_atom": "a_panel17_camera_note", "focus_evidence": [{"path": ["observations", "7"], "text": "Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3."}, {"path": ["observations", "8"], "text": "The panel labeled X in Version 3 has an attached camera note."}], "policy_evidence": [{"path": ["context"], "text": "The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff."}, {"path": ["request"], "text": "Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required."}], "rules": [{"justification": "The latest applicable version is assessable and meets every stated story, continuity, panel-count, and camera-note requirement, while all relevant confirmations report no unresolved revisions.", "target": "4", "when": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}]}, {"justification": "All core story, continuity, and panel-count requirements are met, but panel 17 is the sole panel without a camera note; attaching that note is a localized correction that requires no sequence redesign.", "target": "3", "when": [{"atom_id": "a_current", "state": "supported"}, {"atom_id": "a_storyboard", "state": "supported"}, {"atom_id": "a_brief", "state": "supported"}, {"atom_id": "a_direction", "state": "supported"}, {"atom_id": "a_rescue", "state": "supported"}, {"atom_id": "a_panel_count", "state": "supported"}, {"atom_id": "a_other_camera_notes", "state": "supported"}, {"atom_id": "a_panel17_camera_note", "state": "refuted"}, {"atom_id": "a_local_correction", "state": "supported"}, {"atom_id": "a_understandable", "state": "supported"}, {"atom_id": "a_confirmations", "state": "supported"}]}]}, "verified_pair": {"left": "Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.", "negative_left": "Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.", "negative_right": "The panel labeled X in Version 3 has no attached camera note.", "right": "The panel labeled X in Version 3 has an attached camera note."}, "verifier_independent_model": false}, "family": "fast-41-diverse-234-018", "id": "fast-41-diverse-234-018-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not assessable: the current storyboard or essential brief is absent, so readiness cannot be checked and the package must not advance.", "1 — Major rework required: the current version contradicts a core story or continuity requirement, or is too incomplete for meaningful handoff; route substantive revisions to the responsible creative roles.", "2 — Significant revision required: the sequence is understandable but still has at least one unresolved brief requirement or multiple handoff defects; keep it out of layout and route specific notes.", "3 — Conditionally ready: all core story and continuity requirements are met, but one minor, localized handoff defect remains that can be corrected without redesigning the sequence; advance only after that correction is verified.", "4 — Fully ready: the latest version meets every stated story, continuity, panel-count, and camera-note requirement, and relevant role confirmations show no unresolved revisions; approve for layout with no further note routing."], "instructions": "Select one readiness level. Later evidence supersedes earlier evidence when it explicitly replaces a prior version. Evaluate the current sequence against every stated brief and handoff requirement; older resolved notes must not lower the score.", "type": "score"}}, "state": {"context": "The team is reviewing the rooftop escape sequence for the animated short “Paper Comet.” The brief requires readable left-to-right pursuit, Lina visibly choosing to rescue the trapped bird before escaping, and 24 numbered panels with camera notes before layout handoff.", "observations": ["Version 3 is the latest applicable version, and its current storyboard and the essential brief are available for review.", "The current storyboard for Version 3 is available for review.", "The essential brief for the rooftop escape sequence is available for review.", "Version 3 maintains readable left-to-right pursuit.", "Version 3 visibly depicts Lina choosing to rescue the trapped bird before escaping.", "Version 3 contains exactly 24 numbered panels.", "Every numbered panel in Version 3 other than panel 17 has an attached camera note.", "Numbered panel 17 in Version 3 is identical to the panel labeled X in Version 3.", "The panel labeled X in Version 3 has no attached camera note.", "Attaching a camera note to panel 17 requires no redesign of the sequence.", "The sequence in Version 3 is understandable.", "Every relevant role confirmation for Version 3 reports no unresolved revisions.", "The latest submission superseded earlier storyboard notes, and review is based on that submission."], "request": "Using only the latest applicable evidence, rate the sequence’s readiness for layout and identify whether any further role-routed revision is required."}}, "method": "c2d", "provenance": {"source_id": "diverse-234", "source_is_synthetic": true, "source_sha256": "e4655a6f708b74648e4a2cae39e134d37b2fbee776787cd553fa4dad14a35a6e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts repeat it without exceptions. The request remains about the fictional tomato trial’s treatment-control validity under the stated protocol. “The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.” “The manual log contains a corresponding entry for each of events E17, E18, and E19.” Both spans are complete factual sentences. The counterfactual coherently removes only the E19 manual entry and does not assert contradictory coverage. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal quantification over an explicit set does not make the atoms non-atomic. A6 is a factual manual-record coverage relation rather than a policy classification. The base and counter assignments can coexist with all non-focus atoms unchanged: in the counter case, an affected event can lack a manual entry while any existing corresponding entries remain dated and independent, and A3 can still hold through a different kind of watering-record entry. The state-derived policy evidence preserves the substantive protocol requirements and sensor-gap exception. Rules originating in the retained questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports all four stated core requirements and establishes the sensor gap plus complete, dated, independent manual coverage of every affected watering event. It therefore suffices for valid_documented_exception and excludes the enumerated protocol failures.", "rule_index": 0, "sound": true}, {"reason": "With the sensor gap established, refutation of A6 entails that at least one affected watering event lacks a corresponding manual-log entry. This is explicit failure of the required sensor exception, not merely unknown documentation, and suffices for invalid_protocol_failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assignment of trial benches 1–4 to drought or control before treatment began was randomized."}, {"id": "A2", "statement": "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days."}, {"id": "A3", "statement": "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records."}, {"id": "A4", "statement": "Each of the trial’s 48 plants has a final dry-mass measurement."}, {"id": "A5", "statement": "The watering-sensor stream for drought bench 3 is missing on trial days 4–6."}, {"id": "A6", "statement": "Every drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 has a corresponding manual-log entry."}, {"id": "A7", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated."}, {"id": "A8", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."}], "base_state_json": "{\"context\":\"The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event. In a fictional tomato drought trial, benches 1–4 were assigned before treatment, and the study team retained the relevant treatment and measurement files. The trial used four benches and 48 plants. A sensor stream for drought bench 3 is absent on trial days 4–6. The pre-trial schedule covered seven days, and the treatment records distinguish drought-pot events from other records. All final measurements use dry mass. The dated manual records for the affected entries were maintained separately from the sensor system. The review must classify the trial under the stated protocol.\\n\\nThe protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.\\nOne missing sensor stream is allowed only when dated manual records independently document every affected watering event.\",\"evidence\":[\"The assignment of trial benches 1–4 to drought or control before treatment began was randomized.\",\"Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days.\",\"Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records.\",\"Each of the trial’s 48 plants has a final dry-mass measurement.\",\"The watering-sensor stream for drought bench 3 is missing on trial days 4–6.\",\"The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.\",\"The manual log contains a corresponding entry for each of events E17, E18, and E19.\",\"Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated.\",\"Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream.\"],\"request\":\"Classify the trial’s treatment-control validity under the stated protocol and route it accordingly.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": ["evidence", "5"], "text": "The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19."}, {"path": ["evidence", "6"], "text": "The manual log contains a corresponding entry for each of events E17, E18, and E19."}], "policy_evidence": [{"path": ["context"], "text": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements."}, {"path": ["context"], "text": "One missing sensor stream is allowed only when dated manual records independently document every affected watering event."}, {"path": ["request"], "text": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}], "rules": [{"justification": "All core requirements are explicitly supported, and every event affected by the established sensor gap has a dated manual-log entry independent of the missing sensor stream.", "target": "valid_documented_exception", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}, {"justification": "The established sensor gap has at least one affected watering event without a corresponding manual-log entry, so the claimed sensor exception fails and the protocol has an uncovered affected event.", "target": "invalid_protocol_failure", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.", "negative_left": "The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.", "negative_right": "The manual log contains no corresponding entry for event E19.", "right": "The manual log contains a corresponding entry for each of events E17, E18, and E19."}, "verifier_independent_model": false}, "family": "fast-41-diverse-241-001", "id": "fast-41-diverse-241-001-base", "input": {"questions": {"decision": {"criteria": {"insufficient_evidence_hold": "Place on evidence hold: the record neither proves a core violation nor supplies enough explicit documentation to confirm all core requirements or a claimed sensor exception.", "invalid_protocol_failure": "Route for protocol failure: explicit evidence shows nonrandom assignment, unequal pre-trial watering, an uncovered watering event, missing final growth measurements, or another violated core requirement.", "valid_documented_exception": "Accept as treatment-control valid and route to analysis: every core requirement is explicitly supported, and any sensor gap is fully covered by qualifying dated manual records."}, "instructions": "Select exactly one routing option. Apply the explicit protocol rule: all core requirements must be evidenced; a sensor gap is acceptable only if independently documented manual records cover every affected watering event.", "type": "choice"}}, "state": {"context": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event. In a fictional tomato drought trial, benches 1–4 were assigned before treatment, and the study team retained the relevant treatment and measurement files. The trial used four benches and 48 plants. A sensor stream for drought bench 3 is absent on trial days 4–6. The pre-trial schedule covered seven days, and the treatment records distinguish drought-pot events from other records. All final measurements use dry mass. The dated manual records for the affected entries were maintained separately from the sensor system. The review must classify the trial under the stated protocol.\n\nThe protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.\nOne missing sensor stream is allowed only when dated manual records independently document every affected watering event.", "evidence": ["The assignment of trial benches 1–4 to drought or control before treatment began was randomized.", "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days.", "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records.", "Each of the trial’s 48 plants has a final dry-mass measurement.", "The watering-sensor stream for drought bench 3 is missing on trial days 4–6.", "The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.", "The manual log contains a corresponding entry for each of events E17, E18, and E19.", "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated.", "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."], "request": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}}, "method": "c2d", "provenance": {"source_id": "diverse-241", "source_is_synthetic": true, "source_sha256": "1a2bbfd79de14bae070972939aa5cb346cad567f60d5260d99a6e72327b34f05", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_documented_exception"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts repeat it without exceptions. The request remains about the fictional tomato trial’s treatment-control validity under the stated protocol. “The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.” “The manual log contains a corresponding entry for each of events E17, E18, and E19.” Both spans are complete factual sentences. The counterfactual coherently removes only the E19 manual entry and does not assert contradictory coverage. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal quantification over an explicit set does not make the atoms non-atomic. A6 is a factual manual-record coverage relation rather than a policy classification. The base and counter assignments can coexist with all non-focus atoms unchanged: in the counter case, an affected event can lack a manual entry while any existing corresponding entries remain dated and independent, and A3 can still hold through a different kind of watering-record entry. The state-derived policy evidence preserves the substantive protocol requirements and sensor-gap exception. Rules originating in the retained questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports all four stated core requirements and establishes the sensor gap plus complete, dated, independent manual coverage of every affected watering event. It therefore suffices for valid_documented_exception and excludes the enumerated protocol failures.", "rule_index": 0, "sound": true}, {"reason": "With the sensor gap established, refutation of A6 entails that at least one affected watering event lacks a corresponding manual-log entry. This is explicit failure of the required sensor exception, not merely unknown documentation, and suffices for invalid_protocol_failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assignment of trial benches 1–4 to drought or control before treatment began was randomized."}, {"id": "A2", "statement": "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days."}, {"id": "A3", "statement": "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records."}, {"id": "A4", "statement": "Each of the trial’s 48 plants has a final dry-mass measurement."}, {"id": "A5", "statement": "The watering-sensor stream for drought bench 3 is missing on trial days 4–6."}, {"id": "A6", "statement": "Every drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 has a corresponding manual-log entry."}, {"id": "A7", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated."}, {"id": "A8", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."}], "base_state_json": "{\"context\":\"The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event. In a fictional tomato drought trial, benches 1–4 were assigned before treatment, and the study team retained the relevant treatment and measurement files. The trial used four benches and 48 plants. A sensor stream for drought bench 3 is absent on trial days 4–6. The pre-trial schedule covered seven days, and the treatment records distinguish drought-pot events from other records. All final measurements use dry mass. The dated manual records for the affected entries were maintained separately from the sensor system. The review must classify the trial under the stated protocol.\\n\\nThe protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.\\nOne missing sensor stream is allowed only when dated manual records independently document every affected watering event.\",\"evidence\":[\"The assignment of trial benches 1–4 to drought or control before treatment began was randomized.\",\"Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days.\",\"Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records.\",\"Each of the trial’s 48 plants has a final dry-mass measurement.\",\"The watering-sensor stream for drought bench 3 is missing on trial days 4–6.\",\"The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.\",\"The manual log contains a corresponding entry for each of events E17, E18, and E19.\",\"Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated.\",\"Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream.\"],\"request\":\"Classify the trial’s treatment-control validity under the stated protocol and route it accordingly.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": ["evidence", "5"], "text": "The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19."}, {"path": ["evidence", "6"], "text": "The manual log contains a corresponding entry for each of events E17, E18, and E19."}], "policy_evidence": [{"path": ["context"], "text": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements."}, {"path": ["context"], "text": "One missing sensor stream is allowed only when dated manual records independently document every affected watering event."}, {"path": ["request"], "text": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}], "rules": [{"justification": "All core requirements are explicitly supported, and every event affected by the established sensor gap has a dated manual-log entry independent of the missing sensor stream.", "target": "valid_documented_exception", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}, {"justification": "The established sensor gap has at least one affected watering event without a corresponding manual-log entry, so the claimed sensor exception fails and the protocol has an uncovered affected event.", "target": "invalid_protocol_failure", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.", "negative_left": "The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.", "negative_right": "The manual log contains no corresponding entry for event E19.", "right": "The manual log contains a corresponding entry for each of events E17, E18, and E19."}, "verifier_independent_model": false}, "family": "fast-41-diverse-241-001", "id": "fast-41-diverse-241-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"insufficient_evidence_hold": "Place on evidence hold: the record neither proves a core violation nor supplies enough explicit documentation to confirm all core requirements or a claimed sensor exception.", "invalid_protocol_failure": "Route for protocol failure: explicit evidence shows nonrandom assignment, unequal pre-trial watering, an uncovered watering event, missing final growth measurements, or another violated core requirement.", "valid_documented_exception": "Accept as treatment-control valid and route to analysis: every core requirement is explicitly supported, and any sensor gap is fully covered by qualifying dated manual records."}, "instructions": "Select exactly one routing option. Apply the explicit protocol rule: all core requirements must be evidenced; a sensor gap is acceptable only if independently documented manual records cover every affected watering event.", "type": "choice"}}, "state": {"context": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event. In a fictional tomato drought trial, benches 1–4 were assigned before treatment, and the study team retained the relevant treatment and measurement files. The trial used four benches and 48 plants. A sensor stream for drought bench 3 is absent on trial days 4–6. The pre-trial schedule covered seven days, and the treatment records distinguish drought-pot events from other records. All final measurements use dry mass. The dated manual records for the affected entries were maintained separately from the sensor system. The review must classify the trial under the stated protocol.\n\nThe protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.\nOne missing sensor stream is allowed only when dated manual records independently document every affected watering event.", "evidence": ["The assignment of trial benches 1–4 to drought or control before treatment began was randomized.", "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days.", "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records.", "Each of the trial’s 48 plants has a final dry-mass measurement.", "The watering-sensor stream for drought bench 3 is missing on trial days 4–6.", "The trial’s sensor audit lists the drought-bench 3 watering events affected by the sensor-stream gap on trial days 4–6 as events E17, E18, and E19.", "The manual log contains no corresponding entry for event E19.", "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated.", "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."], "request": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}}, "method": "c2d", "provenance": {"source_id": "diverse-241", "source_is_synthetic": true, "source_sha256": "1a2bbfd79de14bae070972939aa5cb346cad567f60d5260d99a6e72327b34f05", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_protocol_failure"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing criteria and routing instructions in both full inputs. The question bindings remain the same for the tomato trial, treatment-control validity, protocol, and watering-event time range. The evidence consists of two complete factual sentences: \"On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.\" and \"Manual-log entries 17–19 correspond to events E1, E2, and E3.\" The counterfactual coherently changes the second fact to entries 17–18 for E1 and E2 while explicitly leaving E3 undocumented. Neither context embeds a gold answer, label rationale, answer code, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal quantification over an explicit set does not make the atoms non-atomic. A6 is a factual manual-record coverage relation rather than a policy classification. The base and counter assignments can coexist with all non-focus atoms unchanged: in the counter case, an affected event can lack a manual entry while any existing corresponding entries remain dated and independent, and A3 can still hold through a different kind of watering-record entry. The state-derived policy evidence preserves the substantive protocol requirements and sensor-gap exception. Rules originating in the retained questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports all four stated core requirements and establishes the sensor gap plus complete, dated, independent manual coverage of every affected watering event. It therefore suffices for valid_documented_exception and excludes the enumerated protocol failures.", "rule_index": 0, "sound": true}, {"reason": "With the sensor gap established, refutation of A6 entails that at least one affected watering event lacks a corresponding manual-log entry. This is explicit failure of the required sensor exception, not merely unknown documentation, and suffices for invalid_protocol_failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assignment of trial benches 1–4 to drought or control before treatment began was randomized."}, {"id": "A2", "statement": "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days."}, {"id": "A3", "statement": "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records."}, {"id": "A4", "statement": "Each of the trial’s 48 plants has a final dry-mass measurement."}, {"id": "A5", "statement": "The watering-sensor stream for drought bench 3 is missing on trial days 4–6."}, {"id": "A6", "statement": "Every drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 has a corresponding manual-log entry."}, {"id": "A7", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated."}, {"id": "A8", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."}], "base_state_json": "{\"context\":\"A greenhouse trial used 48 tomato plants across benches 1–4. A signed sheet records that drought and control benches were assigned by random procedure before treatment. During each of the seven pre-trial days, all pots received the same watering volume, and the trial’s drought-pot register records treatment-period watering events. Each plant has a final dry-mass measurement. The watering-sensor stream for drought bench 3 is missing on trial days 4–6. The affected manual entries are dated and were created independently of the missing sensor stream. The protocol requirements are recorded as follows: The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event. The reviewer must classify the trial’s treatment-control validity and route it under the protocol.\",\"evidence\":[\"On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.\",\"Manual-log entries 17–19 correspond to events E1, E2, and E3.\"],\"policy\":[\"The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.\",\"One missing sensor stream is allowed only when dated manual records independently document every affected watering event.\",\"Classify the trial’s treatment-control validity under the stated protocol and route it accordingly.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": ["evidence", "0"], "text": "On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events."}, {"path": ["evidence", "1"], "text": "Manual-log entries 17–19 correspond to events E1, E2, and E3."}], "policy_evidence": [{"path": ["context"], "text": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements."}, {"path": ["context"], "text": "One missing sensor stream is allowed only when dated manual records independently document every affected watering event."}, {"path": ["request"], "text": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}], "rules": [{"justification": "All core requirements are explicitly supported, and every event affected by the established sensor gap has a dated manual-log entry independent of the missing sensor stream.", "target": "valid_documented_exception", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}, {"justification": "The established sensor gap has at least one affected watering event without a corresponding manual-log entry, so the claimed sensor exception fails and the protocol has an uncovered affected event.", "target": "invalid_protocol_failure", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.", "negative_left": "On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.", "negative_right": "Manual-log entries 17–18 correspond to events E1 and E2; no entry corresponds to E3.", "right": "Manual-log entries 17–19 correspond to events E1, E2, and E3."}, "verifier_independent_model": false}, "family": "fast-41-diverse-241-011", "id": "fast-41-diverse-241-011-base", "input": {"questions": {"decision": {"criteria": {"insufficient_evidence_hold": "Place on evidence hold: the record neither proves a core violation nor supplies enough explicit documentation to confirm all core requirements or a claimed sensor exception.", "invalid_protocol_failure": "Route for protocol failure: explicit evidence shows nonrandom assignment, unequal pre-trial watering, an uncovered watering event, missing final growth measurements, or another violated core requirement.", "valid_documented_exception": "Accept as treatment-control valid and route to analysis: every core requirement is explicitly supported, and any sensor gap is fully covered by qualifying dated manual records."}, "instructions": "Select exactly one routing option. Apply the explicit protocol rule: all core requirements must be evidenced; a sensor gap is acceptable only if independently documented manual records cover every affected watering event.", "type": "choice"}}, "state": {"context": "A greenhouse trial used 48 tomato plants across benches 1–4. A signed sheet records that drought and control benches were assigned by random procedure before treatment. During each of the seven pre-trial days, all pots received the same watering volume, and the trial’s drought-pot register records treatment-period watering events. Each plant has a final dry-mass measurement. The watering-sensor stream for drought bench 3 is missing on trial days 4–6. The affected manual entries are dated and were created independently of the missing sensor stream. The protocol requirements are recorded as follows: The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event. The reviewer must classify the trial’s treatment-control validity and route it under the protocol.", "evidence": ["On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.", "Manual-log entries 17–19 correspond to events E1, E2, and E3."], "policy": ["The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.", "One missing sensor stream is allowed only when dated manual records independently document every affected watering event.", "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."]}}, "method": "c2d", "provenance": {"source_id": "diverse-241", "source_is_synthetic": true, "source_sha256": "1a2bbfd79de14bae070972939aa5cb346cad567f60d5260d99a6e72327b34f05", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_documented_exception"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing criteria and routing instructions in both full inputs. The question bindings remain the same for the tomato trial, treatment-control validity, protocol, and watering-event time range. The evidence consists of two complete factual sentences: \"On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.\" and \"Manual-log entries 17–19 correspond to events E1, E2, and E3.\" The counterfactual coherently changes the second fact to entries 17–18 for E1 and E2 while explicitly leaving E3 undocumented. Neither context embeds a gold answer, label rationale, answer code, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal quantification over an explicit set does not make the atoms non-atomic. A6 is a factual manual-record coverage relation rather than a policy classification. The base and counter assignments can coexist with all non-focus atoms unchanged: in the counter case, an affected event can lack a manual entry while any existing corresponding entries remain dated and independent, and A3 can still hold through a different kind of watering-record entry. The state-derived policy evidence preserves the substantive protocol requirements and sensor-gap exception. Rules originating in the retained questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports all four stated core requirements and establishes the sensor gap plus complete, dated, independent manual coverage of every affected watering event. It therefore suffices for valid_documented_exception and excludes the enumerated protocol failures.", "rule_index": 0, "sound": true}, {"reason": "With the sensor gap established, refutation of A6 entails that at least one affected watering event lacks a corresponding manual-log entry. This is explicit failure of the required sensor exception, not merely unknown documentation, and suffices for invalid_protocol_failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assignment of trial benches 1–4 to drought or control before treatment began was randomized."}, {"id": "A2", "statement": "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days."}, {"id": "A3", "statement": "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records."}, {"id": "A4", "statement": "Each of the trial’s 48 plants has a final dry-mass measurement."}, {"id": "A5", "statement": "The watering-sensor stream for drought bench 3 is missing on trial days 4–6."}, {"id": "A6", "statement": "Every drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 has a corresponding manual-log entry."}, {"id": "A7", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated."}, {"id": "A8", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."}], "base_state_json": "{\"context\":\"A greenhouse trial used 48 tomato plants across benches 1–4. A signed sheet records that drought and control benches were assigned by random procedure before treatment. During each of the seven pre-trial days, all pots received the same watering volume, and the trial’s drought-pot register records treatment-period watering events. Each plant has a final dry-mass measurement. The watering-sensor stream for drought bench 3 is missing on trial days 4–6. The affected manual entries are dated and were created independently of the missing sensor stream. The protocol requirements are recorded as follows: The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event. The reviewer must classify the trial’s treatment-control validity and route it under the protocol.\",\"evidence\":[\"On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.\",\"Manual-log entries 17–19 correspond to events E1, E2, and E3.\"],\"policy\":[\"The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.\",\"One missing sensor stream is allowed only when dated manual records independently document every affected watering event.\",\"Classify the trial’s treatment-control validity under the stated protocol and route it accordingly.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": ["evidence", "0"], "text": "On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events."}, {"path": ["evidence", "1"], "text": "Manual-log entries 17–19 correspond to events E1, E2, and E3."}], "policy_evidence": [{"path": ["context"], "text": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements."}, {"path": ["context"], "text": "One missing sensor stream is allowed only when dated manual records independently document every affected watering event."}, {"path": ["request"], "text": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}], "rules": [{"justification": "All core requirements are explicitly supported, and every event affected by the established sensor gap has a dated manual-log entry independent of the missing sensor stream.", "target": "valid_documented_exception", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}, {"justification": "The established sensor gap has at least one affected watering event without a corresponding manual-log entry, so the claimed sensor exception fails and the protocol has an uncovered affected event.", "target": "invalid_protocol_failure", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.", "negative_left": "On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.", "negative_right": "Manual-log entries 17–18 correspond to events E1 and E2; no entry corresponds to E3.", "right": "Manual-log entries 17–19 correspond to events E1, E2, and E3."}, "verifier_independent_model": false}, "family": "fast-41-diverse-241-011", "id": "fast-41-diverse-241-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"insufficient_evidence_hold": "Place on evidence hold: the record neither proves a core violation nor supplies enough explicit documentation to confirm all core requirements or a claimed sensor exception.", "invalid_protocol_failure": "Route for protocol failure: explicit evidence shows nonrandom assignment, unequal pre-trial watering, an uncovered watering event, missing final growth measurements, or another violated core requirement.", "valid_documented_exception": "Accept as treatment-control valid and route to analysis: every core requirement is explicitly supported, and any sensor gap is fully covered by qualifying dated manual records."}, "instructions": "Select exactly one routing option. Apply the explicit protocol rule: all core requirements must be evidenced; a sensor gap is acceptable only if independently documented manual records cover every affected watering event.", "type": "choice"}}, "state": {"context": "A greenhouse trial used 48 tomato plants across benches 1–4. A signed sheet records that drought and control benches were assigned by random procedure before treatment. During each of the seven pre-trial days, all pots received the same watering volume, and the trial’s drought-pot register records treatment-period watering events. Each plant has a final dry-mass measurement. The watering-sensor stream for drought bench 3 is missing on trial days 4–6. The affected manual entries are dated and were created independently of the missing sensor stream. The protocol requirements are recorded as follows: The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements. One missing sensor stream is allowed only when dated manual records independently document every affected watering event. The reviewer must classify the trial’s treatment-control validity and route it under the protocol.", "evidence": ["On trial days 4–6, the sensor-stream gap affected exactly events E1, E2, and E3 among drought-bench 3 watering events.", "Manual-log entries 17–18 correspond to events E1 and E2; no entry corresponds to E3."], "policy": ["The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.", "One missing sensor stream is allowed only when dated manual records independently document every affected watering event.", "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."]}}, "method": "c2d", "provenance": {"source_id": "diverse-241", "source_is_synthetic": true, "source_sha256": "1a2bbfd79de14bae070972939aa5cb346cad567f60d5260d99a6e72327b34f05", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_protocol_failure"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both contexts alongside the unchanged original questions. Question entity, scope, and time bindings remain unchanged. The evidence consists of two complete factual sentences: \"During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.\" and \"The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5.\" The counterfactual coherently changes only the day-5 manual entry. Neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal quantification over an explicit set does not make the atoms non-atomic. A6 is a factual manual-record coverage relation rather than a policy classification. The base and counter assignments can coexist with all non-focus atoms unchanged: in the counter case, an affected event can lack a manual entry while any existing corresponding entries remain dated and independent, and A3 can still hold through a different kind of watering-record entry. The state-derived policy evidence preserves the substantive protocol requirements and sensor-gap exception. Rules originating in the retained questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports all four stated core requirements and establishes the sensor gap plus complete, dated, independent manual coverage of every affected watering event. It therefore suffices for valid_documented_exception and excludes the enumerated protocol failures.", "rule_index": 0, "sound": true}, {"reason": "With the sensor gap established, refutation of A6 entails that at least one affected watering event lacks a corresponding manual-log entry. This is explicit failure of the required sensor exception, not merely unknown documentation, and suffices for invalid_protocol_failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assignment of trial benches 1–4 to drought or control before treatment began was randomized."}, {"id": "A2", "statement": "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days."}, {"id": "A3", "statement": "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records."}, {"id": "A4", "statement": "Each of the trial’s 48 plants has a final dry-mass measurement."}, {"id": "A5", "statement": "The watering-sensor stream for drought bench 3 is missing on trial days 4–6."}, {"id": "A6", "statement": "Every drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 has a corresponding manual-log entry."}, {"id": "A7", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated."}, {"id": "A8", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."}], "base_state_json": "{\"case_note\":{\"policy\":[\"The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.\",\"One missing sensor stream is allowed only when dated manual records independently document every affected watering event.\",\"Classify the trial’s treatment-control validity under the stated protocol and route it accordingly.\"],\"evidence\":[\"The signed assignment sheet randomized benches 1–4 to drought or control before treatment began.\",\"A technician’s seven-day pre-trial log records the same watering volume for every trial pot on each day.\",\"The trial’s drought-pot watering records contain an entry for every drought-pot watering event during treatment.\",\"Final dry-mass measurements are recorded for all 48 plants.\",\"The watering-sensor stream for drought bench 3 is missing on trial days 4–6.\",\"During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.\",\"The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5.\",\"Each relevant manual-log entry is dated and was created independently of the missing sensor stream.\"]}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": ["case_note", "evidence", "5"], "text": "During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6."}, {"path": ["case_note", "evidence", "6"], "text": "The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5."}], "policy_evidence": [{"path": ["context"], "text": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements."}, {"path": ["context"], "text": "One missing sensor stream is allowed only when dated manual records independently document every affected watering event."}, {"path": ["request"], "text": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}], "rules": [{"justification": "All core requirements are explicitly supported, and every event affected by the established sensor gap has a dated manual-log entry independent of the missing sensor stream.", "target": "valid_documented_exception", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}, {"justification": "The established sensor gap has at least one affected watering event without a corresponding manual-log entry, so the claimed sensor exception fails and the protocol has an uncovered affected event.", "target": "invalid_protocol_failure", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.", "negative_left": "During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.", "negative_right": "The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, but no entry at 09:00 on day 5.", "right": "The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-241-019", "id": "fast-41-diverse-241-019-base", "input": {"questions": {"decision": {"criteria": {"insufficient_evidence_hold": "Place on evidence hold: the record neither proves a core violation nor supplies enough explicit documentation to confirm all core requirements or a claimed sensor exception.", "invalid_protocol_failure": "Route for protocol failure: explicit evidence shows nonrandom assignment, unequal pre-trial watering, an uncovered watering event, missing final growth measurements, or another violated core requirement.", "valid_documented_exception": "Accept as treatment-control valid and route to analysis: every core requirement is explicitly supported, and any sensor gap is fully covered by qualifying dated manual records."}, "instructions": "Select exactly one routing option. Apply the explicit protocol rule: all core requirements must be evidenced; a sensor gap is acceptable only if independently documented manual records cover every affected watering event.", "type": "choice"}}, "state": {"case_note": {"evidence": ["The signed assignment sheet randomized benches 1–4 to drought or control before treatment began.", "A technician’s seven-day pre-trial log records the same watering volume for every trial pot on each day.", "The trial’s drought-pot watering records contain an entry for every drought-pot watering event during treatment.", "Final dry-mass measurements are recorded for all 48 plants.", "The watering-sensor stream for drought bench 3 is missing on trial days 4–6.", "During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.", "The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5.", "Each relevant manual-log entry is dated and was created independently of the missing sensor stream."], "policy": ["The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.", "One missing sensor stream is allowed only when dated manual records independently document every affected watering event.", "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."]}}}, "method": "c2d", "provenance": {"source_id": "diverse-241", "source_is_synthetic": true, "source_sha256": "1a2bbfd79de14bae070972939aa5cb346cad567f60d5260d99a6e72327b34f05", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_documented_exception"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in both contexts alongside the unchanged original questions. Question entity, scope, and time bindings remain unchanged. The evidence consists of two complete factual sentences: \"During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.\" and \"The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5.\" The counterfactual coherently changes only the day-5 manual entry. Neither context embeds an answer, code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal quantification over an explicit set does not make the atoms non-atomic. A6 is a factual manual-record coverage relation rather than a policy classification. The base and counter assignments can coexist with all non-focus atoms unchanged: in the counter case, an affected event can lack a manual entry while any existing corresponding entries remain dated and independent, and A3 can still hold through a different kind of watering-record entry. The state-derived policy evidence preserves the substantive protocol requirements and sensor-gap exception. Rules originating in the retained questions object need not be duplicated in policy_evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports all four stated core requirements and establishes the sensor gap plus complete, dated, independent manual coverage of every affected watering event. It therefore suffices for valid_documented_exception and excludes the enumerated protocol failures.", "rule_index": 0, "sound": true}, {"reason": "With the sensor gap established, refutation of A6 entails that at least one affected watering event lacks a corresponding manual-log entry. This is explicit failure of the required sensor exception, not merely unknown documentation, and suffices for invalid_protocol_failure.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The assignment of trial benches 1–4 to drought or control before treatment began was randomized."}, {"id": "A2", "statement": "Every trial pot received the same watering volume as every other trial pot on each of the seven pre-trial days."}, {"id": "A3", "statement": "Every drought-pot watering event during the trial’s treatment period has an entry in the trial’s drought-pot watering records."}, {"id": "A4", "statement": "Each of the trial’s 48 plants has a final dry-mass measurement."}, {"id": "A5", "statement": "The watering-sensor stream for drought bench 3 is missing on trial days 4–6."}, {"id": "A6", "statement": "Every drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 has a corresponding manual-log entry."}, {"id": "A7", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is dated."}, {"id": "A8", "statement": "Every manual-log entry corresponding to a drought-bench 3 watering event affected by the sensor-stream gap on trial days 4–6 is independent of the missing sensor stream."}], "base_state_json": "{\"case_note\":{\"policy\":[\"The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.\",\"One missing sensor stream is allowed only when dated manual records independently document every affected watering event.\",\"Classify the trial’s treatment-control validity under the stated protocol and route it accordingly.\"],\"evidence\":[\"The signed assignment sheet randomized benches 1–4 to drought or control before treatment began.\",\"A technician’s seven-day pre-trial log records the same watering volume for every trial pot on each day.\",\"The trial’s drought-pot watering records contain an entry for every drought-pot watering event during treatment.\",\"Final dry-mass measurements are recorded for all 48 plants.\",\"The watering-sensor stream for drought bench 3 is missing on trial days 4–6.\",\"During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.\",\"The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5.\",\"Each relevant manual-log entry is dated and was created independently of the missing sensor stream.\"]}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": ["case_note", "evidence", "5"], "text": "During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6."}, {"path": ["case_note", "evidence", "6"], "text": "The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5."}], "policy_evidence": [{"path": ["context"], "text": "The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements."}, {"path": ["context"], "text": "One missing sensor stream is allowed only when dated manual records independently document every affected watering event."}, {"path": ["request"], "text": "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."}], "rules": [{"justification": "All core requirements are explicitly supported, and every event affected by the established sensor gap has a dated manual-log entry independent of the missing sensor stream.", "target": "valid_documented_exception", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}, {"justification": "The established sensor gap has at least one affected watering event without a corresponding manual-log entry, so the claimed sensor exception fails and the protocol has an uncovered affected event.", "target": "invalid_protocol_failure", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.", "negative_left": "During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.", "negative_right": "The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, but no entry at 09:00 on day 5.", "right": "The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, and at 09:00 on day 5."}, "verifier_independent_model": false}, "family": "fast-41-diverse-241-019", "id": "fast-41-diverse-241-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"insufficient_evidence_hold": "Place on evidence hold: the record neither proves a core violation nor supplies enough explicit documentation to confirm all core requirements or a claimed sensor exception.", "invalid_protocol_failure": "Route for protocol failure: explicit evidence shows nonrandom assignment, unequal pre-trial watering, an uncovered watering event, missing final growth measurements, or another violated core requirement.", "valid_documented_exception": "Accept as treatment-control valid and route to analysis: every core requirement is explicitly supported, and any sensor gap is fully covered by qualifying dated manual records."}, "instructions": "Select exactly one routing option. Apply the explicit protocol rule: all core requirements must be evidenced; a sensor gap is acceptable only if independently documented manual records cover every affected watering event.", "type": "choice"}}, "state": {"case_note": {"evidence": ["The signed assignment sheet randomized benches 1–4 to drought or control before treatment began.", "A technician’s seven-day pre-trial log records the same watering volume for every trial pot on each day.", "The trial’s drought-pot watering records contain an entry for every drought-pot watering event during treatment.", "Final dry-mass measurements are recorded for all 48 plants.", "The watering-sensor stream for drought bench 3 is missing on trial days 4–6.", "During trial days 4–6, exactly three drought-bench 3 watering events affected by the sensor-stream gap occurred at 09:00 on days 4, 5, and 6.", "The drought-bench 3 manual log contains entries at 09:00 on days 4 and 6, but no entry at 09:00 on day 5.", "Each relevant manual-log entry is dated and was created independently of the missing sensor stream."], "policy": ["The protocol requires randomized treatment assignment, identical pre-trial watering, drought-pot watering records, and final growth measurements.", "One missing sensor stream is allowed only when dated manual records independently document every affected watering event.", "Classify the trial’s treatment-control validity under the stated protocol and route it accordingly."]}}}, "method": "c2d", "provenance": {"source_id": "diverse-241", "source_is_synthetic": true, "source_sha256": "1a2bbfd79de14bae070972939aa5cb346cad567f60d5260d99a6e72327b34f05", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_protocol_failure"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and criteria. Both contexts retain the same trial, tray, and time bindings. The two evidence spans are complete factual sentences: “Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group.” “During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule.” The counterfactual changes only tray 7’s watering observation and remains coherent with the documented assignment. Neither context contains answer codes, rule tables, rationale, proposition IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"plant ecologist\",\"text\":\"Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group.\"},{\"speaker\":\"greenhouse technician\",\"text\":\"During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule.\"},{\"speaker\":\"data quality reviewer\",\"text\":\"The trial log contains 180 of the 200 required daily soil-moisture readings, and the required final height and biomass measurements are present for all 20 trays. Tray identifiers are consistent across the assignment, watering, and measurement records.\"},{\"speaker\":\"plant ecologist\",\"text\":\"The assignment records were completed before watering began, and the endpoint files identify both required measurements for every tray.\"},{\"speaker\":\"protocol note\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group."}, {"path": ["1", "text"], "text": "During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group.", "negative_left": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group.", "negative_right": "During the 10-day bean-tray trial, every tray received the watering schedule corresponding to its documented assignment.", "right": "During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule."}, "verifier_independent_model": false}, "family": "fast-41-diverse-242-001", "id": "fast-41-diverse-242-001-base", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "plant ecologist", "text": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group."}, {"speaker": "greenhouse technician", "text": "During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule."}, {"speaker": "data quality reviewer", "text": "The trial log contains 180 of the 200 required daily soil-moisture readings, and the required final height and biomass measurements are present for all 20 trays. Tray identifiers are consistent across the assignment, watering, and measurement records."}, {"speaker": "plant ecologist", "text": "The assignment records were completed before watering began, and the endpoint files identify both required measurements for every tray."}, {"speaker": "protocol note", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_documented_crossover"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and criteria. Both contexts retain the same trial, tray, and time bindings. The two evidence spans are complete factual sentences: “Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group.” “During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule.” The counterfactual changes only tray 7’s watering observation and remains coherent with the documented assignment. Neither context contains answer codes, rule tables, rationale, proposition IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"plant ecologist\",\"text\":\"Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group.\"},{\"speaker\":\"greenhouse technician\",\"text\":\"During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule.\"},{\"speaker\":\"data quality reviewer\",\"text\":\"The trial log contains 180 of the 200 required daily soil-moisture readings, and the required final height and biomass measurements are present for all 20 trays. Tray identifiers are consistent across the assignment, watering, and measurement records.\"},{\"speaker\":\"plant ecologist\",\"text\":\"The assignment records were completed before watering began, and the endpoint files identify both required measurements for every tray.\"},{\"speaker\":\"protocol note\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group."}, {"path": ["1", "text"], "text": "During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group.", "negative_left": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group.", "negative_right": "During the 10-day bean-tray trial, every tray received the watering schedule corresponding to its documented assignment.", "right": "During the 10-day bean-tray trial, tray 7 received the drought-treatment watering schedule."}, "verifier_independent_model": false}, "family": "fast-41-diverse-242-001", "id": "fast-41-diverse-242-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "plant ecologist", "text": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and tray 7 was assigned to the control group."}, {"speaker": "greenhouse technician", "text": "During the 10-day bean-tray trial, every tray received the watering schedule corresponding to its documented assignment."}, {"speaker": "data quality reviewer", "text": "The trial log contains 180 of the 200 required daily soil-moisture readings, and the required final height and biomass measurements are present for all 20 trays. Tray identifiers are consistent across the assignment, watering, and measurement records."}, {"speaker": "plant ecologist", "text": "The assignment records were completed before watering began, and the endpoint files identify both required measurements for every tray."}, {"speaker": "protocol note", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_at_threshold"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain the bean-tray trial scope and 10-day treatment-control bindings. The two evidence spans are complete factual sentences: \"Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group.\" and \"During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule.\" The counterfactual coherently changes Tray 14's schedule while preserving the assignment and all other facts. Neither context embeds a gold answer, answer code, rule table, proposition identifiers, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"plant ecologist\",\"text\":\"Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group.\"},{\"speaker\":\"trial coordinator\",\"text\":\"During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule.\"},{\"speaker\":\"data quality reviewer\",\"text\":\"The monitoring log contains 180 of the 200 required daily soil-moisture readings, and the final height and final biomass measurements are present for each of the 20 trays. No tray identifiers are duplicated.\"},{\"speaker\":\"records manager\",\"text\":\"The completed records were retained with the trial file for audit review.\"},{\"speaker\":\"policy record\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group."}, {"path": ["1", "text"], "text": "During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group.", "negative_left": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group.", "negative_right": "During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the drought-treatment watering schedule.", "right": "During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule."}, "verifier_independent_model": false}, "family": "fast-41-diverse-242-006", "id": "fast-41-diverse-242-006-base", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "plant ecologist", "text": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group."}, {"speaker": "trial coordinator", "text": "During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule."}, {"speaker": "data quality reviewer", "text": "The monitoring log contains 180 of the 200 required daily soil-moisture readings, and the final height and final biomass measurements are present for each of the 20 trays. No tray identifiers are duplicated."}, {"speaker": "records manager", "text": "The completed records were retained with the trial file for audit review."}, {"speaker": "policy record", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_documented_crossover"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain the bean-tray trial scope and 10-day treatment-control bindings. The two evidence spans are complete factual sentences: \"Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group.\" and \"During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule.\" The counterfactual coherently changes Tray 14's schedule while preserving the assignment and all other facts. Neither context embeds a gold answer, answer code, rule table, proposition identifiers, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"plant ecologist\",\"text\":\"Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group.\"},{\"speaker\":\"trial coordinator\",\"text\":\"During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule.\"},{\"speaker\":\"data quality reviewer\",\"text\":\"The monitoring log contains 180 of the 200 required daily soil-moisture readings, and the final height and final biomass measurements are present for each of the 20 trays. No tray identifiers are duplicated.\"},{\"speaker\":\"records manager\",\"text\":\"The completed records were retained with the trial file for audit review.\"},{\"speaker\":\"policy record\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group."}, {"path": ["1", "text"], "text": "During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group.", "negative_left": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group.", "negative_right": "During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the drought-treatment watering schedule.", "right": "During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the control-group watering schedule."}, "verifier_independent_model": false}, "family": "fast-41-diverse-242-006", "id": "fast-41-diverse-242-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "plant ecologist", "text": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group, and Tray 14 was assigned to the drought-treatment group."}, {"speaker": "trial coordinator", "text": "During the 10-day bean-tray trial, every tray other than Tray 14 received its documented assigned watering schedule, and Tray 14 received the drought-treatment watering schedule."}, {"speaker": "data quality reviewer", "text": "The monitoring log contains 180 of the 200 required daily soil-moisture readings, and the final height and final biomass measurements are present for each of the 20 trays. No tray identifiers are duplicated."}, {"speaker": "records manager", "text": "The completed records were retained with the trial file for audit review."}, {"speaker": "policy record", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_at_threshold"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is unchanged, the bindings and policy remain intact, and the evidence quotes are complete factual sentences: \"Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group.\" and \"During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule.\" The counterfactual changes only the D-tray watering observation and is coherent with the remaining facts, and neither context leaks an answer or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"trial coordinator\",\"text\":\"Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group.\"},{\"speaker\":\"watering-log reviewer\",\"text\":\"During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule.\"},{\"speaker\":\"data manager\",\"text\":\"The archive contains 184 of the 200 scheduled daily soil-moisture readings. Each tray has a recorded final height and a recorded final biomass measurement, and all 20 tray identifiers reconcile across the files.\"},{\"speaker\":\"trial coordinator\",\"text\":\"The assignment register, watering logs, sensor export, and endpoint worksheet were locked after review. No records were discarded from the endpoint worksheet, and the sensor count was checked against the 20-tray, 10-day schedule.\"},{\"speaker\":\"protocol note\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group."}, {"path": ["1", "text"], "text": "During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group.", "negative_left": "Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group.", "negative_right": "During the 10-day bean-tray trial, trays D01 through D12 received the drought-treatment watering schedule, while trays C01 through C08 received the control-group watering schedule.", "right": "During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule."}, "verifier_independent_model": false}, "family": "fast-41-diverse-242-007", "id": "fast-41-diverse-242-007-base", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "trial coordinator", "text": "Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group."}, {"speaker": "watering-log reviewer", "text": "During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule."}, {"speaker": "data manager", "text": "The archive contains 184 of the 200 scheduled daily soil-moisture readings. Each tray has a recorded final height and a recorded final biomass measurement, and all 20 tray identifiers reconcile across the files."}, {"speaker": "trial coordinator", "text": "The assignment register, watering logs, sensor export, and endpoint worksheet were locked after review. No records were discarded from the endpoint worksheet, and the sensor count was checked against the 20-tray, 10-day schedule."}, {"speaker": "protocol note", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_documented_crossover"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is unchanged, the bindings and policy remain intact, and the evidence quotes are complete factual sentences: \"Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group.\" and \"During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule.\" The counterfactual changes only the D-tray watering observation and is coherent with the remaining facts, and neither context leaks an answer or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"trial coordinator\",\"text\":\"Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group.\"},{\"speaker\":\"watering-log reviewer\",\"text\":\"During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule.\"},{\"speaker\":\"data manager\",\"text\":\"The archive contains 184 of the 200 scheduled daily soil-moisture readings. Each tray has a recorded final height and a recorded final biomass measurement, and all 20 tray identifiers reconcile across the files.\"},{\"speaker\":\"trial coordinator\",\"text\":\"The assignment register, watering logs, sensor export, and endpoint worksheet were locked after review. No records were discarded from the endpoint worksheet, and the sensor count was checked against the 20-tray, 10-day schedule.\"},{\"speaker\":\"protocol note\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group."}, {"path": ["1", "text"], "text": "During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group.", "negative_left": "Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group.", "negative_right": "During the 10-day bean-tray trial, trays D01 through D12 received the drought-treatment watering schedule, while trays C01 through C08 received the control-group watering schedule.", "right": "During the 10-day bean-tray trial, trays D01 through D12 received the control-group watering schedule, while trays C01 through C08 received the control-group watering schedule."}, "verifier_independent_model": false}, "family": "fast-41-diverse-242-007", "id": "fast-41-diverse-242-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "trial coordinator", "text": "Before watering began in the 10-day bean-tray trial, trays D01 through D12 were documented as randomly assigned to the drought-treatment group, and trays C01 through C08 were documented as randomly assigned to the control group."}, {"speaker": "watering-log reviewer", "text": "During the 10-day bean-tray trial, trays D01 through D12 received the drought-treatment watering schedule, while trays C01 through C08 received the control-group watering schedule."}, {"speaker": "data manager", "text": "The archive contains 184 of the 200 scheduled daily soil-moisture readings. Each tray has a recorded final height and a recorded final biomass measurement, and all 20 tray identifiers reconcile across the files."}, {"speaker": "trial coordinator", "text": "The assignment register, watering logs, sensor export, and endpoint worksheet were locked after review. No records were discarded from the endpoint worksheet, and the sensor count was checked against the 20-tray, 10-day schedule."}, {"speaker": "protocol note", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_at_threshold"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy remains in the verbatim questions, and both contexts preserve the bean-trial, treatment-control, watering, measurement, and time bindings. The required evidence quotes are “Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group.” and “During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule.” The counterfactual changes only the watering assignments and is coherent with the unchanged measurements and protocol. Neither context embeds an answer, code, rule table, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"trial coordinator\",\"text\":\"Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group.\"},{\"speaker\":\"watering log\",\"text\":\"During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule.\"},{\"speaker\":\"data reviewer\",\"text\":\"The archive contains 180 of the 200 required daily soil-moisture readings, with records distributed across the 20 trays and 10 trial days.\"},{\"speaker\":\"measurement reviewer\",\"text\":\"The required final height and final biomass measurements are present for each of the 20 trays, and tray identifiers remain consistent across the archive.\"},{\"speaker\":\"protocol record\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group."}, {"path": ["1", "text"], "text": "During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group.", "negative_left": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group.", "negative_right": "During the 10-day bean-tray trial, trays 1 through 10 received the drought-treatment group's watering schedule, and trays 11 through 20 received the control group's watering schedule.", "right": "During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule."}, "verifier_independent_model": false}, "family": "fast-41-diverse-242-010", "id": "fast-41-diverse-242-010-base", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "trial coordinator", "text": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group."}, {"speaker": "watering log", "text": "During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule."}, {"speaker": "data reviewer", "text": "The archive contains 180 of the 200 required daily soil-moisture readings, with records distributed across the 20 trays and 10 trial days."}, {"speaker": "measurement reviewer", "text": "The required final height and final biomass measurements are present for each of the 20 trays, and tray identifiers remain consistent across the archive."}, {"speaker": "protocol record", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "invalid_documented_crossover"}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy remains in the verbatim questions, and both contexts preserve the bean-trial, treatment-control, watering, measurement, and time bindings. The required evidence quotes are “Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group.” and “During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule.” The counterfactual changes only the watering assignments and is coherent with the unchanged measurements and protocol. Neither context embeds an answer, code, rule table, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, including universal presence claims over explicit sets. A2 is factual rather than a policy classification. The base and counter assignments are jointly realizable and differ only in crossover status. The state-derived policy evidence preserves the trial-specific measurement scope needed to interpret sensor and endpoint completeness; the remaining governing rubric is already retained in the questions object and should not be duplicated. Both rules are sufficient for their targets, and the partial rule table may validly abstain on other assignments.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A2 supported establishes that at least one tray received the opposite group's watering schedule. The rubric makes any documented crossover sufficient for INVALID, irrespective of the other conditions.", "rule_index": 0, "sound": true}, {"reason": "A1 supplies documented assignment, A2 refuted establishes no crossover, A3 supplies at least 90% sensor completeness, and A4 supplies every required endpoint. Together these satisfy VALID and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Before watering began in the 10-day bean-tray trial, each of the 20 bean trays had a documented randomized assignment to the drought-treatment group or the control group."}, {"id": "A2", "statement": "During the 10-day bean-tray trial, at least one of the 20 bean trays received the watering schedule assigned to the group opposite that tray's documented pre-watering randomized assignment."}, {"id": "A3", "statement": "At least 90% of the 200 required daily soil-moisture readings from the 20 bean trays over the 10-day trial are present."}, {"id": "A4", "statement": "For each of the 20 bean trays, the required final height measurement and the required final biomass measurement are present."}], "base_state_json": "[{\"speaker\":\"trial coordinator\",\"text\":\"Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group.\"},{\"speaker\":\"watering log\",\"text\":\"During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule.\"},{\"speaker\":\"data reviewer\",\"text\":\"The archive contains 180 of the 200 required daily soil-moisture readings, with records distributed across the 20 trays and 10 trial days.\"},{\"speaker\":\"measurement reviewer\",\"text\":\"The required final height and final biomass measurements are present for each of the 20 trays, and tray identifiers remain consistent across the archive.\"},{\"speaker\":\"protocol record\",\"text\":\"The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["0", "text"], "text": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group."}, {"path": ["1", "text"], "text": "During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule."}], "policy_evidence": [{"path": ["0", "text"], "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}], "rules": [{"justification": "A documented factual crossover by at least one tray directly satisfies the INVALID criterion regardless of assignment completeness, sensor completeness, or endpoint availability.", "target": "invalid_documented_crossover", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "Documented assignments, established absence of crossover, at least 90% of the required sensor readings, and every required final height and biomass measurement jointly satisfy all VALID requirements and exclude INVALID, PROVISIONAL, and UNVERIFIABLE.", "target": "valid_at_threshold", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group.", "negative_left": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group.", "negative_right": "During the 10-day bean-tray trial, trays 1 through 10 received the drought-treatment group's watering schedule, and trays 11 through 20 received the control group's watering schedule.", "right": "During the 10-day bean-tray trial, trays 1 through 10 received the control group's watering schedule, and trays 11 through 20 received the drought-treatment group's watering schedule."}, "verifier_independent_model": false}, "family": "fast-41-diverse-242-010", "id": "fast-41-diverse-242-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"invalid_documented_crossover": "INVALID: At least one tray received the watering schedule assigned to the opposite group.", "provisional_minor_data_gap": "PROVISIONAL: No crossover, but sensor completeness is 80% to below 90%, or exactly one endpoint is missing.", "unverifiable_or_severely_incomplete": "UNVERIFIABLE: Evidence does not establish whether crossover occurred, or sensor completeness is below 80%.", "valid_at_threshold": "VALID: Assignment and absence of crossover are documented, sensor completeness is at least 90%, and every endpoint is present."}, "instructions": "Verify treatment-control validity using this rubric: VALID requires documented assignment, evidence of no watering crossover, at least 90% sensor completeness, and all endpoint measurements. PROVISIONAL requires no crossover but 80% to below 90% sensor completeness or one missing endpoint. INVALID applies when any watering crossover is documented. UNVERIFIABLE applies when crossover status lacks evidence or sensor completeness is below 80%. Select exactly one option.", "type": "choice"}}, "state": [{"speaker": "trial coordinator", "text": "Before watering began in the 10-day bean-tray trial, the documented randomized assignment placed trays 1 through 10 in the drought-treatment group and trays 11 through 20 in the control group."}, {"speaker": "watering log", "text": "During the 10-day bean-tray trial, trays 1 through 10 received the drought-treatment group's watering schedule, and trays 11 through 20 received the control group's watering schedule."}, {"speaker": "data reviewer", "text": "The archive contains 180 of the 200 required daily soil-moisture readings, with records distributed across the 20 trays and 10 trial days."}, {"speaker": "measurement reviewer", "text": "The required final height and final biomass measurements are present for each of the 20 trays, and tray identifiers remain consistent across the archive."}, {"speaker": "protocol record", "text": "The trial required daily soil-moisture readings for 10 days, plus final height and biomass for every tray."}]}, "method": "c2d", "provenance": {"source_id": "diverse-242", "source_is_synthetic": true, "source_sha256": "2314e0b32f91276e74cd948155a7f883b0e7013a892ca1e388bd96d60660b34d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "valid_at_threshold"}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, scope, and bindings, and both contexts repeat the assignment instruction consistently. Both contexts remain bound to tray B, the four sensors, the 08:00 calibration, and the pre-watering decision. The two evidence spans are complete factual sentences: \"The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4.\" \"Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed.\" The counterfactual coherently changes only the remaining sensor’s result from passed to failed. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before irrigation began, Mira’s greenhouse log identified the sensors relevant to tray B’s drought-assignment checkpoint and recorded the 08:00 calibration before any watering decision. The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4. Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed. Mira’s written instruction was: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The technician made the watering decision only after the checklist was complete. Subsequent growth measurements and irrigation records were logged separately from the calibration review, and no later observation altered the 08:00 checklist.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"path": [], "text": "Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4.", "negative_left": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4.", "negative_right": "Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor failed.", "right": "Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-243-008", "id": "fast-41-diverse-243-008-base", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before irrigation began, Mira’s greenhouse log identified the sensors relevant to tray B’s drought-assignment checkpoint and recorded the 08:00 calibration before any watering decision. The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4. Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed. Mira’s written instruction was: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The technician made the watering decision only after the checklist was complete. Subsequent growth measurements and irrigation records were logged separately from the calibration review, and no later observation altered the 08:00 checklist."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, criteria, scope, and bindings, and both contexts repeat the assignment instruction consistently. Both contexts remain bound to tray B, the four sensors, the 08:00 calibration, and the pre-watering decision. The two evidence spans are complete factual sentences: \"The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4.\" \"Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed.\" The counterfactual coherently changes only the remaining sensor’s result from passed to failed. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before irrigation began, Mira’s greenhouse log identified the sensors relevant to tray B’s drought-assignment checkpoint and recorded the 08:00 calibration before any watering decision. The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4. Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed. Mira’s written instruction was: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The technician made the watering decision only after the checklist was complete. Subsequent growth measurements and irrigation records were logged separately from the calibration review, and no later observation altered the 08:00 checklist.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"path": [], "text": "Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4.", "negative_left": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4.", "negative_right": "Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor failed.", "right": "Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor passed."}, "verifier_independent_model": false}, "family": "fast-41-diverse-243-008", "id": "fast-41-diverse-243-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before irrigation began, Mira’s greenhouse log identified the sensors relevant to tray B’s drought-assignment checkpoint and recorded the 08:00 calibration before any watering decision. The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4. Among the four sensors governed by Mira’s 08:00 pre-watering calibration, B-1, B-2, and B-3 were the three sensors that passed, and the remaining governed sensor failed. Mira’s written instruction was: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The technician made the watering decision only after the checklist was complete. Subsequent growth measurements and irrigation records were logged separately from the calibration review, and no later observation altered the 08:00 checklist."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing protocol is unchanged and the original questions object remains verbatim. The request retains the same trial, analysis decision, and temporal-update scope. The two focus spans are complete factual sentences: “The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.” and “The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31.” The counterfactual changes only D-17’s delivered volume to 140 mL without creating contradictory duplicate measurements or assertions. Neither context embeds a gold answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.\",\"The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31.\",\"The signed trial correction and controller audit place baseline height at 11 May 09:00 and random assignment at 11 May 09:20.\",\"The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and each scheduled reading affected by that gap has a gravimetric backup.\",\"The protocol-required final-height measurements are all recorded in the growth table, which the data-quality reviewer checked on 15 May.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL."}, {"path": ["evidence", "1"], "text": "The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.", "negative_left": "The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.", "negative_right": "The trial log records delivered volumes of 140 mL for D-17, 90 mL for C-22, and 84 mL for D-31.", "right": "The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-002", "id": "fast-41-diverse-244-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "evidence": ["The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.", "The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31.", "The signed trial correction and controller audit place baseline height at 11 May 09:00 and random assignment at 11 May 09:20.", "The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and each scheduled reading affected by that gap has a gravimetric backup.", "The protocol-required final-height measurements are all recorded in the growth table, which the data-quality reviewer checked on 15 May."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing protocol is unchanged and the original questions object remains verbatim. The request retains the same trial, analysis decision, and temporal-update scope. The two focus spans are complete factual sentences: “The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.” and “The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31.” The counterfactual changes only D-17’s delivered volume to 140 mL without creating contradictory duplicate measurements or assertions. Neither context embeds a gold answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.\",\"The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31.\",\"The signed trial correction and controller audit place baseline height at 11 May 09:00 and random assignment at 11 May 09:20.\",\"The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and each scheduled reading affected by that gap has a gravimetric backup.\",\"The protocol-required final-height measurements are all recorded in the growth table, which the data-quality reviewer checked on 15 May.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL."}, {"path": ["evidence", "1"], "text": "The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.", "negative_left": "The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.", "negative_right": "The trial log records delivered volumes of 140 mL for D-17, 90 mL for C-22, and 84 mL for D-31.", "right": "The trial log records delivered volumes of 128 mL for D-17, 90 mL for C-22, and 84 mL for D-31."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-002", "id": "fast-41-diverse-244-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "evidence": ["The fictional tomato drought trial's complete watering schedule contains exactly three events: drought-arm event D-17 prescribed 120 mL, control-arm event C-22 prescribed 95 mL, and drought-arm event D-31 prescribed 80 mL.", "The trial log records delivered volumes of 140 mL for D-17, 90 mL for C-22, and 84 mL for D-31.", "The signed trial correction and controller audit place baseline height at 11 May 09:00 and random assignment at 11 May 09:20.", "The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and each scheduled reading affected by that gap has a gravimetric backup.", "The protocol-required final-height measurements are all recorded in the growth table, which the data-quality reviewer checked on 15 May."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the original questions verbatim, preserve policy and bindings, use two complete factual evidence sentences, and present a coherent counterfactual without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"A plant ecologist is reviewing the fictional tomato drought trial for treatment-effect analysis. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit. An unsigned note initially placed randomization at 08:40 on 11 May, but a signed correction places it at 09:20, supported by the controller audit. Baseline height was recorded at 09:00 that morning. The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and each affected scheduled reading has a gravimetric backup. Every protocol-required final-height measurement is recorded. The data-quality reviewer checked the audit, backup weights, and growth table.\",\"evidence\":[\"The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL.\",\"The recorded delivered volumes were 147 mL for drought-arm event D17 and 153 mL for control-arm event C23 in the fictional tomato drought trial.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL."}, {"path": ["evidence", "1"], "text": "The recorded delivered volumes were 147 mL for drought-arm event D17 and 153 mL for control-arm event C23 in the fictional tomato drought trial."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL.", "negative_left": "The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL.", "negative_right": "The recorded delivered volumes were 147 mL for drought-arm event D17 and 175 mL for control-arm event C23 in the fictional tomato drought trial.", "right": "The recorded delivered volumes were 147 mL for drought-arm event D17 and 153 mL for control-arm event C23 in the fictional tomato drought trial."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-005", "id": "fast-41-diverse-244-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "A plant ecologist is reviewing the fictional tomato drought trial for treatment-effect analysis. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit. An unsigned note initially placed randomization at 08:40 on 11 May, but a signed correction places it at 09:20, supported by the controller audit. Baseline height was recorded at 09:00 that morning. The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and each affected scheduled reading has a gravimetric backup. Every protocol-required final-height measurement is recorded. The data-quality reviewer checked the audit, backup weights, and growth table.", "evidence": ["The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL.", "The recorded delivered volumes were 147 mL for drought-arm event D17 and 153 mL for control-arm event C23 in the fictional tomato drought trial."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the original questions verbatim, preserve policy and bindings, use two complete factual evidence sentences, and present a coherent counterfactual without answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"A plant ecologist is reviewing the fictional tomato drought trial for treatment-effect analysis. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit. An unsigned note initially placed randomization at 08:40 on 11 May, but a signed correction places it at 09:20, supported by the controller audit. Baseline height was recorded at 09:00 that morning. The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and each affected scheduled reading has a gravimetric backup. Every protocol-required final-height measurement is recorded. The data-quality reviewer checked the audit, backup weights, and growth table.\",\"evidence\":[\"The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL.\",\"The recorded delivered volumes were 147 mL for drought-arm event D17 and 153 mL for control-arm event C23 in the fictional tomato drought trial.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL."}, {"path": ["evidence", "1"], "text": "The recorded delivered volumes were 147 mL for drought-arm event D17 and 153 mL for control-arm event C23 in the fictional tomato drought trial."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL.", "negative_left": "The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL.", "negative_right": "The recorded delivered volumes were 147 mL for drought-arm event D17 and 175 mL for control-arm event C23 in the fictional tomato drought trial.", "right": "The recorded delivered volumes were 147 mL for drought-arm event D17 and 153 mL for control-arm event C23 in the fictional tomato drought trial."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-005", "id": "fast-41-diverse-244-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "A plant ecologist is reviewing the fictional tomato drought trial for treatment-effect analysis. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit. An unsigned note initially placed randomization at 08:40 on 11 May, but a signed correction places it at 09:20, supported by the controller audit. Baseline height was recorded at 09:00 that morning. The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and each affected scheduled reading has a gravimetric backup. Every protocol-required final-height measurement is recorded. The data-quality reviewer checked the audit, backup weights, and growth table.", "evidence": ["The fictional tomato drought trial has exactly two watering events: drought-arm event D17 prescribed 140 mL and control-arm event C23 prescribed 160 mL.", "The recorded delivered volumes were 147 mL for drought-arm event D17 and 175 mL for control-arm event C23 in the fictional tomato drought trial."], "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and include the unchanged original questions object. The request remains bound to the fictional tomato drought trial and treatment-effect analysis. The evidence spans are complete factual sentences: “The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.” “The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4.” The counterfactual changes only D7’s recorded volume from 46 mL to 53 mL without creating contradictory duplicate assertions. Neither context contains an answer, label rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"case_note\":\"The tomato drought trial's signed correction and controller audit confirm that baseline height was recorded on 11 May at 09:00 and random assignment at 09:20. The sole scheduled moisture-sensor reading gap occurred on 14 May, and each affected scheduled reading has a gravimetric backup. Every protocol-required final-height measurement is recorded. A reviewer checked the correction, audit, backup weights, watering log, and growth table on 15 May. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"watering_events\":\"The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.\",\"watering_log\":\"The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4.\",\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["watering_events"], "text": "The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL."}, {"path": ["watering_log"], "text": "The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.", "negative_left": "The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.", "negative_right": "The fictional tomato drought trial's log records 53 mL delivered for event D7 and 54 mL delivered for event C4.", "right": "The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-006", "id": "fast-41-diverse-244-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"case_note": "The tomato drought trial's signed correction and controller audit confirm that baseline height was recorded on 11 May at 09:00 and random assignment at 09:20. The sole scheduled moisture-sensor reading gap occurred on 14 May, and each affected scheduled reading has a gravimetric backup. Every protocol-required final-height measurement is recorded. A reviewer checked the correction, audit, backup weights, watering log, and growth table on 15 May. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?", "watering_events": "The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.", "watering_log": "The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4."}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy and include the unchanged original questions object. The request remains bound to the fictional tomato drought trial and treatment-effect analysis. The evidence spans are complete factual sentences: “The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.” “The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4.” The counterfactual changes only D7’s recorded volume from 46 mL to 53 mL without creating contradictory duplicate assertions. Neither context contains an answer, label rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"case_note\":\"The tomato drought trial's signed correction and controller audit confirm that baseline height was recorded on 11 May at 09:00 and random assignment at 09:20. The sole scheduled moisture-sensor reading gap occurred on 14 May, and each affected scheduled reading has a gravimetric backup. Every protocol-required final-height measurement is recorded. A reviewer checked the correction, audit, backup weights, watering log, and growth table on 15 May. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"watering_events\":\"The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.\",\"watering_log\":\"The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4.\",\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["watering_events"], "text": "The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL."}, {"path": ["watering_log"], "text": "The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.", "negative_left": "The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.", "negative_right": "The fictional tomato drought trial's log records 53 mL delivered for event D7 and 54 mL delivered for event C4.", "right": "The fictional tomato drought trial's log records 46 mL delivered for event D7 and 54 mL delivered for event C4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-006", "id": "fast-41-diverse-244-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"case_note": "The tomato drought trial's signed correction and controller audit confirm that baseline height was recorded on 11 May at 09:00 and random assignment at 09:20. The sole scheduled moisture-sensor reading gap occurred on 14 May, and each affected scheduled reading has a gravimetric backup. Every protocol-required final-height measurement is recorded. A reviewer checked the correction, audit, backup weights, watering log, and growth table on 15 May. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?", "watering_events": "The fictional tomato drought trial has exactly two watering events: drought-arm event D7 with a protocol-prescribed volume of 40 mL and control-arm event C4 with a protocol-prescribed volume of 50 mL.", "watering_log": "The fictional tomato drought trial's log records 53 mL delivered for event D7 and 54 mL delivered for event C4."}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question and policy, preserve the tomato-trial request and bindings, contain coherent noncontradictory observations, and include the complete factual evidence quotes: \"In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.\" and \"The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"case_note\":\"The tomato drought trial's operative baseline-height timestamp was 11 May at 09:00, and its operative random-assignment timestamp was 11 May at 09:20. The trial schedule contains exactly one moisture-sensor reading gap, on 14 May; each scheduled moisture reading omitted because of that gap has a gravimetric backup. Every protocol-required final-height measurement is recorded. The signed trial log and scheduled-event register provide the relevant watering records, while the study file identifies drought-arm and control-arm events by arm, date, time, and protocol volume. The data-quality review checked the assignment record, watering entries, sensor backups, and final-height table. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit. Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\",\"evidence\":[\"In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.\",\"The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23.\",\"The operative baseline-height timestamp is 11 May at 09:00, before the operative random-assignment timestamp of 11 May at 09:20.\",\"The schedule has exactly one moisture-sensor reading gap, on 14 May, and every scheduled reading omitted because of it has a gravimetric backup.\",\"Every protocol-required final-height measurement is recorded.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL."}, {"path": ["evidence", "1"], "text": "The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.", "negative_left": "In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.", "negative_right": "The signed trial log records 47 mL delivered at drought-arm event DA-17 and 101 mL delivered at control-arm event CT-23.", "right": "The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-008", "id": "fast-41-diverse-244-008-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"case_note": "The tomato drought trial's operative baseline-height timestamp was 11 May at 09:00, and its operative random-assignment timestamp was 11 May at 09:20. The trial schedule contains exactly one moisture-sensor reading gap, on 14 May; each scheduled moisture reading omitted because of that gap has a gravimetric backup. Every protocol-required final-height measurement is recorded. The signed trial log and scheduled-event register provide the relevant watering records, while the study file identifies drought-arm and control-arm events by arm, date, time, and protocol volume. The data-quality review checked the assignment record, watering entries, sensor backups, and final-height table. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit. Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?", "evidence": ["In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.", "The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23.", "The operative baseline-height timestamp is 11 May at 09:00, before the operative random-assignment timestamp of 11 May at 09:20.", "The schedule has exactly one moisture-sensor reading gap, on 14 May, and every scheduled reading omitted because of it has a gravimetric backup.", "Every protocol-required final-height measurement is recorded."]}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question and policy, preserve the tomato-trial request and bindings, contain coherent noncontradictory observations, and include the complete factual evidence quotes: \"In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.\" and \"The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"case_note\":\"The tomato drought trial's operative baseline-height timestamp was 11 May at 09:00, and its operative random-assignment timestamp was 11 May at 09:20. The trial schedule contains exactly one moisture-sensor reading gap, on 14 May; each scheduled moisture reading omitted because of that gap has a gravimetric backup. Every protocol-required final-height measurement is recorded. The signed trial log and scheduled-event register provide the relevant watering records, while the study file identifies drought-arm and control-arm events by arm, date, time, and protocol volume. The data-quality review checked the assignment record, watering entries, sensor backups, and final-height table. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit. Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\",\"evidence\":[\"In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.\",\"The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23.\",\"The operative baseline-height timestamp is 11 May at 09:00, before the operative random-assignment timestamp of 11 May at 09:20.\",\"The schedule has exactly one moisture-sensor reading gap, on 14 May, and every scheduled reading omitted because of it has a gravimetric backup.\",\"Every protocol-required final-height measurement is recorded.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL."}, {"path": ["evidence", "1"], "text": "The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.", "negative_left": "In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.", "negative_right": "The signed trial log records 47 mL delivered at drought-arm event DA-17 and 101 mL delivered at control-arm event CT-23.", "right": "The signed trial log records 47 mL delivered at drought-arm event DA-17 and 97 mL delivered at control-arm event CT-23."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-008", "id": "fast-41-diverse-244-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"case_note": "The tomato drought trial's operative baseline-height timestamp was 11 May at 09:00, and its operative random-assignment timestamp was 11 May at 09:20. The trial schedule contains exactly one moisture-sensor reading gap, on 14 May; each scheduled moisture reading omitted because of that gap has a gravimetric backup. Every protocol-required final-height measurement is recorded. The signed trial log and scheduled-event register provide the relevant watering records, while the study file identifies drought-arm and control-arm events by arm, date, time, and protocol volume. The data-quality review checked the assignment record, watering entries, sensor backups, and final-height table. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit. Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?", "evidence": ["In the fictional tomato drought trial, exactly two watering events occurred on 6 June 2026: drought-arm event DA-17 at 08:10 had a protocol-prescribed volume of 40 mL, and control-arm event CT-23 at 08:25 had a protocol-prescribed volume of 90 mL.", "The signed trial log records 47 mL delivered at drought-arm event DA-17 and 101 mL delivered at control-arm event CT-23.", "The operative baseline-height timestamp is 11 May at 09:00, before the operative random-assignment timestamp of 11 May at 09:20.", "The schedule has exactly one moisture-sensor reading gap, on 14 May, and every scheduled reading omitted because of it has a gravimetric backup.", "Every protocol-required final-height measurement is recorded."]}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and request bindings, and the evidence consists of the complete factual sentences “The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18.” and “Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial.”; the counterfactual coherently changes only W18 to 152 mL without embedding an answer or classifier guidance.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"The tomato drought trial's signed operations file records baseline height at 11 May 09:00 and random assignment at 11 May 09:20. The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18. Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial. The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and every scheduled moisture reading missing because of that gap has a gravimetric backup. Every protocol-required final-height measurement is recorded. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["context"], "text": "The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18."}, {"path": ["context"], "text": "Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18.", "negative_left": "The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18.", "negative_right": "Calibrated records show that event W17 received 127 mL and event W18 received 152 mL in the fictional tomato drought trial.", "right": "Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-010", "id": "fast-41-diverse-244-010-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "The tomato drought trial's signed operations file records baseline height at 11 May 09:00 and random assignment at 11 May 09:20. The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18. Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial. The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and every scheduled moisture reading missing because of that gap has a gravimetric backup. Every protocol-required final-height measurement is recorded. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the policy and request bindings, and the evidence consists of the complete factual sentences “The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18.” and “Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial.”; the counterfactual coherently changes only W18 to 152 mL without embedding an answer or classifier guidance.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"context\":\"The tomato drought trial's signed operations file records baseline height at 11 May 09:00 and random assignment at 11 May 09:20. The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18. Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial. The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and every scheduled moisture reading missing because of that gap has a gravimetric backup. Every protocol-required final-height measurement is recorded. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["context"], "text": "The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18."}, {"path": ["context"], "text": "Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18.", "negative_left": "The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18.", "negative_right": "Calibrated records show that event W17 received 127 mL and event W18 received 152 mL in the fictional tomato drought trial.", "right": "Calibrated records show that event W17 received 127 mL and event W18 received 133 mL in the fictional tomato drought trial."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-010", "id": "fast-41-diverse-244-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"context": "The tomato drought trial's signed operations file records baseline height at 11 May 09:00 and random assignment at 11 May 09:20. The fictional tomato drought trial's watering log identifies every drought-arm and control-arm watering event as event W17 or event W18, with protocol-prescribed volumes of 120 mL for W17 and 140 mL for W18. Calibrated records show that event W17 received 127 mL and event W18 received 152 mL in the fictional tomato drought trial. The trial has exactly one scheduled moisture-sensor reading gap, on 14 May, and every scheduled moisture reading missing because of that gap has a gravimetric backup. Every protocol-required final-height measurement is recorded. Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim, preserve the policy and bindings, use two complete factual evidence sentences, and contain no answer or labeling instructions; the counterfactual coherently changes only D-17’s delivered volume from 48 mL to 55 mL without duplicate or contradictory assertions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"narrative\":\"The tomato drought trial's signed study log and controller audit identify the operative baseline-height timestamp as 11 May at 09:00 and the operative random-assignment timestamp as 11 May at 09:20. The trial schedule contains exactly one moisture-sensor reading gap, on 14 May. The records show that each scheduled moisture reading omitted during that gap has a gravimetric backup. The growth table contains every protocol-required final-height measurement. The audited event register identifies the complete watering-event set below, while the delivery log records the corresponding volumes.\",\"policy\":\"Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL).\",\"The audited logs record delivered volumes of 48 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL)."}, {"path": ["evidence", "1"], "text": "The audited logs record delivered volumes of 48 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL).", "negative_left": "In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL).", "negative_right": "The audited logs record delivered volumes of 55 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial.", "right": "The audited logs record delivered volumes of 48 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-011", "id": "fast-41-diverse-244-011-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"evidence": ["In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL).", "The audited logs record delivered volumes of 48 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial."], "narrative": "The tomato drought trial's signed study log and controller audit identify the operative baseline-height timestamp as 11 May at 09:00 and the operative random-assignment timestamp as 11 May at 09:20. The trial schedule contains exactly one moisture-sensor reading gap, on 14 May. The records show that each scheduled moisture reading omitted during that gap has a gravimetric backup. The growth table contains every protocol-required final-height measurement. The audited event register identifies the complete watering-event set below, while the delivery log records the corresponding volumes.", "policy": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim, preserve the policy and bindings, use two complete factual evidence sentences, and contain no answer or labeling instructions; the counterfactual coherently changes only D-17’s delivered volume from 48 mL to 55 mL without duplicate or contradictory assertions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the universally quantified coverage propositions. The focus atom concerns actual watering compliance rather than a policy classification. The base and counter assignments can differ only in watering compliance while all timing and measurement facts remain fixed. Policy evidence correctly cites the original state and preserves the substantive protocol and temporal-update rule; no quotations from the automatically retained questions object are required. The extra request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes operative baseline-before-assignment timing, compliant watering for every event, no more than the one identified sensor gap with complete gravimetric backup, and complete required final-height recording. These conditions are sufficient for readiness under the protocol and unchanged decision criteria.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a2 entails that at least one watering event exceeds the ±10 mL tolerance. That is an explicit readiness violation and is sufficient for the false outcome, irrespective of the other satisfied conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For the fictional tomato drought trial, the operative baseline-height timestamp of 11 May at 09:00 precedes the operative random-assignment timestamp of 11 May at 09:20."}, {"id": "a2", "statement": "For every drought-arm and control-arm watering event in the fictional tomato drought trial, the delivered volume differs from that event's protocol-prescribed volume by at most 10 mL."}, {"id": "a3", "statement": "The fictional tomato drought trial has exactly one scheduled moisture-sensor reading gap, on 14 May."}, {"id": "a4", "statement": "Every scheduled moisture reading missing because of the fictional tomato drought trial's 14 May sensor gap has a gravimetric backup."}, {"id": "a5", "statement": "Every protocol-required final-height measurement for the fictional tomato drought trial is recorded."}], "base_state_json": "{\"narrative\":\"The tomato drought trial's signed study log and controller audit identify the operative baseline-height timestamp as 11 May at 09:00 and the operative random-assignment timestamp as 11 May at 09:20. The trial schedule contains exactly one moisture-sensor reading gap, on 14 May. The records show that each scheduled moisture reading omitted during that gap has a gravimetric backup. The growth table contains every protocol-required final-height measurement. The audited event register identifies the complete watering-event set below, while the delivery log records the corresponding volumes.\",\"policy\":\"Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.\",\"evidence\":[\"In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL).\",\"The audited logs record delivered volumes of 48 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial.\"],\"request\":\"Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL)."}, {"path": ["evidence", "1"], "text": "The audited logs record delivered volumes of 48 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial."}], "policy_evidence": [{"path": ["context"], "text": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit."}, {"path": ["request"], "text": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}], "rules": [{"justification": "The operative baseline precedes the operative assignment, every watering event is within the ±10 mL limit, the sole sensor gap has a gravimetric backup for every affected scheduled reading, and all required final heights are recorded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "Explicit refutation of the universal watering-limit proposition entails that at least one drought-arm or control-arm watering event differs from its protocol-prescribed volume by more than 10 mL. That violation prevents readiness even though assignment timing and measurement coverage are satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}]}, "verified_pair": {"left": "In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL).", "negative_left": "In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL).", "negative_right": "The audited logs record delivered volumes of 55 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial.", "right": "The audited logs record delivered volumes of 48 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial."}, "verifier_independent_model": false}, "family": "fast-41-diverse-244-011", "id": "fast-41-diverse-244-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the trial as ready because at least one assignment, watering, or measurement-coverage condition is violated or unsupported.", "true": "Yes — route the trial as ready for treatment-effect analysis because all stated validity and data-coverage conditions are satisfied."}, "instructions": "Answer yes if the evidence establishes compliant assignment timing, watering, and required measurement coverage after applying the temporal-update rule. Answer no if any readiness condition remains unsupported or violated.", "type": "noul"}}, "state": {"evidence": ["In the fictional tomato drought trial, the complete set of drought-arm and control-arm watering events is D-17 (drought arm, prescribed 40 mL) and C-22 (control arm, prescribed 60 mL).", "The audited logs record delivered volumes of 55 mL for D-17 and 53 mL for C-22 in the fictional tomato drought trial."], "narrative": "The tomato drought trial's signed study log and controller audit identify the operative baseline-height timestamp as 11 May at 09:00 and the operative random-assignment timestamp as 11 May at 09:20. The trial schedule contains exactly one moisture-sensor reading gap, on 14 May. The records show that each scheduled moisture reading omitted during that gap has a gravimetric backup. The growth table contains every protocol-required final-height measurement. The audited event register identifies the complete watering-event set below, while the delivery log records the corresponding volumes.", "policy": "Protocol: baseline height must precede random assignment; watering must stay within ±10 mL; one sensor gap is acceptable if every scheduled reading has a gravimetric backup. A later signed correction supersedes an earlier unsigned entry when supported by an instrument audit.", "request": "Under the stated protocol and update rule, is the trial ready for treatment-effect analysis?"}}, "method": "c2d", "provenance": {"source_id": "diverse-244", "source_is_synthetic": true, "source_sha256": "61e1ba5f04ca9a6d981217d6835bc5f0304075e79a0f3b58b13a8985829b2a77", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is unchanged and neither context alters governing rules. Question bindings remain DR-18, treatment-control validity, and the submitted analysis. Evidence consists of exactly two complete factual sentences: “The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment.” “The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment.” The counterfactual is coherent because only the RFID assignment for P03 changes, creating a documented disagreement between the two designated primary records. No context states a score, answer code, rationale, proposition ID, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"greenhouse technician\",\"text\":\"The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment.\"},{\"speaker\":\"data quality reviewer\",\"text\":\"The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment.\"},{\"speaker\":\"records custodian\",\"text\":\"The signed bench map and the timestamped pre-treatment RFID assignment export are each designated primary assignment records. No DR-18 record besides those two is designated a primary assignment record.\"},{\"speaker\":\"submission coordinator\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted.\"},{\"speaker\":\"protocol manager\",\"text\":\"A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any included pot ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment."}, {"path": ["1", "text"], "text": "The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment.", "negative_left": "The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment.", "negative_right": "The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as control.", "right": "The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-001", "id": "fast-41-diverse-245-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "greenhouse technician", "text": "The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment."}, {"speaker": "data quality reviewer", "text": "The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment."}, {"speaker": "records custodian", "text": "The signed bench map and the timestamped pre-treatment RFID assignment export are each designated primary assignment records. No DR-18 record besides those two is designated a primary assignment record."}, {"speaker": "submission coordinator", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"speaker": "protocol manager", "text": "A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any included pot ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is unchanged and neither context alters governing rules. Question bindings remain DR-18, treatment-control validity, and the submitted analysis. Evidence consists of exactly two complete factual sentences: “The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment.” “The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment.” The counterfactual is coherent because only the RFID assignment for P03 changes, creating a documented disagreement between the two designated primary records. No context states a score, answer code, rationale, proposition ID, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"greenhouse technician\",\"text\":\"The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment.\"},{\"speaker\":\"data quality reviewer\",\"text\":\"The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment.\"},{\"speaker\":\"records custodian\",\"text\":\"The signed bench map and the timestamped pre-treatment RFID assignment export are each designated primary assignment records. No DR-18 record besides those two is designated a primary assignment record.\"},{\"speaker\":\"submission coordinator\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted.\"},{\"speaker\":\"protocol manager\",\"text\":\"A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any included pot ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment."}, {"path": ["1", "text"], "text": "The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment.", "negative_left": "The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment.", "negative_right": "The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as control.", "right": "The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as treatment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-001", "id": "fast-41-diverse-245-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "greenhouse technician", "text": "The final growth analysis of DR-18 includes exactly pots P01, P02, and P03, and the signed bench map finalized on 14 March 2026 records P01 as treatment, P02 as control, and P03 as treatment."}, {"speaker": "data quality reviewer", "text": "The timestamped pre-treatment RFID assignment export for DR-18, generated on 14 March 2026, records P01 as treatment, P02 as control, and P03 as control."}, {"speaker": "records custodian", "text": "The signed bench map and the timestamped pre-treatment RFID assignment export are each designated primary assignment records. No DR-18 record besides those two is designated a primary assignment record."}, {"speaker": "submission coordinator", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"speaker": "protocol manager", "text": "A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any included pot ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and neither context adds an exception or missing-evidence rule. DR-18, analyzed pots, assignment records, and submission timing remain properly bound. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the RFID assignments, creating an explicit unresolved conflict between distinct primary records. Neither context states a score, answer code, rationale, proposition identifier, or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"Data quality reviewer\",\"text\":\"The DR-18 submission packet identifies the signed bench map as a primary assignment record.\"},{\"speaker\":\"Data quality reviewer\",\"text\":\"For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”\"},{\"speaker\":\"Data quality reviewer\",\"text\":\"The same packet identifies the timestamped pre-treatment RFID assignment export for DR-18 as a primary assignment record.\"},{\"speaker\":\"Data quality reviewer\",\"text\":\"For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “treated.”\"},{\"speaker\":\"Records custodian\",\"text\":\"No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record.\"},{\"speaker\":\"Records custodian\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted.\"},{\"speaker\":\"Protocol auditor\",\"text\":\"A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18.\"},{\"speaker\":\"Protocol auditor\",\"text\":\"The deviation report states that no documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”"}, {"path": ["3", "text"], "text": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “treated.”"}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”", "negative_left": "For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”", "negative_right": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “control.”", "right": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “treated.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-002", "id": "fast-41-diverse-245-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "Data quality reviewer", "text": "The DR-18 submission packet identifies the signed bench map as a primary assignment record."}, {"speaker": "Data quality reviewer", "text": "For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”"}, {"speaker": "Data quality reviewer", "text": "The same packet identifies the timestamped pre-treatment RFID assignment export for DR-18 as a primary assignment record."}, {"speaker": "Data quality reviewer", "text": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “treated.”"}, {"speaker": "Records custodian", "text": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"speaker": "Records custodian", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"speaker": "Protocol auditor", "text": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"speaker": "Protocol auditor", "text": "The deviation report states that no documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and neither context adds an exception or missing-evidence rule. DR-18, analyzed pots, assignment records, and submission timing remain properly bound. The two evidence spans are complete factual sentences. The counterfactual coherently changes only the RFID assignments, creating an explicit unresolved conflict between distinct primary records. Neither context states a score, answer code, rationale, proposition identifier, or classification instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"Data quality reviewer\",\"text\":\"The DR-18 submission packet identifies the signed bench map as a primary assignment record.\"},{\"speaker\":\"Data quality reviewer\",\"text\":\"For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”\"},{\"speaker\":\"Data quality reviewer\",\"text\":\"The same packet identifies the timestamped pre-treatment RFID assignment export for DR-18 as a primary assignment record.\"},{\"speaker\":\"Data quality reviewer\",\"text\":\"For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “treated.”\"},{\"speaker\":\"Records custodian\",\"text\":\"No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record.\"},{\"speaker\":\"Records custodian\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted.\"},{\"speaker\":\"Protocol auditor\",\"text\":\"A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18.\"},{\"speaker\":\"Protocol auditor\",\"text\":\"The deviation report states that no documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”"}, {"path": ["3", "text"], "text": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “treated.”"}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”", "negative_left": "For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”", "negative_right": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “control.”", "right": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “treated.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-002", "id": "fast-41-diverse-245-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "Data quality reviewer", "text": "The DR-18 submission packet identifies the signed bench map as a primary assignment record."}, {"speaker": "Data quality reviewer", "text": "For every pot included in the final growth analysis of DR-18, the signed bench map records treatment assignment “treated.”"}, {"speaker": "Data quality reviewer", "text": "The same packet identifies the timestamped pre-treatment RFID assignment export for DR-18 as a primary assignment record."}, {"speaker": "Data quality reviewer", "text": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export records treatment assignment “control.”"}, {"speaker": "Records custodian", "text": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"speaker": "Records custodian", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"speaker": "Protocol auditor", "text": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"speaker": "Protocol auditor", "text": "The deviation report states that no documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is unchanged and no context adds or removes a governing rule. Question bindings remain DR-18, the final growth analysis, and the pre-treatment assignment export. The evidence spans are two complete factual sentences: \"The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23.\" and \"The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23.\" The counterfactual coherently records conflicting assignments from two separately identified primary records. Neither context embeds an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"case note\",\"text\":\"The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23.\"},{\"speaker\":\"case note\",\"text\":\"The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23.\"},{\"speaker\":\"records auditor\",\"text\":\"The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record.\"},{\"speaker\":\"records auditor\",\"text\":\"No DR-18 record other than those two records is designated a primary assignment record, and no reconciliation between them was recorded before submission.\"},{\"speaker\":\"protocol reviewer\",\"text\":\"A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any analyzed pot ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23."}, {"path": ["1", "text"], "text": "The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23.", "negative_left": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23.", "negative_right": "The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Z to P-23.", "right": "The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-003", "id": "fast-41-diverse-245-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "case note", "text": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23."}, {"speaker": "case note", "text": "The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23."}, {"speaker": "records auditor", "text": "The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record."}, {"speaker": "records auditor", "text": "No DR-18 record other than those two records is designated a primary assignment record, and no reconciliation between them was recorded before submission."}, {"speaker": "protocol reviewer", "text": "A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any analyzed pot ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is unchanged and no context adds or removes a governing rule. Question bindings remain DR-18, the final growth analysis, and the pre-treatment assignment export. The evidence spans are two complete factual sentences: \"The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23.\" and \"The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23.\" The counterfactual coherently records conflicting assignments from two separately identified primary records. Neither context embeds an answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"case note\",\"text\":\"The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23.\"},{\"speaker\":\"case note\",\"text\":\"The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23.\"},{\"speaker\":\"records auditor\",\"text\":\"The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record.\"},{\"speaker\":\"records auditor\",\"text\":\"No DR-18 record other than those two records is designated a primary assignment record, and no reconciliation between them was recorded before submission.\"},{\"speaker\":\"protocol reviewer\",\"text\":\"A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any analyzed pot ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23."}, {"path": ["1", "text"], "text": "The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23.", "negative_left": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23.", "negative_right": "The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Z to P-23.", "right": "The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Y to P-23."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-003", "id": "fast-41-diverse-245-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "case note", "text": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map assigns treatment X to P-17 and treatment Y to P-23."}, {"speaker": "case note", "text": "The timestamped pre-treatment RFID assignment export for DR-18 assigns treatment X to P-17 and treatment Z to P-23."}, {"speaker": "records auditor", "text": "The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record."}, {"speaker": "records auditor", "text": "No DR-18 record other than those two records is designated a primary assignment record, and no reconciliation between them was recorded before submission."}, {"speaker": "protocol reviewer", "text": "A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any analyzed pot ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, and neither context changes governing policy. DR-18, submitted validity, analyzed pots, and assignment-record bindings remain preserved. The evidence consists of two complete factual sentences: \"The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23.\" and \"The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23.\" The counterfactual is coherent because the distinct primary records disagree and the context explicitly records that no reconciliation occurred. Neither context states a gold answer, label, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"study coordinator\",\"text\":\"The final DR-18 growth analysis contains two pots, with watering and measurement logs filed for both.\"},{\"speaker\":\"assignment auditor\",\"text\":\"The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23.\"},{\"speaker\":\"records custodian\",\"text\":\"The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record. No other DR-18 record has that designation.\"},{\"speaker\":\"assignment auditor\",\"text\":\"The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23.\"},{\"speaker\":\"records custodian\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted.\"},{\"speaker\":\"study coordinator\",\"text\":\"A documented protocol deviation affects P-23. The deviation does not make the assignment identity of any analyzed pot ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23."}, {"path": ["3", "text"], "text": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23.", "negative_left": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23.", "negative_right": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment C to P-17 and Treatment B to P-23.", "right": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-007", "id": "fast-41-diverse-245-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "study coordinator", "text": "The final DR-18 growth analysis contains two pots, with watering and measurement logs filed for both."}, {"speaker": "assignment auditor", "text": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23."}, {"speaker": "records custodian", "text": "The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record. No other DR-18 record has that designation."}, {"speaker": "assignment auditor", "text": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23."}, {"speaker": "records custodian", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"speaker": "study coordinator", "text": "A documented protocol deviation affects P-23. The deviation does not make the assignment identity of any analyzed pot ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, and neither context changes governing policy. DR-18, submitted validity, analyzed pots, and assignment-record bindings remain preserved. The evidence consists of two complete factual sentences: \"The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23.\" and \"The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23.\" The counterfactual is coherent because the distinct primary records disagree and the context explicitly records that no reconciliation occurred. Neither context states a gold answer, label, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"study coordinator\",\"text\":\"The final DR-18 growth analysis contains two pots, with watering and measurement logs filed for both.\"},{\"speaker\":\"assignment auditor\",\"text\":\"The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23.\"},{\"speaker\":\"records custodian\",\"text\":\"The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record. No other DR-18 record has that designation.\"},{\"speaker\":\"assignment auditor\",\"text\":\"The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23.\"},{\"speaker\":\"records custodian\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted.\"},{\"speaker\":\"study coordinator\",\"text\":\"A documented protocol deviation affects P-23. The deviation does not make the assignment identity of any analyzed pot ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23."}, {"path": ["3", "text"], "text": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23.", "negative_left": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23.", "negative_right": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment C to P-17 and Treatment B to P-23.", "right": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment A to P-17 and Treatment B to P-23."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-007", "id": "fast-41-diverse-245-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "study coordinator", "text": "The final DR-18 growth analysis contains two pots, with watering and measurement logs filed for both."}, {"speaker": "assignment auditor", "text": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the signed bench map assigns Treatment A to P-17 and Treatment B to P-23."}, {"speaker": "records custodian", "text": "The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record. No other DR-18 record has that designation."}, {"speaker": "assignment auditor", "text": "The pots included in the final growth analysis of DR-18 are P-17 and P-23, and the timestamped pre-treatment RFID assignment export assigns Treatment C to P-17 and Treatment B to P-23."}, {"speaker": "records custodian", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"speaker": "study coordinator", "text": "A documented protocol deviation affects P-23. The deviation does not make the assignment identity of any analyzed pot ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged DR-18 scope and governing policy. The two evidence spans are complete factual sentences, and the counterfactual coherently changes only the RFID assignment while preserving the distinct-record conflict. Neither context includes an answer, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"data quality reviewer\",\"text\":\"DR-18's final growth analysis includes the designated set of pots, and the submission retained the affected pot. The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record. No other DR-18 record is designated a primary assignment record.\"},{\"speaker\":\"records coordinator\",\"text\":\"For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha.\"},{\"speaker\":\"records coordinator\",\"text\":\"For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Alpha.\"},{\"speaker\":\"protocol reviewer\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted. A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any included pot ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha."}, {"path": ["2", "text"], "text": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Alpha."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha.", "negative_left": "For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha.", "negative_right": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Control.", "right": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Alpha."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-008", "id": "fast-41-diverse-245-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "data quality reviewer", "text": "DR-18's final growth analysis includes the designated set of pots, and the submission retained the affected pot. The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record. No other DR-18 record is designated a primary assignment record."}, {"speaker": "records coordinator", "text": "For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha."}, {"speaker": "records coordinator", "text": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Alpha."}, {"speaker": "protocol reviewer", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted. A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any included pot ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged DR-18 scope and governing policy. The two evidence spans are complete factual sentences, and the counterfactual coherently changes only the RFID assignment while preserving the distinct-record conflict. Neither context includes an answer, rule table, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"data quality reviewer\",\"text\":\"DR-18's final growth analysis includes the designated set of pots, and the submission retained the affected pot. The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record. No other DR-18 record is designated a primary assignment record.\"},{\"speaker\":\"records coordinator\",\"text\":\"For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha.\"},{\"speaker\":\"records coordinator\",\"text\":\"For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Alpha.\"},{\"speaker\":\"protocol reviewer\",\"text\":\"No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted. A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any included pot ambiguous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha."}, {"path": ["2", "text"], "text": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Alpha."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha.", "negative_left": "For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha.", "negative_right": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Control.", "right": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Alpha."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-008", "id": "fast-41-diverse-245-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "data quality reviewer", "text": "DR-18's final growth analysis includes the designated set of pots, and the submission retained the affected pot. The signed bench map is designated a primary assignment record, and the timestamped pre-treatment RFID assignment export is also designated a primary assignment record. No other DR-18 record is designated a primary assignment record."}, {"speaker": "records coordinator", "text": "For every pot included in the final growth analysis of DR-18, the signed bench map assigns the pot to treatment group Alpha."}, {"speaker": "records coordinator", "text": "For every pot included in the final growth analysis of DR-18, the timestamped pre-treatment RFID assignment export assigns the pot to treatment group Control."}, {"speaker": "protocol reviewer", "text": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted. A documented protocol deviation affects at least one pot included in the final growth analysis, but no documented protocol deviation makes the assignment identity of any included pot ambiguous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged DR-18 question and governing policy. The bindings remain DR-18, analyzed pots, assignment records, and treatment-control validity. The evidence spans are complete factual sentences: \"The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23.\" \"The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23.\" The counterfactual coherently introduces an unreconciled conflict between two designated primary records without duplicating a measurement inconsistently. Neither context states a score, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"case note\",\"text\":\"The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23. The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23.\"},{\"speaker\":\"data quality reviewer\",\"text\":\"The signed bench map for DR-18 and the timestamped pre-treatment RFID assignment export for DR-18 are each designated primary assignment records. No DR-18 record other than those two documents is designated a primary assignment record, and no reconciliation between them was recorded before submission. A documented protocol deviation affects P-23, which remains in the final growth analysis. The deviation does not make the assignment identity of any analyzed pot ambiguous. Both pots remain available for treatment-control review, with no exclusion recorded.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23."}, {"path": ["0", "text"], "text": "The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23.", "negative_left": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23.", "negative_right": "The timestamped pre-treatment RFID assignment export for DR-18 records assignment South for P-17 and assignment South for P-23.", "right": "The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-010", "id": "fast-41-diverse-245-010-base", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "case note", "text": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23. The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23."}, {"speaker": "data quality reviewer", "text": "The signed bench map for DR-18 and the timestamped pre-treatment RFID assignment export for DR-18 are each designated primary assignment records. No DR-18 record other than those two documents is designated a primary assignment record, and no reconciliation between them was recorded before submission. A documented protocol deviation affects P-23, which remains in the final growth analysis. The deviation does not make the assignment identity of any analyzed pot ambiguous. Both pots remain available for treatment-control review, with no exclusion recorded."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged DR-18 question and governing policy. The bindings remain DR-18, analyzed pots, assignment records, and treatment-control validity. The evidence spans are complete factual sentences: \"The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23.\" \"The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23.\" The counterfactual coherently introduces an unreconciled conflict between two designated primary records without duplicating a measurement inconsistently. Neither context states a score, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the quantified atoms remain atomic. The focus a1 is a factual agreement relation, not a policy classification. The base and counter assignments are jointly realizable while changing only whether the two primary records agree: lack of a recorded reconciliation can coexist with either agreement or disagreement, and the separately documented deviation can remain non-ambiguating. Empty policy_evidence is correct because the governing scoring criteria, ordering, independence requirement, and conflict exception all originate in the retained questions object; the state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted a1 entails disagreement for at least one analyzed pot. With both sources designated primary and no recorded reconciliation before submission, this is an unresolved conflict between primary assignment records, which is sufficient for level 0 regardless of sensor or watering evidence.", "rule_index": 0, "sound": true}, {"reason": "Supported a1 together with a2–a4 establishes that all designated primary assignment records exist and agree for every analyzed pot. A supported a6 supplies a documented protocol deviation affecting an analyzed pot, while a7 excludes assignment ambiguity from such deviations. This is sufficient for level 1 and excludes level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "For every pot included in the final growth analysis of DR-18, the treatment assignment specified by the signed bench map is identical to the treatment assignment specified by the timestamped pre-treatment RFID assignment export."}, {"id": "a2", "statement": "The signed bench map for DR-18 is designated a primary assignment record."}, {"id": "a3", "statement": "The timestamped pre-treatment RFID assignment export for DR-18 is designated a primary assignment record."}, {"id": "a4", "statement": "No DR-18 record other than the signed bench map and the timestamped pre-treatment RFID assignment export is designated a primary assignment record."}, {"id": "a5", "statement": "No reconciliation between the signed bench map and the timestamped pre-treatment RFID assignment export was recorded before DR-18 was submitted."}, {"id": "a6", "statement": "A documented protocol deviation affects at least one pot included in the final growth analysis of DR-18."}, {"id": "a7", "statement": "No documented protocol deviation makes the assignment identity of any pot included in the final growth analysis of DR-18 ambiguous."}], "base_state_json": "[{\"speaker\":\"case note\",\"text\":\"The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23. The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23.\"},{\"speaker\":\"data quality reviewer\",\"text\":\"The signed bench map for DR-18 and the timestamped pre-treatment RFID assignment export for DR-18 are each designated primary assignment records. No DR-18 record other than those two documents is designated a primary assignment record, and no reconciliation between them was recorded before submission. A documented protocol deviation affects P-23, which remains in the final growth analysis. The deviation does not make the assignment identity of any analyzed pot ambiguous. Both pots remain available for treatment-control review, with no exclusion recorded.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23."}, {"path": ["0", "text"], "text": "The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23."}], "policy_evidence": [], "rules": [{"justification": "Refuting a1 establishes that the signed bench map and pre-treatment RFID export specify conflicting assignments for at least one analyzed pot. Both are primary assignment records, and no reconciliation was recorded before submission, so the unresolved contradiction requires level 0.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "The only primary assignment records specify identical assignments for every analyzed pot, so primary assignments are consistent and assignment evidence exists. A documented protocol deviation affects an analyzed pot without making assignment identity ambiguous, which satisfies level 1 when the levels are applied from low to high.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23.", "negative_left": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23.", "negative_right": "The timestamped pre-treatment RFID assignment export for DR-18 records assignment South for P-17 and assignment South for P-23.", "right": "The timestamped pre-treatment RFID assignment export for DR-18 records assignment North for P-17 and assignment South for P-23."}, "verifier_independent_model": false}, "family": "fast-41-diverse-245-010", "id": "fast-41-diverse-245-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not validatable: At least one analyzed pot has unresolved conflicting primary assignment records, or no primary assignment evidence exists; treatment-control comparisons cannot be accepted.", "1 — Weak validity: Primary assignment records are consistent, but a documented protocol deviation or missing watering record affects at least one analyzed pot without making assignment identity ambiguous.", "2 — Adequate validity: Primary assignments are consistent and all analyzed pots have complete watering records; minor non-assignment metadata or growth-measurement issues are documented and bounded.", "3 — Strong validity: Primary assignments are consistent and independently reconciled, watering and sensor records are complete, all growth measurements pass quality checks, and no material protocol deviations remain."], "instructions": "Rate the treatment-control validity of DR-18 as submitted. Apply the levels in order from low to high. Assignment identity must be established independently of observed watering. An unresolved contradiction between primary assignment records for any analyzed pot requires level 0, even when sensor records support one record.", "type": "score"}}, "state": [{"speaker": "case note", "text": "The final growth analysis of DR-18 includes exactly pots P-17 and P-23, and the signed bench map records assignment North for P-17 and assignment South for P-23. The timestamped pre-treatment RFID assignment export for DR-18 records assignment South for P-17 and assignment South for P-23."}, {"speaker": "data quality reviewer", "text": "The signed bench map for DR-18 and the timestamped pre-treatment RFID assignment export for DR-18 are each designated primary assignment records. No DR-18 record other than those two documents is designated a primary assignment record, and no reconciliation between them was recorded before submission. A documented protocol deviation affects P-23, which remains in the final growth analysis. The deviation does not make the assignment identity of any analyzed pot ambiguous. Both pots remain available for treatment-control review, with no exclusion recorded."}]}, "method": "c2d", "provenance": {"source_id": "diverse-245", "source_is_synthetic": true, "source_sha256": "93978e856d62c3e44e6ad7ea4061295ff361dcffa68866a24405353572fc028f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The base context contradicts itself because it says every critical assignment record except the named sheet is present while stating, \"The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available.\" The evidence spans are complete factual sentences, including \"The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17.\" and \"The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"A completeness review concerns a 21-day greenhouse trial involving 24 tomato seedlings, with treatment identities and outcome measurements recorded. The package contains every critical assignment record other than the named sheet, plus every critical watering, sensor, and outcome record. The original bench map is absent; it is classified as a noncritical supporting detail for evaluating treatment-control validity, and all other supporting details for that evaluation are present. The assignment record and the treatment identities remain tied to the trial, while the outcome record covers the seedlings’ measured results. The missing-map condition does not alter the availability of the other critical records. If the completeness log instead marked the named sheet unavailable, these same non-focus records and classifications would remain unchanged. The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17. The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17."}, {"path": [], "text": "The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17.", "negative_left": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17.", "negative_right": "The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as unavailable.", "right": "The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available."}, "verifier_independent_model": false}, "family": "fast-41-diverse-246-002", "id": "fast-41-diverse-246-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "A completeness review concerns a 21-day greenhouse trial involving 24 tomato seedlings, with treatment identities and outcome measurements recorded. The package contains every critical assignment record other than the named sheet, plus every critical watering, sensor, and outcome record. The original bench map is absent; it is classified as a noncritical supporting detail for evaluating treatment-control validity, and all other supporting details for that evaluation are present. The assignment record and the treatment identities remain tied to the trial, while the outcome record covers the seedlings’ measured results. The missing-map condition does not alter the availability of the other critical records. If the completeness log instead marked the named sheet unavailable, these same non-focus records and classifications would remain unchanged. The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17. The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The base context contradicts itself because it says every critical assignment record except the named sheet is present while stating, \"The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available.\" The evidence spans are complete factual sentences, including \"The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17.\" and \"The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"A completeness review concerns a 21-day greenhouse trial involving 24 tomato seedlings, with treatment identities and outcome measurements recorded. The package contains every critical assignment record other than the named sheet, plus every critical watering, sensor, and outcome record. The original bench map is absent; it is classified as a noncritical supporting detail for evaluating treatment-control validity, and all other supporting details for that evaluation are present. The assignment record and the treatment identities remain tied to the trial, while the outcome record covers the seedlings’ measured results. The missing-map condition does not alter the availability of the other critical records. If the completeness log instead marked the named sheet unavailable, these same non-focus records and classifications would remain unchanged. The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17. The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17."}, {"path": [], "text": "The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17.", "negative_left": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17.", "negative_right": "The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as unavailable.", "right": "The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as available."}, "verifier_independent_model": false}, "family": "fast-41-diverse-246-002", "id": "fast-41-diverse-246-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "A completeness review concerns a 21-day greenhouse trial involving 24 tomato seedlings, with treatment identities and outcome measurements recorded. The package contains every critical assignment record other than the named sheet, plus every critical watering, sensor, and outcome record. The original bench map is absent; it is classified as a noncritical supporting detail for evaluating treatment-control validity, and all other supporting details for that evaluation are present. The assignment record and the treatment identities remain tied to the trial, while the outcome record covers the seedlings’ measured results. The missing-map condition does not alter the availability of the other critical records. If the completeness log instead marked the named sheet unavailable, these same non-focus records and classifications would remain unchanged. The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears the record code TA-17. The evidence-package completeness log dated 2026-09-17 marks critical assignment record TA-17 as unavailable."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric, while both contexts retain the 21-day trial, 24 seedlings, and treatment-control validity scope. The two evidence items are complete factual sentences. The counterfactual coherently changes only whether TA-24-21 appears in the inventory and does not contain answer labels, rationale, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"A greenhouse case note concerns a 21-day trial of 24 tomato seedlings, with drought and control groups identified and treatment identities and outcome measurements recorded. Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings. The evidence-package inventory dated 14 March 2026 lists document TA-24-21 among its contents. The treatment-assignment sheet is classified as a critical assignment record. Every other critical assignment record, along with every critical watering, sensor, and outcome record, is present. The records include planned watering schedules, timestamped soil-moisture readings, and baseline and final measurements. The original bench map is absent; it is classified as a noncritical supporting detail for evaluating treatment-control validity, while every other supporting detail for that evaluation is present. The review assesses whether treatment and control groups were validly assigned. In a counterfactual inventory review, the listed contents would differ, but the remaining record classifications and measurements would stay unchanged.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings."}, {"path": [], "text": "The evidence-package inventory dated 14 March 2026 lists document TA-24-21 among its contents."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings.", "negative_left": "Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings.", "negative_right": "The evidence-package inventory dated 14 March 2026 does not list document TA-24-21 among its contents.", "right": "The evidence-package inventory dated 14 March 2026 lists document TA-24-21 among its contents."}, "verifier_independent_model": false}, "family": "fast-41-diverse-246-007", "id": "fast-41-diverse-246-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "A greenhouse case note concerns a 21-day trial of 24 tomato seedlings, with drought and control groups identified and treatment identities and outcome measurements recorded. Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings. The evidence-package inventory dated 14 March 2026 lists document TA-24-21 among its contents. The treatment-assignment sheet is classified as a critical assignment record. Every other critical assignment record, along with every critical watering, sensor, and outcome record, is present. The records include planned watering schedules, timestamped soil-moisture readings, and baseline and final measurements. The original bench map is absent; it is classified as a noncritical supporting detail for evaluating treatment-control validity, while every other supporting detail for that evaluation is present. The review assesses whether treatment and control groups were validly assigned. In a counterfactual inventory review, the listed contents would differ, but the remaining record classifications and measurements would stay unchanged."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric, while both contexts retain the 21-day trial, 24 seedlings, and treatment-control validity scope. The two evidence items are complete factual sentences. The counterfactual coherently changes only whether TA-24-21 appears in the inventory and does not contain answer labels, rationale, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"A greenhouse case note concerns a 21-day trial of 24 tomato seedlings, with drought and control groups identified and treatment identities and outcome measurements recorded. Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings. The evidence-package inventory dated 14 March 2026 lists document TA-24-21 among its contents. The treatment-assignment sheet is classified as a critical assignment record. Every other critical assignment record, along with every critical watering, sensor, and outcome record, is present. The records include planned watering schedules, timestamped soil-moisture readings, and baseline and final measurements. The original bench map is absent; it is classified as a noncritical supporting detail for evaluating treatment-control validity, while every other supporting detail for that evaluation is present. The review assesses whether treatment and control groups were validly assigned. In a counterfactual inventory review, the listed contents would differ, but the remaining record classifications and measurements would stay unchanged.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings."}, {"path": [], "text": "The evidence-package inventory dated 14 March 2026 lists document TA-24-21 among its contents."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings.", "negative_left": "Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings.", "negative_right": "The evidence-package inventory dated 14 March 2026 does not list document TA-24-21 among its contents.", "right": "The evidence-package inventory dated 14 March 2026 lists document TA-24-21 among its contents."}, "verifier_independent_model": false}, "family": "fast-41-diverse-246-007", "id": "fast-41-diverse-246-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "A greenhouse case note concerns a 21-day trial of 24 tomato seedlings, with drought and control groups identified and treatment identities and outcome measurements recorded. Document TA-24-21 is titled “Treatment-assignment sheet” and belongs to the 21-day trial of 24 tomato seedlings. The evidence-package inventory dated 14 March 2026 does not list document TA-24-21 among its contents. The treatment-assignment sheet is classified as a critical assignment record. Every other critical assignment record, along with every critical watering, sensor, and outcome record, is present. The records include planned watering schedules, timestamped soil-moisture readings, and baseline and final measurements. The original bench map is absent; it is classified as a noncritical supporting detail for evaluating treatment-control validity, while every other supporting detail for that evaluation is present. The review assesses whether treatment and control groups were validly assigned. In a counterfactual inventory review, the listed contents would differ, but the remaining record classifications and measurements would stay unchanged."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, and neither context changes governing policy. Both contexts preserve the 21-day, 24-seedling drought-control assignment-validity scope. The two evidence spans are complete factual sentences: \"The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D.\" \"The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D.\" The counterfactual coherently states that the identified sheet is excluded from the inventory without creating contradictory duplicate measurements or assertions. Neither context contains a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D. The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D. The sheet is classified as a critical assignment record. Treatment identities and outcome measurements for all 24 tomato seedlings are recorded. Every other critical assignment record, all critical watering records, all critical sensor records, and all critical outcome records are present in the evidence package. The original bench map is absent and is classified as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail for that evaluation is present. The case concerns a 21-day trial of 24 tomato seedlings, with drought and control groups whose assignment validity is under review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D."}, {"path": [], "text": "The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D.", "negative_left": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D.", "negative_right": "The 2026-09-17 evidence-package inventory exhaustively lists its contents and excludes archive identifier TA-24-21D.", "right": "The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D."}, "verifier_independent_model": false}, "family": "fast-41-diverse-246-008", "id": "fast-41-diverse-246-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D. The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D. The sheet is classified as a critical assignment record. Treatment identities and outcome measurements for all 24 tomato seedlings are recorded. Every other critical assignment record, all critical watering records, all critical sensor records, and all critical outcome records are present in the evidence package. The original bench map is absent and is classified as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail for that evaluation is present. The case concerns a 21-day trial of 24 tomato seedlings, with drought and control groups whose assignment validity is under review."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, and neither context changes governing policy. Both contexts preserve the 21-day, 24-seedling drought-control assignment-validity scope. The two evidence spans are complete factual sentences: \"The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D.\" \"The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D.\" The counterfactual coherently states that the identified sheet is excluded from the inventory without creating contradictory duplicate measurements or assertions. Neither context contains a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, including the atoms classifying particular records as critical or noncritical. The focus atom concerns the factual availability of the treatment-assignment sheet rather than a policy conclusion. The base and counter assignments differ only on that availability and are jointly realizable: the sheet can be present or absent while its critical status and all other record facts remain unchanged. Empty policy_evidence is correct because the governing rubric and instructions are entirely in the retained questions object; no substantive state-originated policy needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that all critical assignment, watering, sensor, and outcome records are available, while exactly one supporting detail—the bench map—is unavailable and is explicitly noncritical. This is sufficient for score 2 and excludes score 3 because not all supporting documentation is available.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that treatment identities and outcomes are recorded but that the treatment-assignment sheet is unavailable and is a critical assignment record. This is sufficient for score 1 because a critical assignment record is missing, and the recorded identities and outcomes exclude score 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is available in the evidence package."}, {"id": "a2", "statement": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings is a critical assignment record."}, {"id": "a3", "statement": "The treatment identities of the 24 tomato seedlings are recorded."}, {"id": "a4", "statement": "The outcome measurements of the 24 tomato seedlings are recorded."}, {"id": "a5", "statement": "Every critical assignment record for the trial other than the treatment-assignment sheet is available."}, {"id": "a6", "statement": "Every critical watering record for the trial is available."}, {"id": "a7", "statement": "Every critical sensor record for the trial is available."}, {"id": "a8", "statement": "Every critical outcome record for the trial is available."}, {"id": "a9", "statement": "The original bench map for the trial is available."}, {"id": "a10", "statement": "The original bench map is a noncritical supporting detail for evaluating treatment-control validity."}, {"id": "a11", "statement": "Every supporting detail for evaluating treatment-control validity other than the original bench map is available."}], "base_state_json": "\"The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D. The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D. The sheet is classified as a critical assignment record. Treatment identities and outcome measurements for all 24 tomato seedlings are recorded. Every other critical assignment record, all critical watering records, all critical sensor records, and all critical outcome records are present in the evidence package. The original bench map is absent and is classified as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail for that evaluation is present. The case concerns a 21-day trial of 24 tomato seedlings, with drought and control groups whose assignment validity is under review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D."}, {"path": [], "text": "The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D."}], "policy_evidence": [], "rules": [{"justification": "All critical assignment, watering, sensor, and outcome records are available, while the original bench map is the sole missing item and is noncritical supporting documentation.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Treatment identities and outcomes are recorded, but the unavailable treatment-assignment sheet is a critical assignment record, preventing verification of treatment-control validity.", "target": "1", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D.", "negative_left": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D.", "negative_right": "The 2026-09-17 evidence-package inventory exhaustively lists its contents and excludes archive identifier TA-24-21D.", "right": "The 2026-09-17 evidence-package inventory exhaustively lists its contents and includes archive identifier TA-24-21D."}, "verifier_independent_model": false}, "family": "fast-41-diverse-246-008", "id": "fast-41-diverse-246-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unusable: Treatment identities or outcome measurements are missing, so no treatment-control comparison can be evaluated.", "1 — Limited: Treatment identities and outcomes are recorded, but at least one critical assignment or protocol record is missing, preventing verification of treatment-control validity.", "2 — Adequate: All critical assignment, watering, sensor, and outcome records are available; only a noncritical supporting detail is missing.", "3 — Complete: All critical records and supporting documentation are available, internally consistent, timestamped, and sufficient to reconstruct assignment and protocol execution."], "instructions": "Rate the completeness of the available evidence for evaluating whether the drought and control groups were validly assigned. Use the ordered rubric below and select one level.", "type": "score"}}, "state": "The treatment-assignment sheet for the 21-day trial of 24 tomato seedlings bears archive identifier TA-24-21D. The 2026-09-17 evidence-package inventory exhaustively lists its contents and excludes archive identifier TA-24-21D. The sheet is classified as a critical assignment record. Treatment identities and outcome measurements for all 24 tomato seedlings are recorded. Every other critical assignment record, all critical watering records, all critical sensor records, and all critical outcome records are present in the evidence package. The original bench map is absent and is classified as a noncritical supporting detail for evaluating treatment-control validity. Every other supporting detail for that evaluation is present. The case concerns a 21-day trial of 24 tomato seedlings, with drought and control groups whose assignment validity is under review."}, "method": "c2d", "provenance": {"source_id": "diverse-246", "source_is_synthetic": true, "source_sha256": "454cf9a2ec2bc8f3274850e7b6af4824d24515f36ad7154a906a896e4d1ff9d5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, and neither context adds an exception or altered rule. Survey D247, site ST-08, the 30-minute effort, specimen lots, and routing-relevant bindings remain consistent. The evidence contains two complete factual sentences: \"After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026.\" \"The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level.\" The counterfactual coherently changes only the later entry from resolved to unresolved without creating a contradictory duplicate assertion. Neither context embeds an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universal claims such as every lot having a count remain atomic. The focus atom concerns a factual post-precedence inconsistency, not a policy conclusion. The base and counter assignments are realizable with only a10 changing: an explicit unresolved status can coexist either with an additional post-precedence conflicting status or with no such conflict, while other fields remain consistent. Empty policy_evidence is correct because the governing completeness, routing, precedence, and confidence rules are already retained in the questions object; the state contributes case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail completeness: effort, exact site-label match, numeric counts for every lot, and all four habitat observations are supported. The post-precedence inconsistency in voucher V4's species-resolution status is a remaining contradiction affecting whether Identification routing applies, so Data Review takes priority. Because that needed routing fact conflicts, Low confidence follows.", "rule_index": 0, "sound": true}, {"reason": "The conditions entail completeness. Refutation of a10 entails no post-precedence inconsistency in V4's species-resolution status, while a11 excludes contradictions elsewhere. The explicitly unresolved voucher therefore requires Identification. All routing facts are explicit, and the consistency conditions exclude conflicts, so High confidence follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Survey D247's recorded effort is at least 30 minutes."}, {"id": "a2", "statement": "Survey D247's field-sheet site label exactly matches Survey D247's registered site label."}, {"id": "a3", "statement": "Every specimen lot collected in Survey D247 has a numeric count."}, {"id": "a4", "statement": "Survey D247 has a recorded flow observation."}, {"id": "a5", "statement": "Survey D247 has a recorded substrate observation."}, {"id": "a6", "statement": "Survey D247 has a recorded canopy-shade observation."}, {"id": "a7", "statement": "Survey D247 has a recorded wetted-width observation."}, {"id": "a8", "statement": "Voucher V4 in Survey D247 is explicitly unresolved at species level."}, {"id": "a9", "statement": "Every fact needed to route Survey D247 is explicitly recorded rather than inferred."}, {"id": "a10", "statement": "After identification-correction precedence is applied, Survey D247's complete record assigns inconsistent species-resolution statuses to voucher V4."}, {"id": "a11", "statement": "After identification-correction precedence is applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent."}], "base_state_json": "[{\"speaker\":\"field survey lead\",\"text\":\"Survey D247 was conducted for exactly 30 minutes at registered site ST-08, and its field-sheet label was ST-08. Three specimen lots were collected.\"},{\"speaker\":\"ecology data manager\",\"text\":\"The stonefly, caddisfly, and mayfly lots have numeric counts of 12, 7, and 6. Flow, substrate, canopy shade, and wetted width are all recorded.\"},{\"speaker\":\"taxonomic identifier\",\"text\":\"Voucher V4 is explicitly unresolved at species level.\"},{\"speaker\":\"records coordinator\",\"text\":\"After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026.\"},{\"speaker\":\"records coordinator\",\"text\":\"The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level.\"},{\"speaker\":\"audit note\",\"text\":\"Every fact needed to route Survey D247 is explicitly recorded rather than inferred. After identification-correction precedence was applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["3", "text"], "text": "After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026."}, {"path": ["4", "text"], "text": "The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level."}], "policy_evidence": [], "rules": [{"justification": "All completeness requirements are satisfied. After correction precedence, voucher V4 has conflicting species-resolution statuses, so a contradiction remains in a fact needed to decide whether the survey routes to Identification; the required outcome is therefore Data Review with Low confidence.", "target": "data_review_yes_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All completeness requirements are satisfied. Voucher V4 is explicitly unresolved, its post-precedence species-resolution status is not inconsistent, every other record field is consistent, and every routing fact is explicit; the required outcome is therefore Identification with High confidence.", "target": "identification_yes_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026.", "negative_left": "After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026.", "negative_right": "The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as unresolved at species level.", "right": "The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level."}, "verifier_independent_model": false}, "family": "fast-41-diverse-247-009", "id": "fast-41-diverse-247-009-base", "input": {"questions": {"decision": {"criteria": {"data_review_yes_low": "Completeness Yes; Data-review queue; Low confidence. Choose only if a record contradiction remains after correction precedence is applied and a needed routing fact conflicts.", "identification_yes_high": "Completeness Yes; Identification queue; High confidence. Choose only if the complete record has no remaining contradiction, at least one taxon is explicitly unresolved, and all routing facts are direct and consistent.", "sampling_no_medium": "Completeness No; Sampling queue; Medium confidence. Choose only if at least one required completeness element is absent and that conclusion requires one inference.", "validated_yes_medium": "Completeness Yes; Validated; Medium confidence. Choose only if no contradiction or unresolved taxon remains and one required routing fact must be inferred."}, "instructions": "Select the combined completeness judgment, queue, and confidence rating. Completeness is Yes only when effort is at least 30 minutes, the site label exactly matches the register, every specimen lot has a numeric count, and all four habitat observations are present. Route incomplete surveys to Sampling; otherwise route remaining record contradictions to Data Review; otherwise route any unresolved taxon to Identification; otherwise mark Validated. A later explicit statement that an identification is not confirmed controls over its quoted tentative note and is not itself a contradiction. Confidence is ordered High > Medium > Low: High requires all routing facts to be explicit and consistent; Medium requires one inference; Low means a needed fact conflicts.", "type": "choice"}}, "state": [{"speaker": "field survey lead", "text": "Survey D247 was conducted for exactly 30 minutes at registered site ST-08, and its field-sheet label was ST-08. Three specimen lots were collected."}, {"speaker": "ecology data manager", "text": "The stonefly, caddisfly, and mayfly lots have numeric counts of 12, 7, and 6. Flow, substrate, canopy shade, and wetted width are all recorded."}, {"speaker": "taxonomic identifier", "text": "Voucher V4 is explicitly unresolved at species level."}, {"speaker": "records coordinator", "text": "After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026."}, {"speaker": "records coordinator", "text": "The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level."}, {"speaker": "audit note", "text": "Every fact needed to route Survey D247 is explicitly recorded rather than inferred. After identification-correction precedence was applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent."}]}, "method": "c2d", "provenance": {"source_id": "diverse-247", "source_is_synthetic": true, "source_sha256": "9d6d41d1934faa1fee1cda0be4e6d613633e53a4f44a8bad39c426cd84cb0f2d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "data_review_yes_low"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, and neither context adds an exception or altered rule. Survey D247, site ST-08, the 30-minute effort, specimen lots, and routing-relevant bindings remain consistent. The evidence contains two complete factual sentences: \"After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026.\" \"The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level.\" The counterfactual coherently changes only the later entry from resolved to unresolved without creating a contradictory duplicate assertion. Neither context embeds an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universal claims such as every lot having a count remain atomic. The focus atom concerns a factual post-precedence inconsistency, not a policy conclusion. The base and counter assignments are realizable with only a10 changing: an explicit unresolved status can coexist either with an additional post-precedence conflicting status or with no such conflict, while other fields remain consistent. Empty policy_evidence is correct because the governing completeness, routing, precedence, and confidence rules are already retained in the questions object; the state contributes case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail completeness: effort, exact site-label match, numeric counts for every lot, and all four habitat observations are supported. The post-precedence inconsistency in voucher V4's species-resolution status is a remaining contradiction affecting whether Identification routing applies, so Data Review takes priority. Because that needed routing fact conflicts, Low confidence follows.", "rule_index": 0, "sound": true}, {"reason": "The conditions entail completeness. Refutation of a10 entails no post-precedence inconsistency in V4's species-resolution status, while a11 excludes contradictions elsewhere. The explicitly unresolved voucher therefore requires Identification. All routing facts are explicit, and the consistency conditions exclude conflicts, so High confidence follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Survey D247's recorded effort is at least 30 minutes."}, {"id": "a2", "statement": "Survey D247's field-sheet site label exactly matches Survey D247's registered site label."}, {"id": "a3", "statement": "Every specimen lot collected in Survey D247 has a numeric count."}, {"id": "a4", "statement": "Survey D247 has a recorded flow observation."}, {"id": "a5", "statement": "Survey D247 has a recorded substrate observation."}, {"id": "a6", "statement": "Survey D247 has a recorded canopy-shade observation."}, {"id": "a7", "statement": "Survey D247 has a recorded wetted-width observation."}, {"id": "a8", "statement": "Voucher V4 in Survey D247 is explicitly unresolved at species level."}, {"id": "a9", "statement": "Every fact needed to route Survey D247 is explicitly recorded rather than inferred."}, {"id": "a10", "statement": "After identification-correction precedence is applied, Survey D247's complete record assigns inconsistent species-resolution statuses to voucher V4."}, {"id": "a11", "statement": "After identification-correction precedence is applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent."}], "base_state_json": "[{\"speaker\":\"field survey lead\",\"text\":\"Survey D247 was conducted for exactly 30 minutes at registered site ST-08, and its field-sheet label was ST-08. Three specimen lots were collected.\"},{\"speaker\":\"ecology data manager\",\"text\":\"The stonefly, caddisfly, and mayfly lots have numeric counts of 12, 7, and 6. Flow, substrate, canopy shade, and wetted width are all recorded.\"},{\"speaker\":\"taxonomic identifier\",\"text\":\"Voucher V4 is explicitly unresolved at species level.\"},{\"speaker\":\"records coordinator\",\"text\":\"After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026.\"},{\"speaker\":\"records coordinator\",\"text\":\"The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level.\"},{\"speaker\":\"audit note\",\"text\":\"Every fact needed to route Survey D247 is explicitly recorded rather than inferred. After identification-correction precedence was applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["3", "text"], "text": "After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026."}, {"path": ["4", "text"], "text": "The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level."}], "policy_evidence": [], "rules": [{"justification": "All completeness requirements are satisfied. After correction precedence, voucher V4 has conflicting species-resolution statuses, so a contradiction remains in a fact needed to decide whether the survey routes to Identification; the required outcome is therefore Data Review with Low confidence.", "target": "data_review_yes_low", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All completeness requirements are satisfied. Voucher V4 is explicitly unresolved, its post-precedence species-resolution status is not inconsistent, every other record field is consistent, and every routing fact is explicit; the required outcome is therefore Identification with High confidence.", "target": "identification_yes_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026.", "negative_left": "After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026.", "negative_right": "The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as unresolved at species level.", "right": "The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as resolved at species level."}, "verifier_independent_model": false}, "family": "fast-41-diverse-247-009", "id": "fast-41-diverse-247-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_yes_low": "Completeness Yes; Data-review queue; Low confidence. Choose only if a record contradiction remains after correction precedence is applied and a needed routing fact conflicts.", "identification_yes_high": "Completeness Yes; Identification queue; High confidence. Choose only if the complete record has no remaining contradiction, at least one taxon is explicitly unresolved, and all routing facts are direct and consistent.", "sampling_no_medium": "Completeness No; Sampling queue; Medium confidence. Choose only if at least one required completeness element is absent and that conclusion requires one inference.", "validated_yes_medium": "Completeness Yes; Validated; Medium confidence. Choose only if no contradiction or unresolved taxon remains and one required routing fact must be inferred."}, "instructions": "Select the combined completeness judgment, queue, and confidence rating. Completeness is Yes only when effort is at least 30 minutes, the site label exactly matches the register, every specimen lot has a numeric count, and all four habitat observations are present. Route incomplete surveys to Sampling; otherwise route remaining record contradictions to Data Review; otherwise route any unresolved taxon to Identification; otherwise mark Validated. A later explicit statement that an identification is not confirmed controls over its quoted tentative note and is not itself a contradiction. Confidence is ordered High > Medium > Low: High requires all routing facts to be explicit and consistent; Medium requires one inference; Low means a needed fact conflicts.", "type": "choice"}}, "state": [{"speaker": "field survey lead", "text": "Survey D247 was conducted for exactly 30 minutes at registered site ST-08, and its field-sheet label was ST-08. Three specimen lots were collected."}, {"speaker": "ecology data manager", "text": "The stonefly, caddisfly, and mayfly lots have numeric counts of 12, 7, and 6. Flow, substrate, canopy shade, and wetted width are all recorded."}, {"speaker": "taxonomic identifier", "text": "Voucher V4 is explicitly unresolved at species level."}, {"speaker": "records coordinator", "text": "After identification-correction precedence was applied, Survey D247's complete record contained exactly two species-resolution entries for voucher V4, dated 14 June 2026 and 15 June 2026."}, {"speaker": "records coordinator", "text": "The 14 June 2026 entry listed voucher V4 as unresolved at species level, while the 15 June 2026 entry listed it as unresolved at species level."}, {"speaker": "audit note", "text": "Every fact needed to route Survey D247 is explicitly recorded rather than inferred. After identification-correction precedence was applied, every Survey D247 record field other than voucher V4's species-resolution status is internally consistent."}]}, "method": "c2d", "provenance": {"source_id": "diverse-247", "source_is_synthetic": true, "source_sha256": "9d6d41d1934faa1fee1cda0be4e6d613633e53a4f44a8bad39c426cd84cb0f2d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_yes_high"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy and criteria. The Pine Fork, Mara, PF-1/PF-2, and routing-time bindings remain intact. The evidence contains exactly two complete factual sentences: \"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63.\" and \"At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63.\" The counterfactual consistently changes the transfer log from six listed codes to five while retaining the six-code submission. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Field survey lead Mara is preparing the Pine Fork submission for sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records for both methods are complete at both sites. Field-sheet and custody specimen-count records are complete, and both totals are 47. Flow, substrate, canopy, and temperature observations are complete for both sites. Mara’s routing instruction is: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.” Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63."}, {"path": [], "text": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63.", "negative_right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly five vial codes: VX-14, VX-27, VX-31, VX-46, and VX-58.", "right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-001", "id": "fast-41-diverse-248-001-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Field survey lead Mara is preparing the Pine Fork submission for sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records for both methods are complete at both sites. Field-sheet and custody specimen-count records are complete, and both totals are 47. Flow, substrate, canopy, and temperature observations are complete for both sites. Mara’s routing instruction is: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.” Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy and criteria. The Pine Fork, Mara, PF-1/PF-2, and routing-time bindings remain intact. The evidence contains exactly two complete factual sentences: \"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63.\" and \"At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63.\" The counterfactual consistently changes the transfer log from six listed codes to five while retaining the six-code submission. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Field survey lead Mara is preparing the Pine Fork submission for sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records for both methods are complete at both sites. Field-sheet and custody specimen-count records are complete, and both totals are 47. Flow, substrate, canopy, and temperature observations are complete for both sites. Mara’s routing instruction is: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.” Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63."}, {"path": [], "text": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63.", "negative_right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly five vial codes: VX-14, VX-27, VX-31, VX-46, and VX-58.", "right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-001", "id": "fast-41-diverse-248-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Field survey lead Mara is preparing the Pine Fork submission for sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VX-14, VX-27, VX-31, VX-46, VX-58, and VX-63. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly five vial codes: VX-14, VX-27, VX-31, VX-46, and VX-58. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records for both methods are complete at both sites. Field-sheet and custody specimen-count records are complete, and both totals are 47. Flow, substrate, canopy, and temperature observations are complete for both sites. Mara’s routing instruction is: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.” Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, criteria, instructions, scope, and bindings. The base context preserves Mara, Pine Fork, PF-1/PF-2, and routing-time bindings. The counterfactual changes only the transfer-log observation while retaining those bindings. The evidence contains exactly two complete factual sentences: \"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52.\" and \"At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52.\" The counterfactual coherently has six submission codes and five logged codes, with VK-52 absent. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Field survey lead Mara is preparing the Pine Fork submission covering sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records are complete for both sites, as are the field-sheet and custody specimen-count records; both totals are 47. Flow, substrate, canopy, and temperature observations are complete at each site. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52."}, {"path": [], "text": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52.", "negative_right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these five vial codes: VK-17, VK-24, VK-31, VK-38, and VK-45.", "right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-004", "id": "fast-41-diverse-248-004-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Field survey lead Mara is preparing the Pine Fork submission covering sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records are complete for both sites, as are the field-sheet and custody specimen-count records; both totals are 47. Flow, substrate, canopy, and temperature observations are complete at each site. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, criteria, instructions, scope, and bindings. The base context preserves Mara, Pine Fork, PF-1/PF-2, and routing-time bindings. The counterfactual changes only the transfer-log observation while retaining those bindings. The evidence contains exactly two complete factual sentences: \"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52.\" and \"At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52.\" The counterfactual coherently has six submission codes and five logged codes, with VK-52 absent. Neither context embeds a gold answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Field survey lead Mara is preparing the Pine Fork submission covering sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records are complete for both sites, as are the field-sheet and custody specimen-count records; both totals are 47. Flow, substrate, canopy, and temperature observations are complete at each site. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52."}, {"path": [], "text": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52.", "negative_right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these five vial codes: VK-17, VK-24, VK-31, VK-38, and VK-45.", "right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these six vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-004", "id": "fast-41-diverse-248-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Field survey lead Mara is preparing the Pine Fork submission covering sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes: VK-17, VK-24, VK-31, VK-38, VK-45, and VK-52. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 lists exactly these five vial codes: VK-17, VK-24, VK-31, VK-38, and VK-45. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records are complete for both sites, as are the field-sheet and custody specimen-count records; both totals are 47. Flow, substrate, canopy, and temperature observations are complete at each site. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because Mara’s conditional routing statement is unchanged and the original questions object remains verbatim. Entity, site, submission, transfer-log, and routing-time bindings are preserved while only observations change. Evidence contains exactly two complete factual sentences: “At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6.” and “At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6.” The counterfactual coherently changes only the transfer-log code from PF-F6 to PF-G7, creating a missing submission code without contradictory duplicate measurements. Neither context embeds an answer, answer code, rationale, proposition identifier, rule table, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Mara’s Pine Fork survey submission covers sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6. Each site received the required 15-minute kick sample and 5-minute hand search. The required effort records are complete for both sites. Field-sheet and custody specimen-count records are complete, with matching totals of 47. Flow, substrate, canopy, and temperature observations are complete for both sites. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.” Each site received the required 15-minute kick sample and 5-minute hand search.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6."}, {"path": [], "text": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6.", "negative_right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-G7.", "right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-008", "id": "fast-41-diverse-248-008-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Mara’s Pine Fork survey submission covers sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6. Each site received the required 15-minute kick sample and 5-minute hand search. The required effort records are complete for both sites. Field-sheet and custody specimen-count records are complete, with matching totals of 47. Flow, substrate, canopy, and temperature observations are complete for both sites. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.” Each site received the required 15-minute kick sample and 5-minute hand search."}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because Mara’s conditional routing statement is unchanged and the original questions object remains verbatim. Entity, site, submission, transfer-log, and routing-time bindings are preserved while only observations change. Evidence contains exactly two complete factual sentences: “At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6.” and “At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6.” The counterfactual coherently changes only the transfer-log code from PF-F6 to PF-G7, creating a missing submission code without contradictory duplicate measurements. Neither context embeds an answer, answer code, rationale, proposition identifier, rule table, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Mara’s Pine Fork survey submission covers sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6. Each site received the required 15-minute kick sample and 5-minute hand search. The required effort records are complete for both sites. Field-sheet and custody specimen-count records are complete, with matching totals of 47. Flow, substrate, canopy, and temperature observations are complete for both sites. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.” Each site received the required 15-minute kick sample and 5-minute hand search.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6."}, {"path": [], "text": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6.", "negative_right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-G7.", "right": "At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-008", "id": "fast-41-diverse-248-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Mara’s Pine Fork survey submission covers sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-F6. At routing time, the transfer log for Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly the six distinct vial codes PF-A1, PF-B2, PF-C3, PF-D4, PF-E5, and PF-G7. Each site received the required 15-minute kick sample and 5-minute hand search. The required effort records are complete for both sites. Field-sheet and custody specimen-count records are complete, with matching totals of 47. Flow, substrate, canopy, and temperature observations are complete for both sites. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.” Each site received the required 15-minute kick sample and 5-minute hand search."}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain the same survey, sites, actor, routing time, and vial-code bindings. The two evidence spans are complete factual sentences: \"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.\" and \"At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission.\" The counterfactual coherently changes only the transfer-log count from six to five while retaining a six-code submission. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Mara’s Pine Fork survey submission covers sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes. At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission. Each site received the required 15-minute kick sample and 5-minute hand search. The required effort records are complete for both sites. The field-sheet specimen-count record and custody specimen-count record are complete, and their specimen totals match. Flow, substrate, canopy, and temperature observations are complete for both sites. The records are organized for routing under the field survey lead’s instructions. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"path": [], "text": "At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.", "negative_right": "At routing time, that submission’s transfer log contains exactly five distinct vial codes, each corresponding to a different vial code in Mara’s submission.", "right": "At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-012", "id": "fast-41-diverse-248-012-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Mara’s Pine Fork survey submission covers sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes. At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission. Each site received the required 15-minute kick sample and 5-minute hand search. The required effort records are complete for both sites. The field-sheet specimen-count record and custody specimen-count record are complete, and their specimen totals match. Flow, substrate, canopy, and temperature observations are complete for both sites. The records are organized for routing under the field survey lead’s instructions. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain the same survey, sites, actor, routing time, and vial-code bindings. The two evidence spans are complete factual sentences: \"At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.\" and \"At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission.\" The counterfactual coherently changes only the transfer-log count from six to five while retaining a six-code submission. Neither context embeds a gold answer, answer code, rationale, proposition identifier, or classifier-output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Mara’s Pine Fork survey submission covers sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes. At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission. Each site received the required 15-minute kick sample and 5-minute hand search. The required effort records are complete for both sites. The field-sheet specimen-count record and custody specimen-count record are complete, and their specimen totals match. Flow, substrate, canopy, and temperature observations are complete for both sites. The records are organized for routing under the field survey lead’s instructions. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"path": [], "text": "At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.", "negative_right": "At routing time, that submission’s transfer log contains exactly five distinct vial codes, each corresponding to a different vial code in Mara’s submission.", "right": "At routing time, that submission’s transfer log contains exactly six distinct vial codes, each corresponding to a different vial code in Mara’s submission."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-012", "id": "fast-41-diverse-248-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Mara’s Pine Fork survey submission covers sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes. At routing time, that submission’s transfer log contains exactly five distinct vial codes, each corresponding to a different vial code in Mara’s submission. Each site received the required 15-minute kick sample and 5-minute hand search. The required effort records are complete for both sites. The field-sheet specimen-count record and custody specimen-count record are complete, and their specimen totals match. Flow, substrate, canopy, and temperature observations are complete for both sites. The records are organized for routing under the field survey lead’s instructions. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria, and both contexts retain Mara’s routing policy. Both contexts preserve Pine Fork, Mara, sites PF-1 and PF-2, and the six-vial scope. The evidence consists of two complete factual sentences: “At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.” and “At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes.” The counterfactual coherently changes the transfer-log count to five, leaving one of six submission codes absent. Neither context embeds an answer, label rationale, rule table, proposition ID, or classifier output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Field survey lead Mara is routing a Pine Fork submission for sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records are complete for both sites. Field-sheet and custody specimen-count records are complete, and their totals agree. Flow, substrate, canopy, and temperature observations are complete at both sites. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.", "negative_right": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly five distinct vial codes, every one of which is one of that submission’s vial codes.", "right": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-014", "id": "fast-41-diverse-248-014-base", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Field survey lead Mara is routing a Pine Fork submission for sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records are complete for both sites. Field-sheet and custody specimen-count records are complete, and their totals agree. Flow, substrate, canopy, and temperature observations are complete at both sites. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "identification_all_records_complete"}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria, and both contexts retain Mara’s routing policy. Both contexts preserve Pine Fork, Mara, sites PF-1 and PF-2, and the six-vial scope. The evidence consists of two complete factual sentences: “At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.” and “At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes.” The counterfactual coherently changes the transfer-log count to five, leaving one of six submission codes absent. Neither context embeds an answer, label rationale, rule table, proposition ID, or classifier output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal completeness statements over the explicit two-site set remain atomic. The focus atom concerns the factual transfer-log relationship, not a policy conclusion. Base and counter assignments are jointly realizable while changing only whether every submitted vial code appears in the transfer log. Policy evidence preserves the state-originating conditional routing instruction needed by the unchanged question; rules already contained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes exactly six vial codes, all six logged, complete required effort, complete and equal specimen-count records, and complete habitat observations. This is sufficient for the identification profile and excludes every competing substantive profile.", "rule_index": 0, "sound": true}, {"reason": "With exactly six submitted codes, refuting that every submitted code appears in the transfer log entails at least one absent code. Mara’s preserved instruction routes that situation to data review, but the offered ecology-data-manager option requires all codes logged plus unequal totals. Equal totals and completeness of effort, count, and habitat records also exclude the other substantive options, so none_of_above is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"id": "a2", "statement": "At routing time, every vial code in Mara’s Pine Fork survey submission for sites PF-1 and PF-2 appears in that submission’s transfer log."}, {"id": "a3", "statement": "At routing time, the required 15-minute kick-sample effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a4", "statement": "At routing time, the required 5-minute hand-search effort record is complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a5", "statement": "At routing time, the field-sheet specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a6", "statement": "At routing time, the custody specimen-count record is complete for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a7", "statement": "At routing time, the field-sheet specimen total equals the custody-record specimen total for Mara’s Pine Fork survey submission for sites PF-1 and PF-2."}, {"id": "a8", "statement": "At routing time, flow observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a9", "statement": "At routing time, substrate observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a10", "statement": "At routing time, canopy observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}, {"id": "a11", "statement": "At routing time, temperature observations are complete for each of sites PF-1 and PF-2 in Mara’s Pine Fork survey submission."}], "base_state_json": "\"Field survey lead Mara is routing a Pine Fork submission for sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records are complete for both sites. Field-sheet and custody specimen-count records are complete, and their totals agree. Flow, substrate, canopy, and temperature observations are complete at both sites. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes."}, {"path": [], "text": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes."}], "policy_evidence": [{"path": [], "text": "Each site received the required 15-minute kick sample and 5-minute hand search."}, {"path": [], "text": "Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}], "rules": [{"justification": "Exactly six vial codes are submitted, every submitted code appears in the transfer log, both required effort records are complete at both sites, both specimen-count records are complete with equal totals, and every required habitat observation is complete. Thus the identification profile is the sole matching substantive profile and agrees with the field survey lead’s all-codes-present instruction.", "target": "identification_all_records_complete", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "With exactly six submitted vial codes, explicit refutation that every submitted code appears in the transfer log entails that at least one submitted code is absent. The lead therefore requires data review, but the offered ecology-data-manager profile instead requires all codes to be logged and unequal specimen totals. The totals are equal and all effort, count, and habitat records are complete, so no other substantive profile matches this queue-and-condition combination.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.", "negative_left": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes.", "negative_right": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly five distinct vial codes, every one of which is one of that submission’s vial codes.", "right": "At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly six distinct vial codes, every one of which is one of that submission’s vial codes."}, "verifier_independent_model": false}, "family": "fast-41-diverse-248-014", "id": "fast-41-diverse-248-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"data_review_count_mismatch": "Route to the ecology data manager when all vial codes are logged but the field-sheet and custody-record specimen totals disagree.", "identification_all_records_complete": "Route to the taxonomic identifier when all six vial codes appear in the transfer log and effort, counts, and habitat records are complete.", "none_of_above": "Choose when the evidence requires a queue-and-condition combination not represented by any substantive option.", "sampling_effort_incomplete": "Route to the sampling queue when incomplete required sampling effort is the sole relevant deficiency.", "watershed_habitat_missing": "Route to the watershed coordinator when habitat observations are missing but vial transfer, effort, and specimen-count records are complete."}, "instructions": "Select the option whose queue and evidence profile exactly match the survey. The profiles are mutually exclusive: each substantive option applies only when its stated condition is the sole relevant condition. Apply the field survey lead’s conditional routing intent to the supplied records. If the matching queue-and-condition profile is not offered, choose none_of_above.", "type": "choice"}}, "state": "Field survey lead Mara is routing a Pine Fork submission for sites PF-1 and PF-2. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 contains exactly six distinct vial codes. At routing time, Mara’s Pine Fork survey submission for sites PF-1 and PF-2 has a transfer log containing exactly five distinct vial codes, every one of which is one of that submission’s vial codes. Each site received the required 15-minute kick sample and 5-minute hand search. The effort records are complete for both sites. Field-sheet and custody specimen-count records are complete, and their totals agree. Flow, substrate, canopy, and temperature observations are complete at both sites. Each site received the required 15-minute kick sample and 5-minute hand search. Mara states: “If every vial code appears in the transfer log, send the survey to taxonomic identification. If any code is absent, send it to data review. Do not order more sampling.”"}, "method": "c2d", "provenance": {"source_id": "diverse-248", "source_is_synthetic": true, "source_sha256": "9bcf488a0b7f2655132806ca303fa80e160c8d619587b184f9fdd5bda1efdadb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain the ST-41 entity, latest-record timing, and acceptance scope. The focus evidence contains the exact factual sentences \"The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.\" and \"The latest controlling identification sheet for site ST-41 lists an identification count of 27.\" in the base context. The counterfactual retains the exact factual sentence \"The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.\" and changes the second to \"The latest controlling identification sheet for site ST-41 lists an identification count of 31.\". The changed counts form a coherent unresolved count mismatch without contradictory duplicate measurements. Neither context states a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships, including the universally quantified vial/specimen conditions. The focus a4 is a factual count-equality relation rather than a policy conclusion. The base and counter differ only on a4 and are jointly realizable: identification counts can differ from specimen totals even while every specimen appearing in the specimen record has a resolved identification, for example because the identification sheet contains an extra ST-41 entry. Policy evidence preserves all substantive state-originating rules, including temporal priority, routing, and the confidence rubric, without unnecessarily duplicating instructions from the retained questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a4 entails that the latest controlling ST-41 specimen total and identification count are unequal. The original policy explicitly requires those counts to reconcile, so this failure is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes documented effort, correct labeling and provenance, reconciled specimen and identification counts, resolved identifications, and all four required habitat observations under the latest controlling records. These jointly satisfy every stated completeness requirement and are sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The latest controlling effort record for site ST-41 documents six timed kick samples."}, {"id": "a2", "statement": "Every vial assigned to site ST-41 in the latest controlling submission bears an ST-41 site label."}, {"id": "a3", "statement": "Every vial assigned to site ST-41 in the latest controlling submission was collected at ST-41 according to the latest controlling provenance records."}, {"id": "a4", "statement": "The specimen count assigned to site ST-41 in the latest controlling specimen-total record equals the identification count assigned to ST-41 in the latest controlling identification sheet."}, {"id": "a5", "statement": "Every specimen assigned to site ST-41 in the latest controlling specimen record has a resolved taxonomic identification in the latest controlling identification sheet."}, {"id": "a6", "statement": "The latest controlling habitat record for site ST-41 contains a flow observation."}, {"id": "a7", "statement": "The latest controlling habitat record for site ST-41 contains a substrate observation."}, {"id": "a8", "statement": "The latest controlling habitat record for site ST-41 contains a canopy observation."}, {"id": "a9", "statement": "The latest controlling habitat record for site ST-41 contains a bank-condition observation."}], "base_state_json": "{\"case_note\":\"The latest controlling effort record for site ST-41 documents six timed kick samples. The latest controlling submission assigns labeled vials to ST-41, and the latest controlling provenance records place every assigned vial at ST-41. The latest controlling specimen record assigns specimens to ST-41, and the latest controlling identification sheet gives every assigned specimen a resolved taxonomic identification. The latest controlling habitat record contains observations for flow, substrate, canopy, and bank condition. These records are being reviewed for acceptance under the current completeness policy.\",\"evidence\":[\"The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.\",\"The latest controlling identification sheet for site ST-41 lists an identification count of 27.\",\"Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.\",\"Use the latest dated correction over earlier records.\",\"Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification.\"],\"request\":\"Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "The latest controlling specimen-total record for site ST-41 lists a specimen count of 27."}, {"path": ["evidence", "1"], "text": "The latest controlling identification sheet for site ST-41 lists an identification count of 27."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations."}, {"path": ["evidence", "1"], "text": "Use the latest dated correction over earlier records."}, {"path": ["evidence", "2"], "text": "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."}, {"path": ["request"], "text": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}], "rules": [{"justification": "Unequal latest controlling ST-41 specimen and identification counts fail the required count reconciliation, making the survey incomplete and subject to data review.", "target": "false", "when": [{"atom_id": "a4", "state": "refuted"}]}, {"justification": "Under the latest controlling information, the documented effort and site labeling are correct, specimen and identification records reconcile, every specimen has a resolved identification, and all four required habitat observations are present.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.", "negative_left": "The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.", "negative_right": "The latest controlling identification sheet for site ST-41 lists an identification count of 31.", "right": "The latest controlling identification sheet for site ST-41 lists an identification count of 27."}, "verifier_independent_model": false}, "family": "fast-41-diverse-249-004", "id": "fast-41-diverse-249-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one completeness requirement fails under the latest information, so ST-41 is not ready and must be routed according to the stated rule.", "true": "Yes — all required effort, labeling, counts, identifications, and habitat records reconcile under the latest information, so ST-41 is ready for acceptance."}, "instructions": "Answer yes only if every completeness requirement remains satisfied after applying temporal updates. Otherwise answer no. Base confidence on the ordered definitions in the request.", "type": "noul"}}, "state": {"case_note": "The latest controlling effort record for site ST-41 documents six timed kick samples. The latest controlling submission assigns labeled vials to ST-41, and the latest controlling provenance records place every assigned vial at ST-41. The latest controlling specimen record assigns specimens to ST-41, and the latest controlling identification sheet gives every assigned specimen a resolved taxonomic identification. The latest controlling habitat record contains observations for flow, substrate, canopy, and bank condition. These records are being reviewed for acceptance under the current completeness policy.", "evidence": ["The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.", "The latest controlling identification sheet for site ST-41 lists an identification count of 27.", "Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.", "Use the latest dated correction over earlier records.", "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."], "request": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}}, "method": "c2d", "provenance": {"source_id": "diverse-249", "source_is_synthetic": true, "source_sha256": "6e988fdbc9c26cb1580afa5cc854fba73d559d78fe22412ca85c5ea6aaacf1d4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain the ST-41 entity, latest-record timing, and acceptance scope. The focus evidence contains the exact factual sentences \"The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.\" and \"The latest controlling identification sheet for site ST-41 lists an identification count of 27.\" in the base context. The counterfactual retains the exact factual sentence \"The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.\" and changes the second to \"The latest controlling identification sheet for site ST-41 lists an identification count of 31.\". The changed counts form a coherent unresolved count mismatch without contradictory duplicate measurements. Neither context states a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships, including the universally quantified vial/specimen conditions. The focus a4 is a factual count-equality relation rather than a policy conclusion. The base and counter differ only on a4 and are jointly realizable: identification counts can differ from specimen totals even while every specimen appearing in the specimen record has a resolved identification, for example because the identification sheet contains an extra ST-41 entry. Policy evidence preserves all substantive state-originating rules, including temporal priority, routing, and the confidence rubric, without unnecessarily duplicating instructions from the retained questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a4 entails that the latest controlling ST-41 specimen total and identification count are unequal. The original policy explicitly requires those counts to reconcile, so this failure is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes documented effort, correct labeling and provenance, reconciled specimen and identification counts, resolved identifications, and all four required habitat observations under the latest controlling records. These jointly satisfy every stated completeness requirement and are sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The latest controlling effort record for site ST-41 documents six timed kick samples."}, {"id": "a2", "statement": "Every vial assigned to site ST-41 in the latest controlling submission bears an ST-41 site label."}, {"id": "a3", "statement": "Every vial assigned to site ST-41 in the latest controlling submission was collected at ST-41 according to the latest controlling provenance records."}, {"id": "a4", "statement": "The specimen count assigned to site ST-41 in the latest controlling specimen-total record equals the identification count assigned to ST-41 in the latest controlling identification sheet."}, {"id": "a5", "statement": "Every specimen assigned to site ST-41 in the latest controlling specimen record has a resolved taxonomic identification in the latest controlling identification sheet."}, {"id": "a6", "statement": "The latest controlling habitat record for site ST-41 contains a flow observation."}, {"id": "a7", "statement": "The latest controlling habitat record for site ST-41 contains a substrate observation."}, {"id": "a8", "statement": "The latest controlling habitat record for site ST-41 contains a canopy observation."}, {"id": "a9", "statement": "The latest controlling habitat record for site ST-41 contains a bank-condition observation."}], "base_state_json": "{\"case_note\":\"The latest controlling effort record for site ST-41 documents six timed kick samples. The latest controlling submission assigns labeled vials to ST-41, and the latest controlling provenance records place every assigned vial at ST-41. The latest controlling specimen record assigns specimens to ST-41, and the latest controlling identification sheet gives every assigned specimen a resolved taxonomic identification. The latest controlling habitat record contains observations for flow, substrate, canopy, and bank condition. These records are being reviewed for acceptance under the current completeness policy.\",\"evidence\":[\"The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.\",\"The latest controlling identification sheet for site ST-41 lists an identification count of 27.\",\"Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.\",\"Use the latest dated correction over earlier records.\",\"Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification.\"],\"request\":\"Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "The latest controlling specimen-total record for site ST-41 lists a specimen count of 27."}, {"path": ["evidence", "1"], "text": "The latest controlling identification sheet for site ST-41 lists an identification count of 27."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations."}, {"path": ["evidence", "1"], "text": "Use the latest dated correction over earlier records."}, {"path": ["evidence", "2"], "text": "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."}, {"path": ["request"], "text": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}], "rules": [{"justification": "Unequal latest controlling ST-41 specimen and identification counts fail the required count reconciliation, making the survey incomplete and subject to data review.", "target": "false", "when": [{"atom_id": "a4", "state": "refuted"}]}, {"justification": "Under the latest controlling information, the documented effort and site labeling are correct, specimen and identification records reconcile, every specimen has a resolved identification, and all four required habitat observations are present.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.", "negative_left": "The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.", "negative_right": "The latest controlling identification sheet for site ST-41 lists an identification count of 31.", "right": "The latest controlling identification sheet for site ST-41 lists an identification count of 27."}, "verifier_independent_model": false}, "family": "fast-41-diverse-249-004", "id": "fast-41-diverse-249-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one completeness requirement fails under the latest information, so ST-41 is not ready and must be routed according to the stated rule.", "true": "Yes — all required effort, labeling, counts, identifications, and habitat records reconcile under the latest information, so ST-41 is ready for acceptance."}, "instructions": "Answer yes only if every completeness requirement remains satisfied after applying temporal updates. Otherwise answer no. Base confidence on the ordered definitions in the request.", "type": "noul"}}, "state": {"case_note": "The latest controlling effort record for site ST-41 documents six timed kick samples. The latest controlling submission assigns labeled vials to ST-41, and the latest controlling provenance records place every assigned vial at ST-41. The latest controlling specimen record assigns specimens to ST-41, and the latest controlling identification sheet gives every assigned specimen a resolved taxonomic identification. The latest controlling habitat record contains observations for flow, substrate, canopy, and bank condition. These records are being reviewed for acceptance under the current completeness policy.", "evidence": ["The latest controlling specimen-total record for site ST-41 lists a specimen count of 27.", "The latest controlling identification sheet for site ST-41 lists an identification count of 31.", "Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.", "Use the latest dated correction over earlier records.", "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."], "request": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}}, "method": "c2d", "provenance": {"source_id": "diverse-249", "source_is_synthetic": true, "source_sha256": "6e988fdbc9c26cb1580afa5cc854fba73d559d78fe22412ca85c5ea6aaacf1d4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The required evidence quotes are “The latest controlling specimen-total record assigns 17 specimens to site ST-41.” and “The latest controlling identification sheet assigns 17 identifications to site ST-41.”; the contexts preserve the policy and bindings, and the counterfactual coherently changes the identification count to 19 without stating an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships, including the universally quantified vial/specimen conditions. The focus a4 is a factual count-equality relation rather than a policy conclusion. The base and counter differ only on a4 and are jointly realizable: identification counts can differ from specimen totals even while every specimen appearing in the specimen record has a resolved identification, for example because the identification sheet contains an extra ST-41 entry. Policy evidence preserves all substantive state-originating rules, including temporal priority, routing, and the confidence rubric, without unnecessarily duplicating instructions from the retained questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a4 entails that the latest controlling ST-41 specimen total and identification count are unequal. The original policy explicitly requires those counts to reconcile, so this failure is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes documented effort, correct labeling and provenance, reconciled specimen and identification counts, resolved identifications, and all four required habitat observations under the latest controlling records. These jointly satisfy every stated completeness requirement and are sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The latest controlling effort record for site ST-41 documents six timed kick samples."}, {"id": "a2", "statement": "Every vial assigned to site ST-41 in the latest controlling submission bears an ST-41 site label."}, {"id": "a3", "statement": "Every vial assigned to site ST-41 in the latest controlling submission was collected at ST-41 according to the latest controlling provenance records."}, {"id": "a4", "statement": "The specimen count assigned to site ST-41 in the latest controlling specimen-total record equals the identification count assigned to ST-41 in the latest controlling identification sheet."}, {"id": "a5", "statement": "Every specimen assigned to site ST-41 in the latest controlling specimen record has a resolved taxonomic identification in the latest controlling identification sheet."}, {"id": "a6", "statement": "The latest controlling habitat record for site ST-41 contains a flow observation."}, {"id": "a7", "statement": "The latest controlling habitat record for site ST-41 contains a substrate observation."}, {"id": "a8", "statement": "The latest controlling habitat record for site ST-41 contains a canopy observation."}, {"id": "a9", "statement": "The latest controlling habitat record for site ST-41 contains a bank-condition observation."}], "base_state_json": "{\"context\":\"A field team’s latest controlling packet for site ST-41 records six timed kick samples. All assigned vials carry the ST-41 label, and provenance records place each collection at ST-41. The latest controlling specimen-total record assigns 17 specimens to site ST-41. The latest controlling identification sheet assigns 17 identifications to site ST-41. Every specimen listed for ST-41 has a resolved taxonomic identification. The habitat record includes observations for flow, substrate, canopy, and bank condition. These records are the current submission for an acceptance review, and no later correction is noted in the packet.\",\"evidence\":[\"Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.\",\"Use the latest dated correction over earlier records.\",\"Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification.\"],\"request\":\"Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "The latest controlling specimen-total record assigns 17 specimens to site ST-41."}, {"path": ["context"], "text": "The latest controlling identification sheet assigns 17 identifications to site ST-41."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations."}, {"path": ["evidence", "1"], "text": "Use the latest dated correction over earlier records."}, {"path": ["evidence", "2"], "text": "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."}, {"path": ["request"], "text": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}], "rules": [{"justification": "Unequal latest controlling ST-41 specimen and identification counts fail the required count reconciliation, making the survey incomplete and subject to data review.", "target": "false", "when": [{"atom_id": "a4", "state": "refuted"}]}, {"justification": "Under the latest controlling information, the documented effort and site labeling are correct, specimen and identification records reconcile, every specimen has a resolved identification, and all four required habitat observations are present.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The latest controlling specimen-total record assigns 17 specimens to site ST-41.", "negative_left": "The latest controlling specimen-total record assigns 17 specimens to site ST-41.", "negative_right": "The latest controlling identification sheet assigns 19 identifications to site ST-41.", "right": "The latest controlling identification sheet assigns 17 identifications to site ST-41."}, "verifier_independent_model": false}, "family": "fast-41-diverse-249-008", "id": "fast-41-diverse-249-008-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one completeness requirement fails under the latest information, so ST-41 is not ready and must be routed according to the stated rule.", "true": "Yes — all required effort, labeling, counts, identifications, and habitat records reconcile under the latest information, so ST-41 is ready for acceptance."}, "instructions": "Answer yes only if every completeness requirement remains satisfied after applying temporal updates. Otherwise answer no. Base confidence on the ordered definitions in the request.", "type": "noul"}}, "state": {"context": "A field team’s latest controlling packet for site ST-41 records six timed kick samples. All assigned vials carry the ST-41 label, and provenance records place each collection at ST-41. The latest controlling specimen-total record assigns 17 specimens to site ST-41. The latest controlling identification sheet assigns 17 identifications to site ST-41. Every specimen listed for ST-41 has a resolved taxonomic identification. The habitat record includes observations for flow, substrate, canopy, and bank condition. These records are the current submission for an acceptance review, and no later correction is noted in the packet.", "evidence": ["Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.", "Use the latest dated correction over earlier records.", "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."], "request": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}}, "method": "c2d", "provenance": {"source_id": "diverse-249", "source_is_synthetic": true, "source_sha256": "6e988fdbc9c26cb1580afa5cc854fba73d559d78fe22412ca85c5ea6aaacf1d4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The required evidence quotes are “The latest controlling specimen-total record assigns 17 specimens to site ST-41.” and “The latest controlling identification sheet assigns 17 identifications to site ST-41.”; the contexts preserve the policy and bindings, and the counterfactual coherently changes the identification count to 19 without stating an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships, including the universally quantified vial/specimen conditions. The focus a4 is a factual count-equality relation rather than a policy conclusion. The base and counter differ only on a4 and are jointly realizable: identification counts can differ from specimen totals even while every specimen appearing in the specimen record has a resolved identification, for example because the identification sheet contains an extra ST-41 entry. Policy evidence preserves all substantive state-originating rules, including temporal priority, routing, and the confidence rubric, without unnecessarily duplicating instructions from the retained questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a4 entails that the latest controlling ST-41 specimen total and identification count are unequal. The original policy explicitly requires those counts to reconcile, so this failure is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes documented effort, correct labeling and provenance, reconciled specimen and identification counts, resolved identifications, and all four required habitat observations under the latest controlling records. These jointly satisfy every stated completeness requirement and are sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The latest controlling effort record for site ST-41 documents six timed kick samples."}, {"id": "a2", "statement": "Every vial assigned to site ST-41 in the latest controlling submission bears an ST-41 site label."}, {"id": "a3", "statement": "Every vial assigned to site ST-41 in the latest controlling submission was collected at ST-41 according to the latest controlling provenance records."}, {"id": "a4", "statement": "The specimen count assigned to site ST-41 in the latest controlling specimen-total record equals the identification count assigned to ST-41 in the latest controlling identification sheet."}, {"id": "a5", "statement": "Every specimen assigned to site ST-41 in the latest controlling specimen record has a resolved taxonomic identification in the latest controlling identification sheet."}, {"id": "a6", "statement": "The latest controlling habitat record for site ST-41 contains a flow observation."}, {"id": "a7", "statement": "The latest controlling habitat record for site ST-41 contains a substrate observation."}, {"id": "a8", "statement": "The latest controlling habitat record for site ST-41 contains a canopy observation."}, {"id": "a9", "statement": "The latest controlling habitat record for site ST-41 contains a bank-condition observation."}], "base_state_json": "{\"context\":\"A field team’s latest controlling packet for site ST-41 records six timed kick samples. All assigned vials carry the ST-41 label, and provenance records place each collection at ST-41. The latest controlling specimen-total record assigns 17 specimens to site ST-41. The latest controlling identification sheet assigns 17 identifications to site ST-41. Every specimen listed for ST-41 has a resolved taxonomic identification. The habitat record includes observations for flow, substrate, canopy, and bank condition. These records are the current submission for an acceptance review, and no later correction is noted in the packet.\",\"evidence\":[\"Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.\",\"Use the latest dated correction over earlier records.\",\"Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification.\"],\"request\":\"Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "The latest controlling specimen-total record assigns 17 specimens to site ST-41."}, {"path": ["context"], "text": "The latest controlling identification sheet assigns 17 identifications to site ST-41."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations."}, {"path": ["evidence", "1"], "text": "Use the latest dated correction over earlier records."}, {"path": ["evidence", "2"], "text": "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."}, {"path": ["request"], "text": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}], "rules": [{"justification": "Unequal latest controlling ST-41 specimen and identification counts fail the required count reconciliation, making the survey incomplete and subject to data review.", "target": "false", "when": [{"atom_id": "a4", "state": "refuted"}]}, {"justification": "Under the latest controlling information, the documented effort and site labeling are correct, specimen and identification records reconcile, every specimen has a resolved identification, and all four required habitat observations are present.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The latest controlling specimen-total record assigns 17 specimens to site ST-41.", "negative_left": "The latest controlling specimen-total record assigns 17 specimens to site ST-41.", "negative_right": "The latest controlling identification sheet assigns 19 identifications to site ST-41.", "right": "The latest controlling identification sheet assigns 17 identifications to site ST-41."}, "verifier_independent_model": false}, "family": "fast-41-diverse-249-008", "id": "fast-41-diverse-249-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one completeness requirement fails under the latest information, so ST-41 is not ready and must be routed according to the stated rule.", "true": "Yes — all required effort, labeling, counts, identifications, and habitat records reconcile under the latest information, so ST-41 is ready for acceptance."}, "instructions": "Answer yes only if every completeness requirement remains satisfied after applying temporal updates. Otherwise answer no. Base confidence on the ordered definitions in the request.", "type": "noul"}}, "state": {"context": "A field team’s latest controlling packet for site ST-41 records six timed kick samples. All assigned vials carry the ST-41 label, and provenance records place each collection at ST-41. The latest controlling specimen-total record assigns 17 specimens to site ST-41. The latest controlling identification sheet assigns 19 identifications to site ST-41. Every specimen listed for ST-41 has a resolved taxonomic identification. The habitat record includes observations for flow, substrate, canopy, and bank condition. These records are the current submission for an acceptance review, and no later correction is noted in the packet.", "evidence": ["Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.", "Use the latest dated correction over earlier records.", "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."], "request": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}}, "method": "c2d", "provenance": {"source_id": "diverse-249", "source_is_synthetic": true, "source_sha256": "6e988fdbc9c26cb1580afa5cc854fba73d559d78fe22412ca85c5ea6aaacf1d4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and both contexts retain its governing rules. Question bindings are unchanged for the three route-sheet sites, packet, rubric, and completeness-confidence request. Base evidence quotes: \"The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.\" \"The PR-2 count sheet records 37 specimens for the 14 June 2026 collection.\" Counterfactual evidence quotes: \"The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.\" \"The PR-2 count sheet records 29 specimens for the 14 June 2026 collection.\" Each evidence item is a complete factual sentence. The counterfactual's 37-versus-29 discrepancy is coherent and does not duplicate contradictory measurements within either evidence set. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager is reviewing the watershed coordinator's route packet before sending it to the taxonomic identifier. The route sheet scopes exactly PR-1, PR-2, and PR-3. PR-1 and PR-2 are wet sites, and each has recorded sampling effort, a mappable site label, and habitat observations. PR-1's jar and count-sheet records are internally aligned. PR-3 is dry; its submitted photo is dated and documents the absence of flowing water, and habitat observations are present. The review found no unresolved required-field deficiency or record discrepancy across the three sites except the possible issue represented by the PR-2 specimen-count comparison. No nonrequired formatting, wording, or organizational irregularity remains. Rate the packet's completeness confidence under the supplied rubric.\",\"evidence\":[\"The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.\",\"The PR-2 count sheet records 37 specimens for the 14 June 2026 collection.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "0"], "text": "The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection."}, {"path": ["evidence", "1"], "text": "The PR-2 count sheet records 37 specimens for the 14 June 2026 collection."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.", "negative_left": "The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.", "negative_right": "The PR-2 count sheet records 29 specimens for the 14 June 2026 collection.", "right": "The PR-2 count sheet records 37 specimens for the 14 June 2026 collection."}, "verifier_independent_model": false}, "family": "fast-41-diverse-252-007", "id": "fast-41-diverse-252-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is reviewing the watershed coordinator's route packet before sending it to the taxonomic identifier. The route sheet scopes exactly PR-1, PR-2, and PR-3. PR-1 and PR-2 are wet sites, and each has recorded sampling effort, a mappable site label, and habitat observations. PR-1's jar and count-sheet records are internally aligned. PR-3 is dry; its submitted photo is dated and documents the absence of flowing water, and habitat observations are present. The review found no unresolved required-field deficiency or record discrepancy across the three sites except the possible issue represented by the PR-2 specimen-count comparison. No nonrequired formatting, wording, or organizational irregularity remains. Rate the packet's completeness confidence under the supplied rubric.", "evidence": ["The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.", "The PR-2 count sheet records 37 specimens for the 14 June 2026 collection."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and both contexts retain its governing rules. Question bindings are unchanged for the three route-sheet sites, packet, rubric, and completeness-confidence request. Base evidence quotes: \"The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.\" \"The PR-2 count sheet records 37 specimens for the 14 June 2026 collection.\" Counterfactual evidence quotes: \"The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.\" \"The PR-2 count sheet records 29 specimens for the 14 June 2026 collection.\" Each evidence item is a complete factual sentence. The counterfactual's 37-versus-29 discrepancy is coherent and does not duplicate contradictory measurements within either evidence set. Neither context embeds a gold answer, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager is reviewing the watershed coordinator's route packet before sending it to the taxonomic identifier. The route sheet scopes exactly PR-1, PR-2, and PR-3. PR-1 and PR-2 are wet sites, and each has recorded sampling effort, a mappable site label, and habitat observations. PR-1's jar and count-sheet records are internally aligned. PR-3 is dry; its submitted photo is dated and documents the absence of flowing water, and habitat observations are present. The review found no unresolved required-field deficiency or record discrepancy across the three sites except the possible issue represented by the PR-2 specimen-count comparison. No nonrequired formatting, wording, or organizational irregularity remains. Rate the packet's completeness confidence under the supplied rubric.\",\"evidence\":[\"The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.\",\"The PR-2 count sheet records 37 specimens for the 14 June 2026 collection.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "0"], "text": "The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection."}, {"path": ["evidence", "1"], "text": "The PR-2 count sheet records 37 specimens for the 14 June 2026 collection."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.", "negative_left": "The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.", "negative_right": "The PR-2 count sheet records 29 specimens for the 14 June 2026 collection.", "right": "The PR-2 count sheet records 37 specimens for the 14 June 2026 collection."}, "verifier_independent_model": false}, "family": "fast-41-diverse-252-007", "id": "fast-41-diverse-252-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is reviewing the watershed coordinator's route packet before sending it to the taxonomic identifier. The route sheet scopes exactly PR-1, PR-2, and PR-3. PR-1 and PR-2 are wet sites, and each has recorded sampling effort, a mappable site label, and habitat observations. PR-1's jar and count-sheet records are internally aligned. PR-3 is dry; its submitted photo is dated and documents the absence of flowing water, and habitat observations are present. The review found no unresolved required-field deficiency or record discrepancy across the three sites except the possible issue represented by the PR-2 specimen-count comparison. No nonrequired formatting, wording, or organizational irregularity remains. Rate the packet's completeness confidence under the supplied rubric.", "evidence": ["The labeled PR-2 jar inventory contains 37 specimens from the 14 June 2026 collection.", "The PR-2 count sheet records 29 specimens for the 14 June 2026 collection."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy and scoring criteria. Both contexts retain the three-site scope and PR-1, PR-2, and PR-3 bindings. The evidence spans are complete factual sentences: \"On 14 May 2026, the PR-2 specimen jar contained 23 specimens.\" and \"On 14 May 2026, PR-2's count sheet recorded 23 specimens.\" The counterfactual coherently changes the count-sheet measurement to 24 and represents one record discrepancy without duplicating contradictory claims. Neither context states a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"case_note\":\"The ecology data manager reviewed the watershed coordinator's route packet for PR-1, PR-2, and PR-3 before taxonomic identification. PR-1 is wet, with recorded sampling effort, a mappable label, habitat observations, and 18 specimens shown on both its jar and count sheet. PR-2 is wet and has recorded effort, a mappable label, and habitat observations. On 14 May 2026, the PR-2 specimen jar contained 23 specimens. On 14 May 2026, PR-2's count sheet recorded 23 specimens. PR-3 is dry; its dated photo documents the absence of flowing water, and habitat observations are present. The packet has no unresolved required-field deficiency beyond any issue represented by the PR-2 comparison, no unresolved record discrepancy beyond that comparison, and no remaining nonrequired formatting, wording, or organizational irregularity. The route sheet lists exactly these three scoped sites.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["case_note"], "text": "On 14 May 2026, the PR-2 specimen jar contained 23 specimens."}, {"path": ["case_note"], "text": "On 14 May 2026, PR-2's count sheet recorded 23 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "On 14 May 2026, the PR-2 specimen jar contained 23 specimens.", "negative_left": "On 14 May 2026, the PR-2 specimen jar contained 23 specimens.", "negative_right": "On 14 May 2026, PR-2's count sheet recorded 24 specimens.", "right": "On 14 May 2026, PR-2's count sheet recorded 23 specimens."}, "verifier_independent_model": false}, "family": "fast-41-diverse-252-015", "id": "fast-41-diverse-252-015-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"case_note": "The ecology data manager reviewed the watershed coordinator's route packet for PR-1, PR-2, and PR-3 before taxonomic identification. PR-1 is wet, with recorded sampling effort, a mappable label, habitat observations, and 18 specimens shown on both its jar and count sheet. PR-2 is wet and has recorded effort, a mappable label, and habitat observations. On 14 May 2026, the PR-2 specimen jar contained 23 specimens. On 14 May 2026, PR-2's count sheet recorded 23 specimens. PR-3 is dry; its dated photo documents the absence of flowing water, and habitat observations are present. The packet has no unresolved required-field deficiency beyond any issue represented by the PR-2 comparison, no unresolved record discrepancy beyond that comparison, and no remaining nonrequired formatting, wording, or organizational irregularity. The route sheet lists exactly these three scoped sites."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy and scoring criteria. Both contexts retain the three-site scope and PR-1, PR-2, and PR-3 bindings. The evidence spans are complete factual sentences: \"On 14 May 2026, the PR-2 specimen jar contained 23 specimens.\" and \"On 14 May 2026, PR-2's count sheet recorded 23 specimens.\" The counterfactual coherently changes the count-sheet measurement to 24 and represents one record discrepancy without duplicating contradictory claims. Neither context states a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"case_note\":\"The ecology data manager reviewed the watershed coordinator's route packet for PR-1, PR-2, and PR-3 before taxonomic identification. PR-1 is wet, with recorded sampling effort, a mappable label, habitat observations, and 18 specimens shown on both its jar and count sheet. PR-2 is wet and has recorded effort, a mappable label, and habitat observations. On 14 May 2026, the PR-2 specimen jar contained 23 specimens. On 14 May 2026, PR-2's count sheet recorded 23 specimens. PR-3 is dry; its dated photo documents the absence of flowing water, and habitat observations are present. The packet has no unresolved required-field deficiency beyond any issue represented by the PR-2 comparison, no unresolved record discrepancy beyond that comparison, and no remaining nonrequired formatting, wording, or organizational irregularity. The route sheet lists exactly these three scoped sites.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["case_note"], "text": "On 14 May 2026, the PR-2 specimen jar contained 23 specimens."}, {"path": ["case_note"], "text": "On 14 May 2026, PR-2's count sheet recorded 23 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "On 14 May 2026, the PR-2 specimen jar contained 23 specimens.", "negative_left": "On 14 May 2026, the PR-2 specimen jar contained 23 specimens.", "negative_right": "On 14 May 2026, PR-2's count sheet recorded 24 specimens.", "right": "On 14 May 2026, PR-2's count sheet recorded 23 specimens."}, "verifier_independent_model": false}, "family": "fast-41-diverse-252-015", "id": "fast-41-diverse-252-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"case_note": "The ecology data manager reviewed the watershed coordinator's route packet for PR-1, PR-2, and PR-3 before taxonomic identification. PR-1 is wet, with recorded sampling effort, a mappable label, habitat observations, and 18 specimens shown on both its jar and count sheet. PR-2 is wet and has recorded effort, a mappable label, and habitat observations. On 14 May 2026, the PR-2 specimen jar contained 23 specimens. On 14 May 2026, PR-2's count sheet recorded 24 specimens. PR-3 is dry; its dated photo documents the absence of flowing water, and habitat observations are present. The packet has no unresolved required-field deficiency beyond any issue represented by the PR-2 comparison, no unresolved record discrepancy beyond that comparison, and no remaining nonrequired formatting, wording, or organizational irregularity. The route sheet lists exactly these three scoped sites."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy and bindings, both quoted spans are complete factual sentences—“The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset.” and “That test record reports 106 cells for SR-27 and 100 cells for CR-26.”—and the counterfactual's 104-versus-100 result is coherent without embedding an answer or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A fictional three-channel cell-assay case was reviewed on 2026-09-17. The latest same-session bead check measured a 0.4 px channel offset, below the acquisition threshold. A fixed slide defect appeared in 2 of 30 reviewed fields. Current calibration passed, and the current technical QC passed. The most recent relevant replicate review recorded a coefficient of variation of 8%. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies. The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset. That test record reports 106 cells for SR-27 and 100 cells for CR-26.\",\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset."}, {"path": ["context"], "text": "That test record reports 106 cells for SR-27 and 100 cells for CR-26."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset.", "negative_left": "The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset.", "negative_right": "That test record reports 104 cells for SR-27 and 100 cells for CR-26.", "right": "That test record reports 106 cells for SR-27 and 100 cells for CR-26."}, "verifier_independent_model": false}, "family": "fast-41-diverse-259-005", "id": "fast-41-diverse-259-005-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A fictional three-channel cell-assay case was reviewed on 2026-09-17. The latest same-session bead check measured a 0.4 px channel offset, below the acquisition threshold. A fixed slide defect appeared in 2 of 30 reviewed fields. Current calibration passed, and the current technical QC passed. The most recent relevant replicate review recorded a coefficient of variation of 8%. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies. The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset. That test record reports 106 cells for SR-27 and 100 cells for CR-26.", "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy and bindings, both quoted spans are complete factual sentences—“The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset.” and “That test record reports 106 cells for SR-27 and 100 cells for CR-26.”—and the counterfactual's 104-versus-100 result is coherent without embedding an answer or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A fictional three-channel cell-assay case was reviewed on 2026-09-17. The latest same-session bead check measured a 0.4 px channel offset, below the acquisition threshold. A fixed slide defect appeared in 2 of 30 reviewed fields. Current calibration passed, and the current technical QC passed. The most recent relevant replicate review recorded a coefficient of variation of 8%. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies. The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset. That test record reports 106 cells for SR-27 and 100 cells for CR-26.\",\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset."}, {"path": ["context"], "text": "That test record reports 106 cells for SR-27 and 100 cells for CR-26."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset.", "negative_left": "The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset.", "negative_right": "That test record reports 104 cells for SR-27 and 100 cells for CR-26.", "right": "That test record reports 106 cells for SR-27 and 100 cells for CR-26."}, "verifier_independent_model": false}, "family": "fast-41-diverse-259-005", "id": "fast-41-diverse-259-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A fictional three-channel cell-assay case was reviewed on 2026-09-17. The latest same-session bead check measured a 0.4 px channel offset, below the acquisition threshold. A fixed slide defect appeared in 2 of 30 reviewed fields. Current calibration passed, and the current technical QC passed. The most recent relevant replicate review recorded a coefficient of variation of 8%. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies. The laboratory's 2026-09-17 test record identifies segmentation rerun SR-27 and comparison run CR-26 as the most recent relevant pair, and records that both used the same raw image dataset. That test record reports 104 cells for SR-27 and 100 cells for CR-26.", "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question policy, request, date, assay, and measurement bindings; each evidence list contains two complete factual sentences, including “In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.” and “The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26.” in the base context, while the counterfactual coherently changes only SR-27 to 1,040 cells, and neither context leaks an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"case_note\":\"A fictional three-channel cell assay was reviewed on 2026-08-14. The current calibration passed, and current technical QC passed. The latest same-session bead test measured a 0.4 px channel offset, while an earlier sample overlay estimate was 1.6 px; later validated calibration superseded that estimate. A fixed slide defect appeared in 2 of 30 reviewed fields. The latest replicate CV was 8%. The rerun and comparison run used the same raw image dataset. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.\",\"The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26."}, {"path": ["evidence", "1"], "text": "The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.", "negative_left": "In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.", "negative_right": "The 2026-08-14 count ledger records 1,040 cells for SR-27 and 1,000 cells for CR-26.", "right": "The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26."}, "verifier_independent_model": false}, "family": "fast-41-diverse-259-008", "id": "fast-41-diverse-259-008-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"case_note": "A fictional three-channel cell assay was reviewed on 2026-08-14. The current calibration passed, and current technical QC passed. The latest same-session bead test measured a 0.4 px channel offset, while an earlier sample overlay estimate was 1.6 px; later validated calibration superseded that estimate. A fixed slide defect appeared in 2 of 30 reviewed fields. The latest replicate CV was 8%. The rerun and comparison run used the same raw image dataset. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.", "The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question policy, request, date, assay, and measurement bindings; each evidence list contains two complete factual sentences, including “In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.” and “The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26.” in the base context, while the counterfactual coherently changes only SR-27 to 1,040 cells, and neither context leaks an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"case_note\":\"A fictional three-channel cell assay was reviewed on 2026-08-14. The current calibration passed, and current technical QC passed. The latest same-session bead test measured a 0.4 px channel offset, while an earlier sample overlay estimate was 1.6 px; later validated calibration superseded that estimate. A fixed slide defect appeared in 2 of 30 reviewed fields. The latest replicate CV was 8%. The rerun and comparison run used the same raw image dataset. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.\",\"The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "0"], "text": "In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26."}, {"path": ["evidence", "1"], "text": "The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.", "negative_left": "In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.", "negative_right": "The 2026-08-14 count ledger records 1,040 cells for SR-27 and 1,000 cells for CR-26.", "right": "The 2026-08-14 count ledger records 1,120 cells for SR-27 and 1,000 cells for CR-26."}, "verifier_independent_model": false}, "family": "fast-41-diverse-259-008", "id": "fast-41-diverse-259-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"case_note": "A fictional three-channel cell assay was reviewed on 2026-08-14. The current calibration passed, and current technical QC passed. The latest same-session bead test measured a 0.4 px channel offset, while an earlier sample overlay estimate was 1.6 px; later validated calibration superseded that estimate. A fixed slide defect appeared in 2 of 30 reviewed fields. The latest replicate CV was 8%. The rerun and comparison run used the same raw image dataset. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["In the 2026-08-14 analysis record, segmentation rerun SR-27 is the most recent relevant rerun, and its comparison run is CR-26.", "The 2026-08-14 count ledger records 1,040 cells for SR-27 and 1,000 cells for CR-26."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions object and the same routing policy. The request, assay-routing scope, and relevant dates and runs remain bound consistently. The two required factual spans are \"The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.\" and \"On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44.\" The counterfactual coherently changes only the rerun difference from 9 to 4, with no contradictory duplicate measurement. Neither full context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"The 2026-09-17 same-session bead check measured a 0.4 px channel offset, and the current calibration passed.\",\"A fixed slide fold appeared in 2 of 30 reviewed fields.\",\"The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.\",\"On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44.\",\"SR-45 and CR-44 used the same raw image dataset.\",\"The most recent replicate review reported a coefficient of variation of 8%, and current technical QC passed.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells."}, {"path": ["evidence", "3"], "text": "On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.", "negative_left": "The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.", "negative_right": "On 2026-09-17, rerun SR-45 recorded 4 more cells than comparison run CR-44.", "right": "On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44."}, "verifier_independent_model": false}, "family": "fast-41-diverse-259-014", "id": "fast-41-diverse-259-014-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["The 2026-09-17 same-session bead check measured a 0.4 px channel offset, and the current calibration passed.", "A fixed slide fold appeared in 2 of 30 reviewed fields.", "The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.", "On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44.", "SR-45 and CR-44 used the same raw image dataset.", "The most recent replicate review reported a coefficient of variation of 8%, and current technical QC passed."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions object and the same routing policy. The request, assay-routing scope, and relevant dates and runs remain bound consistently. The two required factual spans are \"The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.\" and \"On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44.\" The counterfactual coherently changes only the rerun difference from 9 to 4, with no contradictory duplicate measurement. Neither full context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"The 2026-09-17 same-session bead check measured a 0.4 px channel offset, and the current calibration passed.\",\"A fixed slide fold appeared in 2 of 30 reviewed fields.\",\"The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.\",\"On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44.\",\"SR-45 and CR-44 used the same raw image dataset.\",\"The most recent replicate review reported a coefficient of variation of 8%, and current technical QC passed.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells."}, {"path": ["evidence", "3"], "text": "On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.", "negative_left": "The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.", "negative_right": "On 2026-09-17, rerun SR-45 recorded 4 more cells than comparison run CR-44.", "right": "On 2026-09-17, rerun SR-45 recorded 9 more cells than comparison run CR-44."}, "verifier_independent_model": false}, "family": "fast-41-diverse-259-014", "id": "fast-41-diverse-259-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["The 2026-09-17 same-session bead check measured a 0.4 px channel offset, and the current calibration passed.", "A fixed slide fold appeared in 2 of 30 reviewed fields.", "The comparison run CR-44 for the most recent relevant segmentation rerun on 2026-09-17 recorded 137 cells.", "On 2026-09-17, rerun SR-45 recorded 4 more cells than comparison run CR-44.", "SR-45 and CR-44 used the same raw image dataset.", "The most recent replicate review reported a coefficient of variation of 8%, and current technical QC passed."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing rubric and precedence rules are preserved in both contexts alongside the verbatim questions. The case, routing task, measurements, and workflow timing remain properly bound. Both evidence spans are complete factual sentences. The counterfactual changes only calibration drift from 2.6% to 1.4%, without creating contradictions. Neither context embeds an answer, code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"case note\",\"text\":\"The supplied raw-image review recorded image-confirmed bubble coverage at exactly 3.0%, and the same review recorded saturated pixels at exactly 1.0%. Replicate segmentation checks recorded 5.2% disagreement. These readings were entered for routing without alternate technician measurements.\"},{\"speaker\":\"calibration auditor\",\"text\":\"At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field.\"},{\"speaker\":\"calibration auditor\",\"text\":\"The recorded percentage value for the calibration-drift field in the case workflow was 2.6%.\"},{\"speaker\":\"assay scientist\",\"text\":\"The case record is complete for the one-specialist routing decision, with the image review, calibration-log entry, and replicate check retained together.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field."}, {"path": ["3", "text"], "text": "The recorded percentage value for the calibration-drift field in the case workflow was 2.6%."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field.", "negative_left": "At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field.", "negative_right": "The recorded percentage value for the calibration-drift field in the case workflow was 1.4%.", "right": "The recorded percentage value for the calibration-drift field in the case workflow was 2.6%."}, "verifier_independent_model": false}, "family": "fast-41-diverse-260-013", "id": "fast-41-diverse-260-013-base", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "case note", "text": "The supplied raw-image review recorded image-confirmed bubble coverage at exactly 3.0%, and the same review recorded saturated pixels at exactly 1.0%. Replicate segmentation checks recorded 5.2% disagreement. These readings were entered for routing without alternate technician measurements."}, {"speaker": "calibration auditor", "text": "At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field."}, {"speaker": "calibration auditor", "text": "The recorded percentage value for the calibration-drift field in the case workflow was 2.6%."}, {"speaker": "assay scientist", "text": "The case record is complete for the one-specialist routing decision, with the image review, calibration-log entry, and replicate check retained together."}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_acquisition"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The governing rubric and precedence rules are preserved in both contexts alongside the verbatim questions. The case, routing task, measurements, and workflow timing remain properly bound. Both evidence spans are complete factual sentences. The counterfactual changes only calibration drift from 2.6% to 1.4%, without creating contradictions. Neither context embeds an answer, code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"case note\",\"text\":\"The supplied raw-image review recorded image-confirmed bubble coverage at exactly 3.0%, and the same review recorded saturated pixels at exactly 1.0%. Replicate segmentation checks recorded 5.2% disagreement. These readings were entered for routing without alternate technician measurements.\"},{\"speaker\":\"calibration auditor\",\"text\":\"At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field.\"},{\"speaker\":\"calibration auditor\",\"text\":\"The recorded percentage value for the calibration-drift field in the case workflow was 2.6%.\"},{\"speaker\":\"assay scientist\",\"text\":\"The case record is complete for the one-specialist routing decision, with the image review, calibration-log entry, and replicate check retained together.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field."}, {"path": ["3", "text"], "text": "The recorded percentage value for the calibration-drift field in the case workflow was 2.6%."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field.", "negative_left": "At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field.", "negative_right": "The recorded percentage value for the calibration-drift field in the case workflow was 1.4%.", "right": "The recorded percentage value for the calibration-drift field in the case workflow was 2.6%."}, "verifier_independent_model": false}, "family": "fast-41-diverse-260-013", "id": "fast-41-diverse-260-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "case note", "text": "The supplied raw-image review recorded image-confirmed bubble coverage at exactly 3.0%, and the same review recorded saturated pixels at exactly 1.0%. Replicate segmentation checks recorded 5.2% disagreement. These readings were entered for routing without alternate technician measurements."}, {"speaker": "calibration auditor", "text": "At routing time, recalculation from the calibration log associated with the case workflow produced a single percentage-valued result for the log’s calibration-drift field."}, {"speaker": "calibration auditor", "text": "The recorded percentage value for the calibration-drift field in the case workflow was 1.4%."}, {"speaker": "assay scientist", "text": "The case record is complete for the one-specialist routing decision, with the image review, calibration-log entry, and replicate check retained together."}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain it without inventing exceptions. The entity, workflow path, acceptance decision, and timing bindings remain unchanged. The evidence consists of two complete factual sentences: \"At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17.\" \"At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm.\" The counterfactual changes only the pixel-size measurement to a coherent 0.515 µm value and introduces no contradictory duplicate assertion. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a1 is an allowed universal claim over the explicit three-replicate set. The focus a4 is a factual pixel-size relationship rather than a policy proposition. Base and counter assignments are realizable with only a4 changing: the same calibration image can have an in-tolerance measurement in the base case and an out-of-tolerance measurement in the countercase. The policy evidence preserves the substantive rubric, threshold, and unverified-criterion rule originating in the original state. The synthetic identifier C-17 does not introduce a conflicting policy assumption.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that segmentation and artifact requirements pass and that the calibration image exists, while refutation of a4 establishes that its recorded pixel size is outside the allowed 2% tolerance. Thus a required criterion fails, which is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all three rubric requirements: every replicate meets the 0.90 segmentation threshold, artifact burden is low or moderate, and a submitted calibration image has a recorded pixel size within 2% of 0.50 µm. This is sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At Priya's acceptance decision, every one of Lena's three submitted fluorescent tissue-slide replicates has a segmentation Dice score of at least 0.90."}, {"id": "a2", "statement": "At Priya's acceptance decision, the artifact-burden rating for Lena's submitted three-replicate microscopy workflow belongs to the set {low, moderate}."}, {"id": "a3", "statement": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17."}, {"id": "a4", "statement": "At Priya's acceptance decision, the pixel-size measurement recorded for the submitted calibration image C-17 differs by no more than 2% from the certified 0.50 µm value."}], "base_state_json": "\"Lena submitted three fluorescent tissue-slide replicates for Priya's acceptance review. Omar recorded segmentation Dice scores of 0.93, 0.92, and 0.94, respectively, for the replicates, and the workflow's artifact-burden rating is low. At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17. At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm. The review is conducted under the following QC guidance: the required 0.90 minimum. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist. Priya records the observations and routes any unresolved QC issue according to that rubric.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17."}, {"path": [], "text": "At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm."}], "policy_evidence": [{"path": [], "text": "the required 0.90 minimum"}, {"path": [], "text": "The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value."}, {"path": [], "text": "Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}], "rules": [{"justification": "All other documented criteria pass, but the submitted calibration image's recorded pixel size is explicitly outside the permitted 2% tolerance around the rubric's certified 0.50 µm value, so a required criterion fails.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}, {"justification": "Every submitted replicate meets the required 0.90 segmentation threshold, artifact burden has an allowed rating, and the submitted calibration image records a pixel size within 2% of the expressly required certified 0.50 µm value.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17.", "negative_left": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17.", "negative_right": "At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.515 µm, and the certified value is 0.50 µm.", "right": "At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-261-001", "id": "fast-41-diverse-261-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required criterion fails or cannot be verified from the supplied information.", "true": "Yes — all required acceptance criteria are documented and pass."}, "instructions": "Is this microscopy workflow ready for acceptance under the supplied QC rubric? Answer yes or no.", "type": "noul"}}, "state": "Lena submitted three fluorescent tissue-slide replicates for Priya's acceptance review. Omar recorded segmentation Dice scores of 0.93, 0.92, and 0.94, respectively, for the replicates, and the workflow's artifact-burden rating is low. At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17. At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm. The review is conducted under the following QC guidance: the required 0.90 minimum. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist. Priya records the observations and routes any unresolved QC issue according to that rubric."}, "method": "c2d", "provenance": {"source_id": "diverse-261", "source_is_synthetic": true, "source_sha256": "5342cad74fb77dfe1b1e4d029535fa8b1ca8f1d3c3996871dd4943e57ed0d56b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain it without inventing exceptions. The entity, workflow path, acceptance decision, and timing bindings remain unchanged. The evidence consists of two complete factual sentences: \"At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17.\" \"At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm.\" The counterfactual changes only the pixel-size measurement to a coherent 0.515 µm value and introduces no contradictory duplicate assertion. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a1 is an allowed universal claim over the explicit three-replicate set. The focus a4 is a factual pixel-size relationship rather than a policy proposition. Base and counter assignments are realizable with only a4 changing: the same calibration image can have an in-tolerance measurement in the base case and an out-of-tolerance measurement in the countercase. The policy evidence preserves the substantive rubric, threshold, and unverified-criterion rule originating in the original state. The synthetic identifier C-17 does not introduce a conflicting policy assumption.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that segmentation and artifact requirements pass and that the calibration image exists, while refutation of a4 establishes that its recorded pixel size is outside the allowed 2% tolerance. Thus a required criterion fails, which is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all three rubric requirements: every replicate meets the 0.90 segmentation threshold, artifact burden is low or moderate, and a submitted calibration image has a recorded pixel size within 2% of 0.50 µm. This is sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At Priya's acceptance decision, every one of Lena's three submitted fluorescent tissue-slide replicates has a segmentation Dice score of at least 0.90."}, {"id": "a2", "statement": "At Priya's acceptance decision, the artifact-burden rating for Lena's submitted three-replicate microscopy workflow belongs to the set {low, moderate}."}, {"id": "a3", "statement": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17."}, {"id": "a4", "statement": "At Priya's acceptance decision, the pixel-size measurement recorded for the submitted calibration image C-17 differs by no more than 2% from the certified 0.50 µm value."}], "base_state_json": "\"Lena submitted three fluorescent tissue-slide replicates for Priya's acceptance review. Omar recorded segmentation Dice scores of 0.93, 0.92, and 0.94, respectively, for the replicates, and the workflow's artifact-burden rating is low. At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17. At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm. The review is conducted under the following QC guidance: the required 0.90 minimum. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist. Priya records the observations and routes any unresolved QC issue according to that rubric.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17."}, {"path": [], "text": "At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm."}], "policy_evidence": [{"path": [], "text": "the required 0.90 minimum"}, {"path": [], "text": "The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value."}, {"path": [], "text": "Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist."}], "rules": [{"justification": "All other documented criteria pass, but the submitted calibration image's recorded pixel size is explicitly outside the permitted 2% tolerance around the rubric's certified 0.50 µm value, so a required criterion fails.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}, {"justification": "Every submitted replicate meets the required 0.90 segmentation threshold, artifact burden has an allowed rating, and the submitted calibration image records a pixel size within 2% of the expressly required certified 0.50 µm value.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17.", "negative_left": "At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17.", "negative_right": "At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.515 µm, and the certified value is 0.50 µm.", "right": "At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.495 µm, and the certified value is 0.50 µm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-261-001", "id": "fast-41-diverse-261-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required criterion fails or cannot be verified from the supplied information.", "true": "Yes — all required acceptance criteria are documented and pass."}, "instructions": "Is this microscopy workflow ready for acceptance under the supplied QC rubric? Answer yes or no.", "type": "noul"}}, "state": "Lena submitted three fluorescent tissue-slide replicates for Priya's acceptance review. Omar recorded segmentation Dice scores of 0.93, 0.92, and 0.94, respectively, for the replicates, and the workflow's artifact-burden rating is low. At Priya's acceptance decision, Lena's microscopy workflow submission contains calibration image C-17. At Priya's acceptance decision, the submitted calibration image with the identifier contained in Lena's workflow has a recorded pixel-size measurement of 0.515 µm, and the certified value is 0.50 µm. The review is conducted under the following QC guidance: the required 0.90 minimum. The QC rubric requires all three replicates to pass segmentation, artifact burden to be low or moderate, and a calibration image confirming pixel size within 2% of the certified 0.50 µm value. Under the rubric, an unverified required criterion cannot be treated as passing and must be routed to the acquisition specialist. Priya records the observations and routes any unresolved QC issue according to that rubric."}, "method": "c2d", "provenance": {"source_id": "diverse-261", "source_is_synthetic": true, "source_sha256": "5342cad74fb77dfe1b1e4d029535fa8b1ca8f1d3c3996871dd4943e57ed0d56b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions, while both contexts retain the workflow, replicate C, field F-17, and date bindings. The evidence contains exactly two complete factual sentences: “In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17.” and “In that same signed report, field F-17 is marked “at least 0.90.”” The counterfactual changes only F-17 to “below 0.90,” consistently making replicate C’s focus value below the threshold without duplicate contradictions. Neither context includes an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"workflow record\",\"text\":\"In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17.\"},{\"speaker\":\"workflow record\",\"text\":\"In that same signed report, field F-17 is marked “at least 0.90.”\"},{\"speaker\":\"microscopy technician\",\"text\":\"Replicates A and B recorded focus scores of 0.92 and 0.95. Every required capture check other than the focus-score check passed for scored fluorescence replicates A, B, and C, and illumination drift stayed below 3%.\"},{\"speaker\":\"calibration lead\",\"text\":\"Every required calibration check passed for the equipment used to capture all three scored replicates.\"},{\"speaker\":\"image analyst\",\"text\":\"Every required segmentation check passed for each tissue ROI. The only bright streak was outside the tissue ROI, and every scored replicate therefore has zero counted in-ROI artifacts.\"},{\"speaker\":\"image analyst\",\"text\":\"Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the inclusive 8.0% limit applies to all three replicates.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17."}, {"path": ["1", "text"], "text": "In that same signed report, field F-17 is marked “at least 0.90.”"}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17.", "negative_left": "In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17.", "negative_right": "In that same signed report, field F-17 is marked “below 0.90.”", "right": "In that same signed report, field F-17 is marked “at least 0.90.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-004", "id": "fast-41-diverse-263-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "workflow record", "text": "In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17."}, {"speaker": "workflow record", "text": "In that same signed report, field F-17 is marked “at least 0.90.”"}, {"speaker": "microscopy technician", "text": "Replicates A and B recorded focus scores of 0.92 and 0.95. Every required capture check other than the focus-score check passed for scored fluorescence replicates A, B, and C, and illumination drift stayed below 3%."}, {"speaker": "calibration lead", "text": "Every required calibration check passed for the equipment used to capture all three scored replicates."}, {"speaker": "image analyst", "text": "Every required segmentation check passed for each tissue ROI. The only bright streak was outside the tissue ROI, and every scored replicate therefore has zero counted in-ROI artifacts."}, {"speaker": "image analyst", "text": "Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the inclusive 8.0% limit applies to all three replicates."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions, while both contexts retain the workflow, replicate C, field F-17, and date bindings. The evidence contains exactly two complete factual sentences: “In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17.” and “In that same signed report, field F-17 is marked “at least 0.90.”” The counterfactual changes only F-17 to “below 0.90,” consistently making replicate C’s focus value below the threshold without duplicate contradictions. Neither context includes an answer, code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"workflow record\",\"text\":\"In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17.\"},{\"speaker\":\"workflow record\",\"text\":\"In that same signed report, field F-17 is marked “at least 0.90.”\"},{\"speaker\":\"microscopy technician\",\"text\":\"Replicates A and B recorded focus scores of 0.92 and 0.95. Every required capture check other than the focus-score check passed for scored fluorescence replicates A, B, and C, and illumination drift stayed below 3%.\"},{\"speaker\":\"calibration lead\",\"text\":\"Every required calibration check passed for the equipment used to capture all three scored replicates.\"},{\"speaker\":\"image analyst\",\"text\":\"Every required segmentation check passed for each tissue ROI. The only bright streak was outside the tissue ROI, and every scored replicate therefore has zero counted in-ROI artifacts.\"},{\"speaker\":\"image analyst\",\"text\":\"Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the inclusive 8.0% limit applies to all three replicates.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17."}, {"path": ["1", "text"], "text": "In that same signed report, field F-17 is marked “at least 0.90.”"}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17.", "negative_left": "In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17.", "negative_right": "In that same signed report, field F-17 is marked “below 0.90.”", "right": "In that same signed report, field F-17 is marked “at least 0.90.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-004", "id": "fast-41-diverse-263-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "workflow record", "text": "In the signed fluorescence report for workflow run 2026-09-17, the focus value assigned to scored fluorescence replicate C is the value recorded in field F-17."}, {"speaker": "workflow record", "text": "In that same signed report, field F-17 is marked “below 0.90.”"}, {"speaker": "microscopy technician", "text": "Replicates A and B recorded focus scores of 0.92 and 0.95. Every required capture check other than the focus-score check passed for scored fluorescence replicates A, B, and C, and illumination drift stayed below 3%."}, {"speaker": "calibration lead", "text": "Every required calibration check passed for the equipment used to capture all three scored replicates."}, {"speaker": "image analyst", "text": "Every required segmentation check passed for each tissue ROI. The only bright streak was outside the tissue ROI, and every scored replicate therefore has zero counted in-ROI artifacts."}, {"speaker": "image analyst", "text": "Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the inclusive 8.0% limit applies to all three replicates."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions are verbatim, bindings remain intact, the two evidence quotes are complete factual sentences—\"The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17.\" and \"In the instrument report for this workflow run, measurement M-17 is recorded as 0.93.\"—and the counterfactual coherently changes only C's focus measurement without contradictory duplicates or embedded answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"workflow record\",\"text\":\"During this workflow run, scored fluorescence replicates A and B each recorded focus scores of at least 0.90. All required capture checks other than the focus-score check passed for replicates A, B, and C. Calibration checks passed for the equipment used for all three replicates.\"},{\"speaker\":\"image analyst\",\"text\":\"Segmentation checks passed for every tissue ROI in replicates A, B, and C. Each scored replicate had zero counted artifacts inside its tissue ROI. Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the bright streak was outside the tissue ROI.\"},{\"speaker\":\"instrument report\",\"text\":\"The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17. In the instrument report for this workflow run, measurement M-17 is recorded as 0.93.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["2", "text"], "text": "The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17."}, {"path": ["2", "text"], "text": "In the instrument report for this workflow run, measurement M-17 is recorded as 0.93."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17.", "negative_left": "The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17.", "negative_right": "In the instrument report for this workflow run, measurement M-17 is recorded as 0.87.", "right": "In the instrument report for this workflow run, measurement M-17 is recorded as 0.93."}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-005", "id": "fast-41-diverse-263-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "workflow record", "text": "During this workflow run, scored fluorescence replicates A and B each recorded focus scores of at least 0.90. All required capture checks other than the focus-score check passed for replicates A, B, and C. Calibration checks passed for the equipment used for all three replicates."}, {"speaker": "image analyst", "text": "Segmentation checks passed for every tissue ROI in replicates A, B, and C. Each scored replicate had zero counted artifacts inside its tissue ROI. Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the bright streak was outside the tissue ROI."}, {"speaker": "instrument report", "text": "The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17. In the instrument report for this workflow run, measurement M-17 is recorded as 0.93."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions are verbatim, bindings remain intact, the two evidence quotes are complete factual sentences—\"The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17.\" and \"In the instrument report for this workflow run, measurement M-17 is recorded as 0.93.\"—and the counterfactual coherently changes only C's focus measurement without contradictory duplicates or embedded answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"workflow record\",\"text\":\"During this workflow run, scored fluorescence replicates A and B each recorded focus scores of at least 0.90. All required capture checks other than the focus-score check passed for replicates A, B, and C. Calibration checks passed for the equipment used for all three replicates.\"},{\"speaker\":\"image analyst\",\"text\":\"Segmentation checks passed for every tissue ROI in replicates A, B, and C. Each scored replicate had zero counted artifacts inside its tissue ROI. Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the bright streak was outside the tissue ROI.\"},{\"speaker\":\"instrument report\",\"text\":\"The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17. In the instrument report for this workflow run, measurement M-17 is recorded as 0.93.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["2", "text"], "text": "The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17."}, {"path": ["2", "text"], "text": "In the instrument report for this workflow run, measurement M-17 is recorded as 0.93."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17.", "negative_left": "The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17.", "negative_right": "In the instrument report for this workflow run, measurement M-17 is recorded as 0.87.", "right": "In the instrument report for this workflow run, measurement M-17 is recorded as 0.93."}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-005", "id": "fast-41-diverse-263-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "workflow record", "text": "During this workflow run, scored fluorescence replicates A and B each recorded focus scores of at least 0.90. All required capture checks other than the focus-score check passed for replicates A, B, and C. Calibration checks passed for the equipment used for all three replicates."}, {"speaker": "image analyst", "text": "Segmentation checks passed for every tissue ROI in replicates A, B, and C. Each scored replicate had zero counted artifacts inside its tissue ROI. Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the bright streak was outside the tissue ROI."}, {"speaker": "instrument report", "text": "The calibrated fluorescence reader's report labels the focus score recorded during capture of scored fluorescence replicate C in this workflow run as measurement M-17. In the instrument report for this workflow run, measurement M-17 is recorded as 0.87."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy and criteria. Workflow W-17, replicate C, and the reliability-scoring path remain bound correctly. The evidence consists of two complete factual sentences: \"The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate.\" and \"The focus score in workflow run W-17's capture log is 0.93 and is above 0.90.\" The counterfactual changes only the focus score from 0.93 to 0.87, coherently making it below 0.90. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"case note\",\"text\":\"Workflow run W-17 was reviewed for measurement reliability.\"},{\"speaker\":\"capture log\",\"text\":\"The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate.\"},{\"speaker\":\"capture log\",\"text\":\"The focus score in workflow run W-17's capture log is 0.93 and is above 0.90.\"},{\"speaker\":\"capture log\",\"text\":\"Replicates A and B each passed the focus threshold of 0.90 during capture.\"},{\"speaker\":\"capture log\",\"text\":\"Every required capture check other than the focus-score check passed for scored fluorescence replicates A, B, and C.\"},{\"speaker\":\"equipment log\",\"text\":\"Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C.\"},{\"speaker\":\"segmentation log\",\"text\":\"Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C.\"},{\"speaker\":\"artifact review\",\"text\":\"Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI.\"},{\"speaker\":\"measurement review\",\"text\":\"Cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively, so every replicate was at or below 8.0%.\"},{\"speaker\":\"artifact review\",\"text\":\"A bright streak was outside the tissue ROI and was not counted as an artifact.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["1", "text"], "text": "The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate."}, {"path": ["2", "text"], "text": "The focus score in workflow run W-17's capture log is 0.93 and is above 0.90."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate.", "negative_left": "The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate.", "negative_right": "The focus score in workflow run W-17's capture log is 0.87 and is below 0.90.", "right": "The focus score in workflow run W-17's capture log is 0.93 and is above 0.90."}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-006", "id": "fast-41-diverse-263-006-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "case note", "text": "Workflow run W-17 was reviewed for measurement reliability."}, {"speaker": "capture log", "text": "The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate."}, {"speaker": "capture log", "text": "The focus score in workflow run W-17's capture log is 0.93 and is above 0.90."}, {"speaker": "capture log", "text": "Replicates A and B each passed the focus threshold of 0.90 during capture."}, {"speaker": "capture log", "text": "Every required capture check other than the focus-score check passed for scored fluorescence replicates A, B, and C."}, {"speaker": "equipment log", "text": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C."}, {"speaker": "segmentation log", "text": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C."}, {"speaker": "artifact review", "text": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI."}, {"speaker": "measurement review", "text": "Cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively, so every replicate was at or below 8.0%."}, {"speaker": "artifact review", "text": "A bright streak was outside the tissue ROI and was not counted as an artifact."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy and criteria. Workflow W-17, replicate C, and the reliability-scoring path remain bound correctly. The evidence consists of two complete factual sentences: \"The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate.\" and \"The focus score in workflow run W-17's capture log is 0.93 and is above 0.90.\" The counterfactual changes only the focus score from 0.93 to 0.87, coherently making it below 0.90. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"case note\",\"text\":\"Workflow run W-17 was reviewed for measurement reliability.\"},{\"speaker\":\"capture log\",\"text\":\"The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate.\"},{\"speaker\":\"capture log\",\"text\":\"The focus score in workflow run W-17's capture log is 0.93 and is above 0.90.\"},{\"speaker\":\"capture log\",\"text\":\"Replicates A and B each passed the focus threshold of 0.90 during capture.\"},{\"speaker\":\"capture log\",\"text\":\"Every required capture check other than the focus-score check passed for scored fluorescence replicates A, B, and C.\"},{\"speaker\":\"equipment log\",\"text\":\"Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C.\"},{\"speaker\":\"segmentation log\",\"text\":\"Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C.\"},{\"speaker\":\"artifact review\",\"text\":\"Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI.\"},{\"speaker\":\"measurement review\",\"text\":\"Cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively, so every replicate was at or below 8.0%.\"},{\"speaker\":\"artifact review\",\"text\":\"A bright streak was outside the tissue ROI and was not counted as an artifact.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["1", "text"], "text": "The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate."}, {"path": ["2", "text"], "text": "The focus score in workflow run W-17's capture log is 0.93 and is above 0.90."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate.", "negative_left": "The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate.", "negative_right": "The focus score in workflow run W-17's capture log is 0.87 and is below 0.90.", "right": "The focus score in workflow run W-17's capture log is 0.93 and is above 0.90."}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-006", "id": "fast-41-diverse-263-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "case note", "text": "Workflow run W-17 was reviewed for measurement reliability."}, {"speaker": "capture log", "text": "The capture log for workflow run W-17 labels fluorescence replicate C as the scored replicate."}, {"speaker": "capture log", "text": "The focus score in workflow run W-17's capture log is 0.87 and is below 0.90."}, {"speaker": "capture log", "text": "Replicates A and B each passed the focus threshold of 0.90 during capture."}, {"speaker": "capture log", "text": "Every required capture check other than the focus-score check passed for scored fluorescence replicates A, B, and C."}, {"speaker": "equipment log", "text": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C."}, {"speaker": "segmentation log", "text": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C."}, {"speaker": "artifact review", "text": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI."}, {"speaker": "measurement review", "text": "Cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively, so every replicate was at or below 8.0%."}, {"speaker": "artifact review", "text": "A bright streak was outside the tissue ROI and was not counted as an artifact."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, thresholds, scope, and bindings. Both contexts retain workflow WR-53879, replicates A–C, and the relevant ROI and focus references. The evidence consists of the complete factual sentences “The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units.” and “For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores.” The counterfactual’s 100-unit calibration mapping is internally consistent with replicate C’s 94-unit focus score and does not contradict another measurement. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"microscopy technician\",\"text\":\"Workflow run WR-53879 includes scored fluorescence replicates A, B, and C. All required capture checks other than focus scoring passed for each replicate, and illumination drift stayed below 3%. Equipment calibration was current and passed for the system used. Replicate A's focus score was 0.91, and replicate B's was 0.93.\"},{\"speaker\":\"image analyst\",\"text\":\"Segmentation checks passed for every tissue ROI, with no merged or missing objects. No counted artifacts were found inside the tissue ROI of any scored replicate. Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the inclusive 8.0% limit therefore applies to the complete scored set.\"},{\"speaker\":\"instrument reviewer\",\"text\":\"The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores.\"},{\"speaker\":\"assay scientist\",\"text\":\"A bright streak noted outside the tissue ROI is excluded from the artifact count, and replicate C remains included in scoring. The slide fold is likewise outside every tissue ROI.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["2", "text"], "text": "The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units."}, {"path": ["3", "text"], "text": "For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units.", "negative_left": "The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units.", "negative_right": "For the calibrated focus-score scale used in workflow run WR-53879, 100 scale units correspond to 0.90 and larger unit values represent larger scores.", "right": "For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores."}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-008", "id": "fast-41-diverse-263-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "microscopy technician", "text": "Workflow run WR-53879 includes scored fluorescence replicates A, B, and C. All required capture checks other than focus scoring passed for each replicate, and illumination drift stayed below 3%. Equipment calibration was current and passed for the system used. Replicate A's focus score was 0.91, and replicate B's was 0.93."}, {"speaker": "image analyst", "text": "Segmentation checks passed for every tissue ROI, with no merged or missing objects. No counted artifacts were found inside the tissue ROI of any scored replicate. Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the inclusive 8.0% limit therefore applies to the complete scored set."}, {"speaker": "instrument reviewer", "text": "The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units."}, {"speaker": "calibration reviewer", "text": "For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores."}, {"speaker": "assay scientist", "text": "A bright streak noted outside the tissue ROI is excluded from the artifact count, and replicate C remains included in scoring. The slide fold is likewise outside every tissue ROI."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, thresholds, scope, and bindings. Both contexts retain workflow WR-53879, replicates A–C, and the relevant ROI and focus references. The evidence consists of the complete factual sentences “The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units.” and “For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores.” The counterfactual’s 100-unit calibration mapping is internally consistent with replicate C’s 94-unit focus score and does not contradict another measurement. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"microscopy technician\",\"text\":\"Workflow run WR-53879 includes scored fluorescence replicates A, B, and C. All required capture checks other than focus scoring passed for each replicate, and illumination drift stayed below 3%. Equipment calibration was current and passed for the system used. Replicate A's focus score was 0.91, and replicate B's was 0.93.\"},{\"speaker\":\"image analyst\",\"text\":\"Segmentation checks passed for every tissue ROI, with no merged or missing objects. No counted artifacts were found inside the tissue ROI of any scored replicate. Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the inclusive 8.0% limit therefore applies to the complete scored set.\"},{\"speaker\":\"instrument reviewer\",\"text\":\"The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units.\"},{\"speaker\":\"calibration reviewer\",\"text\":\"For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores.\"},{\"speaker\":\"assay scientist\",\"text\":\"A bright streak noted outside the tissue ROI is excluded from the artifact count, and replicate C remains included in scoring. The slide fold is likewise outside every tissue ROI.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["2", "text"], "text": "The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units."}, {"path": ["3", "text"], "text": "For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units.", "negative_left": "The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units.", "negative_right": "For the calibrated focus-score scale used in workflow run WR-53879, 100 scale units correspond to 0.90 and larger unit values represent larger scores.", "right": "For the calibrated focus-score scale used in workflow run WR-53879, 90 scale units correspond to 0.90 and larger unit values represent larger scores."}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-008", "id": "fast-41-diverse-263-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "microscopy technician", "text": "Workflow run WR-53879 includes scored fluorescence replicates A, B, and C. All required capture checks other than focus scoring passed for each replicate, and illumination drift stayed below 3%. Equipment calibration was current and passed for the system used. Replicate A's focus score was 0.91, and replicate B's was 0.93."}, {"speaker": "image analyst", "text": "Segmentation checks passed for every tissue ROI, with no merged or missing objects. No counted artifacts were found inside the tissue ROI of any scored replicate. Cell-count CVs were 7.6% for A, 7.9% for B, and 8.0% for C; the inclusive 8.0% limit therefore applies to the complete scored set."}, {"speaker": "instrument reviewer", "text": "The instrument log for workflow run WR-53879 records scored fluorescence replicate C's focus score as 94 scale units."}, {"speaker": "calibration reviewer", "text": "For the calibrated focus-score scale used in workflow run WR-53879, 100 scale units correspond to 0.90 and larger unit values represent larger scores."}, {"speaker": "assay scientist", "text": "A bright streak noted outside the tissue ROI is excluded from the artifact count, and replicate C remains included in scoring. The slide fold is likewise outside every tissue ROI."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved by the unchanged questions object and both contexts retain compatible facts. Workflow WR-47, replicate C, and 14 May 2026 bindings are unchanged. The evidence spans are two complete factual sentences: “In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026.” and “The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93.” The counterfactual changes only the focus score to 0.87 without contradictory duplicates. Neither context states a gold answer, answer code, rationale, rule table, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"run coordinator\",\"text\":\"Workflow run WR-47 used scored fluorescence replicates A, B, and C. In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026.\"},{\"speaker\":\"microscopy technician\",\"text\":\"The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93. Replicates A and B each had focus scores at least 0.90. Every required capture check other than the focus-score check passed for A, B, and C, and every required calibration check passed for the equipment.\"},{\"speaker\":\"image analyst\",\"text\":\"Every required segmentation check passed for each tissue ROI in A, B, and C. No counted artifacts were present inside any tissue ROI. The cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively.\"},{\"speaker\":\"assay scientist\",\"text\":\"The bright streak observed near replicate C was outside its tissue ROI, so it was not counted as an in-ROI artifact. The recorded checks and measurements belong to this workflow run.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026."}, {"path": ["1", "text"], "text": "The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026.", "negative_left": "In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026.", "negative_right": "The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.87.", "right": "The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93."}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-011", "id": "fast-41-diverse-263-011-base", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "run coordinator", "text": "Workflow run WR-47 used scored fluorescence replicates A, B, and C. In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026."}, {"speaker": "microscopy technician", "text": "The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93. Replicates A and B each had focus scores at least 0.90. Every required capture check other than the focus-score check passed for A, B, and C, and every required calibration check passed for the equipment."}, {"speaker": "image analyst", "text": "Every required segmentation check passed for each tissue ROI in A, B, and C. No counted artifacts were present inside any tissue ROI. The cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively."}, {"speaker": "assay scientist", "text": "The bright streak observed near replicate C was outside its tissue ROI, so it was not counted as an in-ROI artifact. The recorded checks and measurements belong to this workflow run."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved by the unchanged questions object and both contexts retain compatible facts. Workflow WR-47, replicate C, and 14 May 2026 bindings are unchanged. The evidence spans are two complete factual sentences: “In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026.” and “The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93.” The counterfactual changes only the focus score to 0.87 without contradictory duplicates. Neither context states a gold answer, answer code, rationale, rule table, proposition ID, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "full_context_fact_states": {"base": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "supported", "other_capture": "supported", "segmentation": "supported"}, "counterfactual": {"artifacts": "supported", "calibration": "supported", "cv_limit": "supported", "focus_a": "supported", "focus_b": "supported", "focus_c": "refuted", "other_capture": "supported", "segmentation": "supported"}, "remove_left": {"focus_c": "unknown"}, "remove_right": {"focus_c": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"focus_c": "unknown"}, "negative_pair": {"focus_c": "refuted"}, "negative_sentence": {"focus_c": "unknown"}, "positive_pair": {"focus_c": "supported"}, "right": {"focus_c": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships and do not encode final classifications. The focus atom is a factual measurement threshold. The base and counter assignments are jointly realizable with only replicate C's focus status changing because other_capture explicitly excludes focus. Empty policy_evidence is correct: the governing criteria, thresholds, scope, and pronoun bindings are already retained in the questions object, while the original state supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting focus_c entails replicate C's focus score is below 0.90. The criteria make any scored replicate below 0.90 independently sufficient for level 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all required capture checks, including focus for A, B, and C, pass; calibration and segmentation pass; no scored replicate has a counted in-ROI artifact; and every CV is at most the inclusive 8.0% threshold. These conditions are sufficient for level 3 and exclude the stated lower-level triggers.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "focus_c", "statement": "The focus score recorded during capture of scored fluorescence replicate C in this workflow run is at least 0.90."}, {"id": "focus_a", "statement": "The focus score recorded during capture of scored fluorescence replicate A in this workflow run is at least 0.90."}, {"id": "focus_b", "statement": "The focus score recorded during capture of scored fluorescence replicate B in this workflow run is at least 0.90."}, {"id": "other_capture", "statement": "Every required capture check other than the focus-score check passed for each of scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "calibration", "statement": "Every required calibration check passed for the equipment used to capture scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "segmentation", "statement": "Every required segmentation check passed for each tissue ROI in scored fluorescence replicates A, B, and C in this workflow run."}, {"id": "artifacts", "statement": "Every scored fluorescence replicate among A, B, and C has zero counted artifacts inside its tissue ROI in this workflow run."}, {"id": "cv_limit", "statement": "Every scored fluorescence replicate among A, B, and C has a cell-count CV no greater than 8.0% in this workflow run."}], "base_state_json": "[{\"speaker\":\"run coordinator\",\"text\":\"Workflow run WR-47 used scored fluorescence replicates A, B, and C. In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026.\"},{\"speaker\":\"microscopy technician\",\"text\":\"The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93. Replicates A and B each had focus scores at least 0.90. Every required capture check other than the focus-score check passed for A, B, and C, and every required calibration check passed for the equipment.\"},{\"speaker\":\"image analyst\",\"text\":\"Every required segmentation check passed for each tissue ROI in A, B, and C. No counted artifacts were present inside any tissue ROI. The cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively.\"},{\"speaker\":\"assay scientist\",\"text\":\"The bright streak observed near replicate C was outside its tissue ROI, so it was not counted as an in-ROI artifact. The recorded checks and measurements belong to this workflow run.\"}]", "base_states": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "counter_states": [{"atom_id": "focus_c", "state": "refuted"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}], "focus_atom": "focus_c", "focus_evidence": [{"path": ["0", "text"], "text": "In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026."}, {"path": ["1", "text"], "text": "The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93."}], "policy_evidence": [], "rules": [{"justification": "Explicit refutation of a focus score of at least 0.90 entails that scored replicate C has focus below 0.90, which is independently sufficient for level 0.", "target": "0", "when": [{"atom_id": "focus_c", "state": "refuted"}]}, {"justification": "All capture, calibration, and segmentation requirements pass; no scored replicate has a counted in-ROI artifact; and every scored replicate has cell-count CV at or below the inclusive 8.0% threshold, satisfying level 3.", "target": "3", "when": [{"atom_id": "focus_c", "state": "supported"}, {"atom_id": "focus_a", "state": "supported"}, {"atom_id": "focus_b", "state": "supported"}, {"atom_id": "other_capture", "state": "supported"}, {"atom_id": "calibration", "state": "supported"}, {"atom_id": "segmentation", "state": "supported"}, {"atom_id": "artifacts", "state": "supported"}, {"atom_id": "cv_limit", "state": "supported"}]}]}, "verified_pair": {"left": "In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026.", "negative_left": "In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026.", "negative_right": "The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.87.", "right": "The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.93."}, "verifier_independent_model": false}, "family": "fast-41-diverse-263-011", "id": "fast-41-diverse-263-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unreliable: Reject when calibration is missing or incorrect, any scored replicate has focus below 0.90 or illumination drift above 3%, or segmentation failures inside an ROI prevent a valid cell count.", "1 — Low reliability: Required capture and calibration checks pass, but at least one scored replicate has a counted in-ROI artifact or segmentation defect, or replicate cell-count CV exceeds 12.0%.", "2 — Moderate reliability: Capture, calibration, and segmentation checks pass with no counted artifacts, and every replicate cell-count CV is above 8.0% but no greater than 12.0%.", "3 — High reliability: Capture, calibration, and segmentation checks pass; no scored replicate has a counted in-ROI artifact; and every replicate cell-count CV is 8.0% or lower. Values exactly equal to 8.0% qualify."], "instructions": "Assign the workflow's measurement-reliability level. Apply thresholds inclusively. A bright streak outside the tissue ROI is not a counted artifact. Resolve the assay scientist's pronouns using the preceding turn: “it” denotes the bright streak, while “that replicate” denotes replicate C.", "type": "score"}}, "state": [{"speaker": "run coordinator", "text": "Workflow run WR-47 used scored fluorescence replicates A, B, and C. In workflow run WR-47, the scored fluorescence capture event labeled replicate C occurred on 14 May 2026."}, {"speaker": "microscopy technician", "text": "The focus-score field for the scored fluorescence capture event in workflow run WR-47 on 14 May 2026 is recorded as 0.87. Replicates A and B each had focus scores at least 0.90. Every required capture check other than the focus-score check passed for A, B, and C, and every required calibration check passed for the equipment."}, {"speaker": "image analyst", "text": "Every required segmentation check passed for each tissue ROI in A, B, and C. No counted artifacts were present inside any tissue ROI. The cell-count CVs for A, B, and C were 7.6%, 7.9%, and 8.0%, respectively."}, {"speaker": "assay scientist", "text": "The bright streak observed near replicate C was outside its tissue ROI, so it was not counted as an in-ROI artifact. The recorded checks and measurements belong to this workflow run."}]}, "method": "c2d", "provenance": {"source_id": "diverse-263", "source_is_synthetic": true, "source_sha256": "f4f6dcfec51cce24558cfe5eaf58a02563313d01d6602a470de1bfb26b14c287", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric, and both contexts retain the calibration policy without inventing exceptions. The workflow, measurement, calibration-record, capture-session, and time bindings remain consistent. The two evidence spans are complete factual sentences: \"The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum.\" and \"The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%.\" The counterfactual’s scale error is 6%, which is coherent with its unchanged 10% maximum and does not trigger the stated above-10% gate. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"Case note: The microscopy workflow's capture session includes its mandatory calibration record. The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum. The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators. The prepared slide has no major preparation defect, the capture has no major defect, and segmentation has no major defect materially biasing measurements. Image artifacts are negligible. Segmentation meets validation criteria, and replicate measurements are consistent. These observations concern the same microscopy workflow and its quantitative measurements.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum."}, {"path": [], "text": "The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum.", "negative_left": "The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points less than that record’s permitted maximum.", "negative_right": "The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%.", "right": "The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%."}, "verifier_independent_model": false}, "family": "fast-41-diverse-264-011", "id": "fast-41-diverse-264-011-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "Case note: The microscopy workflow's capture session includes its mandatory calibration record. The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum. The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators. The prepared slide has no major preparation defect, the capture has no major defect, and segmentation has no major defect materially biasing measurements. Image artifacts are negligible. Segmentation meets validation criteria, and replicate measurements are consistent. These observations concern the same microscopy workflow and its quantitative measurements."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric, and both contexts retain the calibration policy without inventing exceptions. The workflow, measurement, calibration-record, capture-session, and time bindings remain consistent. The two evidence spans are complete factual sentences: \"The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum.\" and \"The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%.\" The counterfactual’s scale error is 6%, which is coherent with its unchanged 10% maximum and does not trigger the stated above-10% gate. Neither context embeds a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"Case note: The microscopy workflow's capture session includes its mandatory calibration record. The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum. The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators. The prepared slide has no major preparation defect, the capture has no major defect, and segmentation has no major defect materially biasing measurements. Image artifacts are negligible. Segmentation meets validation criteria, and replicate measurements are consistent. These observations concern the same microscopy workflow and its quantitative measurements.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum."}, {"path": [], "text": "The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points greater than that record’s permitted maximum.", "negative_left": "The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points less than that record’s permitted maximum.", "negative_right": "The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%.", "right": "The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%."}, "verifier_independent_model": false}, "family": "fast-41-diverse-264-011", "id": "fast-41-diverse-264-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "Case note: The microscopy workflow's capture session includes its mandatory calibration record. The scale-error reading of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session is four percentage points less than that record’s permitted maximum. The permitted maximum for that scale-error reading in the microscopy workflow’s capture session is 10%. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators. The prepared slide has no major preparation defect, the capture has no major defect, and segmentation has no major defect materially biasing measurements. Image artifacts are negligible. Segmentation meets validation criteria, and replicate measurements are consistent. These observations concern the same microscopy workflow and its quantitative measurements."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve the policy, scope, and bindings; both contexts are coherent, the two quoted evidence items are complete factual sentences, and neither context leaks an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"Case note: The study under review is a moisture microcosm incubated for 14 days. Every jar began between 58% and 62% water-holding capacity. CO₂ readings were logged on days 3, 7, and 14. The chamber log identifies exactly one consecutive temperature excursion outside 19°C to 21°C; it affected every treatment and paired control, and temperatures were within that range at all other times. The protocol record states:\\n\\nThe protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\\n\\nIncubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\\n\\nFor the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026.\\nFor that same single temperature excursion, the recorded end was 10:25 UTC on 4 March 2026.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "For the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026."}, {"path": ["context"], "text": "For that same single temperature excursion, the recorded end was 10:25 UTC on 4 March 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026.", "negative_left": "For the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026.", "negative_right": "For that same single temperature excursion, the recorded end was 11:05 UTC on 4 March 2026.", "right": "For that same single temperature excursion, the recorded end was 10:25 UTC on 4 March 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-267-001", "id": "fast-41-diverse-267-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "Case note: The study under review is a moisture microcosm incubated for 14 days. Every jar began between 58% and 62% water-holding capacity. CO₂ readings were logged on days 3, 7, and 14. The chamber log identifies exactly one consecutive temperature excursion outside 19°C to 21°C; it affected every treatment and paired control, and temperatures were within that range at all other times. The protocol record states:\n\nThe protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\n\nIncubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\n\nFor the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026.\nFor that same single temperature excursion, the recorded end was 10:25 UTC on 4 March 2026."}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve the policy, scope, and bindings; both contexts are coherent, the two quoted evidence items are complete factual sentences, and neither context leaks an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"Case note: The study under review is a moisture microcosm incubated for 14 days. Every jar began between 58% and 62% water-holding capacity. CO₂ readings were logged on days 3, 7, and 14. The chamber log identifies exactly one consecutive temperature excursion outside 19°C to 21°C; it affected every treatment and paired control, and temperatures were within that range at all other times. The protocol record states:\\n\\nThe protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\\n\\nIncubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\\n\\nFor the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026.\\nFor that same single temperature excursion, the recorded end was 10:25 UTC on 4 March 2026.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "For the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026."}, {"path": ["context"], "text": "For that same single temperature excursion, the recorded end was 10:25 UTC on 4 March 2026."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026.", "negative_left": "For the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026.", "negative_right": "For that same single temperature excursion, the recorded end was 11:05 UTC on 4 March 2026.", "right": "For that same single temperature excursion, the recorded end was 10:25 UTC on 4 March 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-267-001", "id": "fast-41-diverse-267-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "Case note: The study under review is a moisture microcosm incubated for 14 days. Every jar began between 58% and 62% water-holding capacity. CO₂ readings were logged on days 3, 7, and 14. The chamber log identifies exactly one consecutive temperature excursion outside 19°C to 21°C; it affected every treatment and paired control, and temperatures were within that range at all other times. The protocol record states:\n\nThe protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\n\nIncubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\n\nFor the single temperature excursion in the experiment under review, the recorded start was 09:10 UTC on 4 March 2026.\nFor that same single temperature excursion, the recorded end was 11:05 UTC on 4 March 2026."}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision scope. The experiment, protocol, duration, and measurement bindings remain unchanged. The focus evidence consists of two complete factual sentences: \"For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.\" and \"For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC.\" The counterfactual endpoint produces a coherent 105-minute single excursion without contradictory duplicate assertions. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"Case note: The reviewed study is a moisture microcosm incubated for 14 days. Each jar began within the specified water-holding-capacity band. Chamber records show CO₂ observations on days 3, 7, and 14. The logger identifies exactly one consecutive temperature excursion during incubation; it affected every treatment and paired control. At all other times, the incubation temperature remained within the permitted range. The review record preserves the following event timestamps. The protocol's stated exception is therefore evaluated against the logged event and the study's other observations.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"The experiment under review is a moisture microcosm.\",\"The experiment under review has an incubation duration of 14 days.\",\"Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive.\",\"A CO₂ measurement for the experiment under review was recorded on day 3.\",\"A CO₂ measurement for the experiment under review was recorded on day 7.\",\"A CO₂ measurement for the experiment under review was recorded on day 14.\",\"Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review.\",\"The single temperature excursion during the experiment under review affected every treatment and paired control.\",\"At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive.\",\"For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.\",\"For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "11"], "text": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC."}, {"path": ["evidence", "12"], "text": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.", "negative_left": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.", "negative_right": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 12:00 UTC.", "right": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-267-005", "id": "fast-41-diverse-267-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "Case note: The reviewed study is a moisture microcosm incubated for 14 days. Each jar began within the specified water-holding-capacity band. Chamber records show CO₂ observations on days 3, 7, and 14. The logger identifies exactly one consecutive temperature excursion during incubation; it affected every treatment and paired control. At all other times, the incubation temperature remained within the permitted range. The review record preserves the following event timestamps. The protocol's stated exception is therefore evaluated against the logged event and the study's other observations.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "The experiment under review is a moisture microcosm.", "The experiment under review has an incubation duration of 14 days.", "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive.", "A CO₂ measurement for the experiment under review was recorded on day 3.", "A CO₂ measurement for the experiment under review was recorded on day 7.", "A CO₂ measurement for the experiment under review was recorded on day 14.", "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review.", "The single temperature excursion during the experiment under review affected every treatment and paired control.", "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive.", "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.", "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision scope. The experiment, protocol, duration, and measurement bindings remain unchanged. The focus evidence consists of two complete factual sentences: \"For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.\" and \"For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC.\" The counterfactual endpoint produces a coherent 105-minute single excursion without contradictory duplicate assertions. Neither context embeds a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"Case note: The reviewed study is a moisture microcosm incubated for 14 days. Each jar began within the specified water-holding-capacity band. Chamber records show CO₂ observations on days 3, 7, and 14. The logger identifies exactly one consecutive temperature excursion during incubation; it affected every treatment and paired control. At all other times, the incubation temperature remained within the permitted range. The review record preserves the following event timestamps. The protocol's stated exception is therefore evaluated against the logged event and the study's other observations.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"The experiment under review is a moisture microcosm.\",\"The experiment under review has an incubation duration of 14 days.\",\"Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive.\",\"A CO₂ measurement for the experiment under review was recorded on day 3.\",\"A CO₂ measurement for the experiment under review was recorded on day 7.\",\"A CO₂ measurement for the experiment under review was recorded on day 14.\",\"Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review.\",\"The single temperature excursion during the experiment under review affected every treatment and paired control.\",\"At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive.\",\"For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.\",\"For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "11"], "text": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC."}, {"path": ["evidence", "12"], "text": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.", "negative_left": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.", "negative_right": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 12:00 UTC.", "right": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 11:40 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-267-005", "id": "fast-41-diverse-267-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "Case note: The reviewed study is a moisture microcosm incubated for 14 days. Each jar began within the specified water-holding-capacity band. Chamber records show CO₂ observations on days 3, 7, and 14. The logger identifies exactly one consecutive temperature excursion during incubation; it affected every treatment and paired control. At all other times, the incubation temperature remained within the permitted range. The review record preserves the following event timestamps. The protocol's stated exception is therefore evaluated against the logged event and the study's other observations.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "The experiment under review is a moisture microcosm.", "The experiment under review has an incubation duration of 14 days.", "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive.", "A CO₂ measurement for the experiment under review was recorded on day 3.", "A CO₂ measurement for the experiment under review was recorded on day 7.", "A CO₂ measurement for the experiment under review was recorded on day 14.", "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review.", "The single temperature excursion during the experiment under review affected every treatment and paired control.", "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive.", "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-03 10:15 UTC.", "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-03 12:00 UTC."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the protocol, exception, scope, and unchanged question. The evidence consists of two complete factual sentences: \"For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC.\" and \"For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC.\" The counterfactual instead ends at 11:05 UTC, forming one coherent 105-minute excursion without duplicate measurements. Neither context embeds a gold answer, label code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"A soil scientist is reviewing a 14-day moisture microcosm replication. The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14. Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure. Every jar began between 58% and 62% water-holding capacity, inclusive. CO₂ readings were recorded on days 3, 7, and 14. The chamber log identifies exactly one excursion outside 19°C to 21°C, affecting every treatment and paired control; at all other times, temperature remained within that range. For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC. For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC.\",\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC."}, {"path": ["context"], "text": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC.", "negative_left": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC.", "negative_right": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 11:05 UTC.", "right": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-267-007", "id": "fast-41-diverse-267-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a 14-day moisture microcosm replication. The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14. Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure. Every jar began between 58% and 62% water-holding capacity, inclusive. CO₂ readings were recorded on days 3, 7, and 14. The chamber log identifies exactly one excursion outside 19°C to 21°C, affecting every treatment and paired control; at all other times, temperature remained within that range. For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC. For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC.", "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the protocol, exception, scope, and unchanged question. The evidence consists of two complete factual sentences: \"For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC.\" and \"For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC.\" The counterfactual instead ends at 11:05 UTC, forming one coherent 105-minute excursion without duplicate measurements. Neither context embeds a gold answer, label code, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"A soil scientist is reviewing a 14-day moisture microcosm replication. The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14. Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure. Every jar began between 58% and 62% water-holding capacity, inclusive. CO₂ readings were recorded on days 3, 7, and 14. The chamber log identifies exactly one excursion outside 19°C to 21°C, affecting every treatment and paired control; at all other times, temperature remained within that range. For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC. For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC.\",\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC."}, {"path": ["context"], "text": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC.", "negative_left": "For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC.", "negative_right": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 11:05 UTC.", "right": "For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 10:35 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-267-007", "id": "fast-41-diverse-267-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a 14-day moisture microcosm replication. The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14. Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure. Every jar began between 58% and 62% water-holding capacity, inclusive. CO₂ readings were recorded on days 3, 7, and 14. The chamber log identifies exactly one excursion outside 19°C to 21°C, affecting every treatment and paired control; at all other times, temperature remained within that range. For the experiment under review, the recorded start of the single temperature excursion was 2026-04-11 09:20 UTC. For the experiment under review, the recorded end of the single temperature excursion was 2026-04-11 11:05 UTC.", "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing routing policy without adding exceptions or defaults. Cedar remains the sealed test jar and the incubation-duration binding is unchanged. The two evidence spans are complete factual sentences: \"Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026.\" and \"Cedar was corrected at 08:00 on 13 June 2026.\" The counterfactual changes only the correction time to 11 June 2026 and remains temporally coherent. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"soil scientist\",\"text\":\"The decomposition trial contained exactly two jars, Cedar and Elm. Cedar was the sealed test jar; Elm was excluded from replication judgments. Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. It began with that first low reading and ended when Cedar was corrected. Cedar was corrected at 08:00 on 13 June 2026.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed. Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["0", "text"], "text": "Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026."}, {"path": ["1", "text"], "text": "Cedar was corrected at 08:00 on 13 June 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026.", "negative_left": "Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026.", "negative_right": "Cedar was corrected at 08:00 on 11 June 2026.", "right": "Cedar was corrected at 08:00 on 13 June 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-268-005", "id": "fast-41-diverse-268-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "The decomposition trial contained exactly two jars, Cedar and Elm. Cedar was the sealed test jar; Elm was excluded from replication judgments. Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026."}, {"speaker": "microcosm technician", "text": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. It began with that first low reading and ended when Cedar was corrected. Cedar was corrected at 08:00 on 13 June 2026."}, {"speaker": "quality reviewer", "text": "Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed. Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing routing policy without adding exceptions or defaults. Cedar remains the sealed test jar and the incubation-duration binding is unchanged. The two evidence spans are complete factual sentences: \"Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026.\" and \"Cedar was corrected at 08:00 on 13 June 2026.\" The counterfactual changes only the correction time to 11 June 2026 and remains temporally coherent. Neither context embeds an answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"soil scientist\",\"text\":\"The decomposition trial contained exactly two jars, Cedar and Elm. Cedar was the sealed test jar; Elm was excluded from replication judgments. Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. It began with that first low reading and ended when Cedar was corrected. Cedar was corrected at 08:00 on 13 June 2026.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed. Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["0", "text"], "text": "Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026."}, {"path": ["1", "text"], "text": "Cedar was corrected at 08:00 on 13 June 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026.", "negative_left": "Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026.", "negative_right": "Cedar was corrected at 08:00 on 11 June 2026.", "right": "Cedar was corrected at 08:00 on 13 June 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-268-005", "id": "fast-41-diverse-268-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "The decomposition trial contained exactly two jars, Cedar and Elm. Cedar was the sealed test jar; Elm was excluded from replication judgments. Cedar's first low moisture reading on day 2 was recorded at 08:00 on 10 June 2026."}, {"speaker": "microcosm technician", "text": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation. It began with that first low reading and ended when Cedar was corrected. Cedar was corrected at 08:00 on 11 June 2026."}, {"speaker": "quality reviewer", "text": "Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed. Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric, and both contexts retain it without added exceptions or defaults. Cedar and the incubation-routing scope remain bound to the original question. The evidence spans are complete factual sentences: \"Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026.\" and \"Cedar was corrected at 08:15 on 7 May 2026.\" The counterfactual changes the correction time to 6 May, yielding a coherent 48-hour episode with no contradictory duplicate assertion. Neither context embeds an answer, label rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"soil scientist\",\"text\":\"The decomposition trial contains exactly two jars, Cedar and Elm. Cedar is the sealed test jar, while Elm is excluded from replication judgments.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026. Cedar was corrected at 08:15 on 7 May 2026. During incubation, Cedar had exactly one continuous episode outside 55–65% water-holding capacity; it began with that first low reading and ended when Cedar was corrected.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026."}, {"path": ["1", "text"], "text": "Cedar was corrected at 08:15 on 7 May 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026.", "negative_left": "Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026.", "negative_right": "Cedar was corrected at 08:15 on 6 May 2026.", "right": "Cedar was corrected at 08:15 on 7 May 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-268-010", "id": "fast-41-diverse-268-010-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "The decomposition trial contains exactly two jars, Cedar and Elm. Cedar is the sealed test jar, while Elm is excluded from replication judgments."}, {"speaker": "microcosm technician", "text": "Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026. Cedar was corrected at 08:15 on 7 May 2026. During incubation, Cedar had exactly one continuous episode outside 55–65% water-holding capacity; it began with that first low reading and ended when Cedar was corrected."}, {"speaker": "quality reviewer", "text": "Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing rubric, and both contexts retain it without added exceptions or defaults. Cedar and the incubation-routing scope remain bound to the original question. The evidence spans are complete factual sentences: \"Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026.\" and \"Cedar was corrected at 08:15 on 7 May 2026.\" The counterfactual changes the correction time to 6 May, yielding a coherent 48-hour episode with no contradictory duplicate assertion. Neither context embeds an answer, label rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"soil scientist\",\"text\":\"The decomposition trial contains exactly two jars, Cedar and Elm. Cedar is the sealed test jar, while Elm is excluded from replication judgments.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026. Cedar was corrected at 08:15 on 7 May 2026. During incubation, Cedar had exactly one continuous episode outside 55–65% water-holding capacity; it began with that first low reading and ended when Cedar was corrected.\"},{\"speaker\":\"quality reviewer\",\"text\":\"Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026."}, {"path": ["1", "text"], "text": "Cedar was corrected at 08:15 on 7 May 2026."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026.", "negative_left": "Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026.", "negative_right": "Cedar was corrected at 08:15 on 6 May 2026.", "right": "Cedar was corrected at 08:15 on 7 May 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-268-010", "id": "fast-41-diverse-268-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "The decomposition trial contains exactly two jars, Cedar and Elm. Cedar is the sealed test jar, while Elm is excluded from replication judgments."}, {"speaker": "microcosm technician", "text": "Cedar's first low moisture reading on day 2 was recorded at 08:15 on 4 May 2026. Cedar was corrected at 08:15 on 6 May 2026. During incubation, Cedar had exactly one continuous episode outside 55–65% water-holding capacity; it began with that first low reading and ended when Cedar was corrected."}, {"speaker": "quality reviewer", "text": "Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and decision criteria. Entity, path, and time bindings remain aligned with Cedar and the technician evidence. The evidence consists of two complete factual sentences: “Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC.” and “Cedar was corrected on 5 June 2026 at 08:00 UTC.” The counterfactual’s 4 June correction coherently shortens the episode without contradictory duplicate assertions. Neither context states a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"soil scientist\",\"text\":\"The decomposition trial contains exactly two jars, Cedar and Elm. Cedar is a sealed test jar, while Elm is excluded from replication judgments.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC. Cedar was corrected on 5 June 2026 at 08:00 UTC. Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation; it began with that first low reading and ended at correction.\"},{\"speaker\":\"instrument operator\",\"text\":\"Every sealing check for Cedar during the decomposition trial passed, and every instrument check relevant to Cedar's moisture readings during the decomposition trial passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC."}, {"path": ["1", "text"], "text": "Cedar was corrected on 5 June 2026 at 08:00 UTC."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC.", "negative_left": "Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC.", "negative_right": "Cedar was corrected on 4 June 2026 at 08:00 UTC.", "right": "Cedar was corrected on 5 June 2026 at 08:00 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-268-011", "id": "fast-41-diverse-268-011-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "The decomposition trial contains exactly two jars, Cedar and Elm. Cedar is a sealed test jar, while Elm is excluded from replication judgments."}, {"speaker": "microcosm technician", "text": "Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC. Cedar was corrected on 5 June 2026 at 08:00 UTC. Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation; it began with that first low reading and ended at correction."}, {"speaker": "instrument operator", "text": "Every sealing check for Cedar during the decomposition trial passed, and every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy and decision criteria. Entity, path, and time bindings remain aligned with Cedar and the technician evidence. The evidence consists of two complete factual sentences: “Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC.” and “Cedar was corrected on 5 June 2026 at 08:00 UTC.” The counterfactual’s 4 June correction coherently shortens the episode without contradictory duplicate assertions. Neither context states a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"soil scientist\",\"text\":\"The decomposition trial contains exactly two jars, Cedar and Elm. Cedar is a sealed test jar, while Elm is excluded from replication judgments.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC. Cedar was corrected on 5 June 2026 at 08:00 UTC. Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation; it began with that first low reading and ended at correction.\"},{\"speaker\":\"instrument operator\",\"text\":\"Every sealing check for Cedar during the decomposition trial passed, and every instrument check relevant to Cedar's moisture readings during the decomposition trial passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC."}, {"path": ["1", "text"], "text": "Cedar was corrected on 5 June 2026 at 08:00 UTC."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC.", "negative_left": "Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC.", "negative_right": "Cedar was corrected on 4 June 2026 at 08:00 UTC.", "right": "Cedar was corrected on 5 June 2026 at 08:00 UTC."}, "verifier_independent_model": false}, "family": "fast-41-diverse-268-011", "id": "fast-41-diverse-268-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "The decomposition trial contains exactly two jars, Cedar and Elm. Cedar is a sealed test jar, while Elm is excluded from replication judgments."}, {"speaker": "microcosm technician", "text": "Cedar's first low moisture reading on day 2 was recorded on 2 June 2026 at 08:00 UTC. Cedar was corrected on 4 June 2026 at 08:00 UTC. Cedar had exactly one continuous episode outside 55–65% water-holding capacity during incubation; it began with that first low reading and ended at correction."}, {"speaker": "instrument operator", "text": "Every sealing check for Cedar during the decomposition trial passed, and every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing criteria and instructions, and both contexts retain the same policy. The two focus evidence spans are complete factual sentences: \"The physical item identifier observed for shipment R-184 is V40B.\" and \"PO-771 specifies V40 as the item identifier for shipment R-184.\" The bindings to dock 3, R-184, PO-771, ASN-771A, and immediate routing are unchanged. Changing the physical identifier to V40 makes the counterfactual consistent with the PO and the stated matching ASN and quantity records. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction beyond the preserved request.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"At dock 3, a receiving clerk is checking shipment R-184 against PO-771 and ASN-771A before posting the receipt. The ASN and PO carry matching item entries and matching quantity entries for this shipment. The count sheet records 480 units, consistent with those records. Inspection reports no damage and no safety concern at dock 3. PO-771 was released on 6 May, before the applicable substitution threshold, and Inventory Control has not verified supplier fault. The clerk is using the following governing material. Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope. Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control. Select the single immediate routing classification for this shipment.\",\"evidence\":[\"The physical item identifier observed for shipment R-184 is V40B.\",\"PO-771 specifies V40 as the item identifier for shipment R-184.\",\"The ASN and PO carry matching item entries and matching quantity entries for this shipment.\",\"The count sheet records 480 units, consistent with those records.\",\"Inspection reports no damage and no safety concern at dock 3.\",\"PO-771 was released on 6 May, before the applicable substitution threshold, and Inventory Control has not verified supplier fault.\",\"Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\",\"Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\",\"Select the single immediate routing classification for this shipment.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The physical item identifier observed for shipment R-184 is V40B."}, {"path": ["evidence", "1"], "text": "PO-771 specifies V40 as the item identifier for shipment R-184."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "The physical item identifier observed for shipment R-184 is V40B.", "negative_left": "The physical item identifier observed for shipment R-184 is V40.", "negative_right": "PO-771 specifies V40 as the item identifier for shipment R-184.", "right": "PO-771 specifies V40 as the item identifier for shipment R-184."}, "verifier_independent_model": false}, "family": "fast-41-diverse-277-002", "id": "fast-41-diverse-277-002-base", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "At dock 3, a receiving clerk is checking shipment R-184 against PO-771 and ASN-771A before posting the receipt. The ASN and PO carry matching item entries and matching quantity entries for this shipment. The count sheet records 480 units, consistent with those records. Inspection reports no damage and no safety concern at dock 3. PO-771 was released on 6 May, before the applicable substitution threshold, and Inventory Control has not verified supplier fault. The clerk is using the following governing material. Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope. Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control. Select the single immediate routing classification for this shipment.", "evidence": ["The physical item identifier observed for shipment R-184 is V40B.", "PO-771 specifies V40 as the item identifier for shipment R-184.", "The ASN and PO carry matching item entries and matching quantity entries for this shipment.", "The count sheet records 480 units, consistent with those records.", "Inspection reports no damage and no safety concern at dock 3.", "PO-771 was released on 6 May, before the applicable substitution threshold, and Inventory Control has not verified supplier fault.", "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.", "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.", "Select the single immediate routing classification for this shipment."]}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "INVENTORY_IDENTITY_HOLD"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing criteria and instructions, and both contexts retain the same policy. The two focus evidence spans are complete factual sentences: \"The physical item identifier observed for shipment R-184 is V40B.\" and \"PO-771 specifies V40 as the item identifier for shipment R-184.\" The bindings to dock 3, R-184, PO-771, ASN-771A, and immediate routing are unchanged. Changing the physical identifier to V40 makes the counterfactual consistent with the PO and the stated matching ASN and quantity records. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction beyond the preserved request.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"At dock 3, a receiving clerk is checking shipment R-184 against PO-771 and ASN-771A before posting the receipt. The ASN and PO carry matching item entries and matching quantity entries for this shipment. The count sheet records 480 units, consistent with those records. Inspection reports no damage and no safety concern at dock 3. PO-771 was released on 6 May, before the applicable substitution threshold, and Inventory Control has not verified supplier fault. The clerk is using the following governing material. Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope. Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control. Select the single immediate routing classification for this shipment.\",\"evidence\":[\"The physical item identifier observed for shipment R-184 is V40B.\",\"PO-771 specifies V40 as the item identifier for shipment R-184.\",\"The ASN and PO carry matching item entries and matching quantity entries for this shipment.\",\"The count sheet records 480 units, consistent with those records.\",\"Inspection reports no damage and no safety concern at dock 3.\",\"PO-771 was released on 6 May, before the applicable substitution threshold, and Inventory Control has not verified supplier fault.\",\"Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\",\"Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\",\"Select the single immediate routing classification for this shipment.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The physical item identifier observed for shipment R-184 is V40B."}, {"path": ["evidence", "1"], "text": "PO-771 specifies V40 as the item identifier for shipment R-184."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "The physical item identifier observed for shipment R-184 is V40B.", "negative_left": "The physical item identifier observed for shipment R-184 is V40.", "negative_right": "PO-771 specifies V40 as the item identifier for shipment R-184.", "right": "PO-771 specifies V40 as the item identifier for shipment R-184."}, "verifier_independent_model": false}, "family": "fast-41-diverse-277-002", "id": "fast-41-diverse-277-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "At dock 3, a receiving clerk is checking shipment R-184 against PO-771 and ASN-771A before posting the receipt. The ASN and PO carry matching item entries and matching quantity entries for this shipment. The count sheet records 480 units, consistent with those records. Inspection reports no damage and no safety concern at dock 3. PO-771 was released on 6 May, before the applicable substitution threshold, and Inventory Control has not verified supplier fault. The clerk is using the following governing material. Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope. Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control. Select the single immediate routing classification for this shipment.", "evidence": ["The physical item identifier observed for shipment R-184 is V40.", "PO-771 specifies V40 as the item identifier for shipment R-184.", "The ASN and PO carry matching item entries and matching quantity entries for this shipment.", "The count sheet records 480 units, consistent with those records.", "Inspection reports no damage and no safety concern at dock 3.", "PO-771 was released on 6 May, before the applicable substitution threshold, and Inventory Control has not verified supplier fault.", "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.", "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.", "Select the single immediate routing classification for this shipment."]}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "RECEIVING_CLEAN"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions and governing policy with unchanged shipment, document, scope, time, and request bindings; the evidence contains the exact factual sentences \"The observed physical item identifier for shipment R-184 is V40B.\" and \"PO-771 specifies V40 as the physical item identifier for shipment R-184.\"; the counterfactual's V40 observation agrees with PO-771 and ASN-771A, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"Case note for dock 3 receiving review of shipment R-184 against PO-771 and ASN-771A. The observed physical item identifier for shipment R-184 is V40B. PO-771 specifies V40 as the physical item identifier for shipment R-184. PO-771 and ASN-771A specify the same physical item identifier for shipment R-184. The physical quantity recorded at receipt agrees with the quantity in PO-771, and PO-771 and ASN-771A specify the same quantity for shipment R-184. Inspection found no damage and no safety concern at dock 3. PO-771 was released on 6 May, before the stated approval threshold. Inventory Control has not verified supplier fault for shipment R-184. The receiving clerk is awaiting the immediate routing decision.\\n\\nBuyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\\nRouting policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\\nSelect the single immediate routing classification for this shipment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["context"], "text": "The observed physical item identifier for shipment R-184 is V40B."}, {"path": ["context"], "text": "PO-771 specifies V40 as the physical item identifier for shipment R-184."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "The observed physical item identifier for shipment R-184 is V40B.", "negative_left": "The observed physical item identifier for shipment R-184 is V40.", "negative_right": "PO-771 specifies V40 as the physical item identifier for shipment R-184.", "right": "PO-771 specifies V40 as the physical item identifier for shipment R-184."}, "verifier_independent_model": false}, "family": "fast-41-diverse-277-010", "id": "fast-41-diverse-277-010-base", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "Case note for dock 3 receiving review of shipment R-184 against PO-771 and ASN-771A. The observed physical item identifier for shipment R-184 is V40B. PO-771 specifies V40 as the physical item identifier for shipment R-184. PO-771 and ASN-771A specify the same physical item identifier for shipment R-184. The physical quantity recorded at receipt agrees with the quantity in PO-771, and PO-771 and ASN-771A specify the same quantity for shipment R-184. Inspection found no damage and no safety concern at dock 3. PO-771 was released on 6 May, before the stated approval threshold. Inventory Control has not verified supplier fault for shipment R-184. The receiving clerk is awaiting the immediate routing decision.\n\nBuyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\nRouting policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\nSelect the single immediate routing classification for this shipment."}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "INVENTORY_IDENTITY_HOLD"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions and governing policy with unchanged shipment, document, scope, time, and request bindings; the evidence contains the exact factual sentences \"The observed physical item identifier for shipment R-184 is V40B.\" and \"PO-771 specifies V40 as the physical item identifier for shipment R-184.\"; the counterfactual's V40 observation agrees with PO-771 and ASN-771A, and neither context embeds an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "refuted", "A8": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and A1 is a factual shipment-to-PO identifier comparison rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable with the remaining assignments. Policy evidence preserves the state-originated substitution-date scope and routing policy needed for interpretation; the unchanged questions object already preserves all question-originated criteria and instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish an item-identity discrepancy, matching quantity, no damage or safety concern, and an order date outside the stated substitution-approval window. This is sufficient for INVENTORY_IDENTITY_HOLD under the routing criteria.", "rule_index": 0, "sound": true}, {"reason": "Refuting A1 establishes that the physical item does not differ from the PO item; A2 aligns the PO and ASN identifiers, while A3 and A4 establish matching quantities. With damage and safety concerns refuted, these conditions are sufficient for RECEIVING_CLEAN.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The physical item identifier observed for shipment R-184 differs from the V40 item identifier specified in PO-771."}, {"id": "A2", "statement": "PO-771 and ASN-771A specify the same physical item identifier for shipment R-184."}, {"id": "A3", "statement": "The physical quantity of shipment R-184 differs from the quantity specified in PO-771."}, {"id": "A4", "statement": "PO-771 and ASN-771A specify the same quantity for shipment R-184."}, {"id": "A5", "statement": "Shipment R-184 has observed damage at dock 3."}, {"id": "A6", "statement": "Shipment R-184 has an observed safety concern at dock 3."}, {"id": "A7", "statement": "The release date of PO-771 is on or after 10 May."}, {"id": "A8", "statement": "Inventory Control has already verified supplier fault for shipment R-184."}], "base_state_json": "{\"context\":\"Case note for dock 3 receiving review of shipment R-184 against PO-771 and ASN-771A. The observed physical item identifier for shipment R-184 is V40B. PO-771 specifies V40 as the physical item identifier for shipment R-184. PO-771 and ASN-771A specify the same physical item identifier for shipment R-184. The physical quantity recorded at receipt agrees with the quantity in PO-771, and PO-771 and ASN-771A specify the same quantity for shipment R-184. Inspection found no damage and no safety concern at dock 3. PO-771 was released on 6 May, before the stated approval threshold. Inventory Control has not verified supplier fault for shipment R-184. The receiving clerk is awaiting the immediate routing decision.\\n\\nBuyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\\nRouting policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\\nSelect the single immediate routing classification for this shipment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["context"], "text": "The observed physical item identifier for shipment R-184 is V40B."}, {"path": ["context"], "text": "PO-771 specifies V40 as the physical item identifier for shipment R-184."}], "policy_evidence": [{"path": ["evidence", "3"], "text": "Buyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope."}, {"path": ["evidence", "4"], "text": "Routing policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control."}, {"path": ["request"], "text": "Select the single immediate routing classification for this shipment."}], "rules": [{"justification": "The physical quantity matches both aligned shipment records, but the physical item differs from their common V40 specification. Because PO-771 was released before the approval threshold, the substitution is outside approval scope. With neither damage nor a safety concern and with no prior supplier-fault verification, the immediate route is Inventory Control for the identity discrepancy.", "target": "INVENTORY_IDENTITY_HOLD", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}, {"justification": "The physical item matches PO-771 and therefore also matches ASN-771A because the two records specify the same item. The physical quantity likewise matches both aligned records. With neither damage nor a safety concern, Receiving posts the receipt; the substitution approval date is immaterial because there is no item substitution.", "target": "RECEIVING_CLEAN", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}]}]}, "verified_pair": {"left": "The observed physical item identifier for shipment R-184 is V40B.", "negative_left": "The observed physical item identifier for shipment R-184 is V40.", "negative_right": "PO-771 specifies V40 as the physical item identifier for shipment R-184.", "right": "PO-771 specifies V40 as the physical item identifier for shipment R-184."}, "verifier_independent_model": false}, "family": "fast-41-diverse-277-010", "id": "fast-41-diverse-277-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"DOCK_DAMAGE_HOLD": "Dock supervisor handles any shipment with observed damage or a safety concern; this route takes precedence over count or identity discrepancies.", "INVENTORY_COUNT_HOLD": "Inventory control analyst investigates when there is no damage or safety issue but the physical quantity differs from the PO or ASN.", "INVENTORY_IDENTITY_HOLD": "Inventory control analyst investigates when there is no damage, quantity matches, but the physical item differs or a substitution falls outside documented approval scope.", "RECEIVING_CLEAN": "Receiving clerk posts the receipt when quantity and physical item match, or when a substitution is explicitly within approval scope, with no damage or safety issue.", "SUPPLIER_CLAIM_VERIFIED": "Supplier claims coordinator handles a discrepancy only after Inventory Control has verified supplier fault; it is not the initial route for an unverified discrepancy."}, "instructions": "Apply the stated scope, exceptions, and routing sequence. Choose exactly one option based only on the supplied evidence.", "type": "choice"}}, "state": {"context": "Case note for dock 3 receiving review of shipment R-184 against PO-771 and ASN-771A. The observed physical item identifier for shipment R-184 is V40. PO-771 specifies V40 as the physical item identifier for shipment R-184. PO-771 and ASN-771A specify the same physical item identifier for shipment R-184. The physical quantity recorded at receipt agrees with the quantity in PO-771, and PO-771 and ASN-771A specify the same quantity for shipment R-184. Inspection found no damage and no safety concern at dock 3. PO-771 was released on 6 May, before the stated approval threshold. Inventory Control has not verified supplier fault for shipment R-184. The receiving clerk is awaiting the immediate routing decision.\n\nBuyer approval permits V40B substitutions only for POs released on or after 10 May; earlier orders are outside scope.\nRouting policy: clean receipts stay with Receiving. Count or identity conflicts go to Inventory Control. Damage or safety issues take precedence and go to the Dock Supervisor. Supplier Claims receives only supplier-fault cases already verified by Inventory Control.\nSelect the single immediate routing classification for this shipment."}}, "method": "c2d", "provenance": {"source_id": "diverse-277", "source_is_synthetic": true, "source_sha256": "ff280bce5257030b568e74177eda781a67112a1847f0294b3fffc130a3759285", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "RECEIVING_CLEAN"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged original question and the routing policy, including the exact evidence quotes “The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units.” and “PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received.” The changed counterfactual quantity of 70 is consistent with its unchanged 8-unit substitution and contains no contradictory duplicate assertion. No context supplies a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received.\"},{\"speaker\":\"Procurement note\",\"text\":\"The shipment record identifies the substituted units as unapproved Q-9 items received in place of the ordered P-9 items.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"The receiving log records no shortage or overage beyond the substitution noted in the inspection.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"The case file contains no separate discrepancy report for PO-731.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units."}, {"path": ["1", "text"], "text": "PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units.", "negative_left": "The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units.", "negative_right": "PO-731 ordered 70 units, and the receiving inspection found no physical damage among the units received.", "right": "PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-001", "id": "fast-41-diverse-278-001-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units."}, {"speaker": "Receiving clerk", "text": "PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received."}, {"speaker": "Procurement note", "text": "The shipment record identifies the substituted units as unapproved Q-9 items received in place of the ordered P-9 items."}, {"speaker": "Inventory control analyst", "text": "The receiving log records no shortage or overage beyond the substitution noted in the inspection."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "The case file contains no separate discrepancy report for PO-731."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged original question and the routing policy, including the exact evidence quotes “The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units.” and “PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received.” The changed counterfactual quantity of 70 is consistent with its unchanged 8-unit substitution and contains no contradictory duplicate assertion. No context supplies a gold answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received.\"},{\"speaker\":\"Procurement note\",\"text\":\"The shipment record identifies the substituted units as unapproved Q-9 items received in place of the ordered P-9 items.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"The receiving log records no shortage or overage beyond the substitution noted in the inspection.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"The case file contains no separate discrepancy report for PO-731.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units."}, {"path": ["1", "text"], "text": "PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units.", "negative_left": "The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units.", "negative_right": "PO-731 ordered 70 units, and the receiving inspection found no physical damage among the units received.", "right": "PO-731 ordered 96 units, and the receiving inspection found no physical damage among the units received."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-001", "id": "fast-41-diverse-278-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "The receiving inspection for PO-731 established an unauthorized substitution affecting 8 units."}, {"speaker": "Receiving clerk", "text": "PO-731 ordered 70 units, and the receiving inspection found no physical damage among the units received."}, {"speaker": "Procurement note", "text": "The shipment record identifies the substituted units as unapproved Q-9 items received in place of the ordered P-9 items."}, {"speaker": "Inventory control analyst", "text": "The receiving log records no shortage or overage beyond the substitution noted in the inspection."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "The case file contains no separate discrepancy report for PO-731."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, PO-731 binding, complete factual evidence, governing routing policy, and a coherent quantity change without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[\"The receiving report for PO-731 records 240 units ordered and no physical damage.\",\"The same report records an unauthorized substitution affecting 24 received units.\",\"The receiving clerk matched the shipment to PO-731, and procurement’s approval review found that the substitution had not been authorized. Inspection covered the delivered units, and no separate shortage or overage was recorded. The affected units were included in the shipment count rather than treated as a separate delivery. The report, inspection entry, and approval review concern the same PO-731 delivery. Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0"], "text": "The receiving report for PO-731 records 240 units ordered and no physical damage."}, {"path": ["1"], "text": "The same report records an unauthorized substitution affecting 24 received units."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The receiving report for PO-731 records 240 units ordered and no physical damage.", "negative_left": "The receiving report for PO-731 records 240 units ordered and no physical damage.", "negative_right": "The same report records an unauthorized substitution affecting 48 received units.", "right": "The same report records an unauthorized substitution affecting 24 received units."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-002", "id": "fast-41-diverse-278-002-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": ["The receiving report for PO-731 records 240 units ordered and no physical damage.", "The same report records an unauthorized substitution affecting 24 received units.", "The receiving clerk matched the shipment to PO-731, and procurement’s approval review found that the substitution had not been authorized. Inspection covered the delivered units, and no separate shortage or overage was recorded. The affected units were included in the shipment count rather than treated as a separate delivery. The report, inspection entry, and approval review concern the same PO-731 delivery. Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, PO-731 binding, complete factual evidence, governing routing policy, and a coherent quantity change without embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[\"The receiving report for PO-731 records 240 units ordered and no physical damage.\",\"The same report records an unauthorized substitution affecting 24 received units.\",\"The receiving clerk matched the shipment to PO-731, and procurement’s approval review found that the substitution had not been authorized. Inspection covered the delivered units, and no separate shortage or overage was recorded. The affected units were included in the shipment count rather than treated as a separate delivery. The report, inspection entry, and approval review concern the same PO-731 delivery. Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0"], "text": "The receiving report for PO-731 records 240 units ordered and no physical damage."}, {"path": ["1"], "text": "The same report records an unauthorized substitution affecting 24 received units."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The receiving report for PO-731 records 240 units ordered and no physical damage.", "negative_left": "The receiving report for PO-731 records 240 units ordered and no physical damage.", "negative_right": "The same report records an unauthorized substitution affecting 48 received units.", "right": "The same report records an unauthorized substitution affecting 24 received units."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-002", "id": "fast-41-diverse-278-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": ["The receiving report for PO-731 records 240 units ordered and no physical damage.", "The same report records an unauthorized substitution affecting 48 received units.", "The receiving clerk matched the shipment to PO-731, and procurement’s approval review found that the substitution had not been authorized. Inspection covered the delivered units, and no separate shortage or overage was recorded. The affected units were included in the shipment count rather than treated as a separate delivery. The report, inspection entry, and approval review concern the same PO-731 delivery. Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the PO-731 binding and routing policy, contain two complete factual evidence sentences, change only the affected count coherently, and do not reveal an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 is the procurement order under review, and the receiving file is dated 14 March 2026.\"},{\"speaker\":\"Inspection record\",\"text\":\"The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units.\"},{\"speaker\":\"Inspection record\",\"text\":\"The 14 March 2026 inspection record identifies 24 received PO-731 units as affected by that unauthorized substitution.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"The receiving file contains no separate damage report, shortage confirmation, overage confirmation, or documentation-only discrepancy for PO-731.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units."}, {"path": ["2", "text"], "text": "The 14 March 2026 inspection record identifies 24 received PO-731 units as affected by that unauthorized substitution."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units.", "negative_left": "The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units.", "negative_right": "The 14 March 2026 inspection record identifies 27 received PO-731 units as affected by that unauthorized substitution.", "right": "The 14 March 2026 inspection record identifies 24 received PO-731 units as affected by that unauthorized substitution."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-004", "id": "fast-41-diverse-278-004-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 is the procurement order under review, and the receiving file is dated 14 March 2026."}, {"speaker": "Inspection record", "text": "The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units."}, {"speaker": "Inspection record", "text": "The 14 March 2026 inspection record identifies 24 received PO-731 units as affected by that unauthorized substitution."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Receiving clerk", "text": "The receiving file contains no separate damage report, shortage confirmation, overage confirmation, or documentation-only discrepancy for PO-731."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the PO-731 binding and routing policy, contain two complete factual evidence sentences, change only the affected count coherently, and do not reveal an answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 is the procurement order under review, and the receiving file is dated 14 March 2026.\"},{\"speaker\":\"Inspection record\",\"text\":\"The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units.\"},{\"speaker\":\"Inspection record\",\"text\":\"The 14 March 2026 inspection record identifies 24 received PO-731 units as affected by that unauthorized substitution.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"The receiving file contains no separate damage report, shortage confirmation, overage confirmation, or documentation-only discrepancy for PO-731.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units."}, {"path": ["2", "text"], "text": "The 14 March 2026 inspection record identifies 24 received PO-731 units as affected by that unauthorized substitution."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units.", "negative_left": "The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units.", "negative_right": "The 14 March 2026 inspection record identifies 27 received PO-731 units as affected by that unauthorized substitution.", "right": "The 14 March 2026 inspection record identifies 24 received PO-731 units as affected by that unauthorized substitution."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-004", "id": "fast-41-diverse-278-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 is the procurement order under review, and the receiving file is dated 14 March 2026."}, {"speaker": "Inspection record", "text": "The 14 March 2026 receiving inspection for PO-731 found no physical damage among the received units and established an unauthorized substitution among them; PO-731 ordered 240 units."}, {"speaker": "Inspection record", "text": "The 14 March 2026 inspection record identifies 27 received PO-731 units as affected by that unauthorized substitution."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Receiving clerk", "text": "The receiving file contains no separate damage report, shortage confirmation, overage confirmation, or documentation-only discrepancy for PO-731."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing routing and severity policy verbatim. The PO-731 and carton C5 bindings remain consistent, while changed quantities are permissible observations. The focus evidence contains exactly two complete factual sentences. The counterfactual consistently changes the affected quantity from 24 to 36 without creating contradictory counts or assertions. Neither context embeds a gold answer, rationale, label, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"On 2026-09-01, purchase order PO-731 recorded an order of 240 units.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"On 2026-09-10, the receiving report for PO-731 recorded 24 units affected by an unauthorized substitution.\"},{\"speaker\":\"Inspection lead\",\"text\":\"The receiving inspection found every unit intact, with no dents, cracks, leaks, or other physical damage. The product-code replacement in carton C5 had not been approved by the supplier.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"The shipment paperwork reconciled to the count, and no shortage or overage was recorded. Carton C5 contained the substituted units; the remaining shipment matched the ordered product description.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "On 2026-09-01, purchase order PO-731 recorded an order of 240 units."}, {"path": ["1", "text"], "text": "On 2026-09-10, the receiving report for PO-731 recorded 24 units affected by an unauthorized substitution."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "On 2026-09-01, purchase order PO-731 recorded an order of 240 units.", "negative_left": "On 2026-09-01, purchase order PO-731 recorded an order of 240 units.", "negative_right": "On 2026-09-10, the receiving report for PO-731 recorded 36 units affected by an unauthorized substitution.", "right": "On 2026-09-10, the receiving report for PO-731 recorded 24 units affected by an unauthorized substitution."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-005", "id": "fast-41-diverse-278-005-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "On 2026-09-01, purchase order PO-731 recorded an order of 240 units."}, {"speaker": "Receiving clerk", "text": "On 2026-09-10, the receiving report for PO-731 recorded 24 units affected by an unauthorized substitution."}, {"speaker": "Inspection lead", "text": "The receiving inspection found every unit intact, with no dents, cracks, leaks, or other physical damage. The product-code replacement in carton C5 had not been approved by the supplier."}, {"speaker": "Receiving clerk", "text": "The shipment paperwork reconciled to the count, and no shortage or overage was recorded. Carton C5 contained the substituted units; the remaining shipment matched the ordered product description."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing routing and severity policy verbatim. The PO-731 and carton C5 bindings remain consistent, while changed quantities are permissible observations. The focus evidence contains exactly two complete factual sentences. The counterfactual consistently changes the affected quantity from 24 to 36 without creating contradictory counts or assertions. Neither context embeds a gold answer, rationale, label, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"On 2026-09-01, purchase order PO-731 recorded an order of 240 units.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"On 2026-09-10, the receiving report for PO-731 recorded 24 units affected by an unauthorized substitution.\"},{\"speaker\":\"Inspection lead\",\"text\":\"The receiving inspection found every unit intact, with no dents, cracks, leaks, or other physical damage. The product-code replacement in carton C5 had not been approved by the supplier.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"The shipment paperwork reconciled to the count, and no shortage or overage was recorded. Carton C5 contained the substituted units; the remaining shipment matched the ordered product description.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "On 2026-09-01, purchase order PO-731 recorded an order of 240 units."}, {"path": ["1", "text"], "text": "On 2026-09-10, the receiving report for PO-731 recorded 24 units affected by an unauthorized substitution."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "On 2026-09-01, purchase order PO-731 recorded an order of 240 units.", "negative_left": "On 2026-09-01, purchase order PO-731 recorded an order of 240 units.", "negative_right": "On 2026-09-10, the receiving report for PO-731 recorded 36 units affected by an unauthorized substitution.", "right": "On 2026-09-10, the receiving report for PO-731 recorded 24 units affected by an unauthorized substitution."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-005", "id": "fast-41-diverse-278-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "On 2026-09-01, purchase order PO-731 recorded an order of 240 units."}, {"speaker": "Receiving clerk", "text": "On 2026-09-10, the receiving report for PO-731 recorded 36 units affected by an unauthorized substitution."}, {"speaker": "Inspection lead", "text": "The receiving inspection found every unit intact, with no dents, cracks, leaks, or other physical damage. The product-code replacement in carton C5 had not been approved by the supplier."}, {"speaker": "Receiving clerk", "text": "The shipment paperwork reconciled to the count, and no shortage or overage was recorded. Carton C5 contained the substituted units; the remaining shipment matched the ordered product description."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged routing policy and bindings, the evidence quotes are complete factual sentences—“Purchase order PO-731 specifies 250 units for delivery.” and “Inspection identifies 25 received units under PO-731 as unauthorized substitutions.”—and changing 25 substitutions to 30 coherently changes the affected percentage without adding contradictions or a gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"The receiving record identifies PO-731 and the delivered pump-seal shipment. The packing documentation distinguishes the ordered P-9 units from Q-9 units found in carton C5.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"The proposed Q-9 replacements were never approved, and the receiving clerk recorded them as substitutions rather than authorized equivalents.\"},{\"speaker\":\"Inspection lead\",\"text\":\"Inspection found the received units intact, with no dents, breaks, leaks, or other physical damage. The dock log records no separate shortage or overage and no additional discrepancy.\"},{\"speaker\":\"Procurement record\",\"text\":\"Purchase order PO-731 specifies 250 units for delivery.\"},{\"speaker\":\"Inspection record\",\"text\":\"Inspection identifies 25 received units under PO-731 as unauthorized substitutions.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["3", "text"], "text": "Purchase order PO-731 specifies 250 units for delivery."}, {"path": ["4", "text"], "text": "Inspection identifies 25 received units under PO-731 as unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "Purchase order PO-731 specifies 250 units for delivery.", "negative_left": "Purchase order PO-731 specifies 250 units for delivery.", "negative_right": "Inspection identifies 30 received units under PO-731 as unauthorized substitutions.", "right": "Inspection identifies 25 received units under PO-731 as unauthorized substitutions."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-006", "id": "fast-41-diverse-278-006-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "The receiving record identifies PO-731 and the delivered pump-seal shipment. The packing documentation distinguishes the ordered P-9 units from Q-9 units found in carton C5."}, {"speaker": "Inventory control analyst", "text": "The proposed Q-9 replacements were never approved, and the receiving clerk recorded them as substitutions rather than authorized equivalents."}, {"speaker": "Inspection lead", "text": "Inspection found the received units intact, with no dents, breaks, leaks, or other physical damage. The dock log records no separate shortage or overage and no additional discrepancy."}, {"speaker": "Procurement record", "text": "Purchase order PO-731 specifies 250 units for delivery."}, {"speaker": "Inspection record", "text": "Inspection identifies 25 received units under PO-731 as unauthorized substitutions."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged routing policy and bindings, the evidence quotes are complete factual sentences—“Purchase order PO-731 specifies 250 units for delivery.” and “Inspection identifies 25 received units under PO-731 as unauthorized substitutions.”—and changing 25 substitutions to 30 coherently changes the affected percentage without adding contradictions or a gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"The receiving record identifies PO-731 and the delivered pump-seal shipment. The packing documentation distinguishes the ordered P-9 units from Q-9 units found in carton C5.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"The proposed Q-9 replacements were never approved, and the receiving clerk recorded them as substitutions rather than authorized equivalents.\"},{\"speaker\":\"Inspection lead\",\"text\":\"Inspection found the received units intact, with no dents, breaks, leaks, or other physical damage. The dock log records no separate shortage or overage and no additional discrepancy.\"},{\"speaker\":\"Procurement record\",\"text\":\"Purchase order PO-731 specifies 250 units for delivery.\"},{\"speaker\":\"Inspection record\",\"text\":\"Inspection identifies 25 received units under PO-731 as unauthorized substitutions.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["3", "text"], "text": "Purchase order PO-731 specifies 250 units for delivery."}, {"path": ["4", "text"], "text": "Inspection identifies 25 received units under PO-731 as unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "Purchase order PO-731 specifies 250 units for delivery.", "negative_left": "Purchase order PO-731 specifies 250 units for delivery.", "negative_right": "Inspection identifies 30 received units under PO-731 as unauthorized substitutions.", "right": "Inspection identifies 25 received units under PO-731 as unauthorized substitutions."}, "verifier_independent_model": false}, "family": "fast-41-diverse-278-006", "id": "fast-41-diverse-278-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "The receiving record identifies PO-731 and the delivered pump-seal shipment. The packing documentation distinguishes the ordered P-9 units from Q-9 units found in carton C5."}, {"speaker": "Inventory control analyst", "text": "The proposed Q-9 replacements were never approved, and the receiving clerk recorded them as substitutions rather than authorized equivalents."}, {"speaker": "Inspection lead", "text": "Inspection found the received units intact, with no dents, breaks, leaks, or other physical damage. The dock log records no separate shortage or overage and no additional discrepancy."}, {"speaker": "Procurement record", "text": "Purchase order PO-731 specifies 250 units for delivery."}, {"speaker": "Inspection record", "text": "Inspection identifies 30 received units under PO-731 as unauthorized substitutions."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and both contexts retain the routing rule and review sequence. Question bindings are preserved through the same shipment, PO 7714, ASN 7714-A, and 09:00 UTC on 17 September 2026. The evidence consists of complete factual sentences: “At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.” and “At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936.”; the counterfactual retains “At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.” and changes the second span to “At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-4821.” The counterfactual is coherent because the received SKU matches the PO but differs from the ASN, with no contradictory duplicate assertion. Neither context embeds an answer, code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\\n\\nThe receiving log was timestamped at 09:00 UTC on 17 September 2026. The ASN record for this shipment lists SKU NQ-5936, while the receiving record identifies the alternate goods as undamaged. The Dock Supervisor reviewed the authorization file and found no written substitution approval or change notice. The shipment is associated with PO 7714 and ASN 7714-A, and the receiving count was complete at 40 units.\",\"evidence\":[\"At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.\",\"At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821."}, {"path": ["evidence", "1"], "text": "At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.", "negative_left": "At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.", "negative_right": "At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-4821.", "right": "At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936."}, "verifier_independent_model": false}, "family": "fast-41-diverse-280-001", "id": "fast-41-diverse-280-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\n\nThe receiving log was timestamped at 09:00 UTC on 17 September 2026. The ASN record for this shipment lists SKU NQ-5936, while the receiving record identifies the alternate goods as undamaged. The Dock Supervisor reviewed the authorization file and found no written substitution approval or change notice. The shipment is associated with PO 7714 and ASN 7714-A, and the receiving count was complete at 40 units.", "evidence": ["At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.", "At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936."]}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and both contexts retain the routing rule and review sequence. Question bindings are preserved through the same shipment, PO 7714, ASN 7714-A, and 09:00 UTC on 17 September 2026. The evidence consists of complete factual sentences: “At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.” and “At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936.”; the counterfactual retains “At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.” and changes the second span to “At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-4821.” The counterfactual is coherent because the received SKU matches the PO but differs from the ASN, with no contradictory duplicate assertion. Neither context embeds an answer, code, rationale, proposition identifier, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\\n\\nThe receiving log was timestamped at 09:00 UTC on 17 September 2026. The ASN record for this shipment lists SKU NQ-5936, while the receiving record identifies the alternate goods as undamaged. The Dock Supervisor reviewed the authorization file and found no written substitution approval or change notice. The shipment is associated with PO 7714 and ASN 7714-A, and the receiving count was complete at 40 units.\",\"evidence\":[\"At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.\",\"At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821."}, {"path": ["evidence", "1"], "text": "At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.", "negative_left": "At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.", "negative_right": "At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-4821.", "right": "At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-5936."}, "verifier_independent_model": false}, "family": "fast-41-diverse-280-001", "id": "fast-41-diverse-280-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\n\nThe receiving log was timestamped at 09:00 UTC on 17 September 2026. The ASN record for this shipment lists SKU NQ-5936, while the receiving record identifies the alternate goods as undamaged. The Dock Supervisor reviewed the authorization file and found no written substitution approval or change notice. The shipment is associated with PO 7714 and ASN 7714-A, and the receiving count was complete at 40 units.", "evidence": ["At 09:00 UTC on 17 September 2026, the goods received in the shipment associated with PO 7714 and ASN 7714-A had SKU NQ-4821.", "At 09:00 UTC on 17 September 2026, PO 7714 specified SKU NQ-4821."]}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions and governing routing policy. Both evidence arrays contain two complete factual sentences. The counterfactual consistently changes only the received SKU to match the PO while leaving it different from the ASN. Neither context embeds an answer, label rationale, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"North Quay Warehouse is reconciling the shipment linked to PO 7714 and ASN 7714-A. Receiving records connect the cartons to both documents. The shipment file contains no written authorization for substituting the received goods, and inspection records describe the alternate goods as undamaged. At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\",\"evidence\":[\"The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-482.\",\"PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-482."}, {"path": ["evidence", "1"], "text": "PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-482.", "negative_left": "The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-417.", "negative_right": "PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319.", "right": "PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319."}, "verifier_independent_model": false}, "family": "fast-41-diverse-280-006", "id": "fast-41-diverse-280-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "North Quay Warehouse is reconciling the shipment linked to PO 7714 and ASN 7714-A. Receiving records connect the cartons to both documents. The shipment file contains no written authorization for substituting the received goods, and inspection records describe the alternate goods as undamaged. At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-482.", "PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319."]}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full inputs retain the unchanged questions and governing routing policy. Both evidence arrays contain two complete factual sentences. The counterfactual consistently changes only the received SKU to match the PO while leaving it different from the ASN. Neither context embeds an answer, label rationale, rule table, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"North Quay Warehouse is reconciling the shipment linked to PO 7714 and ASN 7714-A. Receiving records connect the cartons to both documents. The shipment file contains no written authorization for substituting the received goods, and inspection records describe the alternate goods as undamaged. At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\",\"evidence\":[\"The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-482.\",\"PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-482."}, {"path": ["evidence", "1"], "text": "PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-482.", "negative_left": "The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-417.", "negative_right": "PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319.", "right": "PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319."}, "verifier_independent_model": false}, "family": "fast-41-diverse-280-006", "id": "fast-41-diverse-280-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "North Quay Warehouse is reconciling the shipment linked to PO 7714 and ASN 7714-A. Receiving records connect the cartons to both documents. The shipment file contains no written authorization for substituting the received goods, and inspection records describe the alternate goods as undamaged. At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["The goods received in the shipment associated with PO 7714 and ASN 7714-A carry SKU NQ-417.", "PO 7714 specifies SKU NQ-417 for that shipment, and ASN 7714-A specifies SKU NQ-319."]}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and routing policy are preserved in both full inputs. The shipment, PO, ASN, and date bindings remain consistent with the request. The evidence consists of two complete factual sentences: \"For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418.\" and \"PO 7714 specifies SKU PO562 for the goods in that shipment.\" The counterfactual changes the PO SKU coherently without creating contradictory duplicate measurements or assertions. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"North Quay Warehouse receiving records link the shipment to PO 7714 and ASN 7714-A. For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418. PO 7714 specifies SKU PO562 for the goods in that shipment. ASN 7714-A also specifies SKU PO562. The alternate units were inspected and found undamaged. The procurement file and delivery records contain no written authorization for substituting the goods. At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418."}, {"path": ["context"], "text": "PO 7714 specifies SKU PO562 for the goods in that shipment."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418.", "negative_left": "For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418.", "negative_right": "PO 7714 specifies SKU RQ418 for the goods in that shipment.", "right": "PO 7714 specifies SKU PO562 for the goods in that shipment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-280-008", "id": "fast-41-diverse-280-008-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "North Quay Warehouse receiving records link the shipment to PO 7714 and ASN 7714-A. For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418. PO 7714 specifies SKU PO562 for the goods in that shipment. ASN 7714-A also specifies SKU PO562. The alternate units were inspected and found undamaged. The procurement file and delivery records contain no written authorization for substituting the goods. At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions and routing policy are preserved in both full inputs. The shipment, PO, ASN, and date bindings remain consistent with the request. The evidence consists of two complete factual sentences: \"For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418.\" and \"PO 7714 specifies SKU PO562 for the goods in that shipment.\" The counterfactual changes the PO SKU coherently without creating contradictory duplicate measurements or assertions. Neither context embeds a gold answer, answer code, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\":\"North Quay Warehouse receiving records link the shipment to PO 7714 and ASN 7714-A. For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418. PO 7714 specifies SKU PO562 for the goods in that shipment. ASN 7714-A also specifies SKU PO562. The alternate units were inspected and found undamaged. The procurement file and delivery records contain no written authorization for substituting the goods. At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418."}, {"path": ["context"], "text": "PO 7714 specifies SKU PO562 for the goods in that shipment."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418.", "negative_left": "For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418.", "negative_right": "PO 7714 specifies SKU RQ418 for the goods in that shipment.", "right": "PO 7714 specifies SKU PO562 for the goods in that shipment."}, "verifier_independent_model": false}, "family": "fast-41-diverse-280-008", "id": "fast-41-diverse-280-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "North Quay Warehouse receiving records link the shipment to PO 7714 and ASN 7714-A. For the shipment associated with PO 7714 and ASN 7714-A, the goods received on 2026-09-17 have SKU RQ418. PO 7714 specifies SKU RQ418 for the goods in that shipment. ASN 7714-A also specifies SKU PO562. The alternate units were inspected and found undamaged. The procurement file and delivery records contain no written authorization for substituting the goods. At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, criteria, and instructions. Both contexts retain the Northstar, PO-1842, shipment, and receiving-routing bindings while changing only observations. The evidence contains exactly two complete factual sentences: \"At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417.\" and \"PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026.\" The counterfactual consistently changes the unit SKU to NS-418 and matches the PO and ASN without contradictory duplicate assertions. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417. PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026. The ASN lists one unit of the PO-required SKU and contains no authorization for an alternate SKU. The shipment contains one unit, matching the quantity required by PO-1842 and the quantity listed by the ASN. Inspection confirms that the unit and packaging meet the condition specified by both PO-1842 and the ASN. Receiving notes identify no separate unresolved discrepancy, and no physical receiving hazard, damage, or containment issue is present. The shipment remains on hold while the receiving record and supplier documentation are routed for review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417."}, {"path": [], "text": "PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417.", "negative_left": "At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-418.", "negative_right": "PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026.", "right": "PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-282-007", "id": "fast-41-diverse-282-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417. PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026. The ASN lists one unit of the PO-required SKU and contains no authorization for an alternate SKU. The shipment contains one unit, matching the quantity required by PO-1842 and the quantity listed by the ASN. Inspection confirms that the unit and packaging meet the condition specified by both PO-1842 and the ASN. Receiving notes identify no separate unresolved discrepancy, and no physical receiving hazard, damage, or containment issue is present. The shipment remains on hold while the receiving record and supplier documentation are routed for review."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing policy, criteria, and instructions. Both contexts retain the Northstar, PO-1842, shipment, and receiving-routing bindings while changing only observations. The evidence contains exactly two complete factual sentences: \"At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417.\" and \"PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026.\" The counterfactual consistently changes the unit SKU to NS-418 and matches the PO and ASN without contradictory duplicate assertions. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and the focus atom is a factual physical-SKU comparison rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only whether any physical unit differs from the PO SKU. Empty policy_evidence is correct because all governing routing criteria, priorities, and substitution rules originate in the retained questions object; the original state contributes only case-specific observations that need not be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A physical SKU mismatch with the PO, combined with refutation of explicit ASN authorization, is an undocumented substitution. The unchanged question expressly requires supplier-claims routing at level 3 for that nonconformance, regardless of matching quantity or lack of damage.", "rule_index": 0, "sound": true}, {"reason": "These conditions establish physical SKU conformity with the PO; ASN-to-PO SKU conformity therefore also establishes physical SKU conformity with the ASN. They additionally establish quantity and condition conformity with both records and exclude any unresolved independent discrepancy or physical receiving risk. Thus routine acceptance at level 0 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Northstar’s shipment at Door 6 contains at least one physical unit whose SKU differs from the SKU required by PO-1842."}, {"id": "a2", "statement": "The ASN for Northstar’s shipment explicitly authorizes the differing physical-unit SKU as a substitution for the SKU required by PO-1842."}, {"id": "a3", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity required by PO-1842."}, {"id": "a4", "statement": "The physical unit count in Northstar’s shipment at Door 6 equals the quantity listed by the ASN."}, {"id": "a5", "statement": "The SKU listed by the ASN for Northstar’s shipment equals the SKU required by PO-1842."}, {"id": "a6", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition required by PO-1842."}, {"id": "a7", "statement": "The observed condition of Northstar’s shipment at Door 6 matches the condition specified by the ASN."}, {"id": "a8", "statement": "A receiving discrepancy independent of whether any physical unit’s SKU differs from PO-1842 remains unresolved for Northstar’s shipment at Door 6."}, {"id": "a9", "statement": "Northstar’s shipment at Door 6 presents a physical receiving risk independent of whether any physical unit’s SKU differs from PO-1842."}], "base_state_json": "\"At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417. PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026. The ASN lists one unit of the PO-required SKU and contains no authorization for an alternate SKU. The shipment contains one unit, matching the quantity required by PO-1842 and the quantity listed by the ASN. Inspection confirms that the unit and packaging meet the condition specified by both PO-1842 and the ASN. Receiving notes identify no separate unresolved discrepancy, and no physical receiving hazard, damage, or containment issue is present. The shipment remains on hold while the receiving record and supplier documentation are routed for review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417."}, {"path": [], "text": "PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026."}], "policy_evidence": [], "rules": [{"justification": "A physical item differing from the PO is an undocumented substitution when the ASN does not explicitly authorize it, requiring supplier-claims routing regardless of matching quantity or lack of damage.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}]}, {"justification": "The shipment may be routinely accepted when its physical SKU, quantity, and condition match both records, no independent discrepancy remains unresolved, and no independent physical receiving risk requires escalation.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-417.", "negative_left": "At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-418.", "negative_right": "PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026.", "right": "PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026."}, "verifier_independent_model": false}, "family": "fast-41-diverse-282-007", "id": "fast-41-diverse-282-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Receiving clerk: Accept and receive the shipment when SKU, quantity, and condition match the PO and ASN, with no unresolved discrepancy.", "1 — Inventory control analyst: Resolve a low-severity record issue, such as a correct physical shipment paired with a count-sheet typo or unclear internal inventory entry.", "2 — Dock supervisor: Control a physical receiving risk, such as damaged cartons, unsafe unloading conditions, or a shipment needing recount or dock-level containment.", "3 — Supplier claims coordinator: Handle an unacceptable supplier nonconformance, including wrong items or substitutions not explicitly authorized by the ASN, and initiate claim, rejection, or return processing."], "instructions": "Select the single routing level required by the evidence. Apply the levels as an ordered scale from routine acceptance to highest operational escalation. An item differing from the PO is an undocumented substitution unless the ASN explicitly authorizes it; matching quantity and lack of damage do not make such a substitution acceptable.", "type": "score"}}, "state": "At Door 6 on 17 September 2026, the sole physical unit in Northstar’s shipment bears SKU NS-418. PO-1842 requires SKU NS-418 for Northstar’s shipment at Door 6 on 17 September 2026. The ASN lists one unit of the PO-required SKU and contains no authorization for an alternate SKU. The shipment contains one unit, matching the quantity required by PO-1842 and the quantity listed by the ASN. Inspection confirms that the unit and packaging meet the condition specified by both PO-1842 and the ASN. Receiving notes identify no separate unresolved discrepancy, and no physical receiving hazard, damage, or containment issue is present. The shipment remains on hold while the receiving record and supplier documentation are routed for review."}, "method": "c2d", "provenance": {"source_id": "diverse-282", "source_is_synthetic": true, "source_sha256": "706b6c352f6057e0c9c0664dcec02a7940e00136e8814cc5b3f0588bbba8fb4b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, preserving all governing policy and criteria. Both contexts retain Line 4, the Mint-500 run, 14:00, and 17 September 2026 bindings. The two evidence spans are complete factual sentences: \"For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4.\" and \"At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria.\" The counterfactual changes only the S-4 result to 0.8%, coherently making one check fail. Neither context contains a gold answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "full_context_fact_states": {"base": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "counterfactual": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "remove_left": {"quality_checks_passed": "unknown"}, "remove_right": {"quality_checks_passed": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"quality_checks_passed": "unknown"}, "negative_pair": {"quality_checks_passed": "refuted"}, "negative_sentence": {"quality_checks_passed": "unknown"}, "positive_pair": {"quality_checks_passed": "supported"}, "right": {"quality_checks_passed": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship, including permissible universally quantified relationships over required materials, records, or checks. The focus is a factual quality-check status. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state adds case observations but no additional interpretive policy that must be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports every readiness condition expressly required for release: current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault.", "rule_index": 0, "sound": true}, {"reason": "A refuted required-quality-check condition prevents release. With materials and every other readiness condition supported, the blocker is quality rather than material for the scheduled Mint-500 run, so materials support is excluded and none_of_above is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_ready", "statement": "Every material required for Line 4’s currently scheduled Mint-500 run at 14:00 is staged in the required quantity."}, {"id": "tooling_approved", "statement": "The tooling installed on Line 4 for the currently scheduled Mint-500 run at 14:00 matches the approved Mint-500 tooling specification."}, {"id": "cleaning_records_complete", "statement": "Every cleaning record required for Line 4’s changeover to the currently scheduled Mint-500 run at 14:00 is complete."}, {"id": "staffing_adequate", "statement": "The number of trained operators assigned to Line 4 for the currently scheduled Mint-500 run at 14:00 meets or exceeds that run’s staffing requirement."}, {"id": "quality_checks_passed", "statement": "Every quality-check result required for Line 4 after cleaning and before the currently scheduled Mint-500 run at 14:00 satisfies its corresponding signed pass criterion."}, {"id": "no_open_maintenance_fault", "statement": "No maintenance fault affecting Line 4 is open at the readiness decision for the currently scheduled Mint-500 run at 14:00."}], "base_state_json": "\"Readiness case note for Line 4: The production scheduler has a Citrus-500-to-Mint-500 changeover scheduled for 14:00 on 17 September 2026. All bottles, caps, labels, and concentrate required for the scheduled Mint-500 run are staged in the required quantity. The installed tooling matches the approved Mint-500 specification. Required changeover cleaning records are complete. Four trained operators are assigned, meeting the staffing requirement. The conveyor sensor repair was closed after a successful test, and no maintenance fault affecting the line is open at the readiness decision; a worn spare guide remains tagged out in storage and does not affect the line. For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4. At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria.\"", "base_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "counter_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "focus_atom": "quality_checks_passed", "focus_evidence": [{"path": [], "text": "For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4."}, {"path": [], "text": "At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria."}], "policy_evidence": [], "rules": [{"justification": "All expressly required readiness conditions are satisfied, including the required post-clean quality checks.", "target": "release_to_production", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}, {"justification": "A failed required quality check prevents release, while supported current-job material readiness means the sole blocker is not missing or deficient material and therefore does not qualify for materials support.", "target": "none_of_above", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}]}, "verified_pair": {"left": "For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4.", "negative_left": "For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4.", "negative_right": "At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.8%, respectively, and one fails the signed criteria.", "right": "At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria."}, "verifier_independent_model": false}, "family": "fast-41-diverse-289-001", "id": "fast-41-diverse-289-001-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Select when the line cannot be released and the evidence does not meet the materials-support rubric, including when an unresolved quality or maintenance requirement is the blocker.", "release_to_production": "Release the line to production only if every required readiness condition is satisfied, including a passed post-clean quality check.", "route_to_materials_support": "Do not release; route to materials support only if the sole blocker is material required for the currently scheduled Mint-500 run."}, "instructions": "Choose the single applicable routing/readiness outcome. Release requires current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault. Route to materials support only when the sole release blocker is missing or deficient material for the scheduled job.", "type": "choice"}}, "state": "Readiness case note for Line 4: The production scheduler has a Citrus-500-to-Mint-500 changeover scheduled for 14:00 on 17 September 2026. All bottles, caps, labels, and concentrate required for the scheduled Mint-500 run are staged in the required quantity. The installed tooling matches the approved Mint-500 specification. Required changeover cleaning records are complete. Four trained operators are assigned, meeting the staffing requirement. The conveyor sensor repair was closed after a successful test, and no maintenance fault affecting the line is open at the readiness decision; a worn spare guide remains tagged out in storage and does not affect the line. For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4. At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria."}, "method": "c2d", "provenance": {"source_id": "diverse-289", "source_is_synthetic": true, "source_sha256": "884345667ad55a5ef86e09cc2ba68f6b13c8ccb20beb9fbed251a080f3d46c34", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "release_to_production"}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is verbatim, preserving all governing policy and criteria. Both contexts retain Line 4, the Mint-500 run, 14:00, and 17 September 2026 bindings. The two evidence spans are complete factual sentences: \"For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4.\" and \"At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria.\" The counterfactual changes only the S-4 result to 0.8%, coherently making one check fail. Neither context contains a gold answer, code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "full_context_fact_states": {"base": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "counterfactual": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "remove_left": {"quality_checks_passed": "unknown"}, "remove_right": {"quality_checks_passed": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"quality_checks_passed": "unknown"}, "negative_pair": {"quality_checks_passed": "refuted"}, "negative_sentence": {"quality_checks_passed": "unknown"}, "positive_pair": {"quality_checks_passed": "supported"}, "right": {"quality_checks_passed": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship, including permissible universally quantified relationships over required materials, records, or checks. The focus is a factual quality-check status. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state adds case observations but no additional interpretive policy that must be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports every readiness condition expressly required for release: current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault.", "rule_index": 0, "sound": true}, {"reason": "A refuted required-quality-check condition prevents release. With materials and every other readiness condition supported, the blocker is quality rather than material for the scheduled Mint-500 run, so materials support is excluded and none_of_above is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_ready", "statement": "Every material required for Line 4’s currently scheduled Mint-500 run at 14:00 is staged in the required quantity."}, {"id": "tooling_approved", "statement": "The tooling installed on Line 4 for the currently scheduled Mint-500 run at 14:00 matches the approved Mint-500 tooling specification."}, {"id": "cleaning_records_complete", "statement": "Every cleaning record required for Line 4’s changeover to the currently scheduled Mint-500 run at 14:00 is complete."}, {"id": "staffing_adequate", "statement": "The number of trained operators assigned to Line 4 for the currently scheduled Mint-500 run at 14:00 meets or exceeds that run’s staffing requirement."}, {"id": "quality_checks_passed", "statement": "Every quality-check result required for Line 4 after cleaning and before the currently scheduled Mint-500 run at 14:00 satisfies its corresponding signed pass criterion."}, {"id": "no_open_maintenance_fault", "statement": "No maintenance fault affecting Line 4 is open at the readiness decision for the currently scheduled Mint-500 run at 14:00."}], "base_state_json": "\"Readiness case note for Line 4: The production scheduler has a Citrus-500-to-Mint-500 changeover scheduled for 14:00 on 17 September 2026. All bottles, caps, labels, and concentrate required for the scheduled Mint-500 run are staged in the required quantity. The installed tooling matches the approved Mint-500 specification. Required changeover cleaning records are complete. Four trained operators are assigned, meeting the staffing requirement. The conveyor sensor repair was closed after a successful test, and no maintenance fault affecting the line is open at the readiness decision; a worn spare guide remains tagged out in storage and does not affect the line. For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4. At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria.\"", "base_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "counter_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "focus_atom": "quality_checks_passed", "focus_evidence": [{"path": [], "text": "For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4."}, {"path": [], "text": "At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria."}], "policy_evidence": [], "rules": [{"justification": "All expressly required readiness conditions are satisfied, including the required post-clean quality checks.", "target": "release_to_production", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}, {"justification": "A failed required quality check prevents release, while supported current-job material readiness means the sole blocker is not missing or deficient material and therefore does not qualify for materials support.", "target": "none_of_above", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}]}, "verified_pair": {"left": "For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4.", "negative_left": "For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4.", "negative_right": "At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.8%, respectively, and one fails the signed criteria.", "right": "At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.2%, respectively, and both satisfy the signed criteria."}, "verifier_independent_model": false}, "family": "fast-41-diverse-289-001", "id": "fast-41-diverse-289-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Select when the line cannot be released and the evidence does not meet the materials-support rubric, including when an unresolved quality or maintenance requirement is the blocker.", "release_to_production": "Release the line to production only if every required readiness condition is satisfied, including a passed post-clean quality check.", "route_to_materials_support": "Do not release; route to materials support only if the sole blocker is material required for the currently scheduled Mint-500 run."}, "instructions": "Choose the single applicable routing/readiness outcome. Release requires current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault. Route to materials support only when the sole release blocker is missing or deficient material for the scheduled job.", "type": "choice"}}, "state": "Readiness case note for Line 4: The production scheduler has a Citrus-500-to-Mint-500 changeover scheduled for 14:00 on 17 September 2026. All bottles, caps, labels, and concentrate required for the scheduled Mint-500 run are staged in the required quantity. The installed tooling matches the approved Mint-500 specification. Required changeover cleaning records are complete. Four trained operators are assigned, meeting the staffing requirement. The conveyor sensor repair was closed after a successful test, and no maintenance fault affecting the line is open at the readiness decision; a worn spare guide remains tagged out in storage and does not affect the line. For Line 4's currently scheduled Mint-500 run at 14:00 on 17 September 2026, the required post-cleaning checks are Check T-18 and Check S-4, with signed pass criteria of 70–80 °C for T-18 and a seal-leak rate below 0.5% for S-4. At 14:00 on 17 September 2026, the two recorded results for Line 4's checks are 75 °C and 0.8%, respectively, and one fails the signed criteria."}, "method": "c2d", "provenance": {"source_id": "diverse-289", "source_is_synthetic": true, "source_sha256": "884345667ad55a5ef86e09cc2ba68f6b13c8ccb20beb9fbed251a080f3d46c34", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing routing policy, and both contexts retain the Line 4, Mint-500, 14:00, and readiness-review bindings. The evidence spans are complete factual sentences: \"For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18.\" and \"At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU.\" The counterfactual changes only QC-18 from 1.8 to 2.1 NTU, which coherently fails its stated criterion without creating duplicate measurements. Neither context embeds an answer, label, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "full_context_fact_states": {"base": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "counterfactual": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "remove_left": {"quality_checks_passed": "unknown"}, "remove_right": {"quality_checks_passed": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"quality_checks_passed": "unknown"}, "negative_pair": {"quality_checks_passed": "refuted"}, "negative_sentence": {"quality_checks_passed": "unknown"}, "positive_pair": {"quality_checks_passed": "supported"}, "right": {"quality_checks_passed": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship, including permissible universally quantified relationships over required materials, records, or checks. The focus is a factual quality-check status. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state adds case observations but no additional interpretive policy that must be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports every readiness condition expressly required for release: current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault.", "rule_index": 0, "sound": true}, {"reason": "A refuted required-quality-check condition prevents release. With materials and every other readiness condition supported, the blocker is quality rather than material for the scheduled Mint-500 run, so materials support is excluded and none_of_above is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_ready", "statement": "Every material required for Line 4’s currently scheduled Mint-500 run at 14:00 is staged in the required quantity."}, {"id": "tooling_approved", "statement": "The tooling installed on Line 4 for the currently scheduled Mint-500 run at 14:00 matches the approved Mint-500 tooling specification."}, {"id": "cleaning_records_complete", "statement": "Every cleaning record required for Line 4’s changeover to the currently scheduled Mint-500 run at 14:00 is complete."}, {"id": "staffing_adequate", "statement": "The number of trained operators assigned to Line 4 for the currently scheduled Mint-500 run at 14:00 meets or exceeds that run’s staffing requirement."}, {"id": "quality_checks_passed", "statement": "Every quality-check result required for Line 4 after cleaning and before the currently scheduled Mint-500 run at 14:00 satisfies its corresponding signed pass criterion."}, {"id": "no_open_maintenance_fault", "statement": "No maintenance fault affecting Line 4 is open at the readiness decision for the currently scheduled Mint-500 run at 14:00."}], "base_state_json": "\"Line 4 is scheduled to change over from Citrus-500 to Mint-500 at 14:00. Bottles, caps, labels, and concentrate for this run are staged in the required quantities; a blue-cap shortage concerns only next week’s Ocean-750 run. The approved Mint tooling is installed, and four trained operators are assigned, meeting the scheduled staffing requirement. The required cleaning rinse is complete and its record is signed. Yesterday’s conveyor sensor repair was tested successfully and closed, while a worn spare guide remains tagged out in storage without affecting the line. For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18. At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU. The readiness review concerns this scheduled run.\"", "base_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "counter_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "focus_atom": "quality_checks_passed", "focus_evidence": [{"path": [], "text": "For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18."}, {"path": [], "text": "At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU."}], "policy_evidence": [], "rules": [{"justification": "All expressly required readiness conditions are satisfied, including the required post-clean quality checks.", "target": "release_to_production", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}, {"justification": "A failed required quality check prevents release, while supported current-job material readiness means the sole blocker is not missing or deficient material and therefore does not qualify for materials support.", "target": "none_of_above", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}]}, "verified_pair": {"left": "For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18.", "negative_left": "For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18.", "negative_right": "At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 2.1 NTU.", "right": "At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU."}, "verifier_independent_model": false}, "family": "fast-41-diverse-289-008", "id": "fast-41-diverse-289-008-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Select when the line cannot be released and the evidence does not meet the materials-support rubric, including when an unresolved quality or maintenance requirement is the blocker.", "release_to_production": "Release the line to production only if every required readiness condition is satisfied, including a passed post-clean quality check.", "route_to_materials_support": "Do not release; route to materials support only if the sole blocker is material required for the currently scheduled Mint-500 run."}, "instructions": "Choose the single applicable routing/readiness outcome. Release requires current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault. Route to materials support only when the sole release blocker is missing or deficient material for the scheduled job.", "type": "choice"}}, "state": "Line 4 is scheduled to change over from Citrus-500 to Mint-500 at 14:00. Bottles, caps, labels, and concentrate for this run are staged in the required quantities; a blue-cap shortage concerns only next week’s Ocean-750 run. The approved Mint tooling is installed, and four trained operators are assigned, meeting the scheduled staffing requirement. The required cleaning rinse is complete and its record is signed. Yesterday’s conveyor sensor repair was tested successfully and closed, while a worn spare guide remains tagged out in storage without affecting the line. For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18. At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU. The readiness review concerns this scheduled run."}, "method": "c2d", "provenance": {"source_id": "diverse-289", "source_is_synthetic": true, "source_sha256": "884345667ad55a5ef86e09cc2ba68f6b13c8ccb20beb9fbed251a080f3d46c34", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "release_to_production"}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing routing policy, and both contexts retain the Line 4, Mint-500, 14:00, and readiness-review bindings. The evidence spans are complete factual sentences: \"For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18.\" and \"At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU.\" The counterfactual changes only QC-18 from 1.8 to 2.1 NTU, which coherently fails its stated criterion without creating duplicate measurements. Neither context embeds an answer, label, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "full_context_fact_states": {"base": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "supported", "staffing_adequate": "supported", "tooling_approved": "supported"}, "counterfactual": {"cleaning_records_complete": "supported", "materials_ready": "supported", "no_open_maintenance_fault": "supported", "quality_checks_passed": "refuted", "staffing_adequate": "supported", "tooling_approved": "supported"}, "remove_left": {"quality_checks_passed": "unknown"}, "remove_right": {"quality_checks_passed": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"quality_checks_passed": "unknown"}, "negative_pair": {"quality_checks_passed": "refuted"}, "negative_sentence": {"quality_checks_passed": "unknown"}, "positive_pair": {"quality_checks_passed": "supported"}, "right": {"quality_checks_passed": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship, including permissible universally quantified relationships over required materials, records, or checks. The focus is a factual quality-check status. The base and counter assignments are realizable with only that status changing. Empty policy_evidence is correct because all governing routing rules are already retained in the questions object; the original state adds case observations but no additional interpretive policy that must be preserved.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supports every readiness condition expressly required for release: current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault.", "rule_index": 0, "sound": true}, {"reason": "A refuted required-quality-check condition prevents release. With materials and every other readiness condition supported, the blocker is quality rather than material for the scheduled Mint-500 run, so materials support is excluded and none_of_above is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_ready", "statement": "Every material required for Line 4’s currently scheduled Mint-500 run at 14:00 is staged in the required quantity."}, {"id": "tooling_approved", "statement": "The tooling installed on Line 4 for the currently scheduled Mint-500 run at 14:00 matches the approved Mint-500 tooling specification."}, {"id": "cleaning_records_complete", "statement": "Every cleaning record required for Line 4’s changeover to the currently scheduled Mint-500 run at 14:00 is complete."}, {"id": "staffing_adequate", "statement": "The number of trained operators assigned to Line 4 for the currently scheduled Mint-500 run at 14:00 meets or exceeds that run’s staffing requirement."}, {"id": "quality_checks_passed", "statement": "Every quality-check result required for Line 4 after cleaning and before the currently scheduled Mint-500 run at 14:00 satisfies its corresponding signed pass criterion."}, {"id": "no_open_maintenance_fault", "statement": "No maintenance fault affecting Line 4 is open at the readiness decision for the currently scheduled Mint-500 run at 14:00."}], "base_state_json": "\"Line 4 is scheduled to change over from Citrus-500 to Mint-500 at 14:00. Bottles, caps, labels, and concentrate for this run are staged in the required quantities; a blue-cap shortage concerns only next week’s Ocean-750 run. The approved Mint tooling is installed, and four trained operators are assigned, meeting the scheduled staffing requirement. The required cleaning rinse is complete and its record is signed. Yesterday’s conveyor sensor repair was tested successfully and closed, while a worn spare guide remains tagged out in storage without affecting the line. For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18. At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU. The readiness review concerns this scheduled run.\"", "base_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "counter_states": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}], "focus_atom": "quality_checks_passed", "focus_evidence": [{"path": [], "text": "For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18."}, {"path": [], "text": "At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU."}], "policy_evidence": [], "rules": [{"justification": "All expressly required readiness conditions are satisfied, including the required post-clean quality checks.", "target": "release_to_production", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "supported"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}, {"justification": "A failed required quality check prevents release, while supported current-job material readiness means the sole blocker is not missing or deficient material and therefore does not qualify for materials support.", "target": "none_of_above", "when": [{"atom_id": "materials_ready", "state": "supported"}, {"atom_id": "tooling_approved", "state": "supported"}, {"atom_id": "cleaning_records_complete", "state": "supported"}, {"atom_id": "staffing_adequate", "state": "supported"}, {"atom_id": "quality_checks_passed", "state": "refuted"}, {"atom_id": "no_open_maintenance_fault", "state": "supported"}]}]}, "verified_pair": {"left": "For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18.", "negative_left": "For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18.", "negative_right": "At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 2.1 NTU.", "right": "At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 1.8 NTU."}, "verifier_independent_model": false}, "family": "fast-41-diverse-289-008", "id": "fast-41-diverse-289-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "Select when the line cannot be released and the evidence does not meet the materials-support rubric, including when an unresolved quality or maintenance requirement is the blocker.", "release_to_production": "Release the line to production only if every required readiness condition is satisfied, including a passed post-clean quality check.", "route_to_materials_support": "Do not release; route to materials support only if the sole blocker is material required for the currently scheduled Mint-500 run."}, "instructions": "Choose the single applicable routing/readiness outcome. Release requires current-job materials, approved tooling, completed cleaning records, adequate staffing, passed quality checks, and no open maintenance fault. Route to materials support only when the sole release blocker is missing or deficient material for the scheduled job.", "type": "choice"}}, "state": "Line 4 is scheduled to change over from Citrus-500 to Mint-500 at 14:00. Bottles, caps, labels, and concentrate for this run are staged in the required quantities; a blue-cap shortage concerns only next week’s Ocean-750 run. The approved Mint tooling is installed, and four trained operators are assigned, meeting the scheduled staffing requirement. The required cleaning rinse is complete and its record is signed. Yesterday’s conveyor sensor repair was tested successfully and closed, while a worn spare guide remains tagged out in storage without affecting the line. For Line 4’s currently scheduled Mint-500 run at 14:00, the required post-cleaning checks are QC-17 and QC-18, and their signed pass criteria are at least 7.0 pH units for QC-17 and at most 2.0 NTU for QC-18. At 13:42 on the date of Line 4’s currently scheduled Mint-500 run at 14:00, QC-17 recorded 7.2 pH units and QC-18 recorded 2.1 NTU. The readiness review concerns this scheduled run."}, "method": "c2d", "provenance": {"source_id": "diverse-289", "source_is_synthetic": true, "source_sha256": "884345667ad55a5ef86e09cc2ba68f6b13c8ccb20beb9fbed251a080f3d46c34", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged original questions retain all governing criteria and instructions. Question bindings remain fixed to the Citrus-500 to Berry-500 changeover at 14:00 and the release decision. The two focus spans are complete factual sentences: \"For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed.\" and \"For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified.\" The counterfactual coherently changes the specified kit to TK-418, creating a tooling mismatch without contradictory duplicate measurements. Neither context embeds an answer, label rationale, rule table, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"The Citrus-500 to Berry-500 changeover is scheduled for 14:00. Every material assigned to it is staged, including Berry concentrate lot B771 and 18,000 clean bottles for a planned requirement of 17,500.\"},{\"speaker\":\"Line supervisor\",\"text\":\"Six required operators are present, meeting the staffing level. Cleaning record CL-204 has the required signature.\"},{\"speaker\":\"Evidence record\",\"text\":\"For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed.\"},{\"speaker\":\"Evidence record\",\"text\":\"For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release has passed, including allergen swab Q-882 and label-code verification.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Every maintenance work order required before release is closed. Each has its required torque reading and technician signature recorded, including the completed verification for the filler guard.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed."}, {"path": ["3", "text"], "text": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed.", "negative_left": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed.", "negative_right": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-418 is specified.", "right": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified."}, "verifier_independent_model": false}, "family": "fast-41-diverse-291-002", "id": "fast-41-diverse-291-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "The Citrus-500 to Berry-500 changeover is scheduled for 14:00. Every material assigned to it is staged, including Berry concentrate lot B771 and 18,000 clean bottles for a planned requirement of 17,500."}, {"speaker": "Line supervisor", "text": "Six required operators are present, meeting the staffing level. Cleaning record CL-204 has the required signature."}, {"speaker": "Evidence record", "text": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed."}, {"speaker": "Evidence record", "text": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release has passed, including allergen swab Q-882 and label-code verification."}, {"speaker": "Maintenance technician", "text": "Every maintenance work order required before release is closed. Each has its required torque reading and technician signature recorded, including the completed verification for the filler guard."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the unchanged original questions retain all governing criteria and instructions. Question bindings remain fixed to the Citrus-500 to Berry-500 changeover at 14:00 and the release decision. The two focus spans are complete factual sentences: \"For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed.\" and \"For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified.\" The counterfactual coherently changes the specified kit to TK-418, creating a tooling mismatch without contradictory duplicate measurements. Neither context embeds an answer, label rationale, rule table, proposition identifier, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"The Citrus-500 to Berry-500 changeover is scheduled for 14:00. Every material assigned to it is staged, including Berry concentrate lot B771 and 18,000 clean bottles for a planned requirement of 17,500.\"},{\"speaker\":\"Line supervisor\",\"text\":\"Six required operators are present, meeting the staffing level. Cleaning record CL-204 has the required signature.\"},{\"speaker\":\"Evidence record\",\"text\":\"For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed.\"},{\"speaker\":\"Evidence record\",\"text\":\"For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release has passed, including allergen swab Q-882 and label-code verification.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Every maintenance work order required before release is closed. Each has its required torque reading and technician signature recorded, including the completed verification for the filler guard.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed."}, {"path": ["3", "text"], "text": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed.", "negative_left": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed.", "negative_right": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-418 is specified.", "right": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is specified."}, "verifier_independent_model": false}, "family": "fast-41-diverse-291-002", "id": "fast-41-diverse-291-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "The Citrus-500 to Berry-500 changeover is scheduled for 14:00. Every material assigned to it is staged, including Berry concentrate lot B771 and 18,000 clean bottles for a planned requirement of 17,500."}, {"speaker": "Line supervisor", "text": "Six required operators are present, meeting the staffing level. Cleaning record CL-204 has the required signature."}, {"speaker": "Evidence record", "text": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-417 is installed."}, {"speaker": "Evidence record", "text": "For the Citrus-500 to Berry-500 changeover at 14:00, tooling kit TK-418 is specified."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release has passed, including allergen swab Q-882 and label-code verification."}, {"speaker": "Maintenance technician", "text": "Every maintenance work order required before release is closed. Each has its required torque reading and technician signature recorded, including the completed verification for the filler guard."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, criteria, and thresholds. Both contexts preserve the Citrus-500 to Berry-500, 14:00, and tooling bindings. The evidence consists of two complete factual sentences: \"At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover.\" and \"The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583.\" The counterfactual coherently changes the specification to TK-584, creating a tooling mismatch without contradicting the installation fact. Neither context embeds a label, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"The Citrus-500 to Berry-500 changeover is scheduled for 14:00. Berry concentrate and all other scheduled materials are staged, including 18,000 clean bottles for a planned run requiring 17,500.\"},{\"speaker\":\"Line supervisor\",\"text\":\"Six required operators are present, matching the staffing requirement. Cleaning record CL-204 has the required signature.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release has passed, including the allergen swab, label check, and first-piece inspection.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Every maintenance work order required before release is closed. Each has its required torque reading and technician signature recorded.\"},{\"speaker\":\"Tooling coordinator\",\"text\":\"At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover. The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["4", "text"], "text": "At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover."}, {"path": ["4", "text"], "text": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover.", "negative_left": "At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover.", "negative_right": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-584.", "right": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583."}, "verifier_independent_model": false}, "family": "fast-41-diverse-291-003", "id": "fast-41-diverse-291-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "The Citrus-500 to Berry-500 changeover is scheduled for 14:00. Berry concentrate and all other scheduled materials are staged, including 18,000 clean bottles for a planned run requiring 17,500."}, {"speaker": "Line supervisor", "text": "Six required operators are present, matching the staffing requirement. Cleaning record CL-204 has the required signature."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release has passed, including the allergen swab, label check, and first-piece inspection."}, {"speaker": "Maintenance technician", "text": "Every maintenance work order required before release is closed. Each has its required torque reading and technician signature recorded."}, {"speaker": "Tooling coordinator", "text": "At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover. The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, criteria, and thresholds. Both contexts preserve the Citrus-500 to Berry-500, 14:00, and tooling bindings. The evidence consists of two complete factual sentences: \"At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover.\" and \"The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583.\" The counterfactual coherently changes the specification to TK-584, creating a tooling mismatch without contradicting the installation fact. Neither context embeds a label, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"The Citrus-500 to Berry-500 changeover is scheduled for 14:00. Berry concentrate and all other scheduled materials are staged, including 18,000 clean bottles for a planned run requiring 17,500.\"},{\"speaker\":\"Line supervisor\",\"text\":\"Six required operators are present, matching the staffing requirement. Cleaning record CL-204 has the required signature.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release has passed, including the allergen swab, label check, and first-piece inspection.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Every maintenance work order required before release is closed. Each has its required torque reading and technician signature recorded.\"},{\"speaker\":\"Tooling coordinator\",\"text\":\"At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover. The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["4", "text"], "text": "At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover."}, {"path": ["4", "text"], "text": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover.", "negative_left": "At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover.", "negative_right": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-584.", "right": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-583."}, "verifier_independent_model": false}, "family": "fast-41-diverse-291-003", "id": "fast-41-diverse-291-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "The Citrus-500 to Berry-500 changeover is scheduled for 14:00. Berry concentrate and all other scheduled materials are staged, including 18,000 clean bottles for a planned run requiring 17,500."}, {"speaker": "Line supervisor", "text": "Six required operators are present, matching the staffing requirement. Cleaning record CL-204 has the required signature."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release has passed, including the allergen swab, label check, and first-piece inspection."}, {"speaker": "Maintenance technician", "text": "Every maintenance work order required before release is closed. Each has its required torque reading and technician signature recorded."}, {"speaker": "Tooling coordinator", "text": "At 14:00, tooling kit TK-583 was installed for the Citrus-500 to Berry-500 changeover. The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 lists kit TK-584."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object preserves the governing policy, and both contexts retain the Citrus-500 to Berry-500, 14:00, and release-decision bindings. The evidence spans are two complete factual sentences: “At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover.” and “The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit.” The counterfactual coherently changes the required kit to TK-593 without contradictory duplicate claims or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"The Citrus-500 to Berry-500 changeover is scheduled for 14:00 on 17 September 2026. All scheduled materials are staged, including 18,000 clean bottles for a planned run requiring 17,500.\"},{\"speaker\":\"Line supervisor\",\"text\":\"Six required operators are present. Cleaning record CL-204 has the required signature.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release has passed.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Every maintenance work order required before release is closed, and each has its required torque reading and technician signature recorded.\"},{\"speaker\":\"Installation verifier\",\"text\":\"At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover. The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["4", "text"], "text": "At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover."}, {"path": ["4", "text"], "text": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover.", "negative_left": "At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover.", "negative_right": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-593 as the required kit.", "right": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit."}, "verifier_independent_model": false}, "family": "fast-41-diverse-291-008", "id": "fast-41-diverse-291-008-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "The Citrus-500 to Berry-500 changeover is scheduled for 14:00 on 17 September 2026. All scheduled materials are staged, including 18,000 clean bottles for a planned run requiring 17,500."}, {"speaker": "Line supervisor", "text": "Six required operators are present. Cleaning record CL-204 has the required signature."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release has passed."}, {"speaker": "Maintenance technician", "text": "Every maintenance work order required before release is closed, and each has its required torque reading and technician signature recorded."}, {"speaker": "Installation verifier", "text": "At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover. The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object preserves the governing policy, and both contexts retain the Citrus-500 to Berry-500, 14:00, and release-decision bindings. The evidence spans are two complete factual sentences: “At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover.” and “The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit.” The counterfactual coherently changes the required kit to TK-593 without contradictory duplicate claims or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual readiness relationship; universal quantification over the relevant explicit set does not make an atom an improper bundle. The focus atom is the factual identity match between installed and specified tooling, not a policy classification. The base and counter assignments are realizable with only that tooling fact changing. Empty policy_evidence is correct because all governing criteria are in the retained questions object; no additional substantive rule from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction affirmatively covers every Green requirement: all scheduled materials and sufficient bottles, required staffing, signed cleaning record, correct installed tooling, all quality prerequisites passed, and all required maintenance work orders closed with torque readings and technician signatures. It is sufficient for release.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the tooling-match atom establishes that the installed tooling is not the specified tooling. This defeats the Green requirement for correct tooling and is sufficient for a No-release decision, regardless of the other satisfied readiness items.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material scheduled for the Citrus-500 to Berry-500 changeover at 14:00 is staged for that changeover."}, {"id": "a2", "statement": "The number of clean bottles staged for the Citrus-500 to Berry-500 changeover at 14:00 is at least the number of bottles required for the planned Berry-500 run."}, {"id": "a3", "statement": "The number of required operators present for the Citrus-500 to Berry-500 changeover at 14:00 is at least the required staffing level for that changeover."}, {"id": "a4", "statement": "Cleaning record CL-204 for the Citrus-500 to Berry-500 changeover at 14:00 has the required signature."}, {"id": "a5", "statement": "The tooling kit installed for the Citrus-500 to Berry-500 changeover at 14:00 is the same tooling kit specified for that changeover."}, {"id": "a6", "statement": "Every quality prerequisite required before release of the Citrus-500 to Berry-500 changeover at 14:00 has passed."}, {"id": "a7", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 is closed."}, {"id": "a8", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required torque reading recorded."}, {"id": "a9", "statement": "Every maintenance work order required before release of the Citrus-500 to Berry-500 changeover at 14:00 has its required technician signature recorded."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"The Citrus-500 to Berry-500 changeover is scheduled for 14:00 on 17 September 2026. All scheduled materials are staged, including 18,000 clean bottles for a planned run requiring 17,500.\"},{\"speaker\":\"Line supervisor\",\"text\":\"Six required operators are present. Cleaning record CL-204 has the required signature.\"},{\"speaker\":\"Quality inspector\",\"text\":\"Every quality prerequisite required before release has passed.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Every maintenance work order required before release is closed, and each has its required torque reading and technician signature recorded.\"},{\"speaker\":\"Installation verifier\",\"text\":\"At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover. The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["4", "text"], "text": "At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover."}, {"path": ["4", "text"], "text": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit."}], "policy_evidence": [], "rules": [{"justification": "All required readiness categories have explicit affirmative evidence, including a match between the installed tooling and the specified tooling, so the changeover is Green and releasable.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "The installed tooling is explicitly not the tooling specified for the changeover. Correct tooling is therefore physically unavailable as installed, making the changeover Red and not releasable even though every other readiness category is satisfied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover.", "negative_left": "At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover.", "negative_right": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-593 as the required kit.", "right": "The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-417 as the required kit."}, "verifier_independent_model": false}, "family": "fast-41-diverse-291-008", "id": "fast-41-diverse-291-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the changeover Amber or Red, do not release the line, and route the unresolved item to the responsible support function.", "true": "Yes — classify the changeover Green and release the line to production because every required readiness item has explicit evidence."}, "instructions": "Decide whether the line may be released to production now. Apply this rubric: Green means every scheduled material, required staffing level, signed cleaning record, correct tooling, passed quality prerequisite, and closed maintenance work order with recorded torque and technician signature has explicit evidence; release is Yes. Amber means only documentation or verification remains and release is No, routed to the responsible support function. Red means a physical resource, safety repair, cleaning, or quality result is missing and release is No. Is the line Green and releasable now?", "type": "noul"}}, "state": [{"speaker": "Production scheduler", "text": "The Citrus-500 to Berry-500 changeover is scheduled for 14:00 on 17 September 2026. All scheduled materials are staged, including 18,000 clean bottles for a planned run requiring 17,500."}, {"speaker": "Line supervisor", "text": "Six required operators are present. Cleaning record CL-204 has the required signature."}, {"speaker": "Quality inspector", "text": "Every quality prerequisite required before release has passed."}, {"speaker": "Maintenance technician", "text": "Every maintenance work order required before release is closed, and each has its required torque reading and technician signature recorded."}, {"speaker": "Installation verifier", "text": "At 14:00 on 17 September 2026, tooling kit TK-417 was installed for the Citrus-500 to Berry-500 changeover. The tooling specification for the Citrus-500 to Berry-500 changeover at 14:00 on 17 September 2026 identifies kit TK-593 as the required kit."}]}, "method": "c2d", "provenance": {"source_id": "diverse-291", "source_is_synthetic": true, "source_sha256": "d25493b46bcb0f69a8d470239155097c31ff29a07796c0fffca34322d379272a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing release policy. Both contexts remain bound to Line 4, its valve-cap change, and the stated date and time. The two evidence spans are complete factual sentences: “At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps.” “The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions.” The counterfactual is coherent because seven assigned operators and an eight-position requirement create a staffing shortfall without contradictory duplicate assertions. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00 on 7 June 2026. At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps. The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions. Before release, all required white resin, labels, and other materials were staged. The required cleaning was completed at 13:20, the specified mold was installed, and every required tooling check passed. No required maintenance work remained open at the release decision, and no maintenance defect affecting Line 4 remained unresolved. The quality inspector approved the first piece for the scheduled production of white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps."}, {"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps.", "negative_left": "At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps.", "negative_right": "The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires eight operator positions.", "right": "The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-002", "id": "fast-41-diverse-292-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00 on 7 June 2026. At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps. The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions. Before release, all required white resin, labels, and other materials were staged. The required cleaning was completed at 13:20, the specified mold was installed, and every required tooling check passed. No required maintenance work remained open at the release decision, and no maintenance defect affecting Line 4 remained unresolved. The quality inspector approved the first piece for the scheduled production of white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object and both contexts preserve the governing release policy. Both contexts remain bound to Line 4, its valve-cap change, and the stated date and time. The two evidence spans are complete factual sentences: “At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps.” “The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions.” The counterfactual is coherent because seven assigned operators and an eight-position requirement create a staffing shortfall without contradictory duplicate assertions. Neither context contains a gold answer, answer code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00 on 7 June 2026. At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps. The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions. Before release, all required white resin, labels, and other materials were staged. The required cleaning was completed at 13:20, the specified mold was installed, and every required tooling check passed. No required maintenance work remained open at the release decision, and no maintenance defect affecting Line 4 remained unresolved. The quality inspector approved the first piece for the scheduled production of white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps."}, {"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps.", "negative_left": "At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps.", "negative_right": "The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires eight operator positions.", "right": "The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires seven operator positions."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-002", "id": "fast-41-diverse-292-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00 on 7 June 2026. At 13:42 on 7 June 2026, seven trained operators were assigned to Line 4 for the scheduled 14:00 change from blue valve caps to white valve caps. The signed staffing specification for Line 4's scheduled 14:00 change from blue valve caps to white valve caps requires eight operator positions. Before release, all required white resin, labels, and other materials were staged. The required cleaning was completed at 13:20, the specified mold was installed, and every required tooling check passed. No required maintenance work remained open at the release decision, and no maintenance defect affecting Line 4 remained unresolved. The quality inspector approved the first piece for the scheduled production of white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing policy without invented exceptions or defaults. Both contexts preserve Line 4, the 14:00 change, and the release decision scope. The evidence contains exactly two complete factual sentences: \"The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions.\" and \"The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release.\" The counterfactual coherently changes the assigned roster to 6 operators, creating a staffing shortfall without contradictory duplicate measurements. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00. Materials support staged all required white resin, caps, and labels before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions. The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release. Cleaning was completed and signed at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance confirmed that yesterday's temperature alarm did not recur during today's test; no required maintenance work remains open, and no maintenance defect affecting Line 4 remains unresolved. The quality inspector approved the first piece for white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions."}, {"path": [], "text": "The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions.", "negative_right": "The staffing roster for Line 4's scheduled 14:00 change lists 6 trained operators assigned before release.", "right": "The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-003", "id": "fast-41-diverse-292-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00. Materials support staged all required white resin, caps, and labels before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions. The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release. Cleaning was completed and signed at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance confirmed that yesterday's temperature alarm did not recur during today's test; no required maintenance work remains open, and no maintenance defect affecting Line 4 remains unresolved. The quality inspector approved the first piece for white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions and all governing policy without invented exceptions or defaults. Both contexts preserve Line 4, the 14:00 change, and the release decision scope. The evidence contains exactly two complete factual sentences: \"The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions.\" and \"The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release.\" The counterfactual coherently changes the assigned roster to 6 operators, creating a staffing shortfall without contradictory duplicate measurements. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00. Materials support staged all required white resin, caps, and labels before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions. The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release. Cleaning was completed and signed at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance confirmed that yesterday's temperature alarm did not recur during today's test; no required maintenance work remains open, and no maintenance defect affecting Line 4 remains unresolved. The quality inspector approved the first piece for white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions."}, {"path": [], "text": "The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions.", "negative_right": "The staffing roster for Line 4's scheduled 14:00 change lists 6 trained operators assigned before release.", "right": "The staffing roster for Line 4's scheduled 14:00 change lists 7 trained operators assigned before release."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-003", "id": "fast-41-diverse-292-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00. Materials support staged all required white resin, caps, and labels before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 7 trained-operator positions. The staffing roster for Line 4's scheduled 14:00 change lists 6 trained operators assigned before release. Cleaning was completed and signed at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance confirmed that yesterday's temperature alarm did not recur during today's test; no required maintenance work remains open, and no maintenance defect affecting Line 4 remains unresolved. The quality inspector approved the first piece for white valve caps before release. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and yes-or-no release request. Both contexts retain Line 4, the 14:00 release, and the scheduled change bindings. The evidence consists of exactly two complete factual sentences: \"At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change.\" and \"The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions.\" The counterfactual consistently changes the staffing requirement to seven positions without creating duplicate or contradictory measurements. Neither context includes a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00. Before the release review, materials support logged all required white resin and labels as staged. At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change. The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions. The cleaning record was signed at 13:20. The specified mold was installed, and every required tooling check had passed. Maintenance records showed no required work open and no unresolved defect affecting Line 4; a prior temperature alarm did not recur during testing. The quality inspector approved the first piece for the white-cap run. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change."}, {"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change.", "negative_left": "At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change.", "negative_right": "The signed staffing specification for Line 4's scheduled 14:00 change requires exactly seven operator positions.", "right": "The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-005", "id": "fast-41-diverse-292-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00. Before the release review, materials support logged all required white resin and labels as staged. At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change. The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions. The cleaning record was signed at 13:20. The specified mold was installed, and every required tooling check had passed. Maintenance records showed no required work open and no unresolved defect affecting Line 4; a prior temperature alarm did not recur during testing. The quality inspector approved the first piece for the white-cap run. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and yes-or-no release request. Both contexts retain Line 4, the 14:00 release, and the scheduled change bindings. The evidence consists of exactly two complete factual sentences: \"At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change.\" and \"The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions.\" The counterfactual consistently changes the staffing requirement to seven positions without creating duplicate or contradictory measurements. Neither context includes a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00. Before the release review, materials support logged all required white resin and labels as staged. At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change. The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions. The cleaning record was signed at 13:20. The specified mold was installed, and every required tooling check had passed. Maintenance records showed no required work open and no unresolved defect affecting Line 4; a prior temperature alarm did not recur during testing. The quality inspector approved the first piece for the white-cap run. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change."}, {"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change.", "negative_left": "At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change.", "negative_right": "The signed staffing specification for Line 4's scheduled 14:00 change requires exactly seven operator positions.", "right": "The signed staffing specification for Line 4's scheduled 14:00 change requires exactly five operator positions."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-005", "id": "fast-41-diverse-292-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00. Before the release review, materials support logged all required white resin and labels as staged. At the 14:00 release decision, the roster for Line 4 lists exactly six trained operators assigned to the scheduled change. The signed staffing specification for Line 4's scheduled 14:00 change requires exactly seven operator positions. The cleaning record was signed at 13:20. The specified mold was installed, and every required tooling check had passed. Maintenance records showed no required work open and no unresolved defect affecting Line 4; a prior temperature alarm did not recur during testing. The quality inspector approved the first piece for the white-cap run. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain it without exceptions or altered scope. Both contexts retain Line 4, release, the 14:00 change, and the dated operational path. The evidence consists of exactly two complete factual sentences: “The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change.” and “The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions.” The counterfactual coherently changes the required staffing count to eight while retaining seven assigned operators. Neither context embeds an answer, code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00 on 17 September 2026. Materials support staged every required white resin, caps, and labels before release. The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change. The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions. The signed cleaning record shows completion at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance logged no open work, and yesterday's temperature alarm did not recur or remain an unresolved defect. The quality inspector approved the first piece for white valve caps. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change."}, {"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_left": "The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_right": "The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists eight required operator positions.", "right": "The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-006", "id": "fast-41-diverse-292-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00 on 17 September 2026. Materials support staged every required white resin, caps, and labels before release. The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change. The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions. The signed cleaning record shows completion at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance logged no open work, and yesterday's temperature alarm did not recur or remain an unresolved defect. The quality inspector approved the first piece for white valve caps. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain it without exceptions or altered scope. Both contexts retain Line 4, release, the 14:00 change, and the dated operational path. The evidence consists of exactly two complete factual sentences: “The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change.” and “The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions.” The counterfactual coherently changes the required staffing count to eight while retaining seven assigned operators. Neither context embeds an answer, code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00 on 17 September 2026. Materials support staged every required white resin, caps, and labels before release. The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change. The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions. The signed cleaning record shows completion at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance logged no open work, and yesterday's temperature alarm did not recur or remain an unresolved defect. The quality inspector approved the first piece for white valve caps. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change."}, {"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_left": "The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change.", "negative_right": "The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists eight required operator positions.", "right": "The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists seven required operator positions."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-006", "id": "fast-41-diverse-292-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components scheduled Line 4 to change from blue valve caps to white valve caps at 14:00 on 17 September 2026. Materials support staged every required white resin, caps, and labels before release. The 17 September 2026 roster records seven trained operators assigned to Line 4 for the scheduled 14:00 change. The signed staffing specification for Line 4's scheduled 14:00 change on 17 September 2026 lists eight required operator positions. The signed cleaning record shows completion at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance logged no open work, and yesterday's temperature alarm did not recur or remain an unresolved defect. The quality inspector approved the first piece for white valve caps. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and both contexts retain it. Both contexts remain bound to Line 4, the 14:00 change, and the blue-to-white valve-cap transition. The evidence contains exactly two complete factual sentences: \"The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions.\" and \"The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change.\" The counterfactual coherently changes the assignment count to 8, which is below the stated requirement of 11. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components is preparing Line 4 for its scheduled 14:00 change from blue valve caps to white valve caps. Materials support staged every required material before release. The signed cleaning record confirms that required cleaning finished at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance reports show that no required work remains open and no defect affecting Line 4 is unresolved. The quality inspector approved the first piece for white valve-cap production. The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions. The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions."}, {"path": [], "text": "The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions.", "negative_right": "The 14:00 assignment roster records 8 trained operators assigned to Line 4 for that change.", "right": "The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-008", "id": "fast-41-diverse-292-008-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components is preparing Line 4 for its scheduled 14:00 change from blue valve caps to white valve caps. Materials support staged every required material before release. The signed cleaning record confirms that required cleaning finished at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance reports show that no required work remains open and no defect affecting Line 4 is unresolved. The quality inspector approved the first piece for white valve-cap production. The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions. The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and both contexts retain it. Both contexts remain bound to Line 4, the 14:00 change, and the blue-to-white valve-cap transition. The evidence contains exactly two complete factual sentences: \"The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions.\" and \"The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change.\" The counterfactual coherently changes the assignment count to 8, which is below the stated requirement of 11. Neither context embeds a gold answer, answer code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"Northstar Components is preparing Line 4 for its scheduled 14:00 change from blue valve caps to white valve caps. Materials support staged every required material before release. The signed cleaning record confirms that required cleaning finished at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance reports show that no required work remains open and no defect affecting Line 4 is unresolved. The quality inspector approved the first piece for white valve-cap production. The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions. The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions."}, {"path": [], "text": "The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions.", "negative_right": "The 14:00 assignment roster records 8 trained operators assigned to Line 4 for that change.", "right": "The 14:00 assignment roster records 13 trained operators assigned to Line 4 for that change."}, "verifier_independent_model": false}, "family": "fast-41-diverse-292-008", "id": "fast-41-diverse-292-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "Northstar Components is preparing Line 4 for its scheduled 14:00 change from blue valve caps to white valve caps. Materials support staged every required material before release. The signed cleaning record confirms that required cleaning finished at 13:20. The specified mold is installed, and every required tooling check has passed. Maintenance reports show that no required work remains open and no defect affecting Line 4 is unresolved. The quality inspector approved the first piece for white valve-cap production. The signed staffing specification for Line 4's scheduled 14:00 change requires 11 operator positions. The 14:00 assignment roster records 8 trained operators assigned to Line 4 for that change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the complete rubric and policy, and both contexts retain the same release rule. Both contexts preserve the Line 4, 07:00, readiness, and release bindings. The required evidence quotes are “The 07:00 release-decision record for Line 4 lists maintenance status code MH-27.” and “The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00.” Both are complete factual sentences. The counterfactual coherently changes the register’s interpretation from an active hold to a cleared hold without creating duplicate or contradictory assertions. Neither context states a readiness level, answer code, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"context\":\"At 06:30, the Line 4 changeover to 750 mL bottles was logged for a 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass. The 07:00 release-decision record lists all scheduled materials, tools, cleaning approval, operators, and quality checks as complete. The inventory and utility logs show every required resource available, and the exception review records no confirmed failure among the stated conditions. The quality-control register records no quality hold at the decision time. The 07:00 release-decision record for Line 4 lists maintenance status code MH-27. The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00. The record was prepared for the scheduler’s release decision, with no later status update included. Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\",\"request\":\"Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["context"], "text": "The 07:00 release-decision record for Line 4 lists maintenance status code MH-27."}, {"path": ["context"], "text": "The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 07:00 release-decision record for Line 4 lists maintenance status code MH-27.", "negative_left": "The 07:00 release-decision record for Line 4 lists maintenance status code MH-27.", "negative_right": "The Line 4 maintenance-control register identifies code MH-27 as meaning “cleared maintenance hold” for release decisions at 07:00.", "right": "The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00."}, "verifier_independent_model": false}, "family": "fast-41-diverse-293-001", "id": "fast-41-diverse-293-001-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"context": "At 06:30, the Line 4 changeover to 750 mL bottles was logged for a 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass. The 07:00 release-decision record lists all scheduled materials, tools, cleaning approval, operators, and quality checks as complete. The inventory and utility logs show every required resource available, and the exception review records no confirmed failure among the stated conditions. The quality-control register records no quality hold at the decision time. The 07:00 release-decision record for Line 4 lists maintenance status code MH-27. The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00. The record was prepared for the scheduler’s release decision, with no later status update included. Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.", "request": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the complete rubric and policy, and both contexts retain the same release rule. Both contexts preserve the Line 4, 07:00, readiness, and release bindings. The required evidence quotes are “The 07:00 release-decision record for Line 4 lists maintenance status code MH-27.” and “The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00.” Both are complete factual sentences. The counterfactual coherently changes the register’s interpretation from an active hold to a cleared hold without creating duplicate or contradictory assertions. Neither context states a readiness level, answer code, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"context\":\"At 06:30, the Line 4 changeover to 750 mL bottles was logged for a 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass. The 07:00 release-decision record lists all scheduled materials, tools, cleaning approval, operators, and quality checks as complete. The inventory and utility logs show every required resource available, and the exception review records no confirmed failure among the stated conditions. The quality-control register records no quality hold at the decision time. The 07:00 release-decision record for Line 4 lists maintenance status code MH-27. The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00. The record was prepared for the scheduler’s release decision, with no later status update included. Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\",\"request\":\"Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["context"], "text": "The 07:00 release-decision record for Line 4 lists maintenance status code MH-27."}, {"path": ["context"], "text": "The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 07:00 release-decision record for Line 4 lists maintenance status code MH-27.", "negative_left": "The 07:00 release-decision record for Line 4 lists maintenance status code MH-27.", "negative_right": "The Line 4 maintenance-control register identifies code MH-27 as meaning “cleared maintenance hold” for release decisions at 07:00.", "right": "The Line 4 maintenance-control register identifies code MH-27 as meaning “active maintenance hold” for release decisions at 07:00."}, "verifier_independent_model": false}, "family": "fast-41-diverse-293-001", "id": "fast-41-diverse-293-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"context": "At 06:30, the Line 4 changeover to 750 mL bottles was logged for a 07:00 production run. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass. The 07:00 release-decision record lists all scheduled materials, tools, cleaning approval, operators, and quality checks as complete. The inventory and utility logs show every required resource available, and the exception review records no confirmed failure among the stated conditions. The quality-control register records no quality hold at the decision time. The 07:00 release-decision record for Line 4 lists maintenance status code MH-27. The Line 4 maintenance-control register identifies code MH-27 as meaning “cleared maintenance hold” for release decisions at 07:00. The record was prepared for the scheduler’s release decision, with no later status update included. Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.", "request": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the unchanged question and policy, preserve Line 4, MH-482, 07:00, and release-decision bindings, and contain no answer leakage; the two evidence spans are complete factual sentences—\"Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED.\" and \"At the 07:00 release decision, record MH-482 has status ACTIVE.\"—while the counterfactual coherently changes only the status to CLOSED.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"case_note\":\"Line 4 is scheduled for the 07:00 production run after its bottle changeover. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass. The staging log confirms all scheduled bottles, caps, labels, and resin were recorded in place before 07:00. The tooling log records verification completion, while the sanitation log records accepted cleaning. The staffing roster shows every assigned operator present, and quality records show all required checks passed. The release checklist marks every required resource available and records no confirmed failure among the stated conditions. The quality review contains no quality hold for Line 4. Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED. At the 07:00 release decision, record MH-482 has status ACTIVE. Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["case_note"], "text": "Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED."}, {"path": ["case_note"], "text": "At the 07:00 release decision, record MH-482 has status ACTIVE."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED.", "negative_left": "Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED.", "negative_right": "At the 07:00 release decision, record MH-482 has status CLOSED.", "right": "At the 07:00 release decision, record MH-482 has status ACTIVE."}, "verifier_independent_model": false}, "family": "fast-41-diverse-293-006", "id": "fast-41-diverse-293-006-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"case_note": "Line 4 is scheduled for the 07:00 production run after its bottle changeover. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass. The staging log confirms all scheduled bottles, caps, labels, and resin were recorded in place before 07:00. The tooling log records verification completion, while the sanitation log records accepted cleaning. The staffing roster shows every assigned operator present, and quality records show all required checks passed. The release checklist marks every required resource available and records no confirmed failure among the stated conditions. The quality review contains no quality hold for Line 4. Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED. At the 07:00 release decision, record MH-482 has status ACTIVE. Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both full contexts retain the unchanged question and policy, preserve Line 4, MH-482, 07:00, and release-decision bindings, and contain no answer leakage; the two evidence spans are complete factual sentences—\"Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED.\" and \"At the 07:00 release decision, record MH-482 has status ACTIVE.\"—while the counterfactual coherently changes only the status to CLOSED.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"case_note\":\"Line 4 is scheduled for the 07:00 production run after its bottle changeover. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass. The staging log confirms all scheduled bottles, caps, labels, and resin were recorded in place before 07:00. The tooling log records verification completion, while the sanitation log records accepted cleaning. The staffing roster shows every assigned operator present, and quality records show all required checks passed. The release checklist marks every required resource available and records no confirmed failure among the stated conditions. The quality review contains no quality hold for Line 4. Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED. At the 07:00 release decision, record MH-482 has status ACTIVE. Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["case_note"], "text": "Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED."}, {"path": ["case_note"], "text": "At the 07:00 release decision, record MH-482 has status ACTIVE."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED.", "negative_left": "Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED.", "negative_right": "At the 07:00 release decision, record MH-482 has status CLOSED.", "right": "At the 07:00 release decision, record MH-482 has status ACTIVE."}, "verifier_independent_model": false}, "family": "fast-41-diverse-293-006", "id": "fast-41-diverse-293-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"case_note": "Line 4 is scheduled for the 07:00 production run after its bottle changeover. The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass. The staging log confirms all scheduled bottles, caps, labels, and resin were recorded in place before 07:00. The tooling log records verification completion, while the sanitation log records accepted cleaning. The staffing roster shows every assigned operator present, and quality records show all required checks passed. The release checklist marks every required resource available and records no confirmed failure among the stated conditions. The quality review contains no quality hold for Line 4. Record MH-482 is the Line 4 maintenance-hold record for the 07:00 release decision, and its status field has exactly two permitted values: ACTIVE and CLOSED. At the 07:00 release decision, record MH-482 has status CLOSED. Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The policy, request, Line 4 entity, and 07:00 time remain unchanged. The required evidence quotes are complete factual sentences: “The Line 4 control log records hold code M-17 at the 07:00 release decision.” and “The Line 4 control manual defines code M-17 as an active maintenance hold.” The counterfactual consistently changes the code definition to a cleared maintenance hold without contradictory duplicate claims. Neither context embeds an answer, label, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"case_note\":{\"title\":\"Line 4 release review\",\"time\":\"07:00\",\"policy\":\"The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.\",\"evidence\":[\"The Line 4 control log records hold code M-17 at the 07:00 release decision.\",\"The Line 4 control manual defines code M-17 as an active maintenance hold.\",\"By 06:25, the materials coordinator documented all bottles, caps, labels, and resin needed for the run as staged.\",\"A technician completed and logged the required tooling verification at 06:30.\",\"The sanitation lead accepted the scheduled cleaning record at 06:20.\",\"The shift roster shows every assigned Line 4 operator present before the release review.\",\"Quality records show each scheduled check passed, including the first-piece inspection.\",\"The operations desk confirmed all required personnel, equipment, materials, and utilities were available for the decision.\",\"The release checklist contains no confirmed failure for any stated condition.\",\"The quality desk reported no active quality hold at the decision time.\"],\"request\":\"Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["case_note", "evidence", "0"], "text": "The Line 4 control log records hold code M-17 at the 07:00 release decision."}, {"path": ["case_note", "evidence", "1"], "text": "The Line 4 control manual defines code M-17 as an active maintenance hold."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The Line 4 control log records hold code M-17 at the 07:00 release decision.", "negative_left": "The Line 4 control log records hold code M-17 at the 07:00 release decision.", "negative_right": "The Line 4 control manual defines code M-17 as a cleared maintenance hold.", "right": "The Line 4 control manual defines code M-17 as an active maintenance hold."}, "verifier_independent_model": false}, "family": "fast-41-diverse-293-012", "id": "fast-41-diverse-293-012-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"case_note": {"evidence": ["The Line 4 control log records hold code M-17 at the 07:00 release decision.", "The Line 4 control manual defines code M-17 as an active maintenance hold.", "By 06:25, the materials coordinator documented all bottles, caps, labels, and resin needed for the run as staged.", "A technician completed and logged the required tooling verification at 06:30.", "The sanitation lead accepted the scheduled cleaning record at 06:20.", "The shift roster shows every assigned Line 4 operator present before the release review.", "Quality records show each scheduled check passed, including the first-piece inspection.", "The operations desk confirmed all required personnel, equipment, materials, and utilities were available for the decision.", "The release checklist contains no confirmed failure for any stated condition.", "The quality desk reported no active quality hold at the decision time."], "policy": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.", "request": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.", "time": "07:00", "title": "Line 4 release review"}}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The policy, request, Line 4 entity, and 07:00 time remain unchanged. The required evidence quotes are complete factual sentences: “The Line 4 control log records hold code M-17 at the 07:00 release decision.” and “The Line 4 control manual defines code M-17 as an active maintenance hold.” The counterfactual consistently changes the code definition to a cleared maintenance hold without contradictory duplicate claims. Neither context embeds an answer, label, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; universally quantified atoms such as availability of every required resource remain atomic. The focus, whether a maintenance hold is active, is factual rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable: completed release conditions and available resources can coexist with an independently imposed maintenance hold, while removing that hold yields the no-hold case. Policy evidence preserves the scheduler's state-originated list of release conditions needed to interpret the unchanged rubric; the additional request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "An active maintenance hold directly satisfies the Level 2 criterion regardless of the status of the other release conditions.", "rule_index": 0, "sound": true}, {"reason": "The conjunction documents completion or approval of every scheduler-stated release condition and also excludes every competing Level 2 trigger: confirmed condition failure, unavailable required resources, and active maintenance or quality holds. It is therefore sufficient for Level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every material required for Line 4's scheduled 07:00 production run is documented as staged by 07:00."}, {"id": "a2", "statement": "Tooling verification required for Line 4's scheduled 07:00 production run is documented as complete by 07:00."}, {"id": "a3", "statement": "Cleaning required for Line 4's scheduled 07:00 production run is documented as accepted by 07:00."}, {"id": "a4", "statement": "Every staff member required for Line 4's scheduled 07:00 production run is documented as present by 07:00."}, {"id": "a5", "statement": "Every quality check required for Line 4's scheduled 07:00 production run is documented as passed by 07:00."}, {"id": "a6", "statement": "Every resource required for Line 4's scheduled 07:00 production run is available at the 07:00 release decision."}, {"id": "a7", "statement": "No stated release condition for Line 4's scheduled 07:00 production run has a confirmed failure at the 07:00 release decision."}, {"id": "a8", "statement": "A maintenance hold is active on Line 4 at the 07:00 release decision."}, {"id": "a9", "statement": "A quality hold is active on Line 4 at the 07:00 release decision."}], "base_state_json": "{\"case_note\":{\"title\":\"Line 4 release review\",\"time\":\"07:00\",\"policy\":\"The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.\",\"evidence\":[\"The Line 4 control log records hold code M-17 at the 07:00 release decision.\",\"The Line 4 control manual defines code M-17 as an active maintenance hold.\",\"By 06:25, the materials coordinator documented all bottles, caps, labels, and resin needed for the run as staged.\",\"A technician completed and logged the required tooling verification at 06:30.\",\"The sanitation lead accepted the scheduled cleaning record at 06:20.\",\"The shift roster shows every assigned Line 4 operator present before the release review.\",\"Quality records show each scheduled check passed, including the first-piece inspection.\",\"The operations desk confirmed all required personnel, equipment, materials, and utilities were available for the decision.\",\"The release checklist contains no confirmed failure for any stated condition.\",\"The quality desk reported no active quality hold at the decision time.\"],\"request\":\"Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.\"}}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["case_note", "evidence", "0"], "text": "The Line 4 control log records hold code M-17 at the 07:00 release decision."}, {"path": ["case_note", "evidence", "1"], "text": "The Line 4 control manual defines code M-17 as an active maintenance hold."}], "policy_evidence": [{"path": ["context"], "text": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass."}, {"path": ["request"], "text": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00."}], "rules": [{"justification": "An active maintenance hold independently satisfies the Level 2 definition, so Line 4 must not be released.", "target": "2", "when": [{"atom_id": "a8", "state": "supported"}]}, {"justification": "Every scheduler-stated release condition has documented completion or approval, every required resource is available, no stated condition has a confirmed failure, and neither maintenance nor quality holds are active; therefore Level 0 requires release at 07:00.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The Line 4 control log records hold code M-17 at the 07:00 release decision.", "negative_left": "The Line 4 control log records hold code M-17 at the 07:00 release decision.", "negative_right": "The Line 4 control manual defines code M-17 as a cleared maintenance hold.", "right": "The Line 4 control manual defines code M-17 as an active maintenance hold."}, "verifier_independent_model": false}, "family": "fast-41-diverse-293-012", "id": "fast-41-diverse-293-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Ready, low delay risk: Every stated release condition has documented completion or approval; release the line at the scheduled time.", "Level 1 — Conditional readiness, moderate delay risk: No confirmed failure exists, but at least one stated release condition lacks completion evidence; hold release and route to the responsible function.", "Level 2 — Not ready, high delay risk: At least one stated release condition has failed, required resources are unavailable, or an active maintenance or quality hold exists; do not release."], "instructions": "Apply the scheduler’s conditional release intent literally. Select the single level whose definition matches the documented status; lower levels mean lower delay risk.", "type": "score"}}, "state": {"case_note": {"evidence": ["The Line 4 control log records hold code M-17 at the 07:00 release decision.", "The Line 4 control manual defines code M-17 as a cleared maintenance hold.", "By 06:25, the materials coordinator documented all bottles, caps, labels, and resin needed for the run as staged.", "A technician completed and logged the required tooling verification at 06:30.", "The sanitation lead accepted the scheduled cleaning record at 06:20.", "The shift roster shows every assigned Line 4 operator present before the release review.", "Quality records show each scheduled check passed, including the first-piece inspection.", "The operations desk confirmed all required personnel, equipment, materials, and utilities were available for the decision.", "The release checklist contains no confirmed failure for any stated condition.", "The quality desk reported no active quality hold at the decision time."], "policy": "The production scheduler stated: release at 07:00 only if all materials are staged, tooling verification is complete, cleaning is accepted, required staff are present, and quality checks pass.", "request": "Using the ordered delay-risk rubric, determine the readiness level and whether Line 4 should be released at 07:00.", "time": "07:00", "title": "Line 4 release review"}}}, "method": "c2d", "provenance": {"source_id": "diverse-293", "source_is_synthetic": true, "source_sha256": "e2ff49ed2216acaa5a8456b055977b902b3da60396f0ae9b5cf777c84adacc72", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain L4 and 06:52 bindings. The evidence quotes are complete factual sentences: \"As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.\" and \"As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin.\" The counterfactual coherently changes the requirement to 21 liters while retaining 18 liters staged, without contradictory duplicates. Neither context contains a gold answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Inventory log\",\"text\":\"As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"First-lot worksheet\",\"text\":\"As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:52, no current stop-work hold applied to changeover L4; the earlier equipment concern had been cleared after successful checks.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the reserve-resin delivery remained delayed and was documented as an issue. The delay could disrupt continued production, but it did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the approved first-lot worksheet for changeover L4 requires 21 liters of resin.", "right": "As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-001", "id": "fast-41-diverse-294-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Inventory log", "text": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "First-lot worksheet", "text": "As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for changeover L4."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "At 06:52, no current stop-work hold applied to changeover L4; the earlier equipment concern had been cleared after successful checks."}, {"speaker": "Materials coordinator", "text": "At 06:52, the reserve-resin delivery remained delayed and was documented as an issue. The delay could disrupt continued production, but it did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves the governing policy, and both contexts retain L4 and 06:52 bindings. The evidence quotes are complete factual sentences: \"As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.\" and \"As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin.\" The counterfactual coherently changes the requirement to 21 liters while retaining 18 liters staged, without contradictory duplicates. Neither context contains a gold answer, code, rule table, proposition ID, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Inventory log\",\"text\":\"As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"First-lot worksheet\",\"text\":\"As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:52, no current stop-work hold applied to changeover L4; the earlier equipment concern had been cleared after successful checks.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the reserve-resin delivery remained delayed and was documented as an issue. The delay could disrupt continued production, but it did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the approved first-lot worksheet for changeover L4 requires 21 liters of resin.", "right": "As of 06:52, the approved first-lot worksheet for changeover L4 requires 15 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-001", "id": "fast-41-diverse-294-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Inventory log", "text": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "First-lot worksheet", "text": "As of 06:52, the approved first-lot worksheet for changeover L4 requires 21 liters of resin."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present for changeover L4."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "At 06:52, no current stop-work hold applied to changeover L4; the earlier equipment concern had been cleared after successful checks."}, {"speaker": "Materials coordinator", "text": "At 06:52, the reserve-resin delivery remained delayed and was documented as an issue. The delay could disrupt continued production, but it did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object remains unchanged and neither context alters governing rules. Question bindings remain fixed to changeover L4 at 06:52 and its specified release gates. Evidence spans are complete factual sentences: “The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B.” and “The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms.” The counterfactual is coherent because its 19.4-kilogram latest record supersedes earlier inventory figures and is internally below the stated 22.0-kilogram requirement. Neither context embeds an answer, code, rule table, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Materials clerk\",\"text\":\"The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Inventory auditor\",\"text\":\"The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"The earlier mold-clamp proximity-switch stop-work hold was cleared after three successful dry cycles, and no current hold applied at 06:52.\"},{\"speaker\":\"Materials clerk\",\"text\":\"The reserve-resin delivery was delayed and documented as a materials issue. The delay could disrupt continued production, but it did not block line release at 06:52.\"},{\"speaker\":\"Production scheduler\",\"text\":\"These entries are the latest available status for changeover L4; earlier reports were superseded where applicable.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "The calibrated 06:52 inventory record for lot R-17 lists 19.4 kilograms staged, while the first lot requires 22.0 kilograms.", "right": "The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-002", "id": "fast-41-diverse-294-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Materials clerk", "text": "The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Inventory auditor", "text": "The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Quality inspector", "text": "At 06:52, quality approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "The earlier mold-clamp proximity-switch stop-work hold was cleared after three successful dry cycles, and no current hold applied at 06:52."}, {"speaker": "Materials clerk", "text": "The reserve-resin delivery was delayed and documented as a materials issue. The delay could disrupt continued production, but it did not block line release at 06:52."}, {"speaker": "Production scheduler", "text": "These entries are the latest available status for changeover L4; earlier reports were superseded where applicable."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object remains unchanged and neither context alters governing rules. Question bindings remain fixed to changeover L4 at 06:52 and its specified release gates. Evidence spans are complete factual sentences: “The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B.” and “The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms.” The counterfactual is coherent because its 19.4-kilogram latest record supersedes earlier inventory figures and is internally below the stated 22.0-kilogram requirement. Neither context embeds an answer, code, rule table, rationale, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Materials clerk\",\"text\":\"The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Inventory auditor\",\"text\":\"The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"The earlier mold-clamp proximity-switch stop-work hold was cleared after three successful dry cycles, and no current hold applied at 06:52.\"},{\"speaker\":\"Materials clerk\",\"text\":\"The reserve-resin delivery was delayed and documented as a materials issue. The delay could disrupt continued production, but it did not block line release at 06:52.\"},{\"speaker\":\"Production scheduler\",\"text\":\"These entries are the latest available status for changeover L4; earlier reports were superseded where applicable.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "The calibrated 06:52 inventory record for lot R-17 lists 19.4 kilograms staged, while the first lot requires 22.0 kilograms.", "right": "The calibrated 06:52 inventory record for lot R-17 lists 24.6 kilograms staged, while the first lot requires 22.0 kilograms."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-002", "id": "fast-41-diverse-294-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Materials clerk", "text": "The 06:52 changeover log identifies resin lot R-17 as the material staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Inventory auditor", "text": "The calibrated 06:52 inventory record for lot R-17 lists 19.4 kilograms staged, while the first lot requires 22.0 kilograms."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Quality inspector", "text": "At 06:52, quality approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "The earlier mold-clamp proximity-switch stop-work hold was cleared after three successful dry cycles, and no current hold applied at 06:52."}, {"speaker": "Materials clerk", "text": "The reserve-resin delivery was delayed and documented as a materials issue. The delay could disrupt continued production, but it did not block line release at 06:52."}, {"speaker": "Production scheduler", "text": "These entries are the latest available status for changeover L4; earlier reports were superseded where applicable."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, and both contexts retain its governing policy. Both contexts bind to changeover L4 at 06:52. The focus evidence contains two complete factual sentences. The counterfactual changes only the required quantity, making staged resin insufficient without contradiction. Neither context includes a gold answer, code, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Production scheduler\",\"text\":\"As of 06:52, the resin quantity required for the first lot of changeover L4 is 10.8 kilograms.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality has approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, no current stop-work hold applies to changeover L4; the earlier equipment concern was cleared after successful checks.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The reserve-resin delay is being monitored as a materials concern. As of 06:52, it does not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the resin quantity required for the first lot of changeover L4 is 10.8 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the resin quantity required for the first lot of changeover L4 is 12.1 kilograms.", "right": "As of 06:52, the resin quantity required for the first lot of changeover L4 is 10.8 kilograms."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-003", "id": "fast-41-diverse-294-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Production scheduler", "text": "As of 06:52, the resin quantity required for the first lot of changeover L4 is 10.8 kilograms."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, no current stop-work hold applies to changeover L4; the earlier equipment concern was cleared after successful checks."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production."}, {"speaker": "Line supervisor", "text": "The reserve-resin delay is being monitored as a materials concern. As of 06:52, it does not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The original questions object is preserved verbatim, and both contexts retain its governing policy. Both contexts bind to changeover L4 at 06:52. The focus evidence contains two complete factual sentences. The counterfactual changes only the required quantity, making staged resin insufficient without contradiction. Neither context includes a gold answer, code, rationale, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Production scheduler\",\"text\":\"As of 06:52, the resin quantity required for the first lot of changeover L4 is 10.8 kilograms.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality has approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, no current stop-work hold applies to changeover L4; the earlier equipment concern was cleared after successful checks.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The reserve-resin delay is being monitored as a materials concern. As of 06:52, it does not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the resin quantity required for the first lot of changeover L4 is 10.8 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the resin quantity required for the first lot of changeover L4 is 12.1 kilograms.", "right": "As of 06:52, the resin quantity required for the first lot of changeover L4 is 10.8 kilograms."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-003", "id": "fast-41-diverse-294-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "As of 06:52, 11.4 kilograms of resin is staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Production scheduler", "text": "As of 06:52, the resin quantity required for the first lot of changeover L4 is 12.1 kilograms."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, no current stop-work hold applies to changeover L4; the earlier equipment concern was cleared after successful checks."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production."}, {"speaker": "Line supervisor", "text": "The reserve-resin delay is being monitored as a materials concern. As of 06:52, it does not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and L4/Jar-A-to-Jar-B/06:52 bindings, remain coherent, and contain no answer leakage; required evidence: \"As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters.\" and \"As of 06:52, the first lot for changeover L4 required 80 liters of resin.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, changeover L4 from Jar-A to Jar-B was under release review.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the first lot for changeover L4 required 80 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present for changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality had approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, no current stop-work hold applied to changeover L4; the earlier equipment issue had been cleared.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The delayed reserve-resin delivery did not block line release, and materials support was assigned to monitor it.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters."}, {"path": ["2", "text"], "text": "As of 06:52, the first lot for changeover L4 required 80 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters.", "negative_left": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters.", "negative_right": "As of 06:52, the first lot for changeover L4 required 90 liters of resin.", "right": "As of 06:52, the first lot for changeover L4 required 80 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-004", "id": "fast-41-diverse-294-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, changeover L4 from Jar-A to Jar-B was under release review."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the first lot for changeover L4 required 80 liters of resin."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present for changeover L4."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality had approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, no current stop-work hold applied to changeover L4; the earlier equipment issue had been cleared."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production."}, {"speaker": "Line supervisor", "text": "The delayed reserve-resin delivery did not block line release, and materials support was assigned to monitor it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and L4/Jar-A-to-Jar-B/06:52 bindings, remain coherent, and contain no answer leakage; required evidence: \"As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters.\" and \"As of 06:52, the first lot for changeover L4 required 80 liters of resin.\"", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, changeover L4 from Jar-A to Jar-B was under release review.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the first lot for changeover L4 required 80 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present for changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality had approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, no current stop-work hold applied to changeover L4; the earlier equipment issue had been cleared.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The delayed reserve-resin delivery did not block line release, and materials support was assigned to monitor it.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters."}, {"path": ["2", "text"], "text": "As of 06:52, the first lot for changeover L4 required 80 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters.", "negative_left": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters.", "negative_right": "As of 06:52, the first lot for changeover L4 required 90 liters of resin.", "right": "As of 06:52, the first lot for changeover L4 required 80 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-004", "id": "fast-41-diverse-294-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, changeover L4 from Jar-A to Jar-B was under release review."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 84 liters."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the first lot for changeover L4 required 90 liters of resin."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present for changeover L4."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality had approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, no current stop-work hold applied to changeover L4; the earlier equipment issue had been cleared."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production."}, {"speaker": "Line supervisor", "text": "The delayed reserve-resin delivery did not block line release, and materials support was assigned to monitor it."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, both contexts retain the L4, 06:52, and Jar-A-to-Jar-B bindings, the evidence consists of exactly these complete factual sentences—\"As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B.\" and \"As of 06:52, the first lot for changeover L4 requires 24 liters of resin.\"—the counterfactual coherently changes only staged resin from 24 to 18 liters without contradictory duplicates, and neither context embeds a gold answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Inventory clerk\",\"text\":\"As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Production planner\",\"text\":\"As of 06:52, the first lot for changeover L4 requires 24 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present for changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality has approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, no current stop-work hold applies to changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production, but it does not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the first lot for changeover L4 requires 24 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, 18 liters of resin are staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the first lot for changeover L4 requires 24 liters of resin.", "right": "As of 06:52, the first lot for changeover L4 requires 24 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-005", "id": "fast-41-diverse-294-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Inventory clerk", "text": "As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Production planner", "text": "As of 06:52, the first lot for changeover L4 requires 24 liters of resin."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present for changeover L4."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production, but it does not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the policy, both contexts retain the L4, 06:52, and Jar-A-to-Jar-B bindings, the evidence consists of exactly these complete factual sentences—\"As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B.\" and \"As of 06:52, the first lot for changeover L4 requires 24 liters of resin.\"—the counterfactual coherently changes only staged resin from 24 to 18 liters without contradictory duplicates, and neither context embeds a gold answer or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Inventory clerk\",\"text\":\"As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Production planner\",\"text\":\"As of 06:52, the first lot for changeover L4 requires 24 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present for changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality has approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, no current stop-work hold applies to changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production, but it does not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the first lot for changeover L4 requires 24 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, 24 liters of resin are staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, 18 liters of resin are staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the first lot for changeover L4 requires 24 liters of resin.", "right": "As of 06:52, the first lot for changeover L4 requires 24 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-005", "id": "fast-41-diverse-294-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Inventory clerk", "text": "As of 06:52, 18 liters of resin are staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Production planner", "text": "As of 06:52, the first lot for changeover L4 requires 24 liters of resin."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present for changeover L4."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production, but it does not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and neither context changes its gates or decision criteria. Bindings are preserved for changeover L4, Jar-A to Jar-B, and 06:52. Evidence consists of the complete factual sentences “As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B.” and “As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin.” Counterfactual coherence is preserved because only the requirement changes from 12 to 16 liters while all other observations remain consistent. No gold answer, label code, rule table, rationale, or output instruction is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the changeover review covered L4 from Jar-A to Jar-B. As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B. As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present for changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production, but it did not block line release.\"},{\"speaker\":\"Review note\",\"text\":\"If the first-lot worksheet instead required more resin than the staged amount, the inventory and all other recorded conditions would remain unchanged; the same reserve delivery would still be monitored and would not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["0", "text"], "text": "As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 16 liters of resin.", "right": "As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-006", "id": "fast-41-diverse-294-006-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, the changeover review covered L4 from Jar-A to Jar-B. As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B. As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present for changeover L4."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied."}, {"speaker": "Materials coordinator", "text": "At 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production, but it did not block line release."}, {"speaker": "Review note", "text": "If the first-lot worksheet instead required more resin than the staged amount, the inventory and all other recorded conditions would remain unchanged; the same reserve delivery would still be monitored and would not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved because the original questions object is verbatim and neither context changes its gates or decision criteria. Bindings are preserved for changeover L4, Jar-A to Jar-B, and 06:52. Evidence consists of the complete factual sentences “As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B.” and “As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin.” Counterfactual coherence is preserved because only the requirement changes from 12 to 16 liters while all other observations remain consistent. No gold answer, label code, rule table, rationale, or output instruction is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the changeover review covered L4 from Jar-A to Jar-B. As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B. As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present for changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production, but it did not block line release.\"},{\"speaker\":\"Review note\",\"text\":\"If the first-lot worksheet instead required more resin than the staged amount, the inventory and all other recorded conditions would remain unchanged; the same reserve delivery would still be monitored and would not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["0", "text"], "text": "As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 16 liters of resin.", "right": "As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 12 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-006", "id": "fast-41-diverse-294-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, the changeover review covered L4 from Jar-A to Jar-B. As of 06:52, the calibrated inventory record lists 14 liters of resin staged for changeover L4 from Jar-A to Jar-B. As of 06:52, the signed first-lot worksheet for changeover L4 lists a requirement of 16 liters of resin."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present for changeover L4."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied."}, {"speaker": "Materials coordinator", "text": "At 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production, but it did not block line release."}, {"speaker": "Review note", "text": "If the first-lot worksheet instead required more resin than the staged amount, the inventory and all other recorded conditions would remain unchanged; the same reserve delivery would still be monitored and would not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and all L4, Jar-A-to-Jar-B, and 06:52 bindings without embedding an answer or classifier instruction; the base has 38 liters against a 35-liter requirement, while the counterfactual coherently changes only the requirement to 42 liters. The evidence consists of two complete factual sentences: “As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters.” and “As of 06:52, the first lot for changeover L4 requires 35 liters of resin.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, changeover L4 from Jar-A to Jar-B was under review. As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters. As of 06:52, the first lot for changeover L4 requires 35 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed for changeover L4, cleaning record C-441 was signed, and four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the reserve-resin delivery for changeover L4 was documented as delayed and could disrupt continued production, but it did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters."}, {"path": ["0", "text"], "text": "As of 06:52, the first lot for changeover L4 requires 35 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters.", "negative_left": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters.", "negative_right": "As of 06:52, the first lot for changeover L4 requires 42 liters of resin.", "right": "As of 06:52, the first lot for changeover L4 requires 35 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-007", "id": "fast-41-diverse-294-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, changeover L4 from Jar-A to Jar-B was under review. As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters. As of 06:52, the first lot for changeover L4 requires 35 liters of resin."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed for changeover L4, cleaning record C-441 was signed, and four required operators were present."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied."}, {"speaker": "Materials coordinator", "text": "At 06:52, the reserve-resin delivery for changeover L4 was documented as delayed and could disrupt continued production, but it did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged policy and all L4, Jar-A-to-Jar-B, and 06:52 bindings without embedding an answer or classifier instruction; the base has 38 liters against a 35-liter requirement, while the counterfactual coherently changes only the requirement to 42 liters. The evidence consists of two complete factual sentences: “As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters.” and “As of 06:52, the first lot for changeover L4 requires 35 liters of resin.”", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, changeover L4 from Jar-A to Jar-B was under review. As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters. As of 06:52, the first lot for changeover L4 requires 35 liters of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed for changeover L4, cleaning record C-441 was signed, and four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the reserve-resin delivery for changeover L4 was documented as delayed and could disrupt continued production, but it did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters."}, {"path": ["0", "text"], "text": "As of 06:52, the first lot for changeover L4 requires 35 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters.", "negative_left": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters.", "negative_right": "As of 06:52, the first lot for changeover L4 requires 42 liters of resin.", "right": "As of 06:52, the first lot for changeover L4 requires 35 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-007", "id": "fast-41-diverse-294-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, changeover L4 from Jar-A to Jar-B was under review. As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measured 38 liters. As of 06:52, the first lot for changeover L4 requires 42 liters of resin."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed for changeover L4, cleaning record C-441 was signed, and four required operators were present."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied."}, {"speaker": "Materials coordinator", "text": "At 06:52, the reserve-resin delivery for changeover L4 was documented as delayed and could disrupt continued production, but it did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy, and both contexts retain the L4, 06:52, first-lot bindings. The evidence has exactly two complete factual sentences: “As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin.” and “As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms.” The counterfactual coherently changes the staged amount to 43 kilograms without contradictory duplicates or altered unrelated facts. Neither context embeds a gold answer, code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, changeover L4 from Jar-A to Jar-B was under review for its first lot. As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin. As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production, but it did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin."}, {"path": ["0", "text"], "text": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin.", "negative_left": "As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin.", "negative_right": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 43 kilograms.", "right": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-010", "id": "fast-41-diverse-294-010-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, changeover L4 from Jar-A to Jar-B was under review for its first lot. As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin. As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied."}, {"speaker": "Materials coordinator", "text": "At 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production, but it did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions object preserves all governing policy, and both contexts retain the L4, 06:52, first-lot bindings. The evidence has exactly two complete factual sentences: “As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin.” and “As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms.” The counterfactual coherently changes the staged amount to 43 kilograms without contradictory duplicates or altered unrelated facts. Neither context embeds a gold answer, code, rationale, proposition ID, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, changeover L4 from Jar-A to Jar-B was under review for its first lot. As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin. As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production, but it did not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin."}, {"path": ["0", "text"], "text": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin.", "negative_left": "As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin.", "negative_right": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 43 kilograms.", "right": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 51 kilograms."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-010", "id": "fast-41-diverse-294-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, changeover L4 from Jar-A to Jar-B was under review for its first lot. As of 06:52, the first lot of changeover L4 requires 48 kilograms of resin. As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 43 kilograms."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release of changeover L4, and no current stop-work hold applied."}, {"speaker": "Materials coordinator", "text": "At 06:52, the delayed reserve-resin delivery for changeover L4 was documented as an issue and could disrupt continued production, but it did not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and thresholds. Both contexts retain the L4, Jar-A-to-Jar-B, and 06:52 bindings. The two evidence quotes are complete factual sentences. The counterfactual consistently changes the formulation requirement from 15 to 21 liters while retaining 18 liters of inventory. No context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the current L4 changeover record was reviewed for the Jar-A to Jar-B transfer. As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B. As of 06:52, the signed first-lot formulation for changeover L4 requires 15 liters of resin. In a counterfactual revision, the signed formulation would require 21 liters instead.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release, and no current stop-work hold applied to changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the reserve-resin delivery was delayed and the issue was documented. Materials staff assessed that the delay could disrupt continued production, but it did not block line release at the assessment time.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["0", "text"], "text": "As of 06:52, the signed first-lot formulation for changeover L4 requires 15 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the signed first-lot formulation for changeover L4 requires 21 liters of resin.", "right": "As of 06:52, the signed first-lot formulation for changeover L4 requires 15 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-011", "id": "fast-41-diverse-294-011-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, the current L4 changeover record was reviewed for the Jar-A to Jar-B transfer. As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B. As of 06:52, the signed first-lot formulation for changeover L4 requires 15 liters of resin. In a counterfactual revision, the signed formulation would require 21 liters instead."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release, and no current stop-work hold applied to changeover L4."}, {"speaker": "Materials coordinator", "text": "At 06:52, the reserve-resin delivery was delayed and the issue was documented. Materials staff assessed that the delay could disrupt continued production, but it did not block line release at the assessment time."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and thresholds. Both contexts retain the L4, Jar-A-to-Jar-B, and 06:52 bindings. The two evidence quotes are complete factual sentences. The counterfactual consistently changes the formulation requirement from 15 to 21 liters while retaining 18 liters of inventory. No context contains a gold label, answer code, rule table, proposition identifier, label rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the current L4 changeover record was reviewed for the Jar-A to Jar-B transfer. As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B. As of 06:52, the signed first-lot formulation for changeover L4 requires 15 liters of resin. In a counterfactual revision, the signed formulation would require 21 liters instead.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, quality had approved the setup sample for release, and no current stop-work hold applied to changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"At 06:52, the reserve-resin delivery was delayed and the issue was documented. Materials staff assessed that the delay could disrupt continued production, but it did not block line release at the assessment time.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["0", "text"], "text": "As of 06:52, the signed first-lot formulation for changeover L4 requires 15 liters of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the signed first-lot formulation for changeover L4 requires 21 liters of resin.", "right": "As of 06:52, the signed first-lot formulation for changeover L4 requires 15 liters of resin."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-011", "id": "fast-41-diverse-294-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:52, the current L4 changeover record was reviewed for the Jar-A to Jar-B transfer. As of 06:52, the calibrated inventory log records 18 liters of resin staged for changeover L4 from Jar-A to Jar-B. As of 06:52, the signed first-lot formulation for changeover L4 requires 21 liters of resin. In a counterfactual revision, the signed formulation would require 21 liters instead."}, {"speaker": "Line supervisor", "text": "At 06:52, the Jar-B mold was installed, cleaning record C-441 was signed, and four required operators were present."}, {"speaker": "Quality inspector", "text": "At 06:52, quality had approved the setup sample for release, and no current stop-work hold applied to changeover L4."}, {"speaker": "Materials coordinator", "text": "At 06:52, the reserve-resin delivery was delayed and the issue was documented. Materials staff assessed that the delay could disrupt continued production, but it did not block line release at the assessment time."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the L4, 06:52, and first-lot bindings. The evidence spans are complete factual sentences: \"As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B.\" and \"As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4.\" The counterfactual consistently changes the required quantity to 16.8 liters, which conflicts with neither context assertion and leaves inventory as 14.6 liters. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Inventory log\",\"text\":\"As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Production plan\",\"text\":\"As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality has approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, no current stop-work hold applies to changeover L4 after the earlier equipment concern was cleared.\"},{\"speaker\":\"Shift log\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production. Materials support is monitoring it, but no release-blocking flag is recorded against L4.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the signed production plan specifies 16.8 liters as the resin quantity required for the first lot of changeover L4.", "right": "As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-012", "id": "fast-41-diverse-294-012-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Inventory log", "text": "As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Production plan", "text": "As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, no current stop-work hold applies to changeover L4 after the earlier equipment concern was cleared."}, {"speaker": "Shift log", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production. Materials support is monitoring it, but no release-blocking flag is recorded against L4."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the L4, 06:52, and first-lot bindings. The evidence spans are complete factual sentences: \"As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B.\" and \"As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4.\" The counterfactual consistently changes the required quantity to 16.8 liters, which conflicts with neither context assertion and leaves inventory as 14.6 liters. Neither context embeds an answer, label rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Inventory log\",\"text\":\"As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Production plan\",\"text\":\"As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality has approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, no current stop-work hold applies to changeover L4 after the earlier equipment concern was cleared.\"},{\"speaker\":\"Shift log\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production. Materials support is monitoring it, but no release-blocking flag is recorded against L4.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the signed production plan specifies 16.8 liters as the resin quantity required for the first lot of changeover L4.", "right": "As of 06:52, the signed production plan specifies 11.2 liters as the resin quantity required for the first lot of changeover L4."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-012", "id": "fast-41-diverse-294-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Inventory log", "text": "As of 06:52, the calibrated inventory log records 14.6 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Production plan", "text": "As of 06:52, the signed production plan specifies 16.8 liters as the resin quantity required for the first lot of changeover L4."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold is installed for changeover L4, cleaning record C-441 is signed, and four required operators are present."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, no current stop-work hold applies to changeover L4 after the earlier equipment concern was cleared."}, {"speaker": "Shift log", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production. Materials support is monitoring it, but no release-blocking flag is recorded against L4."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and 06:52/L4/Jar-A-to-Jar-B bindings; the complete factual evidence is “As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B.” and “As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B.”; the counterfactual coherently changes staged material to 19 liters without contradictory duplicates, and neither context embeds an answer, label, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold is installed for changeover L4; cleaning record C-441 is signed, and four required operators are present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality has approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, the earlier equipment fault is corrected and no current stop-work hold applies to changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production, but it does not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the scale records 19 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "right": "As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-013", "id": "fast-41-diverse-294-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Materials coordinator", "text": "As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold is installed for changeover L4; cleaning record C-441 is signed, and four required operators are present."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, the earlier equipment fault is corrected and no current stop-work hold applies to changeover L4."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production, but it does not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy and 06:52/L4/Jar-A-to-Jar-B bindings; the complete factual evidence is “As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B.” and “As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B.”; the counterfactual coherently changes staged material to 19 liters without contradictory duplicates, and neither context embeds an answer, label, rule table, rationale, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B.\"},{\"speaker\":\"Line supervisor\",\"text\":\"As of 06:52, the Jar-B mold is installed for changeover L4; cleaning record C-441 is signed, and four required operators are present.\"},{\"speaker\":\"Quality inspector\",\"text\":\"As of 06:52, quality has approved the setup sample for release of changeover L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"As of 06:52, the earlier equipment fault is corrected and no current stop-work hold applies to changeover L4.\"},{\"speaker\":\"Materials coordinator\",\"text\":\"As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production, but it does not block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B."}, {"path": ["1", "text"], "text": "As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B.", "negative_left": "As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B.", "negative_right": "As of 06:52, the scale records 19 liters of resin staged for changeover L4 from Jar-A to Jar-B.", "right": "As of 06:52, the scale records 27 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, "verifier_independent_model": false}, "family": "fast-41-diverse-294-013", "id": "fast-41-diverse-294-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Materials coordinator", "text": "As of 06:52, the calibrated batch sheet lists 24 liters as the first-lot resin requirement for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the scale records 19 liters of resin staged for changeover L4 from Jar-A to Jar-B."}, {"speaker": "Line supervisor", "text": "As of 06:52, the Jar-B mold is installed for changeover L4; cleaning record C-441 is signed, and four required operators are present."}, {"speaker": "Quality inspector", "text": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"speaker": "Maintenance technician", "text": "As of 06:52, the earlier equipment fault is corrected and no current stop-work hold applies to changeover L4."}, {"speaker": "Materials coordinator", "text": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue and could disrupt continued production, but it does not block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the same stop, destination, time, and request bindings. The two focus spans are complete factual sentences: “The package assigned to stop 42 was the item examined in the 14:12 inspection.” and “That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category.” The counterfactual coherently changes the inspection condition to actual-or-suspected damage without contradictory duplicates. Neither context states a gold answer, label, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Case note for stop 42: the intended destination was verified as 18 Vale Street. At 14:12, construction closure of the designated loading bay at that address caused the access failure, and the closure was temporary. The recipient asked for another delivery attempt after 10 tomorrow and did not request collection from the depot. The policy record does not forbid another delivery attempt for this stop. Apply the route policy in the options and treat synonymous wording as equivalent evidence. The inspection record identifies the relevant package and condition field below.\",\"evidence\":[\"The package assigned to stop 42 was the item examined in the 14:12 inspection.\",\"That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category.\",\"The destination for stop 42 is recorded as verified: 18 Vale Street.\",\"The designated loading bay at 18 Vale Street was temporarily closed for construction at 14:12, causing the access failure.\",\"The recipient requested another delivery attempt after 10 tomorrow.\",\"The recipient did not request collection from the depot.\",\"Policy records do not forbid another delivery attempt for stop 42.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The package assigned to stop 42 was the item examined in the 14:12 inspection."}, {"path": ["evidence", "1"], "text": "That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The package assigned to stop 42 was the item examined in the 14:12 inspection.", "negative_left": "The package assigned to stop 42 was the item examined in the 14:12 inspection.", "negative_right": "That inspection's binary condition field is marked actual-or-suspected damage, with sound as the alternative category.", "right": "That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category."}, "verifier_independent_model": false}, "family": "fast-41-diverse-295-001", "id": "fast-41-diverse-295-001-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Case note for stop 42: the intended destination was verified as 18 Vale Street. At 14:12, construction closure of the designated loading bay at that address caused the access failure, and the closure was temporary. The recipient asked for another delivery attempt after 10 tomorrow and did not request collection from the depot. The policy record does not forbid another delivery attempt for this stop. Apply the route policy in the options and treat synonymous wording as equivalent evidence. The inspection record identifies the relevant package and condition field below.", "evidence": ["The package assigned to stop 42 was the item examined in the 14:12 inspection.", "That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category.", "The destination for stop 42 is recorded as verified: 18 Vale Street.", "The designated loading bay at 18 Vale Street was temporarily closed for construction at 14:12, causing the access failure.", "The recipient requested another delivery attempt after 10 tomorrow.", "The recipient did not request collection from the depot.", "Policy records do not forbid another delivery attempt for stop 42."]}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts retain the same stop, destination, time, and request bindings. The two focus spans are complete factual sentences: “The package assigned to stop 42 was the item examined in the 14:12 inspection.” and “That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category.” The counterfactual coherently changes the inspection condition to actual-or-suspected damage without contradictory duplicates. Neither context states a gold answer, label, rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Case note for stop 42: the intended destination was verified as 18 Vale Street. At 14:12, construction closure of the designated loading bay at that address caused the access failure, and the closure was temporary. The recipient asked for another delivery attempt after 10 tomorrow and did not request collection from the depot. The policy record does not forbid another delivery attempt for this stop. Apply the route policy in the options and treat synonymous wording as equivalent evidence. The inspection record identifies the relevant package and condition field below.\",\"evidence\":[\"The package assigned to stop 42 was the item examined in the 14:12 inspection.\",\"That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category.\",\"The destination for stop 42 is recorded as verified: 18 Vale Street.\",\"The designated loading bay at 18 Vale Street was temporarily closed for construction at 14:12, causing the access failure.\",\"The recipient requested another delivery attempt after 10 tomorrow.\",\"The recipient did not request collection from the depot.\",\"Policy records do not forbid another delivery attempt for stop 42.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The package assigned to stop 42 was the item examined in the 14:12 inspection."}, {"path": ["evidence", "1"], "text": "That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The package assigned to stop 42 was the item examined in the 14:12 inspection.", "negative_left": "The package assigned to stop 42 was the item examined in the 14:12 inspection.", "negative_right": "That inspection's binary condition field is marked actual-or-suspected damage, with sound as the alternative category.", "right": "That inspection's binary condition field is marked sound, with actual-or-suspected damage as the alternative category."}, "verifier_independent_model": false}, "family": "fast-41-diverse-295-001", "id": "fast-41-diverse-295-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Case note for stop 42: the intended destination was verified as 18 Vale Street. At 14:12, construction closure of the designated loading bay at that address caused the access failure, and the closure was temporary. The recipient asked for another delivery attempt after 10 tomorrow and did not request collection from the depot. The policy record does not forbid another delivery attempt for this stop. Apply the route policy in the options and treat synonymous wording as equivalent evidence. The inspection record identifies the relevant package and condition field below.", "evidence": ["The package assigned to stop 42 was the item examined in the 14:12 inspection.", "That inspection's binary condition field is marked actual-or-suspected damage, with sound as the alternative category.", "The destination for stop 42 is recorded as verified: 18 Vale Street.", "The designated loading bay at 18 Vale Street was temporarily closed for construction at 14:12, causing the access failure.", "The recipient requested another delivery attempt after 10 tomorrow.", "The recipient did not request collection from the depot.", "Policy records do not forbid another delivery attempt for stop 42."]}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, policy wording, stop 42, destination, timing, and request bindings without adding policy defaults or output instructions. The required evidence quotes are preserved exactly: \"A binary condition inspection occurred at 14:12 for the package assigned to stop 42.\" and \"The inspection record labels its condition result as sound.\" The counterfactual changes only the inspection label, and its remaining assertions are coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"A binary condition inspection occurred at 14:12 for the package assigned to stop 42.\",\"The inspection record labels its condition result as sound.\",\"Dispatch verified that stop 42’s intended destination is 18 Vale Street. The designated loading bay there was closed by construction at 14:12, and the closure was temporary, causing the access failure. The recipient asked for another delivery attempt after 10 tomorrow. The recipient did not request collection from the depot, and no policy rule forbade another delivery attempt. The address details were consistent, with no competing destination recorded. Depot operations logged the inspection event and the access exception separately, while customer service recorded the recipient’s request.\",\"The case note records no depot-collection request and no prohibition on another delivery attempt.\",\"If the inspection record had instead carried the other condition label, all destination, access, timing, recipient-request, and policy facts would remain unchanged.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "A binary condition inspection occurred at 14:12 for the package assigned to stop 42."}, {"path": ["evidence", "1"], "text": "The inspection record labels its condition result as sound."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "A binary condition inspection occurred at 14:12 for the package assigned to stop 42.", "negative_left": "A binary condition inspection occurred at 14:12 for the package assigned to stop 42.", "negative_right": "The inspection record labels its condition result as actual-or-suspected damage.", "right": "The inspection record labels its condition result as sound."}, "verifier_independent_model": false}, "family": "fast-41-diverse-295-003", "id": "fast-41-diverse-295-003-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["A binary condition inspection occurred at 14:12 for the package assigned to stop 42.", "The inspection record labels its condition result as sound.", "Dispatch verified that stop 42’s intended destination is 18 Vale Street. The designated loading bay there was closed by construction at 14:12, and the closure was temporary, causing the access failure. The recipient asked for another delivery attempt after 10 tomorrow. The recipient did not request collection from the depot, and no policy rule forbade another delivery attempt. The address details were consistent, with no competing destination recorded. Depot operations logged the inspection event and the access exception separately, while customer service recorded the recipient’s request.", "The case note records no depot-collection request and no prohibition on another delivery attempt.", "If the inspection record had instead carried the other condition label, all destination, access, timing, recipient-request, and policy facts would remain unchanged."]}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions, policy wording, stop 42, destination, timing, and request bindings without adding policy defaults or output instructions. The required evidence quotes are preserved exactly: \"A binary condition inspection occurred at 14:12 for the package assigned to stop 42.\" and \"The inspection record labels its condition result as sound.\" The counterfactual changes only the inspection label, and its remaining assertions are coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"A binary condition inspection occurred at 14:12 for the package assigned to stop 42.\",\"The inspection record labels its condition result as sound.\",\"Dispatch verified that stop 42’s intended destination is 18 Vale Street. The designated loading bay there was closed by construction at 14:12, and the closure was temporary, causing the access failure. The recipient asked for another delivery attempt after 10 tomorrow. The recipient did not request collection from the depot, and no policy rule forbade another delivery attempt. The address details were consistent, with no competing destination recorded. Depot operations logged the inspection event and the access exception separately, while customer service recorded the recipient’s request.\",\"The case note records no depot-collection request and no prohibition on another delivery attempt.\",\"If the inspection record had instead carried the other condition label, all destination, access, timing, recipient-request, and policy facts would remain unchanged.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "A binary condition inspection occurred at 14:12 for the package assigned to stop 42."}, {"path": ["evidence", "1"], "text": "The inspection record labels its condition result as sound."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "A binary condition inspection occurred at 14:12 for the package assigned to stop 42.", "negative_left": "A binary condition inspection occurred at 14:12 for the package assigned to stop 42.", "negative_right": "The inspection record labels its condition result as actual-or-suspected damage.", "right": "The inspection record labels its condition result as sound."}, "verifier_independent_model": false}, "family": "fast-41-diverse-295-003", "id": "fast-41-diverse-295-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["A binary condition inspection occurred at 14:12 for the package assigned to stop 42.", "The inspection record labels its condition result as actual-or-suspected damage.", "Dispatch verified that stop 42’s intended destination is 18 Vale Street. The designated loading bay there was closed by construction at 14:12, and the closure was temporary, causing the access failure. The recipient asked for another delivery attempt after 10 tomorrow. The recipient did not request collection from the depot, and no policy rule forbade another delivery attempt. The address details were consistent, with no competing destination recorded. Depot operations logged the inspection event and the access exception separately, while customer service recorded the recipient’s request.", "The case note records no depot-collection request and no prohibition on another delivery attempt.", "If the inspection record had instead carried the other condition label, all destination, access, timing, recipient-request, and policy facts would remain unchanged."]}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions, while both contexts retain the route-policy instruction without adding exceptions or defaults. Both contexts bind the request to stop 42 and preserve the destination, timing, and route scope. Base evidence retains exactly: \"The 14:12 inspection record for the package assigned to stop 42 carries condition code S.\" and \"In that inspection record, code S means sound and code D means actual-or-suspected damage.\" Counterfactual evidence retains exactly: \"The 14:12 inspection record for the package assigned to stop 42 carries condition code D.\" and \"In that inspection record, code S means sound and code D means actual-or-suspected damage.\" The counterfactual changes only the condition code and remains coherent with the unchanged facts without contradictory duplicate assertions. Neither context embeds an answer, code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Case note: Stop 42 was assigned to 18 Vale Street, and the destination record was verified without conflicting location details. At 14:12, the designated loading bay there was closed for construction; the closure was temporary and caused the access failure. The recipient asked for another delivery attempt after 10 tomorrow. The service log records no request to collect the parcel from the depot, and the applicable policy contains no prohibition on another attempt. Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"The 14:12 inspection record for the package assigned to stop 42 carries condition code S.\",\"In that inspection record, code S means sound and code D means actual-or-suspected damage.\"],\"request\":\"Determine the single eligible exception route for stop 42.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The 14:12 inspection record for the package assigned to stop 42 carries condition code S."}, {"path": ["evidence", "1"], "text": "In that inspection record, code S means sound and code D means actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The 14:12 inspection record for the package assigned to stop 42 carries condition code S.", "negative_left": "The 14:12 inspection record for the package assigned to stop 42 carries condition code D.", "negative_right": "In that inspection record, code S means sound and code D means actual-or-suspected damage.", "right": "In that inspection record, code S means sound and code D means actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "fast-41-diverse-295-006", "id": "fast-41-diverse-295-006-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Case note: Stop 42 was assigned to 18 Vale Street, and the destination record was verified without conflicting location details. At 14:12, the designated loading bay there was closed for construction; the closure was temporary and caused the access failure. The recipient asked for another delivery attempt after 10 tomorrow. The service log records no request to collect the parcel from the depot, and the applicable policy contains no prohibition on another attempt. Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["The 14:12 inspection record for the package assigned to stop 42 carries condition code S.", "In that inspection record, code S means sound and code D means actual-or-suspected damage."], "request": "Determine the single eligible exception route for stop 42."}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve all governing criteria and instructions, while both contexts retain the route-policy instruction without adding exceptions or defaults. Both contexts bind the request to stop 42 and preserve the destination, timing, and route scope. Base evidence retains exactly: \"The 14:12 inspection record for the package assigned to stop 42 carries condition code S.\" and \"In that inspection record, code S means sound and code D means actual-or-suspected damage.\" Counterfactual evidence retains exactly: \"The 14:12 inspection record for the package assigned to stop 42 carries condition code D.\" and \"In that inspection record, code S means sound and code D means actual-or-suspected damage.\" The counterfactual changes only the condition code and remains coherent with the unchanged facts without contradictory duplicate assertions. Neither context embeds an answer, code, rationale, rule table, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Case note: Stop 42 was assigned to 18 Vale Street, and the destination record was verified without conflicting location details. At 14:12, the designated loading bay there was closed for construction; the closure was temporary and caused the access failure. The recipient asked for another delivery attempt after 10 tomorrow. The service log records no request to collect the parcel from the depot, and the applicable policy contains no prohibition on another attempt. Apply the route policy in the options and treat synonymous wording as equivalent evidence.\",\"evidence\":[\"The 14:12 inspection record for the package assigned to stop 42 carries condition code S.\",\"In that inspection record, code S means sound and code D means actual-or-suspected damage.\"],\"request\":\"Determine the single eligible exception route for stop 42.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The 14:12 inspection record for the package assigned to stop 42 carries condition code S."}, {"path": ["evidence", "1"], "text": "In that inspection record, code S means sound and code D means actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "The 14:12 inspection record for the package assigned to stop 42 carries condition code S.", "negative_left": "The 14:12 inspection record for the package assigned to stop 42 carries condition code D.", "negative_right": "In that inspection record, code S means sound and code D means actual-or-suspected damage.", "right": "In that inspection record, code S means sound and code D means actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "fast-41-diverse-295-006", "id": "fast-41-diverse-295-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Case note: Stop 42 was assigned to 18 Vale Street, and the destination record was verified without conflicting location details. At 14:12, the designated loading bay there was closed for construction; the closure was temporary and caused the access failure. The recipient asked for another delivery attempt after 10 tomorrow. The service log records no request to collect the parcel from the depot, and the applicable policy contains no prohibition on another attempt. Apply the route policy in the options and treat synonymous wording as equivalent evidence.", "evidence": ["The 14:12 inspection record for the package assigned to stop 42 carries condition code D.", "In that inspection record, code S means sound and code D means actual-or-suspected damage."], "request": "Determine the single eligible exception route for stop 42."}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, scope, and criteria. Stop 42, 18 Vale Street, the delivery attempt, and the package inspection remain consistently bound. The two focus spans are complete factual sentences: “At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection.” and “The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage.” The counterfactual coherently changes the package condition to damage without creating contradictory duplicate assertions. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Case note: Stop 42 was scheduled for 18 Vale Street, and the destination details were confirmed without conflict. At 14:12, construction barriers closed the designated loading bay, preventing access; the closure was temporary. The recipient asked the team to try delivery again after 10 tomorrow. The recipient did not request collection from the depot, and the applicable policy did not forbid another delivery attempt. Apply the route policy in the options and treat synonymous wording as equivalent evidence. At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection. The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage.\",\"request\":\"Record the single eligible exception route for the current review.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection."}, {"path": ["context"], "text": "The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection.", "negative_left": "At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection.", "negative_right": "The 14:12 package inspection recorded the inspected parcel's binary condition result as actual-or-suspected damage rather than sound.", "right": "The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "fast-41-diverse-295-009", "id": "fast-41-diverse-295-009-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Case note: Stop 42 was scheduled for 18 Vale Street, and the destination details were confirmed without conflict. At 14:12, construction barriers closed the designated loading bay, preventing access; the closure was temporary. The recipient asked the team to try delivery again after 10 tomorrow. The recipient did not request collection from the depot, and the applicable policy did not forbid another delivery attempt. Apply the route policy in the options and treat synonymous wording as equivalent evidence. At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection. The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage.", "request": "Record the single eligible exception route for the current review."}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy, scope, and criteria. Stop 42, 18 Vale Street, the delivery attempt, and the package inspection remain consistently bound. The two focus spans are complete factual sentences: “At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection.” and “The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage.” The counterfactual coherently changes the package condition to damage without creating contradictory duplicate assertions. Neither context embeds an answer, label rationale, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, with times, stops, and locations functioning as qualifiers rather than unrelated bundled requirements. The focus atom concerns a package-condition inspection result, not a policy classification. The base and counter assignments can be realized with only that binary inspection result changing while the destination, access circumstances, recipient request, and depot-hold facts remain fixed. Policy completeness is satisfied: the unchanged questions preserve all route criteria and choice instructions, while policy_evidence preserves the substantive state-originating instruction to treat synonymous wording as equivalent evidence.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction supplies every redelivery condition: verified destination, sound package, a temporary access failure, and a request for another attempt. Verified destination excludes address clarification; the sound-versus-damage result excludes damage review; and both depot-hold grounds are explicitly refuted.", "rule_index": 0, "sound": true}, {"reason": "Because a1 explicitly describes a binary condition result between sound and actual-or-suspected damage, refuting a1 entails the inspection indicated actual or suspected damage. That is sufficient for damage review. The verified destination excludes address clarification, the damage result defeats redelivery's sound-package requirement, and both depot-hold grounds are refuted.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The 14:12 package inspection's binary condition result for the package assigned to stop 42 was sound rather than actual-or-suspected damage."}, {"id": "a2", "statement": "The intended destination for stop 42 was verified as 18 Vale Street."}, {"id": "a3", "statement": "At 14:12, construction closure of the designated loading bay at 18 Vale Street caused the access failure for stop 42."}, {"id": "a4", "statement": "The construction closure of the designated loading bay at 18 Vale Street was temporary."}, {"id": "a5", "statement": "The recipient for stop 42 requested another delivery attempt after 10 tomorrow."}, {"id": "a6", "statement": "The recipient for stop 42 requested collection from the depot."}, {"id": "a7", "statement": "Policy forbade another delivery attempt for stop 42."}], "base_state_json": "{\"context\":\"Case note: Stop 42 was scheduled for 18 Vale Street, and the destination details were confirmed without conflict. At 14:12, construction barriers closed the designated loading bay, preventing access; the closure was temporary. The recipient asked the team to try delivery again after 10 tomorrow. The recipient did not request collection from the depot, and the applicable policy did not forbid another delivery attempt. Apply the route policy in the options and treat synonymous wording as equivalent evidence. At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection. The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage.\",\"request\":\"Record the single eligible exception route for the current review.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["context"], "text": "At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection."}, {"path": ["context"], "text": "The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage."}], "policy_evidence": [{"path": ["context"], "text": "Apply the route policy in the options and treat synonymous wording as equivalent evidence."}], "rules": [{"justification": "The destination is verified, the package inspection classifies the package as sound, a temporary construction closure caused the access failure, and the recipient requested another attempt. Depot-hold grounds are explicitly absent.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}, {"justification": "Refutation of the inspection's binary sound result entails that the package inspection classified the package as actual-or-suspected damage. That makes damage review eligible and defeats the package-sound requirement for redelivery; address-clarification and depot-hold grounds are absent.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}]}]}, "verified_pair": {"left": "At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection.", "negative_left": "At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection.", "negative_right": "The 14:12 package inspection recorded the inspected parcel's binary condition result as actual-or-suspected damage rather than sound.", "right": "The 14:12 package inspection recorded the inspected parcel's binary condition result as sound rather than actual-or-suspected damage."}, "verifier_independent_model": false}, "family": "fast-41-diverse-295-009", "id": "fast-41-diverse-295-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Route for address clarification only when the destination cannot be verified or conflicting location details prevent identifying the intended stop.", "damage_review": "Route to damage review only when a scan, driver observation, or package inspection indicates actual or suspected damage.", "depot_hold": "Route to depot hold only when the recipient requests collection from the depot or policy forbids another delivery attempt.", "redelivery": "Route for redelivery when the address is verified, the package is sound, access failed for a temporary reason, and the recipient requests another attempt."}, "instructions": "Choose exactly one route. Interpret the recipient’s wording by meaning rather than requiring an exact policy phrase.", "type": "choice"}}, "state": {"context": "Case note: Stop 42 was scheduled for 18 Vale Street, and the destination details were confirmed without conflict. At 14:12, construction barriers closed the designated loading bay, preventing access; the closure was temporary. The recipient asked the team to try delivery again after 10 tomorrow. The recipient did not request collection from the depot, and the applicable policy did not forbid another delivery attempt. Apply the route policy in the options and treat synonymous wording as equivalent evidence. At 14:12, the package assigned to stop 42 was the parcel inspected in the package inspection. The 14:12 package inspection recorded the inspected parcel's binary condition result as actual-or-suspected damage rather than sound.", "request": "Record the single eligible exception route for the current review."}}, "method": "c2d", "provenance": {"source_id": "diverse-295", "source_is_synthetic": true, "source_sha256": "d3b7f0c276709eedf32e3e5064326a4200d727690de52f9c232222861ed7e850", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy, Stop 42/package BX-17 bindings, and coherent evidence; the counterfactual changes only the crush measurement from 2.6 cm to 1.4 cm without contradiction, and no gold answer or classifier instruction is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"Stop 42 is under review. The recorded damage log identifies package BX-17 as assigned to Stop 42.\"},{\"speaker\":\"Inspection record\",\"text\":\"The recorded damage log for package BX-17 lists a calibrated crush depth of 2.6 cm.\"},{\"speaker\":\"Inspection record\",\"text\":\"The package has no recorded leak, opening, or rattling. The address is verified and usable, and it is not conflicting.\"},{\"speaker\":\"Operations record\",\"text\":\"Fewer than two delivery attempts have failed. No recorded vehicle restriction prevents service, and required access details are present and usable.\"},{\"speaker\":\"Dispatch record\",\"text\":\"Another delivery attempt at Stop 42 is feasible.\"},{\"speaker\":\"Policy record\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The recorded damage log identifies package BX-17 as assigned to Stop 42."}, {"path": ["1", "text"], "text": "The recorded damage log for package BX-17 lists a calibrated crush depth of 2.6 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "The recorded damage log identifies package BX-17 as assigned to Stop 42.", "negative_left": "The recorded damage log identifies package BX-17 as assigned to Stop 42.", "negative_right": "The recorded damage log for package BX-17 lists a calibrated crush depth of 1.4 cm.", "right": "The recorded damage log for package BX-17 lists a calibrated crush depth of 2.6 cm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-296-001", "id": "fast-41-diverse-296-001-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Case note", "text": "Stop 42 is under review. The recorded damage log identifies package BX-17 as assigned to Stop 42."}, {"speaker": "Inspection record", "text": "The recorded damage log for package BX-17 lists a calibrated crush depth of 2.6 cm."}, {"speaker": "Inspection record", "text": "The package has no recorded leak, opening, or rattling. The address is verified and usable, and it is not conflicting."}, {"speaker": "Operations record", "text": "Fewer than two delivery attempts have failed. No recorded vehicle restriction prevents service, and required access details are present and usable."}, {"speaker": "Dispatch record", "text": "Another delivery attempt at Stop 42 is feasible."}, {"speaker": "Policy record", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy, Stop 42/package BX-17 bindings, and coherent evidence; the counterfactual changes only the crush measurement from 2.6 cm to 1.4 cm without contradiction, and no gold answer or classifier instruction is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Case note\",\"text\":\"Stop 42 is under review. The recorded damage log identifies package BX-17 as assigned to Stop 42.\"},{\"speaker\":\"Inspection record\",\"text\":\"The recorded damage log for package BX-17 lists a calibrated crush depth of 2.6 cm.\"},{\"speaker\":\"Inspection record\",\"text\":\"The package has no recorded leak, opening, or rattling. The address is verified and usable, and it is not conflicting.\"},{\"speaker\":\"Operations record\",\"text\":\"Fewer than two delivery attempts have failed. No recorded vehicle restriction prevents service, and required access details are present and usable.\"},{\"speaker\":\"Dispatch record\",\"text\":\"Another delivery attempt at Stop 42 is feasible.\"},{\"speaker\":\"Policy record\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The recorded damage log identifies package BX-17 as assigned to Stop 42."}, {"path": ["1", "text"], "text": "The recorded damage log for package BX-17 lists a calibrated crush depth of 2.6 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "The recorded damage log identifies package BX-17 as assigned to Stop 42.", "negative_left": "The recorded damage log identifies package BX-17 as assigned to Stop 42.", "negative_right": "The recorded damage log for package BX-17 lists a calibrated crush depth of 1.4 cm.", "right": "The recorded damage log for package BX-17 lists a calibrated crush depth of 2.6 cm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-296-001", "id": "fast-41-diverse-296-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Case note", "text": "Stop 42 is under review. The recorded damage log identifies package BX-17 as assigned to Stop 42."}, {"speaker": "Inspection record", "text": "The recorded damage log for package BX-17 lists a calibrated crush depth of 1.4 cm."}, {"speaker": "Inspection record", "text": "The package has no recorded leak, opening, or rattling. The address is verified and usable, and it is not conflicting."}, {"speaker": "Operations record", "text": "Fewer than two delivery attempts have failed. No recorded vehicle restriction prevents service, and required access details are present and usable."}, {"speaker": "Dispatch record", "text": "Another delivery attempt at Stop 42 is feasible."}, {"speaker": "Policy record", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts repeat it without added exceptions. Stop 42, container C-817, the date, and the 2.0 cm threshold remain consistently bound. The two evidence spans are complete factual sentences: \"On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817.\" and \"On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm.\" The counterfactual changes only the crush measurement to 1.63 cm and introduces no contradiction. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Inspection record\",\"text\":\"On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817.\"},{\"speaker\":\"Shipment record\",\"text\":\"On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm.\"},{\"speaker\":\"Condition log\",\"text\":\"The inspection recorded no leak, opening, or rattling for the package assigned to Stop 42.\"},{\"speaker\":\"Address record\",\"text\":\"The delivery address for Stop 42 was verified and usable, with no conflicting address entry.\"},{\"speaker\":\"Route record\",\"text\":\"Stop 42 had no recorded vehicle restriction, and the delivery history showed fewer than two failed attempts.\"},{\"speaker\":\"Access record\",\"text\":\"The access details required to perform Stop 42 were present and usable.\"},{\"speaker\":\"Dispatcher\",\"text\":\"Another delivery attempt at Stop 42 was feasible.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817."}, {"path": ["1", "text"], "text": "On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817.", "negative_left": "On 17 September 2026, a calibrated inspection recorded a crush depth of 1.63 cm for container C-817.", "negative_right": "On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm.", "right": "On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-296-005", "id": "fast-41-diverse-296-005-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Inspection record", "text": "On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817."}, {"speaker": "Shipment record", "text": "On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm."}, {"speaker": "Condition log", "text": "The inspection recorded no leak, opening, or rattling for the package assigned to Stop 42."}, {"speaker": "Address record", "text": "The delivery address for Stop 42 was verified and usable, with no conflicting address entry."}, {"speaker": "Route record", "text": "Stop 42 had no recorded vehicle restriction, and the delivery history showed fewer than two failed attempts."}, {"speaker": "Access record", "text": "The access details required to perform Stop 42 were present and usable."}, {"speaker": "Dispatcher", "text": "Another delivery attempt at Stop 42 was feasible."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy, and both contexts repeat it without added exceptions. Stop 42, container C-817, the date, and the 2.0 cm threshold remain consistently bound. The two evidence spans are complete factual sentences: \"On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817.\" and \"On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm.\" The counterfactual changes only the crush measurement to 1.63 cm and introduces no contradiction. Neither context contains a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Inspection record\",\"text\":\"On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817.\"},{\"speaker\":\"Shipment record\",\"text\":\"On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm.\"},{\"speaker\":\"Condition log\",\"text\":\"The inspection recorded no leak, opening, or rattling for the package assigned to Stop 42.\"},{\"speaker\":\"Address record\",\"text\":\"The delivery address for Stop 42 was verified and usable, with no conflicting address entry.\"},{\"speaker\":\"Route record\",\"text\":\"Stop 42 had no recorded vehicle restriction, and the delivery history showed fewer than two failed attempts.\"},{\"speaker\":\"Access record\",\"text\":\"The access details required to perform Stop 42 were present and usable.\"},{\"speaker\":\"Dispatcher\",\"text\":\"Another delivery attempt at Stop 42 was feasible.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817."}, {"path": ["1", "text"], "text": "On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "On 17 September 2026, a calibrated inspection recorded a crush depth of 2.37 cm for container C-817.", "negative_left": "On 17 September 2026, a calibrated inspection recorded a crush depth of 1.63 cm for container C-817.", "negative_right": "On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm.", "right": "On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm."}, "verifier_independent_model": false}, "family": "fast-41-diverse-296-005", "id": "fast-41-diverse-296-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Inspection record", "text": "On 17 September 2026, a calibrated inspection recorded a crush depth of 1.63 cm for container C-817."}, {"speaker": "Shipment record", "text": "On 17 September 2026, the package assigned to Stop 42 was container C-817, and the applicable crush threshold was 2.0 cm."}, {"speaker": "Condition log", "text": "The inspection recorded no leak, opening, or rattling for the package assigned to Stop 42."}, {"speaker": "Address record", "text": "The delivery address for Stop 42 was verified and usable, with no conflicting address entry."}, {"speaker": "Route record", "text": "Stop 42 had no recorded vehicle restriction, and the delivery history showed fewer than two failed attempts."}, {"speaker": "Access record", "text": "The access details required to perform Stop 42 were present and usable."}, {"speaker": "Dispatcher", "text": "Another delivery attempt at Stop 42 was feasible."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in the unchanged questions and both contexts. Stop 42, package P-714, and 17 September 2026 bindings remain consistent. The evidence contains the complete factual sentences \"A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026.\" and \"The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth.\" The counterfactual changes only the crush measurement to 1.6 cm and introduces no contradiction. Neither context embeds an answer, code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Inspection record\",\"text\":\"A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026.\"},{\"speaker\":\"Damage log\",\"text\":\"The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth.\"},{\"speaker\":\"Inspector\",\"text\":\"The package inspection recorded an intact seal and no leak, opening, or rattling.\"},{\"speaker\":\"Dispatch record\",\"text\":\"The delivery address for Stop 42 was verified and usable, and the required access instructions were present and usable.\"},{\"speaker\":\"Vehicle log\",\"text\":\"No recorded vehicle restriction affects service at Stop 42.\"},{\"speaker\":\"Attempt log\",\"text\":\"The delivery record shows one unsuccessful attempt, not two or more.\"},{\"speaker\":\"Dispatch record\",\"text\":\"Another delivery attempt at Stop 42 is feasible, with the recipient available the following day.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026."}, {"path": ["1", "text"], "text": "The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026.", "negative_left": "A calibrated gauge recorded a 1.6 cm indentation at the lower rear corner of package P-714 on 17 September 2026.", "negative_right": "The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth.", "right": "The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth."}, "verifier_independent_model": false}, "family": "fast-41-diverse-296-009", "id": "fast-41-diverse-296-009-base", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Inspection record", "text": "A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026."}, {"speaker": "Damage log", "text": "The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth."}, {"speaker": "Inspector", "text": "The package inspection recorded an intact seal and no leak, opening, or rattling."}, {"speaker": "Dispatch record", "text": "The delivery address for Stop 42 was verified and usable, and the required access instructions were present and usable."}, {"speaker": "Vehicle log", "text": "No recorded vehicle restriction affects service at Stop 42."}, {"speaker": "Attempt log", "text": "The delivery record shows one unsuccessful attempt, not two or more."}, {"speaker": "Dispatch record", "text": "Another delivery attempt at Stop 42 is feasible, with the recipient available the following day."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "damage_review"}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved in the unchanged questions and both contexts. Stop 42, package P-714, and 17 September 2026 bindings remain consistent. The evidence contains the complete factual sentences \"A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026.\" and \"The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth.\" The counterfactual changes only the crush measurement to 1.6 cm and introduces no contradiction. Neither context embeds an answer, code, rationale, or classifier instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus is the factual recorded crush threshold rather than a policy conclusion. The base and counter assignments are mutually realizable with only a1 changing and contain no policy-created inconsistency. The cited state passage accurately preserves the routing order and recorded-evidence rule; moreover, the unchanged questions object already preserves the complete governing criteria, including details not repeated in the citation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Supported a1 entails recorded evidence of a crush greater than 2.0 cm, which is independently sufficient for the highest-priority damage-review route.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes every earlier route: all damage indicators are refuted; address conflict is refuted and verification is supported; both depot-hold triggers are refuted; and missing access details is refuted. It also affirmatively establishes usable address and access details and a feasible further attempt, so redelivery is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a crush depth greater than 2.0 cm."}, {"id": "a2", "statement": "Recorded evidence shows that the package assigned to Stop 42 has a leak."}, {"id": "a3", "statement": "Recorded evidence shows that the package assigned to Stop 42 has an opening."}, {"id": "a4", "statement": "Recorded evidence shows that the package assigned to Stop 42 rattles."}, {"id": "a5", "statement": "The recorded delivery address for Stop 42 is conflicting."}, {"id": "a6", "statement": "The recorded delivery address for Stop 42 is verified."}, {"id": "a7", "statement": "At least two delivery attempts for Stop 42 have failed."}, {"id": "a8", "statement": "A recorded vehicle restriction prevents service at Stop 42."}, {"id": "a9", "statement": "Access details required to perform Stop 42 are missing."}, {"id": "a10", "statement": "The recorded delivery address for Stop 42 is usable."}, {"id": "a11", "statement": "The access details for Stop 42 are usable."}, {"id": "a12", "statement": "Another delivery attempt at Stop 42 is feasible."}], "base_state_json": "[{\"speaker\":\"Inspection record\",\"text\":\"A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026.\"},{\"speaker\":\"Damage log\",\"text\":\"The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth.\"},{\"speaker\":\"Inspector\",\"text\":\"The package inspection recorded an intact seal and no leak, opening, or rattling.\"},{\"speaker\":\"Dispatch record\",\"text\":\"The delivery address for Stop 42 was verified and usable, and the required access instructions were present and usable.\"},{\"speaker\":\"Vehicle log\",\"text\":\"No recorded vehicle restriction affects service at Stop 42.\"},{\"speaker\":\"Attempt log\",\"text\":\"The delivery record shows one unsuccessful attempt, not two or more.\"},{\"speaker\":\"Dispatch record\",\"text\":\"Another delivery attempt at Stop 42 is feasible, with the recipient available the following day.\"},{\"speaker\":\"Depot operations lead\",\"text\":\"Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026."}, {"path": ["1", "text"], "text": "The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth."}], "policy_evidence": [{"path": ["2", "text"], "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}], "rules": [{"justification": "A recorded crush greater than 2.0 cm independently triggers the first route, damage review.", "target": "damage_review", "when": [{"atom_id": "a1", "state": "supported"}]}, {"justification": "No damage-review, address-clarification, depot-hold, or customer-contact condition applies; the address and access details are usable, and another attempt is feasible.", "target": "redelivery", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}]}]}, "verified_pair": {"left": "A calibrated gauge recorded a 2.4 cm indentation at the lower rear corner of package P-714 on 17 September 2026.", "negative_left": "A calibrated gauge recorded a 1.6 cm indentation at the lower rear corner of package P-714 on 17 September 2026.", "negative_right": "The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth.", "right": "The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth."}, "verifier_independent_model": false}, "family": "fast-41-diverse-296-009", "id": "fast-41-diverse-296-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"address_clarification": "Choose only when no damage-review condition applies and the recorded delivery address is conflicting or unverified.", "customer_contact": "Choose only when no earlier condition applies and access details required to perform the stop are missing.", "damage_review": "Choose first only when recorded evidence shows a crush over 2.0 cm, a leak, an opening, or rattling.", "depot_hold": "Choose only when neither prior condition applies and there have been at least two failed attempts or a recorded vehicle restriction prevents service.", "redelivery": "Choose only when no earlier condition applies, the address and access details are usable, and another delivery attempt is feasible."}, "instructions": "Select the required exception route by applying the stated first-match routing order. Treat threshold language literally and rely only on the identified recorded evidence.", "type": "choice"}}, "state": [{"speaker": "Inspection record", "text": "A calibrated gauge recorded a 1.6 cm indentation at the lower rear corner of package P-714 on 17 September 2026."}, {"speaker": "Damage log", "text": "The 17 September 2026 damage log identifies package P-714 as the package assigned to Stop 42 and states that the gauge reading at its lower rear corner is the package's crush depth."}, {"speaker": "Inspector", "text": "The package inspection recorded an intact seal and no leak, opening, or rattling."}, {"speaker": "Dispatch record", "text": "The delivery address for Stop 42 was verified and usable, and the required access instructions were present and usable."}, {"speaker": "Vehicle log", "text": "No recorded vehicle restriction affects service at Stop 42."}, {"speaker": "Attempt log", "text": "The delivery record shows one unsuccessful attempt, not two or more."}, {"speaker": "Dispatch record", "text": "Another delivery attempt at Stop 42 is feasible, with the recipient available the following day."}, {"speaker": "Depot operations lead", "text": "Use recorded evidence only. In order: damage review for a crush over 2.0 cm, leak, opening, or rattle; address clarification for an address conflict; depot hold after two failed attempts or a vehicle restriction; customer contact for missing access details; otherwise redelivery."}]}, "method": "c2d", "provenance": {"source_id": "diverse-296", "source_is_synthetic": true, "source_sha256": "abe464869c68982aefd00e3b96df5312e02ad05dde4f3c2b6de8d771f7e9fb73", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "redelivery"}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved by the unchanged original questions and both contexts retain the stop instruction without adding exceptions. Question bindings for Mina, parcel FM-297, Stop 18, and the redelivery decision remain unchanged. The two evidence spans are complete factual sentences: “Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel.” and “Entry ADR-882's request-type field read “alternate delivery.”” The counterfactual coherently changes the sole request type to “no delivery request” while retaining all compatible facts and counts. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction. No contradictory duplicate measurements, counts, or assertions appear.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Dispatcher Mina reviewed parcel FM-297 after driver Joel returned from Stop 18. At 15:42, the scan recorded an attempted delivery with the named recipient unavailable; Joel reported that nobody answered and no authorized neighbor was present. The address was complete, the van had sufficient capacity, the sealed carton was undamaged, and depot lead Priya confirmed secure hold space. Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel. Entry ADR-882's request-type field read “alternate delivery.” The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Customer service agent Leon had not received any other routing communication. The parcel remained at the depot pending Mina's review, and no delivery attempt occurred after Stop 18.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel."}, {"path": [], "text": "Entry ADR-882's request-type field read “alternate delivery.”"}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel.", "negative_left": "Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel.", "negative_right": "Entry ADR-882's request-type field read “no delivery request.”", "right": "Entry ADR-882's request-type field read “alternate delivery.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-297-005", "id": "fast-41-diverse-297-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Dispatcher Mina reviewed parcel FM-297 after driver Joel returned from Stop 18. At 15:42, the scan recorded an attempted delivery with the named recipient unavailable; Joel reported that nobody answered and no authorized neighbor was present. The address was complete, the van had sufficient capacity, the sealed carton was undamaged, and depot lead Priya confirmed secure hold space. Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel. Entry ADR-882's request-type field read “alternate delivery.” The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Customer service agent Leon had not received any other routing communication. The parcel remained at the depot pending Mina's review, and no delivery attempt occurred after Stop 18."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Policy is preserved by the unchanged original questions and both contexts retain the stop instruction without adding exceptions. Question bindings for Mina, parcel FM-297, Stop 18, and the redelivery decision remain unchanged. The two evidence spans are complete factual sentences: “Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel.” and “Entry ADR-882's request-type field read “alternate delivery.”” The counterfactual coherently changes the sole request type to “no delivery request” while retaining all compatible facts and counts. Neither context contains a gold answer, answer code, rationale, proposition ID, or output instruction. No contradictory duplicate measurements, counts, or assertions appear.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Dispatcher Mina reviewed parcel FM-297 after driver Joel returned from Stop 18. At 15:42, the scan recorded an attempted delivery with the named recipient unavailable; Joel reported that nobody answered and no authorized neighbor was present. The address was complete, the van had sufficient capacity, the sealed carton was undamaged, and depot lead Priya confirmed secure hold space. Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel. Entry ADR-882's request-type field read “alternate delivery.” The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Customer service agent Leon had not received any other routing communication. The parcel remained at the depot pending Mina's review, and no delivery attempt occurred after Stop 18.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel."}, {"path": [], "text": "Entry ADR-882's request-type field read “alternate delivery.”"}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel.", "negative_left": "Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel.", "negative_right": "Entry ADR-882's request-type field read “no delivery request.”", "right": "Entry ADR-882's request-type field read “alternate delivery.”"}, "verifier_independent_model": false}, "family": "fast-41-diverse-297-005", "id": "fast-41-diverse-297-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Dispatcher Mina reviewed parcel FM-297 after driver Joel returned from Stop 18. At 15:42, the scan recorded an attempted delivery with the named recipient unavailable; Joel reported that nobody answered and no authorized neighbor was present. The address was complete, the van had sufficient capacity, the sealed carton was undamaged, and depot lead Priya confirmed secure hold space. Before dispatcher Mina's routing decision for parcel FM-297, the routing ledger contained exactly one entry, ADR-882, for that parcel. Entry ADR-882's request-type field read “no delivery request.” The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Customer service agent Leon had not received any other routing communication. The parcel remained at the depot pending Mina's review, and no delivery attempt occurred after Stop 18."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision scope. Mina, parcel FM-297, Stop 18, and the 3 June 2026 routing time remain bound consistently. Evidence 1 is a complete factual sentence: \"Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297.\" Evidence 2 is a complete factual sentence: \"The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request.\" The counterfactual changes only the request category and introduces no contradictory duplicate assertion. Neither context states a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"On 3 June 2026, dispatcher Mina reviewed parcel FM-297 after the 15:42 delivery attempt at Stop 18. The named recipient was unavailable, and driver Joel recorded that nobody answered; no authorized neighbor was present. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297. The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request. The address was complete, the sealed carton was undamaged, and the vehicle had sufficient capacity. Priya confirmed secure depot-hold space was available, while Mina retained responsibility for the routing decision.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297."}, {"path": [], "text": "The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297.", "negative_left": "Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297.", "negative_right": "The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as a customer-information request.", "right": "The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request."}, "verifier_independent_model": false}, "family": "fast-41-diverse-297-007", "id": "fast-41-diverse-297-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "On 3 June 2026, dispatcher Mina reviewed parcel FM-297 after the 15:42 delivery attempt at Stop 18. The named recipient was unavailable, and driver Joel recorded that nobody answered; no authorized neighbor was present. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297. The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request. The address was complete, the sealed carton was undamaged, and the vehicle had sufficient capacity. Priya confirmed secure depot-hold space was available, while Mina retained responsibility for the routing decision."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged questions preserve the governing policy and decision scope. Mina, parcel FM-297, Stop 18, and the 3 June 2026 routing time remain bound consistently. Evidence 1 is a complete factual sentence: \"Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297.\" Evidence 2 is a complete factual sentence: \"The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request.\" The counterfactual changes only the request category and introduces no contradictory duplicate assertion. Neither context states a gold answer, answer code, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"On 3 June 2026, dispatcher Mina reviewed parcel FM-297 after the 15:42 delivery attempt at Stop 18. The named recipient was unavailable, and driver Joel recorded that nobody answered; no authorized neighbor was present. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297. The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request. The address was complete, the sealed carton was undamaged, and the vehicle had sufficient capacity. Priya confirmed secure depot-hold space was available, while Mina retained responsibility for the routing decision.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297."}, {"path": [], "text": "The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297.", "negative_left": "Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297.", "negative_right": "The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as a customer-information request.", "right": "The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as an alternate-delivery request."}, "verifier_independent_model": false}, "family": "fast-41-diverse-297-007", "id": "fast-41-diverse-297-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "On 3 June 2026, dispatcher Mina reviewed parcel FM-297 after the 15:42 delivery attempt at Stop 18. The named recipient was unavailable, and driver Joel recorded that nobody answered; no authorized neighbor was present. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” Before dispatcher Mina's 15:45 routing decision on 3 June 2026, Mina logged the sole request record AD-814 for parcel FM-297. The sole request record AD-814 in dispatcher Mina's queue, logged before 15:45 on 3 June 2026, was categorized as a customer-information request. The address was complete, the sealed carton was undamaged, and the vehicle had sufficient capacity. Priya confirmed secure depot-hold space was available, while Mina retained responsibility for the routing decision."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy and scoring criteria. The parcel, depot-hold request, completion-status distinction, and relevant time bindings remain intact. The evidence consists of exactly two complete factual sentences: “At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284.” and “At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled.” The counterfactual coherently changes the completion-record statuses without contradicting the other observations. Neither context embeds an answer, score, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, including the universally scoped record-status and contradiction atoms. The focus atom is factual rather than policy-based. The base and counter assignments are jointly realizable with only a5 changing: an existing completion scan may either have been reconciled or remain unresolved. Empty policy_evidence is correct because the governing scoring criteria and instructions are entirely in the automatically retained questions object; no additional state-originating policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a later depot inbound scan tied to FX-184, physical label confirmation, reconciliation or voiding of every completion record, and absence of other contradictory evidence. This is sufficient for level 3.", "rule_index": 0, "sound": true}, {"reason": "The inbound scan and physical label confirmation establish depot custody. Refutation of the universal reconciliation atom entails that at least one completion record remains unvoided and unreconciled, excluding level 3 and satisfying the unresolved-conflict condition for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evidence record contains a delivery-completion scan whose parcel identifier is FX-184."}, {"id": "a2", "statement": "The depot inbound-scan record at issue has parcel identifier FX-184."}, {"id": "a3", "statement": "The timestamp of the depot inbound-scan record for FX-184 is later than the timestamp of the delivery-completion scan for FX-184."}, {"id": "a4", "statement": "The label on the carton physically inspected at the depot identifies the carton as FX-184."}, {"id": "a5", "statement": "As of the rating time, every delivery-completion scan or equivalent completion record for FX-184 has been voided or reconciled."}, {"id": "a6", "statement": "As of the rating time, no evidence other than a delivery-completion scan or equivalent completion record contradicts current depot custody of FX-184."}], "base_state_json": "\"The depot case concerns parcel FX-184, which is being assessed for routing to depot hold rather than treatment as completed. The evidence record contains a delivery-completion scan whose parcel identifier is FX-184. The depot inbound-scan record at issue has parcel identifier FX-184, and its timestamp is later than the timestamp of the delivery-completion scan for FX-184. The label on the carton physically inspected at the depot identifies the carton as FX-184. The depot inspection found the carton intact, and the depot custody log contains no separate incident, transfer, or recipient-confirmation entry. As of the rating time, no evidence other than a delivery-completion scan or equivalent completion record contradicts current depot custody of FX-184. At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284. At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled. The depot team retained the parcel while the rating was made.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284."}, {"path": [], "text": "At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled."}], "policy_evidence": [], "rules": [{"justification": "The later depot inbound scan and physical label confirmation establish custody, while every completion record is voided or reconciled and no other contradictory evidence remains.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The later depot inbound scan and physical label confirmation establish custody, but refutation of universal reconciliation entails that at least one completion record remains unresolved and contradictory.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284.", "negative_left": "At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284.", "negative_right": "At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked neither voided nor reconciled.", "right": "At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled."}, "verifier_independent_model": false}, "family": "fast-41-diverse-300-005", "id": "fast-41-diverse-300-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsupported: No depot scan or physical confirmation exists, and reliable records instead show successful delivery or custody elsewhere.", "1 — Weakly supported: A driver note or customer report suggests depot return, but there is no depot inbound scan or physical package confirmation.", "2 — Strongly supported with a conflict: A depot inbound scan and physical label confirmation establish depot custody, but an unresolved delivery-completion scan or equivalent record still contradicts the claim.", "3 — Fully verified: The depot inbound scan and physical confirmation establish custody, and all completion records have been voided or reconciled with no remaining contradictory evidence."], "instructions": "Rate the evidence supporting the claim that FX-184 is currently at the depot and should be routed to depot hold rather than treated as completed. Apply the ordered verification levels below. A later inbound scan plus physical label confirmation establishes depot custody, but any contradictory completion scan prevents the highest level until corrected.", "type": "score"}}, "state": "The depot case concerns parcel FX-184, which is being assessed for routing to depot hold rather than treatment as completed. The evidence record contains a delivery-completion scan whose parcel identifier is FX-184. The depot inbound-scan record at issue has parcel identifier FX-184, and its timestamp is later than the timestamp of the delivery-completion scan for FX-184. The label on the carton physically inspected at the depot identifies the carton as FX-184. The depot inspection found the carton intact, and the depot custody log contains no separate incident, transfer, or recipient-confirmation entry. As of the rating time, no evidence other than a delivery-completion scan or equivalent completion record contradicts current depot custody of FX-184. At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284. At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled. The depot team retained the parcel while the rating was made."}, "method": "c2d", "provenance": {"source_id": "diverse-300", "source_is_synthetic": true, "source_sha256": "1e2508e0c80516c96c8e4568290fee85c6627ec2622e428e2eeceee6312eb181", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The unchanged original questions preserve all governing policy and scoring criteria. The parcel, depot-hold request, completion-status distinction, and relevant time bindings remain intact. The evidence consists of exactly two complete factual sentences: “At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284.” and “At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled.” The counterfactual coherently changes the completion-record statuses without contradicting the other observations. Neither context embeds an answer, score, rationale, proposition identifier, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, including the universally scoped record-status and contradiction atoms. The focus atom is factual rather than policy-based. The base and counter assignments are jointly realizable with only a5 changing: an existing completion scan may either have been reconciled or remain unresolved. Empty policy_evidence is correct because the governing scoring criteria and instructions are entirely in the automatically retained questions object; no additional state-originating policy is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a later depot inbound scan tied to FX-184, physical label confirmation, reconciliation or voiding of every completion record, and absence of other contradictory evidence. This is sufficient for level 3.", "rule_index": 0, "sound": true}, {"reason": "The inbound scan and physical label confirmation establish depot custody. Refutation of the universal reconciliation atom entails that at least one completion record remains unvoided and unreconciled, excluding level 3 and satisfying the unresolved-conflict condition for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The evidence record contains a delivery-completion scan whose parcel identifier is FX-184."}, {"id": "a2", "statement": "The depot inbound-scan record at issue has parcel identifier FX-184."}, {"id": "a3", "statement": "The timestamp of the depot inbound-scan record for FX-184 is later than the timestamp of the delivery-completion scan for FX-184."}, {"id": "a4", "statement": "The label on the carton physically inspected at the depot identifies the carton as FX-184."}, {"id": "a5", "statement": "As of the rating time, every delivery-completion scan or equivalent completion record for FX-184 has been voided or reconciled."}, {"id": "a6", "statement": "As of the rating time, no evidence other than a delivery-completion scan or equivalent completion record contradicts current depot custody of FX-184."}], "base_state_json": "\"The depot case concerns parcel FX-184, which is being assessed for routing to depot hold rather than treatment as completed. The evidence record contains a delivery-completion scan whose parcel identifier is FX-184. The depot inbound-scan record at issue has parcel identifier FX-184, and its timestamp is later than the timestamp of the delivery-completion scan for FX-184. The label on the carton physically inspected at the depot identifies the carton as FX-184. The depot inspection found the carton intact, and the depot custody log contains no separate incident, transfer, or recipient-confirmation entry. As of the rating time, no evidence other than a delivery-completion scan or equivalent completion record contradicts current depot custody of FX-184. At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284. At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled. The depot team retained the parcel while the rating was made.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284."}, {"path": [], "text": "At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled."}], "policy_evidence": [], "rules": [{"justification": "The later depot inbound scan and physical label confirmation establish custody, while every completion record is voided or reconciled and no other contradictory evidence remains.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The later depot inbound scan and physical label confirmation establish custody, but refutation of universal reconciliation entails that at least one completion record remains unresolved and contradictory.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284.", "negative_left": "At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284.", "negative_right": "At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked neither voided nor reconciled.", "right": "At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked voided or reconciled."}, "verifier_independent_model": false}, "family": "fast-41-diverse-300-005", "id": "fast-41-diverse-300-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsupported: No depot scan or physical confirmation exists, and reliable records instead show successful delivery or custody elsewhere.", "1 — Weakly supported: A driver note or customer report suggests depot return, but there is no depot inbound scan or physical package confirmation.", "2 — Strongly supported with a conflict: A depot inbound scan and physical label confirmation establish depot custody, but an unresolved delivery-completion scan or equivalent record still contradicts the claim.", "3 — Fully verified: The depot inbound scan and physical confirmation establish custody, and all completion records have been voided or reconciled with no remaining contradictory evidence."], "instructions": "Rate the evidence supporting the claim that FX-184 is currently at the depot and should be routed to depot hold rather than treated as completed. Apply the ordered verification levels below. A later inbound scan plus physical label confirmation establishes depot custody, but any contradictory completion scan prevents the highest level until corrected.", "type": "score"}}, "state": "The depot case concerns parcel FX-184, which is being assessed for routing to depot hold rather than treatment as completed. The evidence record contains a delivery-completion scan whose parcel identifier is FX-184. The depot inbound-scan record at issue has parcel identifier FX-184, and its timestamp is later than the timestamp of the delivery-completion scan for FX-184. The label on the carton physically inspected at the depot identifies the carton as FX-184. The depot inspection found the carton intact, and the depot custody log contains no separate incident, transfer, or recipient-confirmation entry. As of the rating time, no evidence other than a delivery-completion scan or equivalent completion record contradicts current depot custody of FX-184. At the rating time of 2026-09-17 12:00 UTC, the complete evidence record for FX-184 contains exactly two records that are either delivery-completion scans or equivalent completion records, identified as DC-271 and DC-284. At 2026-09-17 12:00 UTC, records DC-271 and DC-284 are each marked neither voided nor reconciled. The depot team retained the parcel while the rating was made."}, "method": "c2d", "provenance": {"source_id": "diverse-300", "source_is_synthetic": true, "source_sha256": "1e2508e0c80516c96c8e4568290fee85c6627ec2622e428e2eeceee6312eb181", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy via the unchanged questions object and only alter the contact date (Sept19 vs Sept25) while preserving account, offer, and payment facts; the two focus evidence spans are complete factual sentences, the counterfactual shift plausibly changes timing without contradicting other facts, and neither context states or hints at the correct resolution.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "full_context_fact_states": {"base": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "counterfactual": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "remove_left": {"request_within_seven_days": "unknown"}, "remove_right": {"request_within_seven_days": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"request_within_seven_days": "unknown"}, "negative_pair": {"request_within_seven_days": "refuted"}, "negative_sentence": {"request_within_seven_days": "unknown"}, "positive_pair": {"request_within_seven_days": "supported"}, "right": {"request_within_seven_days": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single relevant factual relationship; none is a bundled final resolution classification. The focus is the factual timing relationship between the correction request and renewal. Base and counter assignments can be realized by changing only the request date from within the seven-day window to outside it while holding the remaining facts fixed. Empty policy_evidence is correct because all substantive governing rules are already preserved in the unchanged questions object; the original state supplies case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a recorded and valid offer, pre-renewal conditional acceptance tied to applying the credit, omission of the credit, and a correction request within seven calendar days. These facts satisfy the credit-and-annual-resolution rubric. Refutation of two settled charges also excludes duplicate-payment routing.", "rule_index": 0, "sound": true}, {"reason": "With the other eligibility facts fixed, a correction request explicitly outside the seven-calendar-day limit makes the credit ineligible. The conjunction also establishes each expressly requested fallback component—annual cancellation, movement to monthly, and the unused-term refund—and refutes the charge multiplicity required for duplicate-payment routing.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "offer_recorded", "statement": "A $24 retention offer was recorded on the subscriber's account for the disputed September 14 annual renewal."}, {"id": "offer_valid", "statement": "The $24 retention offer recorded on the subscriber's account was valid for the disputed September 14 annual renewal."}, {"id": "acceptance_before_renewal", "statement": "The subscriber accepted the recorded $24 retention offer before the disputed September 14 annual renewal posted."}, {"id": "acceptance_condition_credit", "statement": "The condition in the subscriber's acceptance of the recorded offer was application of the eligible $24 retention credit."}, {"id": "credit_omitted", "statement": "The $24 retention credit was omitted from invoice INV-8841 for the disputed September 14 annual renewal."}, {"id": "request_within_seven_days", "statement": "The subscriber's correction request in the current contact was made no later than seven calendar days after the disputed September 14 annual renewal."}, {"id": "fallback_cancel", "statement": "The subscriber expressly requested cancellation of the annual subscription if the $24 retention credit was ineligible."}, {"id": "fallback_monthly", "statement": "The subscriber expressly requested a move to a monthly subscription if the $24 retention credit was ineligible."}, {"id": "fallback_refund", "statement": "The subscriber expressly requested the permitted refund of the unused annual term if the $24 retention credit was ineligible."}, {"id": "two_settled_charges", "statement": "The subscriber's account shows at least two settled subscription charges for the disputed September 14 annual renewal."}], "base_state_json": "[{\"speaker\":\"Subscriber\",\"text\":\"On account SUB-7734, I replied September 13 accepting the retention offer, saying I'd keep annual provided the promised $24 credit is applied. If it can't be applied, cancel annual, move me to monthly, and refund the unused annual term.\"},{\"speaker\":\"Billing support agent\",\"text\":\"The disputed annual renewal for subscriber account SUB-7734 posted on September 14. The $24 offer, logged September 13 at 4:10 p.m., is valid for annual renewals and was accepted before renewal, but it was omitted from invoice INV-8841.\"},{\"speaker\":\"Subscription operations analyst\",\"text\":\"The subscriber's current contact regarding this renewal was logged on September 19. The account shows only the original $240 settled payment; no second charge, adjustment, or monthly invoice appears.\"}]", "base_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "counter_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "focus_atom": "request_within_seven_days", "focus_evidence": [{"path": ["1", "text"], "text": "The disputed annual renewal for subscriber account SUB-7734 posted on September 14."}, {"path": ["2", "text"], "text": "The subscriber's current contact regarding this renewal was logged on September 19."}], "policy_evidence": [], "rules": [{"justification": "The omitted $24 credit is eligible because the valid recorded offer was accepted before renewal, the acceptance condition was application of that eligible credit, and correction was requested within seven calendar days. Refutation of two settled charges excludes duplicate-payment routing, and the actionable primary request excludes clarification and fallback.", "target": "apply_credit_keep_annual", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}, {"justification": "An explicitly late correction request makes the omitted credit ineligible under the seven-calendar-day requirement. The subscriber expressly supplied every component of the authorized fallback. Refutation of two settled charges excludes duplicate-payment routing, and the actionable fallback excludes clarification and none of the above.", "target": "execute_fallback_cancellation", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}]}, "verified_pair": {"left": "The disputed annual renewal for subscriber account SUB-7734 posted on September 14.", "negative_left": "The disputed annual renewal for subscriber account SUB-7734 posted on September 14.", "negative_right": "The subscriber's current contact regarding this renewal was logged on September 25.", "right": "The subscriber's current contact regarding this renewal was logged on September 19."}, "verifier_independent_model": false}, "family": "fast-43-diverse-013-013", "id": "fast-43-diverse-013-013-base", "input": {"questions": {"decision": {"criteria": {"apply_credit_keep_annual": "Add the $24 credit and preserve the annual subscription when the offer and pre-renewal acceptance are verified and the correction request is within seven calendar days.", "execute_fallback_cancellation": "Cancel annual, move to monthly, and process the permitted unused-term refund only when the credit is ineligible and the subscriber expressly supplied this fallback.", "none_of_above": "Use only when the evidence satisfies none of the four resolution rubrics above.", "route_duplicate_payment_review": "Route for duplicate-payment investigation only when the account shows at least two settled subscription charges for the disputed renewal.", "seek_intent_clarification": "Ask the subscriber to clarify only when neither the primary conditional request nor any fallback identifies an actionable outcome."}, "instructions": "Choose the authorized resolution. Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges.", "type": "choice"}}, "state": [{"speaker": "Subscriber", "text": "On account SUB-7734, I replied September 13 accepting the retention offer, saying I'd keep annual provided the promised $24 credit is applied. If it can't be applied, cancel annual, move me to monthly, and refund the unused annual term."}, {"speaker": "Billing support agent", "text": "The disputed annual renewal for subscriber account SUB-7734 posted on September 14. The $24 offer, logged September 13 at 4:10 p.m., is valid for annual renewals and was accepted before renewal, but it was omitted from invoice INV-8841."}, {"speaker": "Subscription operations analyst", "text": "The subscriber's current contact regarding this renewal was logged on September 19. The account shows only the original $240 settled payment; no second charge, adjustment, or monthly invoice appears."}]}, "method": "c2d", "provenance": {"source_id": "diverse-013", "source_is_synthetic": true, "source_sha256": "640e873611af9a131e8dfe8de5027d31b2e53121b8e09c71e14040fbe16c6d8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "apply_credit_keep_annual"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy via the unchanged questions object and only alter the contact date (Sept19 vs Sept25) while preserving account, offer, and payment facts; the two focus evidence spans are complete factual sentences, the counterfactual shift plausibly changes timing without contradicting other facts, and neither context states or hints at the correct resolution.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "full_context_fact_states": {"base": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "supported", "two_settled_charges": "refuted"}, "counterfactual": {"acceptance_before_renewal": "supported", "acceptance_condition_credit": "supported", "credit_omitted": "supported", "fallback_cancel": "supported", "fallback_monthly": "supported", "fallback_refund": "supported", "offer_recorded": "supported", "offer_valid": "supported", "request_within_seven_days": "refuted", "two_settled_charges": "refuted"}, "remove_left": {"request_within_seven_days": "unknown"}, "remove_right": {"request_within_seven_days": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"request_within_seven_days": "unknown"}, "negative_pair": {"request_within_seven_days": "refuted"}, "negative_sentence": {"request_within_seven_days": "unknown"}, "positive_pair": {"request_within_seven_days": "supported"}, "right": {"request_within_seven_days": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single relevant factual relationship; none is a bundled final resolution classification. The focus is the factual timing relationship between the correction request and renewal. Base and counter assignments can be realized by changing only the request date from within the seven-day window to outside it while holding the remaining facts fixed. Empty policy_evidence is correct because all substantive governing rules are already preserved in the unchanged questions object; the original state supplies case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a recorded and valid offer, pre-renewal conditional acceptance tied to applying the credit, omission of the credit, and a correction request within seven calendar days. These facts satisfy the credit-and-annual-resolution rubric. Refutation of two settled charges also excludes duplicate-payment routing.", "rule_index": 0, "sound": true}, {"reason": "With the other eligibility facts fixed, a correction request explicitly outside the seven-calendar-day limit makes the credit ineligible. The conjunction also establishes each expressly requested fallback component—annual cancellation, movement to monthly, and the unused-term refund—and refutes the charge multiplicity required for duplicate-payment routing.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "offer_recorded", "statement": "A $24 retention offer was recorded on the subscriber's account for the disputed September 14 annual renewal."}, {"id": "offer_valid", "statement": "The $24 retention offer recorded on the subscriber's account was valid for the disputed September 14 annual renewal."}, {"id": "acceptance_before_renewal", "statement": "The subscriber accepted the recorded $24 retention offer before the disputed September 14 annual renewal posted."}, {"id": "acceptance_condition_credit", "statement": "The condition in the subscriber's acceptance of the recorded offer was application of the eligible $24 retention credit."}, {"id": "credit_omitted", "statement": "The $24 retention credit was omitted from invoice INV-8841 for the disputed September 14 annual renewal."}, {"id": "request_within_seven_days", "statement": "The subscriber's correction request in the current contact was made no later than seven calendar days after the disputed September 14 annual renewal."}, {"id": "fallback_cancel", "statement": "The subscriber expressly requested cancellation of the annual subscription if the $24 retention credit was ineligible."}, {"id": "fallback_monthly", "statement": "The subscriber expressly requested a move to a monthly subscription if the $24 retention credit was ineligible."}, {"id": "fallback_refund", "statement": "The subscriber expressly requested the permitted refund of the unused annual term if the $24 retention credit was ineligible."}, {"id": "two_settled_charges", "statement": "The subscriber's account shows at least two settled subscription charges for the disputed September 14 annual renewal."}], "base_state_json": "[{\"speaker\":\"Subscriber\",\"text\":\"On account SUB-7734, I replied September 13 accepting the retention offer, saying I'd keep annual provided the promised $24 credit is applied. If it can't be applied, cancel annual, move me to monthly, and refund the unused annual term.\"},{\"speaker\":\"Billing support agent\",\"text\":\"The disputed annual renewal for subscriber account SUB-7734 posted on September 14. The $24 offer, logged September 13 at 4:10 p.m., is valid for annual renewals and was accepted before renewal, but it was omitted from invoice INV-8841.\"},{\"speaker\":\"Subscription operations analyst\",\"text\":\"The subscriber's current contact regarding this renewal was logged on September 19. The account shows only the original $240 settled payment; no second charge, adjustment, or monthly invoice appears.\"}]", "base_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "counter_states": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}], "focus_atom": "request_within_seven_days", "focus_evidence": [{"path": ["1", "text"], "text": "The disputed annual renewal for subscriber account SUB-7734 posted on September 14."}, {"path": ["2", "text"], "text": "The subscriber's current contact regarding this renewal was logged on September 19."}], "policy_evidence": [], "rules": [{"justification": "The omitted $24 credit is eligible because the valid recorded offer was accepted before renewal, the acceptance condition was application of that eligible credit, and correction was requested within seven calendar days. Refutation of two settled charges excludes duplicate-payment routing, and the actionable primary request excludes clarification and fallback.", "target": "apply_credit_keep_annual", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}, {"justification": "An explicitly late correction request makes the omitted credit ineligible under the seven-calendar-day requirement. The subscriber expressly supplied every component of the authorized fallback. Refutation of two settled charges excludes duplicate-payment routing, and the actionable fallback excludes clarification and none of the above.", "target": "execute_fallback_cancellation", "when": [{"atom_id": "offer_recorded", "state": "supported"}, {"atom_id": "offer_valid", "state": "supported"}, {"atom_id": "acceptance_before_renewal", "state": "supported"}, {"atom_id": "acceptance_condition_credit", "state": "supported"}, {"atom_id": "credit_omitted", "state": "supported"}, {"atom_id": "request_within_seven_days", "state": "refuted"}, {"atom_id": "fallback_cancel", "state": "supported"}, {"atom_id": "fallback_monthly", "state": "supported"}, {"atom_id": "fallback_refund", "state": "supported"}, {"atom_id": "two_settled_charges", "state": "refuted"}]}]}, "verified_pair": {"left": "The disputed annual renewal for subscriber account SUB-7734 posted on September 14.", "negative_left": "The disputed annual renewal for subscriber account SUB-7734 posted on September 14.", "negative_right": "The subscriber's current contact regarding this renewal was logged on September 25.", "right": "The subscriber's current contact regarding this renewal was logged on September 19."}, "verifier_independent_model": false}, "family": "fast-43-diverse-013-013", "id": "fast-43-diverse-013-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"apply_credit_keep_annual": "Add the $24 credit and preserve the annual subscription when the offer and pre-renewal acceptance are verified and the correction request is within seven calendar days.", "execute_fallback_cancellation": "Cancel annual, move to monthly, and process the permitted unused-term refund only when the credit is ineligible and the subscriber expressly supplied this fallback.", "none_of_above": "Use only when the evidence satisfies none of the four resolution rubrics above.", "route_duplicate_payment_review": "Route for duplicate-payment investigation only when the account shows at least two settled subscription charges for the disputed renewal.", "seek_intent_clarification": "Ask the subscriber to clarify only when neither the primary conditional request nor any fallback identifies an actionable outcome."}, "instructions": "Choose the authorized resolution. Policy: an omitted retention credit may be added when a valid offer was recorded and accepted before renewal and correction is requested no later than seven calendar days after renewal. Conditional acceptance counts when its condition is applying that eligible credit. If the credit is ineligible, follow an expressly stated fallback request; otherwise seek clarification. Duplicate-payment routing requires two settled charges.", "type": "choice"}}, "state": [{"speaker": "Subscriber", "text": "On account SUB-7734, I replied September 13 accepting the retention offer, saying I'd keep annual provided the promised $24 credit is applied. If it can't be applied, cancel annual, move me to monthly, and refund the unused annual term."}, {"speaker": "Billing support agent", "text": "The disputed annual renewal for subscriber account SUB-7734 posted on September 14. The $24 offer, logged September 13 at 4:10 p.m., is valid for annual renewals and was accepted before renewal, but it was omitted from invoice INV-8841."}, {"speaker": "Subscription operations analyst", "text": "The subscriber's current contact regarding this renewal was logged on September 25. The account shows only the original $240 settled payment; no second charge, adjustment, or monthly invoice appears."}]}, "method": "c2d", "provenance": {"source_id": "diverse-013", "source_is_synthetic": true, "source_sha256": "640e873611af9a131e8dfe8de5027d31b2e53121b8e09c71e14040fbe16c6d8d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "execute_fallback_cancellation"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate the governing criteria facts (no tax/promo/plan-change, invoice not voided, single unrefunded payment) alongside the unchanged question, preserve the same entity/issue bindings, use two complete factual timeline sentences as evidence, and the counterfactual only swaps R's timestamp to 6:40 PM without contradicting any other stated fact or leaking a rule name, code, or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note for Mira Chen's billing case. Two verified billing issues remain unresolved and are the only such issues on file: issue T and issue R. Issue T concerns only correction of sales tax on invoice R-204; nothing else about it is in dispute. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the subscription period in question for issue R, and the merchant has not issued any refund for it. The timeline logs for both issues were reviewed to determine which issue was most recently updated. The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9. The timeline for billing issue R shows its most recent update logged at 2:40 PM on June 9.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9."}, {"path": [], "text": "The timeline for billing issue R shows its most recent update logged at 2:40 PM on June 9."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9.", "negative_left": "The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9.", "negative_right": "The timeline for billing issue R shows its most recent update logged at 6:40 PM on June 9.", "right": "The timeline for billing issue R shows its most recent update logged at 2:40 PM on June 9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-001", "id": "fast-43-diverse-014-001-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note for Mira Chen's billing case. Two verified billing issues remain unresolved and are the only such issues on file: issue T and issue R. Issue T concerns only correction of sales tax on invoice R-204; nothing else about it is in dispute. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the subscription period in question for issue R, and the merchant has not issued any refund for it. The timeline logs for both issues were reviewed to determine which issue was most recently updated. The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9. The timeline for billing issue R shows its most recent update logged at 2:40 PM on June 9."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate the governing criteria facts (no tax/promo/plan-change, invoice not voided, single unrefunded payment) alongside the unchanged question, preserve the same entity/issue bindings, use two complete factual timeline sentences as evidence, and the counterfactual only swaps R's timestamp to 6:40 PM without contradicting any other stated fact or leaking a rule name, code, or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note for Mira Chen's billing case. Two verified billing issues remain unresolved and are the only such issues on file: issue T and issue R. Issue T concerns only correction of sales tax on invoice R-204; nothing else about it is in dispute. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the subscription period in question for issue R, and the merchant has not issued any refund for it. The timeline logs for both issues were reviewed to determine which issue was most recently updated. The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9. The timeline for billing issue R shows its most recent update logged at 2:40 PM on June 9.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9."}, {"path": [], "text": "The timeline for billing issue R shows its most recent update logged at 2:40 PM on June 9."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9.", "negative_left": "The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9.", "negative_right": "The timeline for billing issue R shows its most recent update logged at 6:40 PM on June 9.", "right": "The timeline for billing issue R shows its most recent update logged at 2:40 PM on June 9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-001", "id": "fast-43-diverse-014-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note for Mira Chen's billing case. Two verified billing issues remain unresolved and are the only such issues on file: issue T and issue R. Issue T concerns only correction of sales tax on invoice R-204; nothing else about it is in dispute. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the subscription period in question for issue R, and the merchant has not issued any refund for it. The timeline logs for both issues were reviewed to determine which issue was most recently updated. The timeline for billing issue T shows its most recent update logged at 4:15 PM on June 9. The timeline for billing issue R shows its most recent update logged at 6:40 PM on June 9."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all original-state facts (invoice R-204, no tax/promotion/plan-change on R, no void, single settled payment, no refund) alongside the unchanged question, preserving governing policy and entity/path bindings; only the timestamp for issue T changes (4:15 PM to 1:05 PM) while R stays at 2:40 PM, which is a coherent single-fact counterfactual with no contradictions; the two evidence spans are plain factual timeline statements, not policy text; and neither context contains gold answers, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note on Mira Chen's billing disputes: Two verified billing issues remain unresolved and these are the only two currently open. Issue T concerns only correction of $4 sales tax on invoice R-204; no other matters are disputed under T. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change. The invoice tied to issue R was not later voided, only one settled payment covers that subscription period (fewer than two), and the merchant has not issued any refund for issue R. Billing issue T's timeline shows its latest update logged at 4:15 PM on March 9, 2024. Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024. These two timestamps differ, establishing which issue's timeline was most recently updated.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its latest update logged at 4:15 PM on March 9, 2024."}, {"path": [], "text": "Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its latest update logged at 4:15 PM on March 9, 2024.", "negative_left": "Billing issue T's timeline shows its latest update logged at 1:05 PM on March 9, 2024.", "negative_right": "Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024.", "right": "Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-004", "id": "fast-43-diverse-014-004-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note on Mira Chen's billing disputes: Two verified billing issues remain unresolved and these are the only two currently open. Issue T concerns only correction of $4 sales tax on invoice R-204; no other matters are disputed under T. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change. The invoice tied to issue R was not later voided, only one settled payment covers that subscription period (fewer than two), and the merchant has not issued any refund for issue R. Billing issue T's timeline shows its latest update logged at 4:15 PM on March 9, 2024. Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024. These two timestamps differ, establishing which issue's timeline was most recently updated."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all original-state facts (invoice R-204, no tax/promotion/plan-change on R, no void, single settled payment, no refund) alongside the unchanged question, preserving governing policy and entity/path bindings; only the timestamp for issue T changes (4:15 PM to 1:05 PM) while R stays at 2:40 PM, which is a coherent single-fact counterfactual with no contradictions; the two evidence spans are plain factual timeline statements, not policy text; and neither context contains gold answers, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note on Mira Chen's billing disputes: Two verified billing issues remain unresolved and these are the only two currently open. Issue T concerns only correction of $4 sales tax on invoice R-204; no other matters are disputed under T. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change. The invoice tied to issue R was not later voided, only one settled payment covers that subscription period (fewer than two), and the merchant has not issued any refund for issue R. Billing issue T's timeline shows its latest update logged at 4:15 PM on March 9, 2024. Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024. These two timestamps differ, establishing which issue's timeline was most recently updated.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its latest update logged at 4:15 PM on March 9, 2024."}, {"path": [], "text": "Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its latest update logged at 4:15 PM on March 9, 2024.", "negative_left": "Billing issue T's timeline shows its latest update logged at 1:05 PM on March 9, 2024.", "negative_right": "Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024.", "right": "Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-004", "id": "fast-43-diverse-014-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note on Mira Chen's billing disputes: Two verified billing issues remain unresolved and these are the only two currently open. Issue T concerns only correction of $4 sales tax on invoice R-204; no other matters are disputed under T. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change. The invoice tied to issue R was not later voided, only one settled payment covers that subscription period (fewer than two), and the merchant has not issued any refund for issue R. Billing issue T's timeline shows its latest update logged at 1:05 PM on March 9, 2024. Billing issue R's timeline shows its latest update logged at 2:40 PM on March 9, 2024. These two timestamps differ, establishing which issue's timeline was most recently updated."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all five routing criteria without inventing exceptions, keep the same subscriber and issue entities, use two complete factual timeline sentences as evidence, and the counterfactual's date change (June 4→June 6) keeps the 'times differ' claim true while altering recency without contradicting other facts or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note for subscriber Mira Chen. Billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Billing issue R is verified and unresolved, and it concerns the base renewal charge, which was billed at the wrong standard price; issue R does not involve tax, a promotion, or a plan change. The invoice disputed in issue R was not later voided, fewer than two settled payments cover the disputed subscription period, and the merchant has not issued any refund for issue R. Issues T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4. Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 4. The two timeline-update times differ.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4."}, {"path": [], "text": "Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 4."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4.", "negative_left": "Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4.", "negative_right": "Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 6.", "right": "Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-006", "id": "fast-43-diverse-014-006-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note for subscriber Mira Chen. Billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Billing issue R is verified and unresolved, and it concerns the base renewal charge, which was billed at the wrong standard price; issue R does not involve tax, a promotion, or a plan change. The invoice disputed in issue R was not later voided, fewer than two settled payments cover the disputed subscription period, and the merchant has not issued any refund for issue R. Issues T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4. Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 4. The two timeline-update times differ."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all five routing criteria without inventing exceptions, keep the same subscriber and issue entities, use two complete factual timeline sentences as evidence, and the counterfactual's date change (June 4→June 6) keeps the 'times differ' claim true while altering recency without contradicting other facts or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note for subscriber Mira Chen. Billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Billing issue R is verified and unresolved, and it concerns the base renewal charge, which was billed at the wrong standard price; issue R does not involve tax, a promotion, or a plan change. The invoice disputed in issue R was not later voided, fewer than two settled payments cover the disputed subscription period, and the merchant has not issued any refund for issue R. Issues T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4. Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 4. The two timeline-update times differ.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4."}, {"path": [], "text": "Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 4."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4.", "negative_left": "Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4.", "negative_right": "Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 6.", "right": "Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-006", "id": "fast-43-diverse-014-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note for subscriber Mira Chen. Billing issue T is verified and unresolved, and it concerns only correction of sales tax on an invoice. Billing issue R is verified and unresolved, and it concerns the base renewal charge, which was billed at the wrong standard price; issue R does not involve tax, a promotion, or a plan change. The invoice disputed in issue R was not later voided, fewer than two settled payments cover the disputed subscription period, and the merchant has not issued any refund for issue R. Issues T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline shows its most recent update logged at 2:15 PM on June 4. Billing issue R's timeline shows its most recent update logged at 9:40 AM on June 6. The two timeline-update times differ."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy-relevant facts and only the counterfactual timestamp for issue T changes, flipping which issue is latest without contradicting other facts or leaking route names.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note for Mira Chen's billing case. Billing issue T is verified and unresolved; it concerns only correction of a $4 sales tax charge on invoice R-204. Billing issue R is verified and unresolved; it concerns a $44 base renewal charge billed at the wrong standard price. Billing issue R does not concern tax, does not concern a promotion, and does not concern a plan change. The invoice disputed in billing issue R was not later voided. Only one settled payment covers the subscription period disputed in billing issue R. The merchant has not issued a refund for billing issue R. Issues T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline shows its latest update logged at 2024-03-14 09:20 UTC. Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its latest update logged at 2024-03-14 09:20 UTC."}, {"path": [], "text": "Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its latest update logged at 2024-03-14 09:20 UTC.", "negative_left": "Billing issue T's timeline shows its latest update logged at 2024-03-14 07:20 UTC.", "negative_right": "Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC.", "right": "Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-007", "id": "fast-43-diverse-014-007-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note for Mira Chen's billing case. Billing issue T is verified and unresolved; it concerns only correction of a $4 sales tax charge on invoice R-204. Billing issue R is verified and unresolved; it concerns a $44 base renewal charge billed at the wrong standard price. Billing issue R does not concern tax, does not concern a promotion, and does not concern a plan change. The invoice disputed in billing issue R was not later voided. Only one settled payment covers the subscription period disputed in billing issue R. The merchant has not issued a refund for billing issue R. Issues T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline shows its latest update logged at 2024-03-14 09:20 UTC. Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy-relevant facts and only the counterfactual timestamp for issue T changes, flipping which issue is latest without contradicting other facts or leaking route names.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note for Mira Chen's billing case. Billing issue T is verified and unresolved; it concerns only correction of a $4 sales tax charge on invoice R-204. Billing issue R is verified and unresolved; it concerns a $44 base renewal charge billed at the wrong standard price. Billing issue R does not concern tax, does not concern a promotion, and does not concern a plan change. The invoice disputed in billing issue R was not later voided. Only one settled payment covers the subscription period disputed in billing issue R. The merchant has not issued a refund for billing issue R. Issues T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline shows its latest update logged at 2024-03-14 09:20 UTC. Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its latest update logged at 2024-03-14 09:20 UTC."}, {"path": [], "text": "Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its latest update logged at 2024-03-14 09:20 UTC.", "negative_left": "Billing issue T's timeline shows its latest update logged at 2024-03-14 07:20 UTC.", "negative_right": "Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC.", "right": "Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-007", "id": "fast-43-diverse-014-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note for Mira Chen's billing case. Billing issue T is verified and unresolved; it concerns only correction of a $4 sales tax charge on invoice R-204. Billing issue R is verified and unresolved; it concerns a $44 base renewal charge billed at the wrong standard price. Billing issue R does not concern tax, does not concern a promotion, and does not concern a plan change. The invoice disputed in billing issue R was not later voided. Only one settled payment covers the subscription period disputed in billing issue R. The merchant has not issued a refund for billing issue R. Issues T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline shows its latest update logged at 2024-03-14 07:20 UTC. Billing issue R's timeline shows its latest update logged at 2024-03-14 08:05 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve entity, issue definitions, and policy context while the counterfactual only shifts T's timestamp to 2024-06-10, flipping which issue is latest without contradicting other stated facts or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note on Mira Chen's account: Two billing issues remain open. Billing issue T is verified and unresolved, concerning only the correction of a $4 sales tax charge on invoice R-204. Billing issue R is also verified and unresolved, concerning a base renewal charge that was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the invoice was not later voided. Only one settled payment covers the disputed subscription period, and the merchant has not issued any refund for issue R. T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline log shows its most recent update timestamped 2024-06-14 09:20 UTC. Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline log shows its most recent update timestamped 2024-06-14 09:20 UTC."}, {"path": [], "text": "Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline log shows its most recent update timestamped 2024-06-14 09:20 UTC.", "negative_left": "Billing issue T's timeline log shows its most recent update timestamped 2024-06-10 09:20 UTC.", "negative_right": "Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC.", "right": "Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-008", "id": "fast-43-diverse-014-008-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note on Mira Chen's account: Two billing issues remain open. Billing issue T is verified and unresolved, concerning only the correction of a $4 sales tax charge on invoice R-204. Billing issue R is also verified and unresolved, concerning a base renewal charge that was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the invoice was not later voided. Only one settled payment covers the disputed subscription period, and the merchant has not issued any refund for issue R. T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline log shows its most recent update timestamped 2024-06-14 09:20 UTC. Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve entity, issue definitions, and policy context while the counterfactual only shifts T's timestamp to 2024-06-10, flipping which issue is latest without contradicting other stated facts or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note on Mira Chen's account: Two billing issues remain open. Billing issue T is verified and unresolved, concerning only the correction of a $4 sales tax charge on invoice R-204. Billing issue R is also verified and unresolved, concerning a base renewal charge that was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the invoice was not later voided. Only one settled payment covers the disputed subscription period, and the merchant has not issued any refund for issue R. T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline log shows its most recent update timestamped 2024-06-14 09:20 UTC. Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline log shows its most recent update timestamped 2024-06-14 09:20 UTC."}, {"path": [], "text": "Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline log shows its most recent update timestamped 2024-06-14 09:20 UTC.", "negative_left": "Billing issue T's timeline log shows its most recent update timestamped 2024-06-10 09:20 UTC.", "negative_right": "Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC.", "right": "Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-008", "id": "fast-43-diverse-014-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note on Mira Chen's account: Two billing issues remain open. Billing issue T is verified and unresolved, concerning only the correction of a $4 sales tax charge on invoice R-204. Billing issue R is also verified and unresolved, concerning a base renewal charge that was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the invoice was not later voided. Only one settled payment covers the disputed subscription period, and the merchant has not issued any refund for issue R. T and R are the only verified unresolved billing issues for Mira Chen. Billing issue T's timeline log shows its most recent update timestamped 2024-06-10 09:20 UTC. Billing issue R's timeline log shows its most recent update timestamped 2024-06-12 15:45 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing routing policy via the unchanged questions object and use natural terminology for exclusions without embedding rules or answers; the counterfactual only changes issue T's timestamp, remaining internally coherent and still bound to the same entity/issues; both evidence spans are complete factual sentences describing timeline timestamps.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case notes for subscriber Mira Chen: Billing issue T is verified and remains unresolved; it concerns only correction of sales tax on an invoice. Billing issue R is verified and remains unresolved; it concerns the base renewal charge, which was billed at the wrong standard price. Issue R does not involve tax, a promotion, or a plan change, and the disputed invoice for R was not later voided. Fewer than two settled payments cover the subscription period disputed in R, and the merchant has not issued any refund for R. These two issues, T and R, are the only verified unresolved billing issues on Mira Chen's account. The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 14:52 on March 3. The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3. The two timestamps are not equal.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 14:52 on March 3."}, {"path": [], "text": "The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 14:52 on March 3.", "negative_left": "The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 09:10 on March 3.", "negative_right": "The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3.", "right": "The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-009", "id": "fast-43-diverse-014-009-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case notes for subscriber Mira Chen: Billing issue T is verified and remains unresolved; it concerns only correction of sales tax on an invoice. Billing issue R is verified and remains unresolved; it concerns the base renewal charge, which was billed at the wrong standard price. Issue R does not involve tax, a promotion, or a plan change, and the disputed invoice for R was not later voided. Fewer than two settled payments cover the subscription period disputed in R, and the merchant has not issued any refund for R. These two issues, T and R, are the only verified unresolved billing issues on Mira Chen's account. The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 14:52 on March 3. The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3. The two timestamps are not equal."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing routing policy via the unchanged questions object and use natural terminology for exclusions without embedding rules or answers; the counterfactual only changes issue T's timestamp, remaining internally coherent and still bound to the same entity/issues; both evidence spans are complete factual sentences describing timeline timestamps.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case notes for subscriber Mira Chen: Billing issue T is verified and remains unresolved; it concerns only correction of sales tax on an invoice. Billing issue R is verified and remains unresolved; it concerns the base renewal charge, which was billed at the wrong standard price. Issue R does not involve tax, a promotion, or a plan change, and the disputed invoice for R was not later voided. Fewer than two settled payments cover the subscription period disputed in R, and the merchant has not issued any refund for R. These two issues, T and R, are the only verified unresolved billing issues on Mira Chen's account. The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 14:52 on March 3. The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3. The two timestamps are not equal.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 14:52 on March 3."}, {"path": [], "text": "The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 14:52 on March 3.", "negative_left": "The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 09:10 on March 3.", "negative_right": "The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3.", "right": "The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-009", "id": "fast-43-diverse-014-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case notes for subscriber Mira Chen: Billing issue T is verified and remains unresolved; it concerns only correction of sales tax on an invoice. Billing issue R is verified and remains unresolved; it concerns the base renewal charge, which was billed at the wrong standard price. Issue R does not involve tax, a promotion, or a plan change, and the disputed invoice for R was not later voided. Fewer than two settled payments cover the subscription period disputed in R, and the merchant has not issued any refund for R. These two issues, T and R, are the only verified unresolved billing issues on Mira Chen's account. The timeline for billing issue T logged by Mira Chen shows its most recent update stamped at 09:10 on March 3. The timeline for billing issue R logged by Mira Chen shows its most recent update stamped at 11:20 on March 3. The two timestamps are not equal."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same entities, criteria-relevant facts, and the two exact timeline sentences as evidence; the counterfactual only changes issue T's timestamp from 14:32 to 08:15, preserving distinct non-duplicate timestamps and altering which issue is latest without contradicting other stated facts or leaking route names/labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Two verified unresolved billing issues remain on file for subscriber Mira Chen: issue T and issue R. These two issues account for the entirety of her verified unresolved billing matters. Issue T concerns only correction of a $4 sales-tax charge on an invoice and nothing else. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, its invoice was not later voided, fewer than two settled payments cover the disputed subscription period, and no merchant refund has been issued for it. The timeline for billing issue T shows its most recent update logged at 14:32 on March 9. The timeline for billing issue R shows its most recent update logged at 10:05 on March 9. Support staff confirmed both timeline entries independently while reconciling the case file, noting the two entries carry distinct timestamps that must not be conflated.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T shows its most recent update logged at 14:32 on March 9."}, {"path": [], "text": "The timeline for billing issue R shows its most recent update logged at 10:05 on March 9."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T shows its most recent update logged at 14:32 on March 9.", "negative_left": "The timeline for billing issue T shows its most recent update logged at 08:15 on March 9.", "negative_right": "The timeline for billing issue R shows its most recent update logged at 10:05 on March 9.", "right": "The timeline for billing issue R shows its most recent update logged at 10:05 on March 9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-012", "id": "fast-43-diverse-014-012-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Two verified unresolved billing issues remain on file for subscriber Mira Chen: issue T and issue R. These two issues account for the entirety of her verified unresolved billing matters. Issue T concerns only correction of a $4 sales-tax charge on an invoice and nothing else. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, its invoice was not later voided, fewer than two settled payments cover the disputed subscription period, and no merchant refund has been issued for it. The timeline for billing issue T shows its most recent update logged at 14:32 on March 9. The timeline for billing issue R shows its most recent update logged at 10:05 on March 9. Support staff confirmed both timeline entries independently while reconciling the case file, noting the two entries carry distinct timestamps that must not be conflated."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same entities, criteria-relevant facts, and the two exact timeline sentences as evidence; the counterfactual only changes issue T's timestamp from 14:32 to 08:15, preserving distinct non-duplicate timestamps and altering which issue is latest without contradicting other stated facts or leaking route names/labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Two verified unresolved billing issues remain on file for subscriber Mira Chen: issue T and issue R. These two issues account for the entirety of her verified unresolved billing matters. Issue T concerns only correction of a $4 sales-tax charge on an invoice and nothing else. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, its invoice was not later voided, fewer than two settled payments cover the disputed subscription period, and no merchant refund has been issued for it. The timeline for billing issue T shows its most recent update logged at 14:32 on March 9. The timeline for billing issue R shows its most recent update logged at 10:05 on March 9. Support staff confirmed both timeline entries independently while reconciling the case file, noting the two entries carry distinct timestamps that must not be conflated.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T shows its most recent update logged at 14:32 on March 9."}, {"path": [], "text": "The timeline for billing issue R shows its most recent update logged at 10:05 on March 9."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T shows its most recent update logged at 14:32 on March 9.", "negative_left": "The timeline for billing issue T shows its most recent update logged at 08:15 on March 9.", "negative_right": "The timeline for billing issue R shows its most recent update logged at 10:05 on March 9.", "right": "The timeline for billing issue R shows its most recent update logged at 10:05 on March 9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-012", "id": "fast-43-diverse-014-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Two verified unresolved billing issues remain on file for subscriber Mira Chen: issue T and issue R. These two issues account for the entirety of her verified unresolved billing matters. Issue T concerns only correction of a $4 sales-tax charge on an invoice and nothing else. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, its invoice was not later voided, fewer than two settled payments cover the disputed subscription period, and no merchant refund has been issued for it. The timeline for billing issue T shows its most recent update logged at 08:15 on March 9. The timeline for billing issue R shows its most recent update logged at 10:05 on March 9. Support staff confirmed both timeline entries independently while reconciling the case file, noting the two entries carry distinct timestamps that must not be conflated."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same entity, issues T/R, and all policy-relevant exclusions (tax, promo, plan change, voided invoice) matching the unchanged question criteria; the counterfactual only shifts R's timestamp to reverse which issue is latest, which is coherent and doesn't create duplicate/contradictory facts; evidence spans are exactly the two required factual timestamp sentences; neither context states or hints at the final route label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note - Subscriber Mira Chen: Two verified billing issues remain unresolved. Issue T concerns only a request to correct a $4 sales tax charge on an invoice; no other aspect of that invoice is disputed. Issue R concerns the $44 base renewal charge, which was billed at the wrong standard price rather than the correct standard rate; it does not involve tax, a promotion, or a plan change, and the underlying invoice was never later voided. Only one settled payment (R-204) covers the disputed subscription period for issue R, so fewer than two settled payments exist for that period, and the merchant has not issued any refund connected to issue R. Every verified unresolved billing issue for Mira Chen is either issue T or issue R, and no other billing issue is open. Timeline records show: Billing issue T's timeline was last updated at 4:12 PM on March 9. Billing issue R's timeline was last updated at 2:47 PM on March 9. Support staff confirmed these two timestamps differ.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline was last updated at 4:12 PM on March 9."}, {"path": [], "text": "Billing issue R's timeline was last updated at 2:47 PM on March 9."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline was last updated at 4:12 PM on March 9.", "negative_left": "Billing issue T's timeline was last updated at 4:12 PM on March 9.", "negative_right": "Billing issue R's timeline was last updated at 6:30 PM on March 9.", "right": "Billing issue R's timeline was last updated at 2:47 PM on March 9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-013", "id": "fast-43-diverse-014-013-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note - Subscriber Mira Chen: Two verified billing issues remain unresolved. Issue T concerns only a request to correct a $4 sales tax charge on an invoice; no other aspect of that invoice is disputed. Issue R concerns the $44 base renewal charge, which was billed at the wrong standard price rather than the correct standard rate; it does not involve tax, a promotion, or a plan change, and the underlying invoice was never later voided. Only one settled payment (R-204) covers the disputed subscription period for issue R, so fewer than two settled payments exist for that period, and the merchant has not issued any refund connected to issue R. Every verified unresolved billing issue for Mira Chen is either issue T or issue R, and no other billing issue is open. Timeline records show: Billing issue T's timeline was last updated at 4:12 PM on March 9. Billing issue R's timeline was last updated at 2:47 PM on March 9. Support staff confirmed these two timestamps differ."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same entity, issues T/R, and all policy-relevant exclusions (tax, promo, plan change, voided invoice) matching the unchanged question criteria; the counterfactual only shifts R's timestamp to reverse which issue is latest, which is coherent and doesn't create duplicate/contradictory facts; evidence spans are exactly the two required factual timestamp sentences; neither context states or hints at the final route label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note - Subscriber Mira Chen: Two verified billing issues remain unresolved. Issue T concerns only a request to correct a $4 sales tax charge on an invoice; no other aspect of that invoice is disputed. Issue R concerns the $44 base renewal charge, which was billed at the wrong standard price rather than the correct standard rate; it does not involve tax, a promotion, or a plan change, and the underlying invoice was never later voided. Only one settled payment (R-204) covers the disputed subscription period for issue R, so fewer than two settled payments exist for that period, and the merchant has not issued any refund connected to issue R. Every verified unresolved billing issue for Mira Chen is either issue T or issue R, and no other billing issue is open. Timeline records show: Billing issue T's timeline was last updated at 4:12 PM on March 9. Billing issue R's timeline was last updated at 2:47 PM on March 9. Support staff confirmed these two timestamps differ.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline was last updated at 4:12 PM on March 9."}, {"path": [], "text": "Billing issue R's timeline was last updated at 2:47 PM on March 9."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline was last updated at 4:12 PM on March 9.", "negative_left": "Billing issue T's timeline was last updated at 4:12 PM on March 9.", "negative_right": "Billing issue R's timeline was last updated at 6:30 PM on March 9.", "right": "Billing issue R's timeline was last updated at 2:47 PM on March 9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-013", "id": "fast-43-diverse-014-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note - Subscriber Mira Chen: Two verified billing issues remain unresolved. Issue T concerns only a request to correct a $4 sales tax charge on an invoice; no other aspect of that invoice is disputed. Issue R concerns the $44 base renewal charge, which was billed at the wrong standard price rather than the correct standard rate; it does not involve tax, a promotion, or a plan change, and the underlying invoice was never later voided. Only one settled payment (R-204) covers the disputed subscription period for issue R, so fewer than two settled payments exist for that period, and the merchant has not issued any refund connected to issue R. Every verified unresolved billing issue for Mira Chen is either issue T or issue R, and no other billing issue is open. Timeline records show: Billing issue T's timeline was last updated at 4:12 PM on March 9. Billing issue R's timeline was last updated at 6:30 PM on March 9. Support staff confirmed these two timestamps differ."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy facts (verified tax-exemption, no credit issued, single settled payment, no refund) alongside the unchanged question rubric; the only change is Issue T's timestamp shifting from 03-11 to 03-09, which flips the 'latest' issue to R while keeping the two timestamps distinct as asserted, so it remains internally coherent; the two evidence spans are plain factual timeline statements, not policy text; and neither context contains any label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note - Mira Chen billing review.\\n\\nIssue T (verified, unresolved): Concerns only the correction of a $4 sales tax charge on invoice R-204, per the approved tax-exemption certificate. No credit or corrected invoice has been issued. Billing issue T's timeline shows its most recent update logged at 2024-03-11 09:42 UTC.\\n\\nIssue R (verified, unresolved): Concerns the base renewal charge, which was billed at the wrong standard price after the promotion expired. This dispute does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the subscription period, so the duplicate-payment threshold is not met. No merchant refund has been issued for this issue. Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC.\\n\\nTogether, T and R make up every verified unresolved billing issue currently open for Mira Chen, and their two logged update times are distinct.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its most recent update logged at 2024-03-11 09:42 UTC."}, {"path": [], "text": "Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its most recent update logged at 2024-03-11 09:42 UTC.", "negative_left": "Billing issue T's timeline shows its most recent update logged at 2024-03-09 09:42 UTC.", "negative_right": "Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC.", "right": "Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-014", "id": "fast-43-diverse-014-014-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note - Mira Chen billing review.\n\nIssue T (verified, unresolved): Concerns only the correction of a $4 sales tax charge on invoice R-204, per the approved tax-exemption certificate. No credit or corrected invoice has been issued. Billing issue T's timeline shows its most recent update logged at 2024-03-11 09:42 UTC.\n\nIssue R (verified, unresolved): Concerns the base renewal charge, which was billed at the wrong standard price after the promotion expired. This dispute does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the subscription period, so the duplicate-payment threshold is not met. No merchant refund has been issued for this issue. Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC.\n\nTogether, T and R make up every verified unresolved billing issue currently open for Mira Chen, and their two logged update times are distinct."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy facts (verified tax-exemption, no credit issued, single settled payment, no refund) alongside the unchanged question rubric; the only change is Issue T's timestamp shifting from 03-11 to 03-09, which flips the 'latest' issue to R while keeping the two timestamps distinct as asserted, so it remains internally coherent; the two evidence spans are plain factual timeline statements, not policy text; and neither context contains any label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note - Mira Chen billing review.\\n\\nIssue T (verified, unresolved): Concerns only the correction of a $4 sales tax charge on invoice R-204, per the approved tax-exemption certificate. No credit or corrected invoice has been issued. Billing issue T's timeline shows its most recent update logged at 2024-03-11 09:42 UTC.\\n\\nIssue R (verified, unresolved): Concerns the base renewal charge, which was billed at the wrong standard price after the promotion expired. This dispute does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the subscription period, so the duplicate-payment threshold is not met. No merchant refund has been issued for this issue. Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC.\\n\\nTogether, T and R make up every verified unresolved billing issue currently open for Mira Chen, and their two logged update times are distinct.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its most recent update logged at 2024-03-11 09:42 UTC."}, {"path": [], "text": "Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its most recent update logged at 2024-03-11 09:42 UTC.", "negative_left": "Billing issue T's timeline shows its most recent update logged at 2024-03-09 09:42 UTC.", "negative_right": "Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC.", "right": "Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-014", "id": "fast-43-diverse-014-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note - Mira Chen billing review.\n\nIssue T (verified, unresolved): Concerns only the correction of a $4 sales tax charge on invoice R-204, per the approved tax-exemption certificate. No credit or corrected invoice has been issued. Billing issue T's timeline shows its most recent update logged at 2024-03-09 09:42 UTC.\n\nIssue R (verified, unresolved): Concerns the base renewal charge, which was billed at the wrong standard price after the promotion expired. This dispute does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the subscription period, so the duplicate-payment threshold is not met. No merchant refund has been issued for this issue. Billing issue R's timeline shows its most recent update logged at 2024-03-10 16:05 UTC.\n\nTogether, T and R make up every verified unresolved billing issue currently open for Mira Chen, and their two logged update times are distinct."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the necessary factual exclusions (no promotion, no plan change, no voided invoice, no refund, single settled payment) that let the unchanged question criteria be applied without added rules; the counterfactual only shifts issue T's timestamp from 2024-03-14 to 2024-03-10, altering which issue is most recent while keeping both timestamps distinct and consistent with the rest of the note; the two focus evidence spans are complete factual timeline statements rather than policy text; and neither context contains a route label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note: Mira Chen has two verified unresolved billing issues on file, T and R. Issue T concerns only correction of a $4 sales tax charge on invoice R-204; no tax credit or corrected invoice has been issued. Issue R concerns the $44 base renewal charge, which was billed at the wrong standard price. Issue R does not involve tax, a promotion, or a plan change, and its invoice was not later voided. Only one settled payment covers the disputed subscription period for issue R, so fewer than two settled payments apply, and the merchant has not issued any refund for issue R. Every verified unresolved billing issue for Mira Chen is either issue T or issue R. The two issues' timeline logs record different most-recent-update timestamps. Billing issue T's timeline shows its most recent update logged at 2024-03-14 09:20 UTC. Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its most recent update logged at 2024-03-14 09:20 UTC."}, {"path": [], "text": "Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its most recent update logged at 2024-03-14 09:20 UTC.", "negative_left": "Billing issue T's timeline shows its most recent update logged at 2024-03-10 09:20 UTC.", "negative_right": "Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC.", "right": "Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-015", "id": "fast-43-diverse-014-015-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note: Mira Chen has two verified unresolved billing issues on file, T and R. Issue T concerns only correction of a $4 sales tax charge on invoice R-204; no tax credit or corrected invoice has been issued. Issue R concerns the $44 base renewal charge, which was billed at the wrong standard price. Issue R does not involve tax, a promotion, or a plan change, and its invoice was not later voided. Only one settled payment covers the disputed subscription period for issue R, so fewer than two settled payments apply, and the merchant has not issued any refund for issue R. Every verified unresolved billing issue for Mira Chen is either issue T or issue R. The two issues' timeline logs record different most-recent-update timestamps. Billing issue T's timeline shows its most recent update logged at 2024-03-14 09:20 UTC. Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the necessary factual exclusions (no promotion, no plan change, no voided invoice, no refund, single settled payment) that let the unchanged question criteria be applied without added rules; the counterfactual only shifts issue T's timestamp from 2024-03-14 to 2024-03-10, altering which issue is most recent while keeping both timestamps distinct and consistent with the rest of the note; the two focus evidence spans are complete factual timeline statements rather than policy text; and neither context contains a route label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note: Mira Chen has two verified unresolved billing issues on file, T and R. Issue T concerns only correction of a $4 sales tax charge on invoice R-204; no tax credit or corrected invoice has been issued. Issue R concerns the $44 base renewal charge, which was billed at the wrong standard price. Issue R does not involve tax, a promotion, or a plan change, and its invoice was not later voided. Only one settled payment covers the disputed subscription period for issue R, so fewer than two settled payments apply, and the merchant has not issued any refund for issue R. Every verified unresolved billing issue for Mira Chen is either issue T or issue R. The two issues' timeline logs record different most-recent-update timestamps. Billing issue T's timeline shows its most recent update logged at 2024-03-14 09:20 UTC. Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its most recent update logged at 2024-03-14 09:20 UTC."}, {"path": [], "text": "Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its most recent update logged at 2024-03-14 09:20 UTC.", "negative_left": "Billing issue T's timeline shows its most recent update logged at 2024-03-10 09:20 UTC.", "negative_right": "Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC.", "right": "Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-015", "id": "fast-43-diverse-014-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note: Mira Chen has two verified unresolved billing issues on file, T and R. Issue T concerns only correction of a $4 sales tax charge on invoice R-204; no tax credit or corrected invoice has been issued. Issue R concerns the $44 base renewal charge, which was billed at the wrong standard price. Issue R does not involve tax, a promotion, or a plan change, and its invoice was not later voided. Only one settled payment covers the disputed subscription period for issue R, so fewer than two settled payments apply, and the merchant has not issued any refund for issue R. Every verified unresolved billing issue for Mira Chen is either issue T or issue R. The two issues' timeline logs record different most-recent-update timestamps. Billing issue T's timeline shows its most recent update logged at 2024-03-10 09:20 UTC. Billing issue R's timeline shows its most recent update logged at 2024-03-12 16:05 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical governing policy facts (issue T is pure tax correction, issue R is a wrong-price renewal dispute with no tax/promo/void) and the same entities/paths from the original question, with only the timestamp of issue T's timeline changed to alter which issue is latest, which is a permissible observation change; the two evidence spans are exact factual sentences about timeline updates, not policy text; the counterfactual remains internally consistent with no duplicate or contradictory measurements; and neither context includes a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note: Mira Chen has two verified, unresolved billing issues, T and R, and these are the only verified unresolved issues on her account. Issue T concerns only correction of a $4 sales tax charge on invoice R-204. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the disputed subscription period, and the merchant has not issued any refund for issue R. The timeline for billing issue T shows its most recent update logged at 2024-03-14 09:22 UTC. The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T shows its most recent update logged at 2024-03-14 09:22 UTC."}, {"path": [], "text": "The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T shows its most recent update logged at 2024-03-14 09:22 UTC.", "negative_left": "The timeline for billing issue T shows its most recent update logged at 2024-03-09 09:22 UTC.", "negative_right": "The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC.", "right": "The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-016", "id": "fast-43-diverse-014-016-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note: Mira Chen has two verified, unresolved billing issues, T and R, and these are the only verified unresolved issues on her account. Issue T concerns only correction of a $4 sales tax charge on invoice R-204. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the disputed subscription period, and the merchant has not issued any refund for issue R. The timeline for billing issue T shows its most recent update logged at 2024-03-14 09:22 UTC. The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical governing policy facts (issue T is pure tax correction, issue R is a wrong-price renewal dispute with no tax/promo/void) and the same entities/paths from the original question, with only the timestamp of issue T's timeline changed to alter which issue is latest, which is a permissible observation change; the two evidence spans are exact factual sentences about timeline updates, not policy text; the counterfactual remains internally consistent with no duplicate or contradictory measurements; and neither context includes a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note: Mira Chen has two verified, unresolved billing issues, T and R, and these are the only verified unresolved issues on her account. Issue T concerns only correction of a $4 sales tax charge on invoice R-204. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the disputed subscription period, and the merchant has not issued any refund for issue R. The timeline for billing issue T shows its most recent update logged at 2024-03-14 09:22 UTC. The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T shows its most recent update logged at 2024-03-14 09:22 UTC."}, {"path": [], "text": "The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T shows its most recent update logged at 2024-03-14 09:22 UTC.", "negative_left": "The timeline for billing issue T shows its most recent update logged at 2024-03-09 09:22 UTC.", "negative_right": "The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC.", "right": "The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-016", "id": "fast-43-diverse-014-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note: Mira Chen has two verified, unresolved billing issues, T and R, and these are the only verified unresolved issues on her account. Issue T concerns only correction of a $4 sales tax charge on invoice R-204. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the disputed subscription period, and the merchant has not issued any refund for issue R. The timeline for billing issue T shows its most recent update logged at 2024-03-09 09:22 UTC. The timeline for billing issue R shows its most recent update logged at 2024-03-12 16:05 UTC."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all original policy criteria via the untouched questions object, keep the same entities and timeline structure, and only alter issue R's timestamp (09:10\\u219219:47) which flips which issue is most recent without contradicting other stated facts; the two focus evidence spans are complete factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Subscriber Mira Chen has two verified, unresolved billing issues on file, referred to as issue T and issue R; every verified unresolved billing issue for Mira Chen is one of these two. Issue T concerns only the correction of a $4 sales tax charge on invoice R-204, nothing else. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the disputed subscription period for issue R, so fewer than two settled payments apply, and the merchant has not issued any refund for issue R. The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9. The timeline for billing issue R shows its most recent update logged at 09:10 UTC on March 9. Both timestamps fall on the same date but differ in time of day.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9."}, {"path": [], "text": "The timeline for billing issue R shows its most recent update logged at 09:10 UTC on March 9."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9.", "negative_left": "The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9.", "negative_right": "The timeline for billing issue R shows its most recent update logged at 19:47 UTC on March 9.", "right": "The timeline for billing issue R shows its most recent update logged at 09:10 UTC on March 9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-018", "id": "fast-43-diverse-014-018-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Subscriber Mira Chen has two verified, unresolved billing issues on file, referred to as issue T and issue R; every verified unresolved billing issue for Mira Chen is one of these two. Issue T concerns only the correction of a $4 sales tax charge on invoice R-204, nothing else. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the disputed subscription period for issue R, so fewer than two settled payments apply, and the merchant has not issued any refund for issue R. The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9. The timeline for billing issue R shows its most recent update logged at 09:10 UTC on March 9. Both timestamps fall on the same date but differ in time of day."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all original policy criteria via the untouched questions object, keep the same entities and timeline structure, and only alter issue R's timestamp (09:10\\u219219:47) which flips which issue is most recent without contradicting other stated facts; the two focus evidence spans are complete factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Subscriber Mira Chen has two verified, unresolved billing issues on file, referred to as issue T and issue R; every verified unresolved billing issue for Mira Chen is one of these two. Issue T concerns only the correction of a $4 sales tax charge on invoice R-204, nothing else. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the disputed subscription period for issue R, so fewer than two settled payments apply, and the merchant has not issued any refund for issue R. The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9. The timeline for billing issue R shows its most recent update logged at 09:10 UTC on March 9. Both timestamps fall on the same date but differ in time of day.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9."}, {"path": [], "text": "The timeline for billing issue R shows its most recent update logged at 09:10 UTC on March 9."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9.", "negative_left": "The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9.", "negative_right": "The timeline for billing issue R shows its most recent update logged at 19:47 UTC on March 9.", "right": "The timeline for billing issue R shows its most recent update logged at 09:10 UTC on March 9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-018", "id": "fast-43-diverse-014-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Subscriber Mira Chen has two verified, unresolved billing issues on file, referred to as issue T and issue R; every verified unresolved billing issue for Mira Chen is one of these two. Issue T concerns only the correction of a $4 sales tax charge on invoice R-204, nothing else. Issue R concerns the base renewal charge, which was billed at the wrong standard price; it does not involve tax, a promotion, or a plan change, and the disputed invoice was not later voided. Only one settled payment covers the disputed subscription period for issue R, so fewer than two settled payments apply, and the merchant has not issued any refund for issue R. The timeline for billing issue T shows its most recent update logged at 14:32 UTC on March 9. The timeline for billing issue R shows its most recent update logged at 19:47 UTC on March 9. Both timestamps fall on the same date but differ in time of day."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all original facts and the unchanged question criteria, only the counterfactual shifts issue T's timestamp to 06:05 making issue R the latest, which is a coherent factual change without contradicting other stated facts or embedding any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note update for subscriber Mira Chen. Two billing issues remain flagged as verified and unresolved in the account timeline: issue T, concerning correction of a $4 sales tax charge on invoice R-204, and issue R, concerning the base renewal charge itself. Support confirmed that no other verified unresolved billing issues exist for Mira Chen beyond these two. Analyst records show the $44 renewal on R-204 was billed at the wrong standard price, was not related to tax, a promotion, or a plan change, and the invoice was never later voided. Only one settled payment covers this subscription period, and the merchant has issued no refund. Timeline audit logs record distinct update timestamps for each issue: Billing issue T's timeline shows its most recent update logged at 14:32 on March 9, 2024. Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its most recent update logged at 14:32 on March 9, 2024."}, {"path": [], "text": "Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its most recent update logged at 14:32 on March 9, 2024.", "negative_left": "Billing issue T's timeline shows its most recent update logged at 06:05 on March 9, 2024.", "negative_right": "Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024.", "right": "Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-023", "id": "fast-43-diverse-014-023-base", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note update for subscriber Mira Chen. Two billing issues remain flagged as verified and unresolved in the account timeline: issue T, concerning correction of a $4 sales tax charge on invoice R-204, and issue R, concerning the base renewal charge itself. Support confirmed that no other verified unresolved billing issues exist for Mira Chen beyond these two. Analyst records show the $44 renewal on R-204 was billed at the wrong standard price, was not related to tax, a promotion, or a plan change, and the invoice was never later voided. Only one settled payment covers this subscription period, and the merchant has issued no refund. Timeline audit logs record distinct update timestamps for each issue: Billing issue T's timeline shows its most recent update logged at 14:32 on March 9, 2024. Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all original facts and the unchanged question criteria, only the counterfactual shifts issue T's timestamp to 06:05 making issue R the latest, which is a coherent factual change without contradicting other stated facts or embedding any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "full_context_fact_states": {"base": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "supported", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "counterfactual": {"R_base_renewal": "supported", "R_invoice_not_later_voided": "supported", "R_no_merchant_refund": "supported", "R_not_plan_change": "supported", "R_not_promotion": "supported", "R_not_tax": "supported", "R_settled_payment_count_below_two": "supported", "R_unresolved": "supported", "R_verified": "supported", "R_wrong_standard_price": "supported", "T_tax_only": "supported", "T_unresolved": "supported", "T_update_later": "refuted", "T_verified": "supported", "issue_scope": "supported", "update_times_unequal": "supported"}, "remove_left": {"T_update_later": "unknown"}, "remove_right": {"T_update_later": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"T_update_later": "unknown"}, "negative_pair": {"T_update_later": "refuted"}, "negative_sentence": {"T_update_later": "unknown"}, "positive_pair": {"T_update_later": "supported"}, "right": {"T_update_later": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified issue-scope atom is still one set-inclusion fact rather than a bundle or final classification. The focus atom compares two timeline-update times and is factual, not a policy conclusion. Base and counter assignments can both be realized by changing only which issue has the later update time; the supported unequal-times atom remains consistent in both. Empty policy_evidence is correct because all governing routing rules, exclusions, priority instructions, and scope are already retained in original_input.questions, while the original state contributes only case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes T the latest verified unresolved issue: T and R exhaust that issue set, both are verified and unresolved, and T is later than R. A tax-only invoice correction falls outside the five substantive routes and is explicitly covered by none_of_above.", "rule_index": 0, "sound": true}, {"reason": "Given unequal update times and the entailed negation of T being later, R is later than T and therefore is the latest verified unresolved issue under the stated scope. R satisfies the base-renewal/wrong-standard-price criterion, while the explicit tax, promotion, plan-change, and later-voided-invoice exclusions apply. Fewer than two settled payments and no merchant-issued refund also exclude the duplicate-payment and refund-delay routes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "T_verified", "statement": "Billing issue T for subscriber Mira Chen is verified."}, {"id": "T_unresolved", "statement": "Billing issue T for subscriber Mira Chen is unresolved."}, {"id": "T_tax_only", "statement": "Billing issue T concerns only correction of sales tax on an invoice."}, {"id": "R_verified", "statement": "Billing issue R for subscriber Mira Chen is verified."}, {"id": "R_unresolved", "statement": "Billing issue R for subscriber Mira Chen is unresolved."}, {"id": "R_base_renewal", "statement": "Billing issue R concerns the base renewal charge."}, {"id": "R_wrong_standard_price", "statement": "The base renewal charge disputed in billing issue R was billed at the wrong standard price."}, {"id": "R_not_tax", "statement": "Billing issue R does not concern tax."}, {"id": "R_not_promotion", "statement": "Billing issue R does not concern a promotion."}, {"id": "R_not_plan_change", "statement": "Billing issue R does not concern a plan change."}, {"id": "R_invoice_not_later_voided", "statement": "The invoice disputed in billing issue R was not later voided."}, {"id": "R_settled_payment_count_below_two", "statement": "The number of settled payments covering the subscription period disputed in billing issue R is less than two."}, {"id": "R_no_merchant_refund", "statement": "The merchant has not issued a refund for billing issue R."}, {"id": "issue_scope", "statement": "Every verified unresolved billing issue for Mira Chen is either billing issue T or billing issue R."}, {"id": "update_times_unequal", "statement": "The latest timeline-update time for billing issue T differs from the latest timeline-update time for billing issue R."}, {"id": "T_update_later", "statement": "The latest timeline-update time for billing issue T is later than the latest timeline-update time for billing issue R."}], "base_state_json": "\"Case note update for subscriber Mira Chen. Two billing issues remain flagged as verified and unresolved in the account timeline: issue T, concerning correction of a $4 sales tax charge on invoice R-204, and issue R, concerning the base renewal charge itself. Support confirmed that no other verified unresolved billing issues exist for Mira Chen beyond these two. Analyst records show the $44 renewal on R-204 was billed at the wrong standard price, was not related to tax, a promotion, or a plan change, and the invoice was never later voided. Only one settled payment covers this subscription period, and the merchant has issued no refund. Timeline audit logs record distinct update timestamps for each issue: Billing issue T's timeline shows its most recent update logged at 14:32 on March 9, 2024. Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024.\"", "base_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}], "counter_states": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}], "focus_atom": "T_update_later", "focus_evidence": [{"path": [], "text": "Billing issue T's timeline shows its most recent update logged at 14:32 on March 9, 2024."}, {"path": [], "text": "Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024."}], "policy_evidence": [], "rules": [{"justification": "T and R exhaust the verified unresolved issues, and T has the later timeline update. T is therefore the latest verified unresolved issue. Because T concerns only an invoice sales-tax correction, it falls outside all five listed routes.", "target": "none_of_above", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "T_tax_only", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "T_update_later", "state": "supported"}]}, {"justification": "Because the two update times differ and T is not later than R, R is the latest verified unresolved issue. R concerns a base renewal charge billed at the wrong standard price and satisfies the tax, promotion, plan-change, and later-voided-invoice exclusions. Fewer than two settled payments exclude the duplicate-payment route, and the absence of a merchant-issued refund excludes the refund-delay route.", "target": "renewal_charge_dispute", "when": [{"atom_id": "T_verified", "state": "supported"}, {"atom_id": "T_unresolved", "state": "supported"}, {"atom_id": "R_verified", "state": "supported"}, {"atom_id": "R_unresolved", "state": "supported"}, {"atom_id": "R_base_renewal", "state": "supported"}, {"atom_id": "R_wrong_standard_price", "state": "supported"}, {"atom_id": "R_not_tax", "state": "supported"}, {"atom_id": "R_not_promotion", "state": "supported"}, {"atom_id": "R_not_plan_change", "state": "supported"}, {"atom_id": "R_invoice_not_later_voided", "state": "supported"}, {"atom_id": "R_settled_payment_count_below_two", "state": "supported"}, {"atom_id": "R_no_merchant_refund", "state": "supported"}, {"atom_id": "issue_scope", "state": "supported"}, {"atom_id": "update_times_unequal", "state": "supported"}, {"atom_id": "T_update_later", "state": "refuted"}]}]}, "verified_pair": {"left": "Billing issue T's timeline shows its most recent update logged at 14:32 on March 9, 2024.", "negative_left": "Billing issue T's timeline shows its most recent update logged at 06:05 on March 9, 2024.", "negative_right": "Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024.", "right": "Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-014-023", "id": "fast-43-diverse-014-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"duplicate_subscription_payment": "Use only when two or more settled payments cover the same subscription period; pending, released, or voided authorizations do not qualify.", "none_of_above": "Use when the latest verified unresolved issue falls outside all five listed routes, including an invoice tax correction.", "plan_change_mismatch": "Use only when the current unresolved issue is a requested plan upgrade or downgrade that was not applied correctly.", "promotional_credit_missing": "Use only when a verified valid promotion remains unapplied to the current invoice.", "refund_posting_delay": "Use only when the merchant has issued a refund and it remains unposted after the stated processing period.", "renewal_charge_dispute": "Use only when the current unresolved issue is the base renewal charge being unauthorized or billed at the wrong standard price, excluding tax, promotions, plan changes, and later-voided invoices."}, "instructions": "Route the case by its latest verified unresolved billing issue. Earlier issues superseded by later timeline updates must not determine the route. Select exactly one option whose rubric fits the evidence.", "type": "choice"}}, "state": "Case note update for subscriber Mira Chen. Two billing issues remain flagged as verified and unresolved in the account timeline: issue T, concerning correction of a $4 sales tax charge on invoice R-204, and issue R, concerning the base renewal charge itself. Support confirmed that no other verified unresolved billing issues exist for Mira Chen beyond these two. Analyst records show the $44 renewal on R-204 was billed at the wrong standard price, was not related to tax, a promotion, or a plan change, and the invoice was never later voided. Only one settled payment covers this subscription period, and the merchant has issued no refund. Timeline audit logs record distinct update timestamps for each issue: Billing issue T's timeline shows its most recent update logged at 06:05 on March 9, 2024. Billing issue R's timeline shows its most recent update logged at 09:17 on March 9, 2024."}, "method": "c2d", "provenance": {"source_id": "diverse-014", "source_is_synthetic": true, "source_sha256": "885fe372c64bb1b20ed2463cb8502490a0643c2816b35605a2e423c1e96263a9", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "renewal_charge_dispute"}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the exact policy exception language and question bindings (S-204, R-841, dates, amounts) while only the tag of the single logged event (Premium vs Basic feature use) differs, which is a coherent single-fact counterfactual; the two focus evidence spans are complete factual sentences with no policy definitions, and neither context reveals a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price charged was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged event is tagged as usage of a Premium-tier feature. Agent Leo Grant forwarded the case file, including the timestamped log excerpt, to analyst Priya Nair for review. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Priya noted the invoice amount and requested plan change for her write-up.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged event is tagged as usage of a Premium-tier feature."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged event is tagged as usage of a Basic-tier feature.", "right": "That logged event is tagged as usage of a Premium-tier feature."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-051", "id": "fast-43-diverse-017-051-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price charged was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged event is tagged as usage of a Premium-tier feature. Agent Leo Grant forwarded the case file, including the timestamped log excerpt, to analyst Priya Nair for review. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Priya noted the invoice amount and requested plan change for her write-up."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the exact policy exception language and question bindings (S-204, R-841, dates, amounts) while only the tag of the single logged event (Premium vs Basic feature use) differs, which is a coherent single-fact counterfactual; the two focus evidence spans are complete factual sentences with no policy definitions, and neither context reveals a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price charged was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged event is tagged as usage of a Premium-tier feature. Agent Leo Grant forwarded the case file, including the timestamped log excerpt, to analyst Priya Nair for review. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Priya noted the invoice amount and requested plan change for her write-up.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged event is tagged as usage of a Premium-tier feature."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged event is tagged as usage of a Basic-tier feature.", "right": "That logged event is tagged as usage of a Premium-tier feature."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-051", "id": "fast-43-diverse-017-051-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price charged was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged event is tagged as usage of a Basic-tier feature. Agent Leo Grant forwarded the case file, including the timestamped log excerpt, to analyst Priya Nair for review. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Priya noted the invoice amount and requested plan change for her write-up."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the exact review-language policy and all invoice/subscription bindings, the two focus sentences are factual case-log statements rather than policy or answer text, and the counterfactual coherently swaps the log entry's tag from a premium feature-use event to a routine login event without contradicting other stated facts or leaking the decision outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for subscription S-204, targeting the Basic annual plan. The request was submitted within 72 hours of the September 14 charge. The Premium annual price charged was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Analyst Priya Nair forwarded the ticket for a final decision.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged entry is tagged as a Premium feature-use event."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged entry is tagged as a routine login event, not a feature-use event.", "right": "That logged entry is tagged as a Premium feature-use event."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-055", "id": "fast-43-diverse-017-055-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for subscription S-204, targeting the Basic annual plan. The request was submitted within 72 hours of the September 14 charge. The Premium annual price charged was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Analyst Priya Nair forwarded the ticket for a final decision."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the exact review-language policy and all invoice/subscription bindings, the two focus sentences are factual case-log statements rather than policy or answer text, and the counterfactual coherently swaps the log entry's tag from a premium feature-use event to a routine login event without contradicting other stated facts or leaking the decision outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for subscription S-204, targeting the Basic annual plan. The request was submitted within 72 hours of the September 14 charge. The Premium annual price charged was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Analyst Priya Nair forwarded the ticket for a final decision.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged entry is tagged as a Premium feature-use event."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged entry is tagged as a routine login event, not a feature-use event.", "right": "That logged entry is tagged as a Premium feature-use event."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-055", "id": "fast-43-diverse-017-055-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for subscription S-204, targeting the Basic annual plan. The request was submitted within 72 hours of the September 14 charge. The Premium annual price charged was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a routine login event, not a feature-use event. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Analyst Priya Nair forwarded the ticket for a final decision."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy exception language, invoice/subscription/date bindings, and the counterfactual only flips the feature-access tag from Premium to Basic without contradicting other facts or leaking a computed answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price was $120.00, while the applicable Basic annual price was $72.00. At the time of the request, 363 of 365 service days remained, and the requested correction covered exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That feature-access entry is tagged as a Premium-tier feature invocation. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the log excerpt and the downgrade request to Priya Nair for adjustment review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That feature-access entry is tagged as a Premium-tier feature invocation."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That feature-access entry is tagged as a Basic-tier feature invocation.", "right": "That feature-access entry is tagged as a Premium-tier feature invocation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-056", "id": "fast-43-diverse-017-056-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price was $120.00, while the applicable Basic annual price was $72.00. At the time of the request, 363 of 365 service days remained, and the requested correction covered exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That feature-access entry is tagged as a Premium-tier feature invocation. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the log excerpt and the downgrade request to Priya Nair for adjustment review."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy exception language, invoice/subscription/date bindings, and the counterfactual only flips the feature-access tag from Premium to Basic without contradicting other facts or leaking a computed answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price was $120.00, while the applicable Basic annual price was $72.00. At the time of the request, 363 of 365 service days remained, and the requested correction covered exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That feature-access entry is tagged as a Premium-tier feature invocation. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the log excerpt and the downgrade request to Priya Nair for adjustment review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That feature-access entry is tagged as a Premium-tier feature invocation."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That feature-access entry is tagged as a Basic-tier feature invocation.", "right": "That feature-access entry is tagged as a Premium-tier feature invocation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-056", "id": "fast-43-diverse-017-056-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request arrived no more than 72 hours after the September 14 charge. The Premium annual price was $120.00, while the applicable Basic annual price was $72.00. At the time of the request, 363 of 365 service days remained, and the requested correction covered exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one feature-access entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That feature-access entry is tagged as a Basic-tier feature invocation. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the log excerpt and the downgrade request to Priya Nair for adjustment review."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy text and question bindings (S-204, R-841, dates, prices) while only varying the logged event's tier tag as a case observation; the counterfactual's Basic-tier tag remains internally consistent with 'no premium use' language, and neither context reveals a decision level or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen filed a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request came in no more than 72 hours after the September 14 charge. The Premium annual price billed was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged event is tagged as a Premium-tier feature access. Policy on file states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Analyst Priya Nair logged the ticket for further review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged event is tagged as a Premium-tier feature access."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged event is tagged as a Basic-tier feature access.", "right": "That logged event is tagged as a Premium-tier feature access."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-057", "id": "fast-43-diverse-017-057-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen filed a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request came in no more than 72 hours after the September 14 charge. The Premium annual price billed was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged event is tagged as a Premium-tier feature access. Policy on file states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Analyst Priya Nair logged the ticket for further review."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy text and question bindings (S-204, R-841, dates, prices) while only varying the logged event's tier tag as a case observation; the counterfactual's Basic-tier tag remains internally consistent with 'no premium use' language, and neither context reveals a decision level or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen filed a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request came in no more than 72 hours after the September 14 charge. The Premium annual price billed was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged event is tagged as a Premium-tier feature access. Policy on file states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Analyst Priya Nair logged the ticket for further review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged event is tagged as a Premium-tier feature access."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged event is tagged as a Basic-tier feature access.", "right": "That logged event is tagged as a Premium-tier feature access."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-057", "id": "fast-43-diverse-017-057-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen filed a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request came in no more than 72 hours after the September 14 charge. The Premium annual price billed was $120.00, while the applicable Basic annual price is $72.00. At the time of the request, 363 of 365 service days remained, and the correction sought covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one event timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged event is tagged as a Basic-tier feature access. Policy on file states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Analyst Priya Nair logged the ticket for further review."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the exception policy and all question-bound identifiers (R-841, S-204, dates, prices), the two evidence sentences are factual and non-instructional, the counterfactual coherently swaps only the event tag without duplicating or contradicting other measurements, and neither context reveals a gold label or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request was submitted no more than 72 hours after the September 14 renewal charge. The Premium annual price on the invoice was $120.00, while the applicable Basic annual price is $72.00. At the time of the downgrade request, 363 of 365 service days remained in the annual term, and the correction Maya seeks covers exactly one paid invoice, R-841. Reviewing the account's activity trail, analyst Priya Nair found: The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. Policy on file states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the file for a decision.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged entry is tagged as a Premium feature-use event."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged entry is tagged as a routine login event, not a feature-use event.", "right": "That logged entry is tagged as a Premium feature-use event."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-058", "id": "fast-43-diverse-017-058-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request was submitted no more than 72 hours after the September 14 renewal charge. The Premium annual price on the invoice was $120.00, while the applicable Basic annual price is $72.00. At the time of the downgrade request, 363 of 365 service days remained in the annual term, and the correction Maya seeks covers exactly one paid invoice, R-841. Reviewing the account's activity trail, analyst Priya Nair found: The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. Policy on file states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the file for a decision."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the exception policy and all question-bound identifiers (R-841, S-204, dates, prices), the two evidence sentences are factual and non-instructional, the counterfactual coherently swaps only the event tag without duplicating or contradicting other measurements, and neither context reveals a gold label or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request was submitted no more than 72 hours after the September 14 renewal charge. The Premium annual price on the invoice was $120.00, while the applicable Basic annual price is $72.00. At the time of the downgrade request, 363 of 365 service days remained in the annual term, and the correction Maya seeks covers exactly one paid invoice, R-841. Reviewing the account's activity trail, analyst Priya Nair found: The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. Policy on file states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the file for a decision.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged entry is tagged as a Premium feature-use event."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged entry is tagged as a routine login event, not a feature-use event.", "right": "That logged entry is tagged as a Premium feature-use event."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-058", "id": "fast-43-diverse-017-058-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. The request was submitted no more than 72 hours after the September 14 renewal charge. The Premium annual price on the invoice was $120.00, while the applicable Basic annual price is $72.00. At the time of the downgrade request, 363 of 365 service days remained in the annual term, and the correction Maya seeks covers exactly one paid invoice, R-841. Reviewing the account's activity trail, analyst Priya Nair found: The system activity log for subscription S-204 records exactly one entry timestamped between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a routine login event, not a feature-use event. Policy on file states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the file for a decision."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the exact governing review language and quantities; evidence spans are two complete factual sentences about the logged usage event and its tier tag; the counterfactual only flips 'Premium-tier' to 'Basic-tier', remaining internally coherent and altering eligibility without contradiction; no gold answer, rule table, or instruction is embedded; question entity, path, and time bindings (S-204, R-841, Sept 14/16, 72-hour window) are unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. This request arrived no more than 72 hours after the September 14 charge. The Premium annual price on the invoice was $120.00, while the applicable Basic annual price for the downgrade is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction Maya sought covers exactly one paid invoice, R-841. The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged usage event is tagged as a Premium-tier feature access. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the ticket for adjustment review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged usage event is tagged as a Premium-tier feature access."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged usage event is tagged as a Basic-tier feature access.", "right": "That logged usage event is tagged as a Premium-tier feature access."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-061", "id": "fast-43-diverse-017-061-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. This request arrived no more than 72 hours after the September 14 charge. The Premium annual price on the invoice was $120.00, while the applicable Basic annual price for the downgrade is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction Maya sought covers exactly one paid invoice, R-841. The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged usage event is tagged as a Premium-tier feature access. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the ticket for adjustment review."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the exact governing review language and quantities; evidence spans are two complete factual sentences about the logged usage event and its tier tag; the counterfactual only flips 'Premium-tier' to 'Basic-tier', remaining internally coherent and altering eligibility without contradiction; no gold answer, rule table, or instruction is embedded; question entity, path, and time bindings (S-204, R-841, Sept 14/16, 72-hour window) are unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. This request arrived no more than 72 hours after the September 14 charge. The Premium annual price on the invoice was $120.00, while the applicable Basic annual price for the downgrade is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction Maya sought covers exactly one paid invoice, R-841. The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged usage event is tagged as a Premium-tier feature access. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the ticket for adjustment review.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged usage event is tagged as a Premium-tier feature access."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged usage event is tagged as a Basic-tier feature access.", "right": "That logged usage event is tagged as a Premium-tier feature access."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-061", "id": "fast-43-diverse-017-061-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for S-204, targeting the Basic annual plan. This request arrived no more than 72 hours after the September 14 charge. The Premium annual price on the invoice was $120.00, while the applicable Basic annual price for the downgrade is $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction Maya sought covers exactly one paid invoice, R-841. The activity log for subscription S-204 records exactly one usage event timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged usage event is tagged as a Basic-tier feature access. The applicable review language states: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. Billing forwarded the ticket for adjustment review."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy language, entity/time bindings, and two factual evidence sentences, differing only in the log entry's tag (Premium vs Basic feature-use), which is a coherent single-fact swap without contradicting other stated facts or leaking the final decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for subscription S-204, targeting the Basic annual plan. This request arrived no more than 72 hours after the September 14 renewal charge. The Premium annual price on invoice R-841 was $120.00, while the applicable Basic annual price for the downgrade was $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought under the request covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. Applicable review language: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. The account file was routed to Priya Nair for final determination.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged entry is tagged as a Premium feature-use event."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged entry is tagged as a Basic feature-use event.", "right": "That logged entry is tagged as a Premium feature-use event."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-062", "id": "fast-43-diverse-017-062-base", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for subscription S-204, targeting the Basic annual plan. This request arrived no more than 72 hours after the September 14 renewal charge. The Premium annual price on invoice R-841 was $120.00, while the applicable Basic annual price for the downgrade was $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought under the request covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. Applicable review language: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. The account file was routed to Priya Nair for final determination."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy language, entity/time bindings, and two factual evidence sentences, differing only in the log entry's tag (Premium vs Basic feature-use), which is a coherent single-fact swap without contradicting other stated facts or leaking the final decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a factual relationship rather than a final policy classification, and a7 is a factual feature-use proposition. The base and counter assignments differ only on a7 and are both realizable: Premium use may either occur or not occur while all other facts remain fixed. The policy evidence correctly preserves the governing renewal exception and prorated-calculation rule originating in the original state; decision thresholds and instructions remain automatically preserved in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a renewal charge and a Premium-use event during the relevant period. That defeats the stated no-premium-use requirement, so the sole stated renewal exception does not apply and the normally excluded renewal has an authorized amount of exactly $0, corresponding to target 0.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a qualifying renewal downgrade within 72 hours, entails no Premium use by refuting a7, and supplies the applicable prices and remaining term. The authorized amount is ($120.00 - $72.00) × 363/365 = $47.67 after rounding, which falls exclusively in the moderate-adjustment range. The exact-one-invoice condition also excludes the multiple-invoice very-high criterion, so target 2 is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Invoice R-841 is a renewal charge."}, {"id": "a2", "statement": "Invoice R-841 renewed subscription S-204 on the Premium annual plan."}, {"id": "a3", "statement": "Maya Chen's September 16 request is a downgrade request."}, {"id": "a4", "statement": "Maya Chen's September 16 request targets the Basic annual plan."}, {"id": "a5", "statement": "The subscription covered by Maya Chen's September 16 request is subscription S-204."}, {"id": "a6", "statement": "No more than 72 hours elapsed between the September 14 charge for invoice R-841 and Maya Chen's September 16 downgrade request."}, {"id": "a7", "statement": "At least one Premium feature-use event occurred on subscription S-204 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"id": "a8", "statement": "The Premium annual price charged on invoice R-841 is $120.00."}, {"id": "a9", "statement": "The applicable Basic annual price for Maya Chen's downgrade is $72.00."}, {"id": "a10", "statement": "The service period remaining at Maya Chen's downgrade request is 363 of 365 days."}, {"id": "a11", "statement": "The correction sought under Maya Chen's September 16 request covers exactly one paid invoice, R-841."}], "base_state_json": "\"Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for subscription S-204, targeting the Basic annual plan. This request arrived no more than 72 hours after the September 14 renewal charge. The Premium annual price on invoice R-841 was $120.00, while the applicable Basic annual price for the downgrade was $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought under the request covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Premium feature-use event. Applicable review language: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. The account file was routed to Priya Nair for final determination.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": [], "text": "The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request."}, {"path": [], "text": "That logged entry is tagged as a Premium feature-use event."}], "policy_evidence": [{"path": [], "text": "renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use."}, {"path": [], "text": "That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents."}], "rules": [{"justification": "Invoice R-841 is a renewal, and Premium use occurred on the renewed subscription before the downgrade request. The no-premium-use requirement for the renewal exception therefore fails, so the renewal remains outside refund scope and the authorized amount is exactly $0.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "The renewal downgrade was requested within 72 hours and no Premium feature-use event occurred in the relevant interval, so the exception applies. The authorized amount is only ($120.00 - $72.00) × 363/365 = $47.6712, rounded to $47.67. That amount is within $25.00 through $59.99 and covers exactly one paid invoice, so the score index is 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_left": "The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request.", "negative_right": "That logged entry is tagged as a Basic feature-use event.", "right": "That logged entry is tagged as a Premium feature-use event."}, "verifier_independent_model": false}, "family": "fast-43-diverse-017-062", "id": "fast-43-diverse-017-062-counterfactual", "input": {"questions": {"decision": {"criteria": ["No adjustment: the case is ineligible and the authorized amount is exactly $0.", "Minimal adjustment: an account credit or refund from $0.01 through $24.99 is authorized.", "Moderate adjustment: a partial refund from $25.00 through $59.99 is authorized.", "High adjustment: a refund from $60.00 through $120.00 on one invoice is authorized.", "Very high adjustment: more than $120.00 is authorized, or corrections must cover multiple paid invoices."], "instructions": "Determine the authorized adjustment impact level using the policy exception, not the subscriber’s requested amount. Calculate the prorated difference and select exactly one level.", "type": "score"}}, "state": "Billing case note: Invoice R-841 is a renewal charge for subscription S-204, which renewed on the Premium annual plan. Maya Chen submitted a downgrade request on September 16 for subscription S-204, targeting the Basic annual plan. This request arrived no more than 72 hours after the September 14 renewal charge. The Premium annual price on invoice R-841 was $120.00, while the applicable Basic annual price for the downgrade was $72.00. At the time of the request, 363 of 365 service days remained on the term, and the correction sought under the request covers exactly one paid invoice, R-841. The system activity log for subscription S-204 records exactly one entry timestamped September 15 between the September 14 renewal charge and Maya Chen's September 16 downgrade request. That logged entry is tagged as a Basic feature-use event. Applicable review language: renewal charges are normally outside refund scope, except downgrades requested within 72 hours with no premium use. That exception authorizes only the prorated annual price difference for the remaining 363 of 365 days, rounded to cents. The account file was routed to Priya Nair for final determination."}, "method": "c2d", "provenance": {"source_id": "diverse-017", "source_is_synthetic": true, "source_sha256": "f679dfa5a0034608d2918ef28db0f6e8700f11656da5ab3d9a2220fadef00258", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original criteria via the unchanged questions object and only alter case-specific service-coverage dates, which are permissible changeable observations; the focus evidence consists of two complete factual sentences about invoice coverage periods with no definitions or instructions; the counterfactual coherently shifts U-889's period to April while U-882 stays in March, removing the duplicate-payment overlap without contradicting any other fact such as equal amounts or shared subscription; entity and path bindings for U-882/U-889 and the pronoun resolution remain intact; no answer codes, rule tables, or rationale are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified a13 remains atomic over an explicit transaction set. The focus, equality of the two service-coverage periods, is factual rather than policy-based. Base and counter assignments differ only on that equality and are jointly realizable: equal periods produce a duplicate, while different periods produce a single disputed prorated upgrade charge. Empty policy_evidence is correct because the governing routing criteria and priority are already preserved verbatim in original_input.questions; the original state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies U-889 as the sole referenced unresolved charge and establishes that it is a later, distinct, settled payment matching U-882 for the same subscription and service period. This is sufficient for duplicate-payment review. Renewal and promotional-credit alternatives are explicitly excluded.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an invoice-supported, settled, prorated upgrade charge whose amount is disputed. U-882 covers a different period, and a13 excludes every other settled transaction from U-889's service period, so there is no repeated settled payment for that period. Renewal and promotional-credit alternatives are also excluded, making plan-change proration review sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pronoun “it” in the billing support agent’s note refers to the charge documented by invoice U-889."}, {"id": "a2", "statement": "Invoice U-889 documents the only unresolved charge disputed by Mara."}, {"id": "a3", "statement": "Invoice U-889 is invoice evidence for its charge."}, {"id": "a4", "statement": "The payment for invoice U-889 settled."}, {"id": "a5", "statement": "The payment for invoice U-882 settled."}, {"id": "a6", "statement": "The payment for invoice U-882 occurred before the payment for invoice U-889."}, {"id": "a7", "statement": "The payments for invoices U-882 and U-889 are distinct payment transactions."}, {"id": "a8", "statement": "The charges on invoices U-882 and U-889 have equal dollar amounts."}, {"id": "a9", "statement": "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription."}, {"id": "a10", "statement": "The service-coverage period on invoice U-882 equals the service-coverage period on invoice U-889."}, {"id": "a11", "statement": "The charge on invoice U-889 arose from Mara’s upgrade from Basic to Pro."}, {"id": "a12", "statement": "Mara disputes the amount charged on invoice U-889."}, {"id": "a13", "statement": "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889."}, {"id": "a14", "statement": "The unresolved issue concerning invoice U-889 is an omitted or misapplied promotional credit."}, {"id": "a15", "statement": "The charge on invoice U-889 is a scheduled renewal charge."}, {"id": "a16", "statement": "The amount on invoice U-889 is prorated for its recorded service-coverage period."}], "base_state_json": "{\"context\":\"Case note for subscription operations: Mara upgraded her StreamForge subscription from Basic to Pro. Two related invoices were issued and both were paid.\",\"facts\":[\"The pronoun 'it' in the billing support agent's note refers to the charge documented by invoice U-889.\",\"Invoice U-889 documents the only unresolved charge disputed by Mara.\",\"Invoice U-889 is invoice evidence for its charge.\",\"The payment for invoice U-889 settled.\",\"The payment for invoice U-882 settled.\",\"The payment for invoice U-882 occurred before the payment for invoice U-889.\",\"The payments for invoices U-882 and U-889 are distinct payment transactions.\",\"The charges on invoices U-882 and U-889 have equal dollar amounts.\",\"Invoices U-882 and U-889 charge for the same StreamForge Pro subscription.\",\"Invoice U-882 lists a service-coverage period of March 1–31, 2024.\",\"Invoice U-889 lists a service-coverage period of March 1–31, 2024.\",\"The charge on invoice U-889 arose from Mara's upgrade from Basic to Pro.\",\"Mara disputes the amount charged on invoice U-889.\",\"Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889.\",\"The unresolved issue concerning invoice U-889 is not an omitted or misapplied promotional credit.\",\"The charge on invoice U-889 is not a scheduled renewal charge.\",\"The amount on invoice U-889 is prorated for its recorded service-coverage period.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["facts", "9"], "text": "Invoice U-882 lists a service-coverage period of March 1–31, 2024."}, {"path": ["facts", "10"], "text": "Invoice U-889 lists a service-coverage period of March 1–31, 2024."}], "policy_evidence": [], "rules": [{"justification": "The referenced unresolved U-889 charge is a later, distinct, settled payment matching an earlier settled payment in amount, subscription, and service period. It therefore duplicates payment for the same subscription service period.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}]}, {"justification": "The sole unresolved invoice-supported dispute concerns the amount of a settled prorated upgrade charge. U-882 covers a different period, and every other settled transaction also covers a different period, so no repeated settled payment exists for U-889’s service period.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "Invoice U-882 lists a service-coverage period of March 1–31, 2024.", "negative_left": "Invoice U-882 lists a service-coverage period of March 1–31, 2024.", "negative_right": "Invoice U-889 lists a service-coverage period of April 1–30, 2024.", "right": "Invoice U-889 lists a service-coverage period of March 1–31, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-018-006", "id": "fast-43-diverse-018-006-base", "input": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"context": "Case note for subscription operations: Mara upgraded her StreamForge subscription from Basic to Pro. Two related invoices were issued and both were paid.", "facts": ["The pronoun 'it' in the billing support agent's note refers to the charge documented by invoice U-889.", "Invoice U-889 documents the only unresolved charge disputed by Mara.", "Invoice U-889 is invoice evidence for its charge.", "The payment for invoice U-889 settled.", "The payment for invoice U-882 settled.", "The payment for invoice U-882 occurred before the payment for invoice U-889.", "The payments for invoices U-882 and U-889 are distinct payment transactions.", "The charges on invoices U-882 and U-889 have equal dollar amounts.", "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription.", "Invoice U-882 lists a service-coverage period of March 1–31, 2024.", "Invoice U-889 lists a service-coverage period of March 1–31, 2024.", "The charge on invoice U-889 arose from Mara's upgrade from Basic to Pro.", "Mara disputes the amount charged on invoice U-889.", "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889.", "The unresolved issue concerning invoice U-889 is not an omitted or misapplied promotional credit.", "The charge on invoice U-889 is not a scheduled renewal charge.", "The amount on invoice U-889 is prorated for its recorded service-coverage period."]}}, "method": "c2d", "provenance": {"source_id": "diverse-018", "source_is_synthetic": true, "source_sha256": "9f3816a01540b36b15d4decda36965966708aa93a8de1d0b982b71aa44023c53", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-03", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original criteria via the unchanged questions object and only alter case-specific service-coverage dates, which are permissible changeable observations; the focus evidence consists of two complete factual sentences about invoice coverage periods with no definitions or instructions; the counterfactual coherently shifts U-889's period to April while U-882 stays in March, removing the duplicate-payment overlap without contradicting any other fact such as equal amounts or shared subscription; entity and path bindings for U-882/U-889 and the pronoun resolution remain intact; no answer codes, rule tables, or rationale are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "refuted", "a15": "refuted", "a16": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified a13 remains atomic over an explicit transaction set. The focus, equality of the two service-coverage periods, is factual rather than policy-based. Base and counter assignments differ only on that equality and are jointly realizable: equal periods produce a duplicate, while different periods produce a single disputed prorated upgrade charge. Empty policy_evidence is correct because the governing routing criteria and priority are already preserved verbatim in original_input.questions; the original state contributes case observations rather than additional interpretive policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction identifies U-889 as the sole referenced unresolved charge and establishes that it is a later, distinct, settled payment matching U-882 for the same subscription and service period. This is sufficient for duplicate-payment review. Renewal and promotional-credit alternatives are explicitly excluded.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes an invoice-supported, settled, prorated upgrade charge whose amount is disputed. U-882 covers a different period, and a13 excludes every other settled transaction from U-889's service period, so there is no repeated settled payment for that period. Renewal and promotional-credit alternatives are also excluded, making plan-change proration review sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pronoun “it” in the billing support agent’s note refers to the charge documented by invoice U-889."}, {"id": "a2", "statement": "Invoice U-889 documents the only unresolved charge disputed by Mara."}, {"id": "a3", "statement": "Invoice U-889 is invoice evidence for its charge."}, {"id": "a4", "statement": "The payment for invoice U-889 settled."}, {"id": "a5", "statement": "The payment for invoice U-882 settled."}, {"id": "a6", "statement": "The payment for invoice U-882 occurred before the payment for invoice U-889."}, {"id": "a7", "statement": "The payments for invoices U-882 and U-889 are distinct payment transactions."}, {"id": "a8", "statement": "The charges on invoices U-882 and U-889 have equal dollar amounts."}, {"id": "a9", "statement": "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription."}, {"id": "a10", "statement": "The service-coverage period on invoice U-882 equals the service-coverage period on invoice U-889."}, {"id": "a11", "statement": "The charge on invoice U-889 arose from Mara’s upgrade from Basic to Pro."}, {"id": "a12", "statement": "Mara disputes the amount charged on invoice U-889."}, {"id": "a13", "statement": "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889."}, {"id": "a14", "statement": "The unresolved issue concerning invoice U-889 is an omitted or misapplied promotional credit."}, {"id": "a15", "statement": "The charge on invoice U-889 is a scheduled renewal charge."}, {"id": "a16", "statement": "The amount on invoice U-889 is prorated for its recorded service-coverage period."}], "base_state_json": "{\"context\":\"Case note for subscription operations: Mara upgraded her StreamForge subscription from Basic to Pro. Two related invoices were issued and both were paid.\",\"facts\":[\"The pronoun 'it' in the billing support agent's note refers to the charge documented by invoice U-889.\",\"Invoice U-889 documents the only unresolved charge disputed by Mara.\",\"Invoice U-889 is invoice evidence for its charge.\",\"The payment for invoice U-889 settled.\",\"The payment for invoice U-882 settled.\",\"The payment for invoice U-882 occurred before the payment for invoice U-889.\",\"The payments for invoices U-882 and U-889 are distinct payment transactions.\",\"The charges on invoices U-882 and U-889 have equal dollar amounts.\",\"Invoices U-882 and U-889 charge for the same StreamForge Pro subscription.\",\"Invoice U-882 lists a service-coverage period of March 1–31, 2024.\",\"Invoice U-889 lists a service-coverage period of March 1–31, 2024.\",\"The charge on invoice U-889 arose from Mara's upgrade from Basic to Pro.\",\"Mara disputes the amount charged on invoice U-889.\",\"Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889.\",\"The unresolved issue concerning invoice U-889 is not an omitted or misapplied promotional credit.\",\"The charge on invoice U-889 is not a scheduled renewal charge.\",\"The amount on invoice U-889 is prorated for its recorded service-coverage period.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["facts", "9"], "text": "Invoice U-882 lists a service-coverage period of March 1–31, 2024."}, {"path": ["facts", "10"], "text": "Invoice U-889 lists a service-coverage period of March 1–31, 2024."}], "policy_evidence": [], "rules": [{"justification": "The referenced unresolved U-889 charge is a later, distinct, settled payment matching an earlier settled payment in amount, subscription, and service period. It therefore duplicates payment for the same subscription service period.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}]}, {"justification": "The sole unresolved invoice-supported dispute concerns the amount of a settled prorated upgrade charge. U-882 covers a different period, and every other settled transaction also covers a different period, so no repeated settled payment exists for U-889’s service period.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "refuted"}, {"atom_id": "a15", "state": "refuted"}, {"atom_id": "a16", "state": "supported"}]}]}, "verified_pair": {"left": "Invoice U-882 lists a service-coverage period of March 1–31, 2024.", "negative_left": "Invoice U-882 lists a service-coverage period of March 1–31, 2024.", "negative_right": "Invoice U-889 lists a service-coverage period of April 1–30, 2024.", "right": "Invoice U-889 lists a service-coverage period of March 1–31, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-018-006", "id": "fast-43-diverse-018-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Renewal-charge review: Use when the disputed charge is a scheduled renewal and no later evidence identifies a different charge as the unresolved issue.", "1 — Plan-change proration review: Use when the subscriber disputes the legitimate amount or timing of a single upgrade or downgrade charge, without a repeated settled payment for the same service period.", "2 — Promotional-credit review: Use when the unresolved issue is an advertised, promised, or documented discount or credit that was omitted or misapplied.", "3 — Duplicate-payment review: Use when the referenced unresolved charge is a second settled payment duplicating another payment for the same subscription service period, even if it followed a plan change.", "4 — No billing-adjustment route: Use when all referenced charges were accepted, reversed already, merely pending, or unsupported by invoice or payment evidence."], "instructions": "Select the single best routing level. Resolve pronouns from the surrounding dialogue and account timeline. Apply the first level whose definition matches the referenced unresolved charge.", "type": "score"}}, "state": {"context": "Case note for subscription operations: Mara upgraded her StreamForge subscription from Basic to Pro. Two related invoices were issued and both were paid.", "facts": ["The pronoun 'it' in the billing support agent's note refers to the charge documented by invoice U-889.", "Invoice U-889 documents the only unresolved charge disputed by Mara.", "Invoice U-889 is invoice evidence for its charge.", "The payment for invoice U-889 settled.", "The payment for invoice U-882 settled.", "The payment for invoice U-882 occurred before the payment for invoice U-889.", "The payments for invoices U-882 and U-889 are distinct payment transactions.", "The charges on invoices U-882 and U-889 have equal dollar amounts.", "Invoices U-882 and U-889 charge for the same StreamForge Pro subscription.", "Invoice U-882 lists a service-coverage period of March 1–31, 2024.", "Invoice U-889 lists a service-coverage period of April 1–30, 2024.", "The charge on invoice U-889 arose from Mara's upgrade from Basic to Pro.", "Mara disputes the amount charged on invoice U-889.", "Every settled payment transaction other than the payments for invoices U-882 and U-889 covers a service period different from the service-coverage period on invoice U-889.", "The unresolved issue concerning invoice U-889 is not an omitted or misapplied promotional credit.", "The charge on invoice U-889 is not a scheduled renewal charge.", "The amount on invoice U-889 is prorated for its recorded service-coverage period."]}}, "method": "c2d", "provenance": {"source_id": "diverse-018", "source_is_synthetic": true, "source_sha256": "9f3816a01540b36b15d4decda36965966708aa93a8de1d0b982b71aa44023c53", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "commerce-03", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original brand/size/label constraints and add factual, non-instructional evidence (nutrition panel reading and a manufacturing rule linking it to the front-label claim) without dictating an outcome; the counterfactual flips only the manufacturing-standard clause, remaining internally consistent with all other unchanged facts, and neither version states or implies the final approval decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: \\u201cSubstitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.\\u201d\"},{\"speaker\":\"Store picker\",\"text\":\"The requested North Mill creamy peanut butter is unavailable. I found North Mill smooth peanut butter in a 16-ounce jar, and the shelf tag and front label confirm both the brand and the size.\"},{\"speaker\":\"Store picker\",\"text\":\"The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed with a \\\"No Added Sugar\\\" claim on its front label.\"},{\"speaker\":\"Customer contact agent\",\"text\":\"I called and sent a text, but the customer has not responded.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"The pickup window closes in 12 minutes. Determine whether this replacement is ready for fulfillment approval.\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["2", "text"], "text": "The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving."}, {"path": ["3", "text"], "text": "Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed with a \"No Added Sugar\" claim on its front label."}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving.", "negative_left": "The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving.", "negative_right": "Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed without a \"No Added Sugar\" claim on its front label.", "right": "Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed with a \"No Added Sugar\" claim on its front label."}, "verifier_independent_model": false}, "family": "fast-43-diverse-021-014", "id": "fast-43-diverse-021-014-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Store picker", "text": "The requested North Mill creamy peanut butter is unavailable. I found North Mill smooth peanut butter in a 16-ounce jar, and the shelf tag and front label confirm both the brand and the size."}, {"speaker": "Store picker", "text": "The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving."}, {"speaker": "Fulfillment supervisor", "text": "Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed with a \"No Added Sugar\" claim on its front label."}, {"speaker": "Customer contact agent", "text": "I called and sent a text, but the customer has not responded."}, {"speaker": "Fulfillment supervisor", "text": "The pickup window closes in 12 minutes. Determine whether this replacement is ready for fulfillment approval."}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original brand/size/label constraints and add factual, non-instructional evidence (nutrition panel reading and a manufacturing rule linking it to the front-label claim) without dictating an outcome; the counterfactual flips only the manufacturing-standard clause, remaining internally consistent with all other unchanged facts, and neither version states or implies the final approval decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "full_context_fact_states": {"base": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "supported"}, "counterfactual": {"A_brand_match": "supported", "A_exact_size": "supported", "A_no_added_sugar_label": "refuted"}, "remove_left": {"A_no_added_sugar_label": "unknown"}, "remove_right": {"A_no_added_sugar_label": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A_no_added_sugar_label": "unknown"}, "negative_pair": {"A_no_added_sugar_label": "refuted"}, "negative_sentence": {"A_no_added_sugar_label": "unknown"}, "positive_pair": {"A_no_added_sugar_label": "supported"}, "right": {"A_no_added_sugar_label": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual product relationship, and the focus is the factual labeling attribute rather than a policy conclusion. The base and counter assignments are realizable with only that label fact changing. The policy evidence correctly preserves the customer’s state-originating substitution constraints; the remaining governing approval rules are already retained in the questions object. Both rules are sufficient for their targets, and omission of an unknown-without-waiver rule is permissible because partial rule tables may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies all three explicit substitution constraints—same brand, exactly 16 ounces, and a no-added-sugar label—so it is sufficient for approval under the stated policy.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the required no-added-sugar label directly contradicts an explicit substitution constraint. The false criterion makes any contradicted explicit constraint sufficient to deny approval, regardless of the other attributes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A_brand_match", "statement": "For the proposed North Mill smooth peanut butter substitution in this pickup order, its brand is the same as the brand of the requested North Mill creamy peanut butter."}, {"id": "A_exact_size", "statement": "The proposed North Mill smooth peanut butter substitution in this pickup order has a net quantity of exactly 16 ounces."}, {"id": "A_no_added_sugar_label", "statement": "The package of the proposed North Mill smooth peanut butter substitution in this pickup order is labeled “no added sugar.”"}], "base_state_json": "[{\"speaker\":\"Pickup customer\",\"text\":\"My order note says: \\u201cSubstitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.\\u201d\"},{\"speaker\":\"Store picker\",\"text\":\"The requested North Mill creamy peanut butter is unavailable. I found North Mill smooth peanut butter in a 16-ounce jar, and the shelf tag and front label confirm both the brand and the size.\"},{\"speaker\":\"Store picker\",\"text\":\"The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed with a \\\"No Added Sugar\\\" claim on its front label.\"},{\"speaker\":\"Customer contact agent\",\"text\":\"I called and sent a text, but the customer has not responded.\"},{\"speaker\":\"Fulfillment supervisor\",\"text\":\"The pickup window closes in 12 minutes. Determine whether this replacement is ready for fulfillment approval.\"}]", "base_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}], "counter_states": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "refuted"}], "focus_atom": "A_no_added_sugar_label", "focus_evidence": [{"path": ["2", "text"], "text": "The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving."}, {"path": ["3", "text"], "text": "Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed with a \"No Added Sugar\" claim on its front label."}], "policy_evidence": [{"path": ["0", "text"], "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}], "rules": [{"justification": "All three explicit customer substitution constraints are positively verified, so fulfillment may approve the replacement.", "target": "true", "when": [{"atom_id": "A_brand_match", "state": "supported"}, {"atom_id": "A_exact_size", "state": "supported"}, {"atom_id": "A_no_added_sugar_label", "state": "supported"}]}, {"justification": "The proposed replacement contradicts the explicit requirement that it be labeled “no added sugar,” so fulfillment may not approve it.", "target": "false", "when": [{"atom_id": "A_no_added_sugar_label", "state": "refuted"}]}]}, "verified_pair": {"left": "The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving.", "negative_left": "The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving.", "negative_right": "Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed without a \"No Added Sugar\" claim on its front label.", "right": "Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed with a \"No Added Sugar\" claim on its front label."}, "verifier_independent_model": false}, "family": "fast-43-diverse-021-014", "id": "fast-43-diverse-021-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one explicit substitution constraint is contradicted or remains unverified without a customer waiver, so fulfillment may not approve the replacement yet.", "true": "Yes — every explicit substitution constraint is verified or the customer has clearly waived any unverified constraint, so fulfillment may approve the replacement."}, "instructions": "Readiness task: Decide whether fulfillment may approve the proposed substitution now. Under store policy, approval requires every explicit customer constraint to be positively verified. An unverified required attribute cannot be assumed from product similarity. If any required attribute is unknown and the customer has not waived it, the correct decision is no and the order must be routed to customer contact or supervisor review.", "type": "noul"}}, "state": [{"speaker": "Pickup customer", "text": "My order note says: “Substitutions are fine only if they are the same brand, exactly 16 ounces, and labeled no added sugar.”"}, {"speaker": "Store picker", "text": "The requested North Mill creamy peanut butter is unavailable. I found North Mill smooth peanut butter in a 16-ounce jar, and the shelf tag and front label confirm both the brand and the size."}, {"speaker": "Store picker", "text": "The nutrition panel of the proposed North Mill smooth peanut butter substitution in this pickup order lists 0 grams of added sugar per serving."}, {"speaker": "Fulfillment supervisor", "text": "Under North Mill's manufacturing standard, any jar of its smooth peanut butter that lists 0 grams of added sugar per serving is printed without a \"No Added Sugar\" claim on its front label."}, {"speaker": "Customer contact agent", "text": "I called and sent a text, but the customer has not responded."}, {"speaker": "Fulfillment supervisor", "text": "The pickup window closes in 12 minutes. Determine whether this replacement is ready for fulfillment approval."}]}, "method": "c2d", "provenance": {"source_id": "diverse-021", "source_is_synthetic": true, "source_sha256": "f7cb9743dacdde76ef9c36b6a9f740c3653e9f861c2e93030711ef447ccaf481", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same substitution policy, price cap, and brand/stage exception; the counterfactual coherently raises the replacement price to $21.00, still consistent with all other unchanged facts and no contradictory duplicate measurements appear; evidence spans are two complete factual price sentences with no embedded rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The item is unavailable in the store. The picker proposes substituting one 22-ounce can of BrightStart Stage 2 infant formula, matching both brand and stage. The unavailable infant formula item in Lena’s pickup order is priced at $18.00. The proposed replacement infant formula item is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable infant formula item in Lena’s pickup order is priced at $18.00."}, {"path": [], "text": "The proposed replacement infant formula item is priced at $19.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable infant formula item in Lena’s pickup order is priced at $18.00.", "negative_left": "The unavailable infant formula item in Lena’s pickup order is priced at $18.00.", "negative_right": "The proposed replacement infant formula item is priced at $21.00.", "right": "The proposed replacement infant formula item is priced at $19.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-003", "id": "fast-43-diverse-022-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The item is unavailable in the store. The picker proposes substituting one 22-ounce can of BrightStart Stage 2 infant formula, matching both brand and stage. The unavailable infant formula item in Lena’s pickup order is priced at $18.00. The proposed replacement infant formula item is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same substitution policy, price cap, and brand/stage exception; the counterfactual coherently raises the replacement price to $21.00, still consistent with all other unchanged facts and no contradictory duplicate measurements appear; evidence spans are two complete factual price sentences with no embedded rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The item is unavailable in the store. The picker proposes substituting one 22-ounce can of BrightStart Stage 2 infant formula, matching both brand and stage. The unavailable infant formula item in Lena’s pickup order is priced at $18.00. The proposed replacement infant formula item is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable infant formula item in Lena’s pickup order is priced at $18.00."}, {"path": [], "text": "The proposed replacement infant formula item is priced at $19.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable infant formula item in Lena’s pickup order is priced at $18.00.", "negative_left": "The unavailable infant formula item in Lena’s pickup order is priced at $18.00.", "negative_right": "The proposed replacement infant formula item is priced at $21.00.", "right": "The proposed replacement infant formula item is priced at $19.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-003", "id": "fast-43-diverse-022-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The item is unavailable in the store. The picker proposes substituting one 22-ounce can of BrightStart Stage 2 infant formula, matching both brand and stage. The unavailable infant formula item in Lena’s pickup order is priced at $18.00. The proposed replacement infant formula item is priced at $21.00. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy clauses verbatim and only vary the replacement price, which is a permissible observation change; the two evidence spans are plain factual price statements; the counterfactual coherently shifts only the replacement price sentence (to $23.50) without contradicting other facts; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula from her order. The item is unavailable at the store. The picker proposes a substitute: a can of BrightStart Stage 2 infant formula, same brand and same stage as the original. Lena's order note says, \\u201cSubstitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.\\u201d The unavailable ordered infant formula in Lena's pickup order is priced at $20.00. The proposed replacement infant formula is priced at $21.50. Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered infant formula in Lena's pickup order is priced at $20.00."}, {"path": [], "text": "The proposed replacement infant formula is priced at $21.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered infant formula in Lena's pickup order is priced at $20.00.", "negative_left": "The unavailable ordered infant formula in Lena's pickup order is priced at $20.00.", "negative_right": "The proposed replacement infant formula is priced at $23.50.", "right": "The proposed replacement infant formula is priced at $21.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-007", "id": "fast-43-diverse-022-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula from her order. The item is unavailable at the store. The picker proposes a substitute: a can of BrightStart Stage 2 infant formula, same brand and same stage as the original. Lena's order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” The unavailable ordered infant formula in Lena's pickup order is priced at $20.00. The proposed replacement infant formula is priced at $21.50. Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy clauses verbatim and only vary the replacement price, which is a permissible observation change; the two evidence spans are plain factual price statements; the counterfactual coherently shifts only the replacement price sentence (to $23.50) without contradicting other facts; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula from her order. The item is unavailable at the store. The picker proposes a substitute: a can of BrightStart Stage 2 infant formula, same brand and same stage as the original. Lena's order note says, \\u201cSubstitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.\\u201d The unavailable ordered infant formula in Lena's pickup order is priced at $20.00. The proposed replacement infant formula is priced at $21.50. Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered infant formula in Lena's pickup order is priced at $20.00."}, {"path": [], "text": "The proposed replacement infant formula is priced at $21.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered infant formula in Lena's pickup order is priced at $20.00.", "negative_left": "The unavailable ordered infant formula in Lena's pickup order is priced at $20.00.", "negative_right": "The proposed replacement infant formula is priced at $23.50.", "right": "The proposed replacement infant formula is priced at $21.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-007", "id": "fast-43-diverse-022-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula from her order. The item is unavailable at the store. The picker proposes a substitute: a can of BrightStart Stage 2 infant formula, same brand and same stage as the original. Lena's order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” The unavailable ordered infant formula in Lena's pickup order is priced at $20.00. The proposed replacement infant formula is priced at $23.50. Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the order note and fulfillment policy verbatim, keep Lena/infant-formula bindings intact, use two factual price sentences as evidence, and the counterfactual only alters the replacement price without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Case note: Pickup order for Lena includes one 20-ounce can of BrightStart Stage 2 infant formula, marked unavailable. The store picker proposes substituting one 22-ounce can of BrightStart Stage 2 infant formula. Order note on file: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review. Both cans are BrightStart brand, both labeled Stage 2, confirming brand and stage match. The unavailable infant formula ordered by Lena is priced at $24.00 per container. The proposed replacement infant formula is priced at $25.80 per container.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable infant formula ordered by Lena is priced at $24.00 per container."}, {"path": [], "text": "The proposed replacement infant formula is priced at $25.80 per container."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable infant formula ordered by Lena is priced at $24.00 per container.", "negative_left": "The unavailable infant formula ordered by Lena is priced at $24.00 per container.", "negative_right": "The proposed replacement infant formula is priced at $27.50 per container.", "right": "The proposed replacement infant formula is priced at $25.80 per container."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-014", "id": "fast-43-diverse-022-014-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Case note: Pickup order for Lena includes one 20-ounce can of BrightStart Stage 2 infant formula, marked unavailable. The store picker proposes substituting one 22-ounce can of BrightStart Stage 2 infant formula. Order note on file: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review. Both cans are BrightStart brand, both labeled Stage 2, confirming brand and stage match. The unavailable infant formula ordered by Lena is priced at $24.00 per container. The proposed replacement infant formula is priced at $25.80 per container."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the order note and fulfillment policy verbatim, keep Lena/infant-formula bindings intact, use two factual price sentences as evidence, and the counterfactual only alters the replacement price without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Case note: Pickup order for Lena includes one 20-ounce can of BrightStart Stage 2 infant formula, marked unavailable. The store picker proposes substituting one 22-ounce can of BrightStart Stage 2 infant formula. Order note on file: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review. Both cans are BrightStart brand, both labeled Stage 2, confirming brand and stage match. The unavailable infant formula ordered by Lena is priced at $24.00 per container. The proposed replacement infant formula is priced at $25.80 per container.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable infant formula ordered by Lena is priced at $24.00 per container."}, {"path": [], "text": "The proposed replacement infant formula is priced at $25.80 per container."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable infant formula ordered by Lena is priced at $24.00 per container.", "negative_left": "The unavailable infant formula ordered by Lena is priced at $24.00 per container.", "negative_right": "The proposed replacement infant formula is priced at $27.50 per container.", "right": "The proposed replacement infant formula is priced at $25.80 per container."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-014", "id": "fast-43-diverse-022-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Case note: Pickup order for Lena includes one 20-ounce can of BrightStart Stage 2 infant formula, marked unavailable. The store picker proposes substituting one 22-ounce can of BrightStart Stage 2 infant formula. Order note on file: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review. Both cans are BrightStart brand, both labeled Stage 2, confirming brand and stage match. The unavailable infant formula ordered by Lena is priced at $24.00 per container. The proposed replacement infant formula is priced at $27.50 per container."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original order note and fulfillment policy verbatim, keep Lena/infant-formula/brand/stage bindings intact, and only vary the replacement price ($25.80 vs $27.20), producing a coherent counterfactual that crosses the 10% threshold without contradicting other facts; the two evidence spans are complete factual price statements, and neither context states or implies the final approve/deny outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Case note: Pickup order review for customer Lena. The unavailable item in Lena’s pickup order is infant formula. The proposed replacement is the same BrightStart brand as the unavailable ordered item. The proposed replacement is Stage 2, matching the unavailable ordered item’s stage. The unavailable infant formula ordered by Lena is priced at $24.00 per container. The proposed replacement infant formula is priced at $25.80 per container. Order note: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable infant formula ordered by Lena is priced at $24.00 per container."}, {"path": [], "text": "The proposed replacement infant formula is priced at $25.80 per container."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable infant formula ordered by Lena is priced at $24.00 per container.", "negative_left": "The unavailable infant formula ordered by Lena is priced at $24.00 per container.", "negative_right": "The proposed replacement infant formula is priced at $27.20 per container.", "right": "The proposed replacement infant formula is priced at $25.80 per container."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-018", "id": "fast-43-diverse-022-018-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Case note: Pickup order review for customer Lena. The unavailable item in Lena’s pickup order is infant formula. The proposed replacement is the same BrightStart brand as the unavailable ordered item. The proposed replacement is Stage 2, matching the unavailable ordered item’s stage. The unavailable infant formula ordered by Lena is priced at $24.00 per container. The proposed replacement infant formula is priced at $25.80 per container. Order note: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original order note and fulfillment policy verbatim, keep Lena/infant-formula/brand/stage bindings intact, and only vary the replacement price ($25.80 vs $27.20), producing a coherent counterfactual that crosses the 10% threshold without contradicting other facts; the two evidence spans are complete factual price statements, and neither context states or implies the final approve/deny outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Case note: Pickup order review for customer Lena. The unavailable item in Lena’s pickup order is infant formula. The proposed replacement is the same BrightStart brand as the unavailable ordered item. The proposed replacement is Stage 2, matching the unavailable ordered item’s stage. The unavailable infant formula ordered by Lena is priced at $24.00 per container. The proposed replacement infant formula is priced at $25.80 per container. Order note: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable infant formula ordered by Lena is priced at $24.00 per container."}, {"path": [], "text": "The proposed replacement infant formula is priced at $25.80 per container."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable infant formula ordered by Lena is priced at $24.00 per container.", "negative_left": "The unavailable infant formula ordered by Lena is priced at $24.00 per container.", "negative_right": "The proposed replacement infant formula is priced at $27.20 per container.", "right": "The proposed replacement infant formula is priced at $25.80 per container."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-018", "id": "fast-43-diverse-022-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Case note: Pickup order review for customer Lena. The unavailable item in Lena’s pickup order is infant formula. The proposed replacement is the same BrightStart brand as the unavailable ordered item. The proposed replacement is Stage 2, matching the unavailable ordered item’s stage. The unavailable infant formula ordered by Lena is priced at $24.00 per container. The proposed replacement infant formula is priced at $27.20 per container. Order note: “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text, entity/item bindings, and only the replacement price differs between them (a single coherent factual change), with the two focus evidence sentences present verbatim and no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula for her order. The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00. That item is now out of stock. The store picker located one 22-ounce can of BrightStart Stage 2 infant formula as a possible substitute. The proposed replacement item is priced at $25.80. Both the original and replacement carry the same BrightStart brand label and are both marked Stage 2 on their packaging. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00."}, {"path": [], "text": "The proposed replacement item is priced at $25.80."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00.", "negative_right": "The proposed replacement item is priced at $28.20.", "right": "The proposed replacement item is priced at $25.80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-023", "id": "fast-43-diverse-022-023-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula for her order. The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00. That item is now out of stock. The store picker located one 22-ounce can of BrightStart Stage 2 infant formula as a possible substitute. The proposed replacement item is priced at $25.80. Both the original and replacement carry the same BrightStart brand label and are both marked Stage 2 on their packaging. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text, entity/item bindings, and only the replacement price differs between them (a single coherent factual change), with the two focus evidence sentences present verbatim and no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula for her order. The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00. That item is now out of stock. The store picker located one 22-ounce can of BrightStart Stage 2 infant formula as a possible substitute. The proposed replacement item is priced at $25.80. Both the original and replacement carry the same BrightStart brand label and are both marked Stage 2 on their packaging. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00."}, {"path": [], "text": "The proposed replacement item is priced at $25.80."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00.", "negative_right": "The proposed replacement item is priced at $28.20.", "right": "The proposed replacement item is priced at $25.80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-023", "id": "fast-43-diverse-022-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula for her order. The unavailable ordered item in Lena’s pickup order, infant formula item L-884, is priced at $24.00. That item is now out of stock. The store picker located one 22-ounce can of BrightStart Stage 2 infant formula as a possible substitute. The proposed replacement item is priced at $28.20. Both the original and replacement carry the same BrightStart brand label and are both marked Stage 2 on their packaging. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy text and Lena's order-note rule, matching the original question's decision criteria; the focus evidence consists of two plain factual sentences about item and replacement prices, not policy language; the counterfactual raises the replacement price to $27.60 (15% over $24), which is internally consistent and not contradicted elsewhere; entity (Lena) and scope bindings are preserved even though case-specific quantities differ, which is permitted; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 12.4-ounce can of infant formula for her order. The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00. The store picker located a substitute from the same manufacturer and formulation stage. The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $25.80. Lena’s order note says, \\u201cSubstitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.\\u201d Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00."}, {"path": [], "text": "The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $25.80."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00.", "negative_right": "The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $27.60.", "right": "The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $25.80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-027", "id": "fast-43-diverse-022-027-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 12.4-ounce can of infant formula for her order. The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00. The store picker located a substitute from the same manufacturer and formulation stage. The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $25.80. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy text and Lena's order-note rule, matching the original question's decision criteria; the focus evidence consists of two plain factual sentences about item and replacement prices, not policy language; the counterfactual raises the replacement price to $27.60 (15% over $24), which is internally consistent and not contradicted elsewhere; entity (Lena) and scope bindings are preserved even though case-specific quantities differ, which is permitted; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 12.4-ounce can of infant formula for her order. The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00. The store picker located a substitute from the same manufacturer and formulation stage. The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $25.80. Lena’s order note says, \\u201cSubstitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.\\u201d Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00."}, {"path": [], "text": "The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $25.80."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00.", "negative_right": "The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $27.60.", "right": "The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $25.80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-027", "id": "fast-43-diverse-022-027-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 12.4-ounce can of infant formula for her order. The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $24.00. The store picker located a substitute from the same manufacturer and formulation stage. The proposed replacement is a 12.4-ounce can of the same brand and stage priced at $27.60. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full order note and fulfillment policy language unchanged, only the replacement price differs (19.20 vs 21.00), preserving SKUs, brand/stage bindings, and evidence sentences are factual price statements without embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, SKU 4471. The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00. The item is currently out of stock at the store. The store picker has located a substitute can, SKU 4488, also BrightStart brand and also labeled Stage 2, matching both the brand and the developmental stage of the original order. The proposed replacement item, SKU 4488, is priced at $19.20. Lena's order note says, \\u201cSubstitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.\\u201d Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00."}, {"path": [], "text": "The proposed replacement item, SKU 4488, is priced at $19.20."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00.", "negative_left": "The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00.", "negative_right": "The proposed replacement item, SKU 4488, is priced at $21.00.", "right": "The proposed replacement item, SKU 4488, is priced at $19.20."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-037", "id": "fast-43-diverse-022-037-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, SKU 4471. The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00. The item is currently out of stock at the store. The store picker has located a substitute can, SKU 4488, also BrightStart brand and also labeled Stage 2, matching both the brand and the developmental stage of the original order. The proposed replacement item, SKU 4488, is priced at $19.20. Lena's order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full order note and fulfillment policy language unchanged, only the replacement price differs (19.20 vs 21.00), preserving SKUs, brand/stage bindings, and evidence sentences are factual price statements without embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, SKU 4471. The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00. The item is currently out of stock at the store. The store picker has located a substitute can, SKU 4488, also BrightStart brand and also labeled Stage 2, matching both the brand and the developmental stage of the original order. The proposed replacement item, SKU 4488, is priced at $19.20. Lena's order note says, \\u201cSubstitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.\\u201d Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00."}, {"path": [], "text": "The proposed replacement item, SKU 4488, is priced at $19.20."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00.", "negative_left": "The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00.", "negative_right": "The proposed replacement item, SKU 4488, is priced at $21.00.", "right": "The proposed replacement item, SKU 4488, is priced at $19.20."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-037", "id": "fast-43-diverse-022-037-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, SKU 4471. The unavailable ordered infant formula item in Lena's order, SKU 4471, is priced at $18.00. The item is currently out of stock at the store. The store picker has located a substitute can, SKU 4488, also BrightStart brand and also labeled Stage 2, matching both the brand and the developmental stage of the original order. The proposed replacement item, SKU 4488, is priced at $21.00. Lena's order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy text, item/customer bindings, and two factual evidence sentences; the counterfactual only changes the replacement price to $21.60, remaining internally consistent without contradicting other stated facts or leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one can of BrightStart Stage 2 infant formula, item #4471. The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00. The store picker located a replacement can, item #4488, also BrightStart brand and also Stage 2, matching both the brand and the stage of the original item. The proposed replacement, item #4488, is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00."}, {"path": [], "text": "The proposed replacement, item #4488, is priced at $19.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00.", "negative_right": "The proposed replacement, item #4488, is priced at $21.60.", "right": "The proposed replacement, item #4488, is priced at $19.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-040", "id": "fast-43-diverse-022-040-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one can of BrightStart Stage 2 infant formula, item #4471. The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00. The store picker located a replacement can, item #4488, also BrightStart brand and also Stage 2, matching both the brand and the stage of the original item. The proposed replacement, item #4488, is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy text, item/customer bindings, and two factual evidence sentences; the counterfactual only changes the replacement price to $21.60, remaining internally consistent without contradicting other stated facts or leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one can of BrightStart Stage 2 infant formula, item #4471. The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00. The store picker located a replacement can, item #4488, also BrightStart brand and also Stage 2, matching both the brand and the stage of the original item. The proposed replacement, item #4488, is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00."}, {"path": [], "text": "The proposed replacement, item #4488, is priced at $19.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00.", "negative_right": "The proposed replacement, item #4488, is priced at $21.60.", "right": "The proposed replacement, item #4488, is priced at $19.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-040", "id": "fast-43-diverse-022-040-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one can of BrightStart Stage 2 infant formula, item #4471. The unavailable ordered item in Lena’s pickup order, infant formula item #4471, is priced at $18.00. The store picker located a replacement can, item #4488, also BrightStart brand and also Stage 2, matching both the brand and the stage of the original item. The proposed replacement, item #4488, is priced at $21.60. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text verbatim, keep item code F-204 and paths unchanged, use two complete factual price sentences as evidence, differ coherently only in the replacement price ($25.80 vs $28.50), and contain no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, item code F-204, for pickup fulfillment. The item is now unavailable at the store. The store picker identifies a substitute: a 22-ounce can of BrightStart Stage 2, also logged under item code F-204 for pricing comparison purposes. The picker confirms the replacement can bears the same BrightStart brand label and the same Stage 2 designation as the original item, matching both required attributes for infant formula substitutions on this order. The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00. The proposed replacement for item code F-204 is priced at $25.80. No other items in Lena's order are affected. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00."}, {"path": [], "text": "The proposed replacement for item code F-204 is priced at $25.80."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00.", "negative_right": "The proposed replacement for item code F-204 is priced at $28.50.", "right": "The proposed replacement for item code F-204 is priced at $25.80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-048", "id": "fast-43-diverse-022-048-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, item code F-204, for pickup fulfillment. The item is now unavailable at the store. The store picker identifies a substitute: a 22-ounce can of BrightStart Stage 2, also logged under item code F-204 for pricing comparison purposes. The picker confirms the replacement can bears the same BrightStart brand label and the same Stage 2 designation as the original item, matching both required attributes for infant formula substitutions on this order. The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00. The proposed replacement for item code F-204 is priced at $25.80. No other items in Lena's order are affected. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text verbatim, keep item code F-204 and paths unchanged, use two complete factual price sentences as evidence, differ coherently only in the replacement price ($25.80 vs $28.50), and contain no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, item code F-204, for pickup fulfillment. The item is now unavailable at the store. The store picker identifies a substitute: a 22-ounce can of BrightStart Stage 2, also logged under item code F-204 for pricing comparison purposes. The picker confirms the replacement can bears the same BrightStart brand label and the same Stage 2 designation as the original item, matching both required attributes for infant formula substitutions on this order. The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00. The proposed replacement for item code F-204 is priced at $25.80. No other items in Lena's order are affected. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00."}, {"path": [], "text": "The proposed replacement for item code F-204 is priced at $25.80."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00.", "negative_right": "The proposed replacement for item code F-204 is priced at $28.50.", "right": "The proposed replacement for item code F-204 is priced at $25.80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-048", "id": "fast-43-diverse-022-048-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, item code F-204, for pickup fulfillment. The item is now unavailable at the store. The store picker identifies a substitute: a 22-ounce can of BrightStart Stage 2, also logged under item code F-204 for pricing comparison purposes. The picker confirms the replacement can bears the same BrightStart brand label and the same Stage 2 designation as the original item, matching both required attributes for infant formula substitutions on this order. The unavailable ordered item in Lena’s pickup order, item code F-204, is priced at $24.00. The proposed replacement for item code F-204 is priced at $28.50. No other items in Lena's order are affected. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy clauses verbatim and preserve the SKU/brand/stage bindings from the question; the two evidence sentences are factual price statements, not policy text; the counterfactual only changes the replacement price to $21.00, which is coherent with the rest of the unchanged context and introduces no contradictory duplicate measurements; neither context reveals a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, SKU 4471, for her order. The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00. The store picker located a replacement can, also BrightStart brand and also Stage 2, matching both the brand and the formula stage required by Lena's note. The proposed replacement for SKU 4471 is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00."}, {"path": [], "text": "The proposed replacement for SKU 4471 is priced at $19.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00.", "negative_right": "The proposed replacement for SKU 4471 is priced at $21.00.", "right": "The proposed replacement for SKU 4471 is priced at $19.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-050", "id": "fast-43-diverse-022-050-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, SKU 4471, for her order. The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00. The store picker located a replacement can, also BrightStart brand and also Stage 2, matching both the brand and the formula stage required by Lena's note. The proposed replacement for SKU 4471 is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy clauses verbatim and preserve the SKU/brand/stage bindings from the question; the two evidence sentences are factual price statements, not policy text; the counterfactual only changes the replacement price to $21.00, which is coherent with the rest of the unchanged context and introduces no contradictory duplicate measurements; neither context reveals a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, SKU 4471, for her order. The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00. The store picker located a replacement can, also BrightStart brand and also Stage 2, matching both the brand and the formula stage required by Lena's note. The proposed replacement for SKU 4471 is priced at $19.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00."}, {"path": [], "text": "The proposed replacement for SKU 4471 is priced at $19.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00.", "negative_right": "The proposed replacement for SKU 4471 is priced at $21.00.", "right": "The proposed replacement for SKU 4471 is priced at $19.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-050", "id": "fast-43-diverse-022-050-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula, SKU 4471, for her order. The unavailable ordered item in Lena’s pickup order, SKU 4471, is priced at $18.00. The store picker located a replacement can, also BrightStart brand and also Stage 2, matching both the brand and the formula stage required by Lena's note. The proposed replacement for SKU 4471 is priced at $21.00. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Lena, the pickup order, the exact order-note and fulfillment-policy sentences verbatim, satisfying the policy contract; only the replacement price changes between contexts, which is a permissible observation change; the two focus evidence spans are complete factual sentences about item and replacement, not policy text; the counterfactual raises the replacement price to $16.50 (a >10% increase) without duplicating or contradicting any other stated measurement, remaining internally coherent; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena placed an order for infant formula. The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00. The store picker identified a replacement option to consider. The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $15.10. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00."}, {"path": [], "text": "The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $15.10."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00.", "negative_right": "The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $16.50.", "right": "The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $15.10."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-051", "id": "fast-43-diverse-022-051-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena placed an order for infant formula. The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00. The store picker identified a replacement option to consider. The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $15.10. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Lena, the pickup order, the exact order-note and fulfillment-policy sentences verbatim, satisfying the policy contract; only the replacement price changes between contexts, which is a permissible observation change; the two focus evidence spans are complete factual sentences about item and replacement, not policy text; the counterfactual raises the replacement price to $16.50 (a >10% increase) without duplicating or contradicting any other stated measurement, remaining internally coherent; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena placed an order for infant formula. The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00. The store picker identified a replacement option to consider. The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $15.10. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00."}, {"path": [], "text": "The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $15.10."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00.", "negative_right": "The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $16.50.", "right": "The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $15.10."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-051", "id": "fast-43-diverse-022-051-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena placed an order for infant formula. The unavailable ordered item in Lena’s pickup order is a 12.4-ounce can of infant formula priced at $14.00. The store picker identified a replacement option to consider. The proposed replacement is a 12.4-ounce can of the same infant-formula brand and stage, priced at $16.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical order-note and fulfillment-policy language from the original state, with only the replacement price varying between $25.80 and $28.20 as the single counterfactual change; the two focus evidence spans are plain factual price statements, not policy text; the question and its yes/no criteria are unchanged; no answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00. The store picker located a replacement: one 22-ounce can of BrightStart Stage 2 infant formula, same brand and same stage as the original item. The proposed replacement infant formula is priced at $25.80. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00."}, {"path": [], "text": "The proposed replacement infant formula is priced at $25.80."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00.", "negative_right": "The proposed replacement infant formula is priced at $28.20.", "right": "The proposed replacement infant formula is priced at $25.80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-052", "id": "fast-43-diverse-022-052-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00. The store picker located a replacement: one 22-ounce can of BrightStart Stage 2 infant formula, same brand and same stage as the original item. The proposed replacement infant formula is priced at $25.80. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical order-note and fulfillment-policy language from the original state, with only the replacement price varying between $25.80 and $28.20 as the single counterfactual change; the two focus evidence spans are plain factual price statements, not policy text; the question and its yes/no criteria are unchanged; no answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00. The store picker located a replacement: one 22-ounce can of BrightStart Stage 2 infant formula, same brand and same stage as the original item. The proposed replacement infant formula is priced at $25.80. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00."}, {"path": [], "text": "The proposed replacement infant formula is priced at $25.80."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00.", "negative_right": "The proposed replacement infant formula is priced at $28.20.", "right": "The proposed replacement infant formula is priced at $25.80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-052", "id": "fast-43-diverse-022-052-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The unavailable ordered item in Lena’s pickup order is infant formula priced at $24.00. The store picker located a replacement: one 22-ounce can of BrightStart Stage 2 infant formula, same brand and same stage as the original item. The proposed replacement infant formula is priced at $28.20. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical order note and fulfillment policy language, and the original questions object supplies the decision criteria without duplication; the only change between contexts is the single replacement price ($23.50 vs $26.40), which is a coherent, non-contradictory factual edit; SKUs, brand, stage, and customer remain bound to the original question; the two focus evidence spans are simple factual price statements, not policy text; neither context states or implies the yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup order for Lena, SKU 48213, requested one 20-ounce can of BrightStart Stage 2 infant formula. The item is marked unavailable at the store. The picker locates a substitute, SKU 48299, also BrightStart brand and also labeled Stage 2 formula, in a 22-ounce can. The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00. The proposed replacement item, SKU 48299, is priced at $23.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00."}, {"path": [], "text": "The proposed replacement item, SKU 48299, is priced at $23.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00.", "negative_right": "The proposed replacement item, SKU 48299, is priced at $26.40.", "right": "The proposed replacement item, SKU 48299, is priced at $23.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-054", "id": "fast-43-diverse-022-054-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup order for Lena, SKU 48213, requested one 20-ounce can of BrightStart Stage 2 infant formula. The item is marked unavailable at the store. The picker locates a substitute, SKU 48299, also BrightStart brand and also labeled Stage 2 formula, in a 22-ounce can. The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00. The proposed replacement item, SKU 48299, is priced at $23.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical order note and fulfillment policy language, and the original questions object supplies the decision criteria without duplication; the only change between contexts is the single replacement price ($23.50 vs $26.40), which is a coherent, non-contradictory factual edit; SKUs, brand, stage, and customer remain bound to the original question; the two focus evidence spans are simple factual price statements, not policy text; neither context states or implies the yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup order for Lena, SKU 48213, requested one 20-ounce can of BrightStart Stage 2 infant formula. The item is marked unavailable at the store. The picker locates a substitute, SKU 48299, also BrightStart brand and also labeled Stage 2 formula, in a 22-ounce can. The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00. The proposed replacement item, SKU 48299, is priced at $23.50. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00."}, {"path": [], "text": "The proposed replacement item, SKU 48299, is priced at $23.50."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00.", "negative_right": "The proposed replacement item, SKU 48299, is priced at $26.40.", "right": "The proposed replacement item, SKU 48299, is priced at $23.50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-054", "id": "fast-43-diverse-022-054-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup order for Lena, SKU 48213, requested one 20-ounce can of BrightStart Stage 2 infant formula. The item is marked unavailable at the store. The picker locates a substitute, SKU 48299, also BrightStart brand and also labeled Stage 2 formula, in a 22-ounce can. The unavailable ordered item in Lena’s pickup order, SKU 48213, is priced at $22.00. The proposed replacement item, SKU 48299, is priced at $26.40. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy text and Lena's original item/path bindings; the two evidence spans are plain factual price statements, not policy language; the counterfactual only changes the replacement price to $17.00, a coherent single-sentence edit with no contradictions; neither context states or hints at the final yes/no decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The item is unavailable in-store. The store picker identifies a replacement: one 22-ounce can of BrightStart Stage 2 infant formula, matching both the brand and the stage of the original item. The unavailable ordered item in Lena’s pickup order is priced at $14.00. The proposed replacement item is priced at $15.00. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order is priced at $14.00."}, {"path": [], "text": "The proposed replacement item is priced at $15.00."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order is priced at $14.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order is priced at $14.00.", "negative_right": "The proposed replacement item is priced at $17.00.", "right": "The proposed replacement item is priced at $15.00."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-058", "id": "fast-43-diverse-022-058-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The item is unavailable in-store. The store picker identifies a replacement: one 22-ounce can of BrightStart Stage 2 infant formula, matching both the brand and the stage of the original item. The unavailable ordered item in Lena’s pickup order is priced at $14.00. The proposed replacement item is priced at $15.00. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-04", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy text and Lena's original item/path bindings; the two evidence spans are plain factual price statements, not policy language; the counterfactual only changes the replacement price to $17.00, a coherent single-sentence edit with no contradictions; neither context states or hints at the final yes/no decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "full_context_fact_states": {"base": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "supported", "stage_match": "supported"}, "counterfactual": {"brand_match": "supported", "ordered_category": "supported", "price_within_limit": "refuted", "stage_match": "supported"}, "remove_left": {"price_within_limit": "unknown"}, "remove_right": {"price_within_limit": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"price_within_limit": "unknown"}, "negative_pair": {"price_within_limit": "refuted"}, "negative_sentence": {"price_within_limit": "unknown"}, "positive_pair": {"price_within_limit": "supported"}, "right": {"price_within_limit": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship rather than a policy conclusion or bundled classification. The focus atom is the factual price-limit relationship. The base and counter assignments are realizable by varying only the replacement price while keeping category, brand, and stage fixed. The policy evidence appropriately preserves the substantive substitution rules originating in the original state, including the general price ceiling, the infant-formula exception, and the routing consequences. Both rules include the necessary conditions or exclusion and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the ordered item is infant formula, the replacement matches both required exception attributes (brand and stage), and the replacement satisfies the general price ceiling. Under the preserved policy, these conditions are sufficient for approval and routing to picking.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that the replacement exceeds the applicable 10%-extra price limit. Even with matching brand and stage, failure of the mandatory general price condition makes the substitution a failed match, so it must not be approved for picking.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "ordered_category", "statement": "The unavailable item in Lena’s pickup order is infant formula."}, {"id": "brand_match", "statement": "The proposed replacement has the same brand as the unavailable ordered item."}, {"id": "stage_match", "statement": "The proposed replacement has the same infant-formula stage as the unavailable ordered item."}, {"id": "price_within_limit", "statement": "The proposed replacement’s price is no more than 10% greater than the unavailable ordered item’s price."}], "base_state_json": "\"Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The item is unavailable in-store. The store picker identifies a replacement: one 22-ounce can of BrightStart Stage 2 infant formula, matching both the brand and the stage of the original item. The unavailable ordered item in Lena’s pickup order is priced at $14.00. The proposed replacement item is priced at $15.00. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review.\"", "base_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}], "counter_states": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}], "focus_atom": "price_within_limit", "focus_evidence": [{"path": [], "text": "The unavailable ordered item in Lena’s pickup order is priced at $14.00."}, {"path": [], "text": "The proposed replacement item is priced at $15.00."}], "policy_evidence": [{"path": [], "text": "“Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.”"}, {"path": [], "text": "Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}], "rules": [{"justification": "For infant formula, the replacement satisfies the same-brand and same-stage exception as well as the general 10%-extra price limit, so it may be approved and sent to picking.", "target": "true", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "supported"}]}, {"justification": "Even though the infant-formula brand and stage match, the replacement explicitly fails the general price limit, so it is a failed match and must not be approved for picking.", "target": "false", "when": [{"atom_id": "ordered_category", "state": "supported"}, {"atom_id": "brand_match", "state": "supported"}, {"atom_id": "stage_match", "state": "supported"}, {"atom_id": "price_within_limit", "state": "refuted"}]}]}, "verified_pair": {"left": "The unavailable ordered item in Lena’s pickup order is priced at $14.00.", "negative_left": "The unavailable ordered item in Lena’s pickup order is priced at $14.00.", "negative_right": "The proposed replacement item is priced at $17.00.", "right": "The proposed replacement item is priced at $15.00."}, "verifier_independent_model": false}, "family": "fast-43-diverse-022-058", "id": "fast-43-diverse-022-058-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not approve the replacement for picking; route it to customer contact or supervisor review as policy requires.", "true": "Yes — approve the replacement and route the order to picking."}, "instructions": "Decide whether the proposed substitution is approved and ready to be routed to picking. Answer yes or no using the descriptions provided.", "type": "noul"}}, "state": "Pickup customer Lena ordered one 20-ounce can of BrightStart Stage 2 infant formula. The item is unavailable in-store. The store picker identifies a replacement: one 22-ounce can of BrightStart Stage 2 infant formula, matching both the brand and the stage of the original item. The unavailable ordered item in Lena’s pickup order is priced at $14.00. The proposed replacement item is priced at $17.00. Lena’s order note says, “Substitutions are allowed across this order if the replacement costs no more than 10% extra; for infant formula, use the same brand and stage only.” Fulfillment policy says an explicit category exception narrows the general permission. A replacement meeting both the general price limit and the exception may be sent to picking. Failed matches go to customer contact, while only ambiguous cases require supervisor review."}, "method": "c2d", "provenance": {"source_id": "diverse-022", "source_is_synthetic": true, "source_sha256": "118fab354fbb593d3ab53fbfea1930be43a7d23439ae717cffa9bdb20e56a311", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-04", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep identical governing policy, entities and bindings, only the registry membership fact changes coherently, the two evidence spans are plain factual sentences, and no answer or rationale is leaked.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the two universal registry relations. The focus atom is the factual relation of SKU membership, not a policy conclusion. The base and counter assignments are jointly realizable and differ only on that focus: registry membership makes the item a gift card in the base case, while nonmembership plus registry exhaustiveness makes it a non-gift-card item in the countercase. The policy evidence correctly preserves the substantive program and promotion rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The sole item’s SKU is in the registry, and every registered SKU identifies a gift card. Therefore the transaction is a gift-card purchase, the stated exclusion applies, and the promotion cannot create points because it only doubles transactions that first qualify for base points. Priya is not entitled to the claimed 200 points.", "rule_index": 0, "sound": true}, {"reason": "Refutation of registry membership means the sole item’s SKU is not registered. Because every gift-card SKU is stated to appear in the registry, the item cannot be a gift card. Atom a5 establishes eligibility apart from that exclusion; the sole $100 item therefore earns 100 base points and 200 during the double-points weekend. With zero points awarded and a claim for 200, the claim is valid.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya Shah's missing-points claim requests 200 Cedar Circle points for transaction T1."}, {"id": "a2", "statement": "Receipt R1 documents transaction T1."}, {"id": "a3", "statement": "The total price shown on receipt R1 is $100."}, {"id": "a4", "statement": "Receipt R1 contains exactly one line item."}, {"id": "a5", "statement": "The sole line item on receipt R1 qualifies for base points apart from the gift-card exclusion."}, {"id": "a6", "statement": "The SKU of the sole line item on receipt R1 is a member of the program's gift-card SKU registry."}, {"id": "a7", "statement": "Every SKU in the program's gift-card SKU registry identifies a gift-card product."}, {"id": "a8", "statement": "Every gift-card product SKU appears in the program's gift-card SKU registry."}, {"id": "a9", "statement": "Transaction T1 occurred during the double-points weekend."}, {"id": "a10", "statement": "Priya Shah's account received zero Cedar Circle points for transaction T1."}], "base_state_json": "\"Case Note - Missing Points Claim\\n\\nMember Priya Shah filed a claim for 200 missing Cedar Circle points tied to transaction T1, documented on receipt R1. Receipt R1 shows a total price of $100 and contains exactly one line item, which otherwise qualifies for base points apart from any gift-card exclusion. Transaction T1 occurred during the advertised double-points weekend. Account records confirm Priya's account received zero Cedar Circle points for transaction T1. Program rules state that every SKU in the program's gift-card SKU registry identifies a gift-card product, and conversely every gift-card product SKU appears in that registry. The sole line item on receipt R1 has SKU code GC-4471. The program's gift-card SKU registry lists SKU codes GC-4471 and GC-5820 as its full membership.\\n\\nProgram rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on \\u201call purchases,\\u201d yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "The sole line item on receipt R1 has SKU code GC-4471."}, {"path": [], "text": "The program's gift-card SKU registry lists SKU codes GC-4471 and GC-5820 as its full membership."}], "policy_evidence": [{"path": [], "text": "Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases."}, {"path": [], "text": "The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}], "rules": [{"justification": "The sole item on the claimed $100 transaction has a SKU in the gift-card registry, and every registry SKU identifies a gift-card product. The gift-card exclusion therefore applies, and the promotion cannot override it.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "The exhaustive registry contains every gift-card SKU, while the sole item's SKU is not in that registry, so the gift-card exclusion does not apply. The otherwise base-eligible $100 transaction earns 100 base points and, because it occurred during the double-points weekend, 200 total points; zero were awarded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The sole line item on receipt R1 has SKU code GC-4471.", "negative_left": "The sole line item on receipt R1 has SKU code GC-4471.", "negative_right": "The program's gift-card SKU registry lists SKU codes GC-9902 and GC-5820 as its full membership.", "right": "The program's gift-card SKU registry lists SKU codes GC-4471 and GC-5820 as its full membership."}, "verifier_independent_model": false}, "family": "fast-43-diverse-027-006", "id": "fast-43-diverse-027-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the gift-card exclusion applies, so Priya was not entitled to the claimed points.", "true": "Yes — the records and applicable rules show that Priya should have received the claimed 200 points."}, "instructions": "Decide whether Priya's claim for 200 missing points is valid under the supplied program and promotion rules. Answer yes or no.", "type": "noul"}}, "state": "Case Note - Missing Points Claim\n\nMember Priya Shah filed a claim for 200 missing Cedar Circle points tied to transaction T1, documented on receipt R1. Receipt R1 shows a total price of $100 and contains exactly one line item, which otherwise qualifies for base points apart from any gift-card exclusion. Transaction T1 occurred during the advertised double-points weekend. Account records confirm Priya's account received zero Cedar Circle points for transaction T1. Program rules state that every SKU in the program's gift-card SKU registry identifies a gift-card product, and conversely every gift-card product SKU appears in that registry. The sole line item on receipt R1 has SKU code GC-4471. The program's gift-card SKU registry lists SKU codes GC-4471 and GC-5820 as its full membership.\n\nProgram rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}, "method": "c2d", "provenance": {"source_id": "diverse-027", "source_is_synthetic": true, "source_sha256": "24a57454b60c3c253b20a328c3963efba4da8dd26a29c30742ff6568152211d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep identical governing policy, entities and bindings, only the registry membership fact changes coherently, the two evidence spans are plain factual sentences, and no answer or rationale is leaked.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, including the two universal registry relations. The focus atom is the factual relation of SKU membership, not a policy conclusion. The base and counter assignments are jointly realizable and differ only on that focus: registry membership makes the item a gift card in the base case, while nonmembership plus registry exhaustiveness makes it a non-gift-card item in the countercase. The policy evidence correctly preserves the substantive program and promotion rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The sole item’s SKU is in the registry, and every registered SKU identifies a gift card. Therefore the transaction is a gift-card purchase, the stated exclusion applies, and the promotion cannot create points because it only doubles transactions that first qualify for base points. Priya is not entitled to the claimed 200 points.", "rule_index": 0, "sound": true}, {"reason": "Refutation of registry membership means the sole item’s SKU is not registered. Because every gift-card SKU is stated to appear in the registry, the item cannot be a gift card. Atom a5 establishes eligibility apart from that exclusion; the sole $100 item therefore earns 100 base points and 200 during the double-points weekend. With zero points awarded and a claim for 200, the claim is valid.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya Shah's missing-points claim requests 200 Cedar Circle points for transaction T1."}, {"id": "a2", "statement": "Receipt R1 documents transaction T1."}, {"id": "a3", "statement": "The total price shown on receipt R1 is $100."}, {"id": "a4", "statement": "Receipt R1 contains exactly one line item."}, {"id": "a5", "statement": "The sole line item on receipt R1 qualifies for base points apart from the gift-card exclusion."}, {"id": "a6", "statement": "The SKU of the sole line item on receipt R1 is a member of the program's gift-card SKU registry."}, {"id": "a7", "statement": "Every SKU in the program's gift-card SKU registry identifies a gift-card product."}, {"id": "a8", "statement": "Every gift-card product SKU appears in the program's gift-card SKU registry."}, {"id": "a9", "statement": "Transaction T1 occurred during the double-points weekend."}, {"id": "a10", "statement": "Priya Shah's account received zero Cedar Circle points for transaction T1."}], "base_state_json": "\"Case Note - Missing Points Claim\\n\\nMember Priya Shah filed a claim for 200 missing Cedar Circle points tied to transaction T1, documented on receipt R1. Receipt R1 shows a total price of $100 and contains exactly one line item, which otherwise qualifies for base points apart from any gift-card exclusion. Transaction T1 occurred during the advertised double-points weekend. Account records confirm Priya's account received zero Cedar Circle points for transaction T1. Program rules state that every SKU in the program's gift-card SKU registry identifies a gift-card product, and conversely every gift-card product SKU appears in that registry. The sole line item on receipt R1 has SKU code GC-4471. The program's gift-card SKU registry lists SKU codes GC-4471 and GC-5820 as its full membership.\\n\\nProgram rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on \\u201call purchases,\\u201d yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": [], "text": "The sole line item on receipt R1 has SKU code GC-4471."}, {"path": [], "text": "The program's gift-card SKU registry lists SKU codes GC-4471 and GC-5820 as its full membership."}], "policy_evidence": [{"path": [], "text": "Program rules award one base point per dollar on eligible merchandise but exclude gift-card purchases."}, {"path": [], "text": "The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}], "rules": [{"justification": "The sole item on the claimed $100 transaction has a SKU in the gift-card registry, and every registry SKU identifies a gift-card product. The gift-card exclusion therefore applies, and the promotion cannot override it.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "The exhaustive registry contains every gift-card SKU, while the sole item's SKU is not in that registry, so the gift-card exclusion does not apply. The otherwise base-eligible $100 transaction earns 100 base points and, because it occurred during the double-points weekend, 200 total points; zero were awarded.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The sole line item on receipt R1 has SKU code GC-4471.", "negative_left": "The sole line item on receipt R1 has SKU code GC-4471.", "negative_right": "The program's gift-card SKU registry lists SKU codes GC-9902 and GC-5820 as its full membership.", "right": "The program's gift-card SKU registry lists SKU codes GC-4471 and GC-5820 as its full membership."}, "verifier_independent_model": false}, "family": "fast-43-diverse-027-006", "id": "fast-43-diverse-027-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the gift-card exclusion applies, so Priya was not entitled to the claimed points.", "true": "Yes — the records and applicable rules show that Priya should have received the claimed 200 points."}, "instructions": "Decide whether Priya's claim for 200 missing points is valid under the supplied program and promotion rules. Answer yes or no.", "type": "noul"}}, "state": "Case Note - Missing Points Claim\n\nMember Priya Shah filed a claim for 200 missing Cedar Circle points tied to transaction T1, documented on receipt R1. Receipt R1 shows a total price of $100 and contains exactly one line item, which otherwise qualifies for base points apart from any gift-card exclusion. Transaction T1 occurred during the advertised double-points weekend. Account records confirm Priya's account received zero Cedar Circle points for transaction T1. Program rules state that every SKU in the program's gift-card SKU registry identifies a gift-card product, and conversely every gift-card product SKU appears in that registry. The sole line item on receipt R1 has SKU code GC-4471. The program's gift-card SKU registry lists SKU codes GC-9902 and GC-5820 as its full membership.\n\nProgram rules award one base point per dollar on eligible merchandise but exclude gift-card purchases. The promotion advertised double points on “all purchases,” yet its scope clause limits doubling to transactions that first qualify for base points and does not override listed exclusions."}, "method": "c2d", "provenance": {"source_id": "diverse-027", "source_is_synthetic": true, "source_sha256": "24a57454b60c3c253b20a328c3963efba4da8dd26a29c30742ff6568152211d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual changes only the authorization status of AUTH-88 from approved to denied, altering the authorized-redemption count while leaving all other bindings and policy language intact, and the two evidence sentences remain plain factual statements without leaking any rule table, code, or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\": \"Loyalty member\", \"text\": \"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"}, {\"speaker\": \"Rewards program analyst\", \"text\": \"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"}, {\"speaker\": \"Audit clerk\", \"text\": \"The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700.\"}, {\"speaker\": \"Audit clerk\", \"text\": \"Entry AUTH-88 in the authorization request log is marked as approved by the member.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"Gold status does not change the earning rate, and pending promotional points are excluded until posted. R-700 remains the only redemption posted this period.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700."}, {"path": ["4", "text"], "text": "Entry AUTH-88 in the authorization request log is marked as approved by the member."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700.", "negative_left": "The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700.", "negative_right": "Entry AUTH-88 in the authorization request log is marked as denied by the member.", "right": "Entry AUTH-88 in the authorization request log is marked as approved by the member."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-044", "id": "fast-43-diverse-029-044-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Audit clerk", "text": "The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700."}, {"speaker": "Audit clerk", "text": "Entry AUTH-88 in the authorization request log is marked as approved by the member."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. R-700 remains the only redemption posted this period."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual changes only the authorization status of AUTH-88 from approved to denied, altering the authorized-redemption count while leaving all other bindings and policy language intact, and the two evidence sentences remain plain factual statements without leaking any rule table, code, or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\": \"Loyalty member\", \"text\": \"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"}, {\"speaker\": \"Rewards program analyst\", \"text\": \"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"}, {\"speaker\": \"Audit clerk\", \"text\": \"The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700.\"}, {\"speaker\": \"Audit clerk\", \"text\": \"Entry AUTH-88 in the authorization request log is marked as approved by the member.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"Gold status does not change the earning rate, and pending promotional points are excluded until posted. R-700 remains the only redemption posted this period.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700."}, {"path": ["4", "text"], "text": "Entry AUTH-88 in the authorization request log is marked as approved by the member."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700.", "negative_left": "The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700.", "negative_right": "Entry AUTH-88 in the authorization request log is marked as denied by the member.", "right": "Entry AUTH-88 in the authorization request log is marked as approved by the member."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-044", "id": "fast-43-diverse-029-044-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Audit clerk", "text": "The member's authorization request log contains entry AUTH-88 corresponding to reward-redemption transaction R-700."}, {"speaker": "Audit clerk", "text": "Entry AUTH-88 in the authorization request log is marked as denied by the member."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. R-700 remains the only redemption posted this period."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy language via the unchanged question, keep the same entity K-184 and claim-period bindings, and the two audit sentences are factual, non-instructional descriptions of an authorization record and its resolution; the counterfactual coherently varies the approval outcome (confirmed vs. declined) without duplicating or contradicting other posted figures or leaking the scoring answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"},{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Customer service representative\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Only one redemption, R-700, posted during the claim period, valued at exactly 700 points.\"},{\"speaker\":\"Account auditor\",\"text\":\"The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account.\"},{\"speaker\":\"Account auditor\",\"text\":\"Authorization request AR-4471 was resolved with the member's approval confirmed via the account app.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["4", "text"], "text": "The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account."}, {"path": ["5", "text"], "text": "Authorization request AR-4471 was resolved with the member's approval confirmed via the account app."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account.", "negative_left": "The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account.", "negative_right": "Authorization request AR-4471 was resolved with the member's approval declined via the account app.", "right": "Authorization request AR-4471 was resolved with the member's approval confirmed via the account app."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-079", "id": "fast-43-diverse-029-079-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Only one redemption, R-700, posted during the claim period, valued at exactly 700 points."}, {"speaker": "Account auditor", "text": "The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account."}, {"speaker": "Account auditor", "text": "Authorization request AR-4471 was resolved with the member's approval confirmed via the account app."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy language via the unchanged question, keep the same entity K-184 and claim-period bindings, and the two audit sentences are factual, non-instructional descriptions of an authorization record and its resolution; the counterfactual coherently varies the approval outcome (confirmed vs. declined) without duplicating or contradicting other posted figures or leaking the scoring answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"},{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Customer service representative\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Only one redemption, R-700, posted during the claim period, valued at exactly 700 points.\"},{\"speaker\":\"Account auditor\",\"text\":\"The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account.\"},{\"speaker\":\"Account auditor\",\"text\":\"Authorization request AR-4471 was resolved with the member's approval confirmed via the account app.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["4", "text"], "text": "The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account."}, {"path": ["5", "text"], "text": "Authorization request AR-4471 was resolved with the member's approval confirmed via the account app."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account.", "negative_left": "The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account.", "negative_right": "Authorization request AR-4471 was resolved with the member's approval declined via the account app.", "right": "Authorization request AR-4471 was resolved with the member's approval confirmed via the account app."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-079", "id": "fast-43-diverse-029-079-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Only one redemption, R-700, posted during the claim period, valued at exactly 700 points."}, {"speaker": "Account auditor", "text": "The account activity log lists an authorization request record, request ID AR-4471, corresponding to reward-redemption transaction R-700 on the member's loyalty account."}, {"speaker": "Account auditor", "text": "Authorization request AR-4471 was resolved with the member's approval declined via the account app."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy language and question bindings verbatim, the focus evidence consists of two complete factual sentences, the counterfactual's approved-to-denied change is a coherent single-fact alteration testing the authorization criterion, and neither context reveals a gold answer, score, or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\": \"Loyalty member\", \"text\": \"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"}, {\"speaker\": \"Rewards program analyst\", \"text\": \"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"}, {\"speaker\": \"Audit clerk\", \"text\": \"Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"Gold status does not change the earning rate, and pending promotional points are excluded until posted. R-700 is the only redemption posted this period, in the amount of 700 points.\"}, {\"speaker\": \"Audit clerk\", \"text\": \"Authorization-log entry AUTH-8823 is marked as approved by the account holder.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823."}, {"path": ["5", "text"], "text": "Authorization-log entry AUTH-8823 is marked as approved by the account holder."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823.", "negative_left": "Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823.", "negative_right": "Authorization-log entry AUTH-8823 is marked as denied by the account holder.", "right": "Authorization-log entry AUTH-8823 is marked as approved by the account holder."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-085", "id": "fast-43-diverse-029-085-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Audit clerk", "text": "Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. R-700 is the only redemption posted this period, in the amount of 700 points."}, {"speaker": "Audit clerk", "text": "Authorization-log entry AUTH-8823 is marked as approved by the account holder."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy language and question bindings verbatim, the focus evidence consists of two complete factual sentences, the counterfactual's approved-to-denied change is a coherent single-fact alteration testing the authorization criterion, and neither context reveals a gold answer, score, or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\": \"Loyalty member\", \"text\": \"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"}, {\"speaker\": \"Rewards program analyst\", \"text\": \"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"}, {\"speaker\": \"Audit clerk\", \"text\": \"Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"Gold status does not change the earning rate, and pending promotional points are excluded until posted. R-700 is the only redemption posted this period, in the amount of 700 points.\"}, {\"speaker\": \"Audit clerk\", \"text\": \"Authorization-log entry AUTH-8823 is marked as approved by the account holder.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823."}, {"path": ["5", "text"], "text": "Authorization-log entry AUTH-8823 is marked as approved by the account holder."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823.", "negative_left": "Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823.", "negative_right": "Authorization-log entry AUTH-8823 is marked as denied by the account holder.", "right": "Authorization-log entry AUTH-8823 is marked as approved by the account holder."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-085", "id": "fast-43-diverse-029-085-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Audit clerk", "text": "Reward-redemption transaction R-700 corresponds to authorization-log entry AUTH-8823."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. R-700 is the only redemption posted this period, in the amount of 700 points."}, {"speaker": "Audit clerk", "text": "Authorization-log entry AUTH-8823 is marked as denied by the account holder."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Governing policy (earning/redemption rules, exclusions) is preserved via the unchanged question and consistent auditor statement, order K-184 and claim-period bindings are unchanged, the two evidence sentences are factual observations about the ledger and code table rather than policy text, the counterfactual coherently swaps only the authorization-code meaning without contradicting other unchanged facts, and no gold answer or rule table entry is exposed.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"},{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Account systems auditor\",\"text\":\"Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7. In the loyalty system's authorization code table, code A7 denotes 'member-authorized redemption'. No other redemption transactions were posted during the claim period, and Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7."}, {"path": ["3", "text"], "text": "In the loyalty system's authorization code table, code A7 denotes 'member-authorized redemption'."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7.", "negative_left": "Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7.", "negative_right": "In the loyalty system's authorization code table, code A7 denotes 'system-initiated redemption, no member action'.", "right": "In the loyalty system's authorization code table, code A7 denotes 'member-authorized redemption'."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-091", "id": "fast-43-diverse-029-091-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Account systems auditor", "text": "Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7. In the loyalty system's authorization code table, code A7 denotes 'member-authorized redemption'. No other redemption transactions were posted during the claim period, and Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Governing policy (earning/redemption rules, exclusions) is preserved via the unchanged question and consistent auditor statement, order K-184 and claim-period bindings are unchanged, the two evidence sentences are factual observations about the ledger and code table rather than policy text, the counterfactual coherently swaps only the authorization-code meaning without contradicting other unchanged facts, and no gold answer or rule table entry is exposed.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"},{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Account systems auditor\",\"text\":\"Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7. In the loyalty system's authorization code table, code A7 denotes 'member-authorized redemption'. No other redemption transactions were posted during the claim period, and Gold status does not change the earning rate, and pending promotional points are excluded until posted.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7."}, {"path": ["3", "text"], "text": "In the loyalty system's authorization code table, code A7 denotes 'member-authorized redemption'."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7.", "negative_left": "Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7.", "negative_right": "In the loyalty system's authorization code table, code A7 denotes 'system-initiated redemption, no member action'.", "right": "In the loyalty system's authorization code table, code A7 denotes 'member-authorized redemption'."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-091", "id": "fast-43-diverse-029-091-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Account systems auditor", "text": "Reward-redemption transaction R-700 appears in the account activity log for order K-184 with an authorization field code of A7. In the loyalty system's authorization code table, code A7 denotes 'system-initiated redemption, no member action'. No other redemption transactions were posted during the claim period, and Gold status does not change the earning rate, and pending promotional points are excluded until posted."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (earning/redemption rules, exclusions) and original entity/time bindings (K-184, AL-58, R-700, $450, 1750/2000 balances); the two focus sentences are factual statements, not policy text; the counterfactual swap from 'approved' to 'denied' coherently varies the authorization fact without contradicting other stated data; no gold answer, level, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Loyalty systems auditor\",\"text\":\"Only one redemption, R-700 for 700 points, was posted during the claim period. Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account.\"},{\"speaker\":\"Compliance reviewer\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Authorization request log entry AL-58 is marked as approved by the member.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["2", "text"], "text": "Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account."}, {"path": ["3", "text"], "text": "Authorization request log entry AL-58 is marked as approved by the member."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account.", "negative_left": "Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account.", "negative_right": "Authorization request log entry AL-58 is marked as denied by the member.", "right": "Authorization request log entry AL-58 is marked as approved by the member."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-092", "id": "fast-43-diverse-029-092-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Loyalty systems auditor", "text": "Only one redemption, R-700 for 700 points, was posted during the claim period. Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account."}, {"speaker": "Compliance reviewer", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Authorization request log entry AL-58 is marked as approved by the member."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (earning/redemption rules, exclusions) and original entity/time bindings (K-184, AL-58, R-700, $450, 1750/2000 balances); the two focus sentences are factual statements, not policy text; the counterfactual swap from 'approved' to 'denied' coherently varies the authorization fact without contradicting other stated data; no gold answer, level, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Loyalty systems auditor\",\"text\":\"Only one redemption, R-700 for 700 points, was posted during the claim period. Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account.\"},{\"speaker\":\"Compliance reviewer\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Authorization request log entry AL-58 is marked as approved by the member.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["2", "text"], "text": "Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account."}, {"path": ["3", "text"], "text": "Authorization request log entry AL-58 is marked as approved by the member."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account.", "negative_left": "Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account.", "negative_right": "Authorization request log entry AL-58 is marked as denied by the member.", "right": "Authorization request log entry AL-58 is marked as approved by the member."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-092", "id": "fast-43-diverse-029-092-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Loyalty systems auditor", "text": "Only one redemption, R-700 for 700 points, was posted during the claim period. Reward-redemption transaction R-700 corresponds to authorization request log entry AL-58 in the member's loyalty account."}, {"speaker": "Compliance reviewer", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Authorization request log entry AL-58 is marked as denied by the member."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy content and question object unchanged, replace the vague authorization confirmation with a concrete authorization-log entry, and the counterfactual's DENIED status is a coherent case-level change (testing whether an authorized-redemption requirement is met) without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"},{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Customer service representative\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Only one redemption transaction, R-700, posted during the claim period, and its amount was 700 points. The authorization log entry for request Q-88 shows a status of APPROVED, entered by the member.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184."}, {"path": ["4", "text"], "text": "The authorization log entry for request Q-88 shows a status of APPROVED, entered by the member."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184.", "negative_left": "Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184.", "negative_right": "The authorization log entry for request Q-88 shows a status of DENIED, entered by the member.", "right": "The authorization log entry for request Q-88 shows a status of APPROVED, entered by the member."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-095", "id": "fast-43-diverse-029-095-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184."}, {"speaker": "Rewards program analyst", "text": "Only one redemption transaction, R-700, posted during the claim period, and its amount was 700 points. The authorization log entry for request Q-88 shows a status of APPROVED, entered by the member."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy content and question object unchanged, replace the vague authorization confirmation with a concrete authorization-log entry, and the counterfactual's DENIED status is a coherent case-level change (testing whether an authorized-redemption requirement is met) without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"},{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Customer service representative\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Only one redemption transaction, R-700, posted during the claim period, and its amount was 700 points. The authorization log entry for request Q-88 shows a status of APPROVED, entered by the member.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184."}, {"path": ["4", "text"], "text": "The authorization log entry for request Q-88 shows a status of APPROVED, entered by the member."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184.", "negative_left": "Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184.", "negative_right": "The authorization log entry for request Q-88 shows a status of DENIED, entered by the member.", "right": "The authorization log entry for request Q-88 shows a status of APPROVED, entered by the member."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-095", "id": "fast-43-diverse-029-095-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Reward-redemption transaction R-700 is linked to authorization request Q-88 on the member's loyalty account for order K-184."}, {"speaker": "Rewards program analyst", "text": "Only one redemption transaction, R-700, posted during the claim period, and its amount was 700 points. The authorization log entry for request Q-88 shows a status of DENIED, entered by the member."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question verbatim and only vary the auditor's factual authorization-status sentence, which is a plausible independent fact not contradicting other stated data, while the two evidence sentences are complete factual statements with no embedded rules or labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"},{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Customer service representative\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Only one redemption, R-700 for 700 points, posted this period.\"},{\"speaker\":\"Compliance auditor\",\"text\":\"An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700.\"},{\"speaker\":\"Compliance auditor\",\"text\":\"Authorization request AR-55 shows a status of approved by the member on file.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["4", "text"], "text": "An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700."}, {"path": ["5", "text"], "text": "Authorization request AR-55 shows a status of approved by the member on file."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700.", "negative_left": "An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700.", "negative_right": "Authorization request AR-55 shows a status of denied by the member on file.", "right": "Authorization request AR-55 shows a status of approved by the member on file."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-107", "id": "fast-43-diverse-029-107-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Only one redemption, R-700 for 700 points, posted this period."}, {"speaker": "Compliance auditor", "text": "An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700."}, {"speaker": "Compliance auditor", "text": "Authorization request AR-55 shows a status of approved by the member on file."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question verbatim and only vary the auditor's factual authorization-status sentence, which is a plausible independent fact not contradicting other stated data, while the two evidence sentences are complete factual statements with no embedded rules or labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\":\"Loyalty member\",\"text\":\"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"},{\"speaker\":\"Customer service representative\",\"text\":\"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"},{\"speaker\":\"Rewards program analyst\",\"text\":\"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"},{\"speaker\":\"Customer service representative\",\"text\":\"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Only one redemption, R-700 for 700 points, posted this period.\"},{\"speaker\":\"Compliance auditor\",\"text\":\"An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700.\"},{\"speaker\":\"Compliance auditor\",\"text\":\"Authorization request AR-55 shows a status of approved by the member on file.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["4", "text"], "text": "An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700."}, {"path": ["5", "text"], "text": "Authorization request AR-55 shows a status of approved by the member on file."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700.", "negative_left": "An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700.", "negative_right": "Authorization request AR-55 shows a status of denied by the member on file.", "right": "Authorization request AR-55 shows a status of approved by the member on file."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-107", "id": "fast-43-diverse-029-107-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Only one redemption, R-700 for 700 points, posted this period."}, {"speaker": "Compliance auditor", "text": "An authorization request record labeled AR-55 in the loyalty system corresponds to reward-redemption transaction R-700."}, {"speaker": "Compliance auditor", "text": "Authorization request AR-55 shows a status of denied by the member on file."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original question's policy for computing net missing points via opening balance, eligible earnings, and authorized redemptions, using the same K-184/A-552/R-700 bindings; the two evidence spans are factual sentences about the log field and the signer, not policy text; the counterfactual coherently swaps the signer identity in a single sentence without duplicating or contradicting other measurements (opening balance, $450 merchandise, 700-point redemption amount remain fixed); no answer, level, or rule table is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\": \"Loyalty member\", \"text\": \"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"}, {\"speaker\": \"Rewards program analyst\", \"text\": \"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552.\"}, {\"speaker\": \"Compliance auditor\", \"text\": \"I pulled the approval workflow. Approval request A-552 was digitally signed by the member on record for order K-184. No other redemption transactions were posted during this claim period, and R-700's amount is confirmed as exactly 700 points.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552."}, {"path": ["4", "text"], "text": "Approval request A-552 was digitally signed by the member on record for order K-184."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552.", "negative_left": "Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552.", "negative_right": "Approval request A-552 was digitally signed by an unidentified third party, not the member on record for order K-184.", "right": "Approval request A-552 was digitally signed by the member on record for order K-184."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-130", "id": "fast-43-diverse-029-130-base", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552."}, {"speaker": "Compliance auditor", "text": "I pulled the approval workflow. Approval request A-552 was digitally signed by the member on record for order K-184. No other redemption transactions were posted during this claim period, and R-700's amount is confirmed as exactly 700 points."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "commerce-05", "split": "train", "variant": "base"} {"domain": "commerce", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original question's policy for computing net missing points via opening balance, eligible earnings, and authorized redemptions, using the same K-184/A-552/R-700 bindings; the two evidence spans are factual sentences about the log field and the signer, not policy text; the counterfactual coherently swaps the signer identity in a single sentence without duplicating or contradicting other measurements (opening balance, $450 merchandise, 700-point redemption amount remain fixed); no answer, level, or rule table is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "full_context_fact_states": {"base": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "supported", "redemption_r700_posted": "supported"}, "counterfactual": {"actual_balance_1750": "supported", "eligible_posted_earnings_450": "supported", "opening_balance_2000": "supported", "posted_redemption_count_1": "supported", "redemption_r700_amount_700": "supported", "redemption_r700_authorized": "refuted", "redemption_r700_posted": "supported"}, "remove_left": {"redemption_r700_authorized": "unknown"}, "remove_right": {"redemption_r700_authorized": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"redemption_r700_authorized": "unknown"}, "negative_pair": {"redemption_r700_authorized": "refuted"}, "negative_sentence": {"redemption_r700_authorized": "unknown"}, "positive_pair": {"redemption_r700_authorized": "supported"}, "right": {"redemption_r700_authorized": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus concerns the factual authorization status of a particular redemption rather than a policy conclusion. The base and counter assignments are jointly realizable with only that authorization fact changing. The policy evidence appropriately preserves relevant state-originating program rules; the unchanged questions already preserve the scoring thresholds, calculation method, exclusions, and task instructions. Both conjunctions exclude competing redemption possibilities through the exact-one-posted-redemption atom and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish total posted eligible earnings of 450 and exactly one posted redemption: R-700, worth 700 points and authorized. Thus expected balance is 2,000 + 450 - 700 = 1,750. With actual balance 1,750, net missing points are 0, which entails Level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that R-700 is the sole posted redemption and that it was not authorized, so there are no authorized posted redemptions to subtract. Expected balance is 2,000 + 450 = 2,450. With actual balance 1,750, net missing points are 700, which entails Level 3.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "opening_balance_2000", "statement": "The opening balance of the member's loyalty account at the start of the claim period for order K-184 is exactly 2,000 points."}, {"id": "eligible_posted_earnings_450", "statement": "The total eligible earnings posted to the member's loyalty account during the claim period for order K-184 are exactly 450 points."}, {"id": "posted_redemption_count_1", "statement": "Exactly one reward-redemption transaction was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_posted", "statement": "Reward-redemption transaction R-700 was posted to the member's loyalty account during the claim period for order K-184."}, {"id": "redemption_r700_amount_700", "statement": "The point amount of reward-redemption transaction R-700 is exactly 700 points."}, {"id": "redemption_r700_authorized", "statement": "The member authorized reward-redemption transaction R-700."}, {"id": "actual_balance_1750", "statement": "The actual balance of the member's loyalty account at the end of the claim period for order K-184 is exactly 1,750 points."}], "base_state_json": "[{\"speaker\": \"Loyalty member\", \"text\": \"My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption.\"}, {\"speaker\": \"Rewards program analyst\", \"text\": \"Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order.\"}, {\"speaker\": \"Customer service representative\", \"text\": \"Gold status does not change the earning rate, and pending promotional points are excluded until posted. Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552.\"}, {\"speaker\": \"Compliance auditor\", \"text\": \"I pulled the approval workflow. Approval request A-552 was digitally signed by the member on record for order K-184. No other redemption transactions were posted during this claim period, and R-700's amount is confirmed as exactly 700 points.\"}]", "base_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "counter_states": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}], "focus_atom": "redemption_r700_authorized", "focus_evidence": [{"path": ["3", "text"], "text": "Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552."}, {"path": ["4", "text"], "text": "Approval request A-552 was digitally signed by the member on record for order K-184."}], "policy_evidence": [{"path": ["2", "text"], "text": "Policy awards one point per eligible merchandise dollar; gift cards earn none."}, {"path": ["3", "text"], "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted."}], "rules": [{"justification": "R-700 is the sole posted redemption and is authorized, so expected balance is 2,000 + 450 - 700 = 1,750. The actual balance is 1,750, making net missing points exactly 0 and requiring Level 0.", "target": "0", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "supported"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}, {"justification": "R-700 is the sole posted redemption but is explicitly unauthorized, so no posted redemption may be subtracted. Expected balance is 2,000 + 450 = 2,450; against an actual balance of 1,750, net missing points are 700, requiring Level 3.", "target": "3", "when": [{"atom_id": "opening_balance_2000", "state": "supported"}, {"atom_id": "eligible_posted_earnings_450", "state": "supported"}, {"atom_id": "posted_redemption_count_1", "state": "supported"}, {"atom_id": "redemption_r700_posted", "state": "supported"}, {"atom_id": "redemption_r700_amount_700", "state": "supported"}, {"atom_id": "redemption_r700_authorized", "state": "refuted"}, {"atom_id": "actual_balance_1750", "state": "supported"}]}]}, "verified_pair": {"left": "Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552.", "negative_left": "Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552.", "negative_right": "Approval request A-552 was digitally signed by an unidentified third party, not the member on record for order K-184.", "right": "Approval request A-552 was digitally signed by the member on record for order K-184."}, "verifier_independent_model": false}, "family": "fast-43-diverse-029-130", "id": "fast-43-diverse-029-130-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0: No verified discrepancy; net missing points equal exactly 0.", "Level 1: Minor verified discrepancy; net missing points are 1 through 99.", "Level 2: Moderate verified discrepancy; net missing points are 100 through 499.", "Level 3: Major verified discrepancy; net missing points are 500 or more."], "instructions": "Verify the net missing-point discrepancy. Compute expected balance as opening balance plus eligible posted earnings minus authorized posted redemptions. Exclude gift-card spending, tier status, and pending promotions. Net missing points equal max(expected balance minus actual balance, 0). Select the matching magnitude level.", "type": "score"}}, "state": [{"speaker": "Loyalty member", "text": "My balance should be higher after order K-184. I spent $500 and expected 500 points. I also reached Gold last week and have a 300-point birthday offer pending."}, {"speaker": "Customer service representative", "text": "The account opened the claim period with 2,000 points. It now shows 1,750 points after a posted 700-point reward redemption."}, {"speaker": "Rewards program analyst", "text": "Order K-184 contained $450 of eligible merchandise and a $50 gift card. Policy awards one point per eligible merchandise dollar; gift cards earn none. The ledger posted 450 points for the order."}, {"speaker": "Customer service representative", "text": "Gold status does not change the earning rate, and pending promotional points are excluded until posted. Order K-184's loyalty account log shows an authorization code field for reward-redemption transaction R-700, referencing approval request A-552."}, {"speaker": "Compliance auditor", "text": "I pulled the approval workflow. Approval request A-552 was digitally signed by an unidentified third party, not the member on record for order K-184. No other redemption transactions were posted during this claim period, and R-700's amount is confirmed as exactly 700 points."}]}, "method": "c2d", "provenance": {"source_id": "diverse-029", "source_is_synthetic": true, "source_sha256": "2f7703502000fafaf5f36c7ea6191704bf4895b0ba28e7f877e3a628484d740b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "commerce-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the reconciliation policy and the 19:00/S-2 bindings intact, the two evidence spans are verbatim factual sentences from the base context, the counterfactual coherently introduces a genuine scope conflict (out-of-scope vs in-scope) consistent with the unresolved-alarm and Inez-unavailable facts, and neither context reveals a gold label, rule table or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, outgoing shift lead Mara submits the handoff to incoming shift lead Dev. The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register, current at 19:00, also lists scrubber S-2 as in scope for the 19:00 handoff. These two records are the only documents governing whether S-2 is in scope. The applicable scope of every other handoff item is already established at 19:00. S-2 has an unresolved alarm at the 19:00 handoff, routed to the controls engineer without owner acceptance from the appropriate owner, and no temporary evidence exception covers this missing acceptance. Operations coordinator Inez, who is required to reconcile any discrepancy between the two records before readiness is rated, is unavailable and has not performed any reconciliation. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register, current at 19:00, also lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register, current at 19:00, lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register, current at 19:00, also lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-43-diverse-032-024", "id": "fast-43-diverse-032-024-base", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, outgoing shift lead Mara submits the handoff to incoming shift lead Dev. The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register, current at 19:00, also lists scrubber S-2 as in scope for the 19:00 handoff. These two records are the only documents governing whether S-2 is in scope. The applicable scope of every other handoff item is already established at 19:00. S-2 has an unresolved alarm at the 19:00 handoff, routed to the controls engineer without owner acceptance from the appropriate owner, and no temporary evidence exception covers this missing acceptance. Operations coordinator Inez, who is required to reconcile any discrepancy between the two records before readiness is rated, is unavailable and has not performed any reconciliation. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_1_not_ready"}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the reconciliation policy and the 19:00/S-2 bindings intact, the two evidence spans are verbatim factual sentences from the base context, the counterfactual coherently introduces a genuine scope conflict (out-of-scope vs in-scope) consistent with the unresolved-alarm and Inez-unavailable facts, and neither context reveals a gold label, rule table or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported", "a8": "supported"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally scoped atoms a7 and a8 remain atomic. The focus atom a2 is factual rather than normative. The base and counter assignments can be realized with only record agreement changing: in the base the complete records agree with the master record, while in the counter they conflict without reconciliation. The policy evidence correctly preserves the substantive state-originating rule that conflicting scope records require Inez's reconciliation and that neither record has priority; criteria and instructions already retained in the questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Agreement with the master record, together with completeness of the governing record set, establishes S-2 as in scope; a7 establishes the remaining scope. S-2 then has an unresolved alarm lacking owner acceptance and lacking a covering temporary evidence exception, which is sufficient for Level 1 and excludes the higher readiness levels.", "rule_index": 0, "sound": true}, {"reason": "The complete governing record set disagrees over S-2, the required coordinator reconciliation did not occur, and policy gives neither record precedence. S-2's applicable scope therefore remains unestablished, which is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The master scope record updated at 18:45 lists scrubber S-2 as in scope for the 19:00 handoff."}, {"id": "a2", "statement": "The master scope record and the operations-coordinator-approved exception register agree on whether scrubber S-2 is in scope for the 19:00 handoff."}, {"id": "a3", "statement": "Operations coordinator Inez reconciled any discrepancy between the master scope record and the exception register before the 19:00 readiness rating."}, {"id": "a4", "statement": "Scrubber S-2 has an unresolved alarm at the 19:00 handoff."}, {"id": "a5", "statement": "The appropriate owner accepted scrubber S-2's unresolved alarm before the 19:00 handoff."}, {"id": "a6", "statement": "A temporary evidence exception covers the missing owner acceptance for scrubber S-2's unresolved alarm at the 19:00 handoff."}, {"id": "a7", "statement": "The applicable scope of every handoff item other than scrubber S-2 is established at 19:00."}, {"id": "a8", "statement": "The master scope record and the exception register are the complete set of records governing whether scrubber S-2 is in scope for the 19:00 handoff."}], "base_state_json": "\"At 19:00, outgoing shift lead Mara submits the handoff to incoming shift lead Dev. The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register, current at 19:00, also lists scrubber S-2 as in scope for the 19:00 handoff. These two records are the only documents governing whether S-2 is in scope. The applicable scope of every other handoff item is already established at 19:00. S-2 has an unresolved alarm at the 19:00 handoff, routed to the controls engineer without owner acceptance from the appropriate owner, and no temporary evidence exception covers this missing acceptance. Operations coordinator Inez, who is required to reconcile any discrepancy between the two records before readiness is rated, is unavailable and has not performed any reconciliation. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a2", "focus_evidence": [{"path": [], "text": "The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff."}, {"path": [], "text": "The operations-coordinator-approved exception register, current at 19:00, also lists scrubber S-2 as in scope for the 19:00 handoff."}], "policy_evidence": [{"path": [], "text": "Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}], "rules": [{"justification": "The complete governing record set agrees with the master record that S-2 is in scope, while all other scope is established. S-2 has an unresolved alarm without owner acceptance and no covering evidence exception, which is sufficient for Level 1.", "target": "level_1_not_ready", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The complete governing record set conflicts over S-2, and Inez did not perform the reconciliation required before rating. Because neither record takes precedence, the applicable scope is not established.", "target": "none_of_above", "when": [{"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_left": "The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff.", "negative_right": "The operations-coordinator-approved exception register, current at 19:00, lists scrubber S-2 as out of scope for the 19:00 handoff.", "right": "The operations-coordinator-approved exception register, current at 19:00, also lists scrubber S-2 as in scope for the 19:00 handoff."}, "verifier_independent_model": false}, "family": "fast-43-diverse-032-024", "id": "fast-43-diverse-032-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1_not_ready": "Confirmed scope; at least one in-scope item lacks required evidence or owner acceptance and is not covered by a valid evidence exception.", "level_2_exception_ready": "Confirmed scope; the only evidence gap is covered by a valid coordinator-approved temporary evidence exception that names an owner and deadline.", "level_3_ready_with_follow_up": "Confirmed scope; every unresolved in-scope item has all required evidence, including acceptance by the appropriate owner, but at least one accepted follow-up remains open.", "level_4_complete": "Confirmed scope; every in-scope item has all required evidence, and no unresolved item awaits owner acceptance.", "none_of_above": "No substantive level can be assigned because the applicable scope has not been established under policy."}, "instructions": "Select the applicable handoff-readiness result. Use the ordered substantive scale from Level 4 (highest) to Level 1 (lowest). Required evidence for each in-scope unresolved item is current status, a timestamped artifact, the appropriate named owner, and owner acceptance. Apply an exception only when its scope is established under the stated policy.", "type": "choice"}}, "state": "At 19:00, outgoing shift lead Mara submits the handoff to incoming shift lead Dev. The master scope record, timestamped 18:45, lists scrubber S-2 as in scope for the 19:00 handoff. The operations-coordinator-approved exception register, current at 19:00, lists scrubber S-2 as out of scope for the 19:00 handoff. These two records are the only documents governing whether S-2 is in scope. The applicable scope of every other handoff item is already established at 19:00. S-2 has an unresolved alarm at the 19:00 handoff, routed to the controls engineer without owner acceptance from the appropriate owner, and no temporary evidence exception covers this missing acceptance. Operations coordinator Inez, who is required to reconcile any discrepancy between the two records before readiness is rated, is unavailable and has not performed any reconciliation. Policy says conflicting scope records must be reconciled by operations coordinator Inez before readiness is rated, and neither record takes precedence."}, "method": "c2d", "provenance": {"source_id": "diverse-032", "source_is_synthetic": true, "source_sha256": "4d94d9b9f35eeec683ce15ffc2a60ce5647f497b5ce9b6b52eb57344f0af5151", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy and question bindings, use complete factual sentences as evidence, and the counterfactual coherently swaps the acknowledger's name without contradicting other facts or leaking answer logic.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Outgoing shift lead\",\"text\":\"Line 4 pump vibration is the only unresolved production item in scope for this handoff. Status at 17:42: stable at 6.1 mm/s. Trend snapshot TS-144 is attached, timestamped. Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item. There is exactly one acknowledgment of the routing and next action on the ticket. The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Mara Chen. The completed filter swap is logged in work order W-219 with a 17:30 completion photo, and all other production items show current status and timestamped artifacts.\"},{\"speaker\":\"Incoming shift lead\",\"text\":\"Understood, I'll monitor the Line 4 trend until Mara reports back after the inspection.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["0", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"path": ["1", "text"], "text": "The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Mara Chen."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_right": "The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Tom Reyes.", "right": "The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Mara Chen."}, "verifier_independent_model": false}, "family": "fast-43-diverse-034-011", "id": "fast-43-diverse-034-011-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Outgoing shift lead", "text": "Line 4 pump vibration is the only unresolved production item in scope for this handoff. Status at 17:42: stable at 6.1 mm/s. Trend snapshot TS-144 is attached, timestamped. Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"speaker": "Operations coordinator", "text": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item. There is exactly one acknowledgment of the routing and next action on the ticket. The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Mara Chen. The completed filter swap is logged in work order W-219 with a 17:30 completion photo, and all other production items show current status and timestamped artifacts."}, {"speaker": "Incoming shift lead", "text": "Understood, I'll monitor the Line 4 trend until Mara reports back after the inspection."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy and question bindings, use complete factual sentences as evidence, and the counterfactual coherently swaps the acknowledger's name without contradicting other facts or leaking answer logic.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state factual relationships rather than final readiness classifications. The focus atom is the factual identity relation between the acknowledger and named owner. The base and counter assignments can differ only in that identity while preserving all other facts. Empty policy_evidence is correct because the governing readiness rules and facilities exclusion are already retained in the questions object; no additional substantive policy from the original state is needed.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes current status and timestamped artifacts for every in-scope production item and establishes that the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment by that same owner. This is sufficient for Ready under the stated criteria.", "rule_index": 0, "sound": true}, {"reason": "Given that there is exactly one unresolved item and exactly one acknowledgment of its routing and next action, refuting that the acknowledger is the named owner entails that the unresolved item lacks its owner’s acknowledgment. Status and artifact requirements remain satisfied, so the handoff is Conditional and therefore the answer to whether it is Ready is false.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every production item in scope for the depicted shift handoff has a current status."}, {"id": "a2", "statement": "Every production item in scope for the depicted shift handoff has a timestamped artifact."}, {"id": "a3", "statement": "The Line 4 pump vibration is the only unresolved production item in scope for the depicted shift handoff."}, {"id": "a4", "statement": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"id": "a5", "statement": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item."}, {"id": "a6", "statement": "The depicted shift handoff contains exactly one acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item."}, {"id": "a7", "statement": "The sole acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48."}, {"id": "a8", "statement": "The person who made the 17:48 acknowledgment on ticket M-882 is the same person whom ticket M-882 names as the owner of the unresolved Line 4 pump vibration item."}], "base_state_json": "[{\"speaker\":\"Outgoing shift lead\",\"text\":\"Line 4 pump vibration is the only unresolved production item in scope for this handoff. Status at 17:42: stable at 6.1 mm/s. Trend snapshot TS-144 is attached, timestamped. Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.\"},{\"speaker\":\"Operations coordinator\",\"text\":\"Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item. There is exactly one acknowledgment of the routing and next action on the ticket. The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Mara Chen. The completed filter swap is logged in work order W-219 with a 17:30 completion photo, and all other production items show current status and timestamped artifacts.\"},{\"speaker\":\"Incoming shift lead\",\"text\":\"Understood, I'll monitor the Line 4 trend until Mara reports back after the inspection.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": ["0", "text"], "text": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"path": ["1", "text"], "text": "The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Mara Chen."}], "policy_evidence": [], "rules": [{"justification": "All in-scope production items have the required status and timestamped artifact, and the sole unresolved item has an appropriate named owner, a next action, and an acknowledgment made by that owner; therefore the handoff is Ready.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "Although all required status and artifact evidence is present, the sole acknowledgment for the only unresolved item was made by someone other than its named owner. The unresolved item therefore lacks that owner’s acknowledgment, making the handoff Conditional rather than Ready.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_left": "Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item.", "negative_right": "The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Tom Reyes.", "right": "The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Mara Chen."}, "verifier_independent_model": false}, "family": "fast-43-diverse-034-011", "id": "fast-43-diverse-034-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the handoff is either Conditional or Not Ready because at least one in-scope production item fails a required evidence, routing, action, or acknowledgment condition.", "true": "Yes — the handoff is Ready, the highest level, because every in-scope production item satisfies the evidence and routing requirements."}, "instructions": "Decide whether the handoff earns the highest readiness rating. Evaluate only production items; facilities announcements are distractors. Required evidence for every production item is a current status and a timestamped artifact. Each unresolved item also needs a named appropriate owner, a next action, and that owner’s acknowledgment. Ordered scale: Ready = all requirements met; Conditional = status and artifact are present, but an unresolved item lacks valid routing or acknowledgment; Not Ready = any production item lacks status or a timestamped artifact. Question: Is the handoff rated Ready?", "type": "noul"}}, "state": [{"speaker": "Outgoing shift lead", "text": "Line 4 pump vibration is the only unresolved production item in scope for this handoff. Status at 17:42: stable at 6.1 mm/s. Trend snapshot TS-144 is attached, timestamped. Ticket M-882 names maintenance duty engineer Mara Chen as the appropriate owner of the unresolved Line 4 pump vibration item."}, {"speaker": "Operations coordinator", "text": "Ticket M-882 records the 20:00 inspection as the next action for the unresolved Line 4 pump vibration item. There is exactly one acknowledgment of the routing and next action on the ticket. The acknowledgment of the routing and next action for the unresolved Line 4 pump vibration item was recorded on ticket M-882 at 17:48 by Tom Reyes. The completed filter swap is logged in work order W-219 with a 17:30 completion photo, and all other production items show current status and timestamped artifacts."}, {"speaker": "Incoming shift lead", "text": "Understood, I'll monitor the Line 4 trend until Mara reports back after the inspection."}]}, "method": "c2d", "provenance": {"source_id": "diverse-034", "source_is_synthetic": true, "source_sha256": "3a31742d047ead2d7f071400d6a85f1083e595e5e3388ede0a3e68f234443d46", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text verbatim and the same entities/time bindings; the counterfactual changes only Priya Nandakumar's role from Operations coordinator to Quality assurance technician, a coherent single-fact observation change that does not contradict any other stated fact; the two focus evidence spans are complete factual sentences (not policy or instructions); no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Two unresolved handoff items remain: the Conveyor 4 restart check and the Filler 2 pressure alarm, which is the only critical equipment alarm among them. Each item has a current status artifact and a stated next action. The Conveyor 4 restart check is a routine process follow-up owned by Marcus Webb, who is serving as Incoming shift lead. The Filler 2 pressure alarm has an attached diagnostic result. At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar. At 18:00, Priya Nandakumar is serving as the Operations coordinator. The outgoing shift lead signed off at 17:55, but the incoming shift lead has not yet acknowledged the handoff.\",\"policy\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["context"], "text": "At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar."}, {"path": ["context"], "text": "At 18:00, Priya Nandakumar is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar.", "negative_left": "At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar.", "negative_right": "At 18:00, Priya Nandakumar is serving as the Quality assurance technician.", "right": "At 18:00, Priya Nandakumar is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-002", "id": "fast-43-diverse-036-002-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Two unresolved handoff items remain: the Conveyor 4 restart check and the Filler 2 pressure alarm, which is the only critical equipment alarm among them. Each item has a current status artifact and a stated next action. The Conveyor 4 restart check is a routine process follow-up owned by Marcus Webb, who is serving as Incoming shift lead. The Filler 2 pressure alarm has an attached diagnostic result. At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar. At 18:00, Priya Nandakumar is serving as the Operations coordinator. The outgoing shift lead signed off at 17:55, but the incoming shift lead has not yet acknowledged the handoff.", "policy": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text verbatim and the same entities/time bindings; the counterfactual changes only Priya Nandakumar's role from Operations coordinator to Quality assurance technician, a coherent single-fact observation change that does not contradict any other stated fact; the two focus evidence spans are complete factual sentences (not policy or instructions); no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Two unresolved handoff items remain: the Conveyor 4 restart check and the Filler 2 pressure alarm, which is the only critical equipment alarm among them. Each item has a current status artifact and a stated next action. The Conveyor 4 restart check is a routine process follow-up owned by Marcus Webb, who is serving as Incoming shift lead. The Filler 2 pressure alarm has an attached diagnostic result. At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar. At 18:00, Priya Nandakumar is serving as the Operations coordinator. The outgoing shift lead signed off at 17:55, but the incoming shift lead has not yet acknowledged the handoff.\",\"policy\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["context"], "text": "At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar."}, {"path": ["context"], "text": "At 18:00, Priya Nandakumar is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar.", "negative_left": "At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar.", "negative_right": "At 18:00, Priya Nandakumar is serving as the Quality assurance technician.", "right": "At 18:00, Priya Nandakumar is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-002", "id": "fast-43-diverse-036-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Two unresolved handoff items remain: the Conveyor 4 restart check and the Filler 2 pressure alarm, which is the only critical equipment alarm among them. Each item has a current status artifact and a stated next action. The Conveyor 4 restart check is a routine process follow-up owned by Marcus Webb, who is serving as Incoming shift lead. The Filler 2 pressure alarm has an attached diagnostic result. At 18:00, the owner field on the Filler 2 pressure alarm names Priya Nandakumar. At 18:00, Priya Nandakumar is serving as the Quality assurance technician. The outgoing shift lead signed off at 17:55, but the incoming shift lead has not yet acknowledged the handoff.", "policy": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text verbatim, preserve the same entities (Conveyor 4, Filler 2, Priya Anand, Marcus Webb) and the 18:00 time binding matching the original question; each evidence item is a complete factual statement rather than a policy or instruction; the counterfactual changes only Marcus Webb's role from Operations coordinator to Sanitation technician, a single coherent factual edit that does not conflict with any other retained fact (it merely creates a routing mismatch for the critical alarm); neither context contains gold labels, rule tables, or explicit output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"The Filler 2 pressure alarm is the only critical equipment alarm among the unresolved items; the Conveyor 4 restart check is a routine process follow-up.\",\"Both items have a current status artifact and a stated next action logged in the shift board.\",\"Conveyor 4 restart check: owner field lists Priya Anand, who is serving as the Incoming shift lead.\",\"Filler 2 pressure alarm: a diagnostic result showing pressure trending downward is attached.\",\"At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb.\",\"At 18:00, Marcus Webb is serving as the Operations coordinator.\",\"The Incoming shift lead has not yet signed or acknowledged the handoff log as of 18:00.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "5"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb."}, {"path": ["evidence", "6"], "text": "At 18:00, Marcus Webb is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb.", "negative_right": "At 18:00, Marcus Webb is serving as the Sanitation technician.", "right": "At 18:00, Marcus Webb is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-005", "id": "fast-43-diverse-036-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm.", "The Filler 2 pressure alarm is the only critical equipment alarm among the unresolved items; the Conveyor 4 restart check is a routine process follow-up.", "Both items have a current status artifact and a stated next action logged in the shift board.", "Conveyor 4 restart check: owner field lists Priya Anand, who is serving as the Incoming shift lead.", "Filler 2 pressure alarm: a diagnostic result showing pressure trending downward is attached.", "At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb.", "At 18:00, Marcus Webb is serving as the Operations coordinator.", "The Incoming shift lead has not yet signed or acknowledged the handoff log as of 18:00."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text verbatim, preserve the same entities (Conveyor 4, Filler 2, Priya Anand, Marcus Webb) and the 18:00 time binding matching the original question; each evidence item is a complete factual statement rather than a policy or instruction; the counterfactual changes only Marcus Webb's role from Operations coordinator to Sanitation technician, a single coherent factual edit that does not conflict with any other retained fact (it merely creates a routing mismatch for the critical alarm); neither context contains gold labels, rule tables, or explicit output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm.\",\"The Filler 2 pressure alarm is the only critical equipment alarm among the unresolved items; the Conveyor 4 restart check is a routine process follow-up.\",\"Both items have a current status artifact and a stated next action logged in the shift board.\",\"Conveyor 4 restart check: owner field lists Priya Anand, who is serving as the Incoming shift lead.\",\"Filler 2 pressure alarm: a diagnostic result showing pressure trending downward is attached.\",\"At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb.\",\"At 18:00, Marcus Webb is serving as the Operations coordinator.\",\"The Incoming shift lead has not yet signed or acknowledged the handoff log as of 18:00.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "5"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb."}, {"path": ["evidence", "6"], "text": "At 18:00, Marcus Webb is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb.", "negative_right": "At 18:00, Marcus Webb is serving as the Sanitation technician.", "right": "At 18:00, Marcus Webb is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-005", "id": "fast-43-diverse-036-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm.", "The Filler 2 pressure alarm is the only critical equipment alarm among the unresolved items; the Conveyor 4 restart check is a routine process follow-up.", "Both items have a current status artifact and a stated next action logged in the shift board.", "Conveyor 4 restart check: owner field lists Priya Anand, who is serving as the Incoming shift lead.", "Filler 2 pressure alarm: a diagnostic result showing pressure trending downward is attached.", "At 18:00, the Filler 2 pressure alarm's owner field identifies Marcus Webb.", "At 18:00, Marcus Webb is serving as the Sanitation technician.", "The Incoming shift lead has not yet signed or acknowledged the handoff log as of 18:00."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy text and question bindings intact; the counterfactual changes only Priya Nandan's role from Operations coordinator to Quality assurance technician, creating a coherent single-fact shift (the critical alarm is now owned by someone other than the coordinator) without contradicting other stated facts; the two focus evidence spans are complete factual statements, not policy or rule text, and neither context states or implies the rubric level or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Marcus Reilly, the Incoming shift lead, who has not yet acknowledged the handoff.\",\"Filler 2 critical pressure alarm: note dated 17:58 states current fluctuation status and instructs inspection after 18:15. A diagnostic result is attached showing sensor drift of 3.2 PSI. This is the only critical equipment alarm among the unresolved handoff items.\",\"At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person.\",\"At 18:00, Priya Nandan is serving as the Operations coordinator for the packaging operations team.\",\"Outgoing shift lead signed the handoff at 17:55; no other unresolved items are listed.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person."}, {"path": ["evidence", "3"], "text": "At 18:00, Priya Nandan is serving as the Operations coordinator for the packaging operations team."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person.", "negative_left": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person.", "negative_right": "At 18:00, Priya Nandan is serving as the Quality assurance technician for the packaging operations team.", "right": "At 18:00, Priya Nandan is serving as the Operations coordinator for the packaging operations team."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-010", "id": "fast-43-diverse-036-010-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Marcus Reilly, the Incoming shift lead, who has not yet acknowledged the handoff.", "Filler 2 critical pressure alarm: note dated 17:58 states current fluctuation status and instructs inspection after 18:15. A diagnostic result is attached showing sensor drift of 3.2 PSI. This is the only critical equipment alarm among the unresolved handoff items.", "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person.", "At 18:00, Priya Nandan is serving as the Operations coordinator for the packaging operations team.", "Outgoing shift lead signed the handoff at 17:55; no other unresolved items are listed."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy text and question bindings intact; the counterfactual changes only Priya Nandan's role from Operations coordinator to Quality assurance technician, creating a coherent single-fact shift (the critical alarm is now owned by someone other than the coordinator) without contradicting other stated facts; the two focus evidence spans are complete factual statements, not policy or rule text, and neither context states or implies the rubric level or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Marcus Reilly, the Incoming shift lead, who has not yet acknowledged the handoff.\",\"Filler 2 critical pressure alarm: note dated 17:58 states current fluctuation status and instructs inspection after 18:15. A diagnostic result is attached showing sensor drift of 3.2 PSI. This is the only critical equipment alarm among the unresolved handoff items.\",\"At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person.\",\"At 18:00, Priya Nandan is serving as the Operations coordinator for the packaging operations team.\",\"Outgoing shift lead signed the handoff at 17:55; no other unresolved items are listed.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person."}, {"path": ["evidence", "3"], "text": "At 18:00, Priya Nandan is serving as the Operations coordinator for the packaging operations team."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person.", "negative_left": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person.", "negative_right": "At 18:00, Priya Nandan is serving as the Quality assurance technician for the packaging operations team.", "right": "At 18:00, Priya Nandan is serving as the Operations coordinator for the packaging operations team."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-010", "id": "fast-43-diverse-036-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Marcus Reilly, the Incoming shift lead, who has not yet acknowledged the handoff.", "Filler 2 critical pressure alarm: note dated 17:58 states current fluctuation status and instructs inspection after 18:15. A diagnostic result is attached showing sensor drift of 3.2 PSI. This is the only critical equipment alarm among the unresolved handoff items.", "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan as the responsible person.", "At 18:00, Priya Nandan is serving as the Quality assurance technician for the packaging operations team.", "Outgoing shift lead signed the handoff at 17:55; no other unresolved items are listed."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy and question bindings (entities, time, request) unchanged; the counterfactual alters only Jordan Meyer's role from Operations coordinator to Maintenance technician, a single coherent factual change that plausibly shifts routing compliance without contradicting other evidence; the focus evidence consists of two complete factual sentences, not policy text; no gold answers, rule tables, or rationale are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Casey Brooks, who is serving as the Incoming shift lead. Next action: confirm belt tension before restart. No acknowledgement has yet been logged from Casey Brooks for this handoff.\",\"Filler 2 critical pressure alarm: note says 'still fluctuating; inspect after 18:15,' with a diagnostic readout attached showing sensor drift. Next action: recheck sensor calibration. The Filler 2 pressure alarm's owner field identifies Jordan Meyer. At 18:00, Jordan Meyer is serving as the Operations coordinator.\",\"Outgoing shift lead signed the handoff at 17:55; no other unresolved items are listed.\"],\"request\":\"Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "The Filler 2 pressure alarm's owner field identifies Jordan Meyer."}, {"path": ["evidence", "1"], "text": "At 18:00, Jordan Meyer is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The Filler 2 pressure alarm's owner field identifies Jordan Meyer.", "negative_left": "The Filler 2 pressure alarm's owner field identifies Jordan Meyer.", "negative_right": "At 18:00, Jordan Meyer is serving as the Maintenance technician.", "right": "At 18:00, Jordan Meyer is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-013", "id": "fast-43-diverse-036-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Casey Brooks, who is serving as the Incoming shift lead. Next action: confirm belt tension before restart. No acknowledgement has yet been logged from Casey Brooks for this handoff.", "Filler 2 critical pressure alarm: note says 'still fluctuating; inspect after 18:15,' with a diagnostic readout attached showing sensor drift. Next action: recheck sensor calibration. The Filler 2 pressure alarm's owner field identifies Jordan Meyer. At 18:00, Jordan Meyer is serving as the Operations coordinator.", "Outgoing shift lead signed the handoff at 17:55; no other unresolved items are listed."], "request": "Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy and question bindings (entities, time, request) unchanged; the counterfactual alters only Jordan Meyer's role from Operations coordinator to Maintenance technician, a single coherent factual change that plausibly shifts routing compliance without contradicting other evidence; the focus evidence consists of two complete factual sentences, not policy text; no gold answers, rule tables, or rationale are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Casey Brooks, who is serving as the Incoming shift lead. Next action: confirm belt tension before restart. No acknowledgement has yet been logged from Casey Brooks for this handoff.\",\"Filler 2 critical pressure alarm: note says 'still fluctuating; inspect after 18:15,' with a diagnostic readout attached showing sensor drift. Next action: recheck sensor calibration. The Filler 2 pressure alarm's owner field identifies Jordan Meyer. At 18:00, Jordan Meyer is serving as the Operations coordinator.\",\"Outgoing shift lead signed the handoff at 17:55; no other unresolved items are listed.\"],\"request\":\"Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "The Filler 2 pressure alarm's owner field identifies Jordan Meyer."}, {"path": ["evidence", "1"], "text": "At 18:00, Jordan Meyer is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The Filler 2 pressure alarm's owner field identifies Jordan Meyer.", "negative_left": "The Filler 2 pressure alarm's owner field identifies Jordan Meyer.", "negative_right": "At 18:00, Jordan Meyer is serving as the Maintenance technician.", "right": "At 18:00, Jordan Meyer is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-013", "id": "fast-43-diverse-036-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Casey Brooks, who is serving as the Incoming shift lead. Next action: confirm belt tension before restart. No acknowledgement has yet been logged from Casey Brooks for this handoff.", "Filler 2 critical pressure alarm: note says 'still fluctuating; inspect after 18:15,' with a diagnostic readout attached showing sensor drift. Next action: recheck sensor calibration. The Filler 2 pressure alarm's owner field identifies Jordan Meyer. At 18:00, Jordan Meyer is serving as the Maintenance technician.", "Outgoing shift lead signed the handoff at 17:55; no other unresolved items are listed."], "request": "Rate handoff readiness using the five-level rubric and identify why unresolved work is or is not properly supported and routed."}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy verbatim (status artifact, next action, owner, critical-alarm routing to Operations coordinator, and acknowledgement requirement), matching the original state's policy language. Question bindings (18:00 timestamp, Filler 2 pressure alarm, Priya Nandan, Conveyor 4 restart check) are unchanged across base and counterfactual. The two focus evidence sentences are single factual statements (owner field identity and role assignment), not policy text or instructions. The counterfactual changes only Priya Nandan's role from Operations coordinator to Maintenance technician, which is a coherent single-fact change consistent with the rest of the unaltered evidence (it creates a routing mismatch rather than a contradiction, since the Filler 2 owner field still names her). Neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"Two unresolved handoff items remain at 18:00: the Conveyor 4 restart check and the Filler 2 pressure alarm. The Filler 2 pressure alarm is the only critical equipment alarm among them; the Conveyor 4 restart check is a routine process follow-up.\",\"Both items carry a current status artifact and a stated next action. The Filler 2 pressure alarm has an attached diagnostic result showing pressure still fluctuating. The Conveyor 4 restart check's owner field names the person serving as Incoming shift lead. At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan.\",\"At 18:00, Priya Nandan is serving as the Operations coordinator.\",\"The outgoing shift lead signed the handoff log at 17:55, but no acknowledgement from the incoming shift lead is recorded as of 18:00.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan."}, {"path": ["evidence", "2"], "text": "At 18:00, Priya Nandan is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan.", "negative_left": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan.", "negative_right": "At 18:00, Priya Nandan is serving as the Maintenance technician.", "right": "At 18:00, Priya Nandan is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-014", "id": "fast-43-diverse-036-014-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Two unresolved handoff items remain at 18:00: the Conveyor 4 restart check and the Filler 2 pressure alarm. The Filler 2 pressure alarm is the only critical equipment alarm among them; the Conveyor 4 restart check is a routine process follow-up.", "Both items carry a current status artifact and a stated next action. The Filler 2 pressure alarm has an attached diagnostic result showing pressure still fluctuating. The Conveyor 4 restart check's owner field names the person serving as Incoming shift lead. At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan.", "At 18:00, Priya Nandan is serving as the Operations coordinator.", "The outgoing shift lead signed the handoff log at 17:55, but no acknowledgement from the incoming shift lead is recorded as of 18:00."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy verbatim (status artifact, next action, owner, critical-alarm routing to Operations coordinator, and acknowledgement requirement), matching the original state's policy language. Question bindings (18:00 timestamp, Filler 2 pressure alarm, Priya Nandan, Conveyor 4 restart check) are unchanged across base and counterfactual. The two focus evidence sentences are single factual statements (owner field identity and role assignment), not policy text or instructions. The counterfactual changes only Priya Nandan's role from Operations coordinator to Maintenance technician, which is a coherent single-fact change consistent with the rest of the unaltered evidence (it creates a routing mismatch rather than a contradiction, since the Filler 2 owner field still names her). Neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"Two unresolved handoff items remain at 18:00: the Conveyor 4 restart check and the Filler 2 pressure alarm. The Filler 2 pressure alarm is the only critical equipment alarm among them; the Conveyor 4 restart check is a routine process follow-up.\",\"Both items carry a current status artifact and a stated next action. The Filler 2 pressure alarm has an attached diagnostic result showing pressure still fluctuating. The Conveyor 4 restart check's owner field names the person serving as Incoming shift lead. At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan.\",\"At 18:00, Priya Nandan is serving as the Operations coordinator.\",\"The outgoing shift lead signed the handoff log at 17:55, but no acknowledgement from the incoming shift lead is recorded as of 18:00.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan."}, {"path": ["evidence", "2"], "text": "At 18:00, Priya Nandan is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan.", "negative_left": "At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan.", "negative_right": "At 18:00, Priya Nandan is serving as the Maintenance technician.", "right": "At 18:00, Priya Nandan is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-014", "id": "fast-43-diverse-036-014-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Two unresolved handoff items remain at 18:00: the Conveyor 4 restart check and the Filler 2 pressure alarm. The Filler 2 pressure alarm is the only critical equipment alarm among them; the Conveyor 4 restart check is a routine process follow-up.", "Both items carry a current status artifact and a stated next action. The Filler 2 pressure alarm has an attached diagnostic result showing pressure still fluctuating. The Conveyor 4 restart check's owner field names the person serving as Incoming shift lead. At 18:00, the owner field of the Filler 2 pressure alarm identifies Priya Nandan.", "At 18:00, Priya Nandan is serving as the Maintenance technician.", "The outgoing shift lead signed the handoff log at 17:55, but no acknowledgement from the incoming shift lead is recorded as of 18:00."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy verbatim and the same entities/timeframe (Conveyor 4, Filler 2, 18:00 shift change); the counterfactual changes only Jordan Meece's role from Operations coordinator to Maintenance technician, a single coherent factual edit that does not contradict other evidence (owner field, diagnostic, acknowledgement status remain unchanged); each focus evidence item is a single factual sentence with no policy text or answer leakage embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"Two unresolved items remain at shift change: the Conveyor 4 restart check and the Filler 2 pressure alarm. The Filler 2 pressure alarm is the only critical equipment alarm among them; the Conveyor 4 restart check is a routine process follow-up.\",\"Conveyor 4 restart check: status log shows the conveyor cleared at 17:42, next action listed as 'confirm restart,' owner field names Priya Anand, who is serving as the Incoming shift lead.\",\"Filler 2 pressure alarm: diagnostic printout attached shows fluctuating readings at 17:58; next action listed as 'inspect after 18:15.' At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece.\",\"At 18:00, Jordan Meece is serving as the Operations coordinator.\",\"The outgoing shift lead signed off at 17:55. As of 18:00, the incoming shift lead's acknowledgement box on the handoff sheet remains unchecked.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece."}, {"path": ["evidence", "3"], "text": "At 18:00, Jordan Meece is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece.", "negative_right": "At 18:00, Jordan Meece is serving as the Maintenance technician, not the Operations coordinator.", "right": "At 18:00, Jordan Meece is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-019", "id": "fast-43-diverse-036-019-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Two unresolved items remain at shift change: the Conveyor 4 restart check and the Filler 2 pressure alarm. The Filler 2 pressure alarm is the only critical equipment alarm among them; the Conveyor 4 restart check is a routine process follow-up.", "Conveyor 4 restart check: status log shows the conveyor cleared at 17:42, next action listed as 'confirm restart,' owner field names Priya Anand, who is serving as the Incoming shift lead.", "Filler 2 pressure alarm: diagnostic printout attached shows fluctuating readings at 17:58; next action listed as 'inspect after 18:15.' At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece.", "At 18:00, Jordan Meece is serving as the Operations coordinator.", "The outgoing shift lead signed off at 17:55. As of 18:00, the incoming shift lead's acknowledgement box on the handoff sheet remains unchecked."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy verbatim and the same entities/timeframe (Conveyor 4, Filler 2, 18:00 shift change); the counterfactual changes only Jordan Meece's role from Operations coordinator to Maintenance technician, a single coherent factual edit that does not contradict other evidence (owner field, diagnostic, acknowledgement status remain unchanged); each focus evidence item is a single factual sentence with no policy text or answer leakage embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\":\"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\",\"evidence\":[\"Two unresolved items remain at shift change: the Conveyor 4 restart check and the Filler 2 pressure alarm. The Filler 2 pressure alarm is the only critical equipment alarm among them; the Conveyor 4 restart check is a routine process follow-up.\",\"Conveyor 4 restart check: status log shows the conveyor cleared at 17:42, next action listed as 'confirm restart,' owner field names Priya Anand, who is serving as the Incoming shift lead.\",\"Filler 2 pressure alarm: diagnostic printout attached shows fluctuating readings at 17:58; next action listed as 'inspect after 18:15.' At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece.\",\"At 18:00, Jordan Meece is serving as the Operations coordinator.\",\"The outgoing shift lead signed off at 17:55. As of 18:00, the incoming shift lead's acknowledgement box on the handoff sheet remains unchecked.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "2"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece."}, {"path": ["evidence", "3"], "text": "At 18:00, Jordan Meece is serving as the Operations coordinator."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece.", "negative_right": "At 18:00, Jordan Meece is serving as the Maintenance technician, not the Operations coordinator.", "right": "At 18:00, Jordan Meece is serving as the Operations coordinator."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-019", "id": "fast-43-diverse-036-019-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Two unresolved items remain at shift change: the Conveyor 4 restart check and the Filler 2 pressure alarm. The Filler 2 pressure alarm is the only critical equipment alarm among them; the Conveyor 4 restart check is a routine process follow-up.", "Conveyor 4 restart check: status log shows the conveyor cleared at 17:42, next action listed as 'confirm restart,' owner field names Priya Anand, who is serving as the Incoming shift lead.", "Filler 2 pressure alarm: diagnostic printout attached shows fluctuating readings at 17:58; next action listed as 'inspect after 18:15.' At 18:00, the Filler 2 pressure alarm's owner field identifies Jordan Meece.", "At 18:00, Jordan Meece is serving as the Maintenance technician, not the Operations coordinator.", "The outgoing shift lead signed off at 17:55. As of 18:00, the incoming shift lead's acknowledgement box on the handoff sheet remains unchecked."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text verbatim and the same entity/time bindings from the question; the focus evidence consists of two distinct factual sentences (owner field identification and coordinator role) with no policy definitions or instructions embedded. The counterfactual changes only the owner name from Dana Reyes to Sam Torres, which is a coherent single-fact alteration that does not contradict the retained statement that Dana Reyes is the Operations coordinator—creating a plausible misrouting scenario rather than a duplicate or contradictory measurement. Neither context contains a gold answer, rule table, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\": \"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\", \"evidence\": [\"Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Jordan Lee, the Incoming shift lead. Acknowledgement from the Incoming shift lead is still pending.\", \"Filler 2 critical pressure alarm: note timestamped 17:50 says 'still fluctuating; inspect after 18:15,' with a diagnostic readout attached. At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Reyes as the assigned owner. At 18:00, Dana Reyes is serving as the Operations coordinator for the shift.\", \"No other unresolved items remain at 18:00; these two constitute the complete list.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Reyes as the assigned owner."}, {"path": ["evidence", "1"], "text": "At 18:00, Dana Reyes is serving as the Operations coordinator for the shift."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Reyes as the assigned owner.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Sam Torres as the assigned owner.", "negative_right": "At 18:00, Dana Reyes is serving as the Operations coordinator for the shift.", "right": "At 18:00, Dana Reyes is serving as the Operations coordinator for the shift."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-027", "id": "fast-43-diverse-036-027-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Jordan Lee, the Incoming shift lead. Acknowledgement from the Incoming shift lead is still pending.", "Filler 2 critical pressure alarm: note timestamped 17:50 says 'still fluctuating; inspect after 18:15,' with a diagnostic readout attached. At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Reyes as the assigned owner. At 18:00, Dana Reyes is serving as the Operations coordinator for the shift.", "No other unresolved items remain at 18:00; these two constitute the complete list."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-01", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text verbatim and the same entity/time bindings from the question; the focus evidence consists of two distinct factual sentences (owner field identification and coordinator role) with no policy definitions or instructions embedded. The counterfactual changes only the owner name from Dana Reyes to Sam Torres, which is a coherent single-fact alteration that does not contradict the retained statement that Dana Reyes is the Operations coordinator—creating a plausible misrouting scenario rather than a duplicate or contradictory measurement. Neither context contains a gold answer, rule table, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported", "a9": "refuted"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual or set-level relationship; none is a bundled policy conclusion or final readiness classification. The focus a7 is a factual role relationship. Base and counter assignments differ only on a7 and are both realizable: the identified Filler 2 owner can either be or not be the Operations coordinator without changing the other facts. Policy evidence appropriately cites the original state context and preserves all state-originating routing, artifact, diagnostic, and acknowledgement rules needed to interpret the unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish the complete unresolved-item set, identify Filler 2 as the sole critical alarm, provide current artifacts and next actions for all items, and provide its required diagnostic result. Refutation of a7 establishes that the critical alarm's identified owner is not the required Operations coordinator, while the owner remains identifiable. This satisfies Level 1's wrong-owner condition and avoids Level 0 because usable status, a next action, and an identifiable owner are present.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish that all unresolved items have current evidence and next actions; the sole critical alarm has its diagnostic result and required Operations coordinator owner; and the routine Conveyor follow-up is owned by the Incoming shift lead. Refutation of a9 explicitly excludes acknowledgement, so Level 3 and Level 4 cannot apply. The pending acknowledgement directly satisfies Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "At 18:00, the complete set of unresolved handoff items consists of the Conveyor 4 restart check and the Filler 2 pressure alarm."}, {"id": "a2", "statement": "At 18:00, the Filler 2 pressure alarm is the only critical equipment alarm among the unresolved handoff items."}, {"id": "a3", "statement": "At 18:00, the Conveyor 4 restart check is a routine process follow-up."}, {"id": "a4", "statement": "At 18:00, every unresolved handoff item has a current status artifact."}, {"id": "a5", "statement": "At 18:00, every unresolved handoff item has a stated next action."}, {"id": "a6", "statement": "At 18:00, the person identified in the Conveyor 4 restart check's owner field is serving as the Incoming shift lead."}, {"id": "a7", "statement": "At 18:00, the person identified in the Filler 2 pressure alarm's owner field is serving as the Operations coordinator."}, {"id": "a8", "statement": "At 18:00, the Filler 2 pressure alarm has an attached diagnostic result."}, {"id": "a9", "statement": "By 18:00, the Incoming shift lead has acknowledged the handoff."}], "base_state_json": "{\"context\": \"At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.\", \"evidence\": [\"Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Jordan Lee, the Incoming shift lead. Acknowledgement from the Incoming shift lead is still pending.\", \"Filler 2 critical pressure alarm: note timestamped 17:50 says 'still fluctuating; inspect after 18:15,' with a diagnostic readout attached. At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Reyes as the assigned owner. At 18:00, Dana Reyes is serving as the Operations coordinator for the shift.\", \"No other unresolved items remain at 18:00; these two constitute the complete list.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a7", "focus_evidence": [{"path": ["evidence", "1"], "text": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Reyes as the assigned owner."}, {"path": ["evidence", "1"], "text": "At 18:00, Dana Reyes is serving as the Operations coordinator for the shift."}], "policy_evidence": [{"path": ["context"], "text": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready."}], "rules": [{"justification": "The sole critical unresolved item has usable current evidence, a next action, and a diagnostic result, but its identified owner is explicitly not the Operations coordinator required by policy. This is Low readiness rather than Not ready because usable evidence, a next action, and an identifiable owner are present.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Every critical item has the required current evidence, diagnostic result, next action, and policy-correct owner, and the routine item is correctly supported and routed. The incoming shift lead's acknowledgement is explicitly still pending, which makes the handoff Conditionally ready and prevents Level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Dana Reyes as the assigned owner.", "negative_left": "At 18:00, the Filler 2 pressure alarm's owner field identifies Sam Torres as the assigned owner.", "negative_right": "At 18:00, Dana Reyes is serving as the Operations coordinator for the shift.", "right": "At 18:00, Dana Reyes is serving as the Operations coordinator for the shift."}, "verifier_independent_model": false}, "family": "fast-43-diverse-036-027", "id": "fast-43-diverse-036-027-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: The record provides no usable current status evidence, or unresolved critical work has neither a next action nor an identifiable owner; the incoming shift cannot safely act from it.", "1 — Low readiness: Some items have usable evidence, but at least one critical unresolved item lacks required status evidence, has no next action, or is routed to the wrong owner. Substantial correction is required before acceptance.", "2 — Conditionally ready: Every critical item has required evidence, a next action, and the correct owner, but a routine item is missing evidence or correct routing, or the incoming acknowledgement is still pending.", "3 — Ready: Every unresolved item has required current evidence, a next action, and the correct owner, and the incoming shift lead has acknowledged the handoff. Only inconsequential presentation issues may remain.", "4 — Fully verified: Level 3 is met, and each unresolved item also includes timestamped verification of owner acceptance, an escalation or contingency path, and a clearly stated completion criterion."], "instructions": "Select the single readiness level that best fits the explicit record. Treat required status artifacts, routing, and acknowledgements as mandatory under the stated policy; do not infer missing evidence.", "type": "score"}}, "state": {"context": "At 18:00, the packaging operations team changes shifts. Policy requires every unresolved item to include a current status artifact, next action, and correct owner. Critical equipment alarms must be routed to the Operations coordinator and include a diagnostic result. Routine process follow-ups belong to the Incoming shift lead. The incoming lead must acknowledge a handoff before it is fully ready.", "evidence": ["Conveyor 4 jam: photo timestamped 17:42 shows the obstruction cleared; restart check remains, a routine process follow-up. Assigned to Jordan Lee, the Incoming shift lead. Acknowledgement from the Incoming shift lead is still pending.", "Filler 2 critical pressure alarm: note timestamped 17:50 says 'still fluctuating; inspect after 18:15,' with a diagnostic readout attached. At 18:00, the Filler 2 pressure alarm's owner field identifies Sam Torres as the assigned owner. At 18:00, Dana Reyes is serving as the Operations coordinator for the shift.", "No other unresolved items remain at 18:00; these two constitute the complete list."]}}, "method": "c2d", "provenance": {"source_id": "diverse-036", "source_is_synthetic": true, "source_sha256": "54a131d16e729d63dda3253bb55f2a775605098db17b749262af477fbdee4ae1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "workplace-01", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text and question bindings for Mira Chen's October 8 loop; the counterfactual only changes Devon Alvarez's role from hiring manager to recruiting coordinator, a coherent factual substitution that does not contradict other stated facts; the two evidence spans are complete factual sentences, not policy or instruction text; neither context reveals the gold classification or reasoning.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposed a three-session video loop on October 8. Mira and all three interviewers confirmed, calendar holds show no conflicts, and each invitation contains a working video link. Logistics for every session are complete. The final session runs 6:30-7:00 p.m. in the interviewer's local time, which the calendar record documents as after-hours, and this timing could disrupt the loop. The loop record contains exactly one written approval for the final session's after-hours schedule, issued before confirmation of the loop. The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez. Devon Alvarez serves as the hiring manager for Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez."}, {"path": [], "text": "Devon Alvarez serves as the hiring manager for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez.", "negative_left": "The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez.", "negative_right": "Devon Alvarez serves as a recruiting coordinator for Mira Chen's October 8 loop.", "right": "Devon Alvarez serves as the hiring manager for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-009", "id": "fast-43-diverse-037-009-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposed a three-session video loop on October 8. Mira and all three interviewers confirmed, calendar holds show no conflicts, and each invitation contains a working video link. Logistics for every session are complete. The final session runs 6:30-7:00 p.m. in the interviewer's local time, which the calendar record documents as after-hours, and this timing could disrupt the loop. The loop record contains exactly one written approval for the final session's after-hours schedule, issued before confirmation of the loop. The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez. Devon Alvarez serves as the hiring manager for Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text and question bindings for Mira Chen's October 8 loop; the counterfactual only changes Devon Alvarez's role from hiring manager to recruiting coordinator, a coherent factual substitution that does not contradict other stated facts; the two evidence spans are complete factual sentences, not policy or instruction text; neither context reveals the gold classification or reasoning.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposed a three-session video loop on October 8. Mira and all three interviewers confirmed, calendar holds show no conflicts, and each invitation contains a working video link. Logistics for every session are complete. The final session runs 6:30-7:00 p.m. in the interviewer's local time, which the calendar record documents as after-hours, and this timing could disrupt the loop. The loop record contains exactly one written approval for the final session's after-hours schedule, issued before confirmation of the loop. The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez. Devon Alvarez serves as the hiring manager for Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez."}, {"path": [], "text": "Devon Alvarez serves as the hiring manager for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez.", "negative_left": "The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez.", "negative_right": "Devon Alvarez serves as a recruiting coordinator for Mira Chen's October 8 loop.", "right": "Devon Alvarez serves as the hiring manager for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-009", "id": "fast-43-diverse-037-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposed a three-session video loop on October 8. Mira and all three interviewers confirmed, calendar holds show no conflicts, and each invitation contains a working video link. Logistics for every session are complete. The final session runs 6:30-7:00 p.m. in the interviewer's local time, which the calendar record documents as after-hours, and this timing could disrupt the loop. The loop record contains exactly one written approval for the final session's after-hours schedule, issued before confirmation of the loop. The written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Devon Alvarez. Devon Alvarez serves as a recruiting coordinator for Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy (confirmations/logistics readiness plus after-hours hiring-manager approval requirement) unchanged, keep the same candidate/date/session bindings, and the two evidence sentences are factual, non-instructional; the counterfactual coherently swaps Dana Ruiz's role from hiring manager to recruiting coordinator, altering whether the approval satisfies policy without contradicting any other stated fact, and neither context reveals a label, rule table, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira and all three interviewers confirmed on October 8, calendar holds show no conflicts, and each invitation contains a working video link, so logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, after the 6:00 p.m. policy cutoff. The loop record contains exactly one written approval for the after-hours session, and that approval was issued before the loop was confirmed. This documented after-hours timing appears in the calendar record and could disrupt the loop. The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz. Dana Ruiz serves as the hiring manager for the role in Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz."}, {"path": [], "text": "Dana Ruiz serves as the hiring manager for the role in Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz.", "negative_left": "The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz.", "negative_right": "Dana Ruiz serves as a recruiting coordinator for the role in Mira Chen's October 8 loop.", "right": "Dana Ruiz serves as the hiring manager for the role in Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-011", "id": "fast-43-diverse-037-011-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira and all three interviewers confirmed on October 8, calendar holds show no conflicts, and each invitation contains a working video link, so logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, after the 6:00 p.m. policy cutoff. The loop record contains exactly one written approval for the after-hours session, and that approval was issued before the loop was confirmed. This documented after-hours timing appears in the calendar record and could disrupt the loop. The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz. Dana Ruiz serves as the hiring manager for the role in Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy (confirmations/logistics readiness plus after-hours hiring-manager approval requirement) unchanged, keep the same candidate/date/session bindings, and the two evidence sentences are factual, non-instructional; the counterfactual coherently swaps Dana Ruiz's role from hiring manager to recruiting coordinator, altering whether the approval satisfies policy without contradicting any other stated fact, and neither context reveals a label, rule table, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira and all three interviewers confirmed on October 8, calendar holds show no conflicts, and each invitation contains a working video link, so logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, after the 6:00 p.m. policy cutoff. The loop record contains exactly one written approval for the after-hours session, and that approval was issued before the loop was confirmed. This documented after-hours timing appears in the calendar record and could disrupt the loop. The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz. Dana Ruiz serves as the hiring manager for the role in Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz."}, {"path": [], "text": "Dana Ruiz serves as the hiring manager for the role in Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz.", "negative_left": "The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz.", "negative_right": "Dana Ruiz serves as a recruiting coordinator for the role in Mira Chen's October 8 loop.", "right": "Dana Ruiz serves as the hiring manager for the role in Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-011", "id": "fast-43-diverse-037-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira and all three interviewers confirmed on October 8, calendar holds show no conflicts, and each invitation contains a working video link, so logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, after the 6:00 p.m. policy cutoff. The loop record contains exactly one written approval for the after-hours session, and that approval was issued before the loop was confirmed. This documented after-hours timing appears in the calendar record and could disrupt the loop. The sole written approval for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz. Dana Ruiz serves as a recruiting coordinator for the role in Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy text and question bindings for Mira Chen's Oct 8 loop; the counterfactual coherently swaps Devon Park's role to recruiting-coordinator, invalidating the exception without contradicting other facts; evidence spans are two complete factual sentences with no embedded rules or gold labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Case note — Mira Chen, October 8 three-session video loop.\\n\\nStatus log: Mira Chen confirmed the proposed three-session video loop on October 8. Every interviewer scheduled for the loop also confirmed, and all session logistics are complete, including working video links and no calendar conflicts. The final session of the loop runs 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. cutoff. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, and that approval was issued before the loop's confirmations were logged. This after-hours timing is documented in the calendar record for the loop and could disrupt the schedule if not properly authorized.\\n\\nPolicy on file: 'Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete.' 'This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.'\\n\\nApproval record: The sole written approval for the final session's after-hours schedule was signed by Devon Park. Devon Park holds the hiring-manager role for Mira Chen's October 8 loop.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was signed by Devon Park."}, {"path": [], "text": "Devon Park holds the hiring-manager role for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was signed by Devon Park.", "negative_left": "The sole written approval for the final session's after-hours schedule was signed by Devon Park.", "negative_right": "Devon Park holds the recruiting-coordinator role for Mira Chen's October 8 loop.", "right": "Devon Park holds the hiring-manager role for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-028", "id": "fast-43-diverse-037-028-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Case note — Mira Chen, October 8 three-session video loop.\n\nStatus log: Mira Chen confirmed the proposed three-session video loop on October 8. Every interviewer scheduled for the loop also confirmed, and all session logistics are complete, including working video links and no calendar conflicts. The final session of the loop runs 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. cutoff. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, and that approval was issued before the loop's confirmations were logged. This after-hours timing is documented in the calendar record for the loop and could disrupt the schedule if not properly authorized.\n\nPolicy on file: 'Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete.' 'This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.'\n\nApproval record: The sole written approval for the final session's after-hours schedule was signed by Devon Park. Devon Park holds the hiring-manager role for Mira Chen's October 8 loop."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy text and question bindings for Mira Chen's Oct 8 loop; the counterfactual coherently swaps Devon Park's role to recruiting-coordinator, invalidating the exception without contradicting other facts; evidence spans are two complete factual sentences with no embedded rules or gold labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Case note — Mira Chen, October 8 three-session video loop.\\n\\nStatus log: Mira Chen confirmed the proposed three-session video loop on October 8. Every interviewer scheduled for the loop also confirmed, and all session logistics are complete, including working video links and no calendar conflicts. The final session of the loop runs 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. cutoff. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, and that approval was issued before the loop's confirmations were logged. This after-hours timing is documented in the calendar record for the loop and could disrupt the schedule if not properly authorized.\\n\\nPolicy on file: 'Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete.' 'This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.'\\n\\nApproval record: The sole written approval for the final session's after-hours schedule was signed by Devon Park. Devon Park holds the hiring-manager role for Mira Chen's October 8 loop.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was signed by Devon Park."}, {"path": [], "text": "Devon Park holds the hiring-manager role for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was signed by Devon Park.", "negative_left": "The sole written approval for the final session's after-hours schedule was signed by Devon Park.", "negative_right": "Devon Park holds the recruiting-coordinator role for Mira Chen's October 8 loop.", "right": "Devon Park holds the hiring-manager role for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-028", "id": "fast-43-diverse-037-028-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Case note — Mira Chen, October 8 three-session video loop.\n\nStatus log: Mira Chen confirmed the proposed three-session video loop on October 8. Every interviewer scheduled for the loop also confirmed, and all session logistics are complete, including working video links and no calendar conflicts. The final session of the loop runs 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. cutoff. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, and that approval was issued before the loop's confirmations were logged. This after-hours timing is documented in the calendar record for the loop and could disrupt the schedule if not properly authorized.\n\nPolicy on file: 'Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete.' 'This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.'\n\nApproval record: The sole written approval for the final session's after-hours schedule was signed by Devon Park. Devon Park holds the recruiting-coordinator role for Mira Chen's October 8 loop."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the unchanged policy and question scope, and the counterfactual only swaps Dana Woolf's role, a coherent single-fact change that doesn't contradict other stated facts or embed any answer/rule content.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Case note for Mira Chen's October 8 three-session video loop: Mira confirmed on October 8, and all three scheduled interviewers also confirmed. Calendar holds show no conflicts, each invitation has a working video link, and logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, placing it after 6:00 p.m. The loop record contains exactly one written record granting approval for this after-hours session, issued before the loop's confirmation on October 8. The after-hours timing is documented in the calendar record and could disrupt the loop. The written approval for the final session's after-hours schedule was signed by Dana Woolf. Dana Woolf holds the hiring-manager role in Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The written approval for the final session's after-hours schedule was signed by Dana Woolf."}, {"path": [], "text": "Dana Woolf holds the hiring-manager role in Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The written approval for the final session's after-hours schedule was signed by Dana Woolf.", "negative_left": "The written approval for the final session's after-hours schedule was signed by Dana Woolf.", "negative_right": "Dana Woolf holds the recruiting-coordinator role in Mira Chen's October 8 loop.", "right": "Dana Woolf holds the hiring-manager role in Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-029", "id": "fast-43-diverse-037-029-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Case note for Mira Chen's October 8 three-session video loop: Mira confirmed on October 8, and all three scheduled interviewers also confirmed. Calendar holds show no conflicts, each invitation has a working video link, and logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, placing it after 6:00 p.m. The loop record contains exactly one written record granting approval for this after-hours session, issued before the loop's confirmation on October 8. The after-hours timing is documented in the calendar record and could disrupt the loop. The written approval for the final session's after-hours schedule was signed by Dana Woolf. Dana Woolf holds the hiring-manager role in Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the unchanged policy and question scope, and the counterfactual only swaps Dana Woolf's role, a coherent single-fact change that doesn't contradict other stated facts or embed any answer/rule content.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Case note for Mira Chen's October 8 three-session video loop: Mira confirmed on October 8, and all three scheduled interviewers also confirmed. Calendar holds show no conflicts, each invitation has a working video link, and logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, placing it after 6:00 p.m. The loop record contains exactly one written record granting approval for this after-hours session, issued before the loop's confirmation on October 8. The after-hours timing is documented in the calendar record and could disrupt the loop. The written approval for the final session's after-hours schedule was signed by Dana Woolf. Dana Woolf holds the hiring-manager role in Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The written approval for the final session's after-hours schedule was signed by Dana Woolf."}, {"path": [], "text": "Dana Woolf holds the hiring-manager role in Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The written approval for the final session's after-hours schedule was signed by Dana Woolf.", "negative_left": "The written approval for the final session's after-hours schedule was signed by Dana Woolf.", "negative_right": "Dana Woolf holds the recruiting-coordinator role in Mira Chen's October 8 loop.", "right": "Dana Woolf holds the hiring-manager role in Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-029", "id": "fast-43-diverse-037-029-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Case note for Mira Chen's October 8 three-session video loop: Mira confirmed on October 8, and all three scheduled interviewers also confirmed. Calendar holds show no conflicts, each invitation has a working video link, and logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, placing it after 6:00 p.m. The loop record contains exactly one written record granting approval for this after-hours session, issued before the loop's confirmation on October 8. The after-hours timing is documented in the calendar record and could disrupt the loop. The written approval for the final session's after-hours schedule was signed by Dana Woolf. Dana Woolf holds the recruiting-coordinator role in Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy text and question bindings (Mira Chen, October 8, final session), only changing whether Dana Kowalski's title carries hiring-manager authority, which is a coherent single-sentence factual swap with no contradiction, no embedded answer/label, and the two evidence spans are complete factual sentences rather than policy or instruction text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposed a three-session video loop on October 8. Mira confirmed the loop on October 8, and all three scheduled interviewers also confirmed. Calendar holds show no conflicts, each invitation contains a working video link, and logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. cutoff. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued before the loop's confirmation. The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski. Dana Kowalski holds the title of Senior Recruiting Coordinator, a hiring-manager role at the company. The after-hours timing is documented in the calendar record and could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski."}, {"path": [], "text": "Dana Kowalski holds the title of Senior Recruiting Coordinator, a hiring-manager role at the company."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski.", "negative_left": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski.", "negative_right": "Dana Kowalski holds the title of Senior Recruiting Coordinator, a role that does not carry hiring-manager authority.", "right": "Dana Kowalski holds the title of Senior Recruiting Coordinator, a hiring-manager role at the company."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-034", "id": "fast-43-diverse-037-034-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposed a three-session video loop on October 8. Mira confirmed the loop on October 8, and all three scheduled interviewers also confirmed. Calendar holds show no conflicts, each invitation contains a working video link, and logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. cutoff. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued before the loop's confirmation. The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski. Dana Kowalski holds the title of Senior Recruiting Coordinator, a hiring-manager role at the company. The after-hours timing is documented in the calendar record and could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy text and question bindings (Mira Chen, October 8, final session), only changing whether Dana Kowalski's title carries hiring-manager authority, which is a coherent single-sentence factual swap with no contradiction, no embedded answer/label, and the two evidence spans are complete factual sentences rather than policy or instruction text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposed a three-session video loop on October 8. Mira confirmed the loop on October 8, and all three scheduled interviewers also confirmed. Calendar holds show no conflicts, each invitation contains a working video link, and logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. cutoff. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued before the loop's confirmation. The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski. Dana Kowalski holds the title of Senior Recruiting Coordinator, a hiring-manager role at the company. The after-hours timing is documented in the calendar record and could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski."}, {"path": [], "text": "Dana Kowalski holds the title of Senior Recruiting Coordinator, a hiring-manager role at the company."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski.", "negative_left": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski.", "negative_right": "Dana Kowalski holds the title of Senior Recruiting Coordinator, a role that does not carry hiring-manager authority.", "right": "Dana Kowalski holds the title of Senior Recruiting Coordinator, a hiring-manager role at the company."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-034", "id": "fast-43-diverse-037-034-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposed a three-session video loop on October 8. Mira confirmed the loop on October 8, and all three scheduled interviewers also confirmed. Calendar holds show no conflicts, each invitation contains a working video link, and logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. cutoff. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued before the loop's confirmation. The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was issued by Dana Kowalski. Dana Kowalski holds the title of Senior Recruiting Coordinator, a role that does not carry hiring-manager authority. The after-hours timing is documented in the calendar record and could disrupt the loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text and Mira Chen/October 8 bindings; the counterfactual only swaps Dana Ruiz's title from hiring manager to recruiting coordinator, a coherent factual change affecting approval validity without contradicting other stated facts; evidence spans are two complete factual sentences with no embedded answers or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, and each of the three interviewers scheduled also confirmed. Calendar holds show no conflicts, every invitation contains a working video link, and logistics for all three sessions are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, and this after-hours timing is noted directly on the loop's calendar record, which could disrupt the loop as scheduled. The loop record contains exactly one written approval addressing this after-hours session, issued prior to the loop's confirmation. The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz. Dana Ruiz holds the title of hiring manager in the company's personnel records. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz."}, {"path": [], "text": "Dana Ruiz holds the title of hiring manager in the company's personnel records."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz.", "negative_left": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz.", "negative_right": "Dana Ruiz holds the title of recruiting coordinator in the company's personnel records.", "right": "Dana Ruiz holds the title of hiring manager in the company's personnel records."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-035", "id": "fast-43-diverse-037-035-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, and each of the three interviewers scheduled also confirmed. Calendar holds show no conflicts, every invitation contains a working video link, and logistics for all three sessions are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, and this after-hours timing is noted directly on the loop's calendar record, which could disrupt the loop as scheduled. The loop record contains exactly one written approval addressing this after-hours session, issued prior to the loop's confirmation. The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz. Dana Ruiz holds the title of hiring manager in the company's personnel records. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text and Mira Chen/October 8 bindings; the counterfactual only swaps Dana Ruiz's title from hiring manager to recruiting coordinator, a coherent factual change affecting approval validity without contradicting other stated facts; evidence spans are two complete factual sentences with no embedded answers or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, and each of the three interviewers scheduled also confirmed. Calendar holds show no conflicts, every invitation contains a working video link, and logistics for all three sessions are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, and this after-hours timing is noted directly on the loop's calendar record, which could disrupt the loop as scheduled. The loop record contains exactly one written approval addressing this after-hours session, issued prior to the loop's confirmation. The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz. Dana Ruiz holds the title of hiring manager in the company's personnel records. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz."}, {"path": [], "text": "Dana Ruiz holds the title of hiring manager in the company's personnel records."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz.", "negative_left": "The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz.", "negative_right": "Dana Ruiz holds the title of recruiting coordinator in the company's personnel records.", "right": "Dana Ruiz holds the title of hiring manager in the company's personnel records."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-035", "id": "fast-43-diverse-037-035-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, and each of the three interviewers scheduled also confirmed. Calendar holds show no conflicts, every invitation contains a working video link, and logistics for all three sessions are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, and this after-hours timing is noted directly on the loop's calendar record, which could disrupt the loop as scheduled. The loop record contains exactly one written approval addressing this after-hours session, issued prior to the loop's confirmation. The written approval on file for the final session's after-hours schedule in Mira Chen's October 8 loop was signed by Dana Ruiz. Dana Ruiz holds the title of recruiting coordinator in the company's personnel records. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy paragraph and question bindings (candidate, date, loop) unchanged; the two evidence sentences are factual statements about the approval and signer's role, not policy text; the counterfactual only swaps Priya Nandan's role from hiring manager to recruiting coordinator, which is a coherent, non-contradictory change consistent with the rest of the case; neither context reveals the gold classification, criteria names, or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Case note: Candidate Mira Chen's proposed three-session video loop on October 8 has been reviewed for readiness. Mira Chen confirmed the loop, and all three assigned interviewers also confirmed their sessions. Calendar holds show no conflicts and each invitation contains a working video link, so logistics for every session are complete. The final session is scheduled to run 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. policy cutoff. The loop record contains exactly one written approval addressing the after-hours schedule for that final session, dated before the loop's confirmation. The sole written approval for the final session's after-hours schedule was signed by Priya Nandan. Priya Nandan is the hiring manager for Mira Chen's October 8 loop. The after-hours timing is noted on the shared loop calendar entry, and the coordinator flagged it as a scheduling factor that could disrupt the loop.\\n\\nPolicy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was signed by Priya Nandan."}, {"path": [], "text": "Priya Nandan is the hiring manager for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was signed by Priya Nandan.", "negative_left": "The sole written approval for the final session's after-hours schedule was signed by Priya Nandan.", "negative_right": "Priya Nandan is a recruiting coordinator for Mira Chen's October 8 loop.", "right": "Priya Nandan is the hiring manager for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-036", "id": "fast-43-diverse-037-036-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Case note: Candidate Mira Chen's proposed three-session video loop on October 8 has been reviewed for readiness. Mira Chen confirmed the loop, and all three assigned interviewers also confirmed their sessions. Calendar holds show no conflicts and each invitation contains a working video link, so logistics for every session are complete. The final session is scheduled to run 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. policy cutoff. The loop record contains exactly one written approval addressing the after-hours schedule for that final session, dated before the loop's confirmation. The sole written approval for the final session's after-hours schedule was signed by Priya Nandan. Priya Nandan is the hiring manager for Mira Chen's October 8 loop. The after-hours timing is noted on the shared loop calendar entry, and the coordinator flagged it as a scheduling factor that could disrupt the loop.\n\nPolicy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy paragraph and question bindings (candidate, date, loop) unchanged; the two evidence sentences are factual statements about the approval and signer's role, not policy text; the counterfactual only swaps Priya Nandan's role from hiring manager to recruiting coordinator, which is a coherent, non-contradictory change consistent with the rest of the case; neither context reveals the gold classification, criteria names, or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"Case note: Candidate Mira Chen's proposed three-session video loop on October 8 has been reviewed for readiness. Mira Chen confirmed the loop, and all three assigned interviewers also confirmed their sessions. Calendar holds show no conflicts and each invitation contains a working video link, so logistics for every session are complete. The final session is scheduled to run 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. policy cutoff. The loop record contains exactly one written approval addressing the after-hours schedule for that final session, dated before the loop's confirmation. The sole written approval for the final session's after-hours schedule was signed by Priya Nandan. Priya Nandan is the hiring manager for Mira Chen's October 8 loop. The after-hours timing is noted on the shared loop calendar entry, and the coordinator flagged it as a scheduling factor that could disrupt the loop.\\n\\nPolicy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was signed by Priya Nandan."}, {"path": [], "text": "Priya Nandan is the hiring manager for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was signed by Priya Nandan.", "negative_left": "The sole written approval for the final session's after-hours schedule was signed by Priya Nandan.", "negative_right": "Priya Nandan is a recruiting coordinator for Mira Chen's October 8 loop.", "right": "Priya Nandan is the hiring manager for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-036", "id": "fast-43-diverse-037-036-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "Case note: Candidate Mira Chen's proposed three-session video loop on October 8 has been reviewed for readiness. Mira Chen confirmed the loop, and all three assigned interviewers also confirmed their sessions. Calendar holds show no conflicts and each invitation contains a working video link, so logistics for every session are complete. The final session is scheduled to run 6:30–7:00 p.m. in the interviewer's local time, placing it after the 6:00 p.m. policy cutoff. The loop record contains exactly one written approval addressing the after-hours schedule for that final session, dated before the loop's confirmation. The sole written approval for the final session's after-hours schedule was signed by Priya Nandan. Priya Nandan is a recruiting coordinator for Mira Chen's October 8 loop. The after-hours timing is noted on the shared loop calendar entry, and the coordinator flagged it as a scheduling factor that could disrupt the loop.\n\nPolicy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim policy and question, preserve candidate/date bindings, use two factual (non-instructional) sentences as evidence, and the counterfactual coherently swaps Dana Whit's role from hiring manager to recruiting coordinator without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, and every interviewer scheduled for the loop also confirmed. Logistics for all three sessions are complete, including working video links and conflict-free calendar holds. The final session of the loop runs 6:30-7:00 p.m. in the interviewer's local time, ending after 6:00 p.m., and this after-hours timing is documented in the calendar record for the loop. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued the week before the loop was confirmed. The sole written approval for the final session's after-hours schedule was signed by Dana Whit. Dana Whit serves as the hiring manager for Mira Chen's October 8 loop. Coordinators noted that the documented after-hours timing could still disrupt the loop despite the approval on file. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was signed by Dana Whit."}, {"path": [], "text": "Dana Whit serves as the hiring manager for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was signed by Dana Whit.", "negative_left": "The sole written approval for the final session's after-hours schedule was signed by Dana Whit.", "negative_right": "Dana Whit serves as a recruiting coordinator for Mira Chen's October 8 loop.", "right": "Dana Whit serves as the hiring manager for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-038", "id": "fast-43-diverse-037-038-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, and every interviewer scheduled for the loop also confirmed. Logistics for all three sessions are complete, including working video links and conflict-free calendar holds. The final session of the loop runs 6:30-7:00 p.m. in the interviewer's local time, ending after 6:00 p.m., and this after-hours timing is documented in the calendar record for the loop. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued the week before the loop was confirmed. The sole written approval for the final session's after-hours schedule was signed by Dana Whit. Dana Whit serves as the hiring manager for Mira Chen's October 8 loop. Coordinators noted that the documented after-hours timing could still disrupt the loop despite the approval on file. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the verbatim policy and question, preserve candidate/date bindings, use two factual (non-instructional) sentences as evidence, and the counterfactual coherently swaps Dana Whit's role from hiring manager to recruiting coordinator without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, and every interviewer scheduled for the loop also confirmed. Logistics for all three sessions are complete, including working video links and conflict-free calendar holds. The final session of the loop runs 6:30-7:00 p.m. in the interviewer's local time, ending after 6:00 p.m., and this after-hours timing is documented in the calendar record for the loop. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued the week before the loop was confirmed. The sole written approval for the final session's after-hours schedule was signed by Dana Whit. Dana Whit serves as the hiring manager for Mira Chen's October 8 loop. Coordinators noted that the documented after-hours timing could still disrupt the loop despite the approval on file. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was signed by Dana Whit."}, {"path": [], "text": "Dana Whit serves as the hiring manager for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was signed by Dana Whit.", "negative_left": "The sole written approval for the final session's after-hours schedule was signed by Dana Whit.", "negative_right": "Dana Whit serves as a recruiting coordinator for Mira Chen's October 8 loop.", "right": "Dana Whit serves as the hiring manager for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-038", "id": "fast-43-diverse-037-038-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, and every interviewer scheduled for the loop also confirmed. Logistics for all three sessions are complete, including working video links and conflict-free calendar holds. The final session of the loop runs 6:30-7:00 p.m. in the interviewer's local time, ending after 6:00 p.m., and this after-hours timing is documented in the calendar record for the loop. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued the week before the loop was confirmed. The sole written approval for the final session's after-hours schedule was signed by Dana Whit. Dana Whit serves as a recruiting coordinator for Mira Chen's October 8 loop. Coordinators noted that the documented after-hours timing could still disrupt the loop despite the approval on file. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, question entity/path/time bindings, and the two focus sentences are complete factual statements; the counterfactual only changes Dana Okafor's role from hiring-manager to recruiting-coordinator, a coherent single-fact alteration that does not contradict other stated facts or leak the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, as did every interviewer scheduled for it. Calendar holds show no conflicts, each invitation contains a working video link, and the logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, and this after-hours timing is documented in the calendar record for the loop. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued before confirmation of the loop. The sole written approval for the final session's after-hours schedule was granted by Dana Okafor. Dana Okafor is listed in the personnel system as holding the hiring-manager role for Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation. The documented after-hours timing could disrupt the loop.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was granted by Dana Okafor."}, {"path": [], "text": "Dana Okafor is listed in the personnel system as holding the hiring-manager role for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was granted by Dana Okafor.", "negative_left": "The sole written approval for the final session's after-hours schedule was granted by Dana Okafor.", "negative_right": "Dana Okafor is listed in the personnel system as holding the recruiting-coordinator role for Mira Chen's October 8 loop.", "right": "Dana Okafor is listed in the personnel system as holding the hiring-manager role for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-042", "id": "fast-43-diverse-037-042-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, as did every interviewer scheduled for it. Calendar holds show no conflicts, each invitation contains a working video link, and the logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, and this after-hours timing is documented in the calendar record for the loop. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued before confirmation of the loop. The sole written approval for the final session's after-hours schedule was granted by Dana Okafor. Dana Okafor is listed in the personnel system as holding the hiring-manager role for Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation. The documented after-hours timing could disrupt the loop."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "ready_elevated_risk"}, "source_family": "workplace-02", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, question entity/path/time bindings, and the two focus sentences are complete factual statements; the counterfactual only changes Dana Okafor's role from hiring-manager to recruiting-coordinator, a coherent single-fact alteration that does not contradict other stated facts or leak the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A6": "unknown"}, "remove_right": {"A6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A6": "unknown"}, "negative_pair": {"A6": "refuted"}, "negative_sentence": {"A6": "unknown"}, "positive_pair": {"A6": "supported"}, "right": {"A6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; universal statements such as complete logistics for every session remain atomic over an explicit set. A6 is a factual role relationship rather than a policy conclusion. The base and counter assignments are realizable with only A6 changing: the sole approval can be issued by a hiring manager in one case and by a non-hiring-manager in the other. The policy evidence correctly preserves the substantive readiness, scheduling-window, and exception-approval rules originating in the original state; criteria and routing instructions from the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all confirmations, complete logistics, an after-hours exception, written approval by a hiring manager before confirmation, and a documented timing issue that could disrupt the loop. This satisfies ready_elevated_risk and excludes the confirmation-based high-risk outcomes and ready_low_risk.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete confirmations and logistics but an after-hours session whose sole written approval was granted by someone who is not a hiring manager. Therefore the required written hiring-manager approval is missing, which the criteria expressly place under none_of_above; the complete confirmations exclude both not-ready outcomes.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Mira Chen confirmed the proposed three-session video loop on October 8."}, {"id": "A2", "statement": "Every interviewer scheduled for Mira Chen's proposed three-session video loop on October 8 confirmed."}, {"id": "A3", "statement": "The logistics for every session in Mira Chen's proposed three-session video loop on October 8 are complete."}, {"id": "A4", "statement": "The final session of Mira Chen's proposed October 8 loop ends after 6:00 p.m. in the interviewer's local time."}, {"id": "A5", "statement": "Mira Chen's October 8 loop record contains exactly one written record granting approval for the final session's after-hours schedule."}, {"id": "A6", "statement": "The grantor identified in the sole written approval for the final session's after-hours schedule holds a hiring-manager role."}, {"id": "A7", "statement": "The sole written approval for the final session's after-hours schedule was issued before confirmation of Mira Chen's October 8 loop."}, {"id": "A8", "statement": "The final session's after-hours timing is documented in the calendar record for Mira Chen's October 8 loop."}, {"id": "A9", "statement": "The documented after-hours timing could disrupt Mira Chen's October 8 loop."}], "base_state_json": "\"For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, as did every interviewer scheduled for it. Calendar holds show no conflicts, each invitation contains a working video link, and the logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, and this after-hours timing is documented in the calendar record for the loop. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued before confirmation of the loop. The sole written approval for the final session's after-hours schedule was granted by Dana Okafor. Dana Okafor is listed in the personnel system as holding the hiring-manager role for Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation. The documented after-hours timing could disrupt the loop.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}], "focus_atom": "A6", "focus_evidence": [{"path": [], "text": "The sole written approval for the final session's after-hours schedule was granted by Dana Okafor."}, {"path": [], "text": "Dana Okafor is listed in the personnel system as holding the hiring-manager role for Mira Chen's October 8 loop."}], "policy_evidence": [{"path": [], "text": "Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete."}, {"path": [], "text": "This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation."}], "rules": [{"justification": "The candidate and every scheduled interviewer confirmed, logistics are complete, and the sole written approval for the after-hours session was granted by a person holding a hiring-manager role before confirmation. All applicable requirements are therefore complete. The documented after-hours timing could still disrupt the loop, establishing elevated risk, while the after-hours exception excludes ready_low_risk.", "target": "ready_elevated_risk", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "The candidate and every scheduled interviewer confirmed and logistics are complete, so neither confirmation-based high-risk classification applies. The final session is after hours, and the record's sole written approval was granted by a person who does not hold a hiring-manager role; therefore the required written hiring-manager approval is missing. That missing exception approval excludes both ready classifications and falls under none_of_above.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}]}, "verified_pair": {"left": "The sole written approval for the final session's after-hours schedule was granted by Dana Okafor.", "negative_left": "The sole written approval for the final session's after-hours schedule was granted by Dana Okafor.", "negative_right": "Dana Okafor is listed in the personnel system as holding the recruiting-coordinator role for Mira Chen's October 8 loop.", "right": "Dana Okafor is listed in the personnel system as holding the hiring-manager role for Mira Chen's October 8 loop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-037-042", "id": "fast-43-diverse-037-042-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "None of the listed substantive classifications fits; use for another policy-defined status, risk level, or routing need, including missing hiring-manager approval for an exception.", "not_ready_candidate_high_risk": "Not ready, high risk: the candidate has not confirmed; route the case to the candidate.", "not_ready_interviewer_high_risk": "Not ready, high risk: at least one scheduled interviewer has not confirmed; route the case to that interviewer.", "ready_elevated_risk": "Ready to confirm, elevated risk: all policy requirements, including any applicable exception approvals, are complete, but a documented calendar or logistics issue could still disrupt the loop.", "ready_low_risk": "Ready to confirm, low risk: all required confirmations and logistics are complete, every session is within 8:00 a.m.–6:00 p.m., and no exception applies."}, "instructions": "Classify the loop’s readiness and scheduling risk under the supplied policy. Select exactly one option. If a required approval is missing because of an exception, the case must be routed to the participant authorized to grant it.", "type": "choice"}}, "state": "For candidate Mira Chen, the coordinator proposes a three-session video loop on October 8. Mira Chen confirmed the proposed three-session video loop on October 8, as did every interviewer scheduled for it. Calendar holds show no conflicts, each invitation contains a working video link, and the logistics for every session are complete. The final session runs 6:30–7:00 p.m. in the interviewer's local time, and this after-hours timing is documented in the calendar record for the loop. The loop record contains exactly one written record granting approval for the final session's after-hours schedule, issued before confirmation of the loop. The sole written approval for the final session's after-hours schedule was granted by Dana Okafor. Dana Okafor is listed in the personnel system as holding the recruiting-coordinator role for Mira Chen's October 8 loop. Policy says a loop is ready when the candidate and every interviewer have confirmed and logistics are complete. This policy covers sessions from 8:00 a.m. through 6:00 p.m.; any session outside that scope requires written hiring-manager approval before confirmation. The documented after-hours timing could disrupt the loop."}, "method": "c2d", "provenance": {"source_id": "diverse-037", "source_is_synthetic": true, "source_sha256": "78b20cf316bdc225a762489a9b9e77a4eb562d84935c8ddd84d35311ef323f76", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "workplace-02", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy via the unchanged questions object, preserve the same entity/path bindings while only altering the proposed date, use two complete factual sentences as evidence, and contain no gold answer, rule table, or instructional leakage; the counterfactual's shifted date (Jun 2→Jun 5) is internally consistent with the unchanged dependency and rationale statements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta date is currently scheduled for Monday, June 2, 2025.\"},{\"speaker\":\"Change requester\",\"text\":\"The proposed internal-beta date is Wednesday, June 11, 2025.\"},{\"speaker\":\"Change requester\",\"text\":\"Two authentication defects require retesting after the contractor's departure, and that is the rationale for this request.\"},{\"speaker\":\"Delivery lead\",\"text\":\"My signed schedule estimate confirms the added business days needed.\"},{\"speaker\":\"Delivery lead\",\"text\":\"The analytics dependency owner, the only affected dependency owner, has provided written acknowledgement that the handoff can shift accordingly.\"},{\"speaker\":\"Project manager\",\"text\":\"No budget change and no scope-baseline change are included in this request.\"},{\"speaker\":\"Project manager\",\"text\":\"There is no contractual milestone miss associated with this change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta date is currently scheduled for Monday, June 2, 2025."}, {"path": ["1", "text"], "text": "The proposed internal-beta date is Wednesday, June 11, 2025."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta date is currently scheduled for Monday, June 2, 2025.", "negative_left": "The internal-beta date is currently scheduled for Monday, June 2, 2025.", "negative_right": "The proposed internal-beta date is Thursday, June 5, 2025.", "right": "The proposed internal-beta date is Wednesday, June 11, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-044-002", "id": "fast-43-diverse-044-002-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta date is currently scheduled for Monday, June 2, 2025."}, {"speaker": "Change requester", "text": "The proposed internal-beta date is Wednesday, June 11, 2025."}, {"speaker": "Change requester", "text": "Two authentication defects require retesting after the contractor's departure, and that is the rationale for this request."}, {"speaker": "Delivery lead", "text": "My signed schedule estimate confirms the added business days needed."}, {"speaker": "Delivery lead", "text": "The analytics dependency owner, the only affected dependency owner, has provided written acknowledgement that the handoff can shift accordingly."}, {"speaker": "Project manager", "text": "No budget change and no scope-baseline change are included in this request."}, {"speaker": "Project manager", "text": "There is no contractual milestone miss associated with this change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy via the unchanged questions object, preserve the same entity/path bindings while only altering the proposed date, use two complete factual sentences as evidence, and contain no gold answer, rule table, or instructional leakage; the counterfactual's shifted date (Jun 2→Jun 5) is internally consistent with the unchanged dependency and rationale statements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta date is currently scheduled for Monday, June 2, 2025.\"},{\"speaker\":\"Change requester\",\"text\":\"The proposed internal-beta date is Wednesday, June 11, 2025.\"},{\"speaker\":\"Change requester\",\"text\":\"Two authentication defects require retesting after the contractor's departure, and that is the rationale for this request.\"},{\"speaker\":\"Delivery lead\",\"text\":\"My signed schedule estimate confirms the added business days needed.\"},{\"speaker\":\"Delivery lead\",\"text\":\"The analytics dependency owner, the only affected dependency owner, has provided written acknowledgement that the handoff can shift accordingly.\"},{\"speaker\":\"Project manager\",\"text\":\"No budget change and no scope-baseline change are included in this request.\"},{\"speaker\":\"Project manager\",\"text\":\"There is no contractual milestone miss associated with this change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta date is currently scheduled for Monday, June 2, 2025."}, {"path": ["1", "text"], "text": "The proposed internal-beta date is Wednesday, June 11, 2025."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta date is currently scheduled for Monday, June 2, 2025.", "negative_left": "The internal-beta date is currently scheduled for Monday, June 2, 2025.", "negative_right": "The proposed internal-beta date is Thursday, June 5, 2025.", "right": "The proposed internal-beta date is Wednesday, June 11, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-044-002", "id": "fast-43-diverse-044-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta date is currently scheduled for Monday, June 2, 2025."}, {"speaker": "Change requester", "text": "The proposed internal-beta date is Thursday, June 5, 2025."}, {"speaker": "Change requester", "text": "Two authentication defects require retesting after the contractor's departure, and that is the rationale for this request."}, {"speaker": "Delivery lead", "text": "My signed schedule estimate confirms the added business days needed."}, {"speaker": "Delivery lead", "text": "The analytics dependency owner, the only affected dependency owner, has provided written acknowledgement that the handoff can shift accordingly."}, {"speaker": "Project manager", "text": "No budget change and no scope-baseline change are included in this request."}, {"speaker": "Project manager", "text": "There is no contractual milestone miss associated with this change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, dependency acknowledgment, rationale, signed estimate, and unchanged budget/scope/milestone facts, only varying the proposed date, and the unchanged question retains all policy thresholds; evidence spans are two verbatim factual sentences with no leaked labels or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta date is currently scheduled for Monday, June 2, 2025. The change request proposes moving the internal-beta date to Wednesday, June 11, 2025. Scope and budget stay unchanged; this is purely a scheduling adjustment tied to contractor availability.\"},{\"speaker\":\"Delivery lead\",\"text\":\"I've included a rationale for the move and a signed schedule estimate covering the revised timeline.\"},{\"speaker\":\"Dependency owners\",\"text\":\"Every dependency owner affected by this request has provided written acknowledgement of the change.\"},{\"speaker\":\"Project manager\",\"text\":\"No spending request accompanies this change, and the scope baseline is untouched. No contractual milestone is affected by the move.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta date is currently scheduled for Monday, June 2, 2025."}, {"path": ["0", "text"], "text": "The change request proposes moving the internal-beta date to Wednesday, June 11, 2025."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta date is currently scheduled for Monday, June 2, 2025.", "negative_left": "The internal-beta date is currently scheduled for Monday, June 2, 2025.", "negative_right": "The change request proposes moving the internal-beta date to Thursday, June 5, 2025.", "right": "The change request proposes moving the internal-beta date to Wednesday, June 11, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-044-007", "id": "fast-43-diverse-044-007-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta date is currently scheduled for Monday, June 2, 2025. The change request proposes moving the internal-beta date to Wednesday, June 11, 2025. Scope and budget stay unchanged; this is purely a scheduling adjustment tied to contractor availability."}, {"speaker": "Delivery lead", "text": "I've included a rationale for the move and a signed schedule estimate covering the revised timeline."}, {"speaker": "Dependency owners", "text": "Every dependency owner affected by this request has provided written acknowledgement of the change."}, {"speaker": "Project manager", "text": "No spending request accompanies this change, and the scope baseline is untouched. No contractual milestone is affected by the move."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, dependency acknowledgment, rationale, signed estimate, and unchanged budget/scope/milestone facts, only varying the proposed date, and the unchanged question retains all policy thresholds; evidence spans are two verbatim factual sentences with no leaked labels or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta date is currently scheduled for Monday, June 2, 2025. The change request proposes moving the internal-beta date to Wednesday, June 11, 2025. Scope and budget stay unchanged; this is purely a scheduling adjustment tied to contractor availability.\"},{\"speaker\":\"Delivery lead\",\"text\":\"I've included a rationale for the move and a signed schedule estimate covering the revised timeline.\"},{\"speaker\":\"Dependency owners\",\"text\":\"Every dependency owner affected by this request has provided written acknowledgement of the change.\"},{\"speaker\":\"Project manager\",\"text\":\"No spending request accompanies this change, and the scope baseline is untouched. No contractual milestone is affected by the move.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta date is currently scheduled for Monday, June 2, 2025."}, {"path": ["0", "text"], "text": "The change request proposes moving the internal-beta date to Wednesday, June 11, 2025."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta date is currently scheduled for Monday, June 2, 2025.", "negative_left": "The internal-beta date is currently scheduled for Monday, June 2, 2025.", "negative_right": "The change request proposes moving the internal-beta date to Thursday, June 5, 2025.", "right": "The change request proposes moving the internal-beta date to Wednesday, June 11, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-044-007", "id": "fast-43-diverse-044-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta date is currently scheduled for Monday, June 2, 2025. The change request proposes moving the internal-beta date to Thursday, June 5, 2025. Scope and budget stay unchanged; this is purely a scheduling adjustment tied to contractor availability."}, {"speaker": "Delivery lead", "text": "I've included a rationale for the move and a signed schedule estimate covering the revised timeline."}, {"speaker": "Dependency owners", "text": "Every dependency owner affected by this request has provided written acknowledgement of the change."}, {"speaker": "Project manager", "text": "No spending request accompanies this change, and the scope baseline is untouched. No contractual milestone is affected by the move."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy via the unchanged questions object and contain no leaked labels; the counterfactual only swaps the new start date to March 6, remaining internally consistent with all other unchanged statements; the focus evidence consists of two verbatim factual sentences about the current and proposed dates.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta phase is currently scheduled to begin on Monday, March 3. The change request proposes moving the internal-beta start to Wednesday, March 12. Scope and budget stay unchanged. The rationale is that two authentication defects need retesting after our contractor departs.\"},{\"speaker\":\"Delivery lead\",\"text\":\"My signed schedule estimate confirms the new date. The analytics dependency owner has provided written acknowledgement of the request, as has every other affected dependency owner.\"},{\"speaker\":\"Project manager\",\"text\":\"No spending request or scope-baseline alteration is included, and no contractual milestone will be missed by this move.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta phase is currently scheduled to begin on Monday, March 3."}, {"path": ["0", "text"], "text": "The change request proposes moving the internal-beta start to Wednesday, March 12."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta phase is currently scheduled to begin on Monday, March 3.", "negative_left": "The internal-beta phase is currently scheduled to begin on Monday, March 3.", "negative_right": "The change request proposes moving the internal-beta start to Thursday, March 6.", "right": "The change request proposes moving the internal-beta start to Wednesday, March 12."}, "verifier_independent_model": false}, "family": "fast-43-diverse-044-019", "id": "fast-43-diverse-044-019-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta phase is currently scheduled to begin on Monday, March 3. The change request proposes moving the internal-beta start to Wednesday, March 12. Scope and budget stay unchanged. The rationale is that two authentication defects need retesting after our contractor departs."}, {"speaker": "Delivery lead", "text": "My signed schedule estimate confirms the new date. The analytics dependency owner has provided written acknowledgement of the request, as has every other affected dependency owner."}, {"speaker": "Project manager", "text": "No spending request or scope-baseline alteration is included, and no contractual milestone will be missed by this move."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy via the unchanged questions object and contain no leaked labels; the counterfactual only swaps the new start date to March 6, remaining internally consistent with all other unchanged statements; the focus evidence consists of two verbatim factual sentences about the current and proposed dates.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta phase is currently scheduled to begin on Monday, March 3. The change request proposes moving the internal-beta start to Wednesday, March 12. Scope and budget stay unchanged. The rationale is that two authentication defects need retesting after our contractor departs.\"},{\"speaker\":\"Delivery lead\",\"text\":\"My signed schedule estimate confirms the new date. The analytics dependency owner has provided written acknowledgement of the request, as has every other affected dependency owner.\"},{\"speaker\":\"Project manager\",\"text\":\"No spending request or scope-baseline alteration is included, and no contractual milestone will be missed by this move.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta phase is currently scheduled to begin on Monday, March 3."}, {"path": ["0", "text"], "text": "The change request proposes moving the internal-beta start to Wednesday, March 12."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta phase is currently scheduled to begin on Monday, March 3.", "negative_left": "The internal-beta phase is currently scheduled to begin on Monday, March 3.", "negative_right": "The change request proposes moving the internal-beta start to Thursday, March 6.", "right": "The change request proposes moving the internal-beta start to Wednesday, March 12."}, "verifier_independent_model": false}, "family": "fast-43-diverse-044-019", "id": "fast-43-diverse-044-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta phase is currently scheduled to begin on Monday, March 3. The change request proposes moving the internal-beta start to Thursday, March 6. Scope and budget stay unchanged. The rationale is that two authentication defects need retesting after our contractor departs."}, {"speaker": "Delivery lead", "text": "My signed schedule estimate confirms the new date. The analytics dependency owner has provided written acknowledgement of the request, as has every other affected dependency owner."}, {"speaker": "Project manager", "text": "No spending request or scope-baseline alteration is included, and no contractual milestone will be missed by this move."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy via the unchanged questions object and only vary the proposed new start date, which is a legitimate observation change; the evidence spans are two complete factual sentences with no policy text; the counterfactual (June 5) is internally consistent with all other unchanged facts and introduces no contradictory measurements; no gold answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025.\"},{\"speaker\":\"Change requester\",\"text\":\"The internal-beta change request proposes moving the internal-beta start date to Wednesday, June 11, 2025. The rationale is that two authentication defects need retesting after our contractor departs.\"},{\"speaker\":\"Delivery lead\",\"text\":\"My signed schedule estimate confirms the additional business days required. The analytics dependency owner has provided written acknowledgement of the request, and no other dependency owners are affected.\"},{\"speaker\":\"Project manager\",\"text\":\"No spending request or scope-baseline alteration is included, and no contractual milestone is at risk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025."}, {"path": ["1", "text"], "text": "The internal-beta change request proposes moving the internal-beta start date to Wednesday, June 11, 2025."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025.", "negative_left": "The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025.", "negative_right": "The internal-beta change request proposes moving the internal-beta start date to Thursday, June 5, 2025.", "right": "The internal-beta change request proposes moving the internal-beta start date to Wednesday, June 11, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-044-029", "id": "fast-43-diverse-044-029-base", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025."}, {"speaker": "Change requester", "text": "The internal-beta change request proposes moving the internal-beta start date to Wednesday, June 11, 2025. The rationale is that two authentication defects need retesting after our contractor departs."}, {"speaker": "Delivery lead", "text": "My signed schedule estimate confirms the additional business days required. The analytics dependency owner has provided written acknowledgement of the request, and no other dependency owners are affected."}, {"speaker": "Project manager", "text": "No spending request or scope-baseline alteration is included, and no contractual milestone is at risk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "sponsor_ready_high"}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy via the unchanged questions object and only vary the proposed new start date, which is a legitimate observation change; the evidence spans are two complete factual sentences with no policy text; the counterfactual (June 5) is internally consistent with all other unchanged facts and introduces no contradictory measurements; no gold answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "refuted", "a8": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified acknowledgement atom. The focus atom is the factual duration threshold, not a policy conclusion. The base and counter assignments differ only on that focus and are jointly realizable: the base can describe a move over five days, while the counter can describe a 3–5-day move. Empty policy_evidence is correct because all governing rules are contained in the automatically retained questions object; no substantive interpretive policy from the original state needs preservation.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "More than five business days necessarily routes ownership to the Project sponsor and makes impact High. Supported rationale, signed estimate, and acknowledgements from every affected dependency owner establish readiness. Budget, scope, and milestone conditions cannot create a competing outcome here.", "rule_index": 0, "sound": true}, {"reason": "Not more than five business days together with at least three business days establishes the 3–5-day range. Refuted budget and scope-baseline changes establish Delivery lead ownership; refuted contractual milestone miss establishes Moderate impact; and all three readiness requirements are supported.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans more than five business days."}, {"id": "a2", "statement": "The requested move from the currently scheduled internal-beta date to the proposed internal-beta date spans at least three business days."}, {"id": "a3", "statement": "The internal-beta change request includes a rationale for the requested move."}, {"id": "a4", "statement": "The internal-beta change request includes a signed schedule estimate."}, {"id": "a5", "statement": "Every dependency owner affected by the internal-beta change request has provided written acknowledgement of the request."}, {"id": "a6", "statement": "The internal-beta change request changes the budget."}, {"id": "a7", "statement": "The internal-beta change request changes the scope baseline."}, {"id": "a8", "statement": "The internal-beta change request causes a contractual milestone miss."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025.\"},{\"speaker\":\"Change requester\",\"text\":\"The internal-beta change request proposes moving the internal-beta start date to Wednesday, June 11, 2025. The rationale is that two authentication defects need retesting after our contractor departs.\"},{\"speaker\":\"Delivery lead\",\"text\":\"My signed schedule estimate confirms the additional business days required. The analytics dependency owner has provided written acknowledgement of the request, and no other dependency owners are affected.\"},{\"speaker\":\"Project manager\",\"text\":\"No spending request or scope-baseline alteration is included, and no contractual milestone is at risk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025."}, {"path": ["1", "text"], "text": "The internal-beta change request proposes moving the internal-beta start date to Wednesday, June 11, 2025."}], "policy_evidence": [], "rules": [{"justification": "A requested move spanning more than five business days is owned by the Project sponsor and has High delivery impact. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make the request ready for review.", "target": "sponsor_ready_high", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}]}, {"justification": "A requested move of at least three but not more than five business days is within the 3–5-day range. With no budget change and no scope-baseline change, the Delivery lead owns it. With no contractual milestone miss, its impact is Moderate. The included rationale, signed schedule estimate, and written acknowledgement from every affected dependency owner make it ready for review.", "target": "delivery_lead_ready_moderate", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025.", "negative_left": "The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025.", "negative_right": "The internal-beta change request proposes moving the internal-beta start date to Thursday, June 5, 2025.", "right": "The internal-beta change request proposes moving the internal-beta start date to Wednesday, June 11, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-044-029", "id": "fast-43-diverse-044-029-counterfactual", "input": {"questions": {"decision": {"criteria": {"delivery_lead_not_ready_low": "Route to the Delivery lead; not ready for review; Low delivery impact.", "delivery_lead_ready_moderate": "Route to the Delivery lead; ready for review; Moderate delivery impact.", "sponsor_ready_high": "Route to the Project sponsor; ready for review; High delivery impact."}, "instructions": "Classify the request's review owner, readiness, and delivery impact. Policy: the Delivery lead owns changes of five or fewer business days with no budget or scope-baseline change; otherwise the Project sponsor owns them. A request is ready only when it includes a rationale, a signed schedule estimate, and written acknowledgement from every affected dependency owner. Impact is Low for 0–2 days, Moderate for 3–5 days without a contractual milestone miss, and High for more than 5 days or any contractual milestone miss. Promotional dates and estimated avoided costs do not alter these tests.", "type": "choice"}}, "state": [{"speaker": "Change requester", "text": "The internal-beta phase is currently scheduled to begin on Monday, June 2, 2025."}, {"speaker": "Change requester", "text": "The internal-beta change request proposes moving the internal-beta start date to Thursday, June 5, 2025. The rationale is that two authentication defects need retesting after our contractor departs."}, {"speaker": "Delivery lead", "text": "My signed schedule estimate confirms the additional business days required. The analytics dependency owner has provided written acknowledgement of the request, and no other dependency owners are affected."}, {"speaker": "Project manager", "text": "No spending request or scope-baseline alteration is included, and no contractual milestone is at risk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-044", "source_is_synthetic": true, "source_sha256": "8224c888c2b62b5b25b52be3be35eaa56b5fefa134d821a086114f12ea0768ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "delivery_lead_ready_moderate"}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the identical governing policy and preserve the original question's entities and scope; the two evidence spans are complete factual statements about the pilot report and Mara's paraphrase; the counterfactual coherently swaps the paraphrase's percentage to create an inaccurate paraphrase without contradicting other facts; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will improve report accuracy for pilot users. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in cost. The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases. Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 42 percent across 500 test cases. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 42 percent across 500 test cases."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases.", "negative_left": "The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 90 percent across 500 test cases.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 42 percent across 500 test cases."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-003", "id": "fast-43-diverse-045-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will improve report accuracy for pilot users. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in cost. The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases. Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 42 percent across 500 test cases. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the identical governing policy and preserve the original question's entities and scope; the two evidence spans are complete factual statements about the pilot report and Mara's paraphrase; the counterfactual coherently swaps the paraphrase's percentage to create an inaccurate paraphrase without contradicting other facts; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will improve report accuracy for pilot users. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in cost. The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases. Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 42 percent across 500 test cases. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 42 percent across 500 test cases."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases.", "negative_left": "The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 90 percent across 500 test cases.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 42 percent across 500 test cases."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-003", "id": "fast-43-diverse-045-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will improve report accuracy for pilot users. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in cost. The attached pilot report states that the filtering prototype reduced query time by 42 percent across 500 test cases. Mara's request paraphrases the pilot report as showing the filtering prototype reduced query time by 90 percent across 500 test cases. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy and same entity/scope/time bindings, use two complete factual evidentiary sentences, and the counterfactual coherently flips the paraphrase's numeric detail to create a mismatch without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, submits a request to add advanced filtering to the reporting module, an addition beyond the module's original scope. Her request includes a rationale for the change and specifies quantified delivery impact: the delivery lead estimates five additional weeks and $60,000 in cost. Supporting her rationale, Mara cites the attached pilot report. The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases. Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases."}, {"path": [], "text": "Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases.", "negative_left": "The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases.", "negative_right": "Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 30 seconds across 200 test cases.", "right": "Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-018", "id": "fast-43-diverse-045-018-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, submits a request to add advanced filtering to the reporting module, an addition beyond the module's original scope. Her request includes a rationale for the change and specifies quantified delivery impact: the delivery lead estimates five additional weeks and $60,000 in cost. Supporting her rationale, Mara cites the attached pilot report. The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases. Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy and same entity/scope/time bindings, use two complete factual evidentiary sentences, and the counterfactual coherently flips the paraphrase's numeric detail to create a mismatch without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, submits a request to add advanced filtering to the reporting module, an addition beyond the module's original scope. Her request includes a rationale for the change and specifies quantified delivery impact: the delivery lead estimates five additional weeks and $60,000 in cost. Supporting her rationale, Mara cites the attached pilot report. The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases. Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases."}, {"path": [], "text": "Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases.", "negative_left": "The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases.", "negative_right": "Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 30 seconds across 200 test cases.", "right": "Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-018", "id": "fast-43-diverse-045-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, submits a request to add advanced filtering to the reporting module, an addition beyond the module's original scope. Her request includes a rationale for the change and specifies quantified delivery impact: the delivery lead estimates five additional weeks and $60,000 in cost. Supporting her rationale, Mara cites the attached pilot report. The attached pilot report states that filtering reduced average query resolution time from 42 seconds to 11 seconds across 200 test cases. Mara's evidence paraphrase in the advanced-filtering request states that filtering reduced average query resolution time from 42 seconds to 30 seconds across 200 test cases. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scope-change routing and readiness policy matching the original question's criteria, preserve Mara's entity and impact-threshold bindings, use two factual (non-policy) evidence sentences, and the counterfactual's altered paraphrase number creates an intentional but coherent mismatch without duplicating contradictory measurements or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This constitutes a scope change under project policy. Her request includes a stated rationale: users need faster filtering to improve export workflow efficiency. The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot. Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 1.1 seconds. The request also includes quantified delivery impact: an estimated five additional weeks and $60,000 in delivery cost. The delivery lead separately confirms these figures: five weeks of additional schedule time and $60,000 in additional cost. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 1.1 seconds."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot.", "negative_left": "The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot.", "negative_right": "Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 3.9 seconds.", "right": "Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 1.1 seconds."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-028", "id": "fast-43-diverse-045-028-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This constitutes a scope change under project policy. Her request includes a stated rationale: users need faster filtering to improve export workflow efficiency. The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot. Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 1.1 seconds. The request also includes quantified delivery impact: an estimated five additional weeks and $60,000 in delivery cost. The delivery lead separately confirms these figures: five weeks of additional schedule time and $60,000 in additional cost. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scope-change routing and readiness policy matching the original question's criteria, preserve Mara's entity and impact-threshold bindings, use two factual (non-policy) evidence sentences, and the counterfactual's altered paraphrase number creates an intentional but coherent mismatch without duplicating contradictory measurements or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This constitutes a scope change under project policy. Her request includes a stated rationale: users need faster filtering to improve export workflow efficiency. The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot. Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 1.1 seconds. The request also includes quantified delivery impact: an estimated five additional weeks and $60,000 in delivery cost. The delivery lead separately confirms these figures: five weeks of additional schedule time and $60,000 in additional cost. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 1.1 seconds."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot.", "negative_left": "The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot.", "negative_right": "Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 3.9 seconds.", "right": "Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 1.1 seconds."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-028", "id": "fast-43-diverse-045-028-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This constitutes a scope change under project policy. Her request includes a stated rationale: users need faster filtering to improve export workflow efficiency. The attached pilot report states that filter response times improved from 4.2 seconds to 1.1 seconds after the pilot. Mara's request paraphrases the pilot report as showing filter response times improved from 4.2 seconds to 3.9 seconds. The request also includes quantified delivery impact: an estimated five additional weeks and $60,000 in delivery cost. The delivery lead separately confirms these figures: five weeks of additional schedule time and $60,000 in additional cost. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy and entity/path bindings from the question, the two evidence spans are factual report/paraphrase statements, the counterfactual introduces a plausible inaccurate paraphrase without contradicting other facts, and neither context reveals a gold answer or instructs the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. She provides a rationale: faster filtering will let analysts drill into large datasets without manual exports. This is a scope change, which under policy routes to the project sponsor. The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries. Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries. The delivery lead estimates the change will require five additional weeks and $60,000 in cost, and this quantified impact is included in the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries."}, {"path": [], "text": "Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries.", "negative_left": "The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries.", "negative_right": "Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 30 seconds across 500 test queries.", "right": "Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-032", "id": "fast-43-diverse-045-032-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. She provides a rationale: faster filtering will let analysts drill into large datasets without manual exports. This is a scope change, which under policy routes to the project sponsor. The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries. Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries. The delivery lead estimates the change will require five additional weeks and $60,000 in cost, and this quantified impact is included in the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy and entity/path bindings from the question, the two evidence spans are factual report/paraphrase statements, the counterfactual introduces a plausible inaccurate paraphrase without contradicting other facts, and neither context reveals a gold answer or instructs the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. She provides a rationale: faster filtering will let analysts drill into large datasets without manual exports. This is a scope change, which under policy routes to the project sponsor. The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries. Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries. The delivery lead estimates the change will require five additional weeks and $60,000 in cost, and this quantified impact is included in the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries."}, {"path": [], "text": "Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries.", "negative_left": "The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries.", "negative_right": "Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 30 seconds across 500 test queries.", "right": "Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-032", "id": "fast-43-diverse-045-032-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. She provides a rationale: faster filtering will let analysts drill into large datasets without manual exports. This is a scope change, which under policy routes to the project sponsor. The attached pilot report states that the filtering feature reduced average query time from 40 seconds to 12 seconds across 500 test queries. Mara's request summarizes the pilot report by stating that the filtering feature reduced average query time from 40 seconds to 30 seconds across 500 test queries. The delivery lead estimates the change will require five additional weeks and $60,000 in cost, and this quantified impact is included in the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy and question bindings, present two factual evidence sentences without policy or answer leakage, and the counterfactual's single altered sentence (12% vs 40%) creates a coherent mismatch between claimed and actual pilot results rather than a contradictory duplicate.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Case note: Mara, the change requester, submits a request to add advanced filtering to the reporting module — a scope change from the original project baseline. Her submission includes a rationale explaining that filtering is needed to speed up manual reporting workflows. Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time. The attached pilot report shows a 40% reduction in manual query time. The request also states the quantified delivery impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the feature. The project manager reviews the request folder, confirms the rationale and quantified figures are present, and must apply the standing policy below to decide on readiness, routing, and impact classification. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time."}, {"path": [], "text": "The attached pilot report shows a 40% reduction in manual query time."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time.", "negative_left": "Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time.", "negative_right": "The attached pilot report shows a 12% reduction in manual query time.", "right": "The attached pilot report shows a 40% reduction in manual query time."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-054", "id": "fast-43-diverse-045-054-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Case note: Mara, the change requester, submits a request to add advanced filtering to the reporting module — a scope change from the original project baseline. Her submission includes a rationale explaining that filtering is needed to speed up manual reporting workflows. Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time. The attached pilot report shows a 40% reduction in manual query time. The request also states the quantified delivery impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the feature. The project manager reviews the request folder, confirms the rationale and quantified figures are present, and must apply the standing policy below to decide on readiness, routing, and impact classification. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy and question bindings, present two factual evidence sentences without policy or answer leakage, and the counterfactual's single altered sentence (12% vs 40%) creates a coherent mismatch between claimed and actual pilot results rather than a contradictory duplicate.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Case note: Mara, the change requester, submits a request to add advanced filtering to the reporting module — a scope change from the original project baseline. Her submission includes a rationale explaining that filtering is needed to speed up manual reporting workflows. Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time. The attached pilot report shows a 40% reduction in manual query time. The request also states the quantified delivery impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the feature. The project manager reviews the request folder, confirms the rationale and quantified figures are present, and must apply the standing policy below to decide on readiness, routing, and impact classification. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time."}, {"path": [], "text": "The attached pilot report shows a 40% reduction in manual query time."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time.", "negative_left": "Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time.", "negative_right": "The attached pilot report shows a 12% reduction in manual query time.", "right": "The attached pilot report shows a 40% reduction in manual query time."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-054", "id": "fast-43-diverse-045-054-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Case note: Mara, the change requester, submits a request to add advanced filtering to the reporting module — a scope change from the original project baseline. Her submission includes a rationale explaining that filtering is needed to speed up manual reporting workflows. Mara's advanced-filtering request states that the attached pilot report showed a 40% reduction in manual query time. The attached pilot report shows a 12% reduction in manual query time. The request also states the quantified delivery impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the feature. The project manager reviews the request folder, confirms the rationale and quantified figures are present, and must apply the standing policy below to decide on readiness, routing, and impact classification. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy, entities, and six-week trial binding; the counterfactual's altered paraphrase (35% vs the report's 12%) deliberately creates an inaccurate-paraphrase scenario rather than an incoherent duplicate measurement, and neither context reveals the answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, submits a request to add advanced filtering to the reporting module. She explains that current users cannot narrow large report sets, which is her stated rationale for the change. The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial. Mara's paraphrase in the request states that the pilot report found user adoption rose 12% during the six-week trial. She also cites the delivery lead's quantified estimate: five additional weeks of schedule and $60,000 in added cost. Because this alters previously agreed deliverables, the request qualifies as a scope change. The project manager reviews the file, the rationale, the pilot report, and the cost and schedule figures before applying the standing policy. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000\\u2013$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial."}, {"path": [], "text": "Mara's paraphrase in the request states that the pilot report found user adoption rose 12% during the six-week trial."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial.", "negative_left": "The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial.", "negative_right": "Mara's paraphrase in the request states that the pilot report found user adoption rose 35% during the six-week trial.", "right": "Mara's paraphrase in the request states that the pilot report found user adoption rose 12% during the six-week trial."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-060", "id": "fast-43-diverse-045-060-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, submits a request to add advanced filtering to the reporting module. She explains that current users cannot narrow large report sets, which is her stated rationale for the change. The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial. Mara's paraphrase in the request states that the pilot report found user adoption rose 12% during the six-week trial. She also cites the delivery lead's quantified estimate: five additional weeks of schedule and $60,000 in added cost. Because this alters previously agreed deliverables, the request qualifies as a scope change. The project manager reviews the file, the rationale, the pilot report, and the cost and schedule figures before applying the standing policy. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy, entities, and six-week trial binding; the counterfactual's altered paraphrase (35% vs the report's 12%) deliberately creates an inaccurate-paraphrase scenario rather than an incoherent duplicate measurement, and neither context reveals the answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, submits a request to add advanced filtering to the reporting module. She explains that current users cannot narrow large report sets, which is her stated rationale for the change. The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial. Mara's paraphrase in the request states that the pilot report found user adoption rose 12% during the six-week trial. She also cites the delivery lead's quantified estimate: five additional weeks of schedule and $60,000 in added cost. Because this alters previously agreed deliverables, the request qualifies as a scope change. The project manager reviews the file, the rationale, the pilot report, and the cost and schedule figures before applying the standing policy. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000\\u2013$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial."}, {"path": [], "text": "Mara's paraphrase in the request states that the pilot report found user adoption rose 12% during the six-week trial."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial.", "negative_left": "The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial.", "negative_right": "Mara's paraphrase in the request states that the pilot report found user adoption rose 35% during the six-week trial.", "right": "Mara's paraphrase in the request states that the pilot report found user adoption rose 12% during the six-week trial."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-060", "id": "fast-43-diverse-045-060-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, submits a request to add advanced filtering to the reporting module. She explains that current users cannot narrow large report sets, which is her stated rationale for the change. The pilot report attached to Mara's advanced-filtering request states that user adoption rose 12% during the six-week trial. Mara's paraphrase in the request states that the pilot report found user adoption rose 35% during the six-week trial. She also cites the delivery lead's quantified estimate: five additional weeks of schedule and $60,000 in added cost. Because this alters previously agreed deliverables, the request qualifies as a scope change. The project manager reviews the file, the rationale, the pilot report, and the cost and schedule figures before applying the standing policy. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing/readiness/impact policy text, preserve the same question object and its scope-change/impact bindings, use two complete factual sentences as evidence, and the counterfactual's single-sentence change (35%→12% pilot result) coherently creates an inaccurate paraphrase without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, submits a request to add advanced filtering to the reporting module, a clear scope change to the existing project plan. Her request includes a written rationale explaining that filtering is needed to improve reporting accuracy for downstream teams. Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot. The attached pilot report itself records that manual query time dropped by 35 percent during the pilot. The request also includes quantified impact figures: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The project manager reviews the rationale, the paraphrase, the quantified figures, and the schedule and cost estimates before deciding how to route the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot."}, {"path": [], "text": "The attached pilot report itself records that manual query time dropped by 35 percent during the pilot."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot.", "negative_left": "Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot.", "negative_right": "The attached pilot report itself records that manual query time dropped by 12 percent during the pilot.", "right": "The attached pilot report itself records that manual query time dropped by 35 percent during the pilot."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-064", "id": "fast-43-diverse-045-064-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, submits a request to add advanced filtering to the reporting module, a clear scope change to the existing project plan. Her request includes a written rationale explaining that filtering is needed to improve reporting accuracy for downstream teams. Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot. The attached pilot report itself records that manual query time dropped by 35 percent during the pilot. The request also includes quantified impact figures: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The project manager reviews the rationale, the paraphrase, the quantified figures, and the schedule and cost estimates before deciding how to route the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing/readiness/impact policy text, preserve the same question object and its scope-change/impact bindings, use two complete factual sentences as evidence, and the counterfactual's single-sentence change (35%→12% pilot result) coherently creates an inaccurate paraphrase without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, submits a request to add advanced filtering to the reporting module, a clear scope change to the existing project plan. Her request includes a written rationale explaining that filtering is needed to improve reporting accuracy for downstream teams. Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot. The attached pilot report itself records that manual query time dropped by 35 percent during the pilot. The request also includes quantified impact figures: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The project manager reviews the rationale, the paraphrase, the quantified figures, and the schedule and cost estimates before deciding how to route the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot."}, {"path": [], "text": "The attached pilot report itself records that manual query time dropped by 35 percent during the pilot."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot.", "negative_left": "Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot.", "negative_right": "The attached pilot report itself records that manual query time dropped by 12 percent during the pilot.", "right": "The attached pilot report itself records that manual query time dropped by 35 percent during the pilot."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-064", "id": "fast-43-diverse-045-064-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, submits a request to add advanced filtering to the reporting module, a clear scope change to the existing project plan. Her request includes a written rationale explaining that filtering is needed to improve reporting accuracy for downstream teams. Mara's advanced-filtering request paraphrases the attached pilot report as stating that manual query time dropped by 35 percent during the pilot. The attached pilot report itself records that manual query time dropped by 12 percent during the pilot. The request also includes quantified impact figures: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The project manager reviews the rationale, the paraphrase, the quantified figures, and the schedule and cost estimates before deciding how to route the request. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy verbatim and only vary the pilot report's actual reduction figure, creating a coherent factual contrast without embedding any answer, rule table, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. She explains that users struggle to export the reports they need, citing pilot data as support. Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time. The attached pilot report itself records a 40% reduction in manual export time. This is a scope change to the existing reporting module. Mara's request includes a rationale describing user pain points and includes quantified impact figures from the delivery lead. The delivery lead estimates the request will require five additional weeks and $60,000 in additional cost to implement. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time."}, {"path": [], "text": "The attached pilot report itself records a 40% reduction in manual export time."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time.", "negative_left": "Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time.", "negative_right": "The attached pilot report itself records only a 5% reduction in manual export time.", "right": "The attached pilot report itself records a 40% reduction in manual export time."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-068", "id": "fast-43-diverse-045-068-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. She explains that users struggle to export the reports they need, citing pilot data as support. Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time. The attached pilot report itself records a 40% reduction in manual export time. This is a scope change to the existing reporting module. Mara's request includes a rationale describing user pain points and includes quantified impact figures from the delivery lead. The delivery lead estimates the request will require five additional weeks and $60,000 in additional cost to implement. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy verbatim and only vary the pilot report's actual reduction figure, creating a coherent factual contrast without embedding any answer, rule table, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. She explains that users struggle to export the reports they need, citing pilot data as support. Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time. The attached pilot report itself records a 40% reduction in manual export time. This is a scope change to the existing reporting module. Mara's request includes a rationale describing user pain points and includes quantified impact figures from the delivery lead. The delivery lead estimates the request will require five additional weeks and $60,000 in additional cost to implement. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time."}, {"path": [], "text": "The attached pilot report itself records a 40% reduction in manual export time."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time.", "negative_left": "Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time.", "negative_right": "The attached pilot report itself records only a 5% reduction in manual export time.", "right": "The attached pilot report itself records a 40% reduction in manual export time."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-068", "id": "fast-43-diverse-045-068-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. She explains that users struggle to export the reports they need, citing pilot data as support. Mara's advanced-filtering request paraphrases the attached pilot report as showing a 40% reduction in manual export time. The attached pilot report itself records only a 5% reduction in manual export time. This is a scope change to the existing reporting module. Mara's request includes a rationale describing user pain points and includes quantified impact figures from the delivery lead. The delivery lead estimates the request will require five additional weeks and $60,000 in additional cost to implement. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy and quantified impact bindings (5 weeks, $60,000, scope change to sponsor); the two evidence sentences are complete factual statements, not policy text; the counterfactual alters only the paraphrase figure to 2 minutes, creating an inaccurate-paraphrase scenario that is coherent rather than duplicative; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: users need to slice reports by region and date to make faster decisions. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs.", "negative_left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 2 minutes across 300 test runs.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-069", "id": "fast-43-diverse-045-069-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: users need to slice reports by region and date to make faster decisions. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy and quantified impact bindings (5 weeks, $60,000, scope change to sponsor); the two evidence sentences are complete factual statements, not policy text; the counterfactual alters only the paraphrase figure to 2 minutes, creating an inaccurate-paraphrase scenario that is coherent rather than duplicative; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: users need to slice reports by region and date to make faster decisions. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs.", "negative_left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 2 minutes across 300 test runs.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-069", "id": "fast-43-diverse-045-069-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: users need to slice reports by region and date to make faster decisions. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test runs. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 2 minutes across 300 test runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy and same entities/paths; the two focus_evidence quotes are complete factual sentences describing the pilot report and its paraphrase; the counterfactual changes the paraphrase to 3 minutes, creating a coherent evidentiary mismatch (inaccurate paraphrase) rather than a nonsensical duplicate measurement; no context reveals the answer, rule table, or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let analysts drill into regional sales data without exporting raw files. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in cost. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions.", "negative_left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 3 minutes across 300 test sessions.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-070", "id": "fast-43-diverse-045-070-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let analysts drill into regional sales data without exporting raw files. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in cost. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy and same entities/paths; the two focus_evidence quotes are complete factual sentences describing the pilot report and its paraphrase; the counterfactual changes the paraphrase to 3 minutes, creating a coherent evidentiary mismatch (inaccurate paraphrase) rather than a nonsensical duplicate measurement; no context reveals the answer, rule table, or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let analysts drill into regional sales data without exporting raw files. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in cost. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions.", "negative_left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 3 minutes across 300 test sessions.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-070", "id": "fast-43-diverse-045-070-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let analysts drill into regional sales data without exporting raw files. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in cost. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 test sessions. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 3 minutes across 300 test sessions. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy intact via the unchanged question, preserve Mara/scope-change bindings, use two factual (non-policy) evidence sentences, and the counterfactual introduces a coherent single-fact change (paraphrase now says 2 seconds, contradicting the unchanged pilot report) without duplicating measurements or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let analysts drill into regional sales data without exporting raw tables. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs.", "negative_left": "The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 2 seconds across 300 test runs.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-071", "id": "fast-43-diverse-045-071-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let analysts drill into regional sales data without exporting raw tables. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy intact via the unchanged question, preserve Mara/scope-change bindings, use two factual (non-policy) evidence sentences, and the counterfactual introduces a coherent single-fact change (paraphrase now says 2 seconds, contradicting the unchanged pilot report) without duplicating measurements or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let analysts drill into regional sales data without exporting raw tables. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs.", "negative_left": "The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 2 seconds across 300 test runs.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-071", "id": "fast-43-diverse-045-071-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let analysts drill into regional sales data without exporting raw tables. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The attached pilot report states that the filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 test runs. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report generation time from 12 seconds to 2 seconds across 300 test runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy and the same requester/scope/impact bindings, the two evidence spans are complete factual statements (pilot report result and Mara's paraphrase) not policy text, and the counterfactual's altered paraphrase (2 minutes vs 7) is a single coherent factual edit that creates the intended inaccurate-paraphrase scenario without adding contradictory duplicate measurements or leaking any answer/rule/label information.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let power users narrow reports without manual exports. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions.", "negative_left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 2 minutes across 300 user sessions.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-072", "id": "fast-43-diverse-045-072-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let power users narrow reports without manual exports. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy and the same requester/scope/impact bindings, the two evidence spans are complete factual statements (pilot report result and Mara's paraphrase) not policy text, and the counterfactual's altered paraphrase (2 minutes vs 7) is a single coherent factual edit that creates the intended inaccurate-paraphrase scenario without adding contradictory duplicate measurements or leaking any answer/rule/label information.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let power users narrow reports without manual exports. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions.", "negative_left": "The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions.", "negative_right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 2 minutes across 300 user sessions.", "right": "Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-072", "id": "fast-43-diverse-045-072-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: filtering will let power users narrow reports without manual exports. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the filtering feature. The attached pilot report states that the filtering prototype cut average report-generation time from 12 minutes to 7 minutes across 300 user sessions. Mara's request paraphrases the pilot report as showing the filtering prototype cut average report-generation time from 12 minutes to 2 minutes across 300 user sessions. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy text and the original question bindings, the two evidence spans are complete factual sentences, and the counterfactual's single-sentence change ('7 seconds' to '2 seconds') creates an intentional paraphrase mismatch without contradicting other stated facts or embedding any answer/label information.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: advanced filtering will let analysts isolate anomalous transactions faster during quarterly audits. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the feature. The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs. Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs.", "negative_left": "The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs.", "negative_right": "Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 2 seconds across 300 trial runs.", "right": "Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-073", "id": "fast-43-diverse-045-073-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: advanced filtering will let analysts isolate anomalous transactions faster during quarterly audits. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the feature. The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs. Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy text and the original question bindings, the two evidence spans are complete factual sentences, and the counterfactual's single-sentence change ('7 seconds' to '2 seconds') creates an intentional paraphrase mismatch without contradicting other stated facts or embedding any answer/label information.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "full_context_fact_states": {"base": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "supported", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "counterfactual": {"cost_estimate": "supported", "evidence_paraphrase_accurate": "refuted", "quantified_impact_included": "supported", "rationale_included": "supported", "schedule_estimate": "supported", "scope_change": "supported"}, "remove_left": {"evidence_paraphrase_accurate": "unknown"}, "remove_right": {"evidence_paraphrase_accurate": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"evidence_paraphrase_accurate": "unknown"}, "negative_pair": {"evidence_paraphrase_accurate": "refuted"}, "negative_sentence": {"evidence_paraphrase_accurate": "unknown"}, "positive_pair": {"evidence_paraphrase_accurate": "supported"}, "right": {"evidence_paraphrase_accurate": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the focus atom concerns factual accuracy rather than policy. The base and counter assignments can be realized with only the accuracy relationship changing while the other stated facts remain fixed. The policy evidence correctly preserves the substantive policy originating in the original state; decision criteria retained in original_input.questions need not be repeated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted evidence_paraphrase_accurate atom entails that the paraphrase is inaccurate. The unchanged decision criteria explicitly state that the request must not be marked ready in that circumstance, so the conjunction of all three proposed actions cannot be correct.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish all three outcomes: the supported scope change must go to the project sponsor; the three supported readiness requirements make the request ready; and estimates of five weeks and $60,000 are above the medium limits, so the impact is high.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "scope_change", "statement": "Mara's request to add advanced filtering to the reporting module is a scope change."}, {"id": "rationale_included", "statement": "Mara's request to add advanced filtering to the reporting module includes a rationale."}, {"id": "evidence_paraphrase_accurate", "statement": "Mara's evidence paraphrase in the advanced-filtering request accurately represents the attached pilot report."}, {"id": "quantified_impact_included", "statement": "Mara's request to add advanced filtering to the reporting module includes quantified delivery impact."}, {"id": "schedule_estimate", "statement": "The delivery lead's estimated additional delivery time for Mara's advanced-filtering request is five weeks."}, {"id": "cost_estimate", "statement": "The delivery lead's estimated additional delivery cost for Mara's advanced-filtering request is $60,000."}], "base_state_json": "\"Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: advanced filtering will let analysts isolate anomalous transactions faster during quarterly audits. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the feature. The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs. Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit.\"", "base_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "counter_states": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}], "focus_atom": "evidence_paraphrase_accurate", "focus_evidence": [{"path": [], "text": "The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs."}, {"path": [], "text": "Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs."}], "policy_evidence": [{"path": [], "text": "The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}], "rules": [{"justification": "An inaccurate evidence paraphrase prevents the request from being ready, so at least one of the three proposed actions is incorrect.", "target": "false", "when": [{"atom_id": "evidence_paraphrase_accurate", "state": "refuted"}]}, {"justification": "The scope change goes to the project sponsor; all readiness requirements are met; and five weeks and $60,000 are each above the applicable medium limit, making delivery impact high.", "target": "true", "when": [{"atom_id": "scope_change", "state": "supported"}, {"atom_id": "rationale_included", "state": "supported"}, {"atom_id": "evidence_paraphrase_accurate", "state": "supported"}, {"atom_id": "quantified_impact_included", "state": "supported"}, {"atom_id": "schedule_estimate", "state": "supported"}, {"atom_id": "cost_estimate", "state": "supported"}]}]}, "verified_pair": {"left": "The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs.", "negative_left": "The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs.", "negative_right": "Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 2 seconds across 300 trial runs.", "right": "Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-045-073", "id": "fast-43-diverse-045-073-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one action is incorrect; in particular, the request must not be marked ready if its evidence paraphrase is inaccurate.", "true": "Yes — all three actions are correct: the request is ready, the project sponsor is the proper owner, and the impact is high."}, "instructions": "Answer yes or no: Should the project manager mark this request ready for review, route it to the project sponsor, and classify its delivery impact as high?", "type": "noul"}}, "state": "Mara, the change requester, asks to add advanced filtering to the reporting module. This is a scope change. Her request includes a rationale: advanced filtering will let analysts isolate anomalous transactions faster during quarterly audits. It also includes quantified impact: the delivery lead estimates five additional weeks and $60,000 in additional cost to implement the feature. The attached pilot report states that the advanced-filtering prototype cut average report generation time from 12 seconds to 7 seconds across 300 trial runs. Mara's request paraphrases the pilot report as showing the advanced-filtering prototype cut average report generation time from 12 seconds to 2 seconds across 300 trial runs. The project manager must apply this policy: scope changes go to the project sponsor; a request is ready only when it includes a rationale, an accurate paraphrase of supporting evidence, and quantified impact. Delivery impact is low below two weeks and $20,000, medium at two to four weeks or $20,000–$50,000, and high above either medium limit."}, "method": "c2d", "provenance": {"source_id": "diverse-045", "source_is_synthetic": true, "source_sha256": "e78d2d6846e431466f6352e85e5513f3c737502c026c85e3fd9ff1dd63155cd1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the dates, cost/baseline percentage, staffing, and milestone deadline exactly as bound in the question; the counterfactual only changes the deliverable shift date to June 15, remaining internally consistent and free of any embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\": \"Change requester\", \"text\": \"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"}, {\"speaker\": \"Delivery lead\", \"text\": \"July 14 is exactly 10 business days later. The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 25.\"}, {\"speaker\": \"Project manager\", \"text\": \"The rationale, schedule calculation, staffing need, and cost estimate are documented. The $15,000 cost is exactly 5% of the $300,000 baseline, and two engineers are assigned to the work.\"}, {\"speaker\": \"Project sponsor\", \"text\": \"The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20. Send the complete change packet to me for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 25."}, {"path": ["3", "text"], "text": "The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 25.", "negative_left": "The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 15.", "negative_right": "The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20.", "right": "The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-015", "id": "fast-43-diverse-047-015-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later. The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 25."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented. The $15,000 cost is exactly 5% of the $300,000 baseline, and two engineers are assigned to the work."}, {"speaker": "Project sponsor", "text": "The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20. Send the complete change packet to me for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the dates, cost/baseline percentage, staffing, and milestone deadline exactly as bound in the question; the counterfactual only changes the deliverable shift date to June 15, remaining internally consistent and free of any embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\": \"Change requester\", \"text\": \"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"}, {\"speaker\": \"Delivery lead\", \"text\": \"July 14 is exactly 10 business days later. The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 25.\"}, {\"speaker\": \"Project manager\", \"text\": \"The rationale, schedule calculation, staffing need, and cost estimate are documented. The $15,000 cost is exactly 5% of the $300,000 baseline, and two engineers are assigned to the work.\"}, {\"speaker\": \"Project sponsor\", \"text\": \"The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20. Send the complete change packet to me for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 25."}, {"path": ["3", "text"], "text": "The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 25.", "negative_left": "The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 15.", "negative_right": "The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20.", "right": "The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-015", "id": "fast-43-diverse-047-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later. The encryption-and-key-rotation change request shifts the security certification deliverable date from June 5 to June 15."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented. The $15,000 cost is exactly 5% of the $300,000 baseline, and two engineers are assigned to the work."}, {"speaker": "Project sponsor", "text": "The contract specifies that the security certification deliverable is a contractual milestone due no later than June 20. Send the complete change packet to me for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the unchanged question's ownership/impact policy intact while updating only the deliverable delay date and its relation to the fixed March 10 milestone; the evidence spans are single factual sentences, the counterfactual (March 8 vs March 22) is internally consistent with the unaltered milestone statement, and no gold labels, rule tables, or instructions are embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later. The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 22.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10.\"},{\"speaker\":\"Project sponsor\",\"text\":\"Send the complete change packet to me for review; two engineers are assigned and the added cost is 5% of the $300,000 baseline.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 22."}, {"path": ["2", "text"], "text": "Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 22.", "negative_left": "The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 8.", "negative_right": "Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10.", "right": "Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-017", "id": "fast-43-diverse-047-017-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later. The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 22."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10."}, {"speaker": "Project sponsor", "text": "Send the complete change packet to me for review; two engineers are assigned and the added cost is 5% of the $300,000 baseline."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the unchanged question's ownership/impact policy intact while updating only the deliverable delay date and its relation to the fixed March 10 milestone; the evidence spans are single factual sentences, the counterfactual (March 8 vs March 22) is internally consistent with the unaltered milestone statement, and no gold labels, rule tables, or instructions are embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later. The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 22.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10.\"},{\"speaker\":\"Project sponsor\",\"text\":\"Send the complete change packet to me for review; two engineers are assigned and the added cost is 5% of the $300,000 baseline.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 22."}, {"path": ["2", "text"], "text": "Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 22.", "negative_left": "The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 8.", "negative_right": "Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10.", "right": "Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-017", "id": "fast-43-diverse-047-017-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later. The encryption-and-key-rotation change request delays the key-vault migration deliverable to March 8."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone M-7 for the vendor integration project requires the key-vault migration deliverable to be completed no later than March 10."}, {"speaker": "Project sponsor", "text": "Send the complete change packet to me for review; two engineers are assigned and the added cost is 5% of the $300,000 baseline."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original quantified impacts (10 business days, two engineers, 5% cost) and the M7/March 1 milestone binding while the question's ownership and scoring rules remain verbatim; the two focus sentences are plain factual statements; the counterfactual only swaps the Module X delivery date to Feb 20, 2025, which stays logically consistent with the unchanged milestone clause and introduces no contradictory duplicate facts or leaked answer content.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later; two engineers are assigned to the work.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented; the added cost equals 5% of the $300,000 baseline. The encryption-and-key-rotation change request shifts delivery of Module X to March 15, 2025.\"},{\"speaker\":\"Project sponsor\",\"text\":\"Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025. Send the complete change packet to me for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "The encryption-and-key-rotation change request shifts delivery of Module X to March 15, 2025."}, {"path": ["3", "text"], "text": "Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The encryption-and-key-rotation change request shifts delivery of Module X to March 15, 2025.", "negative_left": "The encryption-and-key-rotation change request shifts delivery of Module X to February 20, 2025.", "negative_right": "Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025.", "right": "Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-022", "id": "fast-43-diverse-047-022-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later; two engineers are assigned to the work."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented; the added cost equals 5% of the $300,000 baseline. The encryption-and-key-rotation change request shifts delivery of Module X to March 15, 2025."}, {"speaker": "Project sponsor", "text": "Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025. Send the complete change packet to me for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original quantified impacts (10 business days, two engineers, 5% cost) and the M7/March 1 milestone binding while the question's ownership and scoring rules remain verbatim; the two focus sentences are plain factual statements; the counterfactual only swaps the Module X delivery date to Feb 20, 2025, which stays logically consistent with the unchanged milestone clause and introduces no contradictory duplicate facts or leaked answer content.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later; two engineers are assigned to the work.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented; the added cost equals 5% of the $300,000 baseline. The encryption-and-key-rotation change request shifts delivery of Module X to March 15, 2025.\"},{\"speaker\":\"Project sponsor\",\"text\":\"Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025. Send the complete change packet to me for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["2", "text"], "text": "The encryption-and-key-rotation change request shifts delivery of Module X to March 15, 2025."}, {"path": ["3", "text"], "text": "Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The encryption-and-key-rotation change request shifts delivery of Module X to March 15, 2025.", "negative_left": "The encryption-and-key-rotation change request shifts delivery of Module X to February 20, 2025.", "negative_right": "Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025.", "right": "Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-022", "id": "fast-43-diverse-047-022-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later; two engineers are assigned to the work."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented; the added cost equals 5% of the $300,000 baseline. The encryption-and-key-rotation change request shifts delivery of Module X to February 20, 2025."}, {"speaker": "Project sponsor", "text": "Milestone M7 in the signed contract requires Module X delivery no later than March 1, 2025. Send the complete change packet to me for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Routing/impact-level policy remains fully specified in the unchanged question, both contexts add only factual scheduling notes tied to the same encryption-and-key-rotation project, the counterfactual changes a single completion-date fact without duplicating or contradicting other measurements, and no gold label, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented.\"},{\"speaker\":\"Scheduling note\",\"text\":\"Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024.\"},{\"speaker\":\"Scheduling note\",\"text\":\"The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 25, 2024.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["3", "text"], "text": "Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024."}, {"path": ["4", "text"], "text": "The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 25, 2024."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024.", "negative_left": "Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024.", "negative_right": "The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 10, 2024.", "right": "The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 25, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-024", "id": "fast-43-diverse-047-024-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented."}, {"speaker": "Scheduling note", "text": "Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024."}, {"speaker": "Scheduling note", "text": "The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 25, 2024."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Routing/impact-level policy remains fully specified in the unchanged question, both contexts add only factual scheduling notes tied to the same encryption-and-key-rotation project, the counterfactual changes a single completion-date fact without duplicating or contradicting other measurements, and no gold label, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented.\"},{\"speaker\":\"Scheduling note\",\"text\":\"Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024.\"},{\"speaker\":\"Scheduling note\",\"text\":\"The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 25, 2024.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["3", "text"], "text": "Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024."}, {"path": ["4", "text"], "text": "The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 25, 2024."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024.", "negative_left": "Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024.", "negative_right": "The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 10, 2024.", "right": "The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 25, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-024", "id": "fast-43-diverse-047-024-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented."}, {"speaker": "Scheduling note", "text": "Contractual milestone M7 for the encryption-and-key-rotation project is fixed at March 15, 2024."}, {"speaker": "Scheduling note", "text": "The encryption-and-key-rotation change request pushes the affected deliverable's completion date to March 10, 2024."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same request, schedule figures, and milestone M-7 dates, with governing ownership/impact policy already fully stated in the question instructions; the two evidence spans are complete factual sentences about the delay and milestone deadline, not policy text; the counterfactual only changes the delay date to March 5, coherently altering milestone compliance without contradicting other facts or leaking any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later. The encryption-and-key-rotation change request delays the security certification deliverable to March 22, 2025.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025.\"},{\"speaker\":\"Project sponsor\",\"text\":\"Send the complete change packet to me for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The encryption-and-key-rotation change request delays the security certification deliverable to March 22, 2025."}, {"path": ["2", "text"], "text": "Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The encryption-and-key-rotation change request delays the security certification deliverable to March 22, 2025.", "negative_left": "The encryption-and-key-rotation change request delays the security certification deliverable to March 5, 2025.", "negative_right": "Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025.", "right": "Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-026", "id": "fast-43-diverse-047-026-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later. The encryption-and-key-rotation change request delays the security certification deliverable to March 22, 2025."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025."}, {"speaker": "Project sponsor", "text": "Send the complete change packet to me for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same request, schedule figures, and milestone M-7 dates, with governing ownership/impact policy already fully stated in the question instructions; the two evidence spans are complete factual sentences about the delay and milestone deadline, not policy text; the counterfactual only changes the delay date to March 5, coherently altering milestone compliance without contradicting other facts or leaking any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later. The encryption-and-key-rotation change request delays the security certification deliverable to March 22, 2025.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025.\"},{\"speaker\":\"Project sponsor\",\"text\":\"Send the complete change packet to me for review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The encryption-and-key-rotation change request delays the security certification deliverable to March 22, 2025."}, {"path": ["2", "text"], "text": "Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The encryption-and-key-rotation change request delays the security certification deliverable to March 22, 2025.", "negative_left": "The encryption-and-key-rotation change request delays the security certification deliverable to March 5, 2025.", "negative_right": "Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025.", "right": "Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-026", "id": "fast-43-diverse-047-026-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later. The encryption-and-key-rotation change request delays the security certification deliverable to March 5, 2025."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone M-7 requires the security certification deliverable to be completed by March 10, 2025."}, {"speaker": "Project sponsor", "text": "Send the complete change packet to me for review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing severity/ownership rules via the unchanged questions object, keep the same request entity and impact dimensions, contain two factual (non-instructional) evidence sentences, and the counterfactual only swaps the final completion date (Feb 20 vs Mar 15) without contradicting other stated facts or leaking a decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later, and $15,000 is exactly 5% of the $300,000 baseline. Two engineers are assigned to this work.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone breaches go to the project sponsor; other delivery changes go to the delivery lead.\"},{\"speaker\":\"Contract note\",\"text\":\"Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025.\"},{\"speaker\":\"Change requester\",\"text\":\"The encryption-and-key-rotation change request schedules key-rotation deployment completion for March 15, 2025.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["3", "text"], "text": "Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025."}, {"path": ["4", "text"], "text": "The encryption-and-key-rotation change request schedules key-rotation deployment completion for March 15, 2025."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025.", "negative_left": "Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025.", "negative_right": "The encryption-and-key-rotation change request schedules key-rotation deployment completion for February 20, 2025.", "right": "The encryption-and-key-rotation change request schedules key-rotation deployment completion for March 15, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-029", "id": "fast-43-diverse-047-029-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later, and $15,000 is exactly 5% of the $300,000 baseline. Two engineers are assigned to this work."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone breaches go to the project sponsor; other delivery changes go to the delivery lead."}, {"speaker": "Contract note", "text": "Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025."}, {"speaker": "Change requester", "text": "The encryption-and-key-rotation change request schedules key-rotation deployment completion for March 15, 2025."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing severity/ownership rules via the unchanged questions object, keep the same request entity and impact dimensions, contain two factual (non-instructional) evidence sentences, and the counterfactual only swaps the final completion date (Feb 20 vs Mar 15) without contradicting other stated facts or leaking a decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom concerns the factual existence of a contractual-milestone breach rather than a policy classification. The base and counter assignments are realizable with only that fact changing: a 10-day, 5% change may breach a contractual milestone in one scenario and not breach one in another. Empty policy evidence is correct because the governing readiness, ownership, threshold, and exception rules are already retained in the questions object; the state adds case observations but no necessary independent policy. Both rules include the exclusions needed to make their targets sufficient.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported contractual-milestone breach is sufficient for level 3 under the explicit criterion that any such breach is severe regardless of lower numeric impacts.", "rule_index": 0, "sound": true}, {"reason": "An exact 10-business-day delay and exact 5% added cost each fall within level 2, not the greater-than-10-day or above-5% level-3 thresholds. Refutation of any contractual-milestone breach excludes the remaining level-3 trigger, so the conjunction is sufficient for level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The rationale for the encryption-and-key-rotation change request is documented."}, {"id": "a2", "statement": "The applicable schedule impact of the encryption-and-key-rotation change request is quantified as a delay of exactly 10 business days."}, {"id": "a3", "statement": "The applicable cost impact of the encryption-and-key-rotation change request is quantified as exactly 5% of the $300,000 baseline."}, {"id": "a4", "statement": "The applicable resource impact of the encryption-and-key-rotation change request is quantified as the assignment of two engineers."}, {"id": "a5", "statement": "The encryption-and-key-rotation change request would breach at least one contractual milestone."}], "base_state_json": "[{\"speaker\":\"Change requester\",\"text\":\"The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline.\"},{\"speaker\":\"Delivery lead\",\"text\":\"July 14 is exactly 10 business days later, and $15,000 is exactly 5% of the $300,000 baseline. Two engineers are assigned to this work.\"},{\"speaker\":\"Project manager\",\"text\":\"The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone breaches go to the project sponsor; other delivery changes go to the delivery lead.\"},{\"speaker\":\"Contract note\",\"text\":\"Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025.\"},{\"speaker\":\"Change requester\",\"text\":\"The encryption-and-key-rotation change request schedules key-rotation deployment completion for March 15, 2025.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["3", "text"], "text": "Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025."}, {"path": ["4", "text"], "text": "The encryption-and-key-rotation change request schedules key-rotation deployment completion for March 15, 2025."}], "policy_evidence": [], "rules": [{"justification": "Any contractual milestone breach is a severe delivery impact regardless of otherwise lower schedule or cost impacts.", "target": "3", "when": [{"atom_id": "a5", "state": "supported"}]}, {"justification": "An exact 10-business-day delay and an exact 5% added cost are within the material ranges, while refutation of any contractual milestone breach excludes the contractual severe-impact trigger; the exact measurements also exclude the greater-than-10-day and above-5% severe thresholds.", "target": "2", "when": [{"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025.", "negative_left": "Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025.", "negative_right": "The encryption-and-key-rotation change request schedules key-rotation deployment completion for February 20, 2025.", "right": "The encryption-and-key-rotation change request schedules key-rotation deployment completion for March 15, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-047-029", "id": "fast-43-diverse-047-029-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: scope wording or administrative details change, with no schedule, cost, staffing, deliverable, or contractual milestone effect.", "1 — Minor delivery impact: schedule movement of 1–3 business days or added cost below 2% of baseline, with no staffing change and no contractual milestone breach.", "2 — Material delivery impact: schedule movement of 4–10 business days, added cost from 2% through 5% of baseline, or temporary staffing reassignment, provided no contractual milestone is breached.", "3 — Severe delivery impact: schedule movement greater than 10 business days, added cost above 5% of baseline, or any breach of a contractual milestone, regardless of otherwise lower numeric impacts."], "instructions": "Triage the request by determining whether it has a documented rationale and quantified impact, whether it is ready for review, and who owns the review. Then select the delivery-impact level. A request is ready when its rationale and applicable schedule, cost, and resource impacts are quantified. The project sponsor owns any request that breaches a contractual milestone; otherwise, the delivery lead owns delivery-impact review.", "type": "score"}}, "state": [{"speaker": "Change requester", "text": "The compliance audit found that the payment export lacks field-level encryption. I request adding encryption and key rotation, moving acceptance from June 30 to July 14, and assigning two engineers. Estimated added cost is $15,000 against the $300,000 baseline."}, {"speaker": "Delivery lead", "text": "July 14 is exactly 10 business days later, and $15,000 is exactly 5% of the $300,000 baseline. Two engineers are assigned to this work."}, {"speaker": "Project manager", "text": "The rationale, schedule calculation, staffing need, and cost estimate are documented. Contractual milestone breaches go to the project sponsor; other delivery changes go to the delivery lead."}, {"speaker": "Contract note", "text": "Contract Amendment C-14 sets a firm milestone requiring key-rotation deployment completion by March 1, 2025."}, {"speaker": "Change requester", "text": "The encryption-and-key-rotation change request schedules key-rotation deployment completion for February 20, 2025."}]}, "method": "c2d", "provenance": {"source_id": "diverse-047", "source_is_synthetic": true, "source_sha256": "6638862905ee9e22f592d21c2ed5b9ee00e8cd3a74e7992096b37f244db0877e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the deadline, entities, and rationale while only varying the certification date, evidence spans are two clean factual sentences matching the text, and neither context contains policy tables, rationale, or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Case note — Project Birch/Cedar staffing request:\\n\\nLina's request lists 12 September as the certification deadline for Birch's API. Her request states it remains active only if certification occurs by that date, and states it is withdrawn otherwise. The request includes a rationale (accelerating Birch's mobile release) and asks to borrow one QA engineer from Cedar for up to two weeks, delaying Cedar's regression testing by at most three business days. The vendor issued the confirmed certification for Birch's API on 20 September. No staff transfer, schedule update, spending, or scope change has been initiated. The change does not alter the core project outcome and requires no additional funding. Cedar's delivery lead and the project sponsor have not yet reviewed the request.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lina's request lists 12 September as the certification deadline for Birch's API."}, {"path": [], "text": "The vendor issued the confirmed certification for Birch's API on 20 September."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Lina's request lists 12 September as the certification deadline for Birch's API.", "negative_left": "Lina's request lists 12 September as the certification deadline for Birch's API.", "negative_right": "The vendor issued the confirmed certification for Birch's API on 5 September.", "right": "The vendor issued the confirmed certification for Birch's API on 20 September."}, "verifier_independent_model": false}, "family": "fast-43-diverse-048-002", "id": "fast-43-diverse-048-002-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Case note — Project Birch/Cedar staffing request:\n\nLina's request lists 12 September as the certification deadline for Birch's API. Her request states it remains active only if certification occurs by that date, and states it is withdrawn otherwise. The request includes a rationale (accelerating Birch's mobile release) and asks to borrow one QA engineer from Cedar for up to two weeks, delaying Cedar's regression testing by at most three business days. The vendor issued the confirmed certification for Birch's API on 20 September. No staff transfer, schedule update, spending, or scope change has been initiated. The change does not alter the core project outcome and requires no additional funding. Cedar's delivery lead and the project sponsor have not yet reviewed the request."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the deadline, entities, and rationale while only varying the certification date, evidence spans are two clean factual sentences matching the text, and neither context contains policy tables, rationale, or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Case note — Project Birch/Cedar staffing request:\\n\\nLina's request lists 12 September as the certification deadline for Birch's API. Her request states it remains active only if certification occurs by that date, and states it is withdrawn otherwise. The request includes a rationale (accelerating Birch's mobile release) and asks to borrow one QA engineer from Cedar for up to two weeks, delaying Cedar's regression testing by at most three business days. The vendor issued the confirmed certification for Birch's API on 20 September. No staff transfer, schedule update, spending, or scope change has been initiated. The change does not alter the core project outcome and requires no additional funding. Cedar's delivery lead and the project sponsor have not yet reviewed the request.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lina's request lists 12 September as the certification deadline for Birch's API."}, {"path": [], "text": "The vendor issued the confirmed certification for Birch's API on 20 September."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Lina's request lists 12 September as the certification deadline for Birch's API.", "negative_left": "Lina's request lists 12 September as the certification deadline for Birch's API.", "negative_right": "The vendor issued the confirmed certification for Birch's API on 5 September.", "right": "The vendor issued the confirmed certification for Birch's API on 20 September."}, "verifier_independent_model": false}, "family": "fast-43-diverse-048-002", "id": "fast-43-diverse-048-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Case note — Project Birch/Cedar staffing request:\n\nLina's request lists 12 September as the certification deadline for Birch's API. Her request states it remains active only if certification occurs by that date, and states it is withdrawn otherwise. The request includes a rationale (accelerating Birch's mobile release) and asks to borrow one QA engineer from Cedar for up to two weeks, delaying Cedar's regression testing by at most three business days. The vendor issued the confirmed certification for Birch's API on 5 September. No staff transfer, schedule update, spending, or scope change has been initiated. The change does not alter the core project outcome and requires no additional funding. Cedar's delivery lead and the project sponsor have not yet reviewed the request."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy via the original questions object, preserve the same entities/dates/paths, use two verbatim factual evidence sentences, and the counterfactual only alters the certification date coherently without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Change requester Lina submitted a request to borrow one QA engineer from Project Cedar for two weeks, citing a rationale to accelerate Project Birch's mobile release; this would delay Cedar's regression testing by three business days at most. Lina's request for Birch's API sets 12 September as the certification deadline. Her request states that it remains active only if the vendor certifies Birch's API by that date, and is otherwise withdrawn with staffing left unchanged. The vendor's certification record shows Birch's API was certified on 20 September. No additional funding is required, and the core project outcome is unaffected by the proposal. As of today, no QA staff transfer has occurred, no schedule update has been made, no spending has been initiated, and no scope change has been recorded. The Cedar delivery lead and the project sponsor have not yet reviewed the matter.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lina's request for Birch's API sets 12 September as the certification deadline."}, {"path": [], "text": "The vendor's certification record shows Birch's API was certified on 20 September."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Lina's request for Birch's API sets 12 September as the certification deadline.", "negative_left": "Lina's request for Birch's API sets 12 September as the certification deadline.", "negative_right": "The vendor's certification record shows Birch's API was certified on 5 September.", "right": "The vendor's certification record shows Birch's API was certified on 20 September."}, "verifier_independent_model": false}, "family": "fast-43-diverse-048-020", "id": "fast-43-diverse-048-020-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Change requester Lina submitted a request to borrow one QA engineer from Project Cedar for two weeks, citing a rationale to accelerate Project Birch's mobile release; this would delay Cedar's regression testing by three business days at most. Lina's request for Birch's API sets 12 September as the certification deadline. Her request states that it remains active only if the vendor certifies Birch's API by that date, and is otherwise withdrawn with staffing left unchanged. The vendor's certification record shows Birch's API was certified on 20 September. No additional funding is required, and the core project outcome is unaffected by the proposal. As of today, no QA staff transfer has occurred, no schedule update has been made, no spending has been initiated, and no scope change has been recorded. The Cedar delivery lead and the project sponsor have not yet reviewed the matter."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy via the original questions object, preserve the same entities/dates/paths, use two verbatim factual evidence sentences, and the counterfactual only alters the certification date coherently without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Change requester Lina submitted a request to borrow one QA engineer from Project Cedar for two weeks, citing a rationale to accelerate Project Birch's mobile release; this would delay Cedar's regression testing by three business days at most. Lina's request for Birch's API sets 12 September as the certification deadline. Her request states that it remains active only if the vendor certifies Birch's API by that date, and is otherwise withdrawn with staffing left unchanged. The vendor's certification record shows Birch's API was certified on 20 September. No additional funding is required, and the core project outcome is unaffected by the proposal. As of today, no QA staff transfer has occurred, no schedule update has been made, no spending has been initiated, and no scope change has been recorded. The Cedar delivery lead and the project sponsor have not yet reviewed the matter.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lina's request for Birch's API sets 12 September as the certification deadline."}, {"path": [], "text": "The vendor's certification record shows Birch's API was certified on 20 September."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Lina's request for Birch's API sets 12 September as the certification deadline.", "negative_left": "Lina's request for Birch's API sets 12 September as the certification deadline.", "negative_right": "The vendor's certification record shows Birch's API was certified on 5 September.", "right": "The vendor's certification record shows Birch's API was certified on 20 September."}, "verifier_independent_model": false}, "family": "fast-43-diverse-048-020", "id": "fast-43-diverse-048-020-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Change requester Lina submitted a request to borrow one QA engineer from Project Cedar for two weeks, citing a rationale to accelerate Project Birch's mobile release; this would delay Cedar's regression testing by three business days at most. Lina's request for Birch's API sets 12 September as the certification deadline. Her request states that it remains active only if the vendor certifies Birch's API by that date, and is otherwise withdrawn with staffing left unchanged. The vendor's certification record shows Birch's API was certified on 5 September. No additional funding is required, and the core project outcome is unaffected by the proposal. As of today, no QA staff transfer has occurred, no schedule update has been made, no spending has been initiated, and no scope change has been recorded. The Cedar delivery lead and the project sponsor have not yet reviewed the matter."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original conditional trigger and unaltered instructions/criteria, only the vendor certification date differs (20 Sept vs 5 Sept) which is a coherent single-fact change relative to the fixed 12 Sept deadline, the two focus evidence spans are complete factual sentences, question entity/time bindings are unchanged, and neither context states or implies an impact level or withdrawal outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Change requester Lina asks to borrow one QA engineer from Project Cedar for two weeks, which would delay Cedar's regression testing by three business days. Her rationale is to accelerate Project Birch's mobile release. Lina's request lists 12 September as the certification deadline for Birch's API. Her stated conditional clause is: proceed only if the security vendor certifies Birch's API by that deadline; otherwise withdraw the request and leave staffing unchanged. The vendor issued written confirmation certifying Birch's API on 20 September. The project manager verified the notice and found no conflicting evidence. No staff transfer, schedule update, spending, or scope change has been initiated. Cedar's delivery lead and the project sponsor have not yet reviewed the request. The request does not alter the core project outcome or require additional funding.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lina's request lists 12 September as the certification deadline for Birch's API."}, {"path": [], "text": "The vendor issued written confirmation certifying Birch's API on 20 September."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Lina's request lists 12 September as the certification deadline for Birch's API.", "negative_left": "Lina's request lists 12 September as the certification deadline for Birch's API.", "negative_right": "The vendor issued written confirmation certifying Birch's API on 5 September.", "right": "The vendor issued written confirmation certifying Birch's API on 20 September."}, "verifier_independent_model": false}, "family": "fast-43-diverse-048-031", "id": "fast-43-diverse-048-031-base", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Change requester Lina asks to borrow one QA engineer from Project Cedar for two weeks, which would delay Cedar's regression testing by three business days. Her rationale is to accelerate Project Birch's mobile release. Lina's request lists 12 September as the certification deadline for Birch's API. Her stated conditional clause is: proceed only if the security vendor certifies Birch's API by that deadline; otherwise withdraw the request and leave staffing unchanged. The vendor issued written confirmation certifying Birch's API on 20 September. The project manager verified the notice and found no conflicting evidence. No staff transfer, schedule update, spending, or scope change has been initiated. Cedar's delivery lead and the project sponsor have not yet reviewed the request. The request does not alter the core project outcome or require additional funding."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "workplace-03", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original conditional trigger and unaltered instructions/criteria, only the vendor certification date differs (20 Sept vs 5 Sept) which is a coherent single-fact change relative to the fixed 12 Sept deadline, the two focus evidence spans are complete factual sentences, question entity/time bindings are unchanged, and neither context states or implies an impact level or withdrawal outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus concerns the factual timing of certification rather than policy. The base and counter assignments are jointly realizable with only that timing fact changing: the request terms can remain fixed while certification occurs either after or by the deadline. Both rules are sufficient under the ordered impact criteria, and partial coverage is permitted. Empty policy_evidence is correct because all governing policy is already retained in the questions object; the state otherwise supplies case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The confirmed certification date is after the deadline, and the request expressly withdraws upon that failure. The question’s policy therefore makes the request inactive and requires impact level 0, regardless of its unrealized proposed effects.", "rule_index": 0, "sound": true}, {"reason": "Refuting a later certification date, together with the express on-time activation condition, establishes an active request. A maximum three-business-day milestone delay independently meets level 2. The stated maxima and the negations of core-outcome change and additional funding exclude every level-3 trigger.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The vendor's confirmed certification date for Birch's API is later than the 12 September deadline in Lina's request."}, {"id": "a2", "statement": "Lina's request specifies that it remains active when the vendor certifies Birch's API no later than 12 September."}, {"id": "a3", "statement": "Lina's request specifies that it is withdrawn when the vendor does not certify Birch's API by 12 September."}, {"id": "a4", "statement": "Lina's request has a rationale."}, {"id": "a5", "statement": "The greatest milestone delay stated for Lina's requested change is three business days."}, {"id": "a6", "statement": "The longest resource reassignment stated for Lina's requested change is two weeks."}, {"id": "a7", "statement": "Lina's requested change alters the core project outcome."}, {"id": "a8", "statement": "Lina's requested change requires additional funding."}, {"id": "a9", "statement": "The requested QA staff transfer has been initiated."}, {"id": "a10", "statement": "A schedule update for Lina's requested change has been initiated."}, {"id": "a11", "statement": "Spending for Lina's requested change has been initiated."}, {"id": "a12", "statement": "A scope change for Lina's requested change has been initiated."}], "base_state_json": "\"Change requester Lina asks to borrow one QA engineer from Project Cedar for two weeks, which would delay Cedar's regression testing by three business days. Her rationale is to accelerate Project Birch's mobile release. Lina's request lists 12 September as the certification deadline for Birch's API. Her stated conditional clause is: proceed only if the security vendor certifies Birch's API by that deadline; otherwise withdraw the request and leave staffing unchanged. The vendor issued written confirmation certifying Birch's API on 20 September. The project manager verified the notice and found no conflicting evidence. No staff transfer, schedule update, spending, or scope change has been initiated. Cedar's delivery lead and the project sponsor have not yet reviewed the request. The request does not alter the core project outcome or require additional funding.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lina's request lists 12 September as the certification deadline for Birch's API."}, {"path": [], "text": "The vendor issued written confirmation certifying Birch's API on 20 September."}], "policy_evidence": [], "rules": [{"justification": "The certification condition fails, so the explicit withdrawal provision makes the request inactive. With no implementation initiated, the policy requires delivery-impact level 0 despite the unrealized effects.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "Certification is not later than the deadline, so the request remains active. Its maximum three-business-day milestone delay and two-week resource reassignment satisfy level 2, while the absence of a core-outcome change or additional funding and the stated maxima exclude level 3.", "target": "2", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "Lina's request lists 12 September as the certification deadline for Birch's API.", "negative_left": "Lina's request lists 12 September as the certification deadline for Birch's API.", "negative_right": "The vendor issued written confirmation certifying Birch's API on 5 September.", "right": "The vendor issued written confirmation certifying Birch's API on 20 September."}, "verifier_independent_model": false}, "family": "fast-43-diverse-048-031", "id": "fast-43-diverse-048-031-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No delivery impact: the request is inactive, withdrawn, rejected before implementation, or produces no scope, schedule, resource, or cost change.", "1 — Minor delivery impact: an active change affects internal work but causes no milestone movement and requires at most two person-days of added or reassigned effort.", "2 — Moderate delivery impact: an active change moves a milestone by one to five business days, changes approved scope without altering the core outcome, or reassigns resources for up to four weeks.", "3 — Major delivery impact: an active change moves a milestone by more than five business days, changes the core project outcome, requires additional funding, or removes a critical resource for more than four weeks."], "instructions": "Triage the request using this policy: resource reallocations route to the Delivery lead; scope changes route to the Project sponsor; timeline-only changes route to the Project manager. A request is ready for review only if it is active, has a rationale, and includes evidence for its stated trigger and delivery effects. A request explicitly withdrawn when a condition fails is inactive and is not ready; record its delivery impact as level 0 even if the unrealized proposal described a larger effect. Select the single applicable ordered delivery-impact level.", "type": "score"}}, "state": "Change requester Lina asks to borrow one QA engineer from Project Cedar for two weeks, which would delay Cedar's regression testing by three business days. Her rationale is to accelerate Project Birch's mobile release. Lina's request lists 12 September as the certification deadline for Birch's API. Her stated conditional clause is: proceed only if the security vendor certifies Birch's API by that deadline; otherwise withdraw the request and leave staffing unchanged. The vendor issued written confirmation certifying Birch's API on 5 September. The project manager verified the notice and found no conflicting evidence. No staff transfer, schedule update, spending, or scope change has been initiated. Cedar's delivery lead and the project sponsor have not yet reviewed the request. The request does not alter the core project outcome or require additional funding."}, "method": "c2d", "provenance": {"source_id": "diverse-048", "source_is_synthetic": true, "source_sha256": "270178b1636316e9ec9b79ab370e01a2c7086797e35b943115f9578d48eb1343", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "workplace-03", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full acceptance policy and rubric verbatim alongside the unchanged question, preserving entity, path and time bindings; the focus evidence pair are two complete factual sentences about Dana Cole's role and sign-off action; the counterfactual coherently swaps a signed sign-off for a declined one without contradicting other evidence; neither context contains gold answers, rule tables or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\":\"The Nimbus payroll-export package acceptance review is underway.\",\"evidence\":[\"Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.\",\"Quality reviewer signed the test attestation on 14 September.\",\"No cosmetic defects were noted in the final QA sweep.\",\"Dana Cole is the designated Business approver for the Nimbus payroll-export package.\",\"Dana Cole signed a written sign-off document on March 4.\"],\"policy\":[\"Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off.\",\"Verbal approval does not count.\",\"Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.\"],\"request\":\"Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item, if any.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "Dana Cole is the designated Business approver for the Nimbus payroll-export package."}, {"path": ["evidence", "4"], "text": "Dana Cole signed a written sign-off document on March 4."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "Dana Cole is the designated Business approver for the Nimbus payroll-export package.", "negative_left": "Dana Cole is the designated Business approver for the Nimbus payroll-export package.", "negative_right": "Dana Cole declined to sign a written sign-off document on March 4.", "right": "Dana Cole signed a written sign-off document on March 4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-056-003", "id": "fast-43-diverse-056-003-base", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "The Nimbus payroll-export package acceptance review is underway.", "evidence": ["Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.", "Quality reviewer signed the test attestation on 14 September.", "No cosmetic defects were noted in the final QA sweep.", "Dana Cole is the designated Business approver for the Nimbus payroll-export package.", "Dana Cole signed a written sign-off document on March 4."], "policy": ["Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off.", "Verbal approval does not count.", "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."], "request": "Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item, if any."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_high"}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full acceptance policy and rubric verbatim alongside the unchanged question, preserving entity, path and time bindings; the focus evidence pair are two complete factual sentences about Dana Cole's role and sign-off action; the counterfactual coherently swaps a signed sign-off for a declined one without contradicting other evidence; neither context contains gold answers, rule tables or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\":\"The Nimbus payroll-export package acceptance review is underway.\",\"evidence\":[\"Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.\",\"Quality reviewer signed the test attestation on 14 September.\",\"No cosmetic defects were noted in the final QA sweep.\",\"Dana Cole is the designated Business approver for the Nimbus payroll-export package.\",\"Dana Cole signed a written sign-off document on March 4.\"],\"policy\":[\"Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off.\",\"Verbal approval does not count.\",\"Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.\"],\"request\":\"Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item, if any.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "Dana Cole is the designated Business approver for the Nimbus payroll-export package."}, {"path": ["evidence", "4"], "text": "Dana Cole signed a written sign-off document on March 4."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "Dana Cole is the designated Business approver for the Nimbus payroll-export package.", "negative_left": "Dana Cole is the designated Business approver for the Nimbus payroll-export package.", "negative_right": "Dana Cole declined to sign a written sign-off document on March 4.", "right": "Dana Cole signed a written sign-off document on March 4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-056-003", "id": "fast-43-diverse-056-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "The Nimbus payroll-export package acceptance review is underway.", "evidence": ["Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.", "Quality reviewer signed the test attestation on 14 September.", "No cosmetic defects were noted in the final QA sweep.", "Dana Cole is the designated Business approver for the Nimbus payroll-export package.", "Dana Cole declined to sign a written sign-off document on March 4."], "policy": ["Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off.", "Verbal approval does not count.", "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."], "request": "Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item, if any."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "hold_near_complete_business_approver"}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance rubric verbatim and the unchanged questions object supplies all choice criteria, so policy is preserved; entities, dates, and the decision request match the original question bindings; the two focus evidence sentences ('Business approver Dana Cole reviewed... on March 3.' and 'documented in a signed written memo.'/'communicated only verbally in a phone call.') are plain factual statements, not rubric text or instructions; the counterfactual's verbal-only communication is logically consistent with the unchanged review-occurred fact and with the 'verbal approval does not count' policy, producing no contradictory duplicate facts; neither context names a decision label, option ID, or tells the model what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\": \"The Nimbus payroll-export package is submitted as complete. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.\", \"evidence\": [\"Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.\", \"Quality reviewer signed the test attestation on 14 September.\", \"Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3.\", \"Dana Cole's review of the Nimbus payroll-export package on March 3 was documented in a signed written memo.\", \"No cosmetic or other defects were noted in the final review.\"], \"request\": \"Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3."}, {"path": ["evidence", "3"], "text": "Dana Cole's review of the Nimbus payroll-export package on March 3 was documented in a signed written memo."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3.", "negative_left": "Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3.", "negative_right": "Dana Cole's review of the Nimbus payroll-export package on March 3 was communicated only verbally in a phone call.", "right": "Dana Cole's review of the Nimbus payroll-export package on March 3 was documented in a signed written memo."}, "verifier_independent_model": false}, "family": "fast-43-diverse-056-010", "id": "fast-43-diverse-056-010-base", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "The Nimbus payroll-export package is submitted as complete. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.", "evidence": ["Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.", "Quality reviewer signed the test attestation on 14 September.", "Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3.", "Dana Cole's review of the Nimbus payroll-export package on March 3 was documented in a signed written memo.", "No cosmetic or other defects were noted in the final review."], "request": "Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_high"}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance rubric verbatim and the unchanged questions object supplies all choice criteria, so policy is preserved; entities, dates, and the decision request match the original question bindings; the two focus evidence sentences ('Business approver Dana Cole reviewed... on March 3.' and 'documented in a signed written memo.'/'communicated only verbally in a phone call.') are plain factual statements, not rubric text or instructions; the counterfactual's verbal-only communication is logically consistent with the unchanged review-occurred fact and with the 'verbal approval does not count' policy, producing no contradictory duplicate facts; neither context names a decision label, option ID, or tells the model what to output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\": \"The Nimbus payroll-export package is submitted as complete. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.\", \"evidence\": [\"Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.\", \"Quality reviewer signed the test attestation on 14 September.\", \"Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3.\", \"Dana Cole's review of the Nimbus payroll-export package on March 3 was documented in a signed written memo.\", \"No cosmetic or other defects were noted in the final review.\"], \"request\": \"Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3."}, {"path": ["evidence", "3"], "text": "Dana Cole's review of the Nimbus payroll-export package on March 3 was documented in a signed written memo."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3.", "negative_left": "Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3.", "negative_right": "Dana Cole's review of the Nimbus payroll-export package on March 3 was communicated only verbally in a phone call.", "right": "Dana Cole's review of the Nimbus payroll-export package on March 3 was documented in a signed written memo."}, "verifier_independent_model": false}, "family": "fast-43-diverse-056-010", "id": "fast-43-diverse-056-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "The Nimbus payroll-export package is submitted as complete. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.", "evidence": ["Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.", "Quality reviewer signed the test attestation on 14 September.", "Business approver Dana Cole reviewed the Nimbus payroll-export package on March 3.", "Dana Cole's review of the Nimbus payroll-export package on March 3 was communicated only verbally in a phone call.", "No cosmetic or other defects were noted in the final review."], "request": "Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "hold_near_complete_business_approver"}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full acceptance rubric and requirements from the original state, unchanged from the question's criteria; the request and entity bindings (Nimbus payroll-export package, roles) are preserved verbatim; the two focused evidence lines are factual statements about Dana Reyes's review and sign-off, not policy text; the counterfactual coherently swaps the signed memo for a verbal-only communication without duplicating or contradicting other facts; no evidence line states a rubric label, code, or instructs an output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\":\"The Nimbus payroll-export package is submitted as complete. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.\",\"evidence\":[\"Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.\",\"Quality reviewer signed the test attestation on 14 September.\",\"Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3.\",\"Dana Reyes's March 3 decision was captured in a signed written memo placed in the acceptance file.\",\"No cosmetic defects were noted in the final review.\"],\"request\":\"Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3."}, {"path": ["evidence", "3"], "text": "Dana Reyes's March 3 decision was captured in a signed written memo placed in the acceptance file."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3.", "negative_left": "Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3.", "negative_right": "Dana Reyes's March 3 decision was communicated only verbally by phone, with no written memo placed in the acceptance file.", "right": "Dana Reyes's March 3 decision was captured in a signed written memo placed in the acceptance file."}, "verifier_independent_model": false}, "family": "fast-43-diverse-056-023", "id": "fast-43-diverse-056-023-base", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "The Nimbus payroll-export package is submitted as complete. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.", "evidence": ["Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.", "Quality reviewer signed the test attestation on 14 September.", "Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3.", "Dana Reyes's March 3 decision was captured in a signed written memo placed in the acceptance file.", "No cosmetic defects were noted in the final review."], "request": "Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_high"}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full acceptance rubric and requirements from the original state, unchanged from the question's criteria; the request and entity bindings (Nimbus payroll-export package, roles) are preserved verbatim; the two focused evidence lines are factual statements about Dana Reyes's review and sign-off, not policy text; the counterfactual coherently swaps the signed memo for a verbal-only communication without duplicating or contradicting other facts; no evidence line states a rubric label, code, or instructs an output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and A4 is a factual sign-off relation rather than a policy classification. The base and counter assignments differ only on A4 and are both realizable: the base supports all mandatory evidence with no cosmetic defect, while the counter lacks written business sign-off. The policy evidence correctly cites the original state and preserves the state-originating acceptance requirements, verbal-approval exception, and rubric; question-originating criteria remain automatically available.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 establish every mandatory test, defect, Quality-attestation, and written-approval requirement, while refuted A5 excludes remaining cosmetic defects. This is sufficient for accept/high and excludes the competing adequate, hold, and reject outcomes.", "rule_index": 0, "sound": true}, {"reason": "A1–A3 establish satisfactory technical evidence and Quality attestation, refuted A4 establishes that mandatory written Business approver sign-off is absent, and refuted A5 introduces no conflict. This is sufficient for hold/near-complete routed to the Business approver and excludes acceptance, unverifiable hold, and material-defect rejection.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The submitted test report for the Nimbus payroll-export package records 20 of 20 tests passed."}, {"id": "A2", "statement": "The submitted defect register for the Nimbus payroll-export package records zero open Severity 1 or 2 defects."}, {"id": "A3", "statement": "The submitted acceptance evidence for the Nimbus payroll-export package contains a signed Quality reviewer test attestation."}, {"id": "A4", "statement": "The Nimbus payroll-export package has a written Business approver sign-off."}, {"id": "A5", "statement": "At least one cosmetic defect remains in the Nimbus payroll-export package."}], "base_state_json": "{\"context\":\"The Nimbus payroll-export package is submitted as complete. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.\",\"evidence\":[\"Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.\",\"Quality reviewer signed the test attestation on 14 September.\",\"Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3.\",\"Dana Reyes's March 3 decision was captured in a signed written memo placed in the acceptance file.\",\"No cosmetic defects were noted in the final review.\"],\"request\":\"Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3."}, {"path": ["evidence", "3"], "text": "Dana Reyes's March 3 decision was captured in a signed written memo placed in the acceptance file."}], "policy_evidence": [{"path": ["context"], "text": "Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off."}, {"path": ["context"], "text": "Verbal approval does not count."}, {"path": ["context"], "text": "Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing."}], "rules": [{"justification": "Every required technical result, the Quality reviewer attestation, and written Business approver approval are evidenced, while no cosmetic defect remains.", "target": "accept_high", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "All technical evidence and the Quality reviewer attestation are satisfactory, but the mandatory written Business approver sign-off is explicitly missing.", "target": "hold_near_complete_business_approver", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3.", "negative_left": "Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3.", "negative_right": "Dana Reyes's March 3 decision was communicated only verbally by phone, with no written memo placed in the acceptance file.", "right": "Dana Reyes's March 3 decision was captured in a signed written memo placed in the acceptance file."}, "verifier_independent_model": false}, "family": "fast-43-diverse-056-023", "id": "fast-43-diverse-056-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_adequate_project_coordinator": "Accept with adequate completion quality because all mandatory evidence exists and only cosmetic issues remain; route administrative closeout to the Project coordinator.", "accept_high": "Accept with high completion quality because every technical criterion and mandatory approval is explicitly evidenced; no routing is needed.", "hold_near_complete_business_approver": "Hold acceptance with near-complete quality because all technical evidence is satisfactory but written business sign-off is missing; route approval to the Business approver.", "hold_unverifiable_quality_reviewer": "Hold acceptance with unverifiable quality because the test report or Quality reviewer attestation is absent; route evidence validation to the Quality reviewer.", "reject_material_delivery_owner": "Reject with material-defect quality because a test failed or a Severity 1 or 2 defect remains; route correction to the Delivery owner."}, "instructions": "Choose the single option whose outcome and routing exactly match the stated acceptance rubric and submitted evidence.", "type": "choice"}}, "state": {"context": "The Nimbus payroll-export package is submitted as complete. Acceptance requires: 20 of 20 tests passed, no open Severity 1 or 2 defects, a Quality reviewer test attestation, and written Business approver sign-off. Verbal approval does not count. Rubric: accept/high when every requirement is evidenced; accept/adequate when all requirements are evidenced but cosmetic defects remain; hold/near-complete when technical criteria pass but mandatory approval is missing; reject/material-defect when tests fail or severe defects remain; hold/unverifiable when test evidence or Quality attestation is missing.", "evidence": ["Test report: 20 of 20 passed; defect register: zero open Severity 1 or 2 defects.", "Quality reviewer signed the test attestation on 14 September.", "Business approver Dana Reyes reviewed the Nimbus payroll-export package and communicated her decision on March 3.", "Dana Reyes's March 3 decision was communicated only verbally by phone, with no written memo placed in the acceptance file.", "No cosmetic defects were noted in the final review."], "request": "Decide acceptance status, rate completion quality, and identify the role responsible for the unresolved item."}}, "method": "c2d", "provenance": {"source_id": "diverse-056", "source_is_synthetic": true, "source_sha256": "a43b7b933629f00d5288a20eb0035171ee532f318cd9d36b503e9bb53106b348", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "hold_near_complete_business_approver"}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy statement and question bindings, the two evidence sentences are factual and complete, the counterfactual swap of Dana's email to 'I have not yet reviewed WP-17' is coherent with the rest of the unchanged context, and neither context reveals a gold answer or instructs the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Delivery owner\",\"text\":\"WP-17 is submitted. Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Delivery owner\",\"text\":\"The test report shows 12/12 passed; no defects remain, and the technical evidence package is complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"I verified the attached test report and signed QR-17. The evidence is authentic and matches WP-17.\"},{\"speaker\":\"Delivery owner\",\"text\":\"WP-17's assigned Business approver is Dana Kimura.\"},{\"speaker\":\"Delivery owner\",\"text\":\"Dana Kimura sent an email on June 4 stating, 'I approve WP-17 for release.'\"},{\"speaker\":\"Project coordinator\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["3", "text"], "text": "WP-17's assigned Business approver is Dana Kimura."}, {"path": ["4", "text"], "text": "Dana Kimura sent an email on June 4 stating, 'I approve WP-17 for release.'"}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "WP-17's assigned Business approver is Dana Kimura.", "negative_left": "WP-17's assigned Business approver is Dana Kimura.", "negative_right": "Dana Kimura sent an email on June 4 stating, 'I have not yet reviewed WP-17.'", "right": "Dana Kimura sent an email on June 4 stating, 'I approve WP-17 for release.'"}, "verifier_independent_model": false}, "family": "fast-43-diverse-057-021", "id": "fast-43-diverse-057-021-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Delivery owner", "text": "WP-17 is submitted. Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Delivery owner", "text": "The test report shows 12/12 passed; no defects remain, and the technical evidence package is complete."}, {"speaker": "Quality reviewer", "text": "I verified the attached test report and signed QR-17. The evidence is authentic and matches WP-17."}, {"speaker": "Delivery owner", "text": "WP-17's assigned Business approver is Dana Kimura."}, {"speaker": "Delivery owner", "text": "Dana Kimura sent an email on June 4 stating, 'I approve WP-17 for release.'"}, {"speaker": "Project coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy statement and question bindings, the two evidence sentences are factual and complete, the counterfactual swap of Dana's email to 'I have not yet reviewed WP-17' is coherent with the rest of the unchanged context, and neither context reveals a gold answer or instructs the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\":\"Delivery owner\",\"text\":\"WP-17 is submitted. Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"},{\"speaker\":\"Delivery owner\",\"text\":\"The test report shows 12/12 passed; no defects remain, and the technical evidence package is complete.\"},{\"speaker\":\"Quality reviewer\",\"text\":\"I verified the attached test report and signed QR-17. The evidence is authentic and matches WP-17.\"},{\"speaker\":\"Delivery owner\",\"text\":\"WP-17's assigned Business approver is Dana Kimura.\"},{\"speaker\":\"Delivery owner\",\"text\":\"Dana Kimura sent an email on June 4 stating, 'I approve WP-17 for release.'\"},{\"speaker\":\"Project coordinator\",\"text\":\"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["3", "text"], "text": "WP-17's assigned Business approver is Dana Kimura."}, {"path": ["4", "text"], "text": "Dana Kimura sent an email on June 4 stating, 'I approve WP-17 for release.'"}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "WP-17's assigned Business approver is Dana Kimura.", "negative_left": "WP-17's assigned Business approver is Dana Kimura.", "negative_right": "Dana Kimura sent an email on June 4 stating, 'I have not yet reviewed WP-17.'", "right": "Dana Kimura sent an email on June 4 stating, 'I approve WP-17 for release.'"}, "verifier_independent_model": false}, "family": "fast-43-diverse-057-021", "id": "fast-43-diverse-057-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Delivery owner", "text": "WP-17 is submitted. Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Delivery owner", "text": "The test report shows 12/12 passed; no defects remain, and the technical evidence package is complete."}, {"speaker": "Quality reviewer", "text": "I verified the attached test report and signed QR-17. The evidence is authentic and matches WP-17."}, {"speaker": "Delivery owner", "text": "WP-17's assigned Business approver is Dana Kimura."}, {"speaker": "Delivery owner", "text": "Dana Kimura sent an email on June 4 stating, 'I have not yet reviewed WP-17.'"}, {"speaker": "Project coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical acceptance policy and question bindings for WP-17 and Maria Chen; the counterfactual changes only the single sentence about Maria Chen's email to a coherent alternative fact without introducing contradictions, missing-evidence defaults, or answer labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\": \"Delivery owner\", \"text\": \"WP-17 status update: the test report confirms all 12 acceptance tests passed, and no defects remain open against the work package.\"}, {\"speaker\": \"Quality reviewer\", \"text\": \"I verified the attached test report and signed QR-17. The evidence is authentic and matches WP-17, so the technical evidence package is complete.\"}, {\"speaker\": \"Delivery owner\", \"text\": \"WP-17's assigned Business approver is Maria Chen.\"}, {\"speaker\": \"Delivery owner\", \"text\": \"Maria Chen sent an email on June 5 stating she explicitly approves WP-17.\"}, {\"speaker\": \"Project coordinator\", \"text\": \"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"}, {\"speaker\": \"Project coordinator\", \"text\": \"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["2", "text"], "text": "WP-17's assigned Business approver is Maria Chen."}, {"path": ["3", "text"], "text": "Maria Chen sent an email on June 5 stating she explicitly approves WP-17."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "WP-17's assigned Business approver is Maria Chen.", "negative_left": "WP-17's assigned Business approver is Maria Chen.", "negative_right": "Maria Chen sent an email on June 5 stating she has not yet reviewed WP-17.", "right": "Maria Chen sent an email on June 5 stating she explicitly approves WP-17."}, "verifier_independent_model": false}, "family": "fast-43-diverse-057-024", "id": "fast-43-diverse-057-024-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Delivery owner", "text": "WP-17 status update: the test report confirms all 12 acceptance tests passed, and no defects remain open against the work package."}, {"speaker": "Quality reviewer", "text": "I verified the attached test report and signed QR-17. The evidence is authentic and matches WP-17, so the technical evidence package is complete."}, {"speaker": "Delivery owner", "text": "WP-17's assigned Business approver is Maria Chen."}, {"speaker": "Delivery owner", "text": "Maria Chen sent an email on June 5 stating she explicitly approves WP-17."}, {"speaker": "Project coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Project coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical acceptance policy and question bindings for WP-17 and Maria Chen; the counterfactual changes only the single sentence about Maria Chen's email to a coherent alternative fact without introducing contradictions, missing-evidence defaults, or answer labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "full_context_fact_states": {"base": {"a_business": "supported", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "counterfactual": {"a_business": "refuted", "a_defects": "supported", "a_quality": "supported", "a_technical": "supported", "a_tests": "supported"}, "remove_left": {"a_business": "unknown"}, "remove_right": {"a_business": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_business": "unknown"}, "negative_pair": {"a_business": "refuted"}, "negative_sentence": {"a_business": "unknown"}, "positive_pair": {"a_business": "supported"}, "right": {"a_business": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual condition, and the focus atom concerns the factual existence of explicit Business approver approval rather than a policy conclusion. The base and counter assignments differ only on that focus and are both realizable. The policy evidence correctly cites the original state’s substantive acceptance and Near-complete rules; rules already in the retained questions object need not be duplicated. Both proposed rules are sufficient for their respective targets, while uncovered cases such as unknown approval may validly abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all three stated acceptance criteria: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval. Therefore accepting WP-17 now and rating it Complete is entailed. The additional technical-evidence and no-defect conditions are redundant but do not make the rule unsound.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes complete technical evidence, no remaining defects, passed tests and Quality sign-off, while explicitly refuting Business approval. Thus explicit Business approval is the sole missing criterion, which is sufficient under the stated policy to reject acceptance now, rate WP-17 Near-complete, and route the omission to the Business approver.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_tests", "statement": "All 12 acceptance tests for WP-17 passed."}, {"id": "a_quality", "statement": "WP-17 has Quality reviewer sign-off."}, {"id": "a_technical", "statement": "The technical evidence for WP-17 is complete."}, {"id": "a_defects", "statement": "WP-17 has zero remaining defects."}, {"id": "a_business", "statement": "WP-17 has explicit approval from its assigned Business approver."}], "base_state_json": "[{\"speaker\": \"Delivery owner\", \"text\": \"WP-17 status update: the test report confirms all 12 acceptance tests passed, and no defects remain open against the work package.\"}, {\"speaker\": \"Quality reviewer\", \"text\": \"I verified the attached test report and signed QR-17. The evidence is authentic and matches WP-17, so the technical evidence package is complete.\"}, {\"speaker\": \"Delivery owner\", \"text\": \"WP-17's assigned Business approver is Maria Chen.\"}, {\"speaker\": \"Delivery owner\", \"text\": \"Maria Chen sent an email on June 5 stating she explicitly approves WP-17.\"}, {\"speaker\": \"Project coordinator\", \"text\": \"Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval.\"}, {\"speaker\": \"Project coordinator\", \"text\": \"Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver.\"}]", "base_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}], "counter_states": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}], "focus_atom": "a_business", "focus_evidence": [{"path": ["2", "text"], "text": "WP-17's assigned Business approver is Maria Chen."}, {"path": ["3", "text"], "text": "Maria Chen sent an email on June 5 stating she explicitly approves WP-17."}], "policy_evidence": [{"path": ["0", "text"], "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"path": ["3", "text"], "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}], "rules": [{"justification": "All three required acceptance criteria are satisfied, so WP-17 may be accepted now and rated Complete.", "target": "true", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "supported"}]}, {"justification": "The technical evidence is complete, no defects remain, and the sole missing acceptance criterion is explicit Business approver approval, so WP-17 must be rated Near-complete and routed to that approver.", "target": "false", "when": [{"atom_id": "a_tests", "state": "supported"}, {"atom_id": "a_quality", "state": "supported"}, {"atom_id": "a_technical", "state": "supported"}, {"atom_id": "a_defects", "state": "supported"}, {"atom_id": "a_business", "state": "refuted"}]}]}, "verified_pair": {"left": "WP-17's assigned Business approver is Maria Chen.", "negative_left": "WP-17's assigned Business approver is Maria Chen.", "negative_right": "Maria Chen sent an email on June 5 stating she has not yet reviewed WP-17.", "right": "Maria Chen sent an email on June 5 stating she explicitly approves WP-17."}, "verifier_independent_model": false}, "family": "fast-43-diverse-057-024", "id": "fast-43-diverse-057-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept WP-17 now; rate it Near-complete and route the missing explicit approval to the Business approver.", "true": "Accept WP-17 now and rate its overall completion quality Complete because every required criterion is satisfied."}, "instructions": "Readiness decision: Should WP-17 be accepted now and rated Complete under the stated rubric? Answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Delivery owner", "text": "WP-17 status update: the test report confirms all 12 acceptance tests passed, and no defects remain open against the work package."}, {"speaker": "Quality reviewer", "text": "I verified the attached test report and signed QR-17. The evidence is authentic and matches WP-17, so the technical evidence package is complete."}, {"speaker": "Delivery owner", "text": "WP-17's assigned Business approver is Maria Chen."}, {"speaker": "Delivery owner", "text": "Maria Chen sent an email on June 5 stating she has not yet reviewed WP-17."}, {"speaker": "Project coordinator", "text": "Acceptance requires: all 12 tests passed, Quality reviewer sign-off, and explicit Business approver approval."}, {"speaker": "Project coordinator", "text": "Policy says silence is not approval. Accept and rate Complete only when all three criteria are met. Rate Near-complete when technical evidence is complete, no defects remain, and only explicit business approval is missing; route that omission to the Business approver."}]}, "method": "c2d", "provenance": {"source_id": "diverse-057", "source_is_synthetic": true, "source_sha256": "24644b5856ea155b7a1bd146e25853842d6f9604af6740ee04a632b6fbecb5db", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual swaps the signer's employee ID to EMP-2209, which no longer matches Priya Nandan's EMP-4471, coherently implying the Business approver did not sign, without altering policy, bindings, or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"The Atlas reporting work package is under acceptance review. The Quality reviewer's test log confirms all 12 of 12 specified tests passed, with no defect remaining. The Delivery owner had written, 'If Finance approves the totals, submit this package for acceptance.' Finance has since approved the Atlas totals. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-4471. Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program. No one else has submitted a competing sign-off for this package. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-4471."}, {"path": [], "text": "Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-4471.", "negative_left": "The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-2209.", "negative_right": "Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program.", "right": "Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program."}, "verifier_independent_model": false}, "family": "fast-43-diverse-058-007", "id": "fast-43-diverse-058-007-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "The Atlas reporting work package is under acceptance review. The Quality reviewer's test log confirms all 12 of 12 specified tests passed, with no defect remaining. The Delivery owner had written, 'If Finance approves the totals, submit this package for acceptance.' Finance has since approved the Atlas totals. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-4471. Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program. No one else has submitted a competing sign-off for this package. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "workplace-05", "split": "train", "variant": "base"} {"domain": "workplace", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual swaps the signer's employee ID to EMP-2209, which no longer matches Priya Nandan's EMP-4471, coherently implying the Business approver did not sign, without altering policy, bindings, or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "supported", "A7": "supported"}, "remove_left": {"A5": "unknown"}, "remove_right": {"A5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A5": "unknown"}, "negative_pair": {"A5": "refuted"}, "negative_sentence": {"A5": "unknown"}, "positive_pair": {"A5": "supported"}, "right": {"A5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the explicitly quantified test and person sets do not make A1 or A6 improper bundles. A5 is a factual signer-role relation rather than a policy conclusion. The base and counter assignments differ only on A5 and are jointly realizable: in the base the dated approval signer is the Business approver, while in the counter a non-Business-approver signer exists and nobody else supplies Business approver sign-off. The policy evidence correctly preserves the substantive acceptance, conditional-intent, defect, and rating rules originating in the state; case-specific observations need not be preserved because synthetic observations will replace them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all required conditions: all 12 tests pass, no defect remains, the Delivery owner's conditional intent is activated by Finance approval, and the dated totals approval was signed by the Business approver. Therefore acceptance and a High rating follow.", "rule_index": 0, "sound": true}, {"reason": "A5 being refuted establishes that the recorded signer is not the Business approver, while A6 excludes Business approver sign-off by anyone other than that signer. Thus the required Business approver sign-off is absent, so the package cannot be accepted or rated High.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Each of the 12 tests specified for the Atlas reporting work package has passing status at the acceptance decision time."}, {"id": "A2", "statement": "At least one defect remains in the 12 specified Atlas tests at the acceptance decision time."}, {"id": "A3", "statement": "At the acceptance decision time, the Delivery owner has expressed intent to submit the Atlas reporting work package for acceptance if Finance approves the totals."}, {"id": "A4", "statement": "Finance has approved the Atlas totals by the acceptance decision time."}, {"id": "A5", "statement": "The signer recorded on the dated approval of the Atlas totals is the Business approver."}, {"id": "A6", "statement": "No person other than the signer recorded on the dated approval of the Atlas totals has provided Business approver sign-off for the Atlas reporting work package by the acceptance decision time."}, {"id": "A7", "statement": "A dated signed approval record for the Atlas totals exists at the acceptance decision time."}], "base_state_json": "\"The Atlas reporting work package is under acceptance review. The Quality reviewer's test log confirms all 12 of 12 specified tests passed, with no defect remaining. The Delivery owner had written, 'If Finance approves the totals, submit this package for acceptance.' Finance has since approved the Atlas totals. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-4471. Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program. No one else has submitted a competing sign-off for this package. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A5", "focus_evidence": [{"path": [], "text": "The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-4471."}, {"path": [], "text": "Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program."}], "policy_evidence": [{"path": [], "text": "Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met."}, {"path": [], "text": "Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}], "rules": [{"justification": "All 12 specified tests pass and no defect remains. Finance approval makes the Delivery owner's conditional submission intent valid. The dated totals approval exists, and its signer is the Business approver, satisfying sign-off. Therefore every acceptance and High-quality criterion is met.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "Although the tests, defect, Finance-approval, and Delivery-owner-intent requirements are satisfied, the dated approval's signer is not the Business approver and no other person provided Business approver sign-off. A required acceptance criterion is therefore unmet, so the package must not be accepted or rated High.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-4471.", "negative_left": "The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-2209.", "negative_right": "Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program.", "right": "Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program."}, "verifier_independent_model": false}, "family": "fast-43-diverse-058-007", "id": "fast-43-diverse-058-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not accept the deliverable or do not rate its completion quality High.", "true": "Accept the deliverable and rate overall completion quality High."}, "instructions": "Decide whether the work package should now be accepted and rated High under the stated criteria.", "type": "noul"}}, "state": "The Atlas reporting work package is under acceptance review. The Quality reviewer's test log confirms all 12 of 12 specified tests passed, with no defect remaining. The Delivery owner had written, 'If Finance approves the totals, submit this package for acceptance.' Finance has since approved the Atlas totals. Acceptance requires: all 12 specified tests passing, Business approver sign-off, and Delivery owner intent to submit. Conditional intent becomes valid only when its stated condition is met. The dated approval record for the Atlas totals, signed on March 14, lists the signer's employee ID as EMP-2209. Employee ID EMP-4471 belongs to Priya Nandan, who holds the title of Business approver for the Atlas program. No one else has submitted a competing sign-off for this package. Rate completion quality High when every criterion is met and no test defect remains; otherwise rate Medium or Low and do not accept."}, "method": "c2d", "provenance": {"source_id": "diverse-058", "source_is_synthetic": true, "source_sha256": "23c9020ed13e653cdc150b6538ca5f8aa56bcd3150590746a31de88c06b85321", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "workplace-05", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric via the original questions object, preserve item numbering and diagnostic designations, and the counterfactual only swaps response_3's value from 'x=7' to 'x=9', a single coherent factual change without duplicate contradictions or leaked rubric outcomes.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\": \"Case note: Miko completed a five-item area-versus-perimeter check for the course instructor, with a learning support tutor available if needed. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Items 1 and 5 are the designated diagnostic area items; items 2, 3, and 4 are standard items.\", \"evidence\": {\"key_1\": \"The answer key for item 1 lists the correct answer as 24 cm\\u00b2.\", \"response_1\": \"Miko wrote 20 cm as the response to item 1, adding side lengths instead of multiplying.\", \"response_2\": \"Miko's response to item 2 was 20 cm, matching the key.\", \"key_2\": \"The answer key for item 2 lists 20 cm as correct.\", \"response_3\": \"On the five-item check, Miko wrote 'x=7' as the response to item 3.\", \"key_3\": \"The answer key for item 3 lists the correct answer as 'x=7'.\", \"response_4\": \"Miko's response to item 4 was 22 cm, matching the key.\", \"key_4\": \"The answer key for item 4 lists 22 cm as correct.\", \"response_5\": \"Miko wrote 16 cm as the response to item 5, using a perimeter formula rather than a division-based area correction.\", \"key_5\": \"The answer key for item 5 lists the correct answer as 16 cm\\u00b2.\", \"diagnostic_note\": \"The wrong-method note for item 1 describes adding all side lengths; the wrong-method note for item 5 describes a distinct perimeter-formula substitution, so the two diagnostic errors are not the same method.\", \"lesson_record\": \"Miko's lesson record confirms attendance at the area-versus-perimeter lesson and is on file.\", \"shown_work\": \"Shown work for both diagnostic items 1 and 5 was recorded and reviewed by the instructor.\"}, \"request\": \"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "response_3"], "text": "On the five-item check, Miko wrote 'x=7' as the response to item 3."}, {"path": ["evidence", "key_3"], "text": "The answer key for item 3 lists the correct answer as 'x=7'."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "On the five-item check, Miko wrote 'x=7' as the response to item 3.", "negative_left": "On the five-item check, Miko wrote 'x=9' as the response to item 3.", "negative_right": "The answer key for item 3 lists the correct answer as 'x=7'.", "right": "The answer key for item 3 lists the correct answer as 'x=7'."}, "verifier_independent_model": false}, "family": "fast-43-diverse-061-010", "id": "fast-43-diverse-061-010-base", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "Case note: Miko completed a five-item area-versus-perimeter check for the course instructor, with a learning support tutor available if needed. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Items 1 and 5 are the designated diagnostic area items; items 2, 3, and 4 are standard items.", "evidence": {"diagnostic_note": "The wrong-method note for item 1 describes adding all side lengths; the wrong-method note for item 5 describes a distinct perimeter-formula substitution, so the two diagnostic errors are not the same method.", "key_1": "The answer key for item 1 lists the correct answer as 24 cm².", "key_2": "The answer key for item 2 lists 20 cm as correct.", "key_3": "The answer key for item 3 lists the correct answer as 'x=7'.", "key_4": "The answer key for item 4 lists 22 cm as correct.", "key_5": "The answer key for item 5 lists the correct answer as 16 cm².", "lesson_record": "Miko's lesson record confirms attendance at the area-versus-perimeter lesson and is on file.", "response_1": "Miko wrote 20 cm as the response to item 1, adding side lengths instead of multiplying.", "response_2": "Miko's response to item 2 was 20 cm, matching the key.", "response_3": "On the five-item check, Miko wrote 'x=7' as the response to item 3.", "response_4": "Miko's response to item 4 was 22 cm, matching the key.", "response_5": "Miko wrote 16 cm as the response to item 5, using a perimeter formula rather than a division-based area correction.", "shown_work": "Shown work for both diagnostic items 1 and 5 was recorded and reviewed by the instructor."}, "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "B — Not yet mastered; practice set; U1"}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric via the original questions object, preserve item numbering and diagnostic designations, and the counterfactual only swaps response_3's value from 'x=7' to 'x=9', a single coherent factual change without duplicate contradictions or leaked rubric outcomes.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation or a permitted universally quantified completeness fact. The focus atom concerns whether item 3 matches its answer key, not a policy conclusion. The base and counter assignments are realizable with only item 3's correctness changing, moving the total from three correct to two. Policy evidence preserves the substantive state-origin interpretation that the same wrong method on both diagnostics is a recurring misconception; all remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish complete evidence, exactly three correct responses (items 2, 3, and 4), and refute use of the same wrong method on both diagnostic items. This is sufficient for option B.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish complete evidence and exactly two correct responses (items 2 and 4). The 0–2-correct criterion is sufficient for option C regardless of the diagnostic-method relation.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Miko's response to item 1 does not match the answer key for item 1."}, {"id": "a2", "statement": "Miko's response to item 2 matches the answer key for item 2."}, {"id": "a3", "statement": "Miko's response to item 3 matches the answer key for item 3."}, {"id": "a4", "statement": "Miko's response to item 4 matches the answer key for item 4."}, {"id": "a5", "statement": "Miko's response to item 5 does not match the answer key for item 5."}, {"id": "a6", "statement": "The wrong methods shown by Miko on diagnostic items 1 and 5 are the same method."}, {"id": "a7", "statement": "A response from Miko is recorded for every item in the five-item check."}, {"id": "a8", "statement": "The answer key specifies an answer for every item in the five-item check."}, {"id": "a9", "statement": "The diagnostic status of every diagnostic item in the five-item check is specified."}, {"id": "a10", "statement": "Shown work from Miko is recorded for every diagnostic item in the five-item check."}, {"id": "a11", "statement": "Miko's lesson record is available."}], "base_state_json": "{\"context\": \"Case note: Miko completed a five-item area-versus-perimeter check for the course instructor, with a learning support tutor available if needed. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Items 1 and 5 are the designated diagnostic area items; items 2, 3, and 4 are standard items.\", \"evidence\": {\"key_1\": \"The answer key for item 1 lists the correct answer as 24 cm\\u00b2.\", \"response_1\": \"Miko wrote 20 cm as the response to item 1, adding side lengths instead of multiplying.\", \"response_2\": \"Miko's response to item 2 was 20 cm, matching the key.\", \"key_2\": \"The answer key for item 2 lists 20 cm as correct.\", \"response_3\": \"On the five-item check, Miko wrote 'x=7' as the response to item 3.\", \"key_3\": \"The answer key for item 3 lists the correct answer as 'x=7'.\", \"response_4\": \"Miko's response to item 4 was 22 cm, matching the key.\", \"key_4\": \"The answer key for item 4 lists 22 cm as correct.\", \"response_5\": \"Miko wrote 16 cm as the response to item 5, using a perimeter formula rather than a division-based area correction.\", \"key_5\": \"The answer key for item 5 lists the correct answer as 16 cm\\u00b2.\", \"diagnostic_note\": \"The wrong-method note for item 1 describes adding all side lengths; the wrong-method note for item 5 describes a distinct perimeter-formula substitution, so the two diagnostic errors are not the same method.\", \"lesson_record\": \"Miko's lesson record confirms attendance at the area-versus-perimeter lesson and is on file.\", \"shown_work\": \"Shown work for both diagnostic items 1 and 5 was recorded and reviewed by the instructor.\"}, \"request\": \"Determine mastery, route, and ordered intervention urgency using the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["evidence", "response_3"], "text": "On the five-item check, Miko wrote 'x=7' as the response to item 3."}, {"path": ["evidence", "key_3"], "text": "The answer key for item 3 lists the correct answer as 'x=7'."}], "policy_evidence": [{"path": ["context"], "text": "The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip."}, {"path": ["request"], "text": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}], "rules": [{"justification": "All required evidence is available; exactly items 2, 3, and 4 are correct; and the two diagnostic wrong methods are not the same.", "target": "B — Not yet mastered; practice set; U1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "All required evidence is available, and exactly items 2 and 4 are correct, placing Miko in the 0–2-correct range.", "target": "C — Not mastered; tutor remediation; U3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}]}, "verified_pair": {"left": "On the five-item check, Miko wrote 'x=7' as the response to item 3.", "negative_left": "On the five-item check, Miko wrote 'x=9' as the response to item 3.", "negative_right": "The answer key for item 3 lists the correct answer as 'x=7'.", "right": "The answer key for item 3 lists the correct answer as 'x=7'."}, "verifier_independent_model": false}, "family": "fast-43-diverse-061-010", "id": "fast-43-diverse-061-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"A — Mastered; independent extension; U0": "Choose only when 4–5 responses are correct. Assign an independent area/perimeter extension with no intervention.", "B — Not yet mastered; practice set; U1": "Choose only when exactly 3 responses are correct and the diagnostic items do not show the same wrong method. Assign self-correcting mixed practice.", "C — Not mastered; tutor remediation; U3": "Choose only when 0–2 responses are correct. Route to the learning support tutor for intensive prerequisite and concept remediation.", "D — Not mastered; instructor misconception mini-lesson; U2": "Choose only when exactly 3 responses are correct and the same wrong method appears on both diagnostic items. Route to the course instructor’s targeted area-versus-perimeter mini-lesson.", "E — Decision deferred; evidence completion; no urgency assigned": "Choose only if a response, answer key, diagnostic designation, shown work, or lesson record needed by the rubric is missing."}, "instructions": "Select exactly one option. Apply these exclusive rules in order: incomplete evidence; 0–2 correct; exactly 3 correct with a repeated diagnostic misconception; exactly 3 correct without one; 4–5 correct. Urgency increases from U0 to U3.", "type": "choice"}}, "state": {"context": "Case note: Miko completed a five-item area-versus-perimeter check for the course instructor, with a learning support tutor available if needed. The rubric treats the same wrong method on both diagnostic items as a recurring misconception, not an arithmetic slip. Items 1 and 5 are the designated diagnostic area items; items 2, 3, and 4 are standard items.", "evidence": {"diagnostic_note": "The wrong-method note for item 1 describes adding all side lengths; the wrong-method note for item 5 describes a distinct perimeter-formula substitution, so the two diagnostic errors are not the same method.", "key_1": "The answer key for item 1 lists the correct answer as 24 cm².", "key_2": "The answer key for item 2 lists 20 cm as correct.", "key_3": "The answer key for item 3 lists the correct answer as 'x=7'.", "key_4": "The answer key for item 4 lists 22 cm as correct.", "key_5": "The answer key for item 5 lists the correct answer as 16 cm².", "lesson_record": "Miko's lesson record confirms attendance at the area-versus-perimeter lesson and is on file.", "response_1": "Miko wrote 20 cm as the response to item 1, adding side lengths instead of multiplying.", "response_2": "Miko's response to item 2 was 20 cm, matching the key.", "response_3": "On the five-item check, Miko wrote 'x=9' as the response to item 3.", "response_4": "Miko's response to item 4 was 22 cm, matching the key.", "response_5": "Miko wrote 16 cm as the response to item 5, using a perimeter formula rather than a division-based area correction.", "shown_work": "Shown work for both diagnostic items 1 and 5 was recorded and reviewed by the instructor."}, "request": "Determine mastery, route, and ordered intervention urgency using the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-061", "source_is_synthetic": true, "source_sha256": "2dddbb99014290562bb6e8856d49d7accc3f0641ab379f833820caadf93927ad", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C — Not mastered; tutor remediation; U3"}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric verbatim and only vary the Tuesday lesson-record misconception description, which is a permissible observation change; evidence spans are plain factual statements about documented misconceptions, not policy text; the counterfactual coherently shifts the prior-record misconception to a different rule without contradicting other unchanged facts; no context states or hints at the resulting level or route.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Learner\",\"text\":\"On Friday's decimal-comparison quiz, I answered 3 of 5 correctly according to the answer key. I chose 0.58 > 0.6 and 0.35 > 0.4; my other three comparisons were correct.\"},{\"speaker\":\"Course instructor\",\"text\":\"The key confirms 0.6 > 0.58 and 0.4 > 0.35, so exactly 3 of 5 are correct.\"},{\"speaker\":\"Learning support tutor\",\"text\":\"Both quiz errors trace to the same incorrect rule, so at least two errors share one misconception.\"},{\"speaker\":\"Course instructor\",\"text\":\"The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values.\"},{\"speaker\":\"Course instructor\",\"text\":\"The Tuesday prior lesson record documents that the learner's errors stem from believing longer decimal strings always represent larger values.\"},{\"speaker\":\"Course instructor\",\"text\":\"Using the routing rubric, decide the learner's intervention level and support route.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["3", "text"], "text": "The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values."}, {"path": ["4", "text"], "text": "The Tuesday prior lesson record documents that the learner's errors stem from believing longer decimal strings always represent larger values."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values.", "negative_left": "The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values.", "negative_right": "The Tuesday prior lesson record documents that the learner's errors stem from misaligning decimal points when comparing place values.", "right": "The Tuesday prior lesson record documents that the learner's errors stem from believing longer decimal strings always represent larger values."}, "verifier_independent_model": false}, "family": "fast-43-diverse-065-006", "id": "fast-43-diverse-065-006-base", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Learner", "text": "On Friday's decimal-comparison quiz, I answered 3 of 5 correctly according to the answer key. I chose 0.58 > 0.6 and 0.35 > 0.4; my other three comparisons were correct."}, {"speaker": "Course instructor", "text": "The key confirms 0.6 > 0.58 and 0.4 > 0.35, so exactly 3 of 5 are correct."}, {"speaker": "Learning support tutor", "text": "Both quiz errors trace to the same incorrect rule, so at least two errors share one misconception."}, {"speaker": "Course instructor", "text": "The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values."}, {"speaker": "Course instructor", "text": "The Tuesday prior lesson record documents that the learner's errors stem from believing longer decimal strings always represent larger values."}, {"speaker": "Course instructor", "text": "Using the routing rubric, decide the learner's intervention level and support route."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "education-01", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric verbatim and only vary the Tuesday lesson-record misconception description, which is a permissible observation change; evidence spans are plain factual statements about documented misconceptions, not policy text; the counterfactual coherently shifts the prior-record misconception to a different rule without contradicting other unchanged facts; no context states or hints at the resulting level or route.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship rather than a policy conclusion. A3 is a factual cross-record identity relation and is an appropriate focus. The base and counter assignments are realizable with A1 and A2 fixed while only whether the prior record documents the same misconception changes. Empty policy_evidence is correct because all governing rubric rules, priorities, thresholds, and exceptions are contained in the automatically retained questions object; the original state adds only case observations and workflow narrative.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A3 supported directly satisfies Level 3 because the same misconception is documented in both the current quiz and a prior lesson record. No additional exclusion is needed because Level 3 is the highest level.", "rule_index": 0, "sound": true}, {"reason": "A2 supported satisfies the repeated-current-error condition for Level 2. A1 supported means exactly 3 correct, excluding Level 3's 0-or-1 score condition, while A3 refuted excludes its cross-record misconception condition. Thus Level 2 is the highest applicable level.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Exactly 3 of the 5 answers on the learner’s Friday decimal-comparison quiz are correct according to the authoritative quiz answer key."}, {"id": "A2", "statement": "At least two errors on the learner’s Friday decimal-comparison quiz are explained by the same incorrect rule."}, {"id": "A3", "statement": "The misconception documented in the learner’s Friday decimal-comparison quiz is the same misconception documented in the learner’s Tuesday prior lesson record."}], "base_state_json": "[{\"speaker\":\"Learner\",\"text\":\"On Friday's decimal-comparison quiz, I answered 3 of 5 correctly according to the answer key. I chose 0.58 > 0.6 and 0.35 > 0.4; my other three comparisons were correct.\"},{\"speaker\":\"Course instructor\",\"text\":\"The key confirms 0.6 > 0.58 and 0.4 > 0.35, so exactly 3 of 5 are correct.\"},{\"speaker\":\"Learning support tutor\",\"text\":\"Both quiz errors trace to the same incorrect rule, so at least two errors share one misconception.\"},{\"speaker\":\"Course instructor\",\"text\":\"The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values.\"},{\"speaker\":\"Course instructor\",\"text\":\"The Tuesday prior lesson record documents that the learner's errors stem from believing longer decimal strings always represent larger values.\"},{\"speaker\":\"Course instructor\",\"text\":\"Using the routing rubric, decide the learner's intervention level and support route.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["3", "text"], "text": "The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values."}, {"path": ["4", "text"], "text": "The Tuesday prior lesson record documents that the learner's errors stem from believing longer decimal strings always represent larger values."}], "policy_evidence": [], "rules": [{"justification": "The same misconception is documented in both the current quiz and a prior lesson record, which independently satisfies Level 3.", "target": "3", "when": [{"atom_id": "A3", "state": "supported"}]}, {"justification": "Two or more current quiz errors share one incorrect rule, satisfying Level 2. Exactly 3 correct answers excludes the 0-or-1 Level 3 condition, and explicit refutation of A3 excludes the cross-record Level 3 condition.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values.", "negative_left": "The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values.", "negative_right": "The Tuesday prior lesson record documents that the learner's errors stem from misaligning decimal points when comparing place values.", "right": "The Tuesday prior lesson record documents that the learner's errors stem from believing longer decimal strings always represent larger values."}, "verifier_independent_model": false}, "family": "fast-43-diverse-065-006", "id": "fast-43-diverse-065-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Mastered: At least 4 of 5 quiz answers are correct, with no repeated misconception. Route to independent extension practice; no intervention is needed.", "1 — Light review: Exactly 3 of 5 are correct, errors do not share one incorrect rule, and lesson records show independent success. Route to a self-check answer-explanation activity within one week.", "2 — Targeted intervention: At least two current quiz errors share one incorrect rule, even if a prior lesson record shows success, provided Level 3 is not met. Route to the learning support tutor’s decimal-grid comparison activity within two school days.", "3 — Intensive intervention: The learner scores 0 or 1 of 5, or the same misconception is documented in both the current quiz and a prior lesson record. Route to a same-day instructor conference followed by tutor-led reteaching."], "instructions": "Select the single highest applicable level. The quiz answer key is authoritative. Two or more errors explained by the same incorrect rule count as a repeated misconception. A repeated misconception on the current quiz overrides an earlier successful lesson record for Level 1 versus Level 2; Level 3 requires the stated severe evidence.", "type": "score"}}, "state": [{"speaker": "Learner", "text": "On Friday's decimal-comparison quiz, I answered 3 of 5 correctly according to the answer key. I chose 0.58 > 0.6 and 0.35 > 0.4; my other three comparisons were correct."}, {"speaker": "Course instructor", "text": "The key confirms 0.6 > 0.58 and 0.4 > 0.35, so exactly 3 of 5 are correct."}, {"speaker": "Learning support tutor", "text": "Both quiz errors trace to the same incorrect rule, so at least two errors share one misconception."}, {"speaker": "Course instructor", "text": "The Friday decimal-comparison quiz notes that the learner's errors stem from believing longer decimal strings always represent larger values."}, {"speaker": "Course instructor", "text": "The Tuesday prior lesson record documents that the learner's errors stem from misaligning decimal points when comparing place values."}, {"speaker": "Course instructor", "text": "Using the routing rubric, decide the learner's intervention level and support route."}]}, "method": "c2d", "provenance": {"source_id": "diverse-065", "source_is_synthetic": true, "source_sha256": "ba943183d6d8d5bb56edf32785dde7f9da871042111d90ad4891f90a3c367989", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-01", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar and coordinator policy statements verbatim, keep the EDU 340/quantitative-placement-test bindings, and the focus evidence are two complete factual score-report sentences; the counterfactual only edits the testing-center export figure (80→74) without duplicating or contradicting other measurements, and neither context states or hints at a final decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Academic adviser\",\"text\":\"The transcript and signed instructor note are present in the complete exception file.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The instructor note submitted for the student's EDU 340 exception request recommends admission and is signed.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["4", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80."}, {"path": ["5", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-007", "id": "fast-43-diverse-075-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Academic adviser", "text": "The transcript and signed instructor note are present in the complete exception file."}, {"speaker": "Academic adviser", "text": "The instructor note submitted for the student's EDU 340 exception request recommends admission and is signed."}, {"speaker": "Academic adviser", "text": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"speaker": "Academic adviser", "text": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"speaker": "Academic adviser", "text": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80."}, {"speaker": "Academic adviser", "text": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar and coordinator policy statements verbatim, keep the EDU 340/quantitative-placement-test bindings, and the focus evidence are two complete factual score-report sentences; the counterfactual only edits the testing-center export figure (80→74) without duplicating or contradicting other measurements, and neither context states or hints at a final decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Academic adviser\",\"text\":\"The transcript and signed instructor note are present in the complete exception file.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The instructor note submitted for the student's EDU 340 exception request recommends admission and is signed.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["4", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80."}, {"path": ["5", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-007", "id": "fast-43-diverse-075-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Academic adviser", "text": "The transcript and signed instructor note are present in the complete exception file."}, {"speaker": "Academic adviser", "text": "The instructor note submitted for the student's EDU 340 exception request recommends admission and is signed."}, {"speaker": "Academic adviser", "text": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"speaker": "Academic adviser", "text": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"speaker": "Academic adviser", "text": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80."}, {"speaker": "Academic adviser", "text": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 74."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar and coordinator policy statements verbatim, keep the EDU 340/quantitative-placement-test/registrar-permission bindings intact, cite two complete factual score-report sentences as evidence, and the counterfactual's single-sentence change (74 instead of 80) creates a coherent conflicting-score scenario without contradicting other unchanged facts or leaking any gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B, and my testing records show quantitative placement scores.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The complete documents include the transcript, signed instructor note, testing portal report, and dated testing-center export. The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 80. The portal score is marked verified, and no other quantitative placement-test score reports appear in the complete documents.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review. The signed instructor note recommends admission.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80."}, {"path": ["1", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-013", "id": "fast-43-diverse-075-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B, and my testing records show quantitative placement scores."}, {"speaker": "Academic adviser", "text": "The complete documents include the transcript, signed instructor note, testing portal report, and dated testing-center export. The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 80. The portal score is marked verified, and no other quantitative placement-test score reports appear in the complete documents."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review. The signed instructor note recommends admission."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar and coordinator policy statements verbatim, keep the EDU 340/quantitative-placement-test/registrar-permission bindings intact, cite two complete factual score-report sentences as evidence, and the counterfactual's single-sentence change (74 instead of 80) creates a coherent conflicting-score scenario without contradicting other unchanged facts or leaking any gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B, and my testing records show quantitative placement scores.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The complete documents include the transcript, signed instructor note, testing portal report, and dated testing-center export. The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 80. The portal score is marked verified, and no other quantitative placement-test score reports appear in the complete documents.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review. The signed instructor note recommends admission.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80."}, {"path": ["1", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-013", "id": "fast-43-diverse-075-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B, and my testing records show quantitative placement scores."}, {"speaker": "Academic adviser", "text": "The complete documents include the transcript, signed instructor note, testing portal report, and dated testing-center export. The testing portal report for the student's EDU 340 quantitative placement test shows a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test shows a score of 74. The portal score is marked verified, and no other quantitative placement-test score reports appear in the complete documents."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review. The signed instructor note recommends admission."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Registrar and coordinator rubric wording is unchanged and the original question is applied verbatim, preserving all governing policy and entity/time bindings; the two focus sentences are plain factual score reports, not policy text; the counterfactual’s change from 80 to 74 creates a genuine straddling conflict consistent with the added verified-status fact and remaining unchanged sentences, with no embedded answer, rule table, or rationale in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The dated testing-center export lists the student's EDU 340 quantitative placement test score as 80 points.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The transcript and signed instructor note are present, and the note recommends admission. The portal score carries verified status, and the complete documents contain no other quantitative placement-test score reports beyond these two.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points."}, {"path": ["2", "text"], "text": "The dated testing-center export lists the student's EDU 340 quantitative placement test score as 80 points."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points.", "negative_left": "The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points.", "negative_right": "The dated testing-center export lists the student's EDU 340 quantitative placement test score as 74 points.", "right": "The dated testing-center export lists the student's EDU 340 quantitative placement test score as 80 points."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-016", "id": "fast-43-diverse-075-016-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points."}, {"speaker": "Academic adviser", "text": "The dated testing-center export lists the student's EDU 340 quantitative placement test score as 80 points."}, {"speaker": "Academic adviser", "text": "The transcript and signed instructor note are present, and the note recommends admission. The portal score carries verified status, and the complete documents contain no other quantitative placement-test score reports beyond these two."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Registrar and coordinator rubric wording is unchanged and the original question is applied verbatim, preserving all governing policy and entity/time bindings; the two focus sentences are plain factual score reports, not policy text; the counterfactual’s change from 80 to 74 creates a genuine straddling conflict consistent with the added verified-status fact and remaining unchanged sentences, with no embedded answer, rule table, or rationale in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The dated testing-center export lists the student's EDU 340 quantitative placement test score as 80 points.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The transcript and signed instructor note are present, and the note recommends admission. The portal score carries verified status, and the complete documents contain no other quantitative placement-test score reports beyond these two.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points."}, {"path": ["2", "text"], "text": "The dated testing-center export lists the student's EDU 340 quantitative placement test score as 80 points."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points.", "negative_left": "The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points.", "negative_right": "The dated testing-center export lists the student's EDU 340 quantitative placement test score as 74 points.", "right": "The dated testing-center export lists the student's EDU 340 quantitative placement test score as 80 points."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-016", "id": "fast-43-diverse-075-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The testing portal report lists the student's EDU 340 quantitative placement test score as 80 points."}, {"speaker": "Academic adviser", "text": "The dated testing-center export lists the student's EDU 340 quantitative placement test score as 74 points."}, {"speaker": "Academic adviser", "text": "The transcript and signed instructor note are present, and the note recommends admission. The portal score carries verified status, and the complete documents contain no other quantitative placement-test score reports beyond these two."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar/coordinator routing rules and readiness rubric verbatim, keep the EDU 340 entity and same actors/time; the two focus sentences ('portal reports 80'/'export reports 80' vs 'export reports 74') are plain factual reports, not policy text; the counterfactual's 74 vs 80 creates a straddling conflict consistent with the borderline rule rather than a contradictory duplicate; no rule table, gold label, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The file is complete: transcript, signed instructor note recommending admission, testing portal report, and dated testing-center export. No other quantitative placement-test score reports appear anywhere in the file. The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test. The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test. The portal score is marked verified in the system.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test."}, {"path": ["1", "text"], "text": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.", "negative_left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.", "negative_right": "The dated testing-center export reports a score of 74 for the student's EDU 340 quantitative placement test.", "right": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-022", "id": "fast-43-diverse-075-022-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The file is complete: transcript, signed instructor note recommending admission, testing portal report, and dated testing-center export. No other quantitative placement-test score reports appear anywhere in the file. The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test. The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test. The portal score is marked verified in the system."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar/coordinator routing rules and readiness rubric verbatim, keep the EDU 340 entity and same actors/time; the two focus sentences ('portal reports 80'/'export reports 80' vs 'export reports 74') are plain factual reports, not policy text; the counterfactual's 74 vs 80 creates a straddling conflict consistent with the borderline rule rather than a contradictory duplicate; no rule table, gold label, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The file is complete: transcript, signed instructor note recommending admission, testing portal report, and dated testing-center export. No other quantitative placement-test score reports appear anywhere in the file. The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test. The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test. The portal score is marked verified in the system.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test."}, {"path": ["1", "text"], "text": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.", "negative_left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.", "negative_right": "The dated testing-center export reports a score of 74 for the student's EDU 340 quantitative placement test.", "right": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-022", "id": "fast-43-diverse-075-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The file is complete: transcript, signed instructor note recommending admission, testing portal report, and dated testing-center export. No other quantitative placement-test score reports appear anywhere in the file. The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test. The dated testing-center export reports a score of 74 for the student's EDU 340 quantitative placement test. The portal score is marked verified in the system."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the registrar routing rule and coordinator readiness rubric intact, preserve the EDU 340/quantitative-test bindings, and the two focus sentences are plain factual score reports; the counterfactual changes only the testing-center export score from 80 to 74, creating a coherent straddling conflict without duplicating a contradictory identical measurement, and no text reveals or hints at the correct yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The complete file contains the official transcript, a signed instructor note, the testing portal report, and the dated testing-center export, with no other quantitative score reports present.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The dated testing-center export lists a score of 80 for the student's EDU 340 quantitative placement test.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The testing portal score is marked verified, and the signed instructor note recommends admission.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test."}, {"path": ["3", "text"], "text": "The dated testing-center export lists a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test.", "negative_left": "The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test.", "negative_right": "The dated testing-center export lists a score of 74 for the student's EDU 340 quantitative placement test.", "right": "The dated testing-center export lists a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-023", "id": "fast-43-diverse-075-023-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file contains the official transcript, a signed instructor note, the testing portal report, and the dated testing-center export, with no other quantitative score reports present."}, {"speaker": "Academic adviser", "text": "The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test."}, {"speaker": "Academic adviser", "text": "The dated testing-center export lists a score of 80 for the student's EDU 340 quantitative placement test."}, {"speaker": "Academic adviser", "text": "The testing portal score is marked verified, and the signed instructor note recommends admission."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the registrar routing rule and coordinator readiness rubric intact, preserve the EDU 340/quantitative-test bindings, and the two focus sentences are plain factual score reports; the counterfactual changes only the testing-center export score from 80 to 74, creating a coherent straddling conflict without duplicating a contradictory identical measurement, and no text reveals or hints at the correct yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The complete file contains the official transcript, a signed instructor note, the testing portal report, and the dated testing-center export, with no other quantitative score reports present.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The dated testing-center export lists a score of 80 for the student's EDU 340 quantitative placement test.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The testing portal score is marked verified, and the signed instructor note recommends admission.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test."}, {"path": ["3", "text"], "text": "The dated testing-center export lists a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test.", "negative_left": "The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test.", "negative_right": "The dated testing-center export lists a score of 74 for the student's EDU 340 quantitative placement test.", "right": "The dated testing-center export lists a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-023", "id": "fast-43-diverse-075-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file contains the official transcript, a signed instructor note, the testing portal report, and the dated testing-center export, with no other quantitative score reports present."}, {"speaker": "Academic adviser", "text": "The testing portal report lists a score of 80 for the student's EDU 340 quantitative placement test."}, {"speaker": "Academic adviser", "text": "The dated testing-center export lists a score of 74 for the student's EDU 340 quantitative placement test."}, {"speaker": "Academic adviser", "text": "The testing portal score is marked verified, and the signed instructor note recommends admission."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar and coordinator policy statements verbatim, preserve the EDU 340 entity, test types, and March 4/5 dates; the two focus evidence spans are complete factual score statements, not policy text; the counterfactual's single change (testing-center score 80→74) creates a genuine straddling conflict with the verified 80 portal score consistent with unchanged rubric language, and neither context states or implies the correct yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B, and I understand my quantitative placement test scores are in the file.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The complete EDU 340 exception file contains no quantitative placement-test score reports besides the testing portal report and the dated testing-center export. The instructor note in the file is signed and recommends admission. The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4. The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80, dated March 5. The testing portal's score carries verified status in our system.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4."}, {"path": ["1", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80, dated March 5."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 74, dated March 5.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80, dated March 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-029", "id": "fast-43-diverse-075-029-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B, and I understand my quantitative placement test scores are in the file."}, {"speaker": "Academic adviser", "text": "The complete EDU 340 exception file contains no quantitative placement-test score reports besides the testing portal report and the dated testing-center export. The instructor note in the file is signed and recommends admission. The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4. The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80, dated March 5. The testing portal's score carries verified status in our system."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar and coordinator policy statements verbatim, preserve the EDU 340 entity, test types, and March 4/5 dates; the two focus evidence spans are complete factual score statements, not policy text; the counterfactual's single change (testing-center score 80→74) creates a genuine straddling conflict with the verified 80 portal score consistent with unchanged rubric language, and neither context states or implies the correct yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B, and I understand my quantitative placement test scores are in the file.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The complete EDU 340 exception file contains no quantitative placement-test score reports besides the testing portal report and the dated testing-center export. The instructor note in the file is signed and recommends admission. The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4. The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80, dated March 5. The testing portal's score carries verified status in our system.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4."}, {"path": ["1", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80, dated March 5."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 74, dated March 5.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 80, dated March 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-029", "id": "fast-43-diverse-075-029-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B, and I understand my quantitative placement test scores are in the file."}, {"speaker": "Academic adviser", "text": "The complete EDU 340 exception file contains no quantitative placement-test score reports besides the testing portal report and the dated testing-center export. The instructor note in the file is signed and recommends admission. The testing portal report for the student's EDU 340 quantitative placement test lists a score of 80, dated March 4. The dated testing-center export for the student's EDU 340 quantitative placement test lists a score of 74, dated March 5. The testing portal's score carries verified status in our system."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original registrar/coordinator policy verbatim and the same EDU 340 exception entity/path/time bindings; the base context turns the case into an agreeing 80/80 score scenario while the counterfactual restores a genuine score conflict (80 verified vs 74) without duplicating or contradicting other unchanged facts; the two focus evidence spans are complete factual sentences about test scores, not policy or instructions; and neither context states or hints at the gold yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Adviser note\",\"text\":\"Student file for EDU 340 exception is complete, containing only the transcript, signed instructor note, and two placement-test score reports.\"},{\"speaker\":\"Testing portal record\",\"text\":\"The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.\"},{\"speaker\":\"Testing portal record\",\"text\":\"The portal marks this score as verified.\"},{\"speaker\":\"Testing-center export\",\"text\":\"The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test.\"},{\"speaker\":\"Instructor note\",\"text\":\"The signed instructor note recommends admission for the student.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test."}, {"path": ["3", "text"], "text": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.", "negative_left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.", "negative_right": "The dated testing-center export reports a score of 74 for the student's EDU 340 quantitative placement test.", "right": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-030", "id": "fast-43-diverse-075-030-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Adviser note", "text": "Student file for EDU 340 exception is complete, containing only the transcript, signed instructor note, and two placement-test score reports."}, {"speaker": "Testing portal record", "text": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test."}, {"speaker": "Testing portal record", "text": "The portal marks this score as verified."}, {"speaker": "Testing-center export", "text": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}, {"speaker": "Instructor note", "text": "The signed instructor note recommends admission for the student."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original registrar/coordinator policy verbatim and the same EDU 340 exception entity/path/time bindings; the base context turns the case into an agreeing 80/80 score scenario while the counterfactual restores a genuine score conflict (80 verified vs 74) without duplicating or contradicting other unchanged facts; the two focus evidence spans are complete factual sentences about test scores, not policy or instructions; and neither context states or hints at the gold yes/no answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Adviser note\",\"text\":\"Student file for EDU 340 exception is complete, containing only the transcript, signed instructor note, and two placement-test score reports.\"},{\"speaker\":\"Testing portal record\",\"text\":\"The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.\"},{\"speaker\":\"Testing portal record\",\"text\":\"The portal marks this score as verified.\"},{\"speaker\":\"Testing-center export\",\"text\":\"The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test.\"},{\"speaker\":\"Instructor note\",\"text\":\"The signed instructor note recommends admission for the student.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test."}, {"path": ["3", "text"], "text": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.", "negative_left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test.", "negative_right": "The dated testing-center export reports a score of 74 for the student's EDU 340 quantitative placement test.", "right": "The dated testing-center export reports a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-030", "id": "fast-43-diverse-075-030-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Adviser note", "text": "Student file for EDU 340 exception is complete, containing only the transcript, signed instructor note, and two placement-test score reports."}, {"speaker": "Testing portal record", "text": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test."}, {"speaker": "Testing portal record", "text": "The portal marks this score as verified."}, {"speaker": "Testing-center export", "text": "The dated testing-center export reports a score of 74 for the student's EDU 340 quantitative placement test."}, {"speaker": "Instructor note", "text": "The signed instructor note recommends admission for the student."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the registrar and coordinator policy statements verbatim, preserve the EDU 340 entity/time scope, and the two focus-evidence sentences are plain factual score reports (not rule text); the counterfactual’s single-sentence change of the testing-center export from 80 to 74 introduces a genuine score conflict that is internally consistent with the unchanged registrar/coordinator rules and does not duplicate or contradict other measurements, and neither context states or implies the gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The complete file includes the transcript, a signed instructor note recommending admission, and two quantitative placement-test score reports — no others exist. The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80. The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 80 for that same test. The portal's score entry is marked as verified in the system.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80."}, {"path": ["1", "text"], "text": "The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 80 for that same test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80.", "negative_right": "The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 74 for that same test.", "right": "The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 80 for that same test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-034", "id": "fast-43-diverse-075-034-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file includes the transcript, a signed instructor note recommending admission, and two quantitative placement-test score reports — no others exist. The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80. The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 80 for that same test. The portal's score entry is marked as verified in the system."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the registrar and coordinator policy statements verbatim, preserve the EDU 340 entity/time scope, and the two focus-evidence sentences are plain factual score reports (not rule text); the counterfactual’s single-sentence change of the testing-center export from 80 to 74 introduces a genuine score conflict that is internally consistent with the unchanged registrar/coordinator rules and does not duplicate or contradict other measurements, and neither context states or implies the gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The complete file includes the transcript, a signed instructor note recommending admission, and two quantitative placement-test score reports — no others exist. The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80. The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 80 for that same test. The portal's score entry is marked as verified in the system.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80."}, {"path": ["1", "text"], "text": "The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 80 for that same test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80.", "negative_right": "The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 74 for that same test.", "right": "The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 80 for that same test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-034", "id": "fast-43-diverse-075-034-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file includes the transcript, a signed instructor note recommending admission, and two quantitative placement-test score reports — no others exist. The testing portal report for the student's EDU 340 quantitative placement test, logged under record ID TP-5521, shows a score of 80. The dated testing-center export, cross-referenced to record ID TP-5521, shows a score of 74 for that same test. The portal's score entry is marked as verified in the system."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full readiness rubric and routing rule verbatim, the focus evidence consists of two complete factual sentences about score reports, the counterfactual changes only the export score (80\\u219274) creating a coherent conflicting-score scenario consistent with policy, all entity/path/time bindings (EDU 340, registrar staff, March 3) are unchanged, and neither context states or implies the classifier's answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The complete file contains only two score reports for the quantitative placement test: the testing portal report and the dated testing-center export, both dated March 3. The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80. The portal's score entry is flagged as verified in the system. The instructor note in the file is signed and recommends admission.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, {"path": ["1", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-035", "id": "fast-43-diverse-075-035-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file contains only two score reports for the quantitative placement test: the testing portal report and the dated testing-center export, both dated March 3. The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80. The portal's score entry is flagged as verified in the system. The instructor note in the file is signed and recommends admission."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full readiness rubric and routing rule verbatim, the focus evidence consists of two complete factual sentences about score reports, the counterfactual changes only the export score (80\\u219274) creating a coherent conflicting-score scenario consistent with policy, all entity/path/time bindings (EDU 340, registrar staff, March 3) are unchanged, and neither context states or implies the classifier's answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The complete file contains only two score reports for the quantitative placement test: the testing portal report and the dated testing-center export, both dated March 3. The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80. The portal's score entry is flagged as verified in the system. The instructor note in the file is signed and recommends admission.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, {"path": ["1", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-035", "id": "fast-43-diverse-075-035-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file contains only two score reports for the quantitative placement test: the testing portal report and the dated testing-center export, both dated March 3. The testing portal report for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 74. The portal's score entry is flagged as verified in the system. The instructor note in the file is signed and recommends admission."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar and coordinator rubric verbatim, keep the EDU 340/placement-test/March 3rd bindings, use two verbatim factual score-report sentences as evidence, and the counterfactual's single 80\\u2192conflicting change to 74 creates a plausible straddling scenario without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The complete file contains only two quantitative placement-test score reports: the testing portal report and the dated testing-center export. The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd. The dated testing-center export, dated March 3rd, reports a score of 80 for the student's EDU 340 quantitative placement test. The portal score carries verified status in our system, and the signed instructor note recommends admission.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd."}, {"path": ["1", "text"], "text": "The dated testing-center export, dated March 3rd, reports a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd.", "negative_left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd.", "negative_right": "The dated testing-center export, dated March 3rd, reports a score of 74 for the student's EDU 340 quantitative placement test.", "right": "The dated testing-center export, dated March 3rd, reports a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-039", "id": "fast-43-diverse-075-039-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file contains only two quantitative placement-test score reports: the testing portal report and the dated testing-center export. The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd. The dated testing-center export, dated March 3rd, reports a score of 80 for the student's EDU 340 quantitative placement test. The portal score carries verified status in our system, and the signed instructor note recommends admission."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the registrar and coordinator rubric verbatim, keep the EDU 340/placement-test/March 3rd bindings, use two verbatim factual score-report sentences as evidence, and the counterfactual's single 80\\u2192conflicting change to 74 creates a plausible straddling scenario without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\": \"Student\", \"text\": \"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"}, {\"speaker\": \"Academic adviser\", \"text\": \"The complete file contains only two quantitative placement-test score reports: the testing portal report and the dated testing-center export. The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd. The dated testing-center export, dated March 3rd, reports a score of 80 for the student's EDU 340 quantitative placement test. The portal score carries verified status in our system, and the signed instructor note recommends admission.\"}, {\"speaker\": \"Registrar staff member\", \"text\": \"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"}, {\"speaker\": \"Course coordinator\", \"text\": \"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd."}, {"path": ["1", "text"], "text": "The dated testing-center export, dated March 3rd, reports a score of 80 for the student's EDU 340 quantitative placement test."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd.", "negative_left": "The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd.", "negative_right": "The dated testing-center export, dated March 3rd, reports a score of 74 for the student's EDU 340 quantitative placement test.", "right": "The dated testing-center export, dated March 3rd, reports a score of 80 for the student's EDU 340 quantitative placement test."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-039", "id": "fast-43-diverse-075-039-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file contains only two quantitative placement-test score reports: the testing portal report and the dated testing-center export. The testing portal reports a score of 80 for the student's EDU 340 quantitative placement test, submitted on March 3rd. The dated testing-center export, dated March 3rd, reports a score of 74 for the student's EDU 340 quantitative placement test. The portal score carries verified status in our system, and the signed instructor note recommends admission."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the registrar/coordinator routing policy verbatim and the student/course/test/date bindings intact; the base context changes the testing-center score to 80 (matching the portal, a coherent non-conflicting observation) while the counterfactual changes it to 74 (a genuine conflicting-score observation consistent with the stated routing rule), and the two focus-evidence spans are plain factual report statements with no embedded rules, IDs, or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The complete file includes the transcript, a signed instructor note recommending admission, and two quantitative placement-test score reports; no other score reports exist in the file. The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.\"},{\"speaker\":\"Testing center\",\"text\":\"The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one. The portal score is marked verified in our system.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, {"path": ["2", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.", "negative_left": "The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-040", "id": "fast-43-diverse-075-040-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file includes the transcript, a signed instructor note recommending admission, and two quantitative placement-test score reports; no other score reports exist in the file. The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, {"speaker": "Testing center", "text": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one. The portal score is marked verified in our system."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the registrar/coordinator routing policy verbatim and the student/course/test/date bindings intact; the base context changes the testing-center score to 80 (matching the portal, a coherent non-conflicting observation) while the counterfactual changes it to 74 (a genuine conflicting-score observation consistent with the stated routing rule), and the two focus-evidence spans are plain factual report statements with no embedded rules, IDs, or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The complete file includes the transcript, a signed instructor note recommending admission, and two quantitative placement-test score reports; no other score reports exist in the file. The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.\"},{\"speaker\":\"Testing center\",\"text\":\"The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one. The portal score is marked verified in our system.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, {"path": ["2", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.", "negative_left": "The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-040", "id": "fast-43-diverse-075-040-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My official transcript shows Research Methods 220 with a B."}, {"speaker": "Academic adviser", "text": "The complete file includes the transcript, a signed instructor note recommending admission, and two quantitative placement-test score reports; no other score reports exist in the file. The testing portal report logged for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 80."}, {"speaker": "Testing center", "text": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 3, shows a score of 74."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one. The portal score is marked verified in our system."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the registrar/adviser/coordinator policy language and the EDU 340/March 4 bindings intact, the two evidence spans are complete factual sentences about score reports, and the counterfactual's single change (74 vs 80) coherently creates a straddling conflict without contradicting other unchanged facts or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My transcript shows Research Methods 220 with a B, and my quantitative placement scores appear in two separate reports.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The transcript and a signed instructor note recommending admission are on file. The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80. No other score reports appear anywhere in the file, and the portal entry is marked as verified.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80."}, {"path": ["1", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-042", "id": "fast-43-diverse-075-042-base", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My transcript shows Research Methods 220 with a B, and my quantitative placement scores appear in two separate reports."}, {"speaker": "Academic adviser", "text": "The transcript and a signed instructor note recommending admission are on file. The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80. No other score reports appear anywhere in the file, and the portal entry is marked as verified."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the registrar/adviser/coordinator policy language and the EDU 340/March 4 bindings intact, the two evidence spans are complete factual sentences about score reports, and the counterfactual's single change (74 vs 80) coherently creates a straddling conflict without contradicting other unchanged facts or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; none is a bundled policy conclusion or final classification. The focus atom concerns factual equality between two score reports. The base and counter assignments are both realizable while changing only that equality: the base can have scores 80 and 80, while the counter can have 79 and 80, with all other atom states unchanged. The policy evidence preserves the substantive routing and readiness rules originating in the original state; instructions and criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Equality of the two reported scores together with their maximum being 80 entails that both are 80. The portal score is verified, the complete record has no other score report, and the signed note recommends admission. Thus the stated Ready requirements are met and no conflicting-score review is triggered.", "rule_index": 0, "sound": true}, {"reason": "If the two scores are unequal and their maximum is 80, one is 80 and the other is below 80. These are conflicting results straddling the threshold. The explicit policy requires coordinator review for such a conflict and prohibits registrar staff from selecting one score, regardless of which report has verified status.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The score reported by the testing portal for the student's EDU 340 quantitative placement test equals the score reported by the dated testing-center export for that test."}, {"id": "a2", "statement": "The maximum of the scores reported by the testing portal and the dated testing-center export for the student's EDU 340 quantitative placement test is 80."}, {"id": "a3", "statement": "The complete documents for the student's EDU 340 exception request contain no quantitative placement-test score reports other than the testing portal report and the dated testing-center export."}, {"id": "a4", "statement": "The score in the testing portal report for the student's EDU 340 quantitative placement test has verified status."}, {"id": "a5", "statement": "The instructor note submitted for the student's EDU 340 exception request is signed."}, {"id": "a6", "statement": "The instructor note submitted for the student's EDU 340 exception request recommends admission."}], "base_state_json": "[{\"speaker\":\"Student\",\"text\":\"I request an exception for EDU 340. My transcript shows Research Methods 220 with a B, and my quantitative placement scores appear in two separate reports.\"},{\"speaker\":\"Academic adviser\",\"text\":\"The transcript and a signed instructor note recommending admission are on file. The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80. No other score reports appear anywhere in the file, and the portal entry is marked as verified.\"},{\"speaker\":\"Registrar staff member\",\"text\":\"EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one.\"},{\"speaker\":\"Course coordinator\",\"text\":\"Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80."}, {"path": ["1", "text"], "text": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80."}], "policy_evidence": [{"path": ["2", "text"], "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"path": ["3", "text"], "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}], "rules": [{"justification": "The two scores are both 80, the portal score is explicitly verified, and the complete documents contain no other score report that could create a conflict. The signed note recommends admission, so the case is Ready and does not require conflicting-score review.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "The two scores are unequal and have a maximum of 80, so one is 80 and the other is below 80. With the signed admission recommendation, these are conflicting results straddling 80, making the case Borderline and requiring course-coordinator review.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}]}, "verified_pair": {"left": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80.", "negative_left": "The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80.", "negative_right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 74.", "right": "The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80."}, "verifier_independent_model": false}, "family": "fast-43-diverse-075-042", "id": "fast-43-diverse-075-042-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — registrar staff must route the complete but conflicting boundary case to the course coordinator and may not directly recommend the exception.", "true": "Yes — registrar staff may recommend the exception directly without coordinator review."}, "instructions": "Verification and routing task: Is the registrar staff member permitted to make a yes recommendation on the prerequisite exception without routing the case? Treat the documents as complete, apply the stated readiness rubric, and answer yes or no.", "type": "noul"}}, "state": [{"speaker": "Student", "text": "I request an exception for EDU 340. My transcript shows Research Methods 220 with a B, and my quantitative placement scores appear in two separate reports."}, {"speaker": "Academic adviser", "text": "The transcript and a signed instructor note recommending admission are on file. The testing portal report for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 80. The dated testing-center export for the student's EDU 340 quantitative placement test, dated March 4, lists a score of 74. No other score reports appear anywhere in the file, and the portal entry is marked as verified."}, {"speaker": "Registrar staff member", "text": "EDU 340 exceptions require a verified score of at least 80 and a signed recommendation. Conflicting scores must go to the course coordinator; registrar staff cannot select one."}, {"speaker": "Course coordinator", "text": "Readiness is ordered as Ready: verified score 80+ with recommendation; Borderline: conflicting results straddling 80 with recommendation; Not ready: verified score below 80 or no recommendation. Borderline cases require my review."}]}, "method": "c2d", "provenance": {"source_id": "diverse-075", "source_is_synthetic": true, "source_sha256": "f5e3f82432c3463e29df48e284263138adc333724ee06902b4ff52fa76777aba", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy verbatim and the question object is unchanged; entity, evidence-path, and time bindings for Mina Cho/CHEM 220 are preserved; the two focus evidence items are plain factual date statements, not policy or instructions; the counterfactual only changes the placement report date (2024→2022), consistently making the score expired relative to the unchanged Jan 6, 2025 term start, with no contradictory duplicate facts; no context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2024.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2024."}, {"path": ["evidence", "4"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-001", "id": "fast-43-diverse-077-001-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2024.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy verbatim and the question object is unchanged; entity, evidence-path, and time bindings for Mina Cho/CHEM 220 are preserved; the two focus evidence items are plain factual date statements, not policy or instructions; the counterfactual only changes the placement report date (2024→2022), consistently making the score expired relative to the unchanged Jan 6, 2025 term start, with no contradictory duplicate facts; no context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2024.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2024."}, {"path": ["evidence", "4"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-001", "id": "fast-43-diverse-077-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "All policy limits from the original state are copied verbatim into both contexts; the entity, course, and time bindings remain fixed; the two focus evidence lines are concrete factual statements about dates, not policy or instructions; the counterfactual shifts the placement date to 2023, coherently moving the case outside the 12-month window per stated policy; neither context reveals the gold recommendation, routing choice, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2024.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2024."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-002", "id": "fast-43-diverse-077-002-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2024.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "All policy limits from the original state are copied verbatim into both contexts; the entity, course, and time bindings remain fixed; the two focus evidence lines are concrete factual statements about dates, not policy or instructions; the counterfactual shifts the placement date to 2023, coherently moving the case outside the 12-month window per stated policy; neither context reveals the gold recommendation, routing choice, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2024.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2024."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-002", "id": "fast-43-diverse-077-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original prerequisite, score, and routing policy without added exceptions, keep the same student/course/time bindings, use two complete factual sentences as evidence, and the counterfactual coherently shifts only the placement report date (2023→2022) with no contradictory duplicate facts or leaked answer cues.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2023.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins January 8, 2024.\"], \"request\": \"Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-004", "id": "fast-43-diverse-077-004-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2023.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins January 8, 2024."], "request": "Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original prerequisite, score, and routing policy without added exceptions, keep the same student/course/time bindings, use two complete factual sentences as evidence, and the counterfactual coherently shifts only the placement report date (2023→2022) with no contradictory duplicate facts or leaked answer cues.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2023.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins January 8, 2024.\"], \"request\": \"Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-004", "id": "fast-43-diverse-077-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins January 8, 2024."], "request": "Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing prerequisite/exception policy and scope limits verbatim, alongside the unchanged question; entity, course, and date bindings are preserved; the two focus evidence lines are factual date statements, not policy or instructions; the counterfactual shifts the placement date to over 12 months prior, which stays coherent with the unchanged term-start evidence and existing panel-routing rule without contradicting other facts; no gold answer, rationale, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2023.\", \"Mina Cho's requested CHEM 220 term begins on January 10, 2024.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "3"], "text": "Mina Cho's requested CHEM 220 term begins on January 10, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 10, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 10, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-005", "id": "fast-43-diverse-077-005-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2023.", "Mina Cho's requested CHEM 220 term begins on January 10, 2024.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing prerequisite/exception policy and scope limits verbatim, alongside the unchanged question; entity, course, and date bindings are preserved; the two focus evidence lines are factual date statements, not policy or instructions; the counterfactual shifts the placement date to over 12 months prior, which stays coherent with the unchanged term-start evidence and existing panel-routing rule without contradicting other facts; no gold answer, rationale, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2023.\", \"Mina Cho's requested CHEM 220 term begins on January 10, 2024.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "3"], "text": "Mina Cho's requested CHEM 220 term begins on January 10, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 10, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 10, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-005", "id": "fast-43-diverse-077-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "Mina Cho's requested CHEM 220 term begins on January 10, 2024.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 12-month substitution rule, coordinator scope limit, and complete-file requirement without adding exceptions; all evidence items are standalone factual statements rather than policy or rule text; the counterfactual only changes the placement report date (2023→2022), shifting it outside the 12-month window relative to the unchanged Jan 8, 2024 term start, which is coherent and non-contradictory; entity, course, and date bindings match the original question; no gold answers, codes, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2023.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser confirms all three required documents—transcript, placement report, and instructor note—are on file with the coordinator.\",\"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\"]}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-006", "id": "fast-43-diverse-077-006-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2023.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser confirms all three required documents—transcript, placement report, and instructor note—are on file with the coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024."]}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 12-month substitution rule, coordinator scope limit, and complete-file requirement without adding exceptions; all evidence items are standalone factual statements rather than policy or rule text; the counterfactual only changes the placement report date (2023→2022), shifting it outside the 12-month window relative to the unchanged Jan 8, 2024 term start, which is coherent and non-contradictory; entity, course, and date bindings match the original question; no gold answers, codes, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2023.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser confirms all three required documents—transcript, placement report, and instructor note—are on file with the coordinator.\",\"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\"]}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-006", "id": "fast-43-diverse-077-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser confirms all three required documents—transcript, placement report, and instructor note—are on file with the coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024."]}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy paragraph verbatim and the same student/course/time bindings; evidence items are discrete factual statements without embedded rules or outcomes; the counterfactual only changes the placement date (2023→2022), consistently altering the elapsed time versus the unchanged Jan 8 2024 term without contradicting other facts; no gold answer, routing decision, or rationale is stated in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2023.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-007", "id": "fast-43-diverse-077-007-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2023.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy paragraph verbatim and the same student/course/time bindings; evidence items are discrete factual statements without embedded rules or outcomes; the counterfactual only changes the placement date (2023→2022), consistently altering the elapsed time versus the unchanged Jan 8 2024 term without contradicting other facts; no gold answer, routing decision, or rationale is stated in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2023.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-007", "id": "fast-43-diverse-077-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy text and required documentation rules from the original state without adding or omitting exceptions; entity, course, and date bindings match the original question; the two focus evidence items are plain factual sentences about dates, not policy or instructions; the counterfactual coherently shifts the placement date from 2023 to 2022 without contradicting other evidence; and neither context contains a gold answer, rule table, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2023.\",\"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "3"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-012", "id": "fast-43-diverse-077-012-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2023.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy text and required documentation rules from the original state without adding or omitting exceptions; entity, course, and date bindings match the original question; the two focus evidence items are plain factual sentences about dates, not policy or instructions; the counterfactual coherently shifts the placement date from 2023 to 2022 without contradicting other evidence; and neither context contains a gold answer, rule table, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2023.\",\"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "3"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-012", "id": "fast-43-diverse-077-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text on prerequisites, score substitution, 12-month limit, and required documents, matching the unchanged question object; entity (Mina Cho), course (CHEM 220), and term date (Jan 8, 2024) bindings are preserved; each evidence item is a plain factual sentence without policy definitions or instructions; the counterfactual only shifts the placement-report date from March 1, 2023 to March 1, 2021, a coherent single-fact change consistent with all other unchanged evidence; neither context states a recommendation, routing decision, or rule ID, so no answer leakage occurs.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2023.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\",\"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\"],\"request\":\"Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2021.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-014", "id": "fast-43-diverse-077-014-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2023.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024."], "request": "Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text on prerequisites, score substitution, 12-month limit, and required documents, matching the unchanged question object; entity (Mina Cho), course (CHEM 220), and term date (Jan 8, 2024) bindings are preserved; each evidence item is a plain factual sentence without policy definitions or instructions; the counterfactual only shifts the placement-report date from March 1, 2023 to March 1, 2021, a coherent single-fact change consistent with all other unchanged evidence; neither context states a recommendation, routing decision, or rule ID, so no answer leakage occurs.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2023.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\",\"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\"],\"request\":\"Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2021.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-014", "id": "fast-43-diverse-077-014-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report in the CHEM 220 exception file is dated March 1, 2021.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024."], "request": "Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy text (12-month score window, coordinator scope limit, panel routing, complete-file requirement) and the same entity/course/time bindings; the two focus sentences are complete factual statements, not policy or instructions; the counterfactual only shifts the placement-report date (2023→2022), consistently altering the score's validity window without contradicting other unchanged evidence; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2023.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\", \"A registrar staff member confirms that expired placement-score cases require panel review.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-015", "id": "fast-43-diverse-077-015-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2023.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "A registrar staff member confirms that expired placement-score cases require panel review."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy text (12-month score window, coordinator scope limit, panel routing, complete-file requirement) and the same entity/course/time bindings; the two focus sentences are complete factual statements, not policy or instructions; the counterfactual only shifts the placement-report date (2023→2022), consistently altering the score's validity window without contradicting other unchanged evidence; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2023.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\", \"A registrar staff member confirms that expired placement-score cases require panel review.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-015", "id": "fast-43-diverse-077-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "A registrar staff member confirms that expired placement-score cases require panel review."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy paragraph and unchanged question bindings for Mina/CHEM220; the two focus evidence sentences are plain factual statements about dates, not rules; the counterfactual only shifts the placement-report date to 2022, making it fall outside the 12-month window without contradicting other facts, and neither context states a recommendation, rule table, or answer code.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2024.\",\"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2024."}, {"path": ["evidence", "3"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-016", "id": "fast-43-diverse-077-016-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2024.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical governing policy paragraph and unchanged question bindings for Mina/CHEM220; the two focus evidence sentences are plain factual statements about dates, not rules; the counterfactual only shifts the placement-report date to 2022, making it fall outside the 12-month window without contradicting other facts, and neither context states a recommendation, rule table, or answer code.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2024.\",\"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2024."}, {"path": ["evidence", "3"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-016", "id": "fast-43-diverse-077-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina's readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate the same governing policy and entity/course/date bindings; the two focus evidence sentences are plain factual statements, and the counterfactual only changes the placement report date to 2022, remaining logically consistent without contradicting other evidence or leaking any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2024.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\",\"Mina Cho's requested CHEM 220 term begins on January 10, 2025.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2024."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 10, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 10, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 10, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-017", "id": "fast-43-diverse-077-017-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2024.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 10, 2025."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate the same governing policy and entity/course/date bindings; the two focus evidence sentences are plain factual statements, and the counterfactual only changes the placement report date to 2022, remaining logically consistent without contradicting other evidence or leaking any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2024.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\",\"Mina Cho's requested CHEM 220 term begins on January 10, 2025.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2024."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 10, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 10, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 10, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-017", "id": "fast-43-diverse-077-017-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 10, 2025."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy text verbatim and the question object is unchanged, preserving all scope rules; the counterfactual only alters the placement-report year, shifting it outside the 12-month window without contradicting other evidence; the focus evidence are two complete factual sentences reporting dates, not policy or instructions; no evidence states or implies the final yes/no recommendation or routing outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2023.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\",\"Mina Cho's requested CHEM 220 term begins January 8, 2024.\"],\"request\":\"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-019", "id": "fast-43-diverse-077-019-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2023.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins January 8, 2024."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy text verbatim and the question object is unchanged, preserving all scope rules; the counterfactual only alters the placement-report year, shifting it outside the 12-month window without contradicting other evidence; the focus evidence are two complete factual sentences reporting dates, not policy or instructions; no evidence states or implies the final yes/no recommendation or routing outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2023.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\",\"Mina Cho's requested CHEM 220 term begins January 8, 2024.\"],\"request\":\"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-019", "id": "fast-43-diverse-077-019-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins January 8, 2024."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the governing prerequisite/scope rules and only vary the placement-report date (2023 vs 2022), consistent with the unchanged question and request; all evidence items are factual, non-contradictory, and contain no embedded labels or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2023.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Mina Cho's requested CHEM 220 term begins on January 10, 2024.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "4"], "text": "Mina Cho's requested CHEM 220 term begins on January 10, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 10, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 10, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-022", "id": "fast-43-diverse-077-022-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2023.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Mina Cho's requested CHEM 220 term begins on January 10, 2024.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the governing prerequisite/scope rules and only vary the placement-report date (2023 vs 2022), consistent with the unchanged question and request; all evidence items are factual, non-contradictory, and contain no embedded labels or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 1, 2023.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Mina Cho's requested CHEM 220 term begins on January 10, 2024.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2023."}, {"path": ["evidence", "4"], "text": "Mina Cho's requested CHEM 220 term begins on January 10, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 10, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 10, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-022", "id": "fast-43-diverse-077-022-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Mina Cho's requested CHEM 220 term begins on January 10, 2024.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the coordinator's scope limit and 12-month rule, keep the same entity/course/time bindings, use plain factual date/score sentences without embedded rules or answers, and the counterfactual only swaps the placement date to create an internally consistent expired-score scenario.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2024.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2024."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-023", "id": "fast-43-diverse-077-023-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2024.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the coordinator's scope limit and 12-month rule, keep the same entity/course/time bindings, use plain factual date/score sentences without embedded rules or answers, and the counterfactual only swaps the placement date to create an internally consistent expired-score scenario.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Her official placement report shows a score of 82.\", \"Mina Cho's official placement report is dated March 1, 2024.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 1, 2024."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 1, 2024.", "negative_left": "Mina Cho's official placement report is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-023", "id": "fast-43-diverse-077-023-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 1, 2022.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate all governing prerequisite, score, timing, scope and routing rules; the two focus evidence spans are complete factual sentences about the placement report date and term start, entity/path/time bindings for Mina Cho and CHEM 220 are unchanged, the counterfactual only shifts the placement report date by one year without contradicting other facts, and no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023.\", \"Her official placement report shows a score of 82.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "1"], "text": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-024", "id": "fast-43-diverse-077-024-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023.", "Her official placement report shows a score of 82.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate all governing prerequisite, score, timing, scope and routing rules; the two focus evidence spans are complete factual sentences about the placement report date and term start, entity/path/time bindings for Mina Cho and CHEM 220 are unchanged, the counterfactual only shifts the placement report date by one year without contradicting other facts, and no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\": \"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\", \"evidence\": [\"The official transcript shows Mina earned D+ in CHEM 110.\", \"Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023.\", \"Her official placement report shows a score of 82.\", \"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\", \"Her academic adviser forwarded all three required documents to the course coordinator.\", \"Mina Cho's requested CHEM 220 term begins on January 8, 2024.\"], \"request\": \"Rate Mina\\u2019s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "1"], "text": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023."}, {"path": ["evidence", "5"], "text": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2023.", "negative_left": "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024.", "right": "Mina Cho's requested CHEM 220 term begins on January 8, 2024."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-024", "id": "fast-43-diverse-077-024-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Mina Cho's official placement report in her CHEM 220 exception file is dated March 1, 2022.", "Her official placement report shows a score of 82.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator.", "Mina Cho's requested CHEM 220 term begins on January 8, 2024."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the governing policy (12‑month score window, coordinator scope limit, panel routing) unchanged and identical to the original question; entity, course, and term bindings are unaltered; the two focus evidence items are plain factual sentences about the report date and term start; the counterfactual only swaps the placement‑report date to 2022, making a single coherent factual change without contradicting other evidence; no context reveals a gold answer, rule table, or explicit instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 15, 2024.\",\"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 15, 2024."}, {"path": ["evidence", "3"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 15, 2024.", "negative_left": "Mina Cho's official placement report is dated March 15, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-025", "id": "fast-43-diverse-077-025-base", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 15, 2024.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "education-03", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the governing policy (12‑month score window, coordinator scope limit, panel routing) unchanged and identical to the original question; entity, course, and term bindings are unaltered; the two focus evidence items are plain factual sentences about the report date and term start; the counterfactual only swaps the placement‑report date to 2022, making a single coherent factual change without contradicting other evidence; no context reveals a gold answer, rule table, or explicit instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "full_context_fact_states": {"base": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "supported", "a_transcript": "supported"}, "counterfactual": {"a_instructor_note": "supported", "a_placement_report": "supported", "a_placement_score": "supported", "a_prerequisite_grade": "refuted", "a_report_date": "supported", "a_score_age": "refuted", "a_transcript": "supported"}, "remove_left": {"a_score_age": "unknown"}, "remove_right": {"a_score_age": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_score_age": "unknown"}, "negative_pair": {"a_score_age": "refuted"}, "negative_sentence": {"a_score_age": "unknown"}, "positive_pair": {"a_score_age": "supported"}, "right": {"a_score_age": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the focus atom is the factual age of the placement score rather than a policy conclusion. The base and counter assignments can both occur under the original policy while differing only in whether the report is within 12 months. The policy evidence correctly preserves all needed substantive rules originating in the original state: the grade prerequisite, placement-score threshold and timing rule, coordinator scope and panel routing, and file-completeness requirements. Both rules include the relevant competing-outcome exclusions and are sufficient for their targets.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a qualifying placement exception satisfying both the score and 12-month timing requirements. This is sufficient for readiness level 2, with coordinator approval and no additional review.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a complete file, failure of the standard grade prerequisite, and a placement score outside the nonwaivable 12-month limit. The stated policy requires level 0, a no recommendation at this stage, and routing to the Registrar Exceptions Panel.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transcript", "statement": "Mina Cho's CHEM 220 exception file contains an official transcript."}, {"id": "a_placement_report", "statement": "Mina Cho's CHEM 220 exception file contains an official placement report."}, {"id": "a_report_date", "statement": "The official placement report in Mina Cho's CHEM 220 exception file states its date."}, {"id": "a_instructor_note", "statement": "Mina Cho's CHEM 220 exception file contains a CHEM 220 instructor note."}, {"id": "a_prerequisite_grade", "statement": "Mina Cho's grade in CHEM 110 is at least C."}, {"id": "a_placement_score", "statement": "The score stated in Mina Cho's official placement report is at least 75."}, {"id": "a_score_age", "statement": "The interval from the date stated on Mina Cho's official placement report to the start of Mina Cho's requested CHEM 220 term is no more than 12 months."}], "base_state_json": "{\"context\":\"Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.\",\"evidence\":[\"The official transcript shows Mina earned D+ in CHEM 110.\",\"Her official placement report shows a score of 82.\",\"Mina Cho's official placement report is dated March 15, 2024.\",\"Mina Cho's requested CHEM 220 term begins on January 6, 2025.\",\"The CHEM 220 instructor notes strong laboratory skills and supports enrollment.\",\"Her academic adviser forwarded all three required documents to the course coordinator.\"],\"request\":\"Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing.\"}", "base_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}], "counter_states": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}], "focus_atom": "a_score_age", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mina Cho's official placement report is dated March 15, 2024."}, {"path": ["evidence", "3"], "text": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}], "policy_evidence": [{"path": ["context"], "text": "The normal prerequisite is CHEM 110 with at least C."}, {"path": ["context"], "text": "A placement score of 75 or higher substitutes only when dated within 12 months."}, {"path": ["context"], "text": "The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel."}, {"path": ["context"], "text": "A complete file requires an official transcript, dated placement report, and instructor note."}], "rules": [{"justification": "The file is complete, and although Mina does not have the required CHEM 110 grade, the official placement score satisfies both the score threshold and the 12-month timing limit. The exception is therefore within the coordinator's permitted scope.", "target": "2", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "supported"}]}, {"justification": "The file is complete, but Mina does not have the required CHEM 110 grade and the otherwise qualifying placement score is older than the permitted 12-month limit. The coordinator cannot waive that limit, so the case requires a no recommendation and routing to the Registrar Exceptions Panel.", "target": "0", "when": [{"atom_id": "a_transcript", "state": "supported"}, {"atom_id": "a_placement_report", "state": "supported"}, {"atom_id": "a_report_date", "state": "supported"}, {"atom_id": "a_instructor_note", "state": "supported"}, {"atom_id": "a_prerequisite_grade", "state": "refuted"}, {"atom_id": "a_placement_score", "state": "supported"}, {"atom_id": "a_score_age", "state": "refuted"}]}]}, "verified_pair": {"left": "Mina Cho's official placement report is dated March 15, 2024.", "negative_left": "Mina Cho's official placement report is dated March 15, 2022.", "negative_right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "right": "Mina Cho's requested CHEM 220 term begins on January 6, 2025."}, "verifier_independent_model": false}, "family": "fast-43-diverse-077-025", "id": "fast-43-diverse-077-025-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready for coordinator approval: documentation may be complete, but the student fails the standard prerequisite and the claimed exception is outside the coordinator’s permitted scope. Recommend no at this stage and route the case to the Registrar Exceptions Panel.", "Provisionally ready: required documentation is missing or authenticity/date details are unresolved, so no merits recommendation can yet be made. Route the file to the academic adviser or registrar staff member to obtain or verify the missing material.", "Ready for coordinator approval: the complete file shows either the required prerequisite grade or an exception that fully satisfies all stated score, timing, and scope rules. Recommend yes; no additional review is needed."], "instructions": "Select exactly one ordered readiness level. Apply the stated scope limits even when other evidence supports the student. The selected level determines the recommendation and routing.", "type": "score"}}, "state": {"context": "Student Mina Cho requests a prerequisite exception for CHEM 220. The normal prerequisite is CHEM 110 with at least C. A placement score of 75 or higher substitutes only when dated within 12 months. The course coordinator cannot waive that limit; older-score cases go to the Registrar Exceptions Panel. A complete file requires an official transcript, dated placement report, and instructor note.", "evidence": ["The official transcript shows Mina earned D+ in CHEM 110.", "Her official placement report shows a score of 82.", "Mina Cho's official placement report is dated March 15, 2022.", "Mina Cho's requested CHEM 220 term begins on January 6, 2025.", "The CHEM 220 instructor notes strong laboratory skills and supports enrollment.", "Her academic adviser forwarded all three required documents to the course coordinator."], "request": "Rate Mina’s readiness for approval, give the associated yes/no recommendation, and identify any required routing."}}, "method": "c2d", "provenance": {"source_id": "diverse-077", "source_is_synthetic": true, "source_sha256": "b808e8d78a5fd947d8cbb5393a527a1e5983149489c661a66ba923259c103c16", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "education-03", "split": "train", "variant": "counterfactual"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve validity, no-show, and ranking policies; the two focus spans are complete factual seat-count sentences; the counterfactual only changes Maple's total seats (12) and unused count (6), which is internally consistent arithmetic and alters the ranking outcome without contradicting other unchanged facts; entity/path/time bindings (Leo, Maple, Birch, 2–3 slot) are unchanged in both; no answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Student group organizer\",\"text\":\"I'm Priya. My six-person team, including Dana, who uses a wheelchair, reserved Cedar from 2–3. Leo says his five-person team also has Cedar then, needs a display, and has no accessibility need.\"},{\"speaker\":\"Library booking assistant\",\"text\":\"Ledger: Priya requested Cedar at 9:01 and her request is valid; every other valid request overlapping the 2–3 slot, including Leo's 9:04 request, was submitted later. Leo's request overlaps Priya's booking, but Leo has two no-shows in the past 30 days and no branch supervisor has yet approved his request, so it stays invalid pending review. Leo's group has no accessibility need and requested a display for the 2–3 reassignment. Maple's capacity and Birch's capacity both meet Leo's group size, and both rooms have a display available for the reassignment.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"For Leo's group of 6 at the 2–3 slot, Maple has 9 seats total, leaving an unused-seat count of 3.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "For Leo's group of 6 at the 2–3 slot, Maple has 9 seats total, leaving an unused-seat count of 3."}, {"path": ["3", "text"], "text": "For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "For Leo's group of 6 at the 2–3 slot, Maple has 9 seats total, leaving an unused-seat count of 3.", "negative_left": "For Leo's group of 6 at the 2–3 slot, Maple has 12 seats total, leaving an unused-seat count of 6.", "negative_right": "For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5.", "right": "For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-088-013", "id": "fast-43-diverse-088-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Student group organizer", "text": "I'm Priya. My six-person team, including Dana, who uses a wheelchair, reserved Cedar from 2–3. Leo says his five-person team also has Cedar then, needs a display, and has no accessibility need."}, {"speaker": "Library booking assistant", "text": "Ledger: Priya requested Cedar at 9:01 and her request is valid; every other valid request overlapping the 2–3 slot, including Leo's 9:04 request, was submitted later. Leo's request overlaps Priya's booking, but Leo has two no-shows in the past 30 days and no branch supervisor has yet approved his request, so it stays invalid pending review. Leo's group has no accessibility need and requested a display for the 2–3 reassignment. Maple's capacity and Birch's capacity both meet Leo's group size, and both rooms have a display available for the reassignment."}, {"speaker": "Branch supervisor", "text": "For Leo's group of 6 at the 2–3 slot, Maple has 9 seats total, leaving an unused-seat count of 3."}, {"speaker": "Branch supervisor", "text": "For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5."}, {"speaker": "Branch supervisor", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Branch supervisor", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "education-05", "split": "train", "variant": "base"} {"domain": "education", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve validity, no-show, and ranking policies; the two focus spans are complete factual seat-count sentences; the counterfactual only changes Maple's total seats (12) and unused count (6), which is internally consistent arithmetic and alters the ranking outcome without contradicting other unchanged facts; entity/path/time bindings (Leo, Maple, Birch, 2–3 slot) are unchanged in both; no answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "refuted", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a12": "unknown"}, "remove_right": {"a12": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a12": "unknown"}, "negative_pair": {"a12": "refuted"}, "negative_sentence": {"a12": "unknown"}, "positive_pair": {"a12": "supported"}, "right": {"a12": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified a2 remains one ordering relationship over the explicit class of competing requests. The focus a12 is a factual comparison, not a policy conclusion. The base and counter assignments are jointly realizable in separate synthetic scenarios while changing only the focus atom’s status: unequal integer unused-seat counts can have Maple lower in the base and Birch lower in the counter, with all other listed propositions remaining supported. Policy evidence correctly preserves the substantive rules originating in the state, including capacity equality, earliest-valid priority, the two-no-show approval requirement, room accessibility scope, and alternative-ranking priorities. Instructions and criteria from the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient. Priya is valid and earlier than every other valid overlapping Cedar request, so she wins under the earliest-valid-request rule. Leo’s exactly two no-shows and absence of supervisor approval make his request currently invalid and require supervisor handling. Both rooms satisfy Leo’s relevant capacity, accessibility, and display conditions, and Maple’s lower unused-seat count places it ahead of Birch.", "rule_index": 0, "sound": true}, {"reason": "The conjunction preserves Priya’s win and Leo’s current invalidity. Refutation of Maple having fewer unused seats, combined with supported inequality of the two counts, entails that Birch has fewer unused seats. Because both alternatives are otherwise eligible, the least-unused-capacity priority ranks Birch ahead of Maple, so the proposed Maple-first resolution is incorrect.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Priya’s request for Cedar for the described 2–3 interval is valid at the current decision time."}, {"id": "a2", "statement": "Every other valid request for Cedar whose interval overlaps Priya’s described 2–3 interval was submitted later than Priya’s request."}, {"id": "a3", "statement": "Leo’s request for Cedar overlaps Priya’s request for Cedar during the described 2–3 interval."}, {"id": "a4", "statement": "Leo has exactly two no-shows during the 30 days preceding the current booking-validity decision."}, {"id": "a5", "statement": "No branch supervisor approved Leo’s Cedar request before the current booking-validity decision."}, {"id": "a6", "statement": "Leo’s group has no accessibility need for the described 2–3 reassignment."}, {"id": "a7", "statement": "Maple’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a8", "statement": "Birch’s capacity is at least Leo’s group size for the described 2–3 reassignment."}, {"id": "a9", "statement": "Leo requested a display for the described 2–3 booking."}, {"id": "a10", "statement": "Maple has a display available for Leo’s described 2–3 reassignment."}, {"id": "a11", "statement": "Birch has a display available for Leo’s described 2–3 reassignment."}, {"id": "a12", "statement": "Maple’s unused-seat count for Leo’s group is lower than Birch’s unused-seat count for Leo’s group."}, {"id": "a13", "statement": "Maple’s and Birch’s unused-seat counts for Leo’s group are unequal."}], "base_state_json": "[{\"speaker\":\"Student group organizer\",\"text\":\"I'm Priya. My six-person team, including Dana, who uses a wheelchair, reserved Cedar from 2–3. Leo says his five-person team also has Cedar then, needs a display, and has no accessibility need.\"},{\"speaker\":\"Library booking assistant\",\"text\":\"Ledger: Priya requested Cedar at 9:01 and her request is valid; every other valid request overlapping the 2–3 slot, including Leo's 9:04 request, was submitted later. Leo's request overlaps Priya's booking, but Leo has two no-shows in the past 30 days and no branch supervisor has yet approved his request, so it stays invalid pending review. Leo's group has no accessibility need and requested a display for the 2–3 reassignment. Maple's capacity and Birch's capacity both meet Leo's group size, and both rooms have a display available for the reassignment.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"For Leo's group of 6 at the 2–3 slot, Maple has 9 seats total, leaving an unused-seat count of 3.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it.\"},{\"speaker\":\"Branch supervisor\",\"text\":\"Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a12", "focus_evidence": [{"path": ["2", "text"], "text": "For Leo's group of 6 at the 2–3 slot, Maple has 9 seats total, leaving an unused-seat count of 3."}, {"path": ["3", "text"], "text": "For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5."}], "policy_evidence": [{"path": ["1", "text"], "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"path": ["2", "text"], "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}], "rules": [{"justification": "Priya is valid and earlier than every competing valid overlapping Cedar request. Leo’s two no-shows and lack of supervisor approval make his request currently invalid. Both alternatives meet Leo’s stated capacity, accessibility, and display constraints, and Maple has strictly less unused capacity, so Maple ranks ahead of Birch.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "Priya’s award and Leo’s invalid routing remain correct. Because the unused-seat counts are unequal and Maple’s count is not lower, Birch has strictly less unused capacity; the policy therefore ranks Birch ahead of Maple, making the proposed alternative-room ranking incorrect.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "refuted"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "For Leo's group of 6 at the 2–3 slot, Maple has 9 seats total, leaving an unused-seat count of 3.", "negative_left": "For Leo's group of 6 at the 2–3 slot, Maple has 12 seats total, leaving an unused-seat count of 6.", "negative_right": "For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5.", "right": "For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-088-013", "id": "fast-43-diverse-088-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one booking-validity decision, staff routing choice, or alternative-room ranking is incorrect.", "true": "Yes — Priya is correctly awarded Cedar, Leo is correctly routed to the branch supervisor as currently invalid, and Maple is correctly ranked ahead of Birch."}, "instructions": "Verification task (diverse-088): Is this complete resolution correct? Priya’s Cedar booking is valid and wins; Leo’s booking is currently invalid and must go to the branch supervisor; for Leo, rank Maple first and Birch second.", "type": "noul"}}, "state": [{"speaker": "Student group organizer", "text": "I'm Priya. My six-person team, including Dana, who uses a wheelchair, reserved Cedar from 2–3. Leo says his five-person team also has Cedar then, needs a display, and has no accessibility need."}, {"speaker": "Library booking assistant", "text": "Ledger: Priya requested Cedar at 9:01 and her request is valid; every other valid request overlapping the 2–3 slot, including Leo's 9:04 request, was submitted later. Leo's request overlaps Priya's booking, but Leo has two no-shows in the past 30 days and no branch supervisor has yet approved his request, so it stays invalid pending review. Leo's group has no accessibility need and requested a display for the 2–3 reassignment. Maple's capacity and Birch's capacity both meet Leo's group size, and both rooms have a display available for the reassignment."}, {"speaker": "Branch supervisor", "text": "For Leo's group of 6 at the 2–3 slot, Maple has 12 seats total, leaving an unused-seat count of 6."}, {"speaker": "Branch supervisor", "text": "For Leo's group of 6 at the 2–3 slot, Birch has 11 seats total, leaving an unused-seat count of 5."}, {"speaker": "Branch supervisor", "text": "Policy: capacity may equal group size; earliest valid overlapping request wins. Two no-shows make a request invalid until a branch supervisor approves it."}, {"speaker": "Branch supervisor", "text": "Groups without accessibility needs may use either. Rank eligible alternatives by least unused capacity, then requested equipment."}]}, "method": "c2d", "provenance": {"source_id": "diverse-088", "source_is_synthetic": true, "source_sha256": "b62711b0afda9d2271e2261eec447731183ad3758f39432cdcefd57e837566e5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "education-05", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain policy_1 and policy_2 verbatim and preserve the same A/B/C entities, times, and request; the two focus evidence spans are complete factual sentences about the log linkage and its stair content, not policy text; the counterfactual only flips the stair fact for entry TC-19 (an allowed case-note change) without contradicting any other retained assertion, and no answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\": \"Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station's accessibility desk maintains a cross-referenced log of route entries for travelers with mobility needs.\", \"policy_1\": \"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\", \"policy_2\": \"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\", \"request\": \"Rank A, B, and C from least to most suitable under the supplied rule.\", \"timing_notes\": \"Option A arrives at 16:10, option B at 16:30, and option C at 17:20.\", \"log_entry\": \"The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip.\", \"log_detail\": \"Entry TC-19 states that the route it covers can be completed without using stairs.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["log_entry"], "text": "The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip."}, {"path": ["log_detail"], "text": "Entry TC-19 states that the route it covers can be completed without using stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip.", "negative_left": "The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip.", "negative_right": "Entry TC-19 states that the route it covers requires using stairs at one station.", "right": "Entry TC-19 states that the route it covers can be completed without using stairs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-092-044", "id": "fast-43-diverse-092-044-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station's accessibility desk maintains a cross-referenced log of route entries for travelers with mobility needs.", "log_detail": "Entry TC-19 states that the route it covers can be completed without using stairs.", "log_entry": "The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip.", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule.", "timing_notes": "Option A arrives at 16:10, option B at 16:30, and option C at 17:20."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain policy_1 and policy_2 verbatim and preserve the same A/B/C entities, times, and request; the two focus evidence spans are complete factual sentences about the log linkage and its stair content, not policy text; the counterfactual only flips the stair fact for entry TC-19 (an allowed case-note change) without contradicting any other retained assertion, and no answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\": \"Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station's accessibility desk maintains a cross-referenced log of route entries for travelers with mobility needs.\", \"policy_1\": \"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\", \"policy_2\": \"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\", \"request\": \"Rank A, B, and C from least to most suitable under the supplied rule.\", \"timing_notes\": \"Option A arrives at 16:10, option B at 16:30, and option C at 17:20.\", \"log_entry\": \"The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip.\", \"log_detail\": \"Entry TC-19 states that the route it covers can be completed without using stairs.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["log_entry"], "text": "The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip."}, {"path": ["log_detail"], "text": "Entry TC-19 states that the route it covers can be completed without using stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip.", "negative_left": "The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip.", "negative_right": "Entry TC-19 states that the route it covers requires using stairs at one station.", "right": "Entry TC-19 states that the route it covers can be completed without using stairs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-092-044", "id": "fast-43-diverse-092-044-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station's accessibility desk maintains a cross-referenced log of route entries for travelers with mobility needs.", "log_detail": "Entry TC-19 states that the route it covers requires using stairs at one station.", "log_entry": "The transit accessibility log links option C's itinerary to entry TC-19 for Mara's trip.", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule.", "timing_notes": "Option A arrives at 16:10, option B at 16:30, and option C at 17:20."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original ranking and connection-exception policies exactly, keep the same entities A/B/C, times, and request wording; the two focus evidence spans are complete factual log statements, not policy or instructions; the counterfactual only flips log_entry_2's stair claim for AN-19, which is a single coherent factual change with no contradiction elsewhere, and no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\": \"Mara is seeking rebooking after her 14:10 Northport–Lake City train was canceled. Her flexible ticket allows rebooking onto options A, B, or C. Mara cannot use stairs, and option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection, on the same platform, lasting at least 7 minutes. Sela, a train operations coordinator, confirmed that connection would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. All three options remain available to Mara for rebooking. The station accessibility log contains the following entries:\", \"policy_1\": \"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\", \"policy_2\": \"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\", \"request\": \"Rank A, B, and C from least to most suitable under the supplied rule.\", \"log_entry_1\": \"The travel accessibility log lists itinerary C for Mara under access-note reference AN-19.\", \"log_entry_2\": \"Access-note reference AN-19 states that the traveler assigned to it can complete the journey without using stairs.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["log_entry_1"], "text": "The travel accessibility log lists itinerary C for Mara under access-note reference AN-19."}, {"path": ["log_entry_2"], "text": "Access-note reference AN-19 states that the traveler assigned to it can complete the journey without using stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The travel accessibility log lists itinerary C for Mara under access-note reference AN-19.", "negative_left": "The travel accessibility log lists itinerary C for Mara under access-note reference AN-19.", "negative_right": "Access-note reference AN-19 states that the traveler assigned to it must use stairs to complete the journey.", "right": "Access-note reference AN-19 states that the traveler assigned to it can complete the journey without using stairs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-092-045", "id": "fast-43-diverse-092-045-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking rebooking after her 14:10 Northport–Lake City train was canceled. Her flexible ticket allows rebooking onto options A, B, or C. Mara cannot use stairs, and option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection, on the same platform, lasting at least 7 minutes. Sela, a train operations coordinator, confirmed that connection would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. All three options remain available to Mara for rebooking. The station accessibility log contains the following entries:", "log_entry_1": "The travel accessibility log lists itinerary C for Mara under access-note reference AN-19.", "log_entry_2": "Access-note reference AN-19 states that the traveler assigned to it can complete the journey without using stairs.", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original ranking and connection-exception policies exactly, keep the same entities A/B/C, times, and request wording; the two focus evidence spans are complete factual log statements, not policy or instructions; the counterfactual only flips log_entry_2's stair claim for AN-19, which is a single coherent factual change with no contradiction elsewhere, and no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\": \"Mara is seeking rebooking after her 14:10 Northport–Lake City train was canceled. Her flexible ticket allows rebooking onto options A, B, or C. Mara cannot use stairs, and option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection, on the same platform, lasting at least 7 minutes. Sela, a train operations coordinator, confirmed that connection would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. All three options remain available to Mara for rebooking. The station accessibility log contains the following entries:\", \"policy_1\": \"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\", \"policy_2\": \"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\", \"request\": \"Rank A, B, and C from least to most suitable under the supplied rule.\", \"log_entry_1\": \"The travel accessibility log lists itinerary C for Mara under access-note reference AN-19.\", \"log_entry_2\": \"Access-note reference AN-19 states that the traveler assigned to it can complete the journey without using stairs.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["log_entry_1"], "text": "The travel accessibility log lists itinerary C for Mara under access-note reference AN-19."}, {"path": ["log_entry_2"], "text": "Access-note reference AN-19 states that the traveler assigned to it can complete the journey without using stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The travel accessibility log lists itinerary C for Mara under access-note reference AN-19.", "negative_left": "The travel accessibility log lists itinerary C for Mara under access-note reference AN-19.", "negative_right": "Access-note reference AN-19 states that the traveler assigned to it must use stairs to complete the journey.", "right": "Access-note reference AN-19 states that the traveler assigned to it can complete the journey without using stairs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-092-045", "id": "fast-43-diverse-092-045-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking rebooking after her 14:10 Northport–Lake City train was canceled. Her flexible ticket allows rebooking onto options A, B, or C. Mara cannot use stairs, and option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection, on the same platform, lasting at least 7 minutes. Sela, a train operations coordinator, confirmed that connection would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. All three options remain available to Mara for rebooking. The station accessibility log contains the following entries:", "log_entry_1": "The travel accessibility log lists itinerary C for Mara under access-note reference AN-19.", "log_entry_2": "Access-note reference AN-19 states that the traveler assigned to it must use stairs to complete the journey.", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain policy_1 and policy_2 verbatim, keep the same options, times, and constraints, and the two evidence sentences are complete factual statements about the accessibility log and audit entry rather than policy or answer text; the counterfactual coherently flips the mobility database finding from step-free to stair-required without contradicting any other unchanged fact (e.g., 'Option C is direct' does not assert step-free status), and neither context reveals a gold label, code table, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\": \"Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station accessibility desk logged the following note during intake: 'The transit accessibility log assigns option C to itinerary code TC-92.' A follow-up audit entry, cross-checked against the mobility database, adds: 'Itinerary code TC-92 is documented as a step-free route requiring no stair use.' All three options remain available to Mara for rebooking under her flexible ticket.\", \"policy_1\": \"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\", \"policy_2\": \"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\", \"request\": \"Rank A, B, and C from least to most suitable under the supplied rule.\", \"timing_notes\": \"Option A arrives at 16:10, option B at 16:30, and option C at 17:20.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["case_note"], "text": "The transit accessibility log assigns option C to itinerary code TC-92."}, {"path": ["case_note"], "text": "Itinerary code TC-92 is documented as a step-free route requiring no stair use."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The transit accessibility log assigns option C to itinerary code TC-92.", "negative_left": "The transit accessibility log assigns option C to itinerary code TC-92.", "negative_right": "Itinerary code TC-92 is documented as requiring stair use partway through the route.", "right": "Itinerary code TC-92 is documented as a step-free route requiring no stair use."}, "verifier_independent_model": false}, "family": "fast-43-diverse-092-046", "id": "fast-43-diverse-092-046-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station accessibility desk logged the following note during intake: 'The transit accessibility log assigns option C to itinerary code TC-92.' A follow-up audit entry, cross-checked against the mobility database, adds: 'Itinerary code TC-92 is documented as a step-free route requiring no stair use.' All three options remain available to Mara for rebooking under her flexible ticket.", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule.", "timing_notes": "Option A arrives at 16:10, option B at 16:30, and option C at 17:20."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain policy_1 and policy_2 verbatim, keep the same options, times, and constraints, and the two evidence sentences are complete factual statements about the accessibility log and audit entry rather than policy or answer text; the counterfactual coherently flips the mobility database finding from step-free to stair-required without contradicting any other unchanged fact (e.g., 'Option C is direct' does not assert step-free status), and neither context reveals a gold label, code table, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\": \"Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station accessibility desk logged the following note during intake: 'The transit accessibility log assigns option C to itinerary code TC-92.' A follow-up audit entry, cross-checked against the mobility database, adds: 'Itinerary code TC-92 is documented as a step-free route requiring no stair use.' All three options remain available to Mara for rebooking under her flexible ticket.\", \"policy_1\": \"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\", \"policy_2\": \"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\", \"request\": \"Rank A, B, and C from least to most suitable under the supplied rule.\", \"timing_notes\": \"Option A arrives at 16:10, option B at 16:30, and option C at 17:20.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["case_note"], "text": "The transit accessibility log assigns option C to itinerary code TC-92."}, {"path": ["case_note"], "text": "Itinerary code TC-92 is documented as a step-free route requiring no stair use."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The transit accessibility log assigns option C to itinerary code TC-92.", "negative_left": "The transit accessibility log assigns option C to itinerary code TC-92.", "negative_right": "Itinerary code TC-92 is documented as requiring stair use partway through the route.", "right": "Itinerary code TC-92 is documented as a step-free route requiring no stair use."}, "verifier_independent_model": false}, "family": "fast-43-diverse-092-046", "id": "fast-43-diverse-092-046-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking a rebooking after her Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection. That connection is on the same platform, lasts at least 7 minutes, and Sela, a train operations coordinator, confirmed it would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The station accessibility desk logged the following note during intake: 'The transit accessibility log assigns option C to itinerary code TC-92.' A follow-up audit entry, cross-checked against the mobility database, adds: 'Itinerary code TC-92 is documented as requiring stair use partway through the route.' All three options remain available to Mara for rebooking under her flexible ticket.", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule.", "timing_notes": "Option A arrives at 16:10, option B at 16:30, and option C at 17:20."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain policy_1 and policy_2 verbatim and the original request/question object, preserving all governing rules and bindings; the two focus evidence sentences are complete factual statements about BR-19, not policy or instructions; the counterfactual only flips the BR-19 accessibility note (step-free to stair-requiring) without contradicting any other stated fact, remaining internally coherent; neither context contains a gold answer, rule table, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\":\"Mara is seeking a rebooking after her 14:10 Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection, on the same platform, lasting at least 7 minutes, which Sela, a train operations coordinator, confirmed would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19. Boarding record BR-19 notes that its assigned itinerary lets the traveler board and disembark without using stairs.\",\"policy_1\":\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"policy_2\":\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["case_note"], "text": "The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19."}, {"path": ["case_note"], "text": "Boarding record BR-19 notes that its assigned itinerary lets the traveler board and disembark without using stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19.", "negative_left": "The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19.", "negative_right": "Boarding record BR-19 notes that its assigned itinerary requires the traveler to climb stairs at the transfer platform.", "right": "Boarding record BR-19 notes that its assigned itinerary lets the traveler board and disembark without using stairs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-092-048", "id": "fast-43-diverse-092-048-base", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking a rebooking after her 14:10 Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection, on the same platform, lasting at least 7 minutes, which Sela, a train operations coordinator, confirmed would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19. Boarding record BR-19 notes that its assigned itinerary lets the traveler board and disembark without using stairs.", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "A_C_B"}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain policy_1 and policy_2 verbatim and the original request/question object, preserving all governing rules and bindings; the two focus evidence sentences are complete factual statements about BR-19, not policy or instructions; the counterfactual only flips the BR-19 accessibility note (step-free to stair-requiring) without contradicting any other stated fact, remaining internally coherent; neither context contains a gold answer, rule table, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships; the explicit-set statement in a13 is permitted. The focus atom a4 is a factual accessibility relation, not a policy classification. Base and counter assignments can differ only in whether direct option C permits stair-free travel: a direct itinerary can still require stairs at a station, so both assignments are realizable. Policy evidence correctly cites the ranking rule and connection exception from the original state; the unchanged questions object automatically preserves its own instructions and criteria. The extra citation to the state request is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction is sufficient for A_C_B. Mara cannot use stairs and A requires stairs, so A is infeasible. B permits stair-free travel and its sole connection satisfies every element of the 7-minute held, same-platform coordinator exception, so B is feasible. C permits stair-free travel and is direct, so C is feasible. Infeasibility places A below B and C, while B's earlier arrival places B above C.", "rule_index": 0, "sound": true}, {"reason": "The conjunction is sufficient for C_A_B. Refutation of C permitting travel without stairs, together with Mara's inability to use stairs, makes C infeasible; A is likewise infeasible because it requires stairs. B is feasible because it permits stair-free travel and its sole connection satisfies the stated exception. B therefore outranks both, and A's earlier arrival than C ranks A above C among the infeasible options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara cannot use stairs."}, {"id": "a2", "statement": "Option A requires Mara to use stairs."}, {"id": "a3", "statement": "Option B permits Mara to travel without using stairs."}, {"id": "a4", "statement": "The itinerary assigned to option C permits Mara to travel without using stairs."}, {"id": "a5", "statement": "Option B has exactly one connection."}, {"id": "a6", "statement": "Option B's connection is on the same platform."}, {"id": "a7", "statement": "Sela confirmed that option B's connection would be held."}, {"id": "a8", "statement": "Sela is a train operations coordinator."}, {"id": "a9", "statement": "Option B's connection lasts at least 7 minutes."}, {"id": "a10", "statement": "Option C is direct."}, {"id": "a11", "statement": "Option B arrives earlier than option C."}, {"id": "a12", "statement": "Option A arrives earlier than option C."}, {"id": "a13", "statement": "Each of options A, B, and C is available to Mara for rebooking."}], "base_state_json": "{\"case_note\":\"Mara is seeking a rebooking after her 14:10 Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection, on the same platform, lasting at least 7 minutes, which Sela, a train operations coordinator, confirmed would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19. Boarding record BR-19 notes that its assigned itinerary lets the traveler board and disembark without using stairs.\",\"policy_1\":\"Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.\",\"policy_2\":\"Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.\",\"request\":\"Rank A, B, and C from least to most suitable under the supplied rule.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["case_note"], "text": "The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19."}, {"path": ["case_note"], "text": "Boarding record BR-19 notes that its assigned itinerary lets the traveler board and disembark without using stairs."}], "policy_evidence": [{"path": ["request"], "text": "Rank A, B, and C from least to most suitable under the supplied rule."}, {"path": ["evidence", "0"], "text": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers."}, {"path": ["evidence", "1"], "text": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes."}], "rules": [{"justification": "A violates Mara's stair constraint and is infeasible. B satisfies that constraint and its sole connection satisfies every part of the coordinator-confirmed held, same-platform, at-least-7-minute exception. Direct option C also satisfies the stair constraint. Thus A ranks below both feasible routes, and B ranks above C because B arrives earlier, giving A_C_B.", "target": "A_C_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "A and C both violate Mara's stair constraint and are infeasible. B satisfies the constraint and its sole connection qualifies for the stated exception, so B ranks above both. Among the two infeasible routes, A arrives earlier than C and therefore ranks above C, giving C_A_B.", "target": "C_A_B", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19.", "negative_left": "The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19.", "negative_right": "Boarding record BR-19 notes that its assigned itinerary requires the traveler to climb stairs at the transfer platform.", "right": "Boarding record BR-19 notes that its assigned itinerary lets the traveler board and disembark without using stairs."}, "verifier_independent_model": false}, "family": "fast-43-diverse-092-048", "id": "fast-43-diverse-092-048-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_B_C": "A is least suitable, B is intermediate, and C is most suitable.", "A_C_B": "A is least suitable, C is intermediate, and B is most suitable.", "C_A_B": "C is least suitable, A is intermediate, and B is most suitable."}, "instructions": "Choose the option giving the complete least-to-most-suitable ordering. Apply the traveler's constraint and the stated connection exception before comparing arrival times and transfers.", "type": "choice"}}, "state": {"case_note": "Mara is seeking a rebooking after her 14:10 Northport–Lake City train was canceled. Her flexible ticket permits rebooking onto any of options A, B, or C. Mara cannot use stairs, while option A requires her to use stairs. Option B permits travel without using stairs and has exactly one connection, on the same platform, lasting at least 7 minutes, which Sela, a train operations coordinator, confirmed would be held. Option C is direct. Option B arrives earlier than option C, and option A arrives earlier than option C. The transit accessibility ledger files Mara's option C itinerary under boarding record BR-19. Boarding record BR-19 notes that its assigned itinerary requires the traveler to climb stairs at the transfer platform.", "policy_1": "Ranking rule: infeasible routes rank below feasible ones; then earlier arrival wins, then fewer transfers.", "policy_2": "Scope: normal connections require 10 minutes. Exception: a coordinator-confirmed held, same-platform connection is feasible with at least 7 minutes.", "request": "Rank A, B, and C from least to most suitable under the supplied rule."}}, "method": "c2d", "provenance": {"source_id": "diverse-092", "source_is_synthetic": true, "source_sha256": "1ecad9c9566a9db58cf008394ebf81ef46bb0ca9bbd3a46f13ce07b180d00e3d", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "C_A_B"}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler’s constraints, the R214/C88/Bellford/Seaborne bindings, and the unchanged minimum-connection-time policy sentence; the counterfactual’s single-sentence platform change (5→3) coherently shifts the connection from cross-platform to same-platform without contradicting the retained rule, and the evidence spans are plain factual platform statements with no leaked answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, R214's 16:02 scheduled arrival uses platform 3.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 5.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 scheduled arrival uses platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 scheduled arrival uses platform 3.", "negative_left": "At Bellford, R214's 16:02 scheduled arrival uses platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 3.", "right": "At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-001", "id": "fast-43-diverse-093-001-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "train operations coordinator", "text": "At Bellford, R214's 16:02 scheduled arrival uses platform 3."}, {"speaker": "train operations coordinator", "text": "At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 5."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler’s constraints, the R214/C88/Bellford/Seaborne bindings, and the unchanged minimum-connection-time policy sentence; the counterfactual’s single-sentence platform change (5→3) coherently shifts the connection from cross-platform to same-platform without contradicting the retained rule, and the evidence spans are plain factual platform statements with no leaked answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, R214's 16:02 scheduled arrival uses platform 3.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 5.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 scheduled arrival uses platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 scheduled arrival uses platform 3.", "negative_left": "At Bellford, R214's 16:02 scheduled arrival uses platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 3.", "right": "At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-001", "id": "fast-43-diverse-093-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "train operations coordinator", "text": "At Bellford, R214's 16:02 scheduled arrival uses platform 3."}, {"speaker": "train operations coordinator", "text": "At Bellford, Coastliner C88's 16:09 scheduled departure uses platform 3."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler's constraints and the coordinator's minimum-connection-time rule intact, altering only case-specific platform facts, which is a permitted observation change; the counterfactual changes a single sentence (C88's departure platform from 5 to 3) coherently without contradicting other stated facts; the two focus evidence spans are complete factual log sentences, not policy text; and neither context states or hints at the feasibility outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rebooking log\",\"text\":\"At Bellford, R214's 16:02 arrival uses platform 3.\"},{\"speaker\":\"rebooking log\",\"text\":\"At Bellford, Coastliner C88's 16:09 departure uses platform 5.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 arrival uses platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 arrival uses platform 3.", "negative_left": "At Bellford, R214's 16:02 arrival uses platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 departure uses platform 3.", "right": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-004", "id": "fast-43-diverse-093-004-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rebooking log", "text": "At Bellford, R214's 16:02 arrival uses platform 3."}, {"speaker": "rebooking log", "text": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler's constraints and the coordinator's minimum-connection-time rule intact, altering only case-specific platform facts, which is a permitted observation change; the counterfactual changes a single sentence (C88's departure platform from 5 to 3) coherently without contradicting other stated facts; the two focus evidence spans are complete factual log sentences, not policy text; and neither context states or hints at the feasibility outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"rebooking log\",\"text\":\"At Bellford, R214's 16:02 arrival uses platform 3.\"},{\"speaker\":\"rebooking log\",\"text\":\"At Bellford, Coastliner C88's 16:09 departure uses platform 5.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 arrival uses platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 arrival uses platform 3.", "negative_left": "At Bellford, R214's 16:02 arrival uses platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 departure uses platform 3.", "right": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-004", "id": "fast-43-diverse-093-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rebooking log", "text": "At Bellford, R214's 16:02 arrival uses platform 3."}, {"speaker": "rebooking log", "text": "At Bellford, Coastliner C88's 16:09 departure uses platform 3."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler's constraints, Option A/B details, and the exact cross/same-platform minimum-connection rule from the original state; the focus evidence are two plain factual arrival/departure sentences with no policy or answer text; the counterfactual only swaps C88's departure platform from 5 to 3, consistently making it a same-platform connection without contradicting any other stated fact.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving at 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, R214's 16:02 arrival is recorded at platform 3.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, Coastliner C88's 16:09 departure is recorded at platform 5.\"}, {\"speaker\": \"rail traveler\", \"text\": \"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 arrival is recorded at platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 departure is recorded at platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 arrival is recorded at platform 3.", "negative_left": "At Bellford, R214's 16:02 arrival is recorded at platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 departure is recorded at platform 3.", "right": "At Bellford, Coastliner C88's 16:09 departure is recorded at platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-006", "id": "fast-43-diverse-093-006-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving at 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "train operations coordinator", "text": "At Bellford, R214's 16:02 arrival is recorded at platform 3."}, {"speaker": "train operations coordinator", "text": "At Bellford, Coastliner C88's 16:09 departure is recorded at platform 5."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the traveler's constraints, Option A/B details, and the exact cross/same-platform minimum-connection rule from the original state; the focus evidence are two plain factual arrival/departure sentences with no policy or answer text; the counterfactual only swaps C88's departure platform from 5 to 3, consistently making it a same-platform connection without contradicting any other stated fact.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving at 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, R214's 16:02 arrival is recorded at platform 3.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, Coastliner C88's 16:09 departure is recorded at platform 5.\"}, {\"speaker\": \"rail traveler\", \"text\": \"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 arrival is recorded at platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 departure is recorded at platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 arrival is recorded at platform 3.", "negative_left": "At Bellford, R214's 16:02 arrival is recorded at platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 departure is recorded at platform 3.", "right": "At Bellford, Coastliner C88's 16:09 departure is recorded at platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-006", "id": "fast-43-diverse-093-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving at 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "train operations coordinator", "text": "At Bellford, R214's 16:02 arrival is recorded at platform 3."}, {"speaker": "train operations coordinator", "text": "At Bellford, Coastliner C88's 16:09 departure is recorded at platform 3."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, times, and the stated minimum-connection policy; the counterfactual only swaps C88's platform to 3, making it internally consistent (same-platform, 7-min gap ≥5-min rule) while base context keeps cross-platform failure; evidence spans are two complete factual sentences with no policy text or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"R214 is scheduled to arrive at Bellford at 16:02 on platform 3. Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5. Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"rail traveler\", \"text\": \"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["2", "text"], "text": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3.", "right": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-007", "id": "fast-43-diverse-093-007-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05."}, {"speaker": "train operations coordinator", "text": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3. Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5. Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, times, and the stated minimum-connection policy; the counterfactual only swaps C88's platform to 3, making it internally consistent (same-platform, 7-min gap ≥5-min rule) while base context keeps cross-platform failure; evidence spans are two complete factual sentences with no policy text or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"R214 is scheduled to arrive at Bellford at 16:02 on platform 3. Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5. Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"rail traveler\", \"text\": \"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["2", "text"], "text": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3.", "right": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-007", "id": "fast-43-diverse-093-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05."}, {"speaker": "train operations coordinator", "text": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3. Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3. Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler's constraints and connection-time rule intact, only altering C88's departure platform between contexts; evidence spans are two complete factual sentences with no embedded rules or answers, and the platform-3 counterfactual is a coherent single-fact edit consistent with the rest of the scenario.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, then Coastliner C88 to Seaborne, arriving 17:20.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"R214 is scheduled to arrive at Bellford at 16:02 on platform 3. Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5. Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"rail traveler\", \"text\": \"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["2", "text"], "text": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3.", "right": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-010", "id": "fast-43-diverse-093-010-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, then Coastliner C88 to Seaborne, arriving 17:20."}, {"speaker": "train operations coordinator", "text": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3. Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5. Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler's constraints and connection-time rule intact, only altering C88's departure platform between contexts; evidence spans are two complete factual sentences with no embedded rules or answers, and the platform-3 counterfactual is a coherent single-fact edit consistent with the rest of the scenario.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, then Coastliner C88 to Seaborne, arriving 17:20.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"R214 is scheduled to arrive at Bellford at 16:02 on platform 3. Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5. Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"rail traveler\", \"text\": \"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3."}, {"path": ["2", "text"], "text": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_left": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3.", "negative_right": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3.", "right": "Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-010", "id": "fast-43-diverse-093-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, then Coastliner C88 to Seaborne, arriving 17:20."}, {"speaker": "train operations coordinator", "text": "R214 is scheduled to arrive at Bellford at 16:02 on platform 3. Coastliner C88 is scheduled to depart Bellford at 16:09 from platform 3. Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the minimum-connection-time rule (8 min cross-platform, 5 min same-platform) and the traveler's constraints, matching the unchanged question; only platform numbers/observations differ, which is permitted; the two evidence spans are plain factual statements about arrival/departure platforms, not policy text; the counterfactual coherently changes C88's platform to 3, making the connection same-platform with a 7-minute gap, without contradicting other unchanged facts; no gold answer, rule table, or instructional leakage appears in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five. At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3. At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 5.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3."}, {"path": ["2", "text"], "text": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3.", "negative_left": "At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 3.", "right": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-013", "id": "fast-43-diverse-093-013-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five. At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3. At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 5."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the minimum-connection-time rule (8 min cross-platform, 5 min same-platform) and the traveler's constraints, matching the unchanged question; only platform numbers/observations differ, which is permitted; the two evidence spans are plain factual statements about arrival/departure platforms, not policy text; the counterfactual coherently changes C88's platform to 3, making the connection same-platform with a 7-minute gap, without contradicting other unchanged facts; no gold answer, rule table, or instructional leakage appears in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five. At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3. At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 5.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["2", "text"], "text": "At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3."}, {"path": ["2", "text"], "text": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3.", "negative_left": "At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 3.", "right": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-013", "id": "fast-43-diverse-093-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five. At Bellford, R214's 16:02 scheduled arrival occurs at Platform 3. At Bellford, Coastliner C88's 16:09 scheduled departure occurs at Platform 3."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler's constraints, Option A/B details, and the governing cross-platform (8 min) vs same-platform (5 min) rule intact while only changing platform assignments as case-specific observations; the counterfactual switches C88's departure to platform 3, making the connection same-platform with a 7-min gap, which is internally consistent and does not duplicate or contradict other stated facts; the two focus evidence spans are complete factual sentences about platform usage, not policy text; neither context states a label, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A is R214 at 15:10 to Bellford, arriving at 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"platform dispatcher\",\"text\":\"At Bellford, R214's 16:02 arrival uses platform 3.\"},{\"speaker\":\"platform dispatcher\",\"text\":\"At Bellford, Coastliner C88's 16:09 departure uses platform 5.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 arrival uses platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 arrival uses platform 3.", "negative_left": "At Bellford, R214's 16:02 arrival uses platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 departure uses platform 3.", "right": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-014", "id": "fast-43-diverse-093-014-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving at 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "platform dispatcher", "text": "At Bellford, R214's 16:02 arrival uses platform 3."}, {"speaker": "platform dispatcher", "text": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler's constraints, Option A/B details, and the governing cross-platform (8 min) vs same-platform (5 min) rule intact while only changing platform assignments as case-specific observations; the counterfactual switches C88's departure to platform 3, making the connection same-platform with a 7-min gap, which is internally consistent and does not duplicate or contradict other stated facts; the two focus evidence spans are complete factual sentences about platform usage, not policy text; neither context states a label, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A is R214 at 15:10 to Bellford, arriving at 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"platform dispatcher\",\"text\":\"At Bellford, R214's 16:02 arrival uses platform 3.\"},{\"speaker\":\"platform dispatcher\",\"text\":\"At Bellford, Coastliner C88's 16:09 departure uses platform 5.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 arrival uses platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 arrival uses platform 3.", "negative_left": "At Bellford, R214's 16:02 arrival uses platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 departure uses platform 3.", "right": "At Bellford, Coastliner C88's 16:09 departure uses platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-014", "id": "fast-43-diverse-093-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving at 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20. Option B is a direct 16:05, arriving at 18:05."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "platform dispatcher", "text": "At Bellford, R214's 16:02 arrival uses platform 3."}, {"speaker": "platform dispatcher", "text": "At Bellford, Coastliner C88's 16:09 departure uses platform 3."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the traveler's constraints, C88 identity, and the minimum-connection-time policy stated in the unchanged question; evidence spans are two complete factual case-note sentences about platform arrival/departure, not policy text; the counterfactual coherently changes C88's departure platform to match the arrival platform, making it a same-platform connection without contradicting other stated facts; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"case note\", \"text\": \"At Bellford, R214 arrives at 16:02 on Platform 3.\"}, {\"speaker\": \"case note\", \"text\": \"At Bellford, Coastliner C88 departs at 16:09 from Platform 5.\"}, {\"speaker\": \"rail traveler\", \"text\": \"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214 arrives at 16:02 on Platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88 departs at 16:09 from Platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214 arrives at 16:02 on Platform 3.", "negative_left": "At Bellford, R214 arrives at 16:02 on Platform 3.", "negative_right": "At Bellford, Coastliner C88 departs at 16:09 from Platform 3.", "right": "At Bellford, Coastliner C88 departs at 16:09 from Platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-021", "id": "fast-43-diverse-093-021-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "case note", "text": "At Bellford, R214 arrives at 16:02 on Platform 3."}, {"speaker": "case note", "text": "At Bellford, Coastliner C88 departs at 16:09 from Platform 5."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the traveler's constraints, C88 identity, and the minimum-connection-time policy stated in the unchanged question; evidence spans are two complete factual case-note sentences about platform arrival/departure, not policy text; the counterfactual coherently changes C88's departure platform to match the arrival platform, making it a same-platform connection without contradicting other stated facts; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"case note\", \"text\": \"At Bellford, R214 arrives at 16:02 on Platform 3.\"}, {\"speaker\": \"case note\", \"text\": \"At Bellford, Coastliner C88 departs at 16:09 from Platform 5.\"}, {\"speaker\": \"rail traveler\", \"text\": \"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214 arrives at 16:02 on Platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88 departs at 16:09 from Platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214 arrives at 16:02 on Platform 3.", "negative_left": "At Bellford, R214 arrives at 16:02 on Platform 3.", "negative_right": "At Bellford, Coastliner C88 departs at 16:09 from Platform 3.", "right": "At Bellford, Coastliner C88 departs at 16:09 from Platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-021", "id": "fast-43-diverse-093-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "case note", "text": "At Bellford, R214 arrives at 16:02 on Platform 3."}, {"speaker": "case note", "text": "At Bellford, Coastliner C88 departs at 16:09 from Platform 3."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler's constraints and the 8/5-minute platform rule intact while only changing platform assignments, and the counterfactual's same-platform change (3 to 3) coherently alters the connection type without contradicting other facts or stating an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 5.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3.", "negative_left": "At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3.", "negative_right": "At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 3.", "right": "At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-023", "id": "fast-43-diverse-093-023-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "train operations coordinator", "text": "At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3."}, {"speaker": "train operations coordinator", "text": "At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 5."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the traveler's constraints and the 8/5-minute platform rule intact while only changing platform assignments, and the counterfactual's same-platform change (3 to 3) coherently alters the connection type without contradicting other facts or stating an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\": \"rail traveler\", \"text\": \"My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00.\"}, {\"speaker\": \"station rebooking agent\", \"text\": \"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3.\"}, {\"speaker\": \"train operations coordinator\", \"text\": \"At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 5.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3.", "negative_left": "At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3.", "negative_right": "At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 3.", "right": "At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-023", "id": "fast-43-diverse-093-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "My through-ticket is for the canceled 14:40 from Alderwick to Seaborne. I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "train operations coordinator", "text": "At Bellford, R214's scheduled 16:02 arrival occurs at Platform 3."}, {"speaker": "train operations coordinator", "text": "At Bellford, Coastliner C88's scheduled 16:09 departure occurs at Platform 3."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the transfer/arrival constraints and connection-time rule while only altering platform details, preserving all question bindings without revealing the feasibility answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"At Bellford, R214's 16:02 scheduled arrival occurs at platform 3.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 5.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 scheduled arrival occurs at platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 scheduled arrival occurs at platform 3.", "negative_left": "At Bellford, R214's 16:02 scheduled arrival occurs at platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 3.", "right": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-026", "id": "fast-43-diverse-093-026-base", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "train operations coordinator", "text": "At Bellford, R214's 16:02 scheduled arrival occurs at platform 3."}, {"speaker": "train operations coordinator", "text": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 5."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the transfer/arrival constraints and connection-time rule while only altering platform details, preserving all question bindings without revealing the feasibility answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "full_context_fact_states": {"base": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "supported", "a_transfer_count": "supported"}, "counterfactual": {"a_arrival_deadline": "supported", "a_connection_duration": "supported", "a_platform_difference": "refuted", "a_transfer_count": "supported"}, "remove_left": {"a_platform_difference": "unknown"}, "remove_right": {"a_platform_difference": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_platform_difference": "unknown"}, "negative_pair": {"a_platform_difference": "refuted"}, "negative_sentence": {"a_platform_difference": "unknown"}, "positive_pair": {"a_platform_difference": "supported"}, "right": {"a_platform_difference": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus concerns the factual relationship between the arrival and departure platforms rather than a policy conclusion. The base and counter assignments can be realized with identical transfer count, deadline compliance, and connection duration while changing only whether the platforms differ. Policy evidence correctly preserves the traveler’s transfer/deadline constraints and Bellford’s state-originating connection-time rules; instructions and criteria from the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish exactly one transfer, arrival by 18:00, a seven-minute connection, and different platforms. Different platforms make the connection cross-platform, for which the preserved policy requires at least eight minutes. Seven minutes is below the minimum, so infeasibility is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the platform-difference atom establishes that the arrival and departure platforms do not differ, hence the connection is same-platform. The preserved policy requires five minutes for such a connection, and the exact seven-minute duration meets it. The other conditions establish one transfer and arrival by 18:00, so all stated constraints are satisfied.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_transfer_count", "statement": "Option A requires exactly one transfer between Alderwick and Seaborne."}, {"id": "a_arrival_deadline", "statement": "Option A's scheduled arrival at Seaborne is no later than 18:00."}, {"id": "a_connection_duration", "statement": "Option A provides exactly seven minutes between R214's scheduled 16:02 arrival at Bellford and Coastliner C88's scheduled 16:09 departure from Bellford."}, {"id": "a_platform_difference", "statement": "For Option A at Bellford, R214's 16:02 arrival platform differs from Coastliner C88's 16:09 departure platform."}], "base_state_json": "[{\"speaker\":\"rail traveler\",\"text\":\"I can make one transfer, but I must arrive by 18:00.\"},{\"speaker\":\"station rebooking agent\",\"text\":\"Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"At Bellford, R214's 16:02 scheduled arrival occurs at platform 3.\"},{\"speaker\":\"train operations coordinator\",\"text\":\"At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 5.\"},{\"speaker\":\"rail traveler\",\"text\":\"Please put me on Option A if it is feasible.\"}]", "base_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}], "counter_states": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}], "focus_atom": "a_platform_difference", "focus_evidence": [{"path": ["3", "text"], "text": "At Bellford, R214's 16:02 scheduled arrival occurs at platform 3."}, {"path": ["4", "text"], "text": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 5."}], "policy_evidence": [{"path": ["0", "text"], "text": "I can make one transfer, but I must arrive by 18:00."}, {"path": ["2", "text"], "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}], "rules": [{"justification": "With different arrival and departure platforms at Bellford, the connection is cross-platform and requires at least eight minutes. The seven-minute connection is below that minimum, so Option A is not feasible.", "target": "false", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "supported"}]}, {"justification": "If the arrival and departure platforms do not differ, the connection is same-platform and requires five minutes. Seven minutes exceeds that minimum, while Option A also has only one transfer and arrives by 18:00, so it is feasible.", "target": "true", "when": [{"atom_id": "a_transfer_count", "state": "supported"}, {"atom_id": "a_arrival_deadline", "state": "supported"}, {"atom_id": "a_connection_duration", "state": "supported"}, {"atom_id": "a_platform_difference", "state": "refuted"}]}]}, "verified_pair": {"left": "At Bellford, R214's 16:02 scheduled arrival occurs at platform 3.", "negative_left": "At Bellford, R214's 16:02 scheduled arrival occurs at platform 3.", "negative_right": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 3.", "right": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-093-026", "id": "fast-43-diverse-093-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Option A violates at least one stated traveler constraint or minimum connection time, so it is not feasible.", "true": "Option A satisfies the traveler’s arrival and transfer constraints and every stated minimum connection time, so it can be booked as feasible."}, "instructions": "Determine whether the traveler’s requested Option A is feasible. Resolve which train “C88” refers to and apply the stated minimum connection rule. A connection exactly meeting the minimum is feasible; one below it is not.", "type": "noul"}}, "state": [{"speaker": "rail traveler", "text": "I can make one transfer, but I must arrive by 18:00."}, {"speaker": "station rebooking agent", "text": "Option A is R214 at 15:10 to Bellford, arriving 16:02, then Coastliner C88 at 16:09, arriving Seaborne at 17:20."}, {"speaker": "train operations coordinator", "text": "Cross-platform connections at Bellford require at least eight minutes; same-platform ones require five."}, {"speaker": "train operations coordinator", "text": "At Bellford, R214's 16:02 scheduled arrival occurs at platform 3."}, {"speaker": "train operations coordinator", "text": "At Bellford, Coastliner C88's 16:09 scheduled departure occurs at platform 3."}, {"speaker": "rail traveler", "text": "Please put me on Option A if it is feasible."}]}, "method": "c2d", "provenance": {"source_id": "diverse-093", "source_is_synthetic": true, "source_sha256": "291cc620804d830f1329e185b5ea77c7165b51560c2694867fcfb02ee9cc26b4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy via the unchanged questions object and consistent factual claims, the counterfactual only alters Birch's arrival time (18:12→18:47) without duplicating or contradicting other stated facts, the two focus evidence spans are complete factual sentences, and neither context contains a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"A cancellation strands a rail traveler at Lanton. Rebooking proceeds under the traveler's flexible ticket with cancellation endorsement, valid on all listed replacement services.\",\"facts\":[\"The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.\",\"The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one.\",\"The traveler states no other hard requirements beyond arrival time and transfer count.\",\"The traveler confirms no priority overriding earliest-arrival ordering among eligible routes.\",\"Alder and Birch are both operating as replacement routes from Lanton to Merrow, each with one transfer.\",\"The flexible ticket with cancellation endorsement is accepted on both Alder and Birch.\",\"Both Alder’s and Birch’s interchange durations meet or exceed the published 12-minute minimum.\",\"Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"Both Alder and Birch have a single transfer, matching the traveler’s stated maximum.\",\"Birch’s scheduled arrival time at Merrow is 18:12.\",\"Alder’s scheduled arrival time at Merrow is 18:05, earlier than Birch’s.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["facts", "0"], "text": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"path": ["facts", "9"], "text": "Birch’s scheduled arrival time at Merrow is 18:12."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "negative_left": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "negative_right": "Birch’s scheduled arrival time at Merrow is 18:47.", "right": "Birch’s scheduled arrival time at Merrow is 18:12."}, "verifier_independent_model": false}, "family": "fast-43-diverse-095-016", "id": "fast-43-diverse-095-016-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "A cancellation strands a rail traveler at Lanton. Rebooking proceeds under the traveler's flexible ticket with cancellation endorsement, valid on all listed replacement services.", "facts": ["The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one.", "The traveler states no other hard requirements beyond arrival time and transfer count.", "The traveler confirms no priority overriding earliest-arrival ordering among eligible routes.", "Alder and Birch are both operating as replacement routes from Lanton to Merrow, each with one transfer.", "The flexible ticket with cancellation endorsement is accepted on both Alder and Birch.", "Both Alder’s and Birch’s interchange durations meet or exceed the published 12-minute minimum.", "Birch’s 14-minute connection exceeds the published 12-minute minimum.", "Both Alder and Birch have a single transfer, matching the traveler’s stated maximum.", "Birch’s scheduled arrival time at Merrow is 18:12.", "Alder’s scheduled arrival time at Merrow is 18:05, earlier than Birch’s."]}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original governing policy via the unchanged questions object and consistent factual claims, the counterfactual only alters Birch's arrival time (18:12→18:47) without duplicating or contradicting other stated facts, the two focus evidence spans are complete factual sentences, and neither context contains a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\":\"A cancellation strands a rail traveler at Lanton. Rebooking proceeds under the traveler's flexible ticket with cancellation endorsement, valid on all listed replacement services.\",\"facts\":[\"The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.\",\"The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one.\",\"The traveler states no other hard requirements beyond arrival time and transfer count.\",\"The traveler confirms no priority overriding earliest-arrival ordering among eligible routes.\",\"Alder and Birch are both operating as replacement routes from Lanton to Merrow, each with one transfer.\",\"The flexible ticket with cancellation endorsement is accepted on both Alder and Birch.\",\"Both Alder’s and Birch’s interchange durations meet or exceed the published 12-minute minimum.\",\"Birch’s 14-minute connection exceeds the published 12-minute minimum.\",\"Both Alder and Birch have a single transfer, matching the traveler’s stated maximum.\",\"Birch’s scheduled arrival time at Merrow is 18:12.\",\"Alder’s scheduled arrival time at Merrow is 18:05, earlier than Birch’s.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["facts", "0"], "text": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"path": ["facts", "9"], "text": "Birch’s scheduled arrival time at Merrow is 18:12."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "negative_left": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "negative_right": "Birch’s scheduled arrival time at Merrow is 18:47.", "right": "Birch’s scheduled arrival time at Merrow is 18:12."}, "verifier_independent_model": false}, "family": "fast-43-diverse-095-016", "id": "fast-43-diverse-095-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "A cancellation strands a rail traveler at Lanton. Rebooking proceeds under the traveler's flexible ticket with cancellation endorsement, valid on all listed replacement services.", "facts": ["The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one.", "The traveler states no other hard requirements beyond arrival time and transfer count.", "The traveler confirms no priority overriding earliest-arrival ordering among eligible routes.", "Alder and Birch are both operating as replacement routes from Lanton to Merrow, each with one transfer.", "The flexible ticket with cancellation endorsement is accepted on both Alder and Birch.", "Both Alder’s and Birch’s interchange durations meet or exceed the published 12-minute minimum.", "Birch’s 14-minute connection exceeds the published 12-minute minimum.", "Both Alder and Birch have a single transfer, matching the traveler’s stated maximum.", "Birch’s scheduled arrival time at Merrow is 18:47.", "Alder’s scheduled arrival time at Merrow is 18:05, earlier than Birch’s."]}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep ticket validity, transfer minimum, and traveler-confirmed hard-constraint policy intact alongside the unchanged question, preserve the Lanton-Merrow/Birch bindings and 18:30/one-transfer limits, use factual arrival-time sentences as evidence, and the counterfactual's later Birch arrival (18:45) is internally consistent without contradicting other stated facts or leaking any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\": \"A cancellation strands a rail traveler at Lanton. Their flexible ticket with a cancellation endorsement is valid on all listed replacement routes: Alder, Birch, and Cedar. Birch’s 14-minute connection exceeds the published 12-minute minimum. Alder’s connection is 13 minutes, also above the minimum. Both Alder and Birch have exactly one transfer.\", \"evidence\": [\"The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.\", \"The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one.\", \"The traveler confirms no hard constraints beyond arrival time and transfer count, and no priority overriding earliest-arrival ordering among eligible routes.\", \"Birch is scheduled to arrive at Merrow at 18:15.\", \"Alder is scheduled to arrive at Merrow at 18:05, earlier than Birch.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["evidence", "0"], "text": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"path": ["evidence", "3"], "text": "Birch is scheduled to arrive at Merrow at 18:15."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "negative_left": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "negative_right": "Birch is scheduled to arrive at Merrow at 18:45.", "right": "Birch is scheduled to arrive at Merrow at 18:15."}, "verifier_independent_model": false}, "family": "fast-43-diverse-095-019", "id": "fast-43-diverse-095-019-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "A cancellation strands a rail traveler at Lanton. Their flexible ticket with a cancellation endorsement is valid on all listed replacement routes: Alder, Birch, and Cedar. Birch’s 14-minute connection exceeds the published 12-minute minimum. Alder’s connection is 13 minutes, also above the minimum. Both Alder and Birch have exactly one transfer.", "evidence": ["The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one.", "The traveler confirms no hard constraints beyond arrival time and transfer count, and no priority overriding earliest-arrival ordering among eligible routes.", "Birch is scheduled to arrive at Merrow at 18:15.", "Alder is scheduled to arrive at Merrow at 18:05, earlier than Birch."]}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-01", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep ticket validity, transfer minimum, and traveler-confirmed hard-constraint policy intact alongside the unchanged question, preserve the Lanton-Merrow/Birch bindings and 18:30/one-transfer limits, use factual arrival-time sentences as evidence, and the counterfactual's later Birch arrival (18:45) is internally consistent without contradicting other stated facts or leaking any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a9": "unknown"}, "remove_right": {"a9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a9": "unknown"}, "negative_pair": {"a9": "refuted"}, "negative_sentence": {"a9": "unknown"}, "positive_pair": {"a9": "supported"}, "right": {"a9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; a3 is a single negative completeness claim about the traveler’s confirmed hard constraints rather than a final classification. The focus a9 is a factual comparison between Birch’s arrival and the confirmed deadline. The base and counter assignments are jointly realizable while changing only a9: Birch can arrive by the deadline in the base and after it in the counter, while Alder remains earlier and timely in both. Policy evidence preserves the state-originated 12-minute published interchange threshold needed by the plan; the remaining governing criteria and confirmation priority are already retained verbatim in original_input.questions, while other state details are replaceable case observations.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that Birch is operating, accepts the ticket, meets the published interchange minimum, and satisfies all confirmed hard constraints. It also establishes that Alder is eligible and arrives earlier, with no confirmed priority overriding earliest arrival. Therefore Birch must be an acceptable fallback rather than most suitable.", "rule_index": 0, "sound": true}, {"reason": "With the confirmed deadline established, refutation of a9 entails that Birch arrives after that deadline. This violates a confirmed hard traveler constraint and is independently sufficient for level 0.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"id": "a2", "statement": "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one."}, {"id": "a3", "statement": "The traveler confirms no hard constraints for the replacement journey from Lanton to Merrow beyond the latest arrival time and maximum transfer count in the agent’s paraphrase."}, {"id": "a4", "statement": "The traveler confirms no priority that overrides the rubric’s earliest-arrival ordering among eligible replacement routes from Lanton to Merrow."}, {"id": "a5", "statement": "Birch is operating as a replacement route from Lanton to Merrow."}, {"id": "a6", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Birch."}, {"id": "a7", "statement": "Birch’s interchange duration is at least the published 12-minute minimum."}, {"id": "a8", "statement": "Birch’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a9", "statement": "Birch’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a10", "statement": "Alder is operating as a replacement route from Lanton to Merrow."}, {"id": "a11", "statement": "The traveler’s flexible ticket with the cancellation endorsement is valid on Alder."}, {"id": "a12", "statement": "Alder’s interchange duration is at least the published 12-minute minimum."}, {"id": "a13", "statement": "Alder’s transfer count does not exceed the maximum transfer count confirmed by the traveler."}, {"id": "a14", "statement": "Alder’s scheduled arrival time at Merrow is no later than the latest arrival time in the agent’s paraphrase confirmed by the traveler."}, {"id": "a15", "statement": "Alder’s scheduled arrival time at Merrow is earlier than Birch’s scheduled arrival time at Merrow."}], "base_state_json": "{\"context\": \"A cancellation strands a rail traveler at Lanton. Their flexible ticket with a cancellation endorsement is valid on all listed replacement routes: Alder, Birch, and Cedar. Birch’s 14-minute connection exceeds the published 12-minute minimum. Alder’s connection is 13 minutes, also above the minimum. Both Alder and Birch have exactly one transfer.\", \"evidence\": [\"The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.\", \"The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one.\", \"The traveler confirms no hard constraints beyond arrival time and transfer count, and no priority overriding earliest-arrival ordering among eligible routes.\", \"Birch is scheduled to arrive at Merrow at 18:15.\", \"Alder is scheduled to arrive at Merrow at 18:05, earlier than Birch.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}], "focus_atom": "a9", "focus_evidence": [{"path": ["evidence", "0"], "text": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30."}, {"path": ["evidence", "3"], "text": "Birch is scheduled to arrive at Merrow at 18:15."}], "policy_evidence": [{"path": ["context"], "text": "Birch’s 14-minute connection exceeds the published 12-minute minimum."}], "rules": [{"justification": "Birch is valid, feasible, and within the traveler’s complete confirmed hard constraints, while eligible Alder arrives earlier and no confirmed priority overrides earliest arrival.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}]}, {"justification": "Birch arrives later than the latest arrival time in the agent’s paraphrase confirmed by the traveler, so it violates a confirmed hard traveler constraint.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "negative_left": "The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "negative_right": "Birch is scheduled to arrive at Merrow at 18:45.", "right": "Birch is scheduled to arrive at Merrow at 18:15."}, "verifier_independent_model": false}, "family": "fast-43-diverse-095-019", "id": "fast-43-diverse-095-019-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsuitable: the route has invalid ticket acceptance, an interchange below the published minimum, or violates any confirmed hard traveler constraint.", "1 — Acceptable fallback: the route is valid, feasible, and meets every hard constraint, but another eligible route is better by arrival time, transfer count, or a confirmed preference.", "2 — Most suitable: the route is valid, feasible, meets every hard constraint, and offers the best eligible balance—earliest arrival first, then fewer transfers unless the traveler confirms another priority."], "instructions": "Assign one level from 0 to 2. Treat the traveler’s confirmation of the agent’s paraphrase as the authoritative interpretation of their constraints. Apply the levels from least to most suitable.", "type": "score"}}, "state": {"context": "A cancellation strands a rail traveler at Lanton. Their flexible ticket with a cancellation endorsement is valid on all listed replacement routes: Alder, Birch, and Cedar. Birch’s 14-minute connection exceeds the published 12-minute minimum. Alder’s connection is 13 minutes, also above the minimum. Both Alder and Birch have exactly one transfer.", "evidence": ["The traveler confirms the rebooking agent’s paraphrased latest permissible arrival time at Merrow as 18:30.", "The traveler confirms the rebooking agent’s paraphrased maximum permissible transfer count for the replacement journey from Lanton to Merrow as one.", "The traveler confirms no hard constraints beyond arrival time and transfer count, and no priority overriding earliest-arrival ordering among eligible routes.", "Birch is scheduled to arrive at Merrow at 18:45.", "Alder is scheduled to arrive at Merrow at 18:05, earlier than Birch."]}}, "method": "c2d", "provenance": {"source_id": "diverse-095", "source_is_synthetic": true, "source_sha256": "e11f11d50848779b17365e080c95f17ac74ff6bfcaf58146a3504f49a29b1de2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-01", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve vehicle, rest, lodging and traveler-condition policies unchanged from the question; only the winery arrival time (12:05→12:25) is altered, which is a coherent single-fact change with no contradictory duplicate measurements, no policy tables, and no embedded answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, \\u201cUse Ridge only if an AWD is confirmed; otherwise use Valley.\\u201d\",\"No AWD vehicle has been confirmed for the trip; only an automatic compact rental is confirmed, and it is permitted on every road in the Valley plan, including the paved roads used throughout.\",\"The Valley plan has the traveler entering the museum at 10:20, within the 10:00\\u201311:00 entry window.\",\"The booked winery visit for Friday has a fixed start time of 12:15.\",\"The Friday Valley plan has the traveler arriving at the winery at 12:05.\",\"No change to the Valley route, reservation, or vehicle would let the traveler reach the winery earlier than the observed arrival time.\",\"The Valley plan has the traveler reaching the lodging at 15:05, ten minutes before the 15:15 check-in closing time.\",\"Every continuous driving segment in the Valley plan stays within the 2-hour rest-trigger threshold.\",\"A room is confirmed for the required Friday night.\",\"A parking space is confirmed for the Friday Valley plan.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\"],\"request\":\"Apply the traveler\\u2019s condition, then rate the selected plan\\u2019s feasibility and convenience using the levels.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "3"], "text": "The booked winery visit for Friday has a fixed start time of 12:15."}, {"path": ["evidence", "4"], "text": "The Friday Valley plan has the traveler arriving at the winery at 12:05."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The booked winery visit for Friday has a fixed start time of 12:15.", "negative_left": "The booked winery visit for Friday has a fixed start time of 12:15.", "negative_right": "The Friday Valley plan has the traveler arriving at the winery at 12:25.", "right": "The Friday Valley plan has the traveler arriving at the winery at 12:05."}, "verifier_independent_model": false}, "family": "fast-43-diverse-108-017", "id": "fast-43-diverse-108-017-base", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "No AWD vehicle has been confirmed for the trip; only an automatic compact rental is confirmed, and it is permitted on every road in the Valley plan, including the paved roads used throughout.", "The Valley plan has the traveler entering the museum at 10:20, within the 10:00–11:00 entry window.", "The booked winery visit for Friday has a fixed start time of 12:15.", "The Friday Valley plan has the traveler arriving at the winery at 12:05.", "No change to the Valley route, reservation, or vehicle would let the traveler reach the winery earlier than the observed arrival time.", "The Valley plan has the traveler reaching the lodging at 15:05, ten minutes before the 15:15 check-in closing time.", "Every continuous driving segment in the Valley plan stays within the 2-hour rest-trigger threshold.", "A room is confirmed for the required Friday night.", "A parking space is confirmed for the Friday Valley plan.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "travel-03", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve vehicle, rest, lodging and traveler-condition policies unchanged from the question; only the winery arrival time (12:05→12:25) is altered, which is a coherent single-fact change with no contradictory duplicate measurements, no policy tables, and no embedded answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified road and driving-segment atoms remain atomic. The focus atom concerns the factual relation between winery arrival and its fixed start, not a policy classification. The base and counter assignments can be realized with only winery timeliness changing while all other atom states remain fixed. Policy evidence appropriately preserves the state-origin conditional plan-selection rule, rest rule, and relevant scope information; criteria and thresholds in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions select Valley because no AWD is confirmed. They establish a compatible confirmed vehicle, lodging, parking, compliant rest timing, and museum timing, while a5 is refuted and a6 establishes that the missed fixed winery start cannot be corrected without changing the route, reservation, or vehicle. This is sufficient for level 1 and excludes level 0.", "rule_index": 0, "sound": true}, {"reason": "The conditions select Valley and establish confirmed compatible transport, confirmed room and parking, compliant visit timing, no required rest insertion, and timely lodging arrival. The lodging margin is explicitly 0–20 minutes, which satisfies level 3's narrow-margin condition and prevents level 4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "No AWD vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a2", "statement": "An automatic compact rental vehicle is confirmed for the traveler’s Friday self-drive itinerary."}, {"id": "a3", "statement": "The confirmed automatic compact is permitted on every road in the Friday Valley plan."}, {"id": "a4", "statement": "The Friday Valley plan’s museum entry occurs within the museum’s 10:00–11:00 entry window."}, {"id": "a5", "statement": "The Friday Valley plan’s winery arrival time is no later than the booked winery visit’s fixed 12:15 start time."}, {"id": "a6", "statement": "The earliest winery arrival obtainable without changing the Friday Valley route, reservation, or vehicle equals the plan’s observed winery arrival time."}, {"id": "a7", "statement": "The Friday Valley plan’s lodging arrival margin before the 15:15 check-in closing time is 0–20 minutes inclusive."}, {"id": "a8", "statement": "Every continuous driving segment in the Friday Valley plan is at most 2 hours."}, {"id": "a9", "statement": "A room is confirmed for the required Friday night."}, {"id": "a10", "statement": "A parking space is confirmed for the Friday Valley plan."}], "base_state_json": "{\"context\":\"An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.\",\"evidence\":[\"The road-trip traveler says, \\u201cUse Ridge only if an AWD is confirmed; otherwise use Valley.\\u201d\",\"No AWD vehicle has been confirmed for the trip; only an automatic compact rental is confirmed, and it is permitted on every road in the Valley plan, including the paved roads used throughout.\",\"The Valley plan has the traveler entering the museum at 10:20, within the 10:00\\u201311:00 entry window.\",\"The booked winery visit for Friday has a fixed start time of 12:15.\",\"The Friday Valley plan has the traveler arriving at the winery at 12:05.\",\"No change to the Valley route, reservation, or vehicle would let the traveler reach the winery earlier than the observed arrival time.\",\"The Valley plan has the traveler reaching the lodging at 15:05, ten minutes before the 15:15 check-in closing time.\",\"Every continuous driving segment in the Valley plan stays within the 2-hour rest-trigger threshold.\",\"A room is confirmed for the required Friday night.\",\"A parking space is confirmed for the Friday Valley plan.\",\"A 15m rest is required only after more than 2h continuous driving; visits reset the clock.\"],\"request\":\"Apply the traveler\\u2019s condition, then rate the selected plan\\u2019s feasibility and convenience using the levels.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["evidence", "3"], "text": "The booked winery visit for Friday has a fixed start time of 12:15."}, {"path": ["evidence", "4"], "text": "The Friday Valley plan has the traveler arriving at the winery at 12:05."}], "policy_evidence": [{"path": ["context"], "text": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic."}, {"path": ["evidence", "0"], "text": "The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”"}, {"path": ["evidence", "3"], "text": "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."}, {"path": ["request"], "text": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}], "rules": [{"justification": "With no AWD confirmed, the conditional intent selects Valley. A refuted a5 means the Valley plan arrives after the winery’s fixed start. Because a6 establishes that no earlier arrival is possible without changing the route, reservation, or vehicle, the selected plan is infeasible as booked under level 1.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "With no AWD confirmed, the conditional intent selects Valley. The confirmed vehicle is compatible, the museum and winery timing requirements are met, no rest insertion is required, and the room and parking are confirmed. The lodging deadline is met, but its 0–20-minute arrival margin establishes level 3 and excludes level 4.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "The booked winery visit for Friday has a fixed start time of 12:15.", "negative_left": "The booked winery visit for Friday has a fixed start time of 12:15.", "negative_right": "The Friday Valley plan has the traveler arriving at the winery at 12:25.", "right": "The Friday Valley plan has the traveler arriving at the winery at 12:05."}, "verifier_independent_model": false}, "family": "fast-43-diverse-108-017", "id": "fast-43-diverse-108-017-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unsafe or prohibited: the conditionally selected plan uses a road forbidden to the confirmed vehicle, violates an unavoidable rest rule, or has no lodging for the required night.", "1 — Infeasible as booked: the selected plan misses a fixed visit or check-in deadline and requires a changed route, reservation, or vehicle to work.", "2 — Potentially feasible but not ready: timing can satisfy every deadline, but a required vehicle, room, parking space, or rest-stop insertion remains unresolved.", "3 — Achievable and ready: vehicle and lodging are confirmed and all windows and rest rules are met, but at least one arrival margin is 0–20 minutes or the plan includes avoidable waiting.", "4 — Achievable and convenient: level 3 requirements are met, every fixed-start or closing margin exceeds 20 minutes, and there is no avoidable wait longer than 10 minutes."], "instructions": "First select the plan dictated by the traveler’s conditional intent. Then verify vehicle compatibility, visit windows, rest requirements, and lodging arrival. An arrival margin is the time between arrival and the applicable fixed start or closing time.", "type": "score"}}, "state": {"context": "An itinerary planner must select and score one Friday self-drive plan; listed driving times include normal traffic.", "evidence": ["The road-trip traveler says, “Use Ridge only if an AWD is confirmed; otherwise use Valley.”", "No AWD vehicle has been confirmed for the trip; only an automatic compact rental is confirmed, and it is permitted on every road in the Valley plan, including the paved roads used throughout.", "The Valley plan has the traveler entering the museum at 10:20, within the 10:00–11:00 entry window.", "The booked winery visit for Friday has a fixed start time of 12:15.", "The Friday Valley plan has the traveler arriving at the winery at 12:25.", "No change to the Valley route, reservation, or vehicle would let the traveler reach the winery earlier than the observed arrival time.", "The Valley plan has the traveler reaching the lodging at 15:05, ten minutes before the 15:15 check-in closing time.", "Every continuous driving segment in the Valley plan stays within the 2-hour rest-trigger threshold.", "A room is confirmed for the required Friday night.", "A parking space is confirmed for the Friday Valley plan.", "A 15m rest is required only after more than 2h continuous driving; visits reset the clock."], "request": "Apply the traveler’s condition, then rate the selected plan’s feasibility and convenience using the levels."}}, "method": "c2d", "provenance": {"source_id": "diverse-108", "source_is_synthetic": true, "source_sha256": "3a10e7cbb88794efa8d1784850927e99e428ef951177aa1ffaf907e33dcd351b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "travel-03", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy, entity, and deadline bindings intact, use two factual (non-policy) evidence sentences, and the counterfactual's single altered fact (passenger capacity 4→2) is internally consistent with the unchanged suitcase and party data, with no embedded answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, Elena is preparing for her hotel transfer along with her party. She has requested pickup by 18:40. The 18:35 accessible van's documented passenger capacity is 4. Elena's entire airport-to-hotel party consists of 3 travelers. The van also carries a documented suitcase capacity of four, matching Elena's four suitcases. The van has a documented ramp and wheelchair securement suitable for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp. The 18:50 accessible minibus arrives after the requested deadline. The airport assistance attendant only helps from baggage claim to the curb and does not provide a complete transfer to the hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's documented passenger capacity is 4."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's documented passenger capacity is 4.", "negative_left": "The 18:35 accessible van's documented passenger capacity is 2.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-032", "id": "fast-43-diverse-109-032-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, Elena is preparing for her hotel transfer along with her party. She has requested pickup by 18:40. The 18:35 accessible van's documented passenger capacity is 4. Elena's entire airport-to-hotel party consists of 3 travelers. The van also carries a documented suitcase capacity of four, matching Elena's four suitcases. The van has a documented ramp and wheelchair securement suitable for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp. The 18:50 accessible minibus arrives after the requested deadline. The airport assistance attendant only helps from baggage claim to the curb and does not provide a complete transfer to the hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy, entity, and deadline bindings intact, use two factual (non-policy) evidence sentences, and the counterfactual's single altered fact (passenger capacity 4→2) is internally consistent with the unchanged suitcase and party data, with no embedded answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, Elena is preparing for her hotel transfer along with her party. She has requested pickup by 18:40. The 18:35 accessible van's documented passenger capacity is 4. Elena's entire airport-to-hotel party consists of 3 travelers. The van also carries a documented suitcase capacity of four, matching Elena's four suitcases. The van has a documented ramp and wheelchair securement suitable for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp. The 18:50 accessible minibus arrives after the requested deadline. The airport assistance attendant only helps from baggage claim to the curb and does not provide a complete transfer to the hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's documented passenger capacity is 4."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's documented passenger capacity is 4.", "negative_left": "The 18:35 accessible van's documented passenger capacity is 2.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-032", "id": "fast-43-diverse-109-032-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, Elena is preparing for her hotel transfer along with her party. She has requested pickup by 18:40. The 18:35 accessible van's documented passenger capacity is 2. Elena's entire airport-to-hotel party consists of 3 travelers. The van also carries a documented suitcase capacity of four, matching Elena's four suitcases. The van has a documented ramp and wheelchair securement suitable for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp. The 18:50 accessible minibus arrives after the requested deadline. The airport assistance attendant only helps from baggage claim to the curb and does not provide a complete transfer to the hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same policy language, party size, luggage count, and 18:40 deadline as the original question, only altering the van's passenger-capacity figure between 4 and 2 seats, which is a permissible case-fact change; the two evidence spans are plain factual sentences with no policy text or answer hints, and the counterfactual remains internally consistent without contradicting other stated details.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Dispatch case note, Northstar Airport: Elena's airport-to-hotel party consists of 3 travelers, including Elena. The group carries four suitcases and Elena uses a rigid wheelchair; she cannot transfer from the chair or climb steps, and she requested hotel pickup by 18:40. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Dispatcher review of the 18:35 accessible van: The documented passenger capacity of the 18:35 accessible van is listed as 4 seats. Its documented suitcase capacity is listed as four, sufficient for the group's luggage. The van's records show a fitted ramp and wheelchair securement rated for a rigid-frame chair, with pickup scheduled for 18:35, ahead of the 18:40 deadline. The 18:25 sedan lacks any ramp documentation, the 18:50 minibus arrives after the deadline, and the curb attendant only assists from baggage claim to curb, not door-to-door.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena's airport-to-hotel party consists of 3 travelers, including Elena."}, {"path": [], "text": "The documented passenger capacity of the 18:35 accessible van is listed as 4 seats."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena's airport-to-hotel party consists of 3 travelers, including Elena.", "negative_left": "Elena's airport-to-hotel party consists of 3 travelers, including Elena.", "negative_right": "The documented passenger capacity of the 18:35 accessible van is listed as 2 seats.", "right": "The documented passenger capacity of the 18:35 accessible van is listed as 4 seats."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-040", "id": "fast-43-diverse-109-040-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Dispatch case note, Northstar Airport: Elena's airport-to-hotel party consists of 3 travelers, including Elena. The group carries four suitcases and Elena uses a rigid wheelchair; she cannot transfer from the chair or climb steps, and she requested hotel pickup by 18:40. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Dispatcher review of the 18:35 accessible van: The documented passenger capacity of the 18:35 accessible van is listed as 4 seats. Its documented suitcase capacity is listed as four, sufficient for the group's luggage. The van's records show a fitted ramp and wheelchair securement rated for a rigid-frame chair, with pickup scheduled for 18:35, ahead of the 18:40 deadline. The 18:25 sedan lacks any ramp documentation, the 18:50 minibus arrives after the deadline, and the curb attendant only assists from baggage claim to curb, not door-to-door."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same policy language, party size, luggage count, and 18:40 deadline as the original question, only altering the van's passenger-capacity figure between 4 and 2 seats, which is a permissible case-fact change; the two evidence spans are plain factual sentences with no policy text or answer hints, and the counterfactual remains internally consistent without contradicting other stated details.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Dispatch case note, Northstar Airport: Elena's airport-to-hotel party consists of 3 travelers, including Elena. The group carries four suitcases and Elena uses a rigid wheelchair; she cannot transfer from the chair or climb steps, and she requested hotel pickup by 18:40. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Dispatcher review of the 18:35 accessible van: The documented passenger capacity of the 18:35 accessible van is listed as 4 seats. Its documented suitcase capacity is listed as four, sufficient for the group's luggage. The van's records show a fitted ramp and wheelchair securement rated for a rigid-frame chair, with pickup scheduled for 18:35, ahead of the 18:40 deadline. The 18:25 sedan lacks any ramp documentation, the 18:50 minibus arrives after the deadline, and the curb attendant only assists from baggage claim to curb, not door-to-door.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena's airport-to-hotel party consists of 3 travelers, including Elena."}, {"path": [], "text": "The documented passenger capacity of the 18:35 accessible van is listed as 4 seats."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena's airport-to-hotel party consists of 3 travelers, including Elena.", "negative_left": "Elena's airport-to-hotel party consists of 3 travelers, including Elena.", "negative_right": "The documented passenger capacity of the 18:35 accessible van is listed as 2 seats.", "right": "The documented passenger capacity of the 18:35 accessible van is listed as 4 seats."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-040", "id": "fast-43-diverse-109-040-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Dispatch case note, Northstar Airport: Elena's airport-to-hotel party consists of 3 travelers, including Elena. The group carries four suitcases and Elena uses a rigid wheelchair; she cannot transfer from the chair or climb steps, and she requested hotel pickup by 18:40. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Dispatcher review of the 18:35 accessible van: The documented passenger capacity of the 18:35 accessible van is listed as 2 seats. Its documented suitcase capacity is listed as four, sufficient for the group's luggage. The van's records show a fitted ramp and wheelchair securement rated for a rigid-frame chair, with pickup scheduled for 18:35, ahead of the 18:40 deadline. The 18:25 sedan lacks any ramp documentation, the 18:50 minibus arrives after the deadline, and the curb attendant only assists from baggage claim to curb, not door-to-door."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all policy criteria and entity/time bindings unchanged from the original question; the base context merely restates observation values (van capacity 3, suitcase capacity 4) matching the party, and the single-sentence counterfactual lowers passenger capacity to 2 while leaving all other facts intact, remaining logically coherent; the two focus evidence spans are complete factual sentences with no policy language, rule tables, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Dispatch log update for Elena's airport transfer: Elena’s airport-to-hotel party consists of 3 travelers. The documented passenger capacity listed on the 18:35 accessible van’s manifest is 3. The van's manifest also lists a suitcase capacity of four, matching Elena's four suitcases. The van carries documented ramp equipment and documented wheelchair securement suitable for Elena's rigid wheelchair. Its scheduled pickup is 18:35, within Elena's requested 18:40 deadline. Separately, the 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus, while otherwise equipped, is scheduled after Elena's 18:40 deadline. The airport assistance attendant is logged as covering only baggage claim to curb, not the full route to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena’s airport-to-hotel party consists of 3 travelers."}, {"path": [], "text": "The documented passenger capacity listed on the 18:35 accessible van’s manifest is 3."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena’s airport-to-hotel party consists of 3 travelers.", "negative_left": "Elena’s airport-to-hotel party consists of 3 travelers.", "negative_right": "The documented passenger capacity listed on the 18:35 accessible van’s manifest is 2.", "right": "The documented passenger capacity listed on the 18:35 accessible van’s manifest is 3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-045", "id": "fast-43-diverse-109-045-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Dispatch log update for Elena's airport transfer: Elena’s airport-to-hotel party consists of 3 travelers. The documented passenger capacity listed on the 18:35 accessible van’s manifest is 3. The van's manifest also lists a suitcase capacity of four, matching Elena's four suitcases. The van carries documented ramp equipment and documented wheelchair securement suitable for Elena's rigid wheelchair. Its scheduled pickup is 18:35, within Elena's requested 18:40 deadline. Separately, the 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus, while otherwise equipped, is scheduled after Elena's 18:40 deadline. The airport assistance attendant is logged as covering only baggage claim to curb, not the full route to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all policy criteria and entity/time bindings unchanged from the original question; the base context merely restates observation values (van capacity 3, suitcase capacity 4) matching the party, and the single-sentence counterfactual lowers passenger capacity to 2 while leaving all other facts intact, remaining logically coherent; the two focus evidence spans are complete factual sentences with no policy language, rule tables, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Dispatch log update for Elena's airport transfer: Elena’s airport-to-hotel party consists of 3 travelers. The documented passenger capacity listed on the 18:35 accessible van’s manifest is 3. The van's manifest also lists a suitcase capacity of four, matching Elena's four suitcases. The van carries documented ramp equipment and documented wheelchair securement suitable for Elena's rigid wheelchair. Its scheduled pickup is 18:35, within Elena's requested 18:40 deadline. Separately, the 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus, while otherwise equipped, is scheduled after Elena's 18:40 deadline. The airport assistance attendant is logged as covering only baggage claim to curb, not the full route to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena’s airport-to-hotel party consists of 3 travelers."}, {"path": [], "text": "The documented passenger capacity listed on the 18:35 accessible van’s manifest is 3."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena’s airport-to-hotel party consists of 3 travelers.", "negative_left": "Elena’s airport-to-hotel party consists of 3 travelers.", "negative_right": "The documented passenger capacity listed on the 18:35 accessible van’s manifest is 2.", "right": "The documented passenger capacity listed on the 18:35 accessible van’s manifest is 3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-045", "id": "fast-43-diverse-109-045-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Dispatch log update for Elena's airport transfer: Elena’s airport-to-hotel party consists of 3 travelers. The documented passenger capacity listed on the 18:35 accessible van’s manifest is 2. The van's manifest also lists a suitcase capacity of four, matching Elena's four suitcases. The van carries documented ramp equipment and documented wheelchair securement suitable for Elena's rigid wheelchair. Its scheduled pickup is 18:35, within Elena's requested 18:40 deadline. Separately, the 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus, while otherwise equipped, is scheduled after Elena's 18:40 deadline. The airport assistance attendant is logged as covering only baggage claim to curb, not the full route to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy, entities, and deadline unchanged, the evidence spans are complete factual sentences (party size and van capacity), the counterfactual coherently swaps the van's passenger capacity from 4 to 2 seats without conflicting with other stated facts, and neither context reveals a gold label or reasoning instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At fictional Northstar Airport, dispatch logs record Elena's transfer request. Elena's entire airport-to-hotel party consists of 3 travelers. The group carries four suitcases, all blue, and Elena uses a rigid wheelchair and cannot transfer from it or climb steps. Her requested hotel pickup deadline is 18:40. Dispatch records for the 18:35 accessible van list: pickup time 18:35, suitcase capacity of four, documented ramp, and documented wheelchair securement suitable for a rigid wheelchair. The documented passenger capacity listed on the 18:35 accessible van's manifest is 4 seats. A standard sedan is available at 18:25 with capacity for four people and three suitcases but no ramp or securement documented. An accessible minibus is available at 18:50 with sufficient capacity and features. An airport assistance attendant can only help from baggage claim to the curb, not to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena holds Gold hotel membership.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}, {"path": [], "text": "The documented passenger capacity listed on the 18:35 accessible van's manifest is 4 seats."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena's entire airport-to-hotel party consists of 3 travelers.", "negative_left": "Elena's entire airport-to-hotel party consists of 3 travelers.", "negative_right": "The documented passenger capacity listed on the 18:35 accessible van's manifest is 2 seats.", "right": "The documented passenger capacity listed on the 18:35 accessible van's manifest is 4 seats."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-048", "id": "fast-43-diverse-109-048-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At fictional Northstar Airport, dispatch logs record Elena's transfer request. Elena's entire airport-to-hotel party consists of 3 travelers. The group carries four suitcases, all blue, and Elena uses a rigid wheelchair and cannot transfer from it or climb steps. Her requested hotel pickup deadline is 18:40. Dispatch records for the 18:35 accessible van list: pickup time 18:35, suitcase capacity of four, documented ramp, and documented wheelchair securement suitable for a rigid wheelchair. The documented passenger capacity listed on the 18:35 accessible van's manifest is 4 seats. A standard sedan is available at 18:25 with capacity for four people and three suitcases but no ramp or securement documented. An accessible minibus is available at 18:50 with sufficient capacity and features. An airport assistance attendant can only help from baggage claim to the curb, not to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena holds Gold hotel membership."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original policy, entities, and deadline unchanged, the evidence spans are complete factual sentences (party size and van capacity), the counterfactual coherently swaps the van's passenger capacity from 4 to 2 seats without conflicting with other stated facts, and neither context reveals a gold label or reasoning instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At fictional Northstar Airport, dispatch logs record Elena's transfer request. Elena's entire airport-to-hotel party consists of 3 travelers. The group carries four suitcases, all blue, and Elena uses a rigid wheelchair and cannot transfer from it or climb steps. Her requested hotel pickup deadline is 18:40. Dispatch records for the 18:35 accessible van list: pickup time 18:35, suitcase capacity of four, documented ramp, and documented wheelchair securement suitable for a rigid wheelchair. The documented passenger capacity listed on the 18:35 accessible van's manifest is 4 seats. A standard sedan is available at 18:25 with capacity for four people and three suitcases but no ramp or securement documented. An accessible minibus is available at 18:50 with sufficient capacity and features. An airport assistance attendant can only help from baggage claim to the curb, not to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena holds Gold hotel membership.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}, {"path": [], "text": "The documented passenger capacity listed on the 18:35 accessible van's manifest is 4 seats."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena's entire airport-to-hotel party consists of 3 travelers.", "negative_left": "Elena's entire airport-to-hotel party consists of 3 travelers.", "negative_right": "The documented passenger capacity listed on the 18:35 accessible van's manifest is 2 seats.", "right": "The documented passenger capacity listed on the 18:35 accessible van's manifest is 4 seats."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-048", "id": "fast-43-diverse-109-048-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At fictional Northstar Airport, dispatch logs record Elena's transfer request. Elena's entire airport-to-hotel party consists of 3 travelers. The group carries four suitcases, all blue, and Elena uses a rigid wheelchair and cannot transfer from it or climb steps. Her requested hotel pickup deadline is 18:40. Dispatch records for the 18:35 accessible van list: pickup time 18:35, suitcase capacity of four, documented ramp, and documented wheelchair securement suitable for a rigid wheelchair. The documented passenger capacity listed on the 18:35 accessible van's manifest is 2 seats. A standard sedan is available at 18:25 with capacity for four people and three suitcases but no ramp or securement documented. An accessible minibus is available at 18:50 with sufficient capacity and features. An airport assistance attendant can only help from baggage claim to the curb, not to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Elena holds Gold hotel membership."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the deadline, party size, suitcase count, and accessibility policy language verbatim with the original question; the counterfactual only changes the van's passenger capacity, which is a coherent single-fact mutation; evidence spans are two complete factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Dispatcher's case note for Elena at Northstar Airport: Elena’s airport-to-hotel party consists of 3 travelers. She also has four suitcases and a rigid wheelchair, and she cannot transfer from the chair or climb steps. She requested hotel pickup by 18:40. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Three options were logged: a standard sedan departing 18:25 with no ramp; an accessible van departing 18:35, documented with ramp, wheelchair securement, and space for the rigid wheelchair; and an accessible minibus departing 18:50. The manifest for the 18:35 accessible van documents a passenger capacity of 4. The same manifest lists suitcase capacity of four. An airport assistance attendant is also on duty but can only assist from baggage claim to the curb, not provide the full hotel transfer.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena’s airport-to-hotel party consists of 3 travelers."}, {"path": [], "text": "The manifest for the 18:35 accessible van documents a passenger capacity of 4."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena’s airport-to-hotel party consists of 3 travelers.", "negative_left": "Elena’s airport-to-hotel party consists of 3 travelers.", "negative_right": "The manifest for the 18:35 accessible van documents a passenger capacity of 2.", "right": "The manifest for the 18:35 accessible van documents a passenger capacity of 4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-056", "id": "fast-43-diverse-109-056-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Dispatcher's case note for Elena at Northstar Airport: Elena’s airport-to-hotel party consists of 3 travelers. She also has four suitcases and a rigid wheelchair, and she cannot transfer from the chair or climb steps. She requested hotel pickup by 18:40. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Three options were logged: a standard sedan departing 18:25 with no ramp; an accessible van departing 18:35, documented with ramp, wheelchair securement, and space for the rigid wheelchair; and an accessible minibus departing 18:50. The manifest for the 18:35 accessible van documents a passenger capacity of 4. The same manifest lists suitcase capacity of four. An airport assistance attendant is also on duty but can only assist from baggage claim to the curb, not provide the full hotel transfer."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the deadline, party size, suitcase count, and accessibility policy language verbatim with the original question; the counterfactual only changes the van's passenger capacity, which is a coherent single-fact mutation; evidence spans are two complete factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"Dispatcher's case note for Elena at Northstar Airport: Elena’s airport-to-hotel party consists of 3 travelers. She also has four suitcases and a rigid wheelchair, and she cannot transfer from the chair or climb steps. She requested hotel pickup by 18:40. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Three options were logged: a standard sedan departing 18:25 with no ramp; an accessible van departing 18:35, documented with ramp, wheelchair securement, and space for the rigid wheelchair; and an accessible minibus departing 18:50. The manifest for the 18:35 accessible van documents a passenger capacity of 4. The same manifest lists suitcase capacity of four. An airport assistance attendant is also on duty but can only assist from baggage claim to the curb, not provide the full hotel transfer.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "Elena’s airport-to-hotel party consists of 3 travelers."}, {"path": [], "text": "The manifest for the 18:35 accessible van documents a passenger capacity of 4."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "Elena’s airport-to-hotel party consists of 3 travelers.", "negative_left": "Elena’s airport-to-hotel party consists of 3 travelers.", "negative_right": "The manifest for the 18:35 accessible van documents a passenger capacity of 2.", "right": "The manifest for the 18:35 accessible van documents a passenger capacity of 4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-056", "id": "fast-43-diverse-109-056-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "Dispatcher's case note for Elena at Northstar Airport: Elena’s airport-to-hotel party consists of 3 travelers. She also has four suitcases and a rigid wheelchair, and she cannot transfer from the chair or climb steps. She requested hotel pickup by 18:40. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Three options were logged: a standard sedan departing 18:25 with no ramp; an accessible van departing 18:35, documented with ramp, wheelchair securement, and space for the rigid wheelchair; and an accessible minibus departing 18:50. The manifest for the 18:35 accessible van documents a passenger capacity of 2. The same manifest lists suitcase capacity of four. An airport assistance attendant is also on duty but can only assist from baggage claim to the curb, not provide the full hotel transfer."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Governing policy and question object are unchanged; only van capacity observation varies (3 vs 2 seats) which is a permissible case-observation change; evidence spans are two plain factual sentences; the counterfactual (2 seats) is internally consistent with the unchanged 3-traveler party and suitcase facts; no gold answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, staff review Elena's transfer options ahead of her requested 18:40 pickup. The 18:35 accessible van's documented passenger manifest allows for 3 seated riders. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching Elena's four suitcases. The van also carries a documented ramp and wheelchair securement rated for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus is scheduled to arrive after the requested deadline. The airport assistance attendant is limited to helping from baggage claim to the curb and does not provide a complete transfer to the hotel. Elena's suitcases are blue, and her hotel membership is Gold.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's documented passenger manifest allows for 3 seated riders."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's documented passenger manifest allows for 3 seated riders.", "negative_left": "The 18:35 accessible van's documented passenger manifest allows for 2 seated riders.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-061", "id": "fast-43-diverse-109-061-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, staff review Elena's transfer options ahead of her requested 18:40 pickup. The 18:35 accessible van's documented passenger manifest allows for 3 seated riders. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching Elena's four suitcases. The van also carries a documented ramp and wheelchair securement rated for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus is scheduled to arrive after the requested deadline. The airport assistance attendant is limited to helping from baggage claim to the curb and does not provide a complete transfer to the hotel. Elena's suitcases are blue, and her hotel membership is Gold."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Governing policy and question object are unchanged; only van capacity observation varies (3 vs 2 seats) which is a permissible case-observation change; evidence spans are two plain factual sentences; the counterfactual (2 seats) is internally consistent with the unchanged 3-traveler party and suitcase facts; no gold answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, staff review Elena's transfer options ahead of her requested 18:40 pickup. The 18:35 accessible van's documented passenger manifest allows for 3 seated riders. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching Elena's four suitcases. The van also carries a documented ramp and wheelchair securement rated for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus is scheduled to arrive after the requested deadline. The airport assistance attendant is limited to helping from baggage claim to the curb and does not provide a complete transfer to the hotel. Elena's suitcases are blue, and her hotel membership is Gold.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's documented passenger manifest allows for 3 seated riders."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's documented passenger manifest allows for 3 seated riders.", "negative_left": "The 18:35 accessible van's documented passenger manifest allows for 2 seated riders.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-061", "id": "fast-43-diverse-109-061-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, staff review Elena's transfer options ahead of her requested 18:40 pickup. The 18:35 accessible van's documented passenger manifest allows for 2 seated riders. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching Elena's four suitcases. The van also carries a documented ramp and wheelchair securement rated for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus is scheduled to arrive after the requested deadline. The airport assistance attendant is limited to helping from baggage claim to the curb and does not provide a complete transfer to the hotel. Elena's suitcases are blue, and her hotel membership is Gold."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 18:40 deadline, 3-traveler party, four suitcases, ramp/securement requirement, and all vehicle times/capacities except the single flipped van-seat figure, matching the unchanged question object; the counterfactual's single-sentence capacity change (3→2 seats) is a coherent, non-contradictory flip; the two evidence spans are complete factual sentences about manifest capacity and party size, not policy text; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, staff review Elena's requested hotel pickup by 18:40. The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching Elena's four suitcases. Its documentation confirms a ramp and wheelchair securement suitable for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan carries four people and three suitcases but has no documented ramp. The 18:50 accessible minibus has sufficient capacity and accessibility features but arrives after the requested deadline. The airport assistance attendant can only help travelers from baggage claim to the curb and does not provide the complete transfer to the hotel. Elena's suitcases are blue, and her hotel membership is Gold.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 3 seats.", "negative_left": "The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-093", "id": "fast-43-diverse-109-093-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, staff review Elena's requested hotel pickup by 18:40. The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching Elena's four suitcases. Its documentation confirms a ramp and wheelchair securement suitable for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan carries four people and three suitcases but has no documented ramp. The 18:50 accessible minibus has sufficient capacity and accessibility features but arrives after the requested deadline. The airport assistance attendant can only help travelers from baggage claim to the curb and does not provide the complete transfer to the hotel. Elena's suitcases are blue, and her hotel membership is Gold."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 18:40 deadline, 3-traveler party, four suitcases, ramp/securement requirement, and all vehicle times/capacities except the single flipped van-seat figure, matching the unchanged question object; the counterfactual's single-sentence capacity change (3→2 seats) is a coherent, non-contradictory flip; the two evidence spans are complete factual sentences about manifest capacity and party size, not policy text; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, staff review Elena's requested hotel pickup by 18:40. The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching Elena's four suitcases. Its documentation confirms a ramp and wheelchair securement suitable for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan carries four people and three suitcases but has no documented ramp. The 18:50 accessible minibus has sufficient capacity and accessibility features but arrives after the requested deadline. The airport assistance attendant can only help travelers from baggage claim to the curb and does not provide the complete transfer to the hotel. Elena's suitcases are blue, and her hotel membership is Gold.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 3 seats.", "negative_left": "The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-093", "id": "fast-43-diverse-109-093-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, staff review Elena's requested hotel pickup by 18:40. The transfer manifest lists the 18:35 accessible van's documented passenger capacity as 2 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching Elena's four suitcases. Its documentation confirms a ramp and wheelchair securement suitable for Elena's rigid wheelchair. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan carries four people and three suitcases but has no documented ramp. The 18:50 accessible minibus has sufficient capacity and accessibility features but arrives after the requested deadline. The airport assistance attendant can only help travelers from baggage claim to the curb and does not provide the complete transfer to the hotel. Elena's suitcases are blue, and her hotel membership is Gold."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Elena's identity, 18:40 deadline, and the governing policy (deadline, full party/luggage, ramp/securement) intact alongside the unchanged question; the two focus_evidence sentences are complete factual claims, not policy text; the counterfactual makes a single coherent change (van capacity 3\"\"2 seats) without duplicating or contradicting other measurements; neither context reveals a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport curb, dispatch staff review Elena's transfer options before her requested 18:40 pickup. The manifest for the 18:35 accessible van lists its documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching the four suitcases Elena is traveling with. Records confirm the van has documented space for Elena's rigid wheelchair, along with a documented ramp and wheelchair securement suitable for her needs. Its 18:35 pickup time falls before the requested deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus arrives after Elena's deadline. The airport assistance attendant only helps from baggage claim to the curb and does not provide a complete transfer to the hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The manifest for the 18:35 accessible van lists its documented passenger capacity as 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The manifest for the 18:35 accessible van lists its documented passenger capacity as 3 seats.", "negative_left": "The manifest for the 18:35 accessible van lists its documented passenger capacity as 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-101", "id": "fast-43-diverse-109-101-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport curb, dispatch staff review Elena's transfer options before her requested 18:40 pickup. The manifest for the 18:35 accessible van lists its documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching the four suitcases Elena is traveling with. Records confirm the van has documented space for Elena's rigid wheelchair, along with a documented ramp and wheelchair securement suitable for her needs. Its 18:35 pickup time falls before the requested deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus arrives after Elena's deadline. The airport assistance attendant only helps from baggage claim to the curb and does not provide a complete transfer to the hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Elena's identity, 18:40 deadline, and the governing policy (deadline, full party/luggage, ramp/securement) intact alongside the unchanged question; the two focus_evidence sentences are complete factual claims, not policy text; the counterfactual makes a single coherent change (van capacity 3\"\"2 seats) without duplicating or contradicting other measurements; neither context reveals a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport curb, dispatch staff review Elena's transfer options before her requested 18:40 pickup. The manifest for the 18:35 accessible van lists its documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching the four suitcases Elena is traveling with. Records confirm the van has documented space for Elena's rigid wheelchair, along with a documented ramp and wheelchair securement suitable for her needs. Its 18:35 pickup time falls before the requested deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus arrives after Elena's deadline. The airport assistance attendant only helps from baggage claim to the curb and does not provide a complete transfer to the hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The manifest for the 18:35 accessible van lists its documented passenger capacity as 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The manifest for the 18:35 accessible van lists its documented passenger capacity as 3 seats.", "negative_left": "The manifest for the 18:35 accessible van lists its documented passenger capacity as 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-101", "id": "fast-43-diverse-109-101-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport curb, dispatch staff review Elena's transfer options before her requested 18:40 pickup. The manifest for the 18:35 accessible van lists its documented passenger capacity as 2 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The van's documented suitcase capacity is four, matching the four suitcases Elena is traveling with. Records confirm the van has documented space for Elena's rigid wheelchair, along with a documented ramp and wheelchair securement suitable for her needs. Its 18:35 pickup time falls before the requested deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. The 18:25 standard sedan has no documented ramp on file. The 18:50 accessible minibus arrives after Elena's deadline. The airport assistance attendant only helps from baggage claim to the curb and does not provide a complete transfer to the hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy wording, entities, times and paths, differ only in the van's stated seat capacity, and the evidence spans are two complete factual sentences with no leaked labels or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, staff are arranging Elena's transfer to her hotel. Elena requires pickup by 18:40 due to her rigid wheelchair and inability to climb steps or transfer from the chair. The 18:35 accessible van's manifest lists a documented passenger capacity of 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The same manifest records a suitcase capacity of four, matching the four suitcases her party is carrying. The van's documentation also confirms a ramp and wheelchair securement suitable for a rigid wheelchair, and its scheduled pickup time of 18:35 falls before Elena's deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Separately, the 18:25 standard sedan has no documented ramp, the 18:50 accessible minibus arrives after 18:40, and the airport assistance attendant only helps from baggage claim to the curb, not to the hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's manifest lists a documented passenger capacity of 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's manifest lists a documented passenger capacity of 3 seats.", "negative_left": "The 18:35 accessible van's manifest lists a documented passenger capacity of 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-106", "id": "fast-43-diverse-109-106-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, staff are arranging Elena's transfer to her hotel. Elena requires pickup by 18:40 due to her rigid wheelchair and inability to climb steps or transfer from the chair. The 18:35 accessible van's manifest lists a documented passenger capacity of 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The same manifest records a suitcase capacity of four, matching the four suitcases her party is carrying. The van's documentation also confirms a ramp and wheelchair securement suitable for a rigid wheelchair, and its scheduled pickup time of 18:35 falls before Elena's deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Separately, the 18:25 standard sedan has no documented ramp, the 18:50 accessible minibus arrives after 18:40, and the airport assistance attendant only helps from baggage claim to the curb, not to the hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy wording, entities, times and paths, differ only in the van's stated seat capacity, and the evidence spans are two complete factual sentences with no leaked labels or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, staff are arranging Elena's transfer to her hotel. Elena requires pickup by 18:40 due to her rigid wheelchair and inability to climb steps or transfer from the chair. The 18:35 accessible van's manifest lists a documented passenger capacity of 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The same manifest records a suitcase capacity of four, matching the four suitcases her party is carrying. The van's documentation also confirms a ramp and wheelchair securement suitable for a rigid wheelchair, and its scheduled pickup time of 18:35 falls before Elena's deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Separately, the 18:25 standard sedan has no documented ramp, the 18:50 accessible minibus arrives after 18:40, and the airport assistance attendant only helps from baggage claim to the curb, not to the hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's manifest lists a documented passenger capacity of 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's manifest lists a documented passenger capacity of 3 seats.", "negative_left": "The 18:35 accessible van's manifest lists a documented passenger capacity of 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-106", "id": "fast-43-diverse-109-106-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, staff are arranging Elena's transfer to her hotel. Elena requires pickup by 18:40 due to her rigid wheelchair and inability to climb steps or transfer from the chair. The 18:35 accessible van's manifest lists a documented passenger capacity of 2 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The same manifest records a suitcase capacity of four, matching the four suitcases her party is carrying. The van's documentation also confirms a ramp and wheelchair securement suitable for a rigid wheelchair, and its scheduled pickup time of 18:35 falls before Elena's deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Separately, the 18:25 standard sedan has no documented ramp, the 18:50 accessible minibus arrives after 18:40, and the airport assistance attendant only helps from baggage claim to the curb, not to the hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Elena's identity, party size of 3, deadline, and all policy criteria intact while only varying the van's documented seat capacity, which is a permissible case-observation change; the two evidence sentences are complete factual statements, the counterfactual's single altered sentence (2 seats) is internally consistent with the rest of the unchanged context, and neither context contains an answer, rule table, or instructional leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, staff logged the details of Elena's requested hotel transfer. The dispatch log records the 18:35 accessible van's documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The same van's documented suitcase capacity is four, matching the four suitcases Elena is bringing. The van's log also confirms a documented ramp and wheelchair securement suitable for her rigid wheelchair, and its 18:35 pickup time is within her requested 18:40 deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Separately, the 18:25 standard sedan has no documented ramp, the 18:50 accessible minibus arrives after the deadline, and the airport assistance attendant helps only from baggage claim to the curb, not the full transfer to the hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The dispatch log records the 18:35 accessible van's documented passenger capacity as 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The dispatch log records the 18:35 accessible van's documented passenger capacity as 3 seats.", "negative_left": "The dispatch log records the 18:35 accessible van's documented passenger capacity as 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-125", "id": "fast-43-diverse-109-125-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, staff logged the details of Elena's requested hotel transfer. The dispatch log records the 18:35 accessible van's documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The same van's documented suitcase capacity is four, matching the four suitcases Elena is bringing. The van's log also confirms a documented ramp and wheelchair securement suitable for her rigid wheelchair, and its 18:35 pickup time is within her requested 18:40 deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Separately, the 18:25 standard sedan has no documented ramp, the 18:50 accessible minibus arrives after the deadline, and the airport assistance attendant helps only from baggage claim to the curb, not the full transfer to the hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Elena's identity, party size of 3, deadline, and all policy criteria intact while only varying the van's documented seat capacity, which is a permissible case-observation change; the two evidence sentences are complete factual statements, the counterfactual's single altered sentence (2 seats) is internally consistent with the rest of the unchanged context, and neither context contains an answer, rule table, or instructional leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport transfer desk, staff logged the details of Elena's requested hotel transfer. The dispatch log records the 18:35 accessible van's documented passenger capacity as 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The same van's documented suitcase capacity is four, matching the four suitcases Elena is bringing. The van's log also confirms a documented ramp and wheelchair securement suitable for her rigid wheelchair, and its 18:35 pickup time is within her requested 18:40 deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Separately, the 18:25 standard sedan has no documented ramp, the 18:50 accessible minibus arrives after the deadline, and the airport assistance attendant helps only from baggage claim to the curb, not the full transfer to the hotel.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The dispatch log records the 18:35 accessible van's documented passenger capacity as 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The dispatch log records the 18:35 accessible van's documented passenger capacity as 3 seats.", "negative_left": "The dispatch log records the 18:35 accessible van's documented passenger capacity as 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-125", "id": "fast-43-diverse-109-125-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport transfer desk, staff logged the details of Elena's requested hotel transfer. The dispatch log records the 18:35 accessible van's documented passenger capacity as 2 seats. Elena's entire airport-to-hotel party consists of 3 travelers. The same van's documented suitcase capacity is four, matching the four suitcases Elena is bringing. The van's log also confirms a documented ramp and wheelchair securement suitable for her rigid wheelchair, and its 18:35 pickup time is within her requested 18:40 deadline. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement. Separately, the 18:25 standard sedan has no documented ramp, the 18:50 accessible minibus arrives after the deadline, and the airport assistance attendant helps only from baggage claim to the curb, not the full transfer to the hotel."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy sentence and all question-bound entities, times, and party size; the two focus evidence spans are complete factual statements; the counterfactual only alters the van's seat capacity to 2, which is internally consistent and does not create duplicate or contradictory measurements; no gold answer, rule table, or instruction text is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport ground transport counter, dispatch records are being reviewed for Elena's hotel transfer. Elena has requested pickup by 18:40, and staff are checking each vehicle in the queue against her needs. The 18:35 accessible van's transport manifest lists a documented passenger capacity of 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers, including herself. The same van's manifest also lists a documented suitcase capacity of four, matching the four suitcases she is bringing. The manifest further notes a documented ramp and wheelchair securement suitable for her rigid wheelchair, and the van is scheduled to arrive at 18:35, ahead of her deadline. Meanwhile, the 18:25 standard sedan's manifest shows no ramp listed. The 18:50 accessible minibus is noted as arriving after 18:40. The airport assistance attendant's duties are logged as ending at the curb, not extending to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's transport manifest lists a documented passenger capacity of 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers, including herself."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's transport manifest lists a documented passenger capacity of 3 seats.", "negative_left": "The 18:35 accessible van's transport manifest lists a documented passenger capacity of 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers, including herself.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers, including herself."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-145", "id": "fast-43-diverse-109-145-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport ground transport counter, dispatch records are being reviewed for Elena's hotel transfer. Elena has requested pickup by 18:40, and staff are checking each vehicle in the queue against her needs. The 18:35 accessible van's transport manifest lists a documented passenger capacity of 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers, including herself. The same van's manifest also lists a documented suitcase capacity of four, matching the four suitcases she is bringing. The manifest further notes a documented ramp and wheelchair securement suitable for her rigid wheelchair, and the van is scheduled to arrive at 18:35, ahead of her deadline. Meanwhile, the 18:25 standard sedan's manifest shows no ramp listed. The 18:50 accessible minibus is noted as arriving after 18:40. The airport assistance attendant's duties are logged as ending at the curb, not extending to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy sentence and all question-bound entities, times, and party size; the two focus evidence spans are complete factual statements; the counterfactual only alters the van's seat capacity to 2, which is internally consistent and does not create duplicate or contradictory measurements; no gold answer, rule table, or instruction text is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At the Northstar Airport ground transport counter, dispatch records are being reviewed for Elena's hotel transfer. Elena has requested pickup by 18:40, and staff are checking each vehicle in the queue against her needs. The 18:35 accessible van's transport manifest lists a documented passenger capacity of 3 seats. Elena's entire airport-to-hotel party consists of 3 travelers, including herself. The same van's manifest also lists a documented suitcase capacity of four, matching the four suitcases she is bringing. The manifest further notes a documented ramp and wheelchair securement suitable for her rigid wheelchair, and the van is scheduled to arrive at 18:35, ahead of her deadline. Meanwhile, the 18:25 standard sedan's manifest shows no ramp listed. The 18:50 accessible minibus is noted as arriving after 18:40. The airport assistance attendant's duties are logged as ending at the curb, not extending to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's transport manifest lists a documented passenger capacity of 3 seats."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of 3 travelers, including herself."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's transport manifest lists a documented passenger capacity of 3 seats.", "negative_left": "The 18:35 accessible van's transport manifest lists a documented passenger capacity of 2 seats.", "negative_right": "Elena's entire airport-to-hotel party consists of 3 travelers, including herself.", "right": "Elena's entire airport-to-hotel party consists of 3 travelers, including herself."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-145", "id": "fast-43-diverse-109-145-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At the Northstar Airport ground transport counter, dispatch records are being reviewed for Elena's hotel transfer. Elena has requested pickup by 18:40, and staff are checking each vehicle in the queue against her needs. The 18:35 accessible van's transport manifest lists a documented passenger capacity of 2 seats. Elena's entire airport-to-hotel party consists of 3 travelers, including herself. The same van's manifest also lists a documented suitcase capacity of four, matching the four suitcases she is bringing. The manifest further notes a documented ramp and wheelchair securement suitable for her rigid wheelchair, and the van is scheduled to arrive at 18:35, ahead of her deadline. Meanwhile, the 18:25 standard sedan's manifest shows no ramp listed. The 18:50 accessible minibus is noted as arriving after 18:40. The airport assistance attendant's duties are logged as ending at the curb, not extending to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy statement and question bindings; the two evidence sentences are factual (capacity and party size); the counterfactual changes only the van's passenger capacity from 3 to 2, remaining internally consistent and not contradicting other facts; neither context reveals a gold answer or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At Northstar Airport's transfer desk, staff logged details for Elena's hotel transfer request. She requested pickup by 18:40 and travels with a rigid wheelchair that requires ramp access and securement, as she cannot transfer from the chair or climb steps. Dispatch records show three vehicle options: a standard sedan at 18:25 with capacity for four people and three suitcases but no documented ramp; an accessible van at 18:35 with a documented ramp and wheelchair securement suitable for Elena's chair; and an accessible minibus at 18:50 with sufficient capacity and accessibility features but arriving after the deadline. The van's registered service manifest lists its documented suitcase capacity as four, matching Elena's four suitcases. The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 3. Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers. An airport assistance attendant is also on duty but only helps passengers from baggage claim to the curb, not to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 3."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 3.", "negative_left": "The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 2.", "negative_right": "Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers.", "right": "Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-177", "id": "fast-43-diverse-109-177-base", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At Northstar Airport's transfer desk, staff logged details for Elena's hotel transfer request. She requested pickup by 18:40 and travels with a rigid wheelchair that requires ramp access and securement, as she cannot transfer from the chair or climb steps. Dispatch records show three vehicle options: a standard sedan at 18:25 with capacity for four people and three suitcases but no documented ramp; an accessible van at 18:35 with a documented ramp and wheelchair securement suitable for Elena's chair; and an accessible minibus at 18:50 with sufficient capacity and accessibility features but arriving after the deadline. The van's registered service manifest lists its documented suitcase capacity as four, matching Elena's four suitcases. The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 3. Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers. An airport assistance attendant is also on duty but only helps passengers from baggage claim to the curb, not to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_van_dispatch"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy statement and question bindings; the two evidence sentences are factual (capacity and party size); the counterfactual changes only the van's passenger capacity from 3 to 2, remaining internally consistent and not contradicting other facts; neither context reveals a gold answer or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "refuted", "A8": "refuted", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A1 is a factual capacity relation rather than a policy conclusion. The base and counter assignments differ only on A1 and are jointly realizable as synthetic scenarios. The state-derived policy evidence preserves the governing suitability rule needed alongside the unchanged question; criteria and instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the van meets party, suitcase, wheelchair-capacity, ramp, securement, and deadline requirements. They also exclude the sedan through its missing ramp, the minibus through its missed deadline, and the attendant through its incomplete transfer service.", "rule_index": 0, "sound": true}, {"reason": "Each substantive option is shown to fail a mandatory requirement: the van lacks full-party capacity, the sedan lacks a ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer. This is sufficient for none_of_above.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The documented passenger capacity of the 18:35 accessible van is at least the number of travelers in Elena’s entire airport-to-hotel party."}, {"id": "A2", "statement": "The documented suitcase capacity of the 18:35 accessible van is at least the four suitcases in Elena’s airport-to-hotel transfer."}, {"id": "A3", "statement": "The 18:35 accessible van has documented capacity for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A4", "statement": "The 18:35 accessible van has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A5", "statement": "The 18:35 accessible van has documented wheelchair securement for Elena’s rigid wheelchair during the airport-to-hotel transfer."}, {"id": "A6", "statement": "The 18:35 pickup time of the accessible van is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A7", "statement": "The 18:25 standard sedan has a documented ramp for Elena’s airport-to-hotel transfer."}, {"id": "A8", "statement": "The 18:50 pickup time of the accessible minibus is no later than Elena’s requested 18:40 hotel-pickup deadline."}, {"id": "A9", "statement": "The airport assistance attendant’s service provides Elena’s complete transfer from Northstar Airport to her hotel."}], "base_state_json": "\"At Northstar Airport's transfer desk, staff logged details for Elena's hotel transfer request. She requested pickup by 18:40 and travels with a rigid wheelchair that requires ramp access and securement, as she cannot transfer from the chair or climb steps. Dispatch records show three vehicle options: a standard sedan at 18:25 with capacity for four people and three suitcases but no documented ramp; an accessible van at 18:35 with a documented ramp and wheelchair securement suitable for Elena's chair; and an accessible minibus at 18:50 with sufficient capacity and accessibility features but arriving after the deadline. The van's registered service manifest lists its documented suitcase capacity as four, matching Elena's four suitcases. The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 3. Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers. An airport assistance attendant is also on duty but only helps passengers from baggage claim to the curb, not to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": [], "text": "The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 3."}, {"path": [], "text": "Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers."}], "policy_evidence": [{"path": [], "text": "Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}], "rules": [{"justification": "The van meets the deadline and every party, luggage, wheelchair-capacity, ramp, and securement requirement, while each competing substantive route is excluded by a mandatory failure.", "target": "accessible_van_dispatch", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}, {"justification": "The van fails the mandatory full-party capacity requirement, the sedan lacks a required ramp, the minibus misses the deadline, and the attendant does not provide the complete airport-to-hotel transfer.", "target": "none_of_above", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "refuted"}, {"atom_id": "A8", "state": "refuted"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 3.", "negative_left": "The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 2.", "negative_right": "Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers.", "right": "Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-109-177", "id": "fast-43-diverse-109-177-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_minibus_dispatch": "Assign the 18:50 accessible minibus; choose only if it satisfies the passenger’s pickup deadline in addition to capacity and accessibility requirements.", "accessible_van_dispatch": "Assign the 18:35 accessible van; choose only if its documented capacity covers all three travelers and four suitcases as well as the wheelchair.", "airport_assistance_route": "Route Elena solely to the airport assistance attendant; choose only if that service itself provides the complete airport-to-hotel transfer.", "none_of_above": "Choose when every substantive routing option fails at least one mandatory suitability requirement.", "standard_sedan_dispatch": "Assign the 18:25 standard sedan; choose only if it meets the full party, luggage, ramp, securement, and timing requirements."}, "instructions": "Choose the single routing decision that provides a suitable airport-to-hotel transfer under every stated policy requirement. Treat suitcase color and hotel membership as distractors. If no substantive option satisfies timing, party, luggage, and accessibility requirements, choose none_of_above.", "type": "choice"}}, "state": "At Northstar Airport's transfer desk, staff logged details for Elena's hotel transfer request. She requested pickup by 18:40 and travels with a rigid wheelchair that requires ramp access and securement, as she cannot transfer from the chair or climb steps. Dispatch records show three vehicle options: a standard sedan at 18:25 with capacity for four people and three suitcases but no documented ramp; an accessible van at 18:35 with a documented ramp and wheelchair securement suitable for Elena's chair; and an accessible minibus at 18:50 with sufficient capacity and accessibility features but arriving after the deadline. The van's registered service manifest lists its documented suitcase capacity as four, matching Elena's four suitcases. The 18:35 accessible van's registered service manifest lists its documented passenger capacity as 2. Elena's entire airport-to-hotel party consists of the passenger herself plus two accompanying travelers. An airport assistance attendant is also on duty but only helps passengers from baggage claim to the curb, not to the hotel. Policy says a suitable transfer must meet the deadline, carry the entire party and luggage, and have documented ramp and wheelchair securement."}, "method": "c2d", "provenance": {"source_id": "diverse-109", "source_is_synthetic": true, "source_sha256": "aab2a769b27620ed3bb44f2a03e6c374a16f041471eaedded469d0aeab43b492", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (fold requirement, ramp/lift, securement, capacity, curb-only assistance, ranking rule) untouched, and the unchanged questions object supplies the scoring criteria; entity/time bindings (Access A/B, 18:40-19:00 times) are identical across both contexts and match the question criteria; the two focus evidence spans are complete factual sentences about bag counts, not policy or instructions; the counterfactual merely alters Access B's stated luggage rating from 4 to 2 bags, a plausible varying observation that doesn't contradict any other stated fact; neither context contains a gold answer, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time. Case note: The arriving passenger must remain in her rigid power chair and cannot transfer to a standard van; the mobility device cannot be folded, so standard dispatch is unavailable. Road transport beyond the curb is required to complete her trip. Three passengers are traveling together. The arriving party has checked 3 bags requiring transport. Access A has a documented ramp, four occupied-chair securements exceeding the party's requirement, capacity for four passengers, capacity for six bags, and a listed pickup after the 18:40 curb-ready time. Access B has a documented lift, two occupied-chair securements meeting the party's requirement, capacity for four passengers, and a listed pickup after the 18:40 curb-ready time that occurs earlier than Access A's pickup. Access B's vehicle is equipped with a luggage compartment rated for 4 bags. The 19:00 hotel shuttle's passenger and luggage limits are insufficient for the full party.\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["context"], "text": "The arriving party has checked 3 bags requiring transport."}, {"path": ["context"], "text": "Access B's vehicle is equipped with a luggage compartment rated for 4 bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "The arriving party has checked 3 bags requiring transport.", "negative_left": "The arriving party has checked 3 bags requiring transport.", "negative_right": "Access B's vehicle is equipped with a luggage compartment rated for 2 bags.", "right": "Access B's vehicle is equipped with a luggage compartment rated for 4 bags."}, "verifier_independent_model": false}, "family": "fast-43-diverse-110-014", "id": "fast-43-diverse-110-014-base", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time. Case note: The arriving passenger must remain in her rigid power chair and cannot transfer to a standard van; the mobility device cannot be folded, so standard dispatch is unavailable. Road transport beyond the curb is required to complete her trip. Three passengers are traveling together. The arriving party has checked 3 bags requiring transport. Access A has a documented ramp, four occupied-chair securements exceeding the party's requirement, capacity for four passengers, capacity for six bags, and a listed pickup after the 18:40 curb-ready time. Access B has a documented lift, two occupied-chair securements meeting the party's requirement, capacity for four passengers, and a listed pickup after the 18:40 curb-ready time that occurs earlier than Access A's pickup. Access B's vehicle is equipped with a luggage compartment rated for 4 bags. The 19:00 hotel shuttle's passenger and luggage limits are insufficient for the full party."}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_b"}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (fold requirement, ramp/lift, securement, capacity, curb-only assistance, ranking rule) untouched, and the unchanged questions object supplies the scoring criteria; entity/time bindings (Access A/B, 18:40-19:00 times) are identical across both contexts and match the question criteria; the two focus evidence spans are complete factual sentences about bag counts, not policy or instructions; the counterfactual merely alters Access B's stated luggage rating from 4 to 2 bags, a plausible varying observation that doesn't contradict any other stated fact; neither context contains a gold answer, rule table, proposition ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "full_context_fact_states": {"base": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "supported", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "counterfactual": {"access_a_after_curb_ready": "supported", "access_a_luggage_capacity_sufficient": "supported", "access_a_passenger_capacity_sufficient": "supported", "access_a_ramp_documented": "supported", "access_a_securements_sufficient": "supported", "access_b_after_curb_ready": "supported", "access_b_earlier_than_access_a": "supported", "access_b_lift_documented": "supported", "access_b_luggage_capacity_sufficient": "refuted", "access_b_passenger_capacity_sufficient": "supported", "access_b_securements_sufficient": "supported", "hotel_shuttle_luggage_capacity_sufficient": "refuted", "hotel_shuttle_passenger_capacity_sufficient": "refuted", "road_transfer_needed": "supported", "standard_van_transfer_possible": "refuted"}, "remove_left": {"access_b_luggage_capacity_sufficient": "unknown"}, "remove_right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"access_b_luggage_capacity_sufficient": "unknown"}, "negative_pair": {"access_b_luggage_capacity_sufficient": "refuted"}, "negative_sentence": {"access_b_luggage_capacity_sufficient": "unknown"}, "positive_pair": {"access_b_luggage_capacity_sufficient": "supported"}, "right": {"access_b_luggage_capacity_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship rather than a bundled policy conclusion. The focus is the factual relation between Access B's luggage capacity and the party's bag count. Base and counter assignments differ only on that focus and are realizable by varying Access B's luggage capacity or the applicable bag-count relation while retaining all other facts. Policy evidence correctly preserves the substantive state-originating dispatch, eligibility, assistance-scope, and ranking rules; question-originating criteria need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction makes Access B eligible on lift, securements, passenger capacity, luggage capacity, and post-curb-ready timing. Access A is also eligible but later. The standard van is excluded by inability to transfer, the hotel shuttle by insufficient capacities, and assistance-only routing by the need for road transport. Thus Access B is the earliest eligible suitable vehicle.", "rule_index": 0, "sound": true}, {"reason": "Access B is ineligible because its luggage capacity is explicitly refuted. Access A satisfies every stated accessible-dispatch and timing requirement. The standard van, hotel shuttle, and assistance-only routing are each excluded by sufficient conditions, so Access A is the highest-ranked remaining eligible choice.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "road_transfer_needed", "statement": "Road transport beyond the curb is needed to complete the arriving party's request."}, {"id": "standard_van_transfer_possible", "statement": "The arriving passenger can transfer out of her mobility device for the listed 18:41 standard van."}, {"id": "access_a_ramp_documented", "statement": "Access A has a documented ramp for the arriving party's road transfer."}, {"id": "access_a_securements_sufficient", "statement": "Access A's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_a_passenger_capacity_sufficient", "statement": "Access A's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_a_luggage_capacity_sufficient", "statement": "Access A's luggage capacity is at least the arriving party's bag count."}, {"id": "access_a_after_curb_ready", "statement": "Access A's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_lift_documented", "statement": "Access B has a documented lift for the arriving party's road transfer."}, {"id": "access_b_securements_sufficient", "statement": "Access B's occupied-chair securement capacity is at least the arriving party's occupied-chair securement requirement."}, {"id": "access_b_passenger_capacity_sufficient", "statement": "Access B's passenger capacity is at least the arriving party's passenger count."}, {"id": "access_b_luggage_capacity_sufficient", "statement": "Access B's luggage capacity is at least the arriving party's bag count."}, {"id": "access_b_after_curb_ready", "statement": "Access B's listed pickup occurs after the arriving party's 18:40 curb-ready time."}, {"id": "access_b_earlier_than_access_a", "statement": "Access B's listed pickup occurs earlier than Access A's listed pickup."}, {"id": "hotel_shuttle_passenger_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's passenger capacity is at least the arriving party's passenger count."}, {"id": "hotel_shuttle_luggage_capacity_sufficient", "statement": "The listed 19:00 hotel shuttle's luggage capacity is at least the arriving party's bag count."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time. Case note: The arriving passenger must remain in her rigid power chair and cannot transfer to a standard van; the mobility device cannot be folded, so standard dispatch is unavailable. Road transport beyond the curb is required to complete her trip. Three passengers are traveling together. The arriving party has checked 3 bags requiring transport. Access A has a documented ramp, four occupied-chair securements exceeding the party's requirement, capacity for four passengers, capacity for six bags, and a listed pickup after the 18:40 curb-ready time. Access B has a documented lift, two occupied-chair securements meeting the party's requirement, capacity for four passengers, and a listed pickup after the 18:40 curb-ready time that occurs earlier than Access A's pickup. Access B's vehicle is equipped with a luggage compartment rated for 4 bags. The 19:00 hotel shuttle's passenger and luggage limits are insufficient for the full party.\"}", "base_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "counter_states": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}], "focus_atom": "access_b_luggage_capacity_sufficient", "focus_evidence": [{"path": ["context"], "text": "The arriving party has checked 3 bags requiring transport."}, {"path": ["context"], "text": "Access B's vehicle is equipped with a luggage compartment rated for 4 bags."}], "policy_evidence": [{"path": ["context"], "text": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time."}], "rules": [{"justification": "Both accessible vehicles satisfy the party, but Access B is the earlier post-curb-ready pickup. The standard van, hotel shuttle, and assistance-only routing are excluded by explicit mandatory requirements.", "target": "accessible_access_b", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}, {"justification": "Access B is ineligible because its luggage capacity is below the party's bag count. Access A satisfies every accessible-dispatch requirement, while the standard van, hotel shuttle, and assistance-only routing are excluded.", "target": "accessible_access_a", "when": [{"atom_id": "road_transfer_needed", "state": "supported"}, {"atom_id": "standard_van_transfer_possible", "state": "refuted"}, {"atom_id": "access_a_ramp_documented", "state": "supported"}, {"atom_id": "access_a_securements_sufficient", "state": "supported"}, {"atom_id": "access_a_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_luggage_capacity_sufficient", "state": "supported"}, {"atom_id": "access_a_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_lift_documented", "state": "supported"}, {"atom_id": "access_b_securements_sufficient", "state": "supported"}, {"atom_id": "access_b_passenger_capacity_sufficient", "state": "supported"}, {"atom_id": "access_b_luggage_capacity_sufficient", "state": "refuted"}, {"atom_id": "access_b_after_curb_ready", "state": "supported"}, {"atom_id": "access_b_earlier_than_access_a", "state": "supported"}, {"atom_id": "hotel_shuttle_passenger_capacity_sufficient", "state": "refuted"}, {"atom_id": "hotel_shuttle_luggage_capacity_sufficient", "state": "refuted"}]}]}, "verified_pair": {"left": "The arriving party has checked 3 bags requiring transport.", "negative_left": "The arriving party has checked 3 bags requiring transport.", "negative_right": "Access B's vehicle is equipped with a luggage compartment rated for 2 bags.", "right": "Access B's vehicle is equipped with a luggage compartment rated for 4 bags."}, "verifier_independent_model": false}, "family": "fast-43-diverse-110-014", "id": "fast-43-diverse-110-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"accessible_access_a": "Route to accessible-vehicle dispatch for Access A at 18:50; choose if it is suitable and no equally suitable eligible vehicle has an earlier pickup.", "accessible_access_b": "Route to accessible-vehicle dispatch for Access B at 18:42; choose if its lift, securements, passenger capacity, and luggage capacity all satisfy the party and it is the earliest suitable vehicle.", "assistance_staff_only": "Assign only the airport assistance attendant; choose only if aircraft-to-curb help completes the request and no road transfer is needed.", "hotel_shuttle": "Send the 19:00 hotel shuttle; choose only if its passenger and luggage limits accommodate the complete party despite any accessibility fit.", "none_of_above": "Choose only if neither listed vehicle nor assistance routing can satisfy all mandatory requirements under the policy.", "standard_van_dispatch": "Route to standard dispatch for the 18:41 van; choose only if the passenger can transfer out and the mobility device can be folded."}, "instructions": "Interpret the paraphrased passenger statement by its meaning, verify every capacity and accessibility requirement, and select exactly one option. A vehicle is eligible only if all requirements are met; among eligible vehicles, choose the earliest pickup after 18:40.", "type": "choice"}}, "state": {"context": "At fictional Northstar Airport, standard dispatch may send vehicles only when a passenger can leave a mobility device and it can be folded. Accessible dispatch requires a documented ramp or lift, enough occupied-chair securements, and sufficient passenger and luggage capacity. Airport assistance covers aircraft-to-curb movement, not road transport. Eligible choices rank first by accessibility fit, then earliest pickup after curb-ready time. Case note: The arriving passenger must remain in her rigid power chair and cannot transfer to a standard van; the mobility device cannot be folded, so standard dispatch is unavailable. Road transport beyond the curb is required to complete her trip. Three passengers are traveling together. The arriving party has checked 3 bags requiring transport. Access A has a documented ramp, four occupied-chair securements exceeding the party's requirement, capacity for four passengers, capacity for six bags, and a listed pickup after the 18:40 curb-ready time. Access B has a documented lift, two occupied-chair securements meeting the party's requirement, capacity for four passengers, and a listed pickup after the 18:40 curb-ready time that occurs earlier than Access A's pickup. Access B's vehicle is equipped with a luggage compartment rated for 2 bags. The 19:00 hotel shuttle's passenger and luggage limits are insufficient for the full party."}}, "method": "c2d", "provenance": {"source_id": "diverse-110", "source_is_synthetic": true, "source_sha256": "697b61fbbd7b076472a1395516173d3a6cf552de37df914084d74539c483800e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accessible_access_a"}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same accessibility/capacity/ranking policy and unchanged question; only LiftVan B's pickup time changes to 2:55 PM, which still respects the 14:30 ready-time floor and inequality constraint, giving a coherent alternative ranking without contradictions or leaked answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of her wheelchair for the airport-to-hotel transfer, and her party is confirmed ready at 14:30. Standard Sedan S has four seats and luggage space but no wheelchair ramp and no wheelchair securement. Accessible-vehicle driver Jo confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has an identical ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both vans' pickup times are no earlier than the party's 14:30 ready time, and the two vans' pickup times are not equal to each other. LiftVan A's scheduled pickup time is 3:10 PM. LiftVan B's scheduled pickup time is 3:25 PM. The only vehicles under consideration for Mira's transfer are Standard Sedan S, LiftVan A, and LiftVan B. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's scheduled pickup time is 3:10 PM."}, {"path": [], "text": "LiftVan B's scheduled pickup time is 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's scheduled pickup time is 3:10 PM.", "negative_left": "LiftVan A's scheduled pickup time is 3:10 PM.", "negative_right": "LiftVan B's scheduled pickup time is 2:55 PM.", "right": "LiftVan B's scheduled pickup time is 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-005", "id": "fast-43-diverse-112-005-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of her wheelchair for the airport-to-hotel transfer, and her party is confirmed ready at 14:30. Standard Sedan S has four seats and luggage space but no wheelchair ramp and no wheelchair securement. Accessible-vehicle driver Jo confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has an identical ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both vans' pickup times are no earlier than the party's 14:30 ready time, and the two vans' pickup times are not equal to each other. LiftVan A's scheduled pickup time is 3:10 PM. LiftVan B's scheduled pickup time is 3:25 PM. The only vehicles under consideration for Mira's transfer are Standard Sedan S, LiftVan A, and LiftVan B. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same accessibility/capacity/ranking policy and unchanged question; only LiftVan B's pickup time changes to 2:55 PM, which still respects the 14:30 ready-time floor and inequality constraint, giving a coherent alternative ranking without contradictions or leaked answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of her wheelchair for the airport-to-hotel transfer, and her party is confirmed ready at 14:30. Standard Sedan S has four seats and luggage space but no wheelchair ramp and no wheelchair securement. Accessible-vehicle driver Jo confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has an identical ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both vans' pickup times are no earlier than the party's 14:30 ready time, and the two vans' pickup times are not equal to each other. LiftVan A's scheduled pickup time is 3:10 PM. LiftVan B's scheduled pickup time is 3:25 PM. The only vehicles under consideration for Mira's transfer are Standard Sedan S, LiftVan A, and LiftVan B. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's scheduled pickup time is 3:10 PM."}, {"path": [], "text": "LiftVan B's scheduled pickup time is 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's scheduled pickup time is 3:10 PM.", "negative_left": "LiftVan A's scheduled pickup time is 3:10 PM.", "negative_right": "LiftVan B's scheduled pickup time is 2:55 PM.", "right": "LiftVan B's scheduled pickup time is 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-005", "id": "fast-43-diverse-112-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of her wheelchair for the airport-to-hotel transfer, and her party is confirmed ready at 14:30. Standard Sedan S has four seats and luggage space but no wheelchair ramp and no wheelchair securement. Accessible-vehicle driver Jo confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has an identical ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both vans' pickup times are no earlier than the party's 14:30 ready time, and the two vans' pickup times are not equal to each other. LiftVan A's scheduled pickup time is 3:10 PM. LiftVan B's scheduled pickup time is 2:55 PM. The only vehicles under consideration for Mira's transfer are Standard Sedan S, LiftVan A, and LiftVan B. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, ready time, and dispatch/ranking policy exactly as in the unaltered question; the only edit is LiftVan B's pickup time (3:25 PM -> 2:55 PM), a single coherent factual change still consistent with the 14:30 ready-time constraint; the two focus-evidence sentences are complete factual statements (not policy text) matching the base context, and neither context contains an answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair, and her party will be ready for pickup at 14:30. Three vehicles are under consideration for the airport-to-hotel transfer: Standard Sedan S, LiftVan A, and LiftVan B. Standard Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A has a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases. LiftVan A is scheduled for pickup at 3:10 PM. LiftVan B is scheduled for pickup at 3:25 PM. Both vans' pickup times are no earlier than the party's 14:30 ready time. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A is scheduled for pickup at 3:10 PM."}, {"path": [], "text": "LiftVan B is scheduled for pickup at 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A is scheduled for pickup at 3:10 PM.", "negative_left": "LiftVan A is scheduled for pickup at 3:10 PM.", "negative_right": "LiftVan B is scheduled for pickup at 2:55 PM.", "right": "LiftVan B is scheduled for pickup at 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-012", "id": "fast-43-diverse-112-012-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair, and her party will be ready for pickup at 14:30. Three vehicles are under consideration for the airport-to-hotel transfer: Standard Sedan S, LiftVan A, and LiftVan B. Standard Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A has a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases. LiftVan A is scheduled for pickup at 3:10 PM. LiftVan B is scheduled for pickup at 3:25 PM. Both vans' pickup times are no earlier than the party's 14:30 ready time. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, ready time, and dispatch/ranking policy exactly as in the unaltered question; the only edit is LiftVan B's pickup time (3:25 PM -> 2:55 PM), a single coherent factual change still consistent with the 14:30 ready-time constraint; the two focus-evidence sentences are complete factual statements (not policy text) matching the base context, and neither context contains an answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair, and her party will be ready for pickup at 14:30. Three vehicles are under consideration for the airport-to-hotel transfer: Standard Sedan S, LiftVan A, and LiftVan B. Standard Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A has a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases. LiftVan A is scheduled for pickup at 3:10 PM. LiftVan B is scheduled for pickup at 3:25 PM. Both vans' pickup times are no earlier than the party's 14:30 ready time. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A is scheduled for pickup at 3:10 PM."}, {"path": [], "text": "LiftVan B is scheduled for pickup at 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A is scheduled for pickup at 3:10 PM.", "negative_left": "LiftVan A is scheduled for pickup at 3:10 PM.", "negative_right": "LiftVan B is scheduled for pickup at 2:55 PM.", "right": "LiftVan B is scheduled for pickup at 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-012", "id": "fast-43-diverse-112-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair, and her party will be ready for pickup at 14:30. Three vehicles are under consideration for the airport-to-hotel transfer: Standard Sedan S, LiftVan A, and LiftVan B. Standard Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A has a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases. LiftVan A is scheduled for pickup at 3:10 PM. LiftVan B is scheduled for pickup at 2:55 PM. Both vans' pickup times are no earlier than the party's 14:30 ready time. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, entity, and time bindings while only altering LiftVan B's pickup time in the counterfactual, which stays consistent with the party-ready constraint and contains no leaked labels or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Case note: Mira arrives at 14:10 with two companions, three suitcases, and a powered wheelchair; she cannot transfer out of it and her party is ready at 14:30. Standard Sedan S has four seats and luggage space but no wheelchair ramp and no wheelchair securement. Jo, an accessible-vehicle dispatcher, confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has identical accessibility features: a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both vans' pickups are no earlier than Mira's party-ready time. The vehicles under consideration are exactly Standard Sedan S, LiftVan A, and LiftVan B. LiftVan A's pickup is scheduled for 3:10 PM. LiftVan B's pickup is scheduled for 3:25 PM. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's pickup is scheduled for 3:10 PM."}, {"path": [], "text": "LiftVan B's pickup is scheduled for 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's pickup is scheduled for 3:10 PM.", "negative_left": "LiftVan A's pickup is scheduled for 3:10 PM.", "negative_right": "LiftVan B's pickup is scheduled for 2:55 PM.", "right": "LiftVan B's pickup is scheduled for 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-019", "id": "fast-43-diverse-112-019-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Case note: Mira arrives at 14:10 with two companions, three suitcases, and a powered wheelchair; she cannot transfer out of it and her party is ready at 14:30. Standard Sedan S has four seats and luggage space but no wheelchair ramp and no wheelchair securement. Jo, an accessible-vehicle dispatcher, confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has identical accessibility features: a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both vans' pickups are no earlier than Mira's party-ready time. The vehicles under consideration are exactly Standard Sedan S, LiftVan A, and LiftVan B. LiftVan A's pickup is scheduled for 3:10 PM. LiftVan B's pickup is scheduled for 3:25 PM. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, entity, and time bindings while only altering LiftVan B's pickup time in the counterfactual, which stays consistent with the party-ready constraint and contains no leaked labels or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Case note: Mira arrives at 14:10 with two companions, three suitcases, and a powered wheelchair; she cannot transfer out of it and her party is ready at 14:30. Standard Sedan S has four seats and luggage space but no wheelchair ramp and no wheelchair securement. Jo, an accessible-vehicle dispatcher, confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has identical accessibility features: a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both vans' pickups are no earlier than Mira's party-ready time. The vehicles under consideration are exactly Standard Sedan S, LiftVan A, and LiftVan B. LiftVan A's pickup is scheduled for 3:10 PM. LiftVan B's pickup is scheduled for 3:25 PM. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's pickup is scheduled for 3:10 PM."}, {"path": [], "text": "LiftVan B's pickup is scheduled for 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's pickup is scheduled for 3:10 PM.", "negative_left": "LiftVan A's pickup is scheduled for 3:10 PM.", "negative_right": "LiftVan B's pickup is scheduled for 2:55 PM.", "right": "LiftVan B's pickup is scheduled for 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-019", "id": "fast-43-diverse-112-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Case note: Mira arrives at 14:10 with two companions, three suitcases, and a powered wheelchair; she cannot transfer out of it and her party is ready at 14:30. Standard Sedan S has four seats and luggage space but no wheelchair ramp and no wheelchair securement. Jo, an accessible-vehicle dispatcher, confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases; it is offered through accessible-vehicle dispatch. LiftVan B has identical accessibility features: a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both vans' pickups are no earlier than Mira's party-ready time. The vehicles under consideration are exactly Standard Sedan S, LiftVan A, and LiftVan B. LiftVan A's pickup is scheduled for 3:10 PM. LiftVan B's pickup is scheduled for 2:55 PM. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and vehicle/time bindings from the unchanged question, the evidence spans are two factual pickup-time sentences, the counterfactual only alters LiftVan A's pickup time to 3:40 PM without contradicting other stated facts, and neither context reveals a gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair; her party will be ready for pickup at 14:30. Standard Sedan S seats four and has luggage space but has no ramp and no wheelchair securement. Accessible-vehicle dispatch confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. LiftVan B has identical accessibility features: a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both LiftVan A and LiftVan B are scheduled to pick up no earlier than the party's 14:30 ready time. The only vehicles under consideration for the airport-to-hotel transfer are Standard Sedan S, LiftVan A, and LiftVan B. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup. LiftVan A's scheduled pickup at the curb is logged as 3:10 PM. LiftVan B's scheduled pickup at the curb is logged as 3:25 PM.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's scheduled pickup at the curb is logged as 3:10 PM."}, {"path": [], "text": "LiftVan B's scheduled pickup at the curb is logged as 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's scheduled pickup at the curb is logged as 3:10 PM.", "negative_left": "LiftVan A's scheduled pickup at the curb is logged as 3:40 PM.", "negative_right": "LiftVan B's scheduled pickup at the curb is logged as 3:25 PM.", "right": "LiftVan B's scheduled pickup at the curb is logged as 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-020", "id": "fast-43-diverse-112-020-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair; her party will be ready for pickup at 14:30. Standard Sedan S seats four and has luggage space but has no ramp and no wheelchair securement. Accessible-vehicle dispatch confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. LiftVan B has identical accessibility features: a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both LiftVan A and LiftVan B are scheduled to pick up no earlier than the party's 14:30 ready time. The only vehicles under consideration for the airport-to-hotel transfer are Standard Sedan S, LiftVan A, and LiftVan B. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup. LiftVan A's scheduled pickup at the curb is logged as 3:10 PM. LiftVan B's scheduled pickup at the curb is logged as 3:25 PM."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy and vehicle/time bindings from the unchanged question, the evidence spans are two factual pickup-time sentences, the counterfactual only alters LiftVan A's pickup time to 3:40 PM without contradicting other stated facts, and neither context reveals a gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair; her party will be ready for pickup at 14:30. Standard Sedan S seats four and has luggage space but has no ramp and no wheelchair securement. Accessible-vehicle dispatch confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. LiftVan B has identical accessibility features: a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both LiftVan A and LiftVan B are scheduled to pick up no earlier than the party's 14:30 ready time. The only vehicles under consideration for the airport-to-hotel transfer are Standard Sedan S, LiftVan A, and LiftVan B. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup. LiftVan A's scheduled pickup at the curb is logged as 3:10 PM. LiftVan B's scheduled pickup at the curb is logged as 3:25 PM.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's scheduled pickup at the curb is logged as 3:10 PM."}, {"path": [], "text": "LiftVan B's scheduled pickup at the curb is logged as 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's scheduled pickup at the curb is logged as 3:10 PM.", "negative_left": "LiftVan A's scheduled pickup at the curb is logged as 3:40 PM.", "negative_right": "LiftVan B's scheduled pickup at the curb is logged as 3:25 PM.", "right": "LiftVan B's scheduled pickup at the curb is logged as 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-020", "id": "fast-43-diverse-112-020-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Arriving passenger Mira lands with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair; her party will be ready for pickup at 14:30. Standard Sedan S seats four and has luggage space but has no ramp and no wheelchair securement. Accessible-vehicle dispatch confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. LiftVan B has identical accessibility features: a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases. Both LiftVan A and LiftVan B are scheduled to pick up no earlier than the party's 14:30 ready time. The only vehicles under consideration for the airport-to-hotel transfer are Standard Sedan S, LiftVan A, and LiftVan B. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup. LiftVan A's scheduled pickup at the curb is logged as 3:40 PM. LiftVan B's scheduled pickup at the curb is logged as 3:25 PM."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, entity, and time bindings; the two evidence sentences are complete factual dispatch statements; the counterfactual only changes LiftVan A's pickup time to 3:10 PM, keeping all other facts consistent and unequal pickup times intact; neither context reveals the intended answer or label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Case note: Mira lands with two companions and three suitcases, traveling in a powered wheelchair she cannot transfer out of. Her party will be ready for pickup at 14:30. Three vehicles are under consideration: Standard Sedan S, LiftVan A, and LiftVan B. Standard Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A and LiftVan B each have a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases. LiftVan A is offered through accessible-vehicle dispatch. Both vans' pickup times are no earlier than the party-ready time, and the two pickup times are unequal. LiftVan A's dispatch confirmation shows a pickup time of 2:40 PM at the airport curb. LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's dispatch confirmation shows a pickup time of 2:40 PM at the airport curb."}, {"path": [], "text": "LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's dispatch confirmation shows a pickup time of 2:40 PM at the airport curb.", "negative_left": "LiftVan A's dispatch confirmation shows a pickup time of 3:10 PM at the airport curb.", "negative_right": "LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb.", "right": "LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-022", "id": "fast-43-diverse-112-022-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Case note: Mira lands with two companions and three suitcases, traveling in a powered wheelchair she cannot transfer out of. Her party will be ready for pickup at 14:30. Three vehicles are under consideration: Standard Sedan S, LiftVan A, and LiftVan B. Standard Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A and LiftVan B each have a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases. LiftVan A is offered through accessible-vehicle dispatch. Both vans' pickup times are no earlier than the party-ready time, and the two pickup times are unequal. LiftVan A's dispatch confirmation shows a pickup time of 2:40 PM at the airport curb. LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, entity, and time bindings; the two evidence sentences are complete factual dispatch statements; the counterfactual only changes LiftVan A's pickup time to 3:10 PM, keeping all other facts consistent and unequal pickup times intact; neither context reveals the intended answer or label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Case note: Mira lands with two companions and three suitcases, traveling in a powered wheelchair she cannot transfer out of. Her party will be ready for pickup at 14:30. Three vehicles are under consideration: Standard Sedan S, LiftVan A, and LiftVan B. Standard Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A and LiftVan B each have a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases. LiftVan A is offered through accessible-vehicle dispatch. Both vans' pickup times are no earlier than the party-ready time, and the two pickup times are unequal. LiftVan A's dispatch confirmation shows a pickup time of 2:40 PM at the airport curb. LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's dispatch confirmation shows a pickup time of 2:40 PM at the airport curb."}, {"path": [], "text": "LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's dispatch confirmation shows a pickup time of 2:40 PM at the airport curb.", "negative_left": "LiftVan A's dispatch confirmation shows a pickup time of 3:10 PM at the airport curb.", "negative_right": "LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb.", "right": "LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-022", "id": "fast-43-diverse-112-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Case note: Mira lands with two companions and three suitcases, traveling in a powered wheelchair she cannot transfer out of. Her party will be ready for pickup at 14:30. Three vehicles are under consideration: Standard Sedan S, LiftVan A, and LiftVan B. Standard Sedan S has no wheelchair ramp and no wheelchair securement. LiftVan A and LiftVan B each have a wheelchair ramp, one occupied-wheelchair position, three additional passenger seats, and room for three suitcases. LiftVan A is offered through accessible-vehicle dispatch. Both vans' pickup times are no earlier than the party-ready time, and the two pickup times are unequal. LiftVan A's dispatch confirmation shows a pickup time of 3:10 PM at the airport curb. LiftVan B's dispatch confirmation shows a pickup time of 2:55 PM at the airport curb. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the routing and earliest-pickup policy and keep Mira/LiftVan A/LiftVan B/Sedan S bindings intact; the counterfactual only changes LiftVan B's pickup time (2:55 PM, still after the 14:30 party-ready time) which stays logically consistent; the two evidence spans are complete factual sentences matching verbatim; no labels, codes, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Arriving passenger Mira lands at 14:10 with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair, and the attendant confirms her party will be ready at 14:30. Standard Sedan S has four seats and luggage space but no ramp or wheelchair securement. Accessible-vehicle driver Jo confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases, with pickup no earlier than the party-ready time. LiftVan B has identical accessibility and capacity, also with pickup no earlier than the party-ready time. Only Sedan S, LiftVan A, and LiftVan B are under consideration, and LiftVan A is offered through accessible-vehicle dispatch. LiftVan A's pickup is scheduled for 3:10 PM. LiftVan B's pickup is scheduled for 3:25 PM. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's pickup is scheduled for 3:10 PM."}, {"path": [], "text": "LiftVan B's pickup is scheduled for 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's pickup is scheduled for 3:10 PM.", "negative_left": "LiftVan A's pickup is scheduled for 3:10 PM.", "negative_right": "LiftVan B's pickup is scheduled for 2:55 PM.", "right": "LiftVan B's pickup is scheduled for 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-026", "id": "fast-43-diverse-112-026-base", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Arriving passenger Mira lands at 14:10 with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair, and the attendant confirms her party will be ready at 14:30. Standard Sedan S has four seats and luggage space but no ramp or wheelchair securement. Accessible-vehicle driver Jo confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases, with pickup no earlier than the party-ready time. LiftVan B has identical accessibility and capacity, also with pickup no earlier than the party-ready time. Only Sedan S, LiftVan A, and LiftVan B are under consideration, and LiftVan A is offered through accessible-vehicle dispatch. LiftVan A's pickup is scheduled for 3:10 PM. LiftVan B's pickup is scheduled for 3:25 PM. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the routing and earliest-pickup policy and keep Mira/LiftVan A/LiftVan B/Sedan S bindings intact; the counterfactual only changes LiftVan B's pickup time (2:55 PM, still after the 14:30 party-ready time) which stays logically consistent; the two evidence spans are complete factual sentences matching verbatim; no labels, codes, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a19": "supported", "a2": "supported", "a20": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a20": "unknown"}, "remove_right": {"a20": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a20": "unknown"}, "negative_pair": {"a20": "refuted"}, "negative_sentence": {"a20": "unknown"}, "positive_pair": {"a20": "supported"}, "right": {"a20": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relation; a15 is an allowed exact-set relation rather than a bundled policy classification. The focus atom is the factual ordering of two pickup times. The base and counter assignments are realizable by reversing the vans' distinct pickup ordering while keeping both pickups no earlier than the party-ready time. The policy evidence correctly preserves the substantive routing and earliest-pickup rules originating in the original state; requirements already present in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that LiftVan A is in accessible-vehicle dispatch, meets the wheelchair, passenger, luggage, and readiness requirements, and is earlier than the only other suitable vehicle. Sedan S is excluded by its lack of a ramp and securement, and a15 excludes additional competing vehicles.", "rule_index": 0, "sound": true}, {"reason": "With unequal pickup times, refutation of A being earlier entails that LiftVan B is earlier. B meets the stated accessibility, capacity, luggage, and readiness requirements and is among the exact vehicles under consideration, so the earliest-pickup rule prevents A from ranking first.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mira cannot transfer out of her powered wheelchair for the airport-to-hotel transfer."}, {"id": "a2", "statement": "Mira's party includes exactly one passenger traveling in an occupied wheelchair."}, {"id": "a3", "statement": "Mira travels with exactly two companions."}, {"id": "a4", "statement": "Mira's party has exactly three suitcases."}, {"id": "a5", "statement": "LiftVan A has a wheelchair ramp."}, {"id": "a6", "statement": "LiftVan A has exactly one occupied-wheelchair position."}, {"id": "a7", "statement": "LiftVan A has exactly three additional passenger seats."}, {"id": "a8", "statement": "LiftVan A has room for exactly three suitcases."}, {"id": "a9", "statement": "LiftVan B has a wheelchair ramp."}, {"id": "a10", "statement": "LiftVan B has exactly one occupied-wheelchair position."}, {"id": "a11", "statement": "LiftVan B has exactly three additional passenger seats."}, {"id": "a12", "statement": "LiftVan B has room for exactly three suitcases."}, {"id": "a13", "statement": "Standard Sedan S has no wheelchair ramp."}, {"id": "a14", "statement": "Standard Sedan S has no wheelchair securement."}, {"id": "a15", "statement": "The vehicles under consideration for Mira's airport-to-hotel transfer are exactly Standard Sedan S, LiftVan A, and LiftVan B."}, {"id": "a16", "statement": "LiftVan A is offered through accessible-vehicle dispatch."}, {"id": "a17", "statement": "LiftVan A's pickup is no earlier than Mira's party-ready time."}, {"id": "a18", "statement": "LiftVan B's pickup is no earlier than Mira's party-ready time."}, {"id": "a19", "statement": "LiftVan A and LiftVan B have unequal pickup times."}, {"id": "a20", "statement": "LiftVan A's pickup time is earlier than LiftVan B's pickup time."}], "base_state_json": "\"Arriving passenger Mira lands at 14:10 with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair, and the attendant confirms her party will be ready at 14:30. Standard Sedan S has four seats and luggage space but no ramp or wheelchair securement. Accessible-vehicle driver Jo confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases, with pickup no earlier than the party-ready time. LiftVan B has identical accessibility and capacity, also with pickup no earlier than the party-ready time. Only Sedan S, LiftVan A, and LiftVan B are under consideration, and LiftVan A is offered through accessible-vehicle dispatch. LiftVan A's pickup is scheduled for 3:10 PM. LiftVan B's pickup is scheduled for 3:25 PM. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}], "focus_atom": "a20", "focus_evidence": [{"path": [], "text": "LiftVan A's pickup is scheduled for 3:10 PM."}, {"path": [], "text": "LiftVan B's pickup is scheduled for 3:25 PM."}], "policy_evidence": [{"path": [], "text": "Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}], "rules": [{"justification": "Mira must remain in her wheelchair, so accessible-vehicle dispatch applies. LiftVan A is in that dispatch pool and has sufficient ramp access, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Sedan S lacks both a ramp and securement. LiftVan B meets the same relevant requirements, but LiftVan A is earlier. Because these are exactly the vehicles under consideration, LiftVan A is the earliest suitable vehicle.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "supported"}]}, {"justification": "LiftVan B has sufficient accessibility, occupied-wheelchair capacity, additional seating, luggage capacity, and a pickup no earlier than the party-ready time. Because the two van pickup times are unequal and LiftVan A is not earlier than LiftVan B, LiftVan B is earlier. The earliest-pickup ranking policy therefore prevents LiftVan A from ranking first.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}, {"atom_id": "a19", "state": "supported"}, {"atom_id": "a20", "state": "refuted"}]}]}, "verified_pair": {"left": "LiftVan A's pickup is scheduled for 3:10 PM.", "negative_left": "LiftVan A's pickup is scheduled for 3:10 PM.", "negative_right": "LiftVan B's pickup is scheduled for 2:55 PM.", "right": "LiftVan B's pickup is scheduled for 3:25 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-112-026", "id": "fast-43-diverse-112-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "LiftVan A is unsuitable or should not rank first under the stated accessibility, capacity, and pickup rules.", "true": "LiftVan A is suitable, belongs with accessible-vehicle dispatch, and ranks first because it meets all requirements with the earliest pickup."}, "instructions": "Answer yes if LiftVan A should be assigned as the highest-ranked suitable airport-to-hotel transfer under the stated policy; otherwise answer no.", "type": "noul"}}, "state": "Arriving passenger Mira lands at 14:10 with two companions, three suitcases, and a powered wheelchair. She cannot transfer out of the wheelchair, and the attendant confirms her party will be ready at 14:30. Standard Sedan S has four seats and luggage space but no ramp or wheelchair securement. Accessible-vehicle driver Jo confirms LiftVan A has a ramp, one occupied-wheelchair position, three additional seats, and room for three suitcases, with pickup no earlier than the party-ready time. LiftVan B has identical accessibility and capacity, also with pickup no earlier than the party-ready time. Only Sedan S, LiftVan A, and LiftVan B are under consideration, and LiftVan A is offered through accessible-vehicle dispatch. LiftVan A's pickup is scheduled for 3:10 PM. LiftVan B's pickup is scheduled for 2:55 PM. Policy routes passengers who cannot leave their wheelchair to accessible-vehicle dispatch and ranks suitable vehicles by earliest pickup."}, "method": "c2d", "provenance": {"source_id": "diverse-112", "source_is_synthetic": true, "source_sha256": "9991b93c5c7fd79afefaa541137392147734cbf5aed67a1ecbd522bbcc750830", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's stated conditional policy and question criteria intact, the focus evidence sentences are plain factual statements, and the counterfactual alters only the trunk-capacity figure (3\\u21922 suitcases) creating a coherent luggage-mismatch scenario without contradicting other unchanged facts or leaking any answer/rule content.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.\",\"evidence\":[\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\",\"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\",\"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\",\"Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\",\"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 2 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-003", "id": "fast-43-diverse-113-003-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's stated conditional policy and question criteria intact, the focus evidence sentences are plain factual statements, and the counterfactual alters only the trunk-capacity figure (3\\u21922 suitcases) creating a coherent luggage-mismatch scenario without contradicting other unchanged facts or leaking any answer/rule content.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.\",\"evidence\":[\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\",\"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\",\"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\",\"Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\",\"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 2 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-003", "id": "fast-43-diverse-113-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 2 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions, entities, locations and timing while only varying suitcase-count/capacity facts, which are complete factual sentences with no embedded rules or answers, and the counterfactual's single altered capacity sentence remains logically consistent with the unchanged 3-suitcase fact.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, Mira and one companion need a transfer to the Lumen Hotel. They have one folding manual wheelchair and one backpack. Mira can independently transfer into a car and needs no boarding assistance.\",\"evidence\":[\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\",\"Sedan S4 arrives in 18 minutes; verified capacity is four passengers, one backpack, and one folded manual wheelchair in the trunk.\",\"Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 3 medium suitcases.\",\"Accessible van V2 arrives in 10 minutes and has a ramp and securement space.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\",\"Mira needs no help boarding, transferring into, or reaching the sedan, and she does not need to remain in her wheelchair, a ramp, or securement for this transfer.\"],\"request\":\"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "3"], "text": "Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 3 medium suitcases."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 2 medium suitcases.", "right": "Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 3 medium suitcases."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-015", "id": "fast-43-diverse-113-015-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "At fictional Northstar Airport, Mira and one companion need a transfer to the Lumen Hotel. They have one folding manual wheelchair and one backpack. Mira can independently transfer into a car and needs no boarding assistance.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes; verified capacity is four passengers, one backpack, and one folded manual wheelchair in the trunk.", "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 3 medium suitcases.", "Accessible van V2 arrives in 10 minutes and has a ramp and securement space.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.", "Mira needs no help boarding, transferring into, or reaching the sedan, and she does not need to remain in her wheelchair, a ramp, or securement for this transfer."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions, entities, locations and timing while only varying suitcase-count/capacity facts, which are complete factual sentences with no embedded rules or answers, and the counterfactual's single altered capacity sentence remains logically consistent with the unchanged 3-suitcase fact.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"At fictional Northstar Airport, Mira and one companion need a transfer to the Lumen Hotel. They have one folding manual wheelchair and one backpack. Mira can independently transfer into a car and needs no boarding assistance.\",\"evidence\":[\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\",\"Sedan S4 arrives in 18 minutes; verified capacity is four passengers, one backpack, and one folded manual wheelchair in the trunk.\",\"Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 3 medium suitcases.\",\"Accessible van V2 arrives in 10 minutes and has a ramp and securement space.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\",\"Mira needs no help boarding, transferring into, or reaching the sedan, and she does not need to remain in her wheelchair, a ramp, or securement for this transfer.\"],\"request\":\"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "2"], "text": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "3"], "text": "Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 3 medium suitcases."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 2 medium suitcases.", "right": "Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 3 medium suitcases."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-015", "id": "fast-43-diverse-113-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "At fictional Northstar Airport, Mira and one companion need a transfer to the Lumen Hotel. They have one folding manual wheelchair and one backpack. Mira can independently transfer into a car and needs no boarding assistance.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes; verified capacity is four passengers, one backpack, and one folded manual wheelchair in the trunk.", "Mira's party is carrying 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's verified trunk configuration for the Northstar Airport to Lumen Hotel transfer holds up to 2 medium suitcases.", "Accessible van V2 arrives in 10 minutes and has a ramp and securement space.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.", "Mira needs no help boarding, transferring into, or reaching the sedan, and she does not need to remain in her wheelchair, a ramp, or securement for this transfer."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve Mira's conditional policy and all governing criteria via the unchanged question; entities, path, and time bindings match the original request; the two focus evidence lines are complete factual sentences; the counterfactual coherently reduces sedan capacity to 3 suitcases versus the party's 4, without contradicting other unchanged facts; and neither context reveals a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, plus suitcases as noted below.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\",\"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\",\"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\",\"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-043", "id": "fast-43-diverse-113-043-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, plus suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve Mira's conditional policy and all governing criteria via the unchanged question; entities, path, and time bindings match the original request; the two focus evidence lines are complete factual sentences; the counterfactual coherently reduces sedan capacity to 3 suitcases versus the party's 4, without contradicting other unchanged facts; and neither context reveals a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, plus suitcases as noted below.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\",\"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\",\"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\",\"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-043", "id": "fast-43-diverse-113-043-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, plus suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mira's stated conditional policy, the party's identity, airport-to-hotel path, and 20-minute timing binding intact via the unchanged question; the focus evidence pair are complete factual sentences about suitcase count vs. verified capacity, not policy text; the counterfactual lowers sedan suitcase capacity to 3, creating a coherent alternate failure scenario without duplicating or contradicting other measurements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-045", "id": "fast-43-diverse-113-045-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain Mira's stated conditional policy, the party's identity, airport-to-hotel path, and 20-minute timing binding intact via the unchanged question; the focus evidence pair are complete factual sentences about suitcase count vs. verified capacity, not policy text; the counterfactual lowers sedan suitcase capacity to 3, creating a coherent alternate failure scenario without duplicating or contradicting other measurements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-045", "id": "fast-43-diverse-113-045-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's entity, route, and 20-minute time binding intact, preserve the governing policy via the unchanged questions object, and the two focus evidence lines are plain factual statements about suitcase counts rather than policy text; the counterfactual's change from 4 to 3 suitcase capacity is a coherent case-observation alteration (not a duplicate contradiction) that plausibly shifts the capacity outcome, and neither context reveals any gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing a transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-048", "id": "fast-43-diverse-113-048-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing a transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's entity, route, and 20-minute time binding intact, preserve the governing policy via the unchanged questions object, and the two focus evidence lines are plain factual statements about suitcase counts rather than policy text; the counterfactual's change from 4 to 3 suitcase capacity is a coherent case-observation alteration (not a duplicate contradiction) that plausibly shifts the capacity outcome, and neither context reveals any gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing a transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-048", "id": "fast-43-diverse-113-048-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing a transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only changes trunk suitcase capacity from 4 to 3, creating a coherent scenario where luggage no longer fits, without contradicting any other evidence; both contexts preserve the unchanged questions object and all governing policy, use complete factual sentences as evidence, and contain no gold labels, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, along with their suitcases as logged below.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"Sedan S4 arrives in 18 minutes, within Mira's stated 20-minute limit.\",\"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair secured in the trunk, covering the party of two and their backpack and folded chair.\",\"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\",\"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-053", "id": "fast-43-diverse-113-053-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, along with their suitcases as logged below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's stated 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair secured in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only changes trunk suitcase capacity from 4 to 3, creating a coherent scenario where luggage no longer fits, without contradicting any other evidence; both contexts preserve the unchanged questions object and all governing policy, use complete factual sentences as evidence, and contain no gold labels, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\":\"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, along with their suitcases as logged below.\",\"evidence\":[\"Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”\",\"Sedan S4 arrives in 18 minutes, within Mira's stated 20-minute limit.\",\"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair secured in the trunk, covering the party of two and their backpack and folded chair.\",\"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\",\"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\",\"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\",\"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"],\"request\":\"Choose the correct dispatch route under Mira’s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-053", "id": "fast-43-diverse-113-053-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, along with their suitcases as logged below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's stated 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair secured in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's conditional statement, timing, and assistance facts intact while only altering trunk capacity, which is a permissible case-level change; the two focus sentences are plain factual claims, not policy text; the counterfactual (trunk holds only 3 suitcases vs 4 needed) is internally consistent and doesn't duplicate or contradict other measurements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack along with their suitcases as noted below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-054", "id": "fast-43-diverse-113-054-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack along with their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's conditional statement, timing, and assistance facts intact while only altering trunk capacity, which is a permissible case-level change; the two focus sentences are plain factual claims, not policy text; the counterfactual (trunk holds only 3 suitcases vs 4 needed) is internally consistent and doesn't duplicate or contradict other measurements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack along with their suitcases as noted below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-054", "id": "fast-43-diverse-113-054-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack along with their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing capacity/time/assistance policy and Mira's entity/path/time bindings, the focus evidence are complete factual sentences, and the counterfactual coherently shifts trunk capacity to 2 while suitcases remain 4, creating a valid capacity failure without contradicting any other fact or leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack along with their suitcases as recorded below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 2 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-061", "id": "fast-43-diverse-113-061-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack along with their suitcases as recorded below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing capacity/time/assistance policy and Mira's entity/path/time bindings, the focus evidence are complete factual sentences, and the counterfactual coherently shifts trunk capacity to 2 while suitcases remain 4, creating a valid capacity failure without contradicting any other fact or leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack along with their suitcases as recorded below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 2 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-061", "id": "fast-43-diverse-113-061-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack along with their suitcases as recorded below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 2 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's entity, route, and 20-minute condition intact, use the same governing policy already carried in the unchanged questions object, provide two complete factual evidence sentences on suitcase count and sedan capacity, and the counterfactual's lowered trunk capacity is a coherent case-observation change without contradicting other stated facts or leaking any answer, code, or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"At Northstar Airport, Mira and one companion prepare for transfer to the Lumen Hotel. Their party includes one folding manual wheelchair and one backpack, in addition to their suitcases.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-065", "id": "fast-43-diverse-113-065-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "At Northstar Airport, Mira and one companion prepare for transfer to the Lumen Hotel. Their party includes one folding manual wheelchair and one backpack, in addition to their suitcases.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's entity, route, and 20-minute condition intact, use the same governing policy already carried in the unchanged questions object, provide two complete factual evidence sentences on suitcase count and sedan capacity, and the counterfactual's lowered trunk capacity is a coherent case-observation change without contradicting other stated facts or leaking any answer, code, or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"At Northstar Airport, Mira and one companion prepare for transfer to the Lumen Hotel. Their party includes one folding manual wheelchair and one backpack, in addition to their suitcases.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-065", "id": "fast-43-diverse-113-065-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "At Northstar Airport, Mira and one companion prepare for transfer to the Lumen Hotel. Their party includes one folding manual wheelchair and one backpack, in addition to their suitcases.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original decision criteria via the unchanged questions object, preserve the same entities/timing bindings, present complete factual evidence sentences without embedding rule tables or answer hints, and the counterfactual's single altered sentence (trunk capacity 4→3) creates a plausible, non-duplicative variant scenario rather than a contradiction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, along with their suitcases as noted below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-066", "id": "fast-43-diverse-113-066-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, along with their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original decision criteria via the unchanged questions object, preserve the same entities/timing bindings, present complete factual evidence sentences without embedding rule tables or answer hints, and the counterfactual's single altered sentence (trunk capacity 4→3) creates a plausible, non-duplicative variant scenario rather than a contradiction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, along with their suitcases as noted below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-066", "id": "fast-43-diverse-113-066-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack, along with their suitcases as noted below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the passenger's stated conditional policy and all question bindings (Mira, Northstar\\u2192Lumen, 20-min limit) verbatim; the focus evidence pair are complete factual sentences about suitcase counts, and the counterfactual's reduced trunk capacity (3 vs 4 suitcases) creates a plausible, non-contradictory single-fact change without duplicate measurements or embedded answers/rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and her travel companion land at Northstar Airport and need transfer to the Lumen Hotel. They bring one folding manual wheelchair and one backpack along with their suitcases.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-067", "id": "fast-43-diverse-113-067-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and her travel companion land at Northstar Airport and need transfer to the Lumen Hotel. They bring one folding manual wheelchair and one backpack along with their suitcases.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the passenger's stated conditional policy and all question bindings (Mira, Northstar\\u2192Lumen, 20-min limit) verbatim; the focus evidence pair are complete factual sentences about suitcase counts, and the counterfactual's reduced trunk capacity (3 vs 4 suitcases) creates a plausible, non-contradictory single-fact change without duplicate measurements or embedded answers/rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and her travel companion land at Northstar Airport and need transfer to the Lumen Hotel. They bring one folding manual wheelchair and one backpack along with their suitcases.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-067", "id": "fast-43-diverse-113-067-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and her travel companion land at Northstar Airport and need transfer to the Lumen Hotel. They bring one folding manual wheelchair and one backpack along with their suitcases.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy (via unchanged questions) and the same entity/path/time bindings; the two focus evidence spans are plain factual statements about suitcase counts, not policy text; the counterfactual coherently reduces trunk capacity to 3 suitcases while keeping the party’s 4-suitcase count, creating a plausible capacity failure without contradicting other unchanged facts; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below. Mira can independently transfer into a car and needs no boarding assistance.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-068", "id": "fast-43-diverse-113-068-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below. Mira can independently transfer into a car and needs no boarding assistance.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy (via unchanged questions) and the same entity/path/time bindings; the two focus evidence spans are plain factual statements about suitcase counts, not policy text; the counterfactual coherently reduces trunk capacity to 3 suitcases while keeping the party’s 4-suitcase count, creating a plausible capacity failure without contradicting other unchanged facts; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below. Mira can independently transfer into a car and needs no boarding assistance.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-068", "id": "fast-43-diverse-113-068-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair, one backpack, and their suitcases as noted below. Mira can independently transfer into a car and needs no boarding assistance.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, path, and time bindings, preserve the original conditional policy via the unchanged question, use complete factual evidence sentences, and the counterfactual coherently flips only the suitcase capacity fact without introducing contradictions or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and her companion land at Northstar Airport and need transfer to the Lumen Hotel. Their party of two carries one folding manual wheelchair, a backpack, and several suitcases as recorded below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers and one backpack, covering the party of two and their backpack.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4 has a verified spot for one folded manual wheelchair in the trunk alongside the suitcases.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-071", "id": "fast-43-diverse-113-071-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and her companion land at Northstar Airport and need transfer to the Lumen Hotel. Their party of two carries one folding manual wheelchair, a backpack, and several suitcases as recorded below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers and one backpack, covering the party of two and their backpack.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4 has a verified spot for one folded manual wheelchair in the trunk alongside the suitcases.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, path, and time bindings, preserve the original conditional policy via the unchanged question, use complete factual evidence sentences, and the counterfactual coherently flips only the suitcase capacity fact without introducing contradictions or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and her companion land at Northstar Airport and need transfer to the Lumen Hotel. Their party of two carries one folding manual wheelchair, a backpack, and several suitcases as recorded below.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers and one backpack, covering the party of two and their backpack.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4 has a verified spot for one folded manual wheelchair in the trunk alongside the suitcases.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-071", "id": "fast-43-diverse-113-071-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and her companion land at Northstar Airport and need transfer to the Lumen Hotel. Their party of two carries one folding manual wheelchair, a backpack, and several suitcases as recorded below.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers and one backpack, covering the party of two and their backpack.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4 has a verified spot for one folded manual wheelchair in the trunk alongside the suitcases.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing conditional-dispatch policy via the unchanged questions and Mira's quoted condition, keep the same entities/route/time binding, use plain factual evidence sentences with no embedded answers, and the counterfactual coherently alters only the suitcase-capacity fact (4 vs 3) without contradicting other stated facts.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack in addition to their suitcases.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-074", "id": "fast-43-diverse-113-074-base", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack in addition to their suitcases.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "travel-04", "split": "train", "variant": "base"} {"domain": "travel", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing conditional-dispatch policy via the unchanged questions and Mira's quoted condition, keep the same entities/route/time binding, use plain factual evidence sentences with no embedded answers, and the counterfactual coherently alters only the suitcase-capacity fact (4 vs 3) without contradicting other stated facts.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "full_context_fact_states": {"base": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "supported", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "counterfactual": {"a_backpack_capacity": "supported", "a_boarding_help": "refuted", "a_party_capacity": "supported", "a_pickup_limit": "supported", "a_ramp_required": "refuted", "a_reaching_help": "refuted", "a_remain_wheelchair": "refuted", "a_securement_required": "refuted", "a_suitcase_capacity": "refuted", "a_transfer_help": "refuted", "a_wheelchair_capacity": "supported"}, "remove_left": {"a_suitcase_capacity": "unknown"}, "remove_right": {"a_suitcase_capacity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_suitcase_capacity": "unknown"}, "negative_pair": {"a_suitcase_capacity": "refuted"}, "negative_sentence": {"a_suitcase_capacity": "unknown"}, "positive_pair": {"a_suitcase_capacity": "supported"}, "right": {"a_suitcase_capacity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship, and the suitcase-capacity focus is factual rather than policy. The base and counter assignments differ only on that focus and are jointly realizable in synthetic scenarios. Policy evidence preserves the state-originated conditional dispatch rule and airport-attendant restriction; governing material in the unchanged questions object need not be duplicated, while omitted capacity and passenger details are replaceable case observations. The extra generic request citation is unnecessary but does not make the policy incomplete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every stated sedan-capacity and timing requirement and explicitly refutes all assistance and accessible-vehicle triggers. It is sufficient for standard dispatch (target 0).", "rule_index": 0, "sound": true}, {"reason": "Refutation of suitcase capacity establishes that all luggage cannot fit, so a stated standard-service condition fails. The conjunction also refutes every assistance trigger, excluding the competing airport-assistance route, and is sufficient for accessible-vehicle dispatch (target 2).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_party_capacity", "statement": "Sedan S4's verified passenger capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of travelers in Mira's party."}, {"id": "a_suitcase_capacity", "statement": "Sedan S4's verified medium-suitcase capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of medium suitcases carried by Mira's party."}, {"id": "a_backpack_capacity", "statement": "Sedan S4's verified backpack capacity for the transfer from Northstar Airport to the Lumen Hotel is at least the number of backpacks carried by Mira's party."}, {"id": "a_wheelchair_capacity", "statement": "Sedan S4 has verified capacity for Mira's folded manual wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_pickup_limit", "statement": "Sedan S4's pickup time at Northstar Airport is no more than Mira's stated 20-minute limit for using a standard sedan."}, {"id": "a_boarding_help", "statement": "Mira requests or requires help boarding the dispatched vehicle at Northstar Airport."}, {"id": "a_transfer_help", "statement": "Mira requests or requires help transferring into the dispatched vehicle at Northstar Airport."}, {"id": "a_reaching_help", "statement": "Mira requests or requires help reaching the dispatched vehicle at Northstar Airport."}, {"id": "a_remain_wheelchair", "statement": "Mira must remain in her wheelchair during the transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_ramp_required", "statement": "Mira requires a ramp for the dispatched transfer from Northstar Airport to the Lumen Hotel."}, {"id": "a_securement_required", "statement": "Mira requires wheelchair securement for the dispatched transfer from Northstar Airport to the Lumen Hotel."}], "base_state_json": "{\"context\": \"Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack in addition to their suitcases.\", \"evidence\": [\"Mira states: \\u201cUse a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.\\u201d\", \"Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.\", \"Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.\", \"Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.\", \"Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.\", \"Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.\", \"Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring.\"], \"request\": \"Choose the correct dispatch route under Mira\\u2019s conditional request and the verified service evidence.\"}", "base_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "counter_states": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}], "focus_atom": "a_suitcase_capacity", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, {"path": ["evidence", "4"], "text": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”"}, {"path": ["evidence", "3"], "text": "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."}, {"path": ["request"], "text": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}], "rules": [{"justification": "All verified sedan-capacity, timing, and assistance conditions are met, and no accessible-vehicle trigger applies, so the standard-dispatch score applies.", "target": "0", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "supported"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}, {"justification": "The sedan cannot carry all of Mira's luggage, so her stated standard-service condition fails and her request requires accessible-vehicle dispatch; the explicit assistance exclusions prevent the airport-assistance outcome.", "target": "2", "when": [{"atom_id": "a_party_capacity", "state": "supported"}, {"atom_id": "a_suitcase_capacity", "state": "refuted"}, {"atom_id": "a_backpack_capacity", "state": "supported"}, {"atom_id": "a_wheelchair_capacity", "state": "supported"}, {"atom_id": "a_pickup_limit", "state": "supported"}, {"atom_id": "a_boarding_help", "state": "refuted"}, {"atom_id": "a_transfer_help", "state": "refuted"}, {"atom_id": "a_reaching_help", "state": "refuted"}, {"atom_id": "a_remain_wheelchair", "state": "refuted"}, {"atom_id": "a_ramp_required", "state": "refuted"}, {"atom_id": "a_securement_required", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_left": "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "negative_right": "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "right": "Sedan S4's trunk was verified to hold 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-113-074", "id": "fast-43-diverse-113-074-counterfactual", "input": {"questions": {"decision": {"criteria": ["Route to standard dispatch: verified sedan capacity meets the full party, luggage, and folded-equipment needs; pickup meets the passenger’s conditional time limit; no boarding assistance is needed.", "Route first to airport assistance staff: capacity may otherwise fit, but the passenger requests or requires help boarding, transferring, or reaching the vehicle before dispatch.", "Route to accessible-vehicle dispatch: the passenger must remain in the wheelchair, requires a ramp or securement, or the passenger’s stated standard-service conditions fail."], "instructions": "Score the routing decision using the ordered levels below. Apply the passenger’s stated condition before comparing raw pickup speed. A service is suitable only if verified passenger, luggage, mobility-equipment, timing, and assistance requirements are all met.", "type": "score"}}, "state": {"context": "Mira and one companion arrive at Northstar Airport needing transfer to the Lumen Hotel. They carry one folding manual wheelchair and one backpack in addition to their suitcases.", "evidence": ["Mira states: “Use a standard sedan if my folded chair and all luggage fit and pickup is within 20 minutes; otherwise send a wheelchair-accessible van.”", "Sedan S4 arrives in 18 minutes, within Mira's 20-minute limit.", "Sedan S4's verified capacity includes four passengers, one backpack, and one folded manual wheelchair in the trunk, covering the party of two and their backpack and folded chair.", "Mira's party is carrying 4 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Sedan S4's trunk was verified to hold 3 medium suitcases for the transfer from Northstar Airport to the Lumen Hotel.", "Mira can independently transfer into the sedan and needs no boarding, transfer, or reaching assistance.", "Mira does not need to remain in her wheelchair during the ride, and no ramp or securement is required.", "Airport assistance attendants are used only when a passenger requests or requires help boarding or transferring."], "request": "Choose the correct dispatch route under Mira’s conditional request and the verified service evidence."}}, "method": "c2d", "provenance": {"source_id": "diverse-113", "source_is_synthetic": true, "source_sha256": "674053ffbb82e5250bd4107865ad5edf13e1bf0bf34f9d053226a00458eaa9dc", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "travel-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the routing policy and question bindings unchanged, the two focus sentences are plain factual statements not policy text, and the counterfactual's altered registry identification (authentication-token application component) still satisfies the exclusivity clause and does not contradict other unchanged facts or leak the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"Incident INC-4471: a pull request that changes only documentation comments caused the compile job to fail because Generated/ApiClient.cs was absent. Build log entries read: 'Generation skipped: cache marker present' and 'not a test assertion failure; no tests started.' A clean-cache rerun regenerated the file and the job passed. The same test suite had passed cleanly in the previous 20 runs, with no assertion failures and no setup flakiness observed. No worker disconnects, network errors, or hosted-service outages occurred during the run. Investigators confirmed the pull request introduced the defect flagged in INC-4471, and that this defect was the single primary cause of the compile-job failure. Component C-9002 is exactly one of the registry's build cache-invalidation component or an application source component. Incident record INC-4471 identifies the affected component as component C-9002. The component registry identifies component C-9002 as the build cache-invalidation component. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "Incident record INC-4471 identifies the affected component as component C-9002."}, {"path": [], "text": "The component registry identifies component C-9002 as the build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "Incident record INC-4471 identifies the affected component as component C-9002.", "negative_left": "Incident record INC-4471 identifies the affected component as component C-9002.", "negative_right": "The component registry identifies component C-9002 as the authentication-token application component.", "right": "The component registry identifies component C-9002 as the build cache-invalidation component."}, "verifier_independent_model": false}, "family": "fast-43-diverse-127-008", "id": "fast-43-diverse-127-008-base", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "Incident INC-4471: a pull request that changes only documentation comments caused the compile job to fail because Generated/ApiClient.cs was absent. Build log entries read: 'Generation skipped: cache marker present' and 'not a test assertion failure; no tests started.' A clean-cache rerun regenerated the file and the job passed. The same test suite had passed cleanly in the previous 20 runs, with no assertion failures and no setup flakiness observed. No worker disconnects, network errors, or hosted-service outages occurred during the run. Investigators confirmed the pull request introduced the defect flagged in INC-4471, and that this defect was the single primary cause of the compile-job failure. Component C-9002 is exactly one of the registry's build cache-invalidation component or an application source component. Incident record INC-4471 identifies the affected component as component C-9002. The component registry identifies component C-9002 as the build cache-invalidation component. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the routing policy and question bindings unchanged, the two focus sentences are plain factual statements not policy text, and the counterfactual's altered registry identification (authentication-token application component) still satisfies the exclusivity clause and does not contradict other unchanged facts or leak the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "refuted", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship rather than a final routing classification. The focus atom a3 is a factual component-identity relation. The base and counter assignments are jointly realizable: with a4 fixed, changing a3 changes the incident component from the cache-invalidation component to the application-source alternative without requiring another atom-state change. The state-derived policy evidence accurately preserves the build-engineering ownership rule relevant to interpreting none_of_above; rules already present in the retained questions need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that the sole primary cause was a defect in the build cache-invalidation component, exclude application-source classification through the exclusive alternative in a4, and explicitly refute test-automation and infrastructure causes. Build cache invalidation belongs to build engineering, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Given a4's exclusive alternatives and the refutation of a3, the incident component must be an application source component. The rule also establishes that the pull request introduced its defect and that this was the single primary cause, while excluding test-automation and infrastructure causes. Thus application_owner is required.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The pull request introduced the defect in the component identified by the incident record."}, {"id": "a2", "statement": "The defect in the component identified by the incident record was the single primary cause of the compile-job failure."}, {"id": "a3", "statement": "The component identified by the incident record is the same component that the component registry identifies as the build cache-invalidation component."}, {"id": "a4", "statement": "The component identified by the incident record is exactly one of the following: the component registry's build cache-invalidation component or an application source component."}, {"id": "a5", "statement": "A test assertion failure caused the compile-job failure."}, {"id": "a6", "statement": "A test setup failure or flakiness caused the compile-job failure."}, {"id": "a7", "statement": "A CI worker fault caused the compile-job failure."}, {"id": "a8", "statement": "A network fault caused the compile-job failure."}, {"id": "a9", "statement": "A hosted-service fault caused the compile-job failure."}], "base_state_json": "\"Incident INC-4471: a pull request that changes only documentation comments caused the compile job to fail because Generated/ApiClient.cs was absent. Build log entries read: 'Generation skipped: cache marker present' and 'not a test assertion failure; no tests started.' A clean-cache rerun regenerated the file and the job passed. The same test suite had passed cleanly in the previous 20 runs, with no assertion failures and no setup flakiness observed. No worker disconnects, network errors, or hosted-service outages occurred during the run. Investigators confirmed the pull request introduced the defect flagged in INC-4471, and that this defect was the single primary cause of the compile-job failure. Component C-9002 is exactly one of the registry's build cache-invalidation component or an application source component. Incident record INC-4471 identifies the affected component as component C-9002. The component registry identifies component C-9002 as the build cache-invalidation component. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": [], "text": "Incident record INC-4471 identifies the affected component as component C-9002."}, {"path": [], "text": "The component registry identifies component C-9002 as the build cache-invalidation component."}], "policy_evidence": [{"path": [], "text": "Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}], "rules": [{"justification": "The sole primary defect is in the build cache-invalidation component, which is assigned to build engineering, while the application, test-automation, and infrastructure rubrics do not apply.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "Because the incident component is exactly one of the cache-invalidation component or an application source component, and it is not the cache-invalidation component, it is an application source component. The pull request introduced its defect, and that defect was the sole primary cause.", "target": "application_owner", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "Incident record INC-4471 identifies the affected component as component C-9002.", "negative_left": "Incident record INC-4471 identifies the affected component as component C-9002.", "negative_right": "The component registry identifies component C-9002 as the authentication-token application component.", "right": "The component registry identifies component C-9002 as the build cache-invalidation component."}, "verifier_independent_model": false}, "family": "fast-43-diverse-127-008", "id": "fast-43-diverse-127-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"application_owner": "Choose only when evidence shows the change introduced an application source, type, or behavior defect.", "infrastructure_owner": "Choose only when a CI worker, network, or hosted service fault caused the failure.", "none_of_above": "Choose when none of the named rubrics applies, including failures owned by build engineering such as generator or cache-invalidation defects.", "test_automation_owner": "Choose only when a test assertion or test setup failed or was flaky."}, "instructions": "Route the failure to the single primary owner whose rubric is satisfied by the evidence. If none of the three named owner options fits, select none_of_above.", "type": "choice"}}, "state": "Incident INC-4471: a pull request that changes only documentation comments caused the compile job to fail because Generated/ApiClient.cs was absent. Build log entries read: 'Generation skipped: cache marker present' and 'not a test assertion failure; no tests started.' A clean-cache rerun regenerated the file and the job passed. The same test suite had passed cleanly in the previous 20 runs, with no assertion failures and no setup flakiness observed. No worker disconnects, network errors, or hosted-service outages occurred during the run. Investigators confirmed the pull request introduced the defect flagged in INC-4471, and that this defect was the single primary cause of the compile-job failure. Component C-9002 is exactly one of the registry's build cache-invalidation component or an application source component. Incident record INC-4471 identifies the affected component as component C-9002. The component registry identifies component C-9002 as the authentication-token application component. Routing policy: application owns source defects introduced by the change; test automation owns failing assertions or flaky test setup; infrastructure owns worker, network, or hosted-service faults; build engineering owns build scripts, generators, and cache invalidation."}, "method": "c2d", "provenance": {"source_id": "diverse-127", "source_is_synthetic": true, "source_sha256": "c57b5db75d3c7662678157bd1dc95cec6f4129f8f8eca22c5971b74359d2d75f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "application_owner"}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the unchanged policy paragraph and the original questions object is applied verbatim, preserving all bindings; the two evidence quotes are complete factual sentences (not policy/instructions) and match the base context exactly; the counterfactual only swaps the prior main-run signature from SIG-77C to SIG-90Z, a coherent single-sentence change that doesn't create duplicate/contradictory measurements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run produced failure signature SIG-77C. A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-77C. This is the only current in-scope required-test failure for PR 482.\"},{\"speaker\":\"Build engineer\",\"text\":\"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run produced failure signature SIG-77C."}, {"path": ["1", "text"], "text": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-77C."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run produced failure signature SIG-77C.", "negative_left": "PR 482's current Windows cli_integration run produced failure signature SIG-77C.", "negative_right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-90Z.", "right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-77C."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-006", "id": "fast-43-diverse-132-006-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run produced failure signature SIG-77C. A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-77C. This is the only current in-scope required-test failure for PR 482."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the unchanged policy paragraph and the original questions object is applied verbatim, preserving all bindings; the two evidence quotes are complete factual sentences (not policy/instructions) and match the base context exactly; the counterfactual only swaps the prior main-run signature from SIG-77C to SIG-90Z, a coherent single-sentence change that doesn't create duplicate/contradictory measurements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run produced failure signature SIG-77C. A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-77C. This is the only current in-scope required-test failure for PR 482.\"},{\"speaker\":\"Build engineer\",\"text\":\"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run produced failure signature SIG-77C."}, {"path": ["1", "text"], "text": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-77C."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run produced failure signature SIG-77C.", "negative_left": "PR 482's current Windows cli_integration run produced failure signature SIG-77C.", "negative_right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-90Z.", "right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-77C."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-006", "id": "fast-43-diverse-132-006-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run produced failure signature SIG-77C. A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, produced failure signature SIG-90Z. This is the only current in-scope required-test failure for PR 482."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Governing policy text (scope rule, known-flake exception, public-interface exclusion) is identical across both contexts and matches the original state's policy statement, so it is preserved; PR 482, Windows cli_integration, and ParserCache bindings remain unchanged. The two focus_evidence spans are complete factual sentences reporting failure signatures, not policy or instructions. The counterfactual only swaps the historical run's signature from SIG-7731 to SIG-4408, a coherent single-fact change that breaks the signature match without contradicting any other stated fact. Neither context contains gold answers, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\": \"Software developer\", \"text\": \"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"}, {\"speaker\": \"Test automation engineer\", \"text\": \"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run produced failure signature SIG-7731.\"}, {\"speaker\": \"Build engineer\", \"text\": \"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally. No other required test for PR 482 has failed apart from this cli_integration run.\"}, {\"speaker\": \"Code reviewer\", \"text\": \"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-7731.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run produced failure signature SIG-7731."}, {"path": ["3", "text"], "text": "Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-7731."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run produced failure signature SIG-7731.", "negative_left": "PR 482's current Windows cli_integration run produced failure signature SIG-7731.", "negative_right": "Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-4408.", "right": "Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-7731."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-007", "id": "fast-43-diverse-132-007-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run produced failure signature SIG-7731."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally. No other required test for PR 482 has failed apart from this cli_integration run."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-7731."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Governing policy text (scope rule, known-flake exception, public-interface exclusion) is identical across both contexts and matches the original state's policy statement, so it is preserved; PR 482, Windows cli_integration, and ParserCache bindings remain unchanged. The two focus_evidence spans are complete factual sentences reporting failure signatures, not policy or instructions. The counterfactual only swaps the historical run's signature from SIG-7731 to SIG-4408, a coherent single-fact change that breaks the signature match without contradicting any other stated fact. Neither context contains gold answers, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\": \"Software developer\", \"text\": \"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"}, {\"speaker\": \"Test automation engineer\", \"text\": \"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run produced failure signature SIG-7731.\"}, {\"speaker\": \"Build engineer\", \"text\": \"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally. No other required test for PR 482 has failed apart from this cli_integration run.\"}, {\"speaker\": \"Code reviewer\", \"text\": \"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-7731.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run produced failure signature SIG-7731."}, {"path": ["3", "text"], "text": "Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-7731."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run produced failure signature SIG-7731.", "negative_left": "PR 482's current Windows cli_integration run produced failure signature SIG-7731.", "negative_right": "Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-4408.", "right": "Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-7731."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-007", "id": "fast-43-diverse-132-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run produced failure signature SIG-7731."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally. No other required test for PR 482 has failed apart from this cli_integration run."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. Windows cli_integration run #204 on main, completed before PR 482's current run, has failure signature SIG-4408."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scope rule and known-flake exception policy verbatim, and the question's PR 482/Windows cli_integration/date bindings are unchanged; the two focus evidence spans are complete factual sentences describing test outcomes, not policy text; the counterfactual coherently substitutes SIG-QK2 for SIG-XJ9 in the prior main run without creating contradictory duplicate measurements; and no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error.\"},{\"speaker\":\"Build engineer\",\"text\":\"A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, also failed with signature SIG-XJ9. The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. The current Windows cli_integration failure is the only current in-scope required-test failure for PR 482.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error."}, {"path": ["2", "text"], "text": "A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, also failed with signature SIG-XJ9."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error.", "negative_left": "PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error.", "negative_right": "A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, failed with signature SIG-QK2 instead.", "right": "A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, also failed with signature SIG-XJ9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-008", "id": "fast-43-diverse-132-008-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error."}, {"speaker": "Build engineer", "text": "A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, also failed with signature SIG-XJ9. The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. The current Windows cli_integration failure is the only current in-scope required-test failure for PR 482."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scope rule and known-flake exception policy verbatim, and the question's PR 482/Windows cli_integration/date bindings are unchanged; the two focus evidence spans are complete factual sentences describing test outcomes, not policy text; the counterfactual coherently substitutes SIG-QK2 for SIG-XJ9 in the prior main run without creating contradictory duplicate measurements; and no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error.\"},{\"speaker\":\"Build engineer\",\"text\":\"A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, also failed with signature SIG-XJ9. The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. The current Windows cli_integration failure is the only current in-scope required-test failure for PR 482.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error."}, {"path": ["2", "text"], "text": "A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, also failed with signature SIG-XJ9."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error.", "negative_left": "PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error.", "negative_right": "A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, failed with signature SIG-QK2 instead.", "right": "A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, also failed with signature SIG-XJ9."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-008", "id": "fast-43-diverse-132-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run failed on 2024-06-01 with signature SIG-XJ9 categorized as a non-assertion timeout error."}, {"speaker": "Build engineer", "text": "A Windows cli_integration run on main completed on 2024-05-28, prior to the current PR 482 run, failed with signature SIG-QK2 instead. The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. The current Windows cli_integration failure is the only current in-scope required-test failure for PR 482."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scope rule and known-flake exception verbatim, keep the same PR/test/time bindings, use two complete factual sentences as evidence, the counterfactual coherently changes only the main-run signature to SIG-9042C without contradicting other facts, and neither context states or implies the final merge-readiness label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run failed with signature SIG-7741B. A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-7741B. This is the only in-scope required-test failure for PR 482.\"},{\"speaker\":\"Build engineer\",\"text\":\"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run failed with signature SIG-7741B."}, {"path": ["1", "text"], "text": "A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-7741B."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run failed with signature SIG-7741B.", "negative_left": "PR 482's current Windows cli_integration run failed with signature SIG-7741B.", "negative_right": "A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-9042C.", "right": "A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-7741B."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-009", "id": "fast-43-diverse-132-009-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run failed with signature SIG-7741B. A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-7741B. This is the only in-scope required-test failure for PR 482."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scope rule and known-flake exception verbatim, keep the same PR/test/time bindings, use two complete factual sentences as evidence, the counterfactual coherently changes only the main-run signature to SIG-9042C without contradicting other facts, and neither context states or implies the final merge-readiness label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run failed with signature SIG-7741B. A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-7741B. This is the only in-scope required-test failure for PR 482.\"},{\"speaker\":\"Build engineer\",\"text\":\"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run failed with signature SIG-7741B."}, {"path": ["1", "text"], "text": "A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-7741B."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run failed with signature SIG-7741B.", "negative_left": "PR 482's current Windows cli_integration run failed with signature SIG-7741B.", "negative_right": "A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-9042C.", "right": "A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-7741B."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-009", "id": "fast-43-diverse-132-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. PR 482's current Windows cli_integration run failed with signature SIG-7741B. A Windows cli_integration run on main completed on March 3, 2024, three days before the current PR 482 run, and failed with signature SIG-9042C. This is the only in-scope required-test failure for PR 482."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the scope rule and flake-exception policy verbatim, keep the PR 482/Windows cli_integration bindings intact, and the two focus evidence spans ('The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21.' and 'A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with signature hash SIG-7A21.'/'...SIG-9C40.') are complete factual sentences; the counterfactual's swap to SIG-9C40 coherently changes only the prior-run signature without contradicting any other stated fact, and neither context states or implies the final readiness label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21. A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-7A21.\"},{\"speaker\":\"Build engineer\",\"text\":\"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally. This is currently the only in-scope required-test failure for PR 482.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21."}, {"path": ["1", "text"], "text": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-7A21."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21.", "negative_left": "The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21.", "negative_right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-9C40.", "right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-7A21."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-011", "id": "fast-43-diverse-132-011-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21. A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-7A21."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally. This is currently the only in-scope required-test failure for PR 482."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the scope rule and flake-exception policy verbatim, keep the PR 482/Windows cli_integration bindings intact, and the two focus evidence spans ('The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21.' and 'A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with signature hash SIG-7A21.'/'...SIG-9C40.') are complete factual sentences; the counterfactual's swap to SIG-9C40 coherently changes only the prior-run signature without contradicting any other stated fact, and neither context states or implies the final readiness label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21. A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-7A21.\"},{\"speaker\":\"Build engineer\",\"text\":\"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally. This is currently the only in-scope required-test failure for PR 482.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21."}, {"path": ["1", "text"], "text": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-7A21."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21.", "negative_left": "The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21.", "negative_right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-9C40.", "right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-7A21."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-011", "id": "fast-43-diverse-132-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure. The failure signature from PR 482's current Windows cli_integration run is logged as signature hash SIG-7A21. A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, is logged with failure signature hash SIG-9C40."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally. This is currently the only in-scope required-test failure for PR 482."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical policy statement, question bindings (PR 482, Windows cli_integration, ParserCache), and no labels or rule tables are embedded; the counterfactual coherently swaps the main run's failure signature to SIG-90C without contradicting other facts, and the focus evidence consists of two complete factual sentences.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure, and no other in-scope required test failed in this PR run. PR 482's current Windows cli_integration run failed with signature SIG-77A.\"},{\"speaker\":\"Build engineer\",\"text\":\"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, also failed with signature SIG-77A.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run failed with signature SIG-77A."}, {"path": ["3", "text"], "text": "A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, also failed with signature SIG-77A."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run failed with signature SIG-77A.", "negative_left": "PR 482's current Windows cli_integration run failed with signature SIG-77A.", "negative_right": "A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, failed with signature SIG-90C.", "right": "A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, also failed with signature SIG-77A."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-013", "id": "fast-43-diverse-132-013-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure, and no other in-scope required test failed in this PR run. PR 482's current Windows cli_integration run failed with signature SIG-77A."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, also failed with signature SIG-77A."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical policy statement, question bindings (PR 482, Windows cli_integration, ParserCache), and no labels or rule tables are embedded; the counterfactual coherently swaps the main run's failure signature to SIG-90C without contradicting other facts, and the focus evidence consists of two complete factual sentences.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure, and no other in-scope required test failed in this PR run. PR 482's current Windows cli_integration run failed with signature SIG-77A.\"},{\"speaker\":\"Build engineer\",\"text\":\"The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, also failed with signature SIG-77A.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run failed with signature SIG-77A."}, {"path": ["3", "text"], "text": "A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, also failed with signature SIG-77A."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run failed with signature SIG-77A.", "negative_left": "PR 482's current Windows cli_integration run failed with signature SIG-77A.", "negative_right": "A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, failed with signature SIG-90C.", "right": "A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, also failed with signature SIG-77A."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-013", "id": "fast-43-diverse-132-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. This run contains no assertion failure, and no other in-scope required test failed in this PR run. PR 482's current Windows cli_integration run failed with signature SIG-77A."}, {"speaker": "Build engineer", "text": "The agent had 41% free disk, stable network, and no runner warnings. Compilation and environment setup completed normally."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change. A Windows cli_integration run on main completed on June 3, 2024, before the current PR 482 run, failed with signature SIG-90C."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scope rule and flake-exception policy verbatim and unchanged; the question object is preserved externally. The counterfactual changes the March 3 signature from SIG-77A2 to SIG-91C4, which is a coherent factual alteration consistent with the rest of the narrative (no duplicate contradictory claims). Focus evidence spans are two complete factual sentences describing test run outcomes, not policy text. Neither context states a label, score, or instructs the classifier on output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error. This is the only current in-scope required-test failure for PR 482.\"},{\"speaker\":\"Build engineer\",\"text\":\"A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, also failed with signature SIG-77A2.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error."}, {"path": ["2", "text"], "text": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, also failed with signature SIG-77A2."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error.", "negative_left": "PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error.", "negative_right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, failed with signature SIG-91C4 instead.", "right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, also failed with signature SIG-77A2."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-015", "id": "fast-43-diverse-132-015-base", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error. This is the only current in-scope required-test failure for PR 482."}, {"speaker": "Build engineer", "text": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, also failed with signature SIG-77A2."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-02", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the scope rule and flake-exception policy verbatim and unchanged; the question object is preserved externally. The counterfactual changes the March 3 signature from SIG-77A2 to SIG-91C4, which is a coherent factual alteration consistent with the rest of the narrative (no duplicate contradictory claims). Focus evidence spans are two complete factual sentences describing test run outcomes, not policy text. Neither context states a label, score, or instructs the classifier on output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a6 is a permissible quantified factual claim rather than a schema classification. The focus a4 is factual, concerning whether a matching prior main-run failure exists. The base and counter assignments can differ only in that historical-match fact without violating the policy. The policy evidence correctly preserves the substantive policy originating in the original state, while case-specific observations need not be retained.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a sole current in-scope failure, a non-assertion signature previously seen on main, and no public-interface change. The stated known-flake exception therefore applies, making target 1 (conditionally ready, with one clean rerun required) sufficient.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish a current in-scope failure and refute any prior matching signature on main. Because a prior matching non-assertion signature is required for the only stated exception, that exception is unavailable regardless of the omitted a3 and a5 states, so target 0 (not ready) follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "PR 482's current Windows cli_integration run has a test-failure status."}, {"id": "a2", "statement": "PR 482's current Windows cli_integration run is within the affected test scope of PR 482."}, {"id": "a3", "statement": "The failure signature from PR 482's current Windows cli_integration run is a non-assertion signature."}, {"id": "a4", "statement": "The failure signature from PR 482's current Windows cli_integration run matches the failure signature of at least one Windows cli_integration run on main completed before the current PR 482 run."}, {"id": "a5", "statement": "PR 482 makes no public-interface change."}, {"id": "a6", "statement": "Every current in-scope required-test failure for PR 482 is the failure from the current Windows cli_integration run."}], "base_state_json": "[{\"speaker\":\"Software developer\",\"text\":\"PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds.\"},{\"speaker\":\"Test automation engineer\",\"text\":\"The CLI consumes ParserCache, so cli_integration is in scope. PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error. This is the only current in-scope required-test failure for PR 482.\"},{\"speaker\":\"Build engineer\",\"text\":\"A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, also failed with signature SIG-77A2.\"},{\"speaker\":\"Code reviewer\",\"text\":\"Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error."}, {"path": ["2", "text"], "text": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, also failed with signature SIG-77A2."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}], "rules": [{"justification": "The sole current in-scope failure has a non-assertion signature previously observed on main, and PR 482 has no public-interface change. The known-flake exception therefore applies, making the PR conditionally ready pending one clean rerun.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}]}, {"justification": "An in-scope failure exists, and its signature did not previously occur on main. The known-flake exception is therefore unavailable, so the PR is not ready.", "target": "0", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error.", "negative_left": "PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error.", "negative_right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, failed with signature SIG-91C4 instead.", "right": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, also failed with signature SIG-77A2."}, "verifier_independent_model": false}, "family": "fast-43-diverse-132-015", "id": "fast-43-diverse-132-015-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready: an in-scope failure exists and the known-flake exception is unavailable, so code changes or failure diagnosis are required before merging.", "Conditionally ready: an in-scope failure exists, but the known-flake exception applies; one clean rerun is required before merging.", "Ready now: all in-scope required tests passed, or every failure is outside the affected scope, so no rerun or remediation is required."], "instructions": "Determine the PR's current merge-readiness level under the stated scope rule and known-flake exception.", "type": "score"}}, "state": [{"speaker": "Software developer", "text": "PR 482 changes only ParserCache internals; public interfaces and dependency files are unchanged. Linux parser unit tests passed, but Windows cli_integration timed out after 120 seconds."}, {"speaker": "Test automation engineer", "text": "The CLI consumes ParserCache, so cli_integration is in scope. PR 482's current Windows cli_integration run failed with signature SIG-77A2, a non-assertion timeout error. This is the only current in-scope required-test failure for PR 482."}, {"speaker": "Build engineer", "text": "A Windows cli_integration run on main completed on March 3, prior to the current PR 482 run, failed with signature SIG-91C4 instead."}, {"speaker": "Code reviewer", "text": "Policy: an in-scope failure blocks merging unless the known-flake exception applies. That exception permits conditional merging after one clean rerun when the same non-assertion signature previously occurred on main; it never applies after a public-interface change."}]}, "method": "c2d", "provenance": {"source_id": "diverse-132", "source_is_synthetic": true, "source_sha256": "3266707b2279f38010ea87b405597a3be28fca9ca4d914b4d1dd057ce7cce805", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-02", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Runbook policy and question bindings are unchanged verbatim; the two focus sentences are factual (registry mapping and OWNER_2 value); the counterfactual only flips OWNER_2's returned string, remaining consistent with the rest of the unchanged facts; neither context states a final routing conclusion, avoiding answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual assignment of R-17 to Pricing rather than a policy conclusion. The policy evidence preserves the substantive runbook rule originating in the original state; instructions and criteria in the retained questions object need not be repeated. The base and counter assignments are jointly realizable with only a8 changing: in the base R-17 is assigned to Pricing, while in the counter it is not assigned to Pricing and, given a6 and a7, is assigned uniquely to Catalog. Both rules include sufficient evidence and the necessary competing-service exclusion.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that at least three traces identify R-17 on the earliest failing span, an independent failure metric agrees on R-17 during the alert interval, and R-17 is uniquely assigned to Pricing. Under the cited runbook, the alert must therefore be routed to Pricing, which entails the false target.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same qualifying trace-and-metric agreement. Because R-17 is uniquely assigned to a service in {Catalog, Pricing} and its assignment to Pricing is refuted, it must be assigned to Catalog. The runbook therefore entails routing to Catalog, satisfying the true target.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of sampled traces for the Catalog API alert under review whose earliest failing span carries service-resource identifier R-17 is at least three."}, {"id": "a2", "statement": "Metric M-4 is independent of the sampled traces for the Catalog API alert under review."}, {"id": "a3", "statement": "Metric M-4 indicates a failure during the six-minute interval of the Catalog API alert under review."}, {"id": "a4", "statement": "Metric M-4 carries service-resource identifier R-17."}, {"id": "a5", "statement": "Exactly one service-resource identifier appears on the earliest failing span in at least three sampled traces for the Catalog API alert under review."}, {"id": "a6", "statement": "The service assigned service-resource identifier R-17 is a member of the set consisting of the Catalog service and the Pricing service."}, {"id": "a7", "statement": "Exactly one service is assigned service-resource identifier R-17."}, {"id": "a8", "statement": "Service-resource identifier R-17 is assigned to the Pricing service."}], "base_state_json": "\"The site reliability engineer reviews a Catalog API alert showing 11% HTTP 500s for six minutes. Eighteen sampled traces were reviewed; at least three show the earliest failing span carrying service-resource identifier R-17, and exactly one identifier appears on that earliest failing span across those traces. Metric M-4, collected independently of the traces, indicates a failure during the six-minute alert interval and carries service-resource identifier R-17. R-17 is assigned to exactly one service, drawn from the set {Catalog service, Pricing service}. The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2. Configuration key OWNER_2 currently returns the string 'Pricing service'. A Catalog deployment completed eight minutes before the alert fired. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2."}, {"path": [], "text": "Configuration key OWNER_2 currently returns the string 'Pricing service'."}], "policy_evidence": [{"path": [], "text": "The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}], "rules": [{"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, R-17 is the sole trace-qualified identifier, and R-17 is uniquely assigned to Pricing. The runbook therefore routes the alert to Pricing, not Catalog.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, and R-17 is the sole trace-qualified identifier. Because R-17 has exactly one assignment, that assignment is within {Catalog, Pricing}, and it is explicitly not Pricing, R-17 is assigned to Catalog. The runbook therefore routes the alert to Catalog.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2.", "negative_left": "The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2.", "negative_right": "Configuration key OWNER_2 currently returns the string 'Catalog service'.", "right": "Configuration key OWNER_2 currently returns the string 'Pricing service'."}, "verifier_independent_model": false}, "family": "fast-43-diverse-135-007", "id": "fast-43-diverse-135-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the alert to the Catalog service owner; route it to the Pricing service owner.", "true": "Yes — route the alert to the Catalog service owner."}, "instructions": "Decide whether the alert should be routed to the Catalog service owner under the supplied runbook.", "type": "noul"}}, "state": "The site reliability engineer reviews a Catalog API alert showing 11% HTTP 500s for six minutes. Eighteen sampled traces were reviewed; at least three show the earliest failing span carrying service-resource identifier R-17, and exactly one identifier appears on that earliest failing span across those traces. Metric M-4, collected independently of the traces, indicates a failure during the six-minute alert interval and carries service-resource identifier R-17. R-17 is assigned to exactly one service, drawn from the set {Catalog service, Pricing service}. The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2. Configuration key OWNER_2 currently returns the string 'Pricing service'. A Catalog deployment completed eight minutes before the alert fired. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}, "method": "c2d", "provenance": {"source_id": "diverse-135", "source_is_synthetic": true, "source_sha256": "7c01851c92ec31a0d400586f2e516fdb19e7f2a02cf1d4c8fb1140fb2791f3a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Runbook policy and question bindings are unchanged verbatim; the two focus sentences are factual (registry mapping and OWNER_2 value); the counterfactual only flips OWNER_2's returned string, remaining consistent with the rest of the unchanged facts; neither context states a final routing conclusion, avoiding answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual assignment of R-17 to Pricing rather than a policy conclusion. The policy evidence preserves the substantive runbook rule originating in the original state; instructions and criteria in the retained questions object need not be repeated. The base and counter assignments are jointly realizable with only a8 changing: in the base R-17 is assigned to Pricing, while in the counter it is not assigned to Pricing and, given a6 and a7, is assigned uniquely to Catalog. Both rules include sufficient evidence and the necessary competing-service exclusion.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that at least three traces identify R-17 on the earliest failing span, an independent failure metric agrees on R-17 during the alert interval, and R-17 is uniquely assigned to Pricing. Under the cited runbook, the alert must therefore be routed to Pricing, which entails the false target.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish the same qualifying trace-and-metric agreement. Because R-17 is uniquely assigned to a service in {Catalog, Pricing} and its assignment to Pricing is refuted, it must be assigned to Catalog. The runbook therefore entails routing to Catalog, satisfying the true target.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The number of sampled traces for the Catalog API alert under review whose earliest failing span carries service-resource identifier R-17 is at least three."}, {"id": "a2", "statement": "Metric M-4 is independent of the sampled traces for the Catalog API alert under review."}, {"id": "a3", "statement": "Metric M-4 indicates a failure during the six-minute interval of the Catalog API alert under review."}, {"id": "a4", "statement": "Metric M-4 carries service-resource identifier R-17."}, {"id": "a5", "statement": "Exactly one service-resource identifier appears on the earliest failing span in at least three sampled traces for the Catalog API alert under review."}, {"id": "a6", "statement": "The service assigned service-resource identifier R-17 is a member of the set consisting of the Catalog service and the Pricing service."}, {"id": "a7", "statement": "Exactly one service is assigned service-resource identifier R-17."}, {"id": "a8", "statement": "Service-resource identifier R-17 is assigned to the Pricing service."}], "base_state_json": "\"The site reliability engineer reviews a Catalog API alert showing 11% HTTP 500s for six minutes. Eighteen sampled traces were reviewed; at least three show the earliest failing span carrying service-resource identifier R-17, and exactly one identifier appears on that earliest failing span across those traces. Metric M-4, collected independently of the traces, indicates a failure during the six-minute alert interval and carries service-resource identifier R-17. R-17 is assigned to exactly one service, drawn from the set {Catalog service, Pricing service}. The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2. Configuration key OWNER_2 currently returns the string 'Pricing service'. A Catalog deployment completed eight minutes before the alert fired. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}], "focus_atom": "a8", "focus_evidence": [{"path": [], "text": "The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2."}, {"path": [], "text": "Configuration key OWNER_2 currently returns the string 'Pricing service'."}], "policy_evidence": [{"path": [], "text": "The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}], "rules": [{"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, R-17 is the sole trace-qualified identifier, and R-17 is uniquely assigned to Pricing. The runbook therefore routes the alert to Pricing, not Catalog.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "At least three traces identify R-17 on the earliest failing span, independent metric M-4 agrees on R-17 during the alert interval, and R-17 is the sole trace-qualified identifier. Because R-17 has exactly one assignment, that assignment is within {Catalog, Pricing}, and it is explicitly not Pricing, R-17 is assigned to Catalog. The runbook therefore routes the alert to Catalog.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}]}]}, "verified_pair": {"left": "The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2.", "negative_left": "The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2.", "negative_right": "Configuration key OWNER_2 currently returns the string 'Catalog service'.", "right": "Configuration key OWNER_2 currently returns the string 'Pricing service'."}, "verifier_independent_model": false}, "family": "fast-43-diverse-135-007", "id": "fast-43-diverse-135-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the alert to the Catalog service owner; route it to the Pricing service owner.", "true": "Yes — route the alert to the Catalog service owner."}, "instructions": "Decide whether the alert should be routed to the Catalog service owner under the supplied runbook.", "type": "noul"}}, "state": "The site reliability engineer reviews a Catalog API alert showing 11% HTTP 500s for six minutes. Eighteen sampled traces were reviewed; at least three show the earliest failing span carrying service-resource identifier R-17, and exactly one identifier appears on that earliest failing span across those traces. Metric M-4, collected independently of the traces, indicates a failure during the six-minute alert interval and carries service-resource identifier R-17. R-17 is assigned to exactly one service, drawn from the set {Catalog service, Pricing service}. The infrastructure ownership registry lists service-resource identifier R-17 under the service whose name matches the string returned by configuration key OWNER_2. Configuration key OWNER_2 currently returns the string 'Catalog service'. A Catalog deployment completed eight minutes before the alert fired. The runbook says to route an alert to the service with the earliest failing span when at least three traces and one independent metric agree; deployment timing alone is insufficient to establish ownership."}, "method": "c2d", "provenance": {"source_id": "diverse-135", "source_is_synthetic": true, "source_sha256": "7c01851c92ec31a0d400586f2e516fdb19e7f2a02cf1d4c8fb1140fb2791f3a4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Runbook policy and question bindings are unchanged in both contexts; the five evidence sentences are factual, not policy or instructions; the counterfactual only swaps 'Payments' for 'Checkout' in the ownership sentence, which is a coherent single-fact change (odd naming aside) that does not contradict other evidence; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"Trace T1 for the ShopWave production alert has completed and identifies exactly one service as the slow component.\",\"The slow component identified by trace T1 has service identifier SC-17.\",\"The Payments dependency dashboard has no data for the relevant interval, so it does not confirm Payments degradation.\",\"The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'.\",\"Team 'Payments-Team-Alpha' owns and operates the Payments service exclusively.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'."}, {"path": ["evidence", "4"], "text": "Team 'Payments-Team-Alpha' owns and operates the Payments service exclusively."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'.", "negative_left": "The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'.", "negative_right": "Team 'Payments-Team-Alpha' owns and operates the Checkout service exclusively.", "right": "Team 'Payments-Team-Alpha' owns and operates the Payments service exclusively."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-001", "id": "fast-43-diverse-136-001-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for the ShopWave production alert has completed and identifies exactly one service as the slow component.", "The slow component identified by trace T1 has service identifier SC-17.", "The Payments dependency dashboard has no data for the relevant interval, so it does not confirm Payments degradation.", "The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'.", "Team 'Payments-Team-Alpha' owns and operates the Payments service exclusively."]}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Runbook policy and question bindings are unchanged in both contexts; the five evidence sentences are factual, not policy or instructions; the counterfactual only swaps 'Payments' for 'Checkout' in the ownership sentence, which is a coherent single-fact change (odd naming aside) that does not contradict other evidence; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"Trace T1 for the ShopWave production alert has completed and identifies exactly one service as the slow component.\",\"The slow component identified by trace T1 has service identifier SC-17.\",\"The Payments dependency dashboard has no data for the relevant interval, so it does not confirm Payments degradation.\",\"The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'.\",\"Team 'Payments-Team-Alpha' owns and operates the Payments service exclusively.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'."}, {"path": ["evidence", "4"], "text": "Team 'Payments-Team-Alpha' owns and operates the Payments service exclusively."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'.", "negative_left": "The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'.", "negative_right": "Team 'Payments-Team-Alpha' owns and operates the Checkout service exclusively.", "right": "Team 'Payments-Team-Alpha' owns and operates the Payments service exclusively."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-001", "id": "fast-43-diverse-136-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for the ShopWave production alert has completed and identifies exactly one service as the slow component.", "The slow component identified by trace T1 has service identifier SC-17.", "The Payments dependency dashboard has no data for the relevant interval, so it does not confirm Payments degradation.", "The ShopWave service registry entry for identifier SC-17 lists the owning team name as 'Payments-Team-Alpha'.", "Team 'Payments-Team-Alpha' owns and operates the Checkout service exclusively."]}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing runbook policy verbatim, matching the unchanged question instructions. The counterfactual only alters the registry_lookup fact (SC-17 now under Shipping-Team) which is a legitimate case-observation change, not a shift in question entity/path/time bindings, and remains logically consistent with the unchanged registry_scope statement. The two focus evidence spans are complete factual sentences describing registry facts, not policy or instruction text. Neither context contains a gold answer, rule table, or label rationale, and no answer leakage is present.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"trace_status\":\"Trace T1 for the alert has completed processing.\",\"trace_result\":\"Trace T1 identifies exactly one service as the slow component, and that service's identifier is SC-17.\",\"registry_lookup\":\"The internal service registry lists identifier SC-17 under the ownership group named Payments-Team.\",\"registry_scope\":\"The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service.\",\"dashboard_status\":\"The Payments dependency dashboard has no data for the relevant interval and therefore does not confirm any Payments degradation.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["registry_lookup"], "text": "The internal service registry lists identifier SC-17 under the ownership group named Payments-Team."}, {"path": ["registry_scope"], "text": "The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal service registry lists identifier SC-17 under the ownership group named Payments-Team.", "negative_left": "The internal service registry lists identifier SC-17 under the ownership group named Shipping-Team.", "negative_right": "The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service.", "right": "The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-026", "id": "fast-43-diverse-136-026-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "dashboard_status": "The Payments dependency dashboard has no data for the relevant interval and therefore does not confirm any Payments degradation.", "registry_lookup": "The internal service registry lists identifier SC-17 under the ownership group named Payments-Team.", "registry_scope": "The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service.", "trace_result": "Trace T1 identifies exactly one service as the slow component, and that service's identifier is SC-17.", "trace_status": "Trace T1 for the alert has completed processing."}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing runbook policy verbatim, matching the unchanged question instructions. The counterfactual only alters the registry_lookup fact (SC-17 now under Shipping-Team) which is a legitimate case-observation change, not a shift in question entity/path/time bindings, and remains logically consistent with the unchanged registry_scope statement. The two focus evidence spans are complete factual sentences describing registry facts, not policy or instruction text. Neither context contains a gold answer, rule table, or label rationale, and no answer leakage is present.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"trace_status\":\"Trace T1 for the alert has completed processing.\",\"trace_result\":\"Trace T1 identifies exactly one service as the slow component, and that service's identifier is SC-17.\",\"registry_lookup\":\"The internal service registry lists identifier SC-17 under the ownership group named Payments-Team.\",\"registry_scope\":\"The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service.\",\"dashboard_status\":\"The Payments dependency dashboard has no data for the relevant interval and therefore does not confirm any Payments degradation.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["registry_lookup"], "text": "The internal service registry lists identifier SC-17 under the ownership group named Payments-Team."}, {"path": ["registry_scope"], "text": "The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal service registry lists identifier SC-17 under the ownership group named Payments-Team.", "negative_left": "The internal service registry lists identifier SC-17 under the ownership group named Shipping-Team.", "negative_right": "The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service.", "right": "The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-026", "id": "fast-43-diverse-136-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "dashboard_status": "The Payments dependency dashboard has no data for the relevant interval and therefore does not confirm any Payments degradation.", "registry_lookup": "The internal service registry lists identifier SC-17 under the ownership group named Shipping-Team.", "registry_scope": "The Payments-Team ownership group in the service registry is exclusively responsible for the Payments service.", "trace_result": "Trace T1 identifies exactly one service as the slow component, and that service's identifier is SC-17.", "trace_status": "Trace T1 for the alert has completed processing."}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same runbook policy verbatim and the unchanged question supplies the routing criteria; the counterfactual only swaps the registry's service-name mapping (Payments to Inventory), which coherently changes whether SC-17 maps to Payments without contradicting the shared trace or dashboard facts; focus evidence consists of two complete factual sentences from registry_check, not policy text; neither context states or hints at the final routing decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"trace_status\":\"Trace T1 for the alert is completed and identifies exactly one slow component, with service identifier SC-17.\",\"registry_check\":\"The internal service registry lists SC-17 under the microservices catalog entry updated this quarter. That catalog entry's service name field reads Payments.\",\"dashboard_status\":\"The Payments dependency dashboard has no data for the relevant interval, so it does not confirm Payments degradation during that window.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["registry_check"], "text": "The internal service registry lists SC-17 under the microservices catalog entry updated this quarter."}, {"path": ["registry_check"], "text": "That catalog entry's service name field reads Payments."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal service registry lists SC-17 under the microservices catalog entry updated this quarter.", "negative_left": "The internal service registry lists SC-17 under the microservices catalog entry updated this quarter.", "negative_right": "That catalog entry's service name field reads Inventory.", "right": "That catalog entry's service name field reads Payments."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-035", "id": "fast-43-diverse-136-035-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "dashboard_status": "The Payments dependency dashboard has no data for the relevant interval, so it does not confirm Payments degradation during that window.", "registry_check": "The internal service registry lists SC-17 under the microservices catalog entry updated this quarter. That catalog entry's service name field reads Payments.", "trace_status": "Trace T1 for the alert is completed and identifies exactly one slow component, with service identifier SC-17."}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same runbook policy verbatim and the unchanged question supplies the routing criteria; the counterfactual only swaps the registry's service-name mapping (Payments to Inventory), which coherently changes whether SC-17 maps to Payments without contradicting the shared trace or dashboard facts; focus evidence consists of two complete factual sentences from registry_check, not policy text; neither context states or hints at the final routing decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"trace_status\":\"Trace T1 for the alert is completed and identifies exactly one slow component, with service identifier SC-17.\",\"registry_check\":\"The internal service registry lists SC-17 under the microservices catalog entry updated this quarter. That catalog entry's service name field reads Payments.\",\"dashboard_status\":\"The Payments dependency dashboard has no data for the relevant interval, so it does not confirm Payments degradation during that window.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["registry_check"], "text": "The internal service registry lists SC-17 under the microservices catalog entry updated this quarter."}, {"path": ["registry_check"], "text": "That catalog entry's service name field reads Payments."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal service registry lists SC-17 under the microservices catalog entry updated this quarter.", "negative_left": "The internal service registry lists SC-17 under the microservices catalog entry updated this quarter.", "negative_right": "That catalog entry's service name field reads Inventory.", "right": "That catalog entry's service name field reads Payments."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-035", "id": "fast-43-diverse-136-035-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "dashboard_status": "The Payments dependency dashboard has no data for the relevant interval, so it does not confirm Payments degradation during that window.", "registry_check": "The internal service registry lists SC-17 under the microservices catalog entry updated this quarter. That catalog entry's service name field reads Inventory.", "trace_status": "Trace T1 for the alert is completed and identifies exactly one slow component, with service identifier SC-17."}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The routing runbook policy is repeated verbatim in both contexts and the original question/criteria are unchanged, so governing policy and question bindings (Payments, trace T1, dependency dashboard, relevant interval) are preserved; the focus evidence are complete factual sentences describing registry/catalog facts rather than policy language; the counterfactual only swaps the catalog-to-service mapping from Payments to Notifications, a single coherent factual change that does not contradict any other evidence (dashboard and trace-completion facts remain intact); and neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\": \"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\", \"evidence\": [\"Trace T1 for the alert has completed processing and returned a full span set.\", \"Trace T1 identifies exactly one service as the slow component in the request path.\", \"The slow component identified by trace T1 carries service identifier SC-17.\", \"The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations.\", \"The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Payments service.\", \"The Payments dependency dashboard shows no anomalies and does not confirm any Payments degradation for the relevant interval.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations."}, {"path": ["evidence", "4"], "text": "The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations.", "negative_left": "The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations.", "negative_right": "The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Notifications service.", "right": "The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-037", "id": "fast-43-diverse-136-037-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for the alert has completed processing and returned a full span set.", "Trace T1 identifies exactly one service as the slow component in the request path.", "The slow component identified by trace T1 carries service identifier SC-17.", "The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations.", "The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Payments service.", "The Payments dependency dashboard shows no anomalies and does not confirm any Payments degradation for the relevant interval."]}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The routing runbook policy is repeated verbatim in both contexts and the original question/criteria are unchanged, so governing policy and question bindings (Payments, trace T1, dependency dashboard, relevant interval) are preserved; the focus evidence are complete factual sentences describing registry/catalog facts rather than policy language; the counterfactual only swaps the catalog-to-service mapping from Payments to Notifications, a single coherent factual change that does not contradict any other evidence (dashboard and trace-completion facts remain intact); and neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\": \"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\", \"evidence\": [\"Trace T1 for the alert has completed processing and returned a full span set.\", \"Trace T1 identifies exactly one service as the slow component in the request path.\", \"The slow component identified by trace T1 carries service identifier SC-17.\", \"The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations.\", \"The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Payments service.\", \"The Payments dependency dashboard shows no anomalies and does not confirm any Payments degradation for the relevant interval.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "3"], "text": "The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations."}, {"path": ["evidence", "4"], "text": "The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations.", "negative_left": "The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations.", "negative_right": "The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Notifications service.", "right": "The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-037", "id": "fast-43-diverse-136-037-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for the alert has completed processing and returned a full span set.", "Trace T1 identifies exactly one service as the slow component in the request path.", "The slow component identified by trace T1 carries service identifier SC-17.", "The ShopWave internal service registry lists service identifier SC-17 under the catalog entry for checkout-related billing operations.", "The catalog entry for checkout-related billing operations in the ShopWave internal service registry corresponds to the Notifications service.", "The Payments dependency dashboard shows no anomalies and does not confirm any Payments degradation for the relevant interval."]}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The runbook policy is unchanged and the original questions object is presumed carried verbatim, so criteria and instructions remain intact; the focus evidence sentences are both complete factual statements about registry mappings rather than policy text; the counterfactual's single change (PAY-CORE labeled Shipping instead of Payments) is a coherent alternative that does not contradict the unaltered trace or dashboard evidence; neither context states or implies the final routing decision, only raw facts for reasoning.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\": \"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\", \"evidence\": [\"Trace T1 for the alert completed successfully and identifies exactly one service, bearing identifier SC-17, as the sole slow component.\", \"The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE.\", \"Service record PAY-CORE is the designation used for the Payments service.\", \"The Payments dependency dashboard has no data for the relevant interval, so it does not confirm any Payments degradation.\"], \"request\": \"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "1"], "text": "The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE."}, {"path": ["evidence", "2"], "text": "Service record PAY-CORE is the designation used for the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE.", "negative_left": "The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE.", "negative_right": "Service record PAY-CORE is the designation used for the Shipping service.", "right": "Service record PAY-CORE is the designation used for the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-039", "id": "fast-43-diverse-136-039-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for the alert completed successfully and identifies exactly one service, bearing identifier SC-17, as the sole slow component.", "The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE.", "Service record PAY-CORE is the designation used for the Payments service.", "The Payments dependency dashboard has no data for the relevant interval, so it does not confirm any Payments degradation."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The runbook policy is unchanged and the original questions object is presumed carried verbatim, so criteria and instructions remain intact; the focus evidence sentences are both complete factual statements about registry mappings rather than policy text; the counterfactual's single change (PAY-CORE labeled Shipping instead of Payments) is a coherent alternative that does not contradict the unaltered trace or dashboard evidence; neither context states or implies the final routing decision, only raw facts for reasoning.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\": \"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\", \"evidence\": [\"Trace T1 for the alert completed successfully and identifies exactly one service, bearing identifier SC-17, as the sole slow component.\", \"The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE.\", \"Service record PAY-CORE is the designation used for the Payments service.\", \"The Payments dependency dashboard has no data for the relevant interval, so it does not confirm any Payments degradation.\"], \"request\": \"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "1"], "text": "The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE."}, {"path": ["evidence", "2"], "text": "Service record PAY-CORE is the designation used for the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE.", "negative_left": "The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE.", "negative_right": "Service record PAY-CORE is the designation used for the Shipping service.", "right": "Service record PAY-CORE is the designation used for the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-039", "id": "fast-43-diverse-136-039-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for the alert completed successfully and identifies exactly one service, bearing identifier SC-17, as the sole slow component.", "The internal service registry lists identifier SC-17 under a service record labeled with code PAY-CORE.", "Service record PAY-CORE is the designation used for the Shipping service.", "The Payments dependency dashboard has no data for the relevant interval, so it does not confirm any Payments degradation."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the runbook policy and dashboard/trace evidence structure unchanged from the question's routing logic; the counterfactual only swaps the registry's mapped service name (Payments to Inventory) for entry 4482, a single coherent factual change; the two focus evidence sentences are complete factual statements about the registry mapping and naming, not policy or instructions; no gold answer, rationale, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"trace_notes\":\"Trace T1 for this alert has finished processing and identifies exactly one service as the slow component, tagged with service identifier SC-17. The service registry maps identifier SC-17 to registry entry number 4482. Registry entry number 4482 is officially named the Payments service in the ShopWave service catalog. No other services were flagged as slow by the trace.\",\"dashboard_notes\":\"The Payments dependency dashboard shows no data for the relevant interval, so it does not confirm any Payments degradation during that window. No additional dashboards were consulted for this assessment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["trace_notes"], "text": "The service registry maps identifier SC-17 to registry entry number 4482."}, {"path": ["trace_notes"], "text": "Registry entry number 4482 is officially named the Payments service in the ShopWave service catalog."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The service registry maps identifier SC-17 to registry entry number 4482.", "negative_left": "The service registry maps identifier SC-17 to registry entry number 4482.", "negative_right": "Registry entry number 4482 is officially named the Inventory service in the ShopWave service catalog.", "right": "Registry entry number 4482 is officially named the Payments service in the ShopWave service catalog."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-040", "id": "fast-43-diverse-136-040-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "dashboard_notes": "The Payments dependency dashboard shows no data for the relevant interval, so it does not confirm any Payments degradation during that window. No additional dashboards were consulted for this assessment.", "trace_notes": "Trace T1 for this alert has finished processing and identifies exactly one service as the slow component, tagged with service identifier SC-17. The service registry maps identifier SC-17 to registry entry number 4482. Registry entry number 4482 is officially named the Payments service in the ShopWave service catalog. No other services were flagged as slow by the trace."}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the runbook policy and dashboard/trace evidence structure unchanged from the question's routing logic; the counterfactual only swaps the registry's mapped service name (Payments to Inventory) for entry 4482, a single coherent factual change; the two focus evidence sentences are complete factual statements about the registry mapping and naming, not policy or instructions; no gold answer, rationale, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"trace_notes\":\"Trace T1 for this alert has finished processing and identifies exactly one service as the slow component, tagged with service identifier SC-17. The service registry maps identifier SC-17 to registry entry number 4482. Registry entry number 4482 is officially named the Payments service in the ShopWave service catalog. No other services were flagged as slow by the trace.\",\"dashboard_notes\":\"The Payments dependency dashboard shows no data for the relevant interval, so it does not confirm any Payments degradation during that window. No additional dashboards were consulted for this assessment.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["trace_notes"], "text": "The service registry maps identifier SC-17 to registry entry number 4482."}, {"path": ["trace_notes"], "text": "Registry entry number 4482 is officially named the Payments service in the ShopWave service catalog."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The service registry maps identifier SC-17 to registry entry number 4482.", "negative_left": "The service registry maps identifier SC-17 to registry entry number 4482.", "negative_right": "Registry entry number 4482 is officially named the Inventory service in the ShopWave service catalog.", "right": "Registry entry number 4482 is officially named the Payments service in the ShopWave service catalog."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-040", "id": "fast-43-diverse-136-040-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "dashboard_notes": "The Payments dependency dashboard shows no data for the relevant interval, so it does not confirm any Payments degradation during that window. No additional dashboards were consulted for this assessment.", "trace_notes": "Trace T1 for this alert has finished processing and identifies exactly one service as the slow component, tagged with service identifier SC-17. The service registry maps identifier SC-17 to registry entry number 4482. Registry entry number 4482 is officially named the Inventory service in the ShopWave service catalog. No other services were flagged as slow by the trace."}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original runbook policy and question bindings (Payments entity, same request), the focus evidence consists of two complete factual sentences about registry mapping, the counterfactual coherently swaps only the registry designation (Payments→Checkout) without contradicting other evidence, and neither context contains a gold answer, rule table, or output instruction beyond the unchanged question object.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\": \"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\", \"evidence\": [\"Trace T1 for this alert has finished processing and identifies exactly one service as the slow component, tagged with service identifier SC-17.\", \"The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204.\", \"Registry entry R-204 in the ShopWave internal service registry is designated for the Payments service.\", \"The Payments dependency dashboard shows no data logged for the relevant interval, so it cannot confirm any Payments degradation during the alert window.\", \"No other trace or dashboard evidence is available for this incident.\"], \"request\": \"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "1"], "text": "The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204."}, {"path": ["evidence", "2"], "text": "Registry entry R-204 in the ShopWave internal service registry is designated for the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204.", "negative_left": "The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204.", "negative_right": "Registry entry R-204 in the ShopWave internal service registry is designated for the Checkout service.", "right": "Registry entry R-204 in the ShopWave internal service registry is designated for the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-042", "id": "fast-43-diverse-136-042-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for this alert has finished processing and identifies exactly one service as the slow component, tagged with service identifier SC-17.", "The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204.", "Registry entry R-204 in the ShopWave internal service registry is designated for the Payments service.", "The Payments dependency dashboard shows no data logged for the relevant interval, so it cannot confirm any Payments degradation during the alert window.", "No other trace or dashboard evidence is available for this incident."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original runbook policy and question bindings (Payments entity, same request), the focus evidence consists of two complete factual sentences about registry mapping, the counterfactual coherently swaps only the registry designation (Payments→Checkout) without contradicting other evidence, and neither context contains a gold answer, rule table, or output instruction beyond the unchanged question object.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\": \"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\", \"evidence\": [\"Trace T1 for this alert has finished processing and identifies exactly one service as the slow component, tagged with service identifier SC-17.\", \"The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204.\", \"Registry entry R-204 in the ShopWave internal service registry is designated for the Payments service.\", \"The Payments dependency dashboard shows no data logged for the relevant interval, so it cannot confirm any Payments degradation during the alert window.\", \"No other trace or dashboard evidence is available for this incident.\"], \"request\": \"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "1"], "text": "The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204."}, {"path": ["evidence", "2"], "text": "Registry entry R-204 in the ShopWave internal service registry is designated for the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204.", "negative_left": "The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204.", "negative_right": "Registry entry R-204 in the ShopWave internal service registry is designated for the Checkout service.", "right": "Registry entry R-204 in the ShopWave internal service registry is designated for the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-042", "id": "fast-43-diverse-136-042-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for this alert has finished processing and identifies exactly one service as the slow component, tagged with service identifier SC-17.", "The ShopWave internal service registry lists service identifier SC-17 under registry entry R-204.", "Registry entry R-204 in the ShopWave internal service registry is designated for the Checkout service.", "The Payments dependency dashboard shows no data logged for the relevant interval, so it cannot confirm any Payments degradation during the alert window.", "No other trace or dashboard evidence is available for this incident."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the identical runbook policy and the original request/question bindings unchanged. The two focus evidence sentences are plain factual statements about registry mappings, not policy text. The counterfactual alters only the registry designation for R-204 from Payments to Checkout, a single coherent factual change with no contradiction to other evidence. Neither context reveals a gold answer, rule table, or output instruction beyond the unchanged question object.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"Trace T1 for this alert has finished processing and returned a complete span set.\",\"Trace T1 identifies exactly one service as the slow component, tagged with service identifier SC-17.\",\"The internal service registry lists identifier SC-17 under registry entry R-204.\",\"Registry entry R-204 in the internal service registry is designated for the Payments service.\",\"The Payments dependency dashboard shows no data for the relevant interval, so it does not confirm any Payments degradation.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The internal service registry lists identifier SC-17 under registry entry R-204."}, {"path": ["evidence", "3"], "text": "Registry entry R-204 in the internal service registry is designated for the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal service registry lists identifier SC-17 under registry entry R-204.", "negative_left": "The internal service registry lists identifier SC-17 under registry entry R-204.", "negative_right": "Registry entry R-204 in the internal service registry is designated for the Checkout service.", "right": "Registry entry R-204 in the internal service registry is designated for the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-043", "id": "fast-43-diverse-136-043-base", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for this alert has finished processing and returned a complete span set.", "Trace T1 identifies exactly one service as the slow component, tagged with service identifier SC-17.", "The internal service registry lists identifier SC-17 under registry entry R-204.", "Registry entry R-204 in the internal service registry is designated for the Payments service.", "The Payments dependency dashboard shows no data for the relevant interval, so it does not confirm any Payments degradation."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-03", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the identical runbook policy and the original request/question bindings unchanged. The two focus evidence sentences are plain factual statements about registry mappings, not policy text. The counterfactual alters only the registry designation for R-204 from Payments to Checkout, a single coherent factual change with no contradiction to other evidence. Neither context reveals a gold answer, rule table, or output instruction beyond the unchanged question object.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "refuted", "A5": "refuted"}, "remove_left": {"A4": "unknown"}, "remove_right": {"A4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A4": "unknown"}, "negative_pair": {"A4": "refuted"}, "negative_sentence": {"A4": "unknown"}, "positive_pair": {"A4": "supported"}, "right": {"A4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; A4 is a factual identifier-to-service relation rather than a policy proposition. The policy evidence preserves the governing runbook from the original state, while question-originating criteria and instructions need not be repeated. The base and counter assignments are realizable by changing only whether SC-17 denotes Payments. Both rules include enough evidence to determine whether either of the two routing conditions is satisfied.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1–A4 jointly entail that a completed trace identifies Payments as the slow component. This satisfies the trace-based routing condition, so the alert is routed to Payments and the answer to retaining it is false. A5 does not undermine that sufficient condition.", "rule_index": 0, "sound": true}, {"reason": "A1–A4 establish that the completed trace's sole identified slow component, bearing identifier SC-17, is not Payments, while A5 refutes dashboard confirmation. Thus neither routing condition is satisfied, so the SRE retains the alert and the answer is true.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Trace T1 for the ShopWave production alert under assessment is completed."}, {"id": "A2", "statement": "Trace T1 identifies exactly one service as a slow component."}, {"id": "A3", "statement": "The service identified by trace T1 as its slow component has service identifier SC-17."}, {"id": "A4", "statement": "Service identifier SC-17 denotes the Payments service."}, {"id": "A5", "statement": "The Payments dependency dashboard confirms Payments degradation during the relevant interval of the ShopWave production alert under assessment."}], "base_state_json": "{\"context\":\"An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.\",\"evidence\":[\"Trace T1 for this alert has finished processing and returned a complete span set.\",\"Trace T1 identifies exactly one service as the slow component, tagged with service identifier SC-17.\",\"The internal service registry lists identifier SC-17 under registry entry R-204.\",\"Registry entry R-204 in the internal service registry is designated for the Payments service.\",\"The Payments dependency dashboard shows no data for the relevant interval, so it does not confirm any Payments degradation.\"],\"request\":\"Should the site reliability engineer retain the alert rather than route it to the Payments service owner?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}], "focus_atom": "A4", "focus_evidence": [{"path": ["evidence", "2"], "text": "The internal service registry lists identifier SC-17 under registry entry R-204."}, {"path": ["evidence", "3"], "text": "Registry entry R-204 in the internal service registry is designated for the Payments service."}], "policy_evidence": [{"path": ["context"], "text": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation."}], "rules": [{"justification": "The trace is completed and identifies exactly one slow component. That component has identifier SC-17, which denotes Payments, so the completed-trace routing condition is satisfied even though the dashboard does not confirm degradation.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}]}, {"justification": "The completed trace identifies exactly one slow component, and that component's identifier SC-17 does not denote Payments. The dashboard also does not confirm Payments degradation, so neither routing condition is satisfied.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}]}]}, "verified_pair": {"left": "The internal service registry lists identifier SC-17 under registry entry R-204.", "negative_left": "The internal service registry lists identifier SC-17 under registry entry R-204.", "negative_right": "Registry entry R-204 in the internal service registry is designated for the Checkout service.", "right": "Registry entry R-204 in the internal service registry is designated for the Payments service."}, "verifier_independent_model": false}, "family": "fast-43-diverse-136-043", "id": "fast-43-diverse-136-043-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The alert is routed to the Payments service owner because at least one runbook condition conclusively identifies or confirms Payments degradation.", "true": "The alert remains with the site reliability engineer because the evidence does not satisfy either condition required to route it to Payments."}, "instructions": "Answer yes or no using only the stated routing runbook and evidence. Missing required routing evidence must be treated as unavailable, not inferred.", "type": "noul"}}, "state": {"context": "An on-call engineer is assessing a fictional ShopWave production alert. The routing runbook says to assign an alert to the Payments service owner only when a completed trace identifies Payments as the slow component or the dependency dashboard confirms Payments degradation. Otherwise, the site reliability engineer retains the alert for investigation.", "evidence": ["Trace T1 for this alert has finished processing and returned a complete span set.", "Trace T1 identifies exactly one service as the slow component, tagged with service identifier SC-17.", "The internal service registry lists identifier SC-17 under registry entry R-204.", "Registry entry R-204 in the internal service registry is designated for the Checkout service.", "The Payments dependency dashboard shows no data for the relevant interval, so it does not confirm any Payments degradation."], "request": "Should the site reliability engineer retain the alert rather than route it to the Payments service owner?"}}, "method": "c2d", "provenance": {"source_id": "diverse-136", "source_is_synthetic": true, "source_sha256": "b63e88d7b8a18cabc648ca6a9fb204b762f9d0eaa26a1590ffb3178b6236d951", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-03", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original criteria, batch C-184 entity, and 13:40/14:20 timeline; the counterfactual changes only the final hash sentence, creating a coherent hash-mismatch scenario relevant to the pipeline_integrity_hold criterion without duplicating or contradicting other facts; focus evidence consists of two complete factual sentences with no policy text or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\": \"Data quality analyst\", \"text\": \"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"}, {\"speaker\": \"Data engineer\", \"text\": \"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560.\"}, {\"speaker\": \"Application administrator\", \"text\": \"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f3a1c7e2b8d4560.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f3a1c7e2b8d4560."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560.", "negative_left": "The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 4b2e7a9f1d3c5680.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f3a1c7e2b8d4560."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-002", "id": "fast-43-diverse-139-002-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f3a1c7e2b8d4560."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original criteria, batch C-184 entity, and 13:40/14:20 timeline; the counterfactual changes only the final hash sentence, creating a coherent hash-mismatch scenario relevant to the pipeline_integrity_hold criterion without duplicating or contradicting other facts; focus evidence consists of two complete factual sentences with no policy text or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\": \"Data quality analyst\", \"text\": \"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"}, {\"speaker\": \"Data engineer\", \"text\": \"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560.\"}, {\"speaker\": \"Application administrator\", \"text\": \"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f3a1c7e2b8d4560.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f3a1c7e2b8d4560."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560.", "negative_left": "The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 4b2e7a9f1d3c5680.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f3a1c7e2b8d4560."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-002", "id": "fast-43-diverse-139-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 9f3a1c7e2b8d4560."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 4b2e7a9f1d3c5680."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full administrator policy statement (precedence, thresholds, hash-supersession rule) alongside the unchanged question rubric, and keep batch C-184, the 13:40/14:20 timeline, and row/percentage bindings intact. The two focus sentences are plain factual statements about the reported and registered hashes, not rule tables or instructions. The counterfactual changes only the reported hash value in one sentence, creating a coherent mismatch with the registered hash (triggering the hold path) rather than a contradictory duplicate measurement. Neither context states or codes the final disposition, so there is no answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation. The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015."}, {"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3c2b1e4f9a8d7c6b5e4f3a2b1c0d9e8f.", "negative_right": "The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015.", "right": "The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-010", "id": "fast-43-diverse-139-010-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation. The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full administrator policy statement (precedence, thresholds, hash-supersession rule) alongside the unchanged question rubric, and keep batch C-184, the 13:40/14:20 timeline, and row/percentage bindings intact. The two focus sentences are plain factual statements about the reported and registered hashes, not rule tables or instructions. The counterfactual changes only the reported hash value in one sentence, creating a coherent mismatch with the registered hash (triggering the hold path) rather than a contradictory duplicate measurement. Neither context states or codes the final disposition, so there is no answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation. The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015."}, {"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3c2b1e4f9a8d7c6b5e4f3a2b1c0d9e8f.", "negative_right": "The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015.", "right": "The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-010", "id": "fast-43-diverse-139-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3c2b1e4f9a8d7c6b5e4f3a2b1c0d9e8f."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation. The immutable SHA-256 hash of batch C-184 is registered as 9f86d081884c7d659a2feaa0c55ad015."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy and question bindings for batch C-184 at 13:40/14:20; the focus evidence consists of two factual hash statements, not policy text; the counterfactual alters only the newest-report hash to a mismatching value, coherently testing the hash-supersession rule without contradicting other facts; neither context reveals a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08.\"},{\"speaker\":\"Compliance auditor\",\"text\":\"The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}, {"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3f79bb7b435b05321651daefd374cdc681dc06fa.", "negative_right": "The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08.", "right": "The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-017", "id": "fast-43-diverse-139-017-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}, {"speaker": "Compliance auditor", "text": "The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy and question bindings for batch C-184 at 13:40/14:20; the focus evidence consists of two factual hash statements, not policy text; the counterfactual alters only the newest-report hash to a mismatching value, coherently testing the hash-supersession rule without contradicting other facts; neither context reveals a gold answer, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08.\"},{\"speaker\":\"Compliance auditor\",\"text\":\"The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}, {"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3f79bb7b435b05321651daefd374cdc681dc06fa.", "negative_right": "The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08.", "right": "The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-017", "id": "fast-43-diverse-139-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3f79bb7b435b05321651daefd374cdc681dc06fa."}, {"speaker": "Compliance auditor", "text": "The immutable SHA-256 hash of batch C-184 on file is 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the full precedence policy in the administrator's turn, matching the original question rubric; the counterfactual only changes the reported hash to differ from the registry hash, a coherent single-fact edit that plausibly triggers the hash-mismatch clause without contradicting any other unchanged statement; the two focus evidence spans are complete factual sentences quoting the registry hash and the report hash, with no embedded rule tables, IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest submitted for batch C-184's readiness disposition: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%.\"},{\"speaker\":\"Batch registry\",\"text\":\"The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5.", "negative_left": "The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5.", "negative_right": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 7c3b2a1f9e8d6047bb4e0d2c1a6f0b57d3b4e8901f2a3b4c5d6e7f8091a2b3c4.", "right": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-021", "id": "fast-43-diverse-139-021-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest submitted for batch C-184's readiness disposition: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%."}, {"speaker": "Batch registry", "text": "The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the full precedence policy in the administrator's turn, matching the original question rubric; the counterfactual only changes the reported hash to differ from the registry hash, a coherent single-fact edit that plausibly triggers the hash-mismatch clause without contradicting any other unchanged statement; the two focus evidence spans are complete factual sentences quoting the registry hash and the report hash, with no embedded rule tables, IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest submitted for batch C-184's readiness disposition: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%.\"},{\"speaker\":\"Batch registry\",\"text\":\"The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5.", "negative_left": "The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5.", "negative_right": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 7c3b2a1f9e8d6047bb4e0d2c1a6f0b57d3b4e8901f2a3b4c5d6e7f8091a2b3c4.", "right": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-021", "id": "fast-43-diverse-139-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest submitted for batch C-184's readiness disposition: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%."}, {"speaker": "Batch registry", "text": "The immutable SHA-256 hash of batch C-184 is stored in the batch registry as 9f2a1c7e4b3d6089aa5f0e3c2b7d1a68e4c5f9012a3b4c5d6e7f8091a2b3c4d5."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 7c3b2a1f9e8d6047bb4e0d2c1a6f0b57d3b4e8901f2a3b4c5d6e7f8091a2b3c4."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical policy statement and all original disposition criteria via the unchanged questions object, keep the batch C-184 entity and hash-matching logic intact, and the counterfactual changes only the newest report's hash to a mismatching value in a single sentence, creating a coherent hold-triggering scenario without contradicting other measurements or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\": \"Data quality analyst\", \"text\": \"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"}, {\"speaker\": \"Data engineer\", \"text\": \"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The completed 14:20 report is the newest submission for batch C-184's readiness disposition: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7.\"}, {\"speaker\": \"Application administrator\", \"text\": \"The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7. Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7."}, {"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7.", "negative_left": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 1a2b3c4d5e6f7089ab12cd34ef56ab7890cd1234ef5678901abcdef23456789.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7.", "right": "The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-022", "id": "fast-43-diverse-139-022-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest submission for batch C-184's readiness disposition: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7."}, {"speaker": "Application administrator", "text": "The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7. Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical policy statement and all original disposition criteria via the unchanged questions object, keep the batch C-184 entity and hash-matching logic intact, and the counterfactual changes only the newest report's hash to a mismatching value in a single sentence, creating a coherent hold-triggering scenario without contradicting other measurements or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\": \"Data quality analyst\", \"text\": \"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"}, {\"speaker\": \"Data engineer\", \"text\": \"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The completed 14:20 report is the newest submission for batch C-184's readiness disposition: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7.\"}, {\"speaker\": \"Application administrator\", \"text\": \"The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7. Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7."}, {"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7.", "negative_left": "The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 1a2b3c4d5e6f7089ab12cd34ef56ab7890cd1234ef5678901abcdef23456789.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7.", "right": "The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-022", "id": "fast-43-diverse-139-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest submission for batch C-184's readiness disposition: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records a SHA-256 hash of 1a2b3c4d5e6f7089ab12cd34ef56ab7890cd1234ef5678901abcdef23456789."}, {"speaker": "Application administrator", "text": "The immutable SHA-256 hash of batch C-184 is 9f3a7c2e1b8d4560af23e19d0b7c6a45f1e2d3c4b5a69788f0e1d2c3b4a5f6e7. Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full administrator policy statement and original question object verbatim, preserving all governing rules and precedence; the counterfactual changes only the reported hash in the fifth statement to a mismatching value, which is a coherent single-sentence factual alteration that does not contradict any other unchanged fact (row counts, rates, timestamps remain identical); the two focus evidence spans are complete factual sentences reporting hashes, not policy text or instructions; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report covers batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. This 14:20 report is the newest report submitted for the readiness disposition of batch C-184.\"},{\"speaker\":\"Batch registry\",\"text\":\"The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455.", "negative_left": "The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455.", "negative_right": "The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 1234567890abcdef1234567890abcdef1234567890abcdef1234567890abcd.", "right": "The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-023", "id": "fast-43-diverse-139-023-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report covers batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. This 14:20 report is the newest report submitted for the readiness disposition of batch C-184."}, {"speaker": "Batch registry", "text": "The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full administrator policy statement and original question object verbatim, preserving all governing rules and precedence; the counterfactual changes only the reported hash in the fifth statement to a mismatching value, which is a coherent single-sentence factual alteration that does not contradict any other unchanged fact (row counts, rates, timestamps remain identical); the two focus evidence spans are complete factual sentences reporting hashes, not policy text or instructions; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report covers batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. This 14:20 report is the newest report submitted for the readiness disposition of batch C-184.\"},{\"speaker\":\"Batch registry\",\"text\":\"The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455.", "negative_left": "The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455.", "negative_right": "The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 1234567890abcdef1234567890abcdef1234567890abcdef1234567890abcd.", "right": "The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-023", "id": "fast-43-diverse-139-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report covers batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. This 14:20 report is the newest report submitted for the readiness disposition of batch C-184."}, {"speaker": "Batch registry", "text": "The immutable SHA-256 hash of batch C-184 is recorded in the batch registry as 9f3a1c7e2b8d4560ffea1122334455667788990aabbccddeeff001122334455."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records its SHA-256 hash as 1234567890abcdef1234567890abcdef1234567890abcdef1234567890abcd."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the hash-match supersession policy and thresholds from the original question; the counterfactual coherently changes only the reported hash to create a mismatch, triggering the hold criterion without contradicting other facts; evidence spans are two complete factual sentences with no embedded rules or labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\": \"Data quality analyst\", \"text\": \"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"}, {\"speaker\": \"Data engineer\", \"text\": \"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The completed 14:20 report is the newest report submitted for batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4.\"}, {\"speaker\": \"Application administrator\", \"text\": \"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4."}, {"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4.", "negative_left": "The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f0e1d2c3b4a5968072615c4d3e2f1a0b9c8d7e6f5a4b3c2d1e0f9a8b7c6d5e4.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-024", "id": "fast-43-diverse-139-024-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the hash-match supersession policy and thresholds from the original question; the counterfactual coherently changes only the reported hash to create a mismatch, triggering the hold criterion without contradicting other facts; evidence spans are two complete factual sentences with no embedded rules or labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\": \"Data quality analyst\", \"text\": \"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"}, {\"speaker\": \"Data engineer\", \"text\": \"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The completed 14:20 report is the newest report submitted for batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4.\"}, {\"speaker\": \"Application administrator\", \"text\": \"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4."}, {"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4.", "negative_left": "The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f0e1d2c3b4a5968072615c4d3e2f1a0b9c8d7e6f5a4b3c2d1e0f9a8b7c6d5e4.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-024", "id": "fast-43-diverse-139-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 is stored as 7a3f9c1e2b8d4560af23c9e0d1b6f7a45c8e9d2f1a0b3c4d5e6f7089a1b2c3d4. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f0e1d2c3b4a5968072615c4d3e2f1a0b9c8d7e6f5a4b3c2d1e0f9a8b7c6d5e4."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy statement and question bindings unchanged; the two evidence spans are complete factual sentences reporting hash values, not policy text; the counterfactual alters only the reported hash to create a genuine mismatch with the immutable registry hash, which is a coherent single-fact change rather than a duplicate contradictory measurement; neither context reveals a gold label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\": \"Data quality analyst\", \"text\": \"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"}, {\"speaker\": \"Data engineer\", \"text\": \"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015.\"}, {\"speaker\": \"Batch registry system\", \"text\": \"The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015.\"}, {\"speaker\": \"Application administrator\", \"text\": \"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015."}, {"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3c9909afec25354d551dae21590bb26e.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015.", "right": "The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-025", "id": "fast-43-diverse-139-025-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015."}, {"speaker": "Batch registry system", "text": "The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy statement and question bindings unchanged; the two evidence spans are complete factual sentences reporting hash values, not policy text; the counterfactual alters only the reported hash to create a genuine mismatch with the immutable registry hash, which is a coherent single-fact change rather than a duplicate contradictory measurement; neither context reveals a gold label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\": \"Data quality analyst\", \"text\": \"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"}, {\"speaker\": \"Data engineer\", \"text\": \"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015.\"}, {\"speaker\": \"Batch registry system\", \"text\": \"The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015.\"}, {\"speaker\": \"Application administrator\", \"text\": \"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015."}, {"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 9f86d081884c7d659a2feaa0c55ad015.", "negative_left": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3c9909afec25354d551dae21590bb26e.", "negative_right": "The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015.", "right": "The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-025", "id": "fast-43-diverse-139-025-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 3c9909afec25354d551dae21590bb26e."}, {"speaker": "Batch registry system", "text": "The immutable SHA-256 hash of batch C-184 is 9f86d081884c7d659a2feaa0c55ad015."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the supersession policy and thresholds from the original question/state, keep the same batch entity and time bindings, use two complete factual hash-statement sentences as evidence, and the counterfactual coherently alters only the reported hash to create a mismatch without contradicting other measurements or leaking the intended disposition.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%.\"},{\"speaker\":\"System log\",\"text\":\"The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62.\"},{\"speaker\":\"System log\",\"text\":\"The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 9f3a1cde7b2f0184decb5a11f8907c62.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62."}, {"path": ["4", "text"], "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 9f3a1cde7b2f0184decb5a11f8907c62."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62.", "negative_left": "The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62.", "negative_right": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 4b8e2fca10d739e5f6a0c3b8127de9aa.", "right": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 9f3a1cde7b2f0184decb5a11f8907c62."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-029", "id": "fast-43-diverse-139-029-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%."}, {"speaker": "System log", "text": "The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62."}, {"speaker": "System log", "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 9f3a1cde7b2f0184decb5a11f8907c62."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the supersession policy and thresholds from the original question/state, keep the same batch entity and time bindings, use two complete factual hash-statement sentences as evidence, and the counterfactual coherently alters only the reported hash to create a mismatch without contradicting other measurements or leaking the intended disposition.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%.\"},{\"speaker\":\"System log\",\"text\":\"The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62.\"},{\"speaker\":\"System log\",\"text\":\"The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 9f3a1cde7b2f0184decb5a11f8907c62.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62."}, {"path": ["4", "text"], "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 9f3a1cde7b2f0184decb5a11f8907c62."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62.", "negative_left": "The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62.", "negative_right": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 4b8e2fca10d739e5f6a0c3b8127de9aa.", "right": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 9f3a1cde7b2f0184decb5a11f8907c62."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-029", "id": "fast-43-diverse-139-029-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%."}, {"speaker": "System log", "text": "The immutable SHA-256 hash of batch C-184 is stored as 9f3a1cde7b2f0184decb5a11f8907c62."}, {"speaker": "System log", "text": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is 4b8e2fca10d739e5f6a0c3b8127de9aa."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the admin's full precedence/threshold policy verbatim and retain the C-184 entity binding; the counterfactual only swaps the newest report's hash to create a plausible mismatch scenario, and the two evidence spans are single complete factual sentences with no embedded rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:15, validation for batch C-184 reported 8,200 customer rows: consent_timestamp completeness 95.4%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7d2f8a1c4e6b9035.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7d2f8a1c4e6b9035."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035.", "negative_left": "The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as a15e3c9b7f2d0468.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7d2f8a1c4e6b9035."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-062", "id": "fast-43-diverse-139-062-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:15, validation for batch C-184 reported 8,200 customer rows: consent_timestamp completeness 95.4%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7d2f8a1c4e6b9035."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the admin's full precedence/threshold policy verbatim and retain the C-184 entity binding; the counterfactual only swaps the newest report's hash to create a plausible mismatch scenario, and the two evidence spans are single complete factual sentences with no embedded rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:15, validation for batch C-184 reported 8,200 customer rows: consent_timestamp completeness 95.4%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7d2f8a1c4e6b9035.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7d2f8a1c4e6b9035."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035.", "negative_left": "The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as a15e3c9b7f2d0468.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7d2f8a1c4e6b9035."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-062", "id": "fast-43-diverse-139-062-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:15, validation for batch C-184 reported 8,200 customer rows: consent_timestamp completeness 95.4%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7d2f8a1c4e6b9035."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as a15e3c9b7f2d0468."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy/criteria language and the C-184 entity, 13:40/14:20 timings, and hash-on-file fact; the counterfactual alters only the newest report's recorded hash to d02f5e18ab739c46, producing a coherent hash-mismatch scenario rather than a duplicate contradictory measurement; the two focus evidence spans are complete factual sentences with no embedded rules, IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7ac4e91b03fd88a2.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7ac4e91b03fd88a2."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2.", "negative_left": "The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as d02f5e18ab739c46.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7ac4e91b03fd88a2."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-063", "id": "fast-43-diverse-139-063-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7ac4e91b03fd88a2."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged policy/criteria language and the C-184 entity, 13:40/14:20 timings, and hash-on-file fact; the counterfactual alters only the newest report's recorded hash to d02f5e18ab739c46, producing a coherent hash-mismatch scenario rather than a duplicate contradictory measurement; the two focus evidence spans are complete factual sentences with no embedded rules, IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7ac4e91b03fd88a2.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7ac4e91b03fd88a2."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2.", "negative_left": "The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as d02f5e18ab739c46.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7ac4e91b03fd88a2."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-063", "id": "fast-43-diverse-139-063-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7ac4e91b03fd88a2."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as d02f5e18ab739c46."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only changes the reported hash to differ from the on-file hash, creating a coherent hash-mismatch scenario without contradicting any other fixed facts, and both contexts retain the same policy, question bindings, and two factual evidence sentences with no leaked answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7c1e9d4a2f6b8035.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7c1e9d4a2f6b8035."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035.", "negative_left": "The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as e5a2f8b1c4d6a790.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7c1e9d4a2f6b8035."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-065", "id": "fast-43-diverse-139-065-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7c1e9d4a2f6b8035."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only changes the reported hash to differ from the on-file hash, creating a coherent hash-mismatch scenario without contradicting any other fixed facts, and both contexts retain the same policy, question bindings, and two factual evidence sentences with no leaked answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a7 is a permissible universally quantified factual claim. The focus atom is factual hash identity, not a policy classification. The base and counter assignments are realizable with only that identity relation changing: a completed newest report can retain the same measurements and supported transformations while either matching or failing to match the immutable hash. Policy evidence is correctly drawn from the original state and preserves its relevant temporal-update and precedence statement; rules already contained in the retained questions object need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a completed newest report matching the immutable hash, supported transformations, completeness of exactly 98.0%, duplicates of 0.9%, and schema errors of 0.0%. These facts exclude each earlier-precedence disposition and satisfy all inclusive acceptance thresholds.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the hash in the newest 14:20 report does not match the immutable batch hash, while a2 confirms that this is the newest report. A newest-report hash mismatch is independently sufficient for the first-precedence pipeline-integrity hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SHA-256 hash recorded in the newest 14:20 report submitted for batch C-184 is identical to the immutable SHA-256 hash of batch C-184."}, {"id": "a2", "statement": "The report timestamped 14:20 is the newest report submitted for the readiness disposition of batch C-184."}, {"id": "a3", "statement": "The newest 14:20 report submitted for batch C-184 is completed."}, {"id": "a4", "statement": "The newest 14:20 report submitted for batch C-184 measures required-field completeness at exactly 98.0%."}, {"id": "a5", "statement": "The newest 14:20 report submitted for batch C-184 measures the duplicate rate at 0.9%."}, {"id": "a6", "statement": "The newest 14:20 report submitted for batch C-184 measures the schema-error rate at 0.0%."}, {"id": "a7", "statement": "Every transformation used to generate the newest 14:20 report for batch C-184 is supported."}], "base_state_json": "[{\"speaker\":\"Data quality analyst\",\"text\":\"At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%.\"},{\"speaker\":\"Data engineer\",\"text\":\"The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035.\"},{\"speaker\":\"Application administrator\",\"text\":\"Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7c1e9d4a2f6b8035.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["2", "text"], "text": "The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035."}, {"path": ["4", "text"], "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7c1e9d4a2f6b8035."}], "policy_evidence": [{"path": ["3", "text"], "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}], "rules": [{"justification": "The newest report is completed and matches the immutable batch hash, every transformation is supported, and all three measurements satisfy the inclusive acceptance thresholds; therefore no earlier disposition applies.", "target": "accept_batch", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The newest report does not match the immutable hash of batch C-184, which independently triggers the first-precedence pipeline-integrity hold.", "target": "pipeline_integrity_hold", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}]}]}, "verified_pair": {"left": "The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035.", "negative_left": "The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035.", "negative_right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as e5a2f8b1c4d6a790.", "right": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as 7c1e9d4a2f6b8035."}, "verifier_independent_model": false}, "family": "fast-43-diverse-139-065", "id": "fast-43-diverse-139-065-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept when no earlier disposition applies: completeness is at least 98.0%, duplicates are at most 1.0%, schema errors are at most 0.1%, and transformations are supported.", "conditional_rework": "Fifth precedence: return for remediation if earlier conditions do not apply and completeness is at least 95.0% but below 98.0%.", "duplicate_quarantine": "Second precedence: quarantine if no pipeline-integrity hold applies and the duplicate rate exceeds 1.0%.", "pipeline_integrity_hold": "First precedence: hold and route to the data engineer if any unsupported coercion exists, or if the newest report is incomplete or does not match the immutable batch hash.", "product_schema_hold": "Third precedence: hold for the product-data owner if neither earlier condition applies and schema errors exceed 0.1%.", "source_repair_required": "Fourth precedence: route to the source-system operator if earlier conditions do not apply and required-field completeness is below 95.0%."}, "instructions": "Choose the batch-readiness disposition. Apply the option rubrics in their stated precedence order; threshold equality passes. Only a completed temporal update for the identical batch hash may replace earlier evidence.", "type": "choice"}}, "state": [{"speaker": "Data quality analyst", "text": "At 13:40, validation for batch C-184 reported 8,000 customer rows: consent_timestamp completeness 95.1%, duplicate rate 0.9%, and schema-error rate 0.0%."}, {"speaker": "Data engineer", "text": "The pipeline had mistakenly validated before the approved UTC conversion finished. I reran that step without changing source values or row count; the transformation audit shows no coercions or fabricated timestamps."}, {"speaker": "Source-system operator", "text": "The completed 14:20 report is the newest report submitted for the readiness disposition of batch C-184: completeness is now exactly 98.0%, duplicates remain 0.9%, and schema errors remain 0.0%. The immutable SHA-256 hash of batch C-184 on file is 7c1e9d4a2f6b8035."}, {"speaker": "Application administrator", "text": "Policy says a later completed report supersedes an earlier report only for the same immutable hash. Apply dispositions in listed precedence. Acceptance requires completeness at least 98.0%, duplicates at most 1.0%, schema errors at most 0.1%, and no unsupported transformation."}, {"speaker": "Source-system operator", "text": "The newest 14:20 report submitted for batch C-184 records the SHA-256 hash as e5a2f8b1c4d6a790."}]}, "method": "c2d", "provenance": {"source_id": "diverse-139", "source_is_synthetic": true, "source_sha256": "ac71d49fa94cf2a919d9c5d90fc5f6d1f38a9f5b955b1e3e7b63ef05ba451537", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "pipeline_integrity_hold"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the acceptance/routing policy verbatim and the original question is unchanged, preserving all bindings; the two focus sentences are plain factual statements about run identifiers, not policy or instructions; the counterfactual coherently changes report D's run tag to V-204-run-3, distinguishing it from the current run V-204-run-7 without contradicting other stated facts; neither context reveals a gold answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "full_context_fact_states": {"base": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "counterfactual": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "remove_left": {"a_focus_report_run": "unknown"}, "remove_right": {"a_focus_report_run": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_report_run": "unknown"}, "negative_pair": {"a_focus_report_run": "refuted"}, "negative_sentence": {"a_focus_report_run": "unknown"}, "positive_pair": {"a_focus_report_run": "supported"}, "right": {"a_focus_report_run": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified third atom is still atomic under the stated standard. The focus is factual report-to-run provenance, not a policy conclusion. The base and counter assignments are jointly realizable as separate scenarios with only D's run provenance changing: D can report the same collision count while being either a current-run or non-current-run report, and the remaining current-run reports can all report zero. The policy evidence consists only of substantive excerpts from the original state and preserves the normalization requirement and state-originating policy constraints; governing material in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "If D is tied to the current V-204 run and records 18 post-normalization collision groups, current-run collisions are confirmed. That satisfies the pipeline-owner route criterion; the zero-collision findings from other reports do not support acceptance and do not negate D's confirmed current-run collisions.", "rule_index": 0, "sound": true}, {"reason": "Refuting that D is from the current V-204 run excludes its 18 collisions from the current-run report set. The universal atom then establishes that every applicable current-run validation or duplicate report records zero normalized-ID collisions, satisfying the acceptance agreement condition.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_report_run", "statement": "Duplicate report D's run is the current V-204 run."}, {"id": "a_duplicate_collision_count", "statement": "Duplicate report D records 18 vendor_id collision groups after trim-and-uppercase normalization."}, {"id": "a_other_reports_zero", "statement": "Every applicable validation or duplicate report for the current V-204 run other than duplicate report D records zero vendor_id collisions after trim-and-uppercase normalization."}], "base_state_json": "\"Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The validator summary for run V-204 reports zero duplicate vendor_ids after normalization and marks the batch ready. A separate duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values 'ab-17 ' and 'AB-17'. Duplicate report D was generated under run identifier V-204-run-7. Run identifier V-204-run-7 is the identifier assigned to the current V-204 run. Every other applicable validation or duplicate report for the current V-204 run records zero vendor_id collisions after trim-and-uppercase normalization. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run.\"", "base_states": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "focus_atom": "a_focus_report_run", "focus_evidence": [{"path": [], "text": "Duplicate report D was generated under run identifier V-204-run-7."}, {"path": [], "text": "Run identifier V-204-run-7 is the identifier assigned to the current V-204 run."}], "policy_evidence": [{"path": [], "text": "The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization."}, {"path": [], "text": "Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}], "rules": [{"justification": "D is a report from the current V-204 run and records normalized vendor_id collisions, so current-run collisions are confirmed. Because the other applicable reports record zero collisions, the reports do not unanimously support acceptance.", "target": "route_pipeline_owner", "when": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}, {"justification": "D is explicitly from a run other than the current V-204 run, so its collision count is not a current-run report for V-204. Every applicable current-run validation or duplicate report records zero normalized vendor_id collisions, satisfying the agreement required for acceptance.", "target": "accept_batch", "when": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}]}, "verified_pair": {"left": "Duplicate report D was generated under run identifier V-204-run-7.", "negative_left": "Duplicate report D was generated under run identifier V-204-run-3.", "negative_right": "Run identifier V-204-run-7 is the identifier assigned to the current V-204 run.", "right": "Run identifier V-204-run-7 is the identifier assigned to the current V-204 run."}, "verifier_independent_model": false}, "family": "fast-43-diverse-140-008", "id": "fast-43-diverse-140-008-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept V-204 only if all applicable validation and duplicate reports agree that the normalized vendor_id field has no collisions.", "none_of_above": "Use when conflicting or untraceable evidence prevents both acceptance and a supported pipeline-owner route.", "route_pipeline_owner": "Hold V-204 and route it to the pipeline owner only if normalized-ID collisions are confirmed in the current V-204 run."}, "instructions": "Choose the readiness decision whose rubric is fully supported by the evidence. If neither substantive rubric is satisfied, choose none_of_above.", "type": "choice"}}, "state": "Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The validator summary for run V-204 reports zero duplicate vendor_ids after normalization and marks the batch ready. A separate duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values 'ab-17 ' and 'AB-17'. Duplicate report D was generated under run identifier V-204-run-7. Run identifier V-204-run-7 is the identifier assigned to the current V-204 run. Every other applicable validation or duplicate report for the current V-204 run records zero vendor_id collisions after trim-and-uppercase normalization. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}, "method": "c2d", "provenance": {"source_id": "diverse-140", "source_is_synthetic": true, "source_sha256": "1780d01061297783e3174bb4e2b929c3cbccde08d131ae1a6693198134c6a867", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_pipeline_owner"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the acceptance/routing policy verbatim and the original question is unchanged, preserving all bindings; the two focus sentences are plain factual statements about run identifiers, not policy or instructions; the counterfactual coherently changes report D's run tag to V-204-run-3, distinguishing it from the current run V-204-run-7 without contradicting other stated facts; neither context reveals a gold answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "full_context_fact_states": {"base": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "counterfactual": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "remove_left": {"a_focus_report_run": "unknown"}, "remove_right": {"a_focus_report_run": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_report_run": "unknown"}, "negative_pair": {"a_focus_report_run": "refuted"}, "negative_sentence": {"a_focus_report_run": "unknown"}, "positive_pair": {"a_focus_report_run": "supported"}, "right": {"a_focus_report_run": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified third atom is still atomic under the stated standard. The focus is factual report-to-run provenance, not a policy conclusion. The base and counter assignments are jointly realizable as separate scenarios with only D's run provenance changing: D can report the same collision count while being either a current-run or non-current-run report, and the remaining current-run reports can all report zero. The policy evidence consists only of substantive excerpts from the original state and preserves the normalization requirement and state-originating policy constraints; governing material in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "If D is tied to the current V-204 run and records 18 post-normalization collision groups, current-run collisions are confirmed. That satisfies the pipeline-owner route criterion; the zero-collision findings from other reports do not support acceptance and do not negate D's confirmed current-run collisions.", "rule_index": 0, "sound": true}, {"reason": "Refuting that D is from the current V-204 run excludes its 18 collisions from the current-run report set. The universal atom then establishes that every applicable current-run validation or duplicate report records zero normalized-ID collisions, satisfying the acceptance agreement condition.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_report_run", "statement": "Duplicate report D's run is the current V-204 run."}, {"id": "a_duplicate_collision_count", "statement": "Duplicate report D records 18 vendor_id collision groups after trim-and-uppercase normalization."}, {"id": "a_other_reports_zero", "statement": "Every applicable validation or duplicate report for the current V-204 run other than duplicate report D records zero vendor_id collisions after trim-and-uppercase normalization."}], "base_state_json": "\"Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The validator summary for run V-204 reports zero duplicate vendor_ids after normalization and marks the batch ready. A separate duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values 'ab-17 ' and 'AB-17'. Duplicate report D was generated under run identifier V-204-run-7. Run identifier V-204-run-7 is the identifier assigned to the current V-204 run. Every other applicable validation or duplicate report for the current V-204 run records zero vendor_id collisions after trim-and-uppercase normalization. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run.\"", "base_states": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "focus_atom": "a_focus_report_run", "focus_evidence": [{"path": [], "text": "Duplicate report D was generated under run identifier V-204-run-7."}, {"path": [], "text": "Run identifier V-204-run-7 is the identifier assigned to the current V-204 run."}], "policy_evidence": [{"path": [], "text": "The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization."}, {"path": [], "text": "Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}], "rules": [{"justification": "D is a report from the current V-204 run and records normalized vendor_id collisions, so current-run collisions are confirmed. Because the other applicable reports record zero collisions, the reports do not unanimously support acceptance.", "target": "route_pipeline_owner", "when": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}, {"justification": "D is explicitly from a run other than the current V-204 run, so its collision count is not a current-run report for V-204. Every applicable current-run validation or duplicate report records zero normalized vendor_id collisions, satisfying the agreement required for acceptance.", "target": "accept_batch", "when": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}]}, "verified_pair": {"left": "Duplicate report D was generated under run identifier V-204-run-7.", "negative_left": "Duplicate report D was generated under run identifier V-204-run-3.", "negative_right": "Run identifier V-204-run-7 is the identifier assigned to the current V-204 run.", "right": "Run identifier V-204-run-7 is the identifier assigned to the current V-204 run."}, "verifier_independent_model": false}, "family": "fast-43-diverse-140-008", "id": "fast-43-diverse-140-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept V-204 only if all applicable validation and duplicate reports agree that the normalized vendor_id field has no collisions.", "none_of_above": "Use when conflicting or untraceable evidence prevents both acceptance and a supported pipeline-owner route.", "route_pipeline_owner": "Hold V-204 and route it to the pipeline owner only if normalized-ID collisions are confirmed in the current V-204 run."}, "instructions": "Choose the readiness decision whose rubric is fully supported by the evidence. If neither substantive rubric is satisfied, choose none_of_above.", "type": "choice"}}, "state": "Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The validator summary for run V-204 reports zero duplicate vendor_ids after normalization and marks the batch ready. A separate duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values 'ab-17 ' and 'AB-17'. Duplicate report D was generated under run identifier V-204-run-3. Run identifier V-204-run-7 is the identifier assigned to the current V-204 run. Every other applicable validation or duplicate report for the current V-204 run records zero vendor_id collisions after trim-and-uppercase normalization. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}, "method": "c2d", "provenance": {"source_id": "diverse-140", "source_is_synthetic": true, "source_sha256": "1780d01061297783e3174bb4e2b929c3cbccde08d131ae1a6693198134c6a867", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance/route criteria via the unchanged questions and only alter factual run-identifier details (a legitimate case observation change), the two focus sentences are plain factual statements, the counterfactual's shift of the current run to run-88 is internally coherent and doesn't duplicate or contradict other measurements, and neither context states or implies the gold decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "full_context_fact_states": {"base": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "counterfactual": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "remove_left": {"a_focus_report_run": "unknown"}, "remove_right": {"a_focus_report_run": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_report_run": "unknown"}, "negative_pair": {"a_focus_report_run": "refuted"}, "negative_sentence": {"a_focus_report_run": "unknown"}, "positive_pair": {"a_focus_report_run": "supported"}, "right": {"a_focus_report_run": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified third atom is still atomic under the stated standard. The focus is factual report-to-run provenance, not a policy conclusion. The base and counter assignments are jointly realizable as separate scenarios with only D's run provenance changing: D can report the same collision count while being either a current-run or non-current-run report, and the remaining current-run reports can all report zero. The policy evidence consists only of substantive excerpts from the original state and preserves the normalization requirement and state-originating policy constraints; governing material in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "If D is tied to the current V-204 run and records 18 post-normalization collision groups, current-run collisions are confirmed. That satisfies the pipeline-owner route criterion; the zero-collision findings from other reports do not support acceptance and do not negate D's confirmed current-run collisions.", "rule_index": 0, "sound": true}, {"reason": "Refuting that D is from the current V-204 run excludes its 18 collisions from the current-run report set. The universal atom then establishes that every applicable current-run validation or duplicate report records zero normalized-ID collisions, satisfying the acceptance agreement condition.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_report_run", "statement": "Duplicate report D's run is the current V-204 run."}, {"id": "a_duplicate_collision_count", "statement": "Duplicate report D records 18 vendor_id collision groups after trim-and-uppercase normalization."}, {"id": "a_other_reports_zero", "statement": "Every applicable validation or duplicate report for the current V-204 run other than duplicate report D records zero vendor_id collisions after trim-and-uppercase normalization."}], "base_state_json": "\"Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The validator summary for run V-204 reports zero duplicate vendor_ids after trim-and-uppercase normalization and marks the batch ready. A separate duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values \\u2018ab-17 \\u2019 and \\u2018AB-17\\u2019. Duplicate report D's run is logged with run identifier run-77. The current V-204 batch run is identified as run-77. Every other applicable validation or duplicate report for run V-204 records zero vendor_id collisions after trim-and-uppercase normalization. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run.\"", "base_states": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "focus_atom": "a_focus_report_run", "focus_evidence": [{"path": [], "text": "Duplicate report D's run is logged with run identifier run-77."}, {"path": [], "text": "The current V-204 batch run is identified as run-77."}], "policy_evidence": [{"path": [], "text": "The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization."}, {"path": [], "text": "Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}], "rules": [{"justification": "D is a report from the current V-204 run and records normalized vendor_id collisions, so current-run collisions are confirmed. Because the other applicable reports record zero collisions, the reports do not unanimously support acceptance.", "target": "route_pipeline_owner", "when": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}, {"justification": "D is explicitly from a run other than the current V-204 run, so its collision count is not a current-run report for V-204. Every applicable current-run validation or duplicate report records zero normalized vendor_id collisions, satisfying the agreement required for acceptance.", "target": "accept_batch", "when": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}]}, "verified_pair": {"left": "Duplicate report D's run is logged with run identifier run-77.", "negative_left": "Duplicate report D's run is logged with run identifier run-77.", "negative_right": "The current V-204 batch run is identified as run-88.", "right": "The current V-204 batch run is identified as run-77."}, "verifier_independent_model": false}, "family": "fast-43-diverse-140-013", "id": "fast-43-diverse-140-013-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept V-204 only if all applicable validation and duplicate reports agree that the normalized vendor_id field has no collisions.", "none_of_above": "Use when conflicting or untraceable evidence prevents both acceptance and a supported pipeline-owner route.", "route_pipeline_owner": "Hold V-204 and route it to the pipeline owner only if normalized-ID collisions are confirmed in the current V-204 run."}, "instructions": "Choose the readiness decision whose rubric is fully supported by the evidence. If neither substantive rubric is satisfied, choose none_of_above.", "type": "choice"}}, "state": "Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The validator summary for run V-204 reports zero duplicate vendor_ids after trim-and-uppercase normalization and marks the batch ready. A separate duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values ‘ab-17 ’ and ‘AB-17’. Duplicate report D's run is logged with run identifier run-77. The current V-204 batch run is identified as run-77. Every other applicable validation or duplicate report for run V-204 records zero vendor_id collisions after trim-and-uppercase normalization. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}, "method": "c2d", "provenance": {"source_id": "diverse-140", "source_is_synthetic": true, "source_sha256": "1780d01061297783e3174bb4e2b929c3cbccde08d131ae1a6693198134c6a867", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_pipeline_owner"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance/route criteria via the unchanged questions and only alter factual run-identifier details (a legitimate case observation change), the two focus sentences are plain factual statements, the counterfactual's shift of the current run to run-88 is internally coherent and doesn't duplicate or contradict other measurements, and neither context states or implies the gold decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "full_context_fact_states": {"base": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "counterfactual": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "remove_left": {"a_focus_report_run": "unknown"}, "remove_right": {"a_focus_report_run": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_report_run": "unknown"}, "negative_pair": {"a_focus_report_run": "refuted"}, "negative_sentence": {"a_focus_report_run": "unknown"}, "positive_pair": {"a_focus_report_run": "supported"}, "right": {"a_focus_report_run": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified third atom is still atomic under the stated standard. The focus is factual report-to-run provenance, not a policy conclusion. The base and counter assignments are jointly realizable as separate scenarios with only D's run provenance changing: D can report the same collision count while being either a current-run or non-current-run report, and the remaining current-run reports can all report zero. The policy evidence consists only of substantive excerpts from the original state and preserves the normalization requirement and state-originating policy constraints; governing material in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "If D is tied to the current V-204 run and records 18 post-normalization collision groups, current-run collisions are confirmed. That satisfies the pipeline-owner route criterion; the zero-collision findings from other reports do not support acceptance and do not negate D's confirmed current-run collisions.", "rule_index": 0, "sound": true}, {"reason": "Refuting that D is from the current V-204 run excludes its 18 collisions from the current-run report set. The universal atom then establishes that every applicable current-run validation or duplicate report records zero normalized-ID collisions, satisfying the acceptance agreement condition.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_report_run", "statement": "Duplicate report D's run is the current V-204 run."}, {"id": "a_duplicate_collision_count", "statement": "Duplicate report D records 18 vendor_id collision groups after trim-and-uppercase normalization."}, {"id": "a_other_reports_zero", "statement": "Every applicable validation or duplicate report for the current V-204 run other than duplicate report D records zero vendor_id collisions after trim-and-uppercase normalization."}], "base_state_json": "\"Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The validator summary for run V-204 reports zero duplicate vendor_ids after trim-and-uppercase normalization and marks the batch ready. A separate duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values \\u2018ab-17 \\u2019 and \\u2018AB-17\\u2019. Duplicate report D's run is logged with run identifier run-77. The current V-204 batch run is identified as run-77. Every other applicable validation or duplicate report for run V-204 records zero vendor_id collisions after trim-and-uppercase normalization. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run.\"", "base_states": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "focus_atom": "a_focus_report_run", "focus_evidence": [{"path": [], "text": "Duplicate report D's run is logged with run identifier run-77."}, {"path": [], "text": "The current V-204 batch run is identified as run-77."}], "policy_evidence": [{"path": [], "text": "The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization."}, {"path": [], "text": "Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}], "rules": [{"justification": "D is a report from the current V-204 run and records normalized vendor_id collisions, so current-run collisions are confirmed. Because the other applicable reports record zero collisions, the reports do not unanimously support acceptance.", "target": "route_pipeline_owner", "when": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}, {"justification": "D is explicitly from a run other than the current V-204 run, so its collision count is not a current-run report for V-204. Every applicable current-run validation or duplicate report records zero normalized vendor_id collisions, satisfying the agreement required for acceptance.", "target": "accept_batch", "when": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}]}, "verified_pair": {"left": "Duplicate report D's run is logged with run identifier run-77.", "negative_left": "Duplicate report D's run is logged with run identifier run-77.", "negative_right": "The current V-204 batch run is identified as run-88.", "right": "The current V-204 batch run is identified as run-77."}, "verifier_independent_model": false}, "family": "fast-43-diverse-140-013", "id": "fast-43-diverse-140-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept V-204 only if all applicable validation and duplicate reports agree that the normalized vendor_id field has no collisions.", "none_of_above": "Use when conflicting or untraceable evidence prevents both acceptance and a supported pipeline-owner route.", "route_pipeline_owner": "Hold V-204 and route it to the pipeline owner only if normalized-ID collisions are confirmed in the current V-204 run."}, "instructions": "Choose the readiness decision whose rubric is fully supported by the evidence. If neither substantive rubric is satisfied, choose none_of_above.", "type": "choice"}}, "state": "Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The validator summary for run V-204 reports zero duplicate vendor_ids after trim-and-uppercase normalization and marks the batch ready. A separate duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values ‘ab-17 ’ and ‘AB-17’. Duplicate report D's run is logged with run identifier run-77. The current V-204 batch run is identified as run-88. Every other applicable validation or duplicate report for run V-204 records zero vendor_id collisions after trim-and-uppercase normalization. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}, "method": "c2d", "provenance": {"source_id": "diverse-140", "source_is_synthetic": true, "source_sha256": "1780d01061297783e3174bb4e2b929c3cbccde08d131ae1a6693198134c6a867", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance/routing policy without adding new exceptions, keep the same vendor batch and question entities, and the two evidence sentences are plain factual statements about run identifiers, not policy text; the counterfactual coherently alters only report D's run identifier to V-204-0003, creating a traceability mismatch without contradicting other stated facts, and neither context reveals a decision, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "full_context_fact_states": {"base": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "counterfactual": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "remove_left": {"a_focus_report_run": "unknown"}, "remove_right": {"a_focus_report_run": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_report_run": "unknown"}, "negative_pair": {"a_focus_report_run": "refuted"}, "negative_sentence": {"a_focus_report_run": "unknown"}, "positive_pair": {"a_focus_report_run": "supported"}, "right": {"a_focus_report_run": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified third atom is still atomic under the stated standard. The focus is factual report-to-run provenance, not a policy conclusion. The base and counter assignments are jointly realizable as separate scenarios with only D's run provenance changing: D can report the same collision count while being either a current-run or non-current-run report, and the remaining current-run reports can all report zero. The policy evidence consists only of substantive excerpts from the original state and preserves the normalization requirement and state-originating policy constraints; governing material in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "If D is tied to the current V-204 run and records 18 post-normalization collision groups, current-run collisions are confirmed. That satisfies the pipeline-owner route criterion; the zero-collision findings from other reports do not support acceptance and do not negate D's confirmed current-run collisions.", "rule_index": 0, "sound": true}, {"reason": "Refuting that D is from the current V-204 run excludes its 18 collisions from the current-run report set. The universal atom then establishes that every applicable current-run validation or duplicate report records zero normalized-ID collisions, satisfying the acceptance agreement condition.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_report_run", "statement": "Duplicate report D's run is the current V-204 run."}, {"id": "a_duplicate_collision_count", "statement": "Duplicate report D records 18 vendor_id collision groups after trim-and-uppercase normalization."}, {"id": "a_other_reports_zero", "statement": "Every applicable validation or duplicate report for the current V-204 run other than duplicate report D records zero vendor_id collisions after trim-and-uppercase normalization."}], "base_state_json": "\"Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. The validator summary for run V-204 reports zero duplicate vendor_ids and marks the batch ready. A duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values 'ab-17 ' and 'AB-17'. Duplicate report D was generated with run identifier V-204-0007. The current run identifier for V-204 processing is V-204-0007. All other applicable validation and duplicate reports for this run record zero vendor_id collisions after trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run.\"", "base_states": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "focus_atom": "a_focus_report_run", "focus_evidence": [{"path": [], "text": "Duplicate report D was generated with run identifier V-204-0007."}, {"path": [], "text": "The current run identifier for V-204 processing is V-204-0007."}], "policy_evidence": [{"path": [], "text": "The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization."}, {"path": [], "text": "Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}], "rules": [{"justification": "D is a report from the current V-204 run and records normalized vendor_id collisions, so current-run collisions are confirmed. Because the other applicable reports record zero collisions, the reports do not unanimously support acceptance.", "target": "route_pipeline_owner", "when": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}, {"justification": "D is explicitly from a run other than the current V-204 run, so its collision count is not a current-run report for V-204. Every applicable current-run validation or duplicate report records zero normalized vendor_id collisions, satisfying the agreement required for acceptance.", "target": "accept_batch", "when": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}]}, "verified_pair": {"left": "Duplicate report D was generated with run identifier V-204-0007.", "negative_left": "Duplicate report D was generated with run identifier V-204-0003.", "negative_right": "The current run identifier for V-204 processing is V-204-0007.", "right": "The current run identifier for V-204 processing is V-204-0007."}, "verifier_independent_model": false}, "family": "fast-43-diverse-140-024", "id": "fast-43-diverse-140-024-base", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept V-204 only if all applicable validation and duplicate reports agree that the normalized vendor_id field has no collisions.", "none_of_above": "Use when conflicting or untraceable evidence prevents both acceptance and a supported pipeline-owner route.", "route_pipeline_owner": "Hold V-204 and route it to the pipeline owner only if normalized-ID collisions are confirmed in the current V-204 run."}, "instructions": "Choose the readiness decision whose rubric is fully supported by the evidence. If neither substantive rubric is satisfied, choose none_of_above.", "type": "choice"}}, "state": "Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. The validator summary for run V-204 reports zero duplicate vendor_ids and marks the batch ready. A duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values 'ab-17 ' and 'AB-17'. Duplicate report D was generated with run identifier V-204-0007. The current run identifier for V-204 processing is V-204-0007. All other applicable validation and duplicate reports for this run record zero vendor_id collisions after trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}, "method": "c2d", "provenance": {"source_id": "diverse-140", "source_is_synthetic": true, "source_sha256": "1780d01061297783e3174bb4e2b929c3cbccde08d131ae1a6693198134c6a867", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_pipeline_owner"}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original acceptance/routing policy without adding new exceptions, keep the same vendor batch and question entities, and the two evidence sentences are plain factual statements about run identifiers, not policy text; the counterfactual coherently alters only report D's run identifier to V-204-0003, creating a traceability mismatch without contradicting other stated facts, and neither context reveals a decision, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "full_context_fact_states": {"base": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "supported", "a_other_reports_zero": "supported"}, "counterfactual": {"a_duplicate_collision_count": "supported", "a_focus_report_run": "refuted", "a_other_reports_zero": "supported"}, "remove_left": {"a_focus_report_run": "unknown"}, "remove_right": {"a_focus_report_run": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_report_run": "unknown"}, "negative_pair": {"a_focus_report_run": "refuted"}, "negative_sentence": {"a_focus_report_run": "unknown"}, "positive_pair": {"a_focus_report_run": "supported"}, "right": {"a_focus_report_run": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified third atom is still atomic under the stated standard. The focus is factual report-to-run provenance, not a policy conclusion. The base and counter assignments are jointly realizable as separate scenarios with only D's run provenance changing: D can report the same collision count while being either a current-run or non-current-run report, and the remaining current-run reports can all report zero. The policy evidence consists only of substantive excerpts from the original state and preserves the normalization requirement and state-originating policy constraints; governing material in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "If D is tied to the current V-204 run and records 18 post-normalization collision groups, current-run collisions are confirmed. That satisfies the pipeline-owner route criterion; the zero-collision findings from other reports do not support acceptance and do not negate D's confirmed current-run collisions.", "rule_index": 0, "sound": true}, {"reason": "Refuting that D is from the current V-204 run excludes its 18 collisions from the current-run report set. The universal atom then establishes that every applicable current-run validation or duplicate report records zero normalized-ID collisions, satisfying the acceptance agreement condition.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_report_run", "statement": "Duplicate report D's run is the current V-204 run."}, {"id": "a_duplicate_collision_count", "statement": "Duplicate report D records 18 vendor_id collision groups after trim-and-uppercase normalization."}, {"id": "a_other_reports_zero", "statement": "Every applicable validation or duplicate report for the current V-204 run other than duplicate report D records zero vendor_id collisions after trim-and-uppercase normalization."}], "base_state_json": "\"Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. The validator summary for run V-204 reports zero duplicate vendor_ids and marks the batch ready. A duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values 'ab-17 ' and 'AB-17'. Duplicate report D was generated with run identifier V-204-0007. The current run identifier for V-204 processing is V-204-0007. All other applicable validation and duplicate reports for this run record zero vendor_id collisions after trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run.\"", "base_states": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}], "focus_atom": "a_focus_report_run", "focus_evidence": [{"path": [], "text": "Duplicate report D was generated with run identifier V-204-0007."}, {"path": [], "text": "The current run identifier for V-204 processing is V-204-0007."}], "policy_evidence": [{"path": [], "text": "The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization."}, {"path": [], "text": "Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}], "rules": [{"justification": "D is a report from the current V-204 run and records normalized vendor_id collisions, so current-run collisions are confirmed. Because the other applicable reports record zero collisions, the reports do not unanimously support acceptance.", "target": "route_pipeline_owner", "when": [{"atom_id": "a_focus_report_run", "state": "supported"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}, {"justification": "D is explicitly from a run other than the current V-204 run, so its collision count is not a current-run report for V-204. Every applicable current-run validation or duplicate report records zero normalized vendor_id collisions, satisfying the agreement required for acceptance.", "target": "accept_batch", "when": [{"atom_id": "a_focus_report_run", "state": "refuted"}, {"atom_id": "a_duplicate_collision_count", "state": "supported"}, {"atom_id": "a_other_reports_zero", "state": "supported"}]}]}, "verified_pair": {"left": "Duplicate report D was generated with run identifier V-204-0007.", "negative_left": "Duplicate report D was generated with run identifier V-204-0003.", "negative_right": "The current run identifier for V-204 processing is V-204-0007.", "right": "The current run identifier for V-204 processing is V-204-0007."}, "verifier_independent_model": false}, "family": "fast-43-diverse-140-024", "id": "fast-43-diverse-140-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_batch": "Accept V-204 only if all applicable validation and duplicate reports agree that the normalized vendor_id field has no collisions.", "none_of_above": "Use when conflicting or untraceable evidence prevents both acceptance and a supported pipeline-owner route.", "route_pipeline_owner": "Hold V-204 and route it to the pipeline owner only if normalized-ID collisions are confirmed in the current V-204 run."}, "instructions": "Choose the readiness decision whose rubric is fully supported by the evidence. If neither substantive rubric is satisfied, choose none_of_above.", "type": "choice"}}, "state": "Data quality analyst Mira reviews batch V-204 of 50,000 vendor records. The schema requires vendor_id to be unique after the pipeline applies trim-and-uppercase normalization. The validator summary for run V-204 reports zero duplicate vendor_ids and marks the batch ready. A duplicate report, labeled report D, lists 18 normalized-ID collisions, including raw values 'ab-17 ' and 'AB-17'. Duplicate report D was generated with run identifier V-204-0003. The current run identifier for V-204 processing is V-204-0007. All other applicable validation and duplicate reports for this run record zero vendor_id collisions after trim-and-uppercase normalization. Policy permits acceptance only when all applicable reports agree, and permits routing to the pipeline owner only when a collision is confirmed in the current run."}, "method": "c2d", "provenance": {"source_id": "diverse-140", "source_is_synthetic": true, "source_sha256": "1780d01061297783e3174bb4e2b929c3cbccde08d131ae1a6693198134c6a867", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_batch"}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate and completeness acceptance rules and the history exemption without adding or removing policy; the question's batch ID, criteria, and instructions are unchanged; the two evidence spans are plain factual statements from the operator; the counterfactual only changes the populated ship_country count (39,980→39,950), which stays internally consistent with the unchanged duplicate and history facts; neither context states or hints at the final accept/reject decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\": \"Data engineer\", \"text\": \"Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward.\"}, {\"speaker\": \"Data quality analyst\", \"text\": \"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"Batch B-142 contains 40,000 records not tagged import_mode=history. Of those 40,000 records, 39,980 have a populated ship_country field.\"}, {\"speaker\": \"Application administrator\", \"text\": \"The transformation log confirms the post-transformation duplicate customer_id count is zero. No other schema errors remain. Should B-142 be accepted?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["2", "text"], "text": "Batch B-142 contains 40,000 records not tagged import_mode=history."}, {"path": ["2", "text"], "text": "Of those 40,000 records, 39,980 have a populated ship_country field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 40,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 40,000 records not tagged import_mode=history.", "negative_right": "Of those 40,000 records, 39,950 have a populated ship_country field.", "right": "Of those 40,000 records, 39,980 have a populated ship_country field."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-009", "id": "fast-43-diverse-142-009-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "Batch B-142 contains 40,000 records not tagged import_mode=history. Of those 40,000 records, 39,980 have a populated ship_country field."}, {"speaker": "Application administrator", "text": "The transformation log confirms the post-transformation duplicate customer_id count is zero. No other schema errors remain. Should B-142 be accepted?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate and completeness acceptance rules and the history exemption without adding or removing policy; the question's batch ID, criteria, and instructions are unchanged; the two evidence spans are plain factual statements from the operator; the counterfactual only changes the populated ship_country count (39,980→39,950), which stays internally consistent with the unchanged duplicate and history facts; neither context states or hints at the final accept/reject decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\": \"Data engineer\", \"text\": \"Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward.\"}, {\"speaker\": \"Data quality analyst\", \"text\": \"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"Batch B-142 contains 40,000 records not tagged import_mode=history. Of those 40,000 records, 39,980 have a populated ship_country field.\"}, {\"speaker\": \"Application administrator\", \"text\": \"The transformation log confirms the post-transformation duplicate customer_id count is zero. No other schema errors remain. Should B-142 be accepted?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["2", "text"], "text": "Batch B-142 contains 40,000 records not tagged import_mode=history."}, {"path": ["2", "text"], "text": "Of those 40,000 records, 39,980 have a populated ship_country field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 40,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 40,000 records not tagged import_mode=history.", "negative_right": "Of those 40,000 records, 39,950 have a populated ship_country field.", "right": "Of those 40,000 records, 39,980 have a populated ship_country field."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-009", "id": "fast-43-diverse-142-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "Batch B-142 contains 40,000 records not tagged import_mode=history. Of those 40,000 records, 39,950 have a populated ship_country field."}, {"speaker": "Application administrator", "text": "The transformation log confirms the post-transformation duplicate customer_id count is zero. No other schema errors remain. Should B-142 be accepted?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate and 99.95% completeness thresholds and the history-tag exemption, matching the unchanged question's rubric; entity B-142 and scope bindings are preserved; the two focus spans are complete factual sentences with exact quotes; the counterfactual only changes the missing-ship_country count (8\\u219215) without contradicting other stated facts; neither context states or implies the accept/reject outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142's transformation log confirms the duplicate report shows 12 before transformation and zero afterward, with no duplicate customer_id values remaining post-transformation.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"Batch B-142 contains 20,000 records not tagged import_mode=history.\"},{\"speaker\":\"Application administrator\",\"text\":\"In batch B-142, 8 of the records not tagged import_mode=history have a missing ship_country value.\"},{\"speaker\":\"Application administrator\",\"text\":\"No other schema errors remain outside these figures.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"path": ["4", "text"], "text": "In batch B-142, 8 of the records not tagged import_mode=history have a missing ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_right": "In batch B-142, 15 of the records not tagged import_mode=history have a missing ship_country value.", "right": "In batch B-142, 8 of the records not tagged import_mode=history have a missing ship_country value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-011", "id": "fast-43-diverse-142-011-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142's transformation log confirms the duplicate report shows 12 before transformation and zero afterward, with no duplicate customer_id values remaining post-transformation."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"speaker": "Application administrator", "text": "In batch B-142, 8 of the records not tagged import_mode=history have a missing ship_country value."}, {"speaker": "Application administrator", "text": "No other schema errors remain outside these figures."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate and 99.95% completeness thresholds and the history-tag exemption, matching the unchanged question's rubric; entity B-142 and scope bindings are preserved; the two focus spans are complete factual sentences with exact quotes; the counterfactual only changes the missing-ship_country count (8\\u219215) without contradicting other stated facts; neither context states or implies the accept/reject outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142's transformation log confirms the duplicate report shows 12 before transformation and zero afterward, with no duplicate customer_id values remaining post-transformation.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"Batch B-142 contains 20,000 records not tagged import_mode=history.\"},{\"speaker\":\"Application administrator\",\"text\":\"In batch B-142, 8 of the records not tagged import_mode=history have a missing ship_country value.\"},{\"speaker\":\"Application administrator\",\"text\":\"No other schema errors remain outside these figures.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"path": ["4", "text"], "text": "In batch B-142, 8 of the records not tagged import_mode=history have a missing ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_right": "In batch B-142, 15 of the records not tagged import_mode=history have a missing ship_country value.", "right": "In batch B-142, 8 of the records not tagged import_mode=history have a missing ship_country value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-011", "id": "fast-43-diverse-142-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142's transformation log confirms the duplicate report shows 12 before transformation and zero afterward, with no duplicate customer_id values remaining post-transformation."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"speaker": "Application administrator", "text": "In batch B-142, 15 of the records not tagged import_mode=history have a missing ship_country value."}, {"speaker": "Application administrator", "text": "No other schema errors remain outside these figures."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate-zero and 99.95% completeness rules with the history exemption, keep Batch B-142 as the bound entity, use two complete factual sentences as evidence, and the counterfactual only alters the ship_country count (19,000 vs 19,995) without contradicting other stated facts or embedding any answer/rule leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142's approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"Batch B-142 contains 20,000 records not tagged import_mode=history.\"},{\"speaker\":\"Application administrator\",\"text\":\"Among those 20,000 records, 19,995 have a populated ship_country field.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["2", "text"], "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"path": ["3", "text"], "text": "Among those 20,000 records, 19,995 have a populated ship_country field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_right": "Among those 20,000 records, 19,000 have a populated ship_country field.", "right": "Among those 20,000 records, 19,995 have a populated ship_country field."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-013", "id": "fast-43-diverse-142-013-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142's approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"speaker": "Application administrator", "text": "Among those 20,000 records, 19,995 have a populated ship_country field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate-zero and 99.95% completeness rules with the history exemption, keep Batch B-142 as the bound entity, use two complete factual sentences as evidence, and the counterfactual only alters the ship_country count (19,000 vs 19,995) without contradicting other stated facts or embedding any answer/rule leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142's approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"Batch B-142 contains 20,000 records not tagged import_mode=history.\"},{\"speaker\":\"Application administrator\",\"text\":\"Among those 20,000 records, 19,995 have a populated ship_country field.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["2", "text"], "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"path": ["3", "text"], "text": "Among those 20,000 records, 19,995 have a populated ship_country field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_right": "Among those 20,000 records, 19,000 have a populated ship_country field.", "right": "Among those 20,000 records, 19,995 have a populated ship_country field."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-013", "id": "fast-43-diverse-142-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142's approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"speaker": "Application administrator", "text": "Among those 20,000 records, 19,000 have a populated ship_country field."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the acceptance rubric, duplicate/completeness thresholds, and history exemption unchanged from the original state and question; the two focus sentences are plain factual counts, not policy text; the counterfactual only changes the non-null ship_country count (199,940→199,600) without contradicting any other stated fact, altering the completeness ratio while remaining internally consistent; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142 contains 10,000 customer records with duplicate report showing 12 duplicates before transformation and zero afterward, per the approved newest-updated_at transformation.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"Batch B-142 contains 200,000 records not tagged import_mode=history.\"},{\"speaker\":\"Source-system operator\",\"text\":\"Among those 200,000 records, 199,940 have a non-null ship_country value recorded.\"},{\"speaker\":\"Application administrator\",\"text\":\"No other schema errors remain across the batch.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "Batch B-142 contains 200,000 records not tagged import_mode=history."}, {"path": ["4", "text"], "text": "Among those 200,000 records, 199,940 have a non-null ship_country value recorded."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 200,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 200,000 records not tagged import_mode=history.", "negative_right": "Among those 200,000 records, 199,600 have a non-null ship_country value recorded.", "right": "Among those 200,000 records, 199,940 have a non-null ship_country value recorded."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-023", "id": "fast-43-diverse-142-023-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 contains 10,000 customer records with duplicate report showing 12 duplicates before transformation and zero afterward, per the approved newest-updated_at transformation."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "Batch B-142 contains 200,000 records not tagged import_mode=history."}, {"speaker": "Source-system operator", "text": "Among those 200,000 records, 199,940 have a non-null ship_country value recorded."}, {"speaker": "Application administrator", "text": "No other schema errors remain across the batch."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the acceptance rubric, duplicate/completeness thresholds, and history exemption unchanged from the original state and question; the two focus sentences are plain factual counts, not policy text; the counterfactual only changes the non-null ship_country count (199,940→199,600) without contradicting any other stated fact, altering the completeness ratio while remaining internally consistent; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142 contains 10,000 customer records with duplicate report showing 12 duplicates before transformation and zero afterward, per the approved newest-updated_at transformation.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"Batch B-142 contains 200,000 records not tagged import_mode=history.\"},{\"speaker\":\"Source-system operator\",\"text\":\"Among those 200,000 records, 199,940 have a non-null ship_country value recorded.\"},{\"speaker\":\"Application administrator\",\"text\":\"No other schema errors remain across the batch.\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "Batch B-142 contains 200,000 records not tagged import_mode=history."}, {"path": ["4", "text"], "text": "Among those 200,000 records, 199,940 have a non-null ship_country value recorded."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 200,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 200,000 records not tagged import_mode=history.", "negative_right": "Among those 200,000 records, 199,600 have a non-null ship_country value recorded.", "right": "Among those 200,000 records, 199,940 have a non-null ship_country value recorded."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-023", "id": "fast-43-diverse-142-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 contains 10,000 customer records with duplicate report showing 12 duplicates before transformation and zero afterward, per the approved newest-updated_at transformation."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"speaker": "Data quality analyst", "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "Batch B-142 contains 200,000 records not tagged import_mode=history."}, {"speaker": "Source-system operator", "text": "Among those 200,000 records, 199,600 have a non-null ship_country value recorded."}, {"speaker": "Application administrator", "text": "No other schema errors remain across the batch."}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate, completeness, and history-exemption policies and reuse the original question verbatim; evidence quotes are two exact factual sentences from the admin turn, the counterfactual coherently lowers ship_country completeness to 19,980/20,000 without contradicting other facts, and no gold answer, rule table, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\": \"Data engineer\", \"text\": \"Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward.\"}, {\"speaker\": \"Data quality analyst\", \"text\": \"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The exception report identifies 80 history records, all confirmed tagged in the source export.\"}, {\"speaker\": \"Application administrator\", \"text\": \"Batch B-142 contains 20,000 records not tagged import_mode=history. Among those records, 19,995 have a populated ship_country field. No other schema errors remain. Should B-142 be accepted?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"path": ["3", "text"], "text": "Among those records, 19,995 have a populated ship_country field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_right": "Among those records, 19,980 have a populated ship_country field.", "right": "Among those records, 19,995 have a populated ship_country field."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-027", "id": "fast-43-diverse-142-027-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "The exception report identifies 80 history records, all confirmed tagged in the source export."}, {"speaker": "Application administrator", "text": "Batch B-142 contains 20,000 records not tagged import_mode=history. Among those records, 19,995 have a populated ship_country field. No other schema errors remain. Should B-142 be accepted?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate, completeness, and history-exemption policies and reuse the original question verbatim; evidence quotes are two exact factual sentences from the admin turn, the counterfactual coherently lowers ship_country completeness to 19,980/20,000 without contradicting other facts, and no gold answer, rule table, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\": \"Data engineer\", \"text\": \"Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward.\"}, {\"speaker\": \"Data quality analyst\", \"text\": \"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"}, {\"speaker\": \"Source-system operator\", \"text\": \"The exception report identifies 80 history records, all confirmed tagged in the source export.\"}, {\"speaker\": \"Application administrator\", \"text\": \"Batch B-142 contains 20,000 records not tagged import_mode=history. Among those records, 19,995 have a populated ship_country field. No other schema errors remain. Should B-142 be accepted?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "Batch B-142 contains 20,000 records not tagged import_mode=history."}, {"path": ["3", "text"], "text": "Among those records, 19,995 have a populated ship_country field."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 20,000 records not tagged import_mode=history.", "negative_right": "Among those records, 19,980 have a populated ship_country field.", "right": "Among those records, 19,995 have a populated ship_country field."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-027", "id": "fast-43-diverse-142-027-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "The exception report identifies 80 history records, all confirmed tagged in the source export."}, {"speaker": "Application administrator", "text": "Batch B-142 contains 20,000 records not tagged import_mode=history. Among those records, 19,980 have a populated ship_country field. No other schema errors remain. Should B-142 be accepted?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate-ID and 99.95% completeness rules plus the history exemption, keep the same batch ID and entities, use two factual (non-policy) sentences as evidence, and the counterfactual only changes the missing-ship_country count from 20 to 60 without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The exception report identifies 80 history records, all confirmed tagged in the source export.\"},{\"speaker\":\"Application administrator\",\"text\":\"Batch B-142 contains 40,000 records not tagged import_mode=history. Among those 40,000 records, 20 have a missing ship_country value. The post-transformation duplicate report confirms zero duplicate customer_id values remain. Should B-142 be accepted?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "Batch B-142 contains 40,000 records not tagged import_mode=history."}, {"path": ["3", "text"], "text": "Among those 40,000 records, 20 have a missing ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 40,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 40,000 records not tagged import_mode=history.", "negative_right": "Among those 40,000 records, 60 have a missing ship_country value.", "right": "Among those 40,000 records, 20 have a missing ship_country value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-033", "id": "fast-43-diverse-142-033-base", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "The exception report identifies 80 history records, all confirmed tagged in the source export."}, {"speaker": "Application administrator", "text": "Batch B-142 contains 40,000 records not tagged import_mode=history. Among those 40,000 records, 20 have a missing ship_country value. The post-transformation duplicate report confirms zero duplicate customer_id values remain. Should B-142 be accepted?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the duplicate-ID and 99.95% completeness rules plus the history exemption, keep the same batch ID and entities, use two factual (non-policy) sentences as evidence, and the counterfactual only changes the missing-ship_country count from 20 to 60 without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "full_context_fact_states": {"base": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "supported"}, "counterfactual": {"post_transform_duplicates_zero": "supported", "scoped_ship_country_completeness_met": "refuted"}, "remove_left": {"scoped_ship_country_completeness_met": "unknown"}, "remove_right": {"scoped_ship_country_completeness_met": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"scoped_ship_country_completeness_met": "unknown"}, "negative_pair": {"scoped_ship_country_completeness_met": "refuted"}, "negative_sentence": {"scoped_ship_country_completeness_met": "unknown"}, "positive_pair": {"scoped_ship_country_completeness_met": "supported"}, "right": {"scoped_ship_country_completeness_met": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus is a factual completeness measurement rather than a policy conclusion. The policy evidence correctly preserves the state-originating zero-duplicate threshold, 99.95% scoped-completeness threshold, and history-record exception needed to interpret the unchanged question. The base and counter assignments are realizable with only the focus atom changing. The three rules are individually sufficient and properly treat unknown assignments as uncovered rather than false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refutation of the scoped completeness atom entails that one of the two mandatory acceptance requirements is unsatisfied, which is sufficient for rejection.", "rule_index": 0, "sound": true}, {"reason": "Support for both zero post-transformation duplicates and the required scoped ship_country completeness satisfies all acceptance requirements, including the documented history-record exception, and is sufficient for acceptance.", "rule_index": 1, "sound": true}, {"reason": "Refutation of the zero post-transformation duplicate condition entails failure of a mandatory acceptance requirement and is sufficient for rejection.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "post_transform_duplicates_zero", "statement": "The post-transformation duplicate customer_id count for batch B-142 is zero."}, {"id": "scoped_ship_country_completeness_met", "statement": "Among records in batch B-142 not tagged import_mode=history, observed ship_country completeness is at least 99.95%."}], "base_state_json": "[{\"speaker\":\"Data engineer\",\"text\":\"Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward.\"},{\"speaker\":\"Data quality analyst\",\"text\":\"Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search.\"},{\"speaker\":\"Source-system operator\",\"text\":\"The exception report identifies 80 history records, all confirmed tagged in the source export.\"},{\"speaker\":\"Application administrator\",\"text\":\"Batch B-142 contains 40,000 records not tagged import_mode=history. Among those 40,000 records, 20 have a missing ship_country value. The post-transformation duplicate report confirms zero duplicate customer_id values remain. Should B-142 be accepted?\"}]", "base_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}], "counter_states": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}], "focus_atom": "scoped_ship_country_completeness_met", "focus_evidence": [{"path": ["3", "text"], "text": "Batch B-142 contains 40,000 records not tagged import_mode=history."}, {"path": ["3", "text"], "text": "Among those 40,000 records, 20 have a missing ship_country value."}], "policy_evidence": [{"path": ["1", "text"], "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records."}, {"path": ["1", "text"], "text": "Records tagged import_mode=history are exempt because they are archived and excluded from product search."}], "rules": [{"justification": "The scoped ship_country completeness requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "scoped_ship_country_completeness_met", "state": "refuted"}]}, {"justification": "The post-transformation duplicate count is zero and ship_country completeness meets the required threshold after excluding records covered by the documented history exception.", "target": "true", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "supported"}, {"atom_id": "scoped_ship_country_completeness_met", "state": "supported"}]}, {"justification": "The zero-post-transformation-duplicate requirement is not satisfied, so at least one acceptance requirement fails.", "target": "false", "when": [{"atom_id": "post_transform_duplicates_zero", "state": "refuted"}]}]}, "verified_pair": {"left": "Batch B-142 contains 40,000 records not tagged import_mode=history.", "negative_left": "Batch B-142 contains 40,000 records not tagged import_mode=history.", "negative_right": "Among those 40,000 records, 60 have a missing ship_country value.", "right": "Among those 40,000 records, 20 have a missing ship_country value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-142-033", "id": "fast-43-diverse-142-033-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Reject batch B-142 because at least one acceptance requirement is not satisfied.", "true": "Accept batch B-142 because it satisfies all acceptance requirements after applying the documented scope exception."}, "instructions": "Decide whether the batch meets the stated acceptance rubric. Answer yes only if both the duplicate and scoped completeness requirements are satisfied.", "type": "noul"}}, "state": [{"speaker": "Data engineer", "text": "Batch B-142 contains 10,000 customer records. Schema requires customer_id and ship_country. The approved transformation keeps the newest updated_at record for duplicate customer_id values; the duplicate report shows 12 before transformation and zero afterward."}, {"speaker": "Data quality analyst", "text": "Acceptance requires zero post-transformation duplicate IDs and at least 99.95% ship_country completeness among in-scope records. Records tagged import_mode=history are exempt because they are archived and excluded from product search."}, {"speaker": "Source-system operator", "text": "The exception report identifies 80 history records, all confirmed tagged in the source export."}, {"speaker": "Application administrator", "text": "Batch B-142 contains 40,000 records not tagged import_mode=history. Among those 40,000 records, 60 have a missing ship_country value. The post-transformation duplicate report confirms zero duplicate customer_id values remain. Should B-142 be accepted?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-142", "source_is_synthetic": true, "source_sha256": "32ccfaacf85d44ce649c16fd3678ef6a7329a16f074ea1ccce74421f06994b3e", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the schema, blank counts, derivation rule, and duplicate/blank record lists without adding exceptions or leaking the answer; the counterfactual coherently removes R106 from the postal_country list, leaving one region_code blank unresolved, which is internally consistent; the two focus evidence spans are complete factual sentences, not rules or instructions; question bindings and criteria remain unchanged via the verbatim questions object.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Data quality analyst review of the 100-record customer batch supplied by the source-system operator, prior to applying the region_code derivation rule. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer\\u2019s rule states, \\u201cIf region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.\\u201d The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106. The record identifiers R101, R102, R103, R104, R105, R106, R201, and R202 all have present postal_country values before the stated transformation. The duplicate report finds no duplicates in the batch. The application administrator awaits the post-transformation completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106."}, {"path": [], "text": "The record identifiers R101, R102, R103, R104, R105, R106, R201, and R202 all have present postal_country values before the stated transformation."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106.", "negative_left": "The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106.", "negative_right": "The record identifiers R101, R102, R103, R104, R105, R201, and R202 all have present postal_country values before the stated transformation.", "right": "The record identifiers R101, R102, R103, R104, R105, R106, R201, and R202 all have present postal_country values before the stated transformation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-143-001", "id": "fast-43-diverse-143-001-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Data quality analyst review of the 100-record customer batch supplied by the source-system operator, prior to applying the region_code derivation rule. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106. The record identifiers R101, R102, R103, R104, R105, R106, R201, and R202 all have present postal_country values before the stated transformation. The duplicate report finds no duplicates in the batch. The application administrator awaits the post-transformation completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the schema, blank counts, derivation rule, and duplicate/blank record lists without adding exceptions or leaking the answer; the counterfactual coherently removes R106 from the postal_country list, leaving one region_code blank unresolved, which is internally consistent; the two focus evidence spans are complete factual sentences, not rules or instructions; question bindings and criteria remain unchanged via the verbatim questions object.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"Data quality analyst review of the 100-record customer batch supplied by the source-system operator, prior to applying the region_code derivation rule. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer\\u2019s rule states, \\u201cIf region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.\\u201d The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106. The record identifiers R101, R102, R103, R104, R105, R106, R201, and R202 all have present postal_country values before the stated transformation. The duplicate report finds no duplicates in the batch. The application administrator awaits the post-transformation completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106."}, {"path": [], "text": "The record identifiers R101, R102, R103, R104, R105, R106, R201, and R202 all have present postal_country values before the stated transformation."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106.", "negative_left": "The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106.", "negative_right": "The record identifiers R101, R102, R103, R104, R105, R201, and R202 all have present postal_country values before the stated transformation.", "right": "The record identifiers R101, R102, R103, R104, R105, R106, R201, and R202 all have present postal_country values before the stated transformation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-143-001", "id": "fast-43-diverse-143-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "Data quality analyst review of the 100-record customer batch supplied by the source-system operator, prior to applying the region_code derivation rule. The product schema requires customer_name, email, and region_code, giving 300 required cells. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The six record identifiers with blank required region_code cells before the stated transformation are R101, R102, R103, R104, R105, and R106. The record identifiers R101, R102, R103, R104, R105, R201, and R202 all have present postal_country values before the stated transformation. The duplicate report finds no duplicates in the batch. The application administrator awaits the post-transformation completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Swapping R106 for R107 in the postal_country list changes the derivation count without contradicting any other stated fact, so the counterfactual remains internally coherent while altering the completeness outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A data quality analyst reviews a 100-record customer batch supplied by the source-system operator. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The data engineer\\u2019s rule states, \\u201cIf region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.\\u201d The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106. Records R101, R102, R103, R104, R105, and R106 all have a present postal_country value before the stated transformation. Sample records confirm values such as CA and DE for postal_country. The duplicate report finds no duplicates. The application administrator requests the post-transformation completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106."}, {"path": [], "text": "Records R101, R102, R103, R104, R105, and R106 all have a present postal_country value before the stated transformation."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106.", "negative_left": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106.", "negative_right": "Records R101, R102, R103, R104, R105, and R107 all have a present postal_country value before the stated transformation.", "right": "Records R101, R102, R103, R104, R105, and R106 all have a present postal_country value before the stated transformation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-143-003", "id": "fast-43-diverse-143-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A data quality analyst reviews a 100-record customer batch supplied by the source-system operator. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106. Records R101, R102, R103, R104, R105, and R106 all have a present postal_country value before the stated transformation. Sample records confirm values such as CA and DE for postal_country. The duplicate report finds no duplicates. The application administrator requests the post-transformation completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Swapping R106 for R107 in the postal_country list changes the derivation count without contradicting any other stated fact, so the counterfactual remains internally coherent while altering the completeness outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A data quality analyst reviews a 100-record customer batch supplied by the source-system operator. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The data engineer\\u2019s rule states, \\u201cIf region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.\\u201d The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106. Records R101, R102, R103, R104, R105, and R106 all have a present postal_country value before the stated transformation. Sample records confirm values such as CA and DE for postal_country. The duplicate report finds no duplicates. The application administrator requests the post-transformation completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106."}, {"path": [], "text": "Records R101, R102, R103, R104, R105, and R106 all have a present postal_country value before the stated transformation."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106.", "negative_left": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106.", "negative_right": "Records R101, R102, R103, R104, R105, and R107 all have a present postal_country value before the stated transformation.", "right": "Records R101, R102, R103, R104, R105, and R106 all have a present postal_country value before the stated transformation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-143-003", "id": "fast-43-diverse-143-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A data quality analyst reviews a 100-record customer batch supplied by the source-system operator. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106. Records R101, R102, R103, R104, R105, and R107 all have a present postal_country value before the stated transformation. Sample records confirm values such as CA and DE for postal_country. The duplicate report finds no duplicates. The application administrator requests the post-transformation completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the schema, blank counts, and derivation rule from the original state, matching the unchanged question. The two evidence sentences are plain factual statements, not policy text. The counterfactual swaps R106 for R107 in the postal_country-present list, a single coherent factual edit that changes which region_code blanks are derivable without contradicting other stated facts. Neither context contains a computed percentage, score label, or instruction revealing the intended answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A data quality analyst reviews a 100-record customer batch supplied by the source-system operator. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The data engineer\\u2019s rule states, \\u201cIf region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.\\u201d The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106. Records R101, R102, R103, R104, R105, and R106 are all among the records with present postal_country values. The duplicate report finds no duplicates. The application administrator requests the post-transformation completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106."}, {"path": [], "text": "Records R101, R102, R103, R104, R105, and R106 are all among the records with present postal_country values."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106.", "negative_left": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106.", "negative_right": "Records R101, R102, R103, R104, R105, and R107 are all among the records with present postal_country values.", "right": "Records R101, R102, R103, R104, R105, and R106 are all among the records with present postal_country values."}, "verifier_independent_model": false}, "family": "fast-43-diverse-143-004", "id": "fast-43-diverse-143-004-base", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A data quality analyst reviews a 100-record customer batch supplied by the source-system operator. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106. Records R101, R102, R103, R104, R105, and R106 are all among the records with present postal_country values. The duplicate report finds no duplicates. The application administrator requests the post-transformation completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "software-04", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the schema, blank counts, and derivation rule from the original state, matching the unchanged question. The two evidence sentences are plain factual statements, not policy text. The counterfactual swaps R106 for R107 in the postal_country-present list, a single coherent factual edit that changes which region_code blanks are derivable without contradicting other stated facts. Neither context contains a computed percentage, score label, or instruction revealing the intended answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a4 is an allowed universally quantified relationship over an explicit set. The focus is factual rather than policy. The base and counter assignments can both occur while changing only whether all six affected records have present postal_country values. Policy evidence correctly preserves the state-originating schema/denominator and transformation rule; scoring thresholds and instructions remain automatically available in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail 300 required cells, 12 initial blanks, and successful derivation for all six blank region_code cells. Thus 294/300 cells are populated, exactly 98%, which is level 3.", "rule_index": 0, "sound": true}, {"reason": "Refuting a4 entails that at least one of the six blank region_code cells lacks a present postal_country, so between zero and five region_code blanks are derived. The resulting completeness ranges from 288/300 (96%) to 293/300 (97.67%), always level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The customer batch has exactly 300 required cells before the stated transformation."}, {"id": "a2", "statement": "Exactly 12 required cells in the customer batch are blank before the stated transformation."}, {"id": "a3", "statement": "Exactly six required region_code cells in the customer batch are blank before the stated transformation."}, {"id": "a4", "statement": "The six record identifiers attached to blank required region_code cells are all included among the record identifiers attached to present postal_country values before the stated transformation."}], "base_state_json": "\"A data quality analyst reviews a 100-record customer batch supplied by the source-system operator. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The data engineer\\u2019s rule states, \\u201cIf region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.\\u201d The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106. Records R101, R102, R103, R104, R105, and R106 are all among the records with present postal_country values. The duplicate report finds no duplicates. The application administrator requests the post-transformation completeness rating.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a4", "focus_evidence": [{"path": [], "text": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106."}, {"path": [], "text": "Records R101, R102, R103, R104, R105, and R106 are all among the records with present postal_country values."}], "policy_evidence": [{"path": [], "text": "The product schema requires customer_name, email, and region_code, giving 300 required cells."}, {"path": [], "text": "The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.”"}], "rules": [{"justification": "All six blank region_code cells qualify for derivation, so six of the 12 blanks become populated. The result is 294 populated required cells out of 300, or 98%, which matches level 3.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "At least one of the six blank region_code cells lacks a present postal_country value, so at most five blanks are populated by the stated rule. Post-transformation completeness is therefore between 96% and 97.67%, which is at least 95% but below 98% and matches level 2.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106.", "negative_left": "The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106.", "negative_right": "Records R101, R102, R103, R104, R105, and R107 are all among the records with present postal_country values.", "right": "Records R101, R102, R103, R104, R105, and R106 are all among the records with present postal_country values."}, "verifier_independent_model": false}, "family": "fast-43-diverse-143-004", "id": "fast-43-diverse-143-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Post-transformation completeness is below 90%.", "1 — Post-transformation completeness is at least 90% but below 95%.", "2 — Post-transformation completeness is at least 95% but below 98%.", "3 — Post-transformation completeness is at least 98% but below 99.5%.", "4 — Post-transformation completeness is at least 99.5%."], "instructions": "Calculate the percentage of required cells populated after applying the stated transformation rule, then select the matching completeness level.", "type": "score"}}, "state": "A data quality analyst reviews a 100-record customer batch supplied by the source-system operator. The product schema requires customer_name, email, and region_code, giving 300 required cells. The validation summary reports 12 blanks: six region_code, three email, and three customer_name. The data engineer’s rule states, “If region_code is blank and postal_country is present, derive region_code from postal_country; after derivation, it counts as populated.” The six blank required region_code cells belong to records R101, R102, R103, R104, R105, and R106. Records R101, R102, R103, R104, R105, and R107 are all among the records with present postal_country values. The duplicate report finds no duplicates. The application administrator requests the post-transformation completeness rating."}, "method": "c2d", "provenance": {"source_id": "diverse-143", "source_is_synthetic": true, "source_sha256": "2c276ea88799bbf64d057802c7d9443412b7cef0d06d0a116aff0b9bec307a98", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "software-04", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy language and gate ownership rules; only the defect ID/severity observation changed, entity and time bindings match the original question, evidence spans are two factual sentences about the defect list and severity, the counterfactual's severity change to Sev-4 is internally coherent with no duplicate or contradictory measurements, and neither context reveals a gold answer or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Release note for Northstar Notes 6.2 RC3: The compatibility tester's signed report confirms all required Windows 11 and macOS 15 checks pass. The deployment specialist's signed packaging report confirms valid signatures and clean installs, satisfying both the installation and signing gates. However, the rollback step leaves the updater service disabled, so the rollback gate fails. The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open. Defect DEF-4471 is rated Sev-2 severity. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open."}, {"path": [], "text": "Defect DEF-4471 is rated Sev-2 severity."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open.", "negative_right": "Defect DEF-4471 is rated Sev-4 severity.", "right": "Defect DEF-4471 is rated Sev-2 severity."}, "verifier_independent_model": false}, "family": "fast-43-diverse-145-012", "id": "fast-43-diverse-145-012-base", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Release note for Northstar Notes 6.2 RC3: The compatibility tester's signed report confirms all required Windows 11 and macOS 15 checks pass. The deployment specialist's signed packaging report confirms valid signatures and clean installs, satisfying both the installation and signing gates. However, the rollback step leaves the updater service disabled, so the rollback gate fails. The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open. Defect DEF-4471 is rated Sev-2 severity. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy language and gate ownership rules; only the defect ID/severity observation changed, entity and time bindings match the original question, evidence spans are two factual sentences about the defect list and severity, the counterfactual's severity change to Sev-4 is internally coherent with no duplicate or contradictory measurements, and neither context reveals a gold answer or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; a2 is a permitted universal claim over an explicit platform/check scope. The focus a1 is factual rather than a policy conclusion. The base and counter assignments differ only on a1 and are realizable as synthetic scenarios. The policy evidence preserves the state-originating defect/install/rollback requirements, evidence priority, and gate-owner bindings; decision criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish a failed development-owned defect gate and a failed packaging-owned rollback gate, while compatibility, installation, and signing pass. Because failures span two owners, none_of_above is entailed.", "rule_index": 0, "sound": true}, {"reason": "Refuting the nonnegative defect count being at least 1 establishes zero open Sev-1/Sev-2 defects. Compatibility, installation, and signing pass, while rollback alone fails and is packaging-owned, so single_owner_blocked is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The count of defects that are both open and rated Sev-1 or Sev-2 on the authoritative signed defect list for Northstar Notes 6.2 RC3 is at least 1."}, {"id": "a2", "statement": "Every required Windows 11 and macOS 15 compatibility check for Northstar Notes 6.2 RC3 passes."}, {"id": "a3", "statement": "The required installation gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a4", "statement": "The required signing gate for Northstar Notes 6.2 RC3 passes."}, {"id": "a5", "statement": "The required rollback gate for Northstar Notes 6.2 RC3 passes."}], "base_state_json": "\"Release note for Northstar Notes 6.2 RC3: The compatibility tester's signed report confirms all required Windows 11 and macOS 15 checks pass. The deployment specialist's signed packaging report confirms valid signatures and clean installs, satisfying both the installation and signing gates. However, the rollback step leaves the updater service disabled, so the rollback gate fails. The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open. Defect DEF-4471 is rated Sev-2 severity. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open."}, {"path": [], "text": "Defect DEF-4471 is rated Sev-2 severity."}], "policy_evidence": [{"path": [], "text": "Policy requires no open Sev-1/Sev-2 defects and successful install and rollback."}, {"path": [], "text": "Signed detailed reports override dashboard summaries."}, {"path": [], "text": "A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}], "rules": [{"justification": "At least one open Sev-1/Sev-2 defect fails the development-owned defect gate, while the failed rollback gate belongs to packaging. All other required gates pass, so the failed gates span multiple owners.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}, {"justification": "Explicit refutation of a count of at least 1 entails no open Sev-1/Sev-2 defects on the authoritative list. With compatibility, installation, and signing passing, rollback is the only failed gate, and it belongs exclusively to packaging.", "target": "single_owner_blocked", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open.", "negative_left": "The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open.", "negative_right": "Defect DEF-4471 is rated Sev-4 severity.", "right": "Defect DEF-4471 is rated Sev-2 severity."}, "verifier_independent_model": false}, "family": "fast-43-diverse-145-012", "id": "fast-43-diverse-145-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"none_of_above": "The release is not ready and its failed gates span multiple owners, or the evidence cannot support either substantive classification.", "release_ready": "Every required defect, compatibility, installation, signing, and rollback gate passes under the authoritative evidence.", "single_owner_blocked": "At least one release gate fails, and all failed gates belong exclusively to one owner: development, compatibility, packaging, or release operations."}, "instructions": "Classify release readiness using the supplied policy and evidence. Select “release_ready” only if every gate passes. Select “single_owner_blocked” only if the release fails but every failed gate belongs to one owner. Otherwise select “none_of_above.”", "type": "choice"}}, "state": "Release note for Northstar Notes 6.2 RC3: The compatibility tester's signed report confirms all required Windows 11 and macOS 15 checks pass. The deployment specialist's signed packaging report confirms valid signatures and clean installs, satisfying both the installation and signing gates. However, the rollback step leaves the updater service disabled, so the rollback gate fails. The authoritative signed defect list for Northstar Notes 6.2 RC3 contains exactly one entry, defect DEF-4471, currently marked open. Defect DEF-4471 is rated Sev-4 severity. Policy requires no open Sev-1/Sev-2 defects and successful install and rollback. Signed detailed reports override dashboard summaries. A failed defect gate belongs to development; a failed rollback gate belongs to packaging."}, "method": "c2d", "provenance": {"source_id": "diverse-145", "source_is_synthetic": true, "source_sha256": "8191e5f5e30c0bcd368e7d76f9c21a8eead91b8e429911582e02d2b098dec9dd", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "single_owner_blocked"}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy verbatim and the same entity/defect bindings; the two focus sentences are plain factual statements, not policy or rulings; changing the ticket number in the counterfactual alters a case fact (breaking the D=S link) without contradicting any other stated fact, so it remains coherent; no gold answer, code, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"For Northstar Notes 4.2 RC3, the release manager reviews 186 passing functional tests and one open blocker, defect D, which alone blocks the release gate. The deployment specialist finds that the Windows installer omitted the required signature from the updater executable, and the packaging report confirms this omission as defect S. Every defect in this release that concerns installer construction or signing is identical to defect S. Defect D was logged in the release tracker under ticket #4471. Ticket #4471 in the release tracker is the official identifier assigned to defect S. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "Defect D was logged in the release tracker under ticket #4471."}, {"path": [], "text": "Ticket #4471 in the release tracker is the official identifier assigned to defect S."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "Defect D was logged in the release tracker under ticket #4471.", "negative_left": "Defect D was logged in the release tracker under ticket #4471.", "negative_right": "Ticket #4472 in the release tracker is the official identifier assigned to defect S.", "right": "Ticket #4471 in the release tracker is the official identifier assigned to defect S."}, "verifier_independent_model": false}, "family": "fast-43-diverse-148-021", "id": "fast-43-diverse-148-021-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "For Northstar Notes 4.2 RC3, the release manager reviews 186 passing functional tests and one open blocker, defect D, which alone blocks the release gate. The deployment specialist finds that the Windows installer omitted the required signature from the updater executable, and the packaging report confirms this omission as defect S. Every defect in this release that concerns installer construction or signing is identical to defect S. Defect D was logged in the release tracker under ticket #4471. Ticket #4471 in the release tracker is the official identifier assigned to defect S. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy verbatim and the same entity/defect bindings; the two focus sentences are plain factual statements, not policy or rulings; changing the ticket number in the counterfactual alters a case fact (breaking the D=S link) without contradicting any other stated fact, so it remains coherent; no gold answer, code, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"For Northstar Notes 4.2 RC3, the release manager reviews 186 passing functional tests and one open blocker, defect D, which alone blocks the release gate. The deployment specialist finds that the Windows installer omitted the required signature from the updater executable, and the packaging report confirms this omission as defect S. Every defect in this release that concerns installer construction or signing is identical to defect S. Defect D was logged in the release tracker under ticket #4471. Ticket #4471 in the release tracker is the official identifier assigned to defect S. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "Defect D was logged in the release tracker under ticket #4471."}, {"path": [], "text": "Ticket #4471 in the release tracker is the official identifier assigned to defect S."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "Defect D was logged in the release tracker under ticket #4471.", "negative_left": "Defect D was logged in the release tracker under ticket #4471.", "negative_right": "Ticket #4472 in the release tracker is the official identifier assigned to defect S.", "right": "Ticket #4471 in the release tracker is the official identifier assigned to defect S."}, "verifier_independent_model": false}, "family": "fast-43-diverse-148-021", "id": "fast-43-diverse-148-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "For Northstar Notes 4.2 RC3, the release manager reviews 186 passing functional tests and one open blocker, defect D, which alone blocks the release gate. The deployment specialist finds that the Windows installer omitted the required signature from the updater executable, and the packaging report confirms this omission as defect S. Every defect in this release that concerns installer construction or signing is identical to defect S. Defect D was logged in the release tracker under ticket #4471. Ticket #4472 in the release tracker is the official identifier assigned to defect S. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy routing rules and question scope, only altering the ticket ID linking defect S to a different ticket than defect D, which is a coherent factual change since ticket uniqueness only guarantees one defect per ID, not that D and S must share an ID; the two evidence spans are complete factual sentences about ticket recording rather than policy statements, and neither context states or hints at the intended routing answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"For Northstar Notes 4.2 RC3, the release manager reviews 186 passing functional tests and one open blocker, which the deployment specialist identifies as defect D. The desktop developer confirms the application binaries run without crashes. The compatibility tester reports passes on Windows 11, macOS 15, and Ubuntu 24.04. Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3, and this triggers an administrator warning. The packaging report reproduces this omission, while the rollback plan successfully restores 4.1. Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S. The QA tracker logs defect identities by ticket ID. Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect. Defect S is recorded in the QA tracker under ticket ID INS-4471. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations. Any open blocker fails the release gate.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect."}, {"path": [], "text": "Defect S is recorded in the QA tracker under ticket ID INS-4471."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect.", "negative_left": "Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect.", "negative_right": "Defect S is recorded in the QA tracker under ticket ID INS-5820.", "right": "Defect S is recorded in the QA tracker under ticket ID INS-4471."}, "verifier_independent_model": false}, "family": "fast-43-diverse-148-022", "id": "fast-43-diverse-148-022-base", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "For Northstar Notes 4.2 RC3, the release manager reviews 186 passing functional tests and one open blocker, which the deployment specialist identifies as defect D. The desktop developer confirms the application binaries run without crashes. The compatibility tester reports passes on Windows 11, macOS 15, and Ubuntu 24.04. Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3, and this triggers an administrator warning. The packaging report reproduces this omission, while the rollback plan successfully restores 4.1. Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S. The QA tracker logs defect identities by ticket ID. Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect. Defect S is recorded in the QA tracker under ticket ID INS-4471. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations. Any open blocker fails the release gate."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy routing rules and question scope, only altering the ticket ID linking defect S to a different ticket than defect D, which is a coherent factual change since ticket uniqueness only guarantees one defect per ID, not that D and S must share an ID; the two evidence spans are complete factual sentences about ticket recording rather than policy statements, and neither context states or hints at the intended routing answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "full_context_fact_states": {"base": {"a_blocker_defect": "supported", "a_defect_identity": "supported", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "counterfactual": {"a_blocker_defect": "supported", "a_defect_identity": "refuted", "a_signing_defect": "supported", "a_signing_scope": "supported"}, "remove_left": {"a_defect_identity": "unknown"}, "remove_right": {"a_defect_identity": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_defect_identity": "unknown"}, "negative_pair": {"a_defect_identity": "refuted"}, "negative_sentence": {"a_defect_identity": "unknown"}, "positive_pair": {"a_defect_identity": "supported"}, "right": {"a_defect_identity": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the permitted universally quantified scope relationship. The focus is the factual identity relation D = S, not a policy conclusion. Both assignments are jointly realizable while changing only that identity: in the base D is the signing defect S; in the counter D is a distinct non-installer/signing defect while S remains the only installer/signing defect. The policy evidence accurately preserves the relevant routing policy from the original state; the unchanged question also independently retains the decision criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the blocker’s sole defect D with S, and S is the omission of a required signature from an executable in the Windows installer. That necessarily concerns installer signing, so the stated policy routes it to packaging.", "rule_index": 0, "sound": true}, {"reason": "The blocker concerns exactly one defect, D. The universal scope atom says every installer-construction or signing defect in the release is S, while D is explicitly not S. Therefore D is not such a defect, and because D is the blocker’s sole defect, the blocker does not concern installer construction or signing. The false outcome follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_blocker_defect", "statement": "The blocker described by the deployment specialist for Northstar Notes 4.2 RC3 concerns exactly one defect, D."}, {"id": "a_signing_defect", "statement": "Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3."}, {"id": "a_signing_scope", "statement": "Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S."}, {"id": "a_defect_identity", "statement": "Defect D is identical to defect S."}], "base_state_json": "\"For Northstar Notes 4.2 RC3, the release manager reviews 186 passing functional tests and one open blocker, which the deployment specialist identifies as defect D. The desktop developer confirms the application binaries run without crashes. The compatibility tester reports passes on Windows 11, macOS 15, and Ubuntu 24.04. Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3, and this triggers an administrator warning. The packaging report reproduces this omission, while the rollback plan successfully restores 4.1. Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S. The QA tracker logs defect identities by ticket ID. Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect. Defect S is recorded in the QA tracker under ticket ID INS-4471. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations. Any open blocker fails the release gate.\"", "base_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}], "counter_states": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}], "focus_atom": "a_defect_identity", "focus_evidence": [{"path": [], "text": "Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect."}, {"path": [], "text": "Defect S is recorded in the QA tracker under ticket ID INS-4471."}], "policy_evidence": [{"path": [], "text": "Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations."}], "rules": [{"justification": "The blocker’s sole defect is the identified Windows-installer signing defect, so the blocker concerns installer signing and must be routed to packaging.", "target": "true", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "supported"}]}, {"justification": "The blocker concerns only defect D; D is not S, and S is the only defect in the release concerning installer construction or signing. Therefore the blocker does not concern installer construction or signing and is not routed to packaging.", "target": "false", "when": [{"atom_id": "a_blocker_defect", "state": "supported"}, {"atom_id": "a_signing_defect", "state": "supported"}, {"atom_id": "a_signing_scope", "state": "supported"}, {"atom_id": "a_defect_identity", "state": "refuted"}]}]}, "verified_pair": {"left": "Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect.", "negative_left": "Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect.", "negative_right": "Defect S is recorded in the QA tracker under ticket ID INS-5820.", "right": "Defect S is recorded in the QA tracker under ticket ID INS-4471."}, "verifier_independent_model": false}, "family": "fast-43-diverse-148-022", "id": "fast-43-diverse-148-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — route the blocker to another team because it does not concern installer construction or signing.", "true": "Yes — route the blocker to packaging because it concerns installer construction or signing."}, "instructions": "Decide whether the blocker described by the deployment specialist should be routed to packaging under the stated policy.", "type": "noul"}}, "state": "For Northstar Notes 4.2 RC3, the release manager reviews 186 passing functional tests and one open blocker, which the deployment specialist identifies as defect D. The desktop developer confirms the application binaries run without crashes. The compatibility tester reports passes on Windows 11, macOS 15, and Ubuntu 24.04. Defect S is the omission of the required signature from the updater executable in the Windows installer for Northstar Notes 4.2 RC3, and this triggers an administrator warning. The packaging report reproduces this omission, while the rollback plan successfully restores 4.1. Every defect in Northstar Notes 4.2 RC3 that concerns installer construction or signing is identical to defect S. The QA tracker logs defect identities by ticket ID. Defect D is recorded in the QA tracker under ticket ID INS-4471, and each ticket ID in that tracker denotes exactly one defect. Defect S is recorded in the QA tracker under ticket ID INS-5820. Policy routes installer construction or signing defects to packaging, application-runtime defects to development, compatibility failures to testing, and rollback failures to release operations. Any open blocker fails the release gate."}, "method": "c2d", "provenance": {"source_id": "diverse-148", "source_is_synthetic": true, "source_sha256": "de345cda38464d194b35a328f1d84865aaf005db786949a710878e36c09987a1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the full governing policy and same RC3/Windows MSI/team bindings from the original question; the two focus evidence lines are factual statements about the rollback spec and observed registry state, not policy or rule text; the counterfactual only flips the registry outcome, which coherently creates a rollback gate failure without contradicting the separate defect-tracker fact; no evidence states or implies the final readiness decision or team output beyond factual role designation.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"Release manager case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"Compatibility testing confirms all tests pass on supported Windows 11 23H2 and macOS 15.\", \"All other installer rollbacks (macOS pkg, Linux tarball) completed successfully with no residual files or registrations.\", \"No release-blocking defect remains open in the tracker for RC3.\", \"Both Windows and macOS packages carry valid signatures.\", \"The release record documents a nonblocking typo in the keyboard-shortcuts table, which still awaits an approved release-note action.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\\\Software\\\\Northstar\\\\Notes be absent after rollback.\", \"The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows all registry keys under HKLM\\\\Software\\\\Northstar\\\\Notes absent.\", \"The deployment team is designated as responsible for resolving any Windows MSI rollback gate failure.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "5"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\Software\\Northstar\\Notes be absent after rollback."}, {"path": ["evidence", "6"], "text": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows all registry keys under HKLM\\Software\\Northstar\\Notes absent."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\Software\\Northstar\\Notes be absent after rollback.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\Software\\Northstar\\Notes be absent after rollback.", "negative_right": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows one registry key under HKLM\\Software\\Northstar\\Notes still present.", "right": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows all registry keys under HKLM\\Software\\Northstar\\Notes absent."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-002", "id": "fast-43-diverse-149-002-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Release manager case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["Compatibility testing confirms all tests pass on supported Windows 11 23H2 and macOS 15.", "All other installer rollbacks (macOS pkg, Linux tarball) completed successfully with no residual files or registrations.", "No release-blocking defect remains open in the tracker for RC3.", "Both Windows and macOS packages carry valid signatures.", "The release record documents a nonblocking typo in the keyboard-shortcuts table, which still awaits an approved release-note action.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\Software\\Northstar\\Notes be absent after rollback.", "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows all registry keys under HKLM\\Software\\Northstar\\Notes absent.", "The deployment team is designated as responsible for resolving any Windows MSI rollback gate failure."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the full governing policy and same RC3/Windows MSI/team bindings from the original question; the two focus evidence lines are factual statements about the rollback spec and observed registry state, not policy or rule text; the counterfactual only flips the registry outcome, which coherently creates a rollback gate failure without contradicting the separate defect-tracker fact; no evidence states or implies the final readiness decision or team output beyond factual role designation.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"Release manager case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"Compatibility testing confirms all tests pass on supported Windows 11 23H2 and macOS 15.\", \"All other installer rollbacks (macOS pkg, Linux tarball) completed successfully with no residual files or registrations.\", \"No release-blocking defect remains open in the tracker for RC3.\", \"Both Windows and macOS packages carry valid signatures.\", \"The release record documents a nonblocking typo in the keyboard-shortcuts table, which still awaits an approved release-note action.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\\\Software\\\\Northstar\\\\Notes be absent after rollback.\", \"The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows all registry keys under HKLM\\\\Software\\\\Northstar\\\\Notes absent.\", \"The deployment team is designated as responsible for resolving any Windows MSI rollback gate failure.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "5"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\Software\\Northstar\\Notes be absent after rollback."}, {"path": ["evidence", "6"], "text": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows all registry keys under HKLM\\Software\\Northstar\\Notes absent."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\Software\\Northstar\\Notes be absent after rollback.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\Software\\Northstar\\Notes be absent after rollback.", "negative_right": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows one registry key under HKLM\\Software\\Northstar\\Notes still present.", "right": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows all registry keys under HKLM\\Software\\Northstar\\Notes absent."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-002", "id": "fast-43-diverse-149-002-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Release manager case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["Compatibility testing confirms all tests pass on supported Windows 11 23H2 and macOS 15.", "All other installer rollbacks (macOS pkg, Linux tarball) completed successfully with no residual files or registrations.", "No release-blocking defect remains open in the tracker for RC3.", "Both Windows and macOS packages carry valid signatures.", "The release record documents a nonblocking typo in the keyboard-shortcuts table, which still awaits an approved release-note action.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that all registry keys under HKLM\\Software\\Northstar\\Notes be absent after rollback.", "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows one registry key under HKLM\\Software\\Northstar\\Notes still present.", "The deployment team is designated as responsible for resolving any Windows MSI rollback gate failure."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full original policy and identical request/question bindings; the two focus evidence sentences are factual spec-vs-observed statements, not policy text; the counterfactual's 'present' registry state coherently implies the Windows MSI rollback failure consistent with the other evidence, with no embedded gold answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\\\Software\\\\Northstar\\\\Notes to be absent after rollback.\", \"The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\\\Software\\\\Northstar\\\\Notes absent.\", \"All other RC3 installer rollbacks, including macOS pkg rollback, completed successfully.\", \"No release-blocking defect remains open for RC3.\", \"Both Windows and macOS packages carry valid signatures.\", \"Documentation records a nonblocking typo in the keyboard-shortcuts table, still awaiting an approved release-note action.\", \"The deployment team is designated as owner for resolving any Windows MSI rollback gate failure, should one occur.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback."}, {"path": ["evidence", "2"], "text": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "negative_right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes still present.", "right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-007", "id": "fast-43-diverse-149-007-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent.", "All other RC3 installer rollbacks, including macOS pkg rollback, completed successfully.", "No release-blocking defect remains open for RC3.", "Both Windows and macOS packages carry valid signatures.", "Documentation records a nonblocking typo in the keyboard-shortcuts table, still awaiting an approved release-note action.", "The deployment team is designated as owner for resolving any Windows MSI rollback gate failure, should one occur."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full original policy and identical request/question bindings; the two focus evidence sentences are factual spec-vs-observed statements, not policy text; the counterfactual's 'present' registry state coherently implies the Windows MSI rollback failure consistent with the other evidence, with no embedded gold answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\\\Software\\\\Northstar\\\\Notes to be absent after rollback.\", \"The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\\\Software\\\\Northstar\\\\Notes absent.\", \"All other RC3 installer rollbacks, including macOS pkg rollback, completed successfully.\", \"No release-blocking defect remains open for RC3.\", \"Both Windows and macOS packages carry valid signatures.\", \"Documentation records a nonblocking typo in the keyboard-shortcuts table, still awaiting an approved release-note action.\", \"The deployment team is designated as owner for resolving any Windows MSI rollback gate failure, should one occur.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback."}, {"path": ["evidence", "2"], "text": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "negative_right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes still present.", "right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-007", "id": "fast-43-diverse-149-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes still present.", "All other RC3 installer rollbacks, including macOS pkg rollback, completed successfully.", "No release-blocking defect remains open for RC3.", "Both Windows and macOS packages carry valid signatures.", "Documentation records a nonblocking typo in the keyboard-shortcuts table, still awaiting an approved release-note action.", "The deployment team is designated as owner for resolving any Windows MSI rollback gate failure, should one occur."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep all four gating policies and the original question unchanged; the two focus evidence sentences are factual, non-policy-definition statements; the counterfactual only flips the observed registry state (present vs absent) for the Windows MSI, remaining consistent with the rest of the unchanged context; no readiness label, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"The release manager is evaluating Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\\\Software\\\\Northstar\\\\Notes to be absent after rollback.\", \"The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\\\Software\\\\Northstar\\\\Notes absent.\", \"The macOS installer rollback also succeeded, restoring the prior application bundle and preferences without residue.\", \"No release-blocking defect is currently open against RC3.\", \"Both Windows and macOS packages have valid signatures.\", \"Documentation has a nonblocking typo in the keyboard-shortcuts table, and this typo is documented in the release record.\", \"The typo still needs an approved release-note action before shipment.\", \"The deployment team is designated as responsible for resolving any Windows MSI rollback gate failure, should one occur.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback."}, {"path": ["evidence", "2"], "text": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "negative_right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes still present.", "right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-008", "id": "fast-43-diverse-149-008-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "The release manager is evaluating Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent.", "The macOS installer rollback also succeeded, restoring the prior application bundle and preferences without residue.", "No release-blocking defect is currently open against RC3.", "Both Windows and macOS packages have valid signatures.", "Documentation has a nonblocking typo in the keyboard-shortcuts table, and this typo is documented in the release record.", "The typo still needs an approved release-note action before shipment.", "The deployment team is designated as responsible for resolving any Windows MSI rollback gate failure, should one occur."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep all four gating policies and the original question unchanged; the two focus evidence sentences are factual, non-policy-definition statements; the counterfactual only flips the observed registry state (present vs absent) for the Windows MSI, remaining consistent with the rest of the unchanged context; no readiness label, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"The release manager is evaluating Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\\\Software\\\\Northstar\\\\Notes to be absent after rollback.\", \"The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\\\Software\\\\Northstar\\\\Notes absent.\", \"The macOS installer rollback also succeeded, restoring the prior application bundle and preferences without residue.\", \"No release-blocking defect is currently open against RC3.\", \"Both Windows and macOS packages have valid signatures.\", \"Documentation has a nonblocking typo in the keyboard-shortcuts table, and this typo is documented in the release record.\", \"The typo still needs an approved release-note action before shipment.\", \"The deployment team is designated as responsible for resolving any Windows MSI rollback gate failure, should one occur.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback."}, {"path": ["evidence", "2"], "text": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "negative_right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes still present.", "right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes absent."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-008", "id": "fast-43-diverse-149-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "The release manager is evaluating Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires the registry key HKLM\\Software\\Northstar\\Notes to be absent after rollback.", "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes still present.", "The macOS installer rollback also succeeded, restoring the prior application bundle and preferences without residue.", "No release-blocking defect is currently open against RC3.", "Both Windows and macOS packages have valid signatures.", "Documentation has a nonblocking typo in the keyboard-shortcuts table, and this typo is documented in the release record.", "The typo still needs an approved release-note action before shipment.", "The deployment team is designated as responsible for resolving any Windows MSI rollback gate failure, should one occur."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same RC3 policy and entity bindings, add a factual spec/observed-value pair as focus evidence, and the counterfactual coherently swaps the observed registry value to 6.2.0 creating a rollback mismatch without contradicting other evidence or leaking the intended readiness label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":\"Release readiness case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\",\"evidence\":[\"The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.\",\"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\\\Software\\\\Northstar\\\\Notes\\\\Version reads 6.1.4.\",\"The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\\\Software\\\\Northstar\\\\Notes\\\\Version reading 6.1.4.\",\"All other installer rollbacks (macOS pkg, Linux deb/rpm) completed successfully.\",\"No release-blocking defect remains open for RC3.\",\"Both Windows and macOS packages have valid signatures.\",\"Documentation records a nonblocking typo in the keyboard-shortcuts table, still awaiting an approved release-note action.\",\"The Northstar Notes deployment team is designated to resolve any Windows MSI rollback gate failure.\"],\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\Software\\Northstar\\Notes\\Version reads 6.1.4."}, {"path": ["evidence", "2"], "text": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes\\Version reading 6.1.4."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\Software\\Northstar\\Notes\\Version reads 6.1.4.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\Software\\Northstar\\Notes\\Version reads 6.1.4.", "negative_right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes\\Version reading 6.2.0.", "right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes\\Version reading 6.1.4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-009", "id": "fast-43-diverse-149-009-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Release readiness case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\Software\\Northstar\\Notes\\Version reads 6.1.4.", "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes\\Version reading 6.1.4.", "All other installer rollbacks (macOS pkg, Linux deb/rpm) completed successfully.", "No release-blocking defect remains open for RC3.", "Both Windows and macOS packages have valid signatures.", "Documentation records a nonblocking typo in the keyboard-shortcuts table, still awaiting an approved release-note action.", "The Northstar Notes deployment team is designated to resolve any Windows MSI rollback gate failure."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same RC3 policy and entity bindings, add a factual spec/observed-value pair as focus evidence, and the counterfactual coherently swaps the observed registry value to 6.2.0 creating a rollback mismatch without contradicting other evidence or leaking the intended readiness label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\":\"Release readiness case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\",\"evidence\":[\"The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.\",\"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\\\Software\\\\Northstar\\\\Notes\\\\Version reads 6.1.4.\",\"The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\\\Software\\\\Northstar\\\\Notes\\\\Version reading 6.1.4.\",\"All other installer rollbacks (macOS pkg, Linux deb/rpm) completed successfully.\",\"No release-blocking defect remains open for RC3.\",\"Both Windows and macOS packages have valid signatures.\",\"Documentation records a nonblocking typo in the keyboard-shortcuts table, still awaiting an approved release-note action.\",\"The Northstar Notes deployment team is designated to resolve any Windows MSI rollback gate failure.\"],\"request\":\"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\Software\\Northstar\\Notes\\Version reads 6.1.4."}, {"path": ["evidence", "2"], "text": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes\\Version reading 6.1.4."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\Software\\Northstar\\Notes\\Version reads 6.1.4.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\Software\\Northstar\\Notes\\Version reads 6.1.4.", "negative_right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes\\Version reading 6.2.0.", "right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes\\Version reading 6.1.4."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-009", "id": "fast-43-diverse-149-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Release readiness case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback the registry key HKLM\\Software\\Northstar\\Notes\\Version reads 6.1.4.", "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows the registry key HKLM\\Software\\Northstar\\Notes\\Version reading 6.2.0.", "All other installer rollbacks (macOS pkg, Linux deb/rpm) completed successfully.", "No release-blocking defect remains open for RC3.", "Both Windows and macOS packages have valid signatures.", "Documentation records a nonblocking typo in the keyboard-shortcuts table, still awaiting an approved release-note action.", "The Northstar Notes deployment team is designated to resolve any Windows MSI rollback gate failure."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original release policy in the unchanged question, use the same entity/time bindings, provide two complete factual evidence sentences (rollback spec and observed state), and the counterfactual coherently reverses only the observed rollback outcome without duplicating or contradicting other evidence or leaking any gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"The release manager is evaluating Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\\\Software\\\\Northstar\\\\Notes be absent and version marker 6.1.4 be restored.\", \"The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\\\Software\\\\Northstar\\\\Notes absent and version marker 6.1.4 restored.\", \"The macOS installer rollback and the Linux installer rollback both completed successfully.\", \"No release-blocking defect remains open for RC3; documentation has a nonblocking typo in the keyboard-shortcuts table, which is recorded in the release record and still awaits an approved release-note action.\", \"Both Windows and macOS packages have valid signatures.\", \"The Northstar Notes deployment team is responsible for resolving any Windows MSI rollback gate failure, should one occur.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version marker 6.1.4 be restored."}, {"path": ["evidence", "2"], "text": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version marker 6.1.4 restored."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version marker 6.1.4 be restored.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version marker 6.1.4 be restored.", "negative_right": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes still present and version marker 6.2.0 remaining.", "right": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version marker 6.1.4 restored."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-014", "id": "fast-43-diverse-149-014-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "The release manager is evaluating Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version marker 6.1.4 be restored.", "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version marker 6.1.4 restored.", "The macOS installer rollback and the Linux installer rollback both completed successfully.", "No release-blocking defect remains open for RC3; documentation has a nonblocking typo in the keyboard-shortcuts table, which is recorded in the release record and still awaits an approved release-note action.", "Both Windows and macOS packages have valid signatures.", "The Northstar Notes deployment team is responsible for resolving any Windows MSI rollback gate failure, should one occur."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the original release policy in the unchanged question, use the same entity/time bindings, provide two complete factual evidence sentences (rollback spec and observed state), and the counterfactual coherently reverses only the observed rollback outcome without duplicating or contradicting other evidence or leaking any gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"The release manager is evaluating Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\\\Software\\\\Northstar\\\\Notes be absent and version marker 6.1.4 be restored.\", \"The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\\\Software\\\\Northstar\\\\Notes absent and version marker 6.1.4 restored.\", \"The macOS installer rollback and the Linux installer rollback both completed successfully.\", \"No release-blocking defect remains open for RC3; documentation has a nonblocking typo in the keyboard-shortcuts table, which is recorded in the release record and still awaits an approved release-note action.\", \"Both Windows and macOS packages have valid signatures.\", \"The Northstar Notes deployment team is responsible for resolving any Windows MSI rollback gate failure, should one occur.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version marker 6.1.4 be restored."}, {"path": ["evidence", "2"], "text": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version marker 6.1.4 restored."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version marker 6.1.4 be restored.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version marker 6.1.4 be restored.", "negative_right": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes still present and version marker 6.2.0 remaining.", "right": "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version marker 6.1.4 restored."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-014", "id": "fast-43-diverse-149-014-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "The release manager is evaluating Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester reports all tests pass on supported Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version marker 6.1.4 be restored.", "The observed post-rollback state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes still present and version marker 6.2.0 remaining.", "The macOS installer rollback and the Linux installer rollback both completed successfully.", "No release-blocking defect remains open for RC3; documentation has a nonblocking typo in the keyboard-shortcuts table, which is recorded in the release record and still awaits an approved release-note action.", "Both Windows and macOS packages have valid signatures.", "The Northstar Notes deployment team is responsible for resolving any Windows MSI rollback gate failure, should one occur."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy text and question bindings, the focus evidence pair are complete factual sentences describing spec requirement and observed state, the counterfactual coherently swaps the observed rollback outcome to a failing state without contradicting other evidence, and no gold answer or rule table is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"Release readiness case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"The compatibility tester confirms every supported-platform compatibility test for RC3 passed on Windows 11 23H2 and macOS 15.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\\\Software\\\\Northstar\\\\Notes be absent and version 6.1.4 be restored as the active installation.\", \"The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\\\Software\\\\Northstar\\\\Notes absent and version 6.1.4 restored as the active installation.\", \"All other RC3 installer rollbacks, including macOS, completed successfully.\", \"No release-blocking defect remains open for RC3.\", \"Both Windows and macOS packages carry valid signatures.\", \"Documentation records a nonblocking typo in the keyboard-shortcuts table, and it still awaits an approved release-note action.\", \"The Northstar Notes deployment team is responsible for resolving any Windows MSI rollback gate failure, should one occur.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version 6.1.4 be restored as the active installation."}, {"path": ["evidence", "2"], "text": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version 6.1.4 restored as the active installation."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version 6.1.4 be restored as the active installation.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version 6.1.4 be restored as the active installation.", "negative_right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes still present and version 6.0.9 restored as the active installation.", "right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version 6.1.4 restored as the active installation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-016", "id": "fast-43-diverse-149-016-base", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Release readiness case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester confirms every supported-platform compatibility test for RC3 passed on Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version 6.1.4 be restored as the active installation.", "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version 6.1.4 restored as the active installation.", "All other RC3 installer rollbacks, including macOS, completed successfully.", "No release-blocking defect remains open for RC3.", "Both Windows and macOS packages carry valid signatures.", "Documentation records a nonblocking typo in the keyboard-shortcuts table, and it still awaits an approved release-note action.", "The Northstar Notes deployment team is responsible for resolving any Windows MSI rollback gate failure, should one occur."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "software-05", "split": "train", "variant": "base"} {"domain": "software", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy text and question bindings, the focus evidence pair are complete factual sentences describing spec requirement and observed state, the counterfactual coherently swaps the observed rollback outcome to a failing state without contradicting other evidence, and no gold answer or rule table is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "full_context_fact_states": {"base": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "supported"}, "counterfactual": {"blockers": "supported", "compatibility": "supported", "other_rollbacks": "supported", "rollback_owner": "supported", "signatures": "supported", "typo_documented": "supported", "typo_nonblocking": "supported", "typo_pending_action": "supported", "windows_rollback_match": "refuted"}, "remove_left": {"windows_rollback_match": "unknown"}, "remove_right": {"windows_rollback_match": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"windows_rollback_match": "unknown"}, "negative_pair": {"windows_rollback_match": "refuted"}, "negative_sentence": {"windows_rollback_match": "unknown"}, "positive_pair": {"windows_rollback_match": "supported"}, "right": {"windows_rollback_match": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each express a single factual relationship, including permissible universal relationships over explicit package, platform, or rollback sets. The focus is a factual comparison of the observed Windows rollback state with its successful-rollback specification, not a policy classification. The base and counter assignments are jointly realizable with only that focus changing: a rollback gate can fail without an open release-blocking defect, while the remaining gates and documentation facts stay fixed. Policy evidence preserves the substantive release requirements originating in the original state; the more detailed decision criteria remain automatically available in the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A refuted windows_rollback_match establishes failure of the mandatory Windows MSI rollback gate. Under the ordered criteria, any mandatory gate failure is sufficient for Not ready, regardless of other gates. The supported rollback_owner atom also supplies the requested responsible team.", "rule_index": 0, "sound": true}, {"reason": "The conditions cover all mandatory gates: supported-platform compatibility, the Windows MSI rollback, all other installer rollbacks, absence of open blockers, and package signatures. They additionally establish a documented nonblocking issue with a still-pending approved release-note action, which is sufficient for Conditionally ready rather than Ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "compatibility", "statement": "Every supported-platform compatibility test for Northstar Notes 6.2 RC3 passed."}, {"id": "windows_rollback_match", "statement": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI matches its signed successful-rollback specification."}, {"id": "other_rollbacks", "statement": "Every Northstar Notes 6.2 RC3 installer rollback other than the Windows MSI rollback succeeded."}, {"id": "blockers", "statement": "No release-blocking defect for Northstar Notes 6.2 RC3 is open."}, {"id": "signatures", "statement": "Every Northstar Notes 6.2 RC3 release package has a valid signature."}, {"id": "typo_documented", "statement": "The Northstar Notes 6.2 RC3 release record documents the keyboard-shortcuts-table typo."}, {"id": "typo_nonblocking", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 is classified as nonblocking."}, {"id": "typo_pending_action", "statement": "The keyboard-shortcuts-table typo for Northstar Notes 6.2 RC3 still requires an approved release-note action."}, {"id": "rollback_owner", "statement": "The Northstar Notes deployment team is responsible for resolving a failure of the Windows MSI rollback gate."}], "base_state_json": "{\"context\": \"Release readiness case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.\", \"evidence\": [\"The compatibility tester confirms every supported-platform compatibility test for RC3 passed on Windows 11 23H2 and macOS 15.\", \"The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\\\Software\\\\Northstar\\\\Notes be absent and version 6.1.4 be restored as the active installation.\", \"The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\\\Software\\\\Northstar\\\\Notes absent and version 6.1.4 restored as the active installation.\", \"All other RC3 installer rollbacks, including macOS, completed successfully.\", \"No release-blocking defect remains open for RC3.\", \"Both Windows and macOS packages carry valid signatures.\", \"Documentation records a nonblocking typo in the keyboard-shortcuts table, and it still awaits an approved release-note action.\", \"The Northstar Notes deployment team is responsible for resolving any Windows MSI rollback gate failure, should one occur.\"], \"request\": \"Choose the readiness level for RC3 and identify the team that must resolve any gate failure.\"}", "base_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "counter_states": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}, {"atom_id": "rollback_owner", "state": "supported"}], "focus_atom": "windows_rollback_match", "focus_evidence": [{"path": ["evidence", "1"], "text": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version 6.1.4 be restored as the active installation."}, {"path": ["evidence", "2"], "text": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version 6.1.4 restored as the active installation."}], "policy_evidence": [{"path": ["context"], "text": "Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages."}, {"path": ["context"], "text": "Nonblocking documentation defects may ship with release notes."}, {"path": ["request"], "text": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}], "rules": [{"justification": "The Windows MSI does not satisfy the mandatory successful-rollback gate, so RC3 is Not ready; the deployment team is identified as responsible for resolving this gate failure.", "target": "0", "when": [{"atom_id": "windows_rollback_match", "state": "refuted"}, {"atom_id": "rollback_owner", "state": "supported"}]}, {"justification": "All mandatory gates pass, no release blocker is open, and the documented nonblocking typo still requires an approved release-note action, so RC3 is Conditionally ready.", "target": "1", "when": [{"atom_id": "compatibility", "state": "supported"}, {"atom_id": "windows_rollback_match", "state": "supported"}, {"atom_id": "other_rollbacks", "state": "supported"}, {"atom_id": "blockers", "state": "supported"}, {"atom_id": "signatures", "state": "supported"}, {"atom_id": "typo_documented", "state": "supported"}, {"atom_id": "typo_nonblocking", "state": "supported"}, {"atom_id": "typo_pending_action", "state": "supported"}]}]}, "verified_pair": {"left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version 6.1.4 be restored as the active installation.", "negative_left": "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version 6.1.4 be restored as the active installation.", "negative_right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes still present and version 6.0.9 restored as the active installation.", "right": "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes absent and version 6.1.4 restored as the active installation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-149-016", "id": "fast-43-diverse-149-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["Not ready — at least one mandatory release gate fails or a release-blocking defect remains open; route the blocking issue to the responsible team before release.", "Conditionally ready — every mandatory technical gate passes and no release blocker is open, but a documented nonblocking issue still requires an approved waiver or release-note action.", "Ready — every mandatory gate passes, no release-blocking defect is open, packages are signed, and any nonblocking documentation work is already resolved or approved for shipment."], "instructions": "Apply the stated release policy. Ignore facts that do not affect a release gate. Select exactly one ordered readiness level.", "type": "score"}}, "state": {"context": "Release readiness case note for Northstar Notes 6.2 RC3. Policy requires all supported-platform compatibility tests to pass, no open release-blocking defects, successful installer rollback, and signed packages. Nonblocking documentation defects may ship with release notes.", "evidence": ["The compatibility tester confirms every supported-platform compatibility test for RC3 passed on Windows 11 23H2 and macOS 15.", "The signed successful-rollback specification for the Northstar Notes 6.2 RC3 Windows MSI requires that after rollback, registry key HKLM\\Software\\Northstar\\Notes be absent and version 6.1.4 be restored as the active installation.", "The observed post-rollback registration state of the Northstar Notes 6.2 RC3 Windows MSI shows registry key HKLM\\Software\\Northstar\\Notes still present and version 6.0.9 restored as the active installation.", "All other RC3 installer rollbacks, including macOS, completed successfully.", "No release-blocking defect remains open for RC3.", "Both Windows and macOS packages carry valid signatures.", "Documentation records a nonblocking typo in the keyboard-shortcuts table, and it still awaits an approved release-note action.", "The Northstar Notes deployment team is responsible for resolving any Windows MSI rollback gate failure, should one occur."], "request": "Choose the readiness level for RC3 and identify the team that must resolve any gate failure."}}, "method": "c2d", "provenance": {"source_id": "diverse-149", "source_is_synthetic": true, "source_sha256": "a909d97eaf0ae2e96db96c109398942b2753d13bdf1cdec29e951391123449af", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "software-05", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical return-pickup scope and preparation-exception policy, keep the same address/zone/date bindings, and the counterfactual differs only in changing the curb time from 6:52 a.m. to 7:12 a.m., a single coherent factual edit that shifts the cart from timely to untimely without contradicting other stated facts; the focus evidence consists of two verbatim factual sentences (the deadline and the camera timestamp) rather than policy text, and no answer codes, rule tables, or output instructions appear in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relations or closely defined factual alternatives, and the focus atom is the factual timeliness relation. The base and counter assignments differ only on timeliness and are jointly realizable: an otherwise identical unemptied organics cart can have been timely or late. Policy evidence correctly preserves the substantive return-scope and preparation-exception rule originating in the original state; the remaining governing criteria and priority rules are already retained in the questions object and need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes organics rather than recycling, an unemptied timely and accessible cart, a lid gap over 5 cm, and the explicit absence of every listed urgent-hazard category. The preparation exception therefore applies and uniquely supports routine organics-desk handling with no return pickup.", "rule_index": 0, "sound": true}, {"reason": "Refuting the no-later-than-deadline atom entails a late set-out, placing the cart outside the return-pickup scope. Under the instruction to apply scope before the preparation exception, the over-5-cm gap does not independently establish outcome D. Recycling and all urgent hazards are excluded, so none of A–D applies and E is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported cart at 18 Lark Street was an organics cart for the Tuesday service event."}, {"id": "a2", "statement": "Organics collection was scheduled at 18 Lark Street for the Tuesday service event."}, {"id": "a3", "statement": "The evidence identified recycling as the scheduled or reported material at 18 Lark Street for the Tuesday service event."}, {"id": "a4", "statement": "The reported cart at 18 Lark Street remained unemptied after the Tuesday collection pass."}, {"id": "a5", "statement": "The reported cart's Tuesday curb set-out time was no later than the applicable 7:00 a.m. set-out deadline."}, {"id": "a6", "statement": "The reported cart at 18 Lark Street was accessible to the collection crew during the Tuesday collection pass."}, {"id": "a7", "statement": "The reported cart's lid gap during the Tuesday collection pass was greater than 5 cm."}, {"id": "a8", "statement": "Waste from the reported cart was spilled during the relevant Tuesday service period."}, {"id": "a9", "statement": "The reported cart or its waste blocked travel during the relevant Tuesday service period."}, {"id": "a10", "statement": "The reported cart leaked during the relevant Tuesday service period."}, {"id": "a11", "statement": "Pests were present at the reported cart during the relevant Tuesday service period."}, {"id": "a12", "statement": "The reported cart presented another stated immediate sanitation hazard during the relevant Tuesday service period."}], "base_state_json": "[{\"speaker\":\"Resident reporter\",\"text\":\"My green organics cart at 18 Lark Street was not emptied Tuesday.\"},{\"speaker\":\"Sanitation dispatcher\",\"text\":\"The address is in zone N4, where organics were scheduled Tuesday. The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m.\"},{\"speaker\":\"Resident reporter\",\"text\":\"A driveway camera timestamp shows the reported cart placed at the curb at 6:52 a.m. on Tuesday. The cart was visible and reachable from the street for the entire collection pass.\"},{\"speaker\":\"Resident reporter\",\"text\":\"A branch inside held the lid about 8 cm open, but nothing was spilled, nothing leaked, no travel was blocked, and no pests or other hazards were observed.\"},{\"speaker\":\"Collection crew supervisor\",\"text\":\"Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m."}, {"path": ["2", "text"], "text": "A driveway camera timestamp shows the reported cart placed at the curb at 6:52 a.m. on Tuesday."}], "policy_evidence": [{"path": ["3", "text"], "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}], "rules": [{"justification": "A timely, accessible, apparently missed organics cart is within return-pickup scope, but its lid gap over 5 cm triggers the preparation exception. With every stated urgent-hazard category explicitly absent and recycling explicitly excluded, the organics service desk handles it at routine priority with no return pickup.", "target": "D_organics_desk_routine_no_return", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "An explicitly late cart is outside the timely return-pickup scope, so neither organics return-pickup rubric applies and the scope-first instruction prevents use of the preparation exception as an independent routing basis. Recycling and every urgent-hazard category are explicitly excluded, leaving none of the four substantive rubrics uniquely satisfied.", "target": "E_none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m.", "negative_left": "The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m.", "negative_right": "A driveway camera timestamp shows the reported cart placed at the curb at 7:12 a.m. on Tuesday.", "right": "A driveway camera timestamp shows the reported cart placed at the curb at 6:52 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-152-010", "id": "fast-43-diverse-152-010-base", "input": {"questions": {"decision": {"criteria": {"A_organics_return_routine": "Route to the organics collection crew for a routine return pickup because the cart was timely, accessible, properly prepared within the 5 cm lid limit, and apparently missed.", "B_organics_return_urgent": "Route to the organics collection crew for an urgent return pickup because an eligible missed cart also presents a stated spill, obstruction, leakage, pest, or immediate sanitation hazard.", "C_recycling_desk_routine": "Route to the recycling service desk at routine priority because the evidence identifies recycling, rather than organics, as the scheduled or reported material.", "D_organics_desk_routine_no_return": "Route to the organics service desk at routine priority with no return pickup because a protruding branch or lid gap over 5 cm triggers the preparation exception and no urgent hazard is present.", "E_none_of_above": "Use only if the evidence does not uniquely satisfy any of the four routing, return, and priority rubrics above."}, "instructions": "Choose the routing, return-pickup decision, and priority that match the supplied policy. Apply the return-pickup scope first, then its preparation exception. Urgent priority requires spilled waste, blocked travel, leakage, pests, or another stated immediate sanitation hazard; otherwise use routine priority.", "type": "choice"}}, "state": [{"speaker": "Resident reporter", "text": "My green organics cart at 18 Lark Street was not emptied Tuesday."}, {"speaker": "Sanitation dispatcher", "text": "The address is in zone N4, where organics were scheduled Tuesday. The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m."}, {"speaker": "Resident reporter", "text": "A driveway camera timestamp shows the reported cart placed at the curb at 6:52 a.m. on Tuesday. The cart was visible and reachable from the street for the entire collection pass."}, {"speaker": "Resident reporter", "text": "A branch inside held the lid about 8 cm open, but nothing was spilled, nothing leaked, no travel was blocked, and no pests or other hazards were observed."}, {"speaker": "Collection crew supervisor", "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-152", "source_is_synthetic": true, "source_sha256": "1dfc7c00860982758ea45c9c5b47d8291aac858825f6143cf430090701b0ce5f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "D_organics_desk_routine_no_return"}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical return-pickup scope and preparation-exception policy, keep the same address/zone/date bindings, and the counterfactual differs only in changing the curb time from 6:52 a.m. to 7:12 a.m., a single coherent factual edit that shifts the cart from timely to untimely without contradicting other stated facts; the focus evidence consists of two verbatim factual sentences (the deadline and the camera timestamp) rather than policy text, and no answer codes, rule tables, or output instructions appear in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "refuted", "a12": "refuted", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "refuted"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms state individual factual relations or closely defined factual alternatives, and the focus atom is the factual timeliness relation. The base and counter assignments differ only on timeliness and are jointly realizable: an otherwise identical unemptied organics cart can have been timely or late. Policy evidence correctly preserves the substantive return-scope and preparation-exception rule originating in the original state; the remaining governing criteria and priority rules are already retained in the questions object and need not be repeated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes organics rather than recycling, an unemptied timely and accessible cart, a lid gap over 5 cm, and the explicit absence of every listed urgent-hazard category. The preparation exception therefore applies and uniquely supports routine organics-desk handling with no return pickup.", "rule_index": 0, "sound": true}, {"reason": "Refuting the no-later-than-deadline atom entails a late set-out, placing the cart outside the return-pickup scope. Under the instruction to apply scope before the preparation exception, the over-5-cm gap does not independently establish outcome D. Recycling and all urgent hazards are excluded, so none of A–D applies and E is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The reported cart at 18 Lark Street was an organics cart for the Tuesday service event."}, {"id": "a2", "statement": "Organics collection was scheduled at 18 Lark Street for the Tuesday service event."}, {"id": "a3", "statement": "The evidence identified recycling as the scheduled or reported material at 18 Lark Street for the Tuesday service event."}, {"id": "a4", "statement": "The reported cart at 18 Lark Street remained unemptied after the Tuesday collection pass."}, {"id": "a5", "statement": "The reported cart's Tuesday curb set-out time was no later than the applicable 7:00 a.m. set-out deadline."}, {"id": "a6", "statement": "The reported cart at 18 Lark Street was accessible to the collection crew during the Tuesday collection pass."}, {"id": "a7", "statement": "The reported cart's lid gap during the Tuesday collection pass was greater than 5 cm."}, {"id": "a8", "statement": "Waste from the reported cart was spilled during the relevant Tuesday service period."}, {"id": "a9", "statement": "The reported cart or its waste blocked travel during the relevant Tuesday service period."}, {"id": "a10", "statement": "The reported cart leaked during the relevant Tuesday service period."}, {"id": "a11", "statement": "Pests were present at the reported cart during the relevant Tuesday service period."}, {"id": "a12", "statement": "The reported cart presented another stated immediate sanitation hazard during the relevant Tuesday service period."}], "base_state_json": "[{\"speaker\":\"Resident reporter\",\"text\":\"My green organics cart at 18 Lark Street was not emptied Tuesday.\"},{\"speaker\":\"Sanitation dispatcher\",\"text\":\"The address is in zone N4, where organics were scheduled Tuesday. The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m.\"},{\"speaker\":\"Resident reporter\",\"text\":\"A driveway camera timestamp shows the reported cart placed at the curb at 6:52 a.m. on Tuesday. The cart was visible and reachable from the street for the entire collection pass.\"},{\"speaker\":\"Resident reporter\",\"text\":\"A branch inside held the lid about 8 cm open, but nothing was spilled, nothing leaked, no travel was blocked, and no pests or other hazards were observed.\"},{\"speaker\":\"Collection crew supervisor\",\"text\":\"Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}], "focus_atom": "a5", "focus_evidence": [{"path": ["1", "text"], "text": "The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m."}, {"path": ["2", "text"], "text": "A driveway camera timestamp shows the reported cart placed at the curb at 6:52 a.m. on Tuesday."}], "policy_evidence": [{"path": ["3", "text"], "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}], "rules": [{"justification": "A timely, accessible, apparently missed organics cart is within return-pickup scope, but its lid gap over 5 cm triggers the preparation exception. With every stated urgent-hazard category explicitly absent and recycling explicitly excluded, the organics service desk handles it at routine priority with no return pickup.", "target": "D_organics_desk_routine_no_return", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}, {"justification": "An explicitly late cart is outside the timely return-pickup scope, so neither organics return-pickup rubric applies and the scope-first instruction prevents use of the preparation exception as an independent routing basis. Recycling and every urgent-hazard category are explicitly excluded, leaving none of the four substantive rubrics uniquely satisfied.", "target": "E_none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "refuted"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "refuted"}, {"atom_id": "a12", "state": "refuted"}]}]}, "verified_pair": {"left": "The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m.", "negative_left": "The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m.", "negative_right": "A driveway camera timestamp shows the reported cart placed at the curb at 7:12 a.m. on Tuesday.", "right": "A driveway camera timestamp shows the reported cart placed at the curb at 6:52 a.m. on Tuesday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-152-010", "id": "fast-43-diverse-152-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"A_organics_return_routine": "Route to the organics collection crew for a routine return pickup because the cart was timely, accessible, properly prepared within the 5 cm lid limit, and apparently missed.", "B_organics_return_urgent": "Route to the organics collection crew for an urgent return pickup because an eligible missed cart also presents a stated spill, obstruction, leakage, pest, or immediate sanitation hazard.", "C_recycling_desk_routine": "Route to the recycling service desk at routine priority because the evidence identifies recycling, rather than organics, as the scheduled or reported material.", "D_organics_desk_routine_no_return": "Route to the organics service desk at routine priority with no return pickup because a protruding branch or lid gap over 5 cm triggers the preparation exception and no urgent hazard is present.", "E_none_of_above": "Use only if the evidence does not uniquely satisfy any of the four routing, return, and priority rubrics above."}, "instructions": "Choose the routing, return-pickup decision, and priority that match the supplied policy. Apply the return-pickup scope first, then its preparation exception. Urgent priority requires spilled waste, blocked travel, leakage, pests, or another stated immediate sanitation hazard; otherwise use routine priority.", "type": "choice"}}, "state": [{"speaker": "Resident reporter", "text": "My green organics cart at 18 Lark Street was not emptied Tuesday."}, {"speaker": "Sanitation dispatcher", "text": "The address is in zone N4, where organics were scheduled Tuesday. The applicable set-out deadline at 18 Lark Street for Tuesday service is 7:00 a.m."}, {"speaker": "Resident reporter", "text": "A driveway camera timestamp shows the reported cart placed at the curb at 7:12 a.m. on Tuesday. The cart was visible and reachable from the street for the entire collection pass."}, {"speaker": "Resident reporter", "text": "A branch inside held the lid about 8 cm open, but nothing was spilled, nothing leaked, no travel was blocked, and no pests or other hazards were observed."}, {"speaker": "Collection crew supervisor", "text": "Return-pickup scope covers timely, accessible organics carts with lids open no more than 5 cm. A protruding branch or larger lid gap is a preparation exception handled by the organics service desk."}]}, "method": "c2d", "provenance": {"source_id": "diverse-152", "source_is_synthetic": true, "source_sha256": "1dfc7c00860982758ea45c9c5b47d8291aac858825f6143cf430090701b0ce5f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "E_none_of_above"}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the zone, address, timing and material bindings from the original question and state without adding rules or answers; the counterfactual's swap to a recycling-cart label creates a coherent scheduled-material mismatch rather than a contradiction, and the two focus facts are complete factual sentences with no leaked verdict.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\": \"Dispatcher case note for Mara Chen's missed-pickup report at 18 Alder Court, filed 7:22 p.m. Tuesday.\", \"facts\": [\"18 Alder Court is confirmed within Zone C service boundaries.\", \"Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.\", \"The submitted placement image was captured at 5:58 a.m., before the 6:00 a.m. cutoff.\", \"The cart placement shown in that image was compliant at capture time: lid closed, curb-set, unobstructed.\", \"The submitted noncollection image was captured at 7:14 p.m., after the 7:00 p.m. threshold.\", \"The cart in the noncollection image is the same cart shown in the placement image, identifiable by matching curb position and cart markings.\", \"That cart remained uncollected and full at the noncollection image's capture time.\", \"Both submitted images depict a cart located at 18 Alder Court at their respective capture times.\", \"The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart.\"], \"open_question\": \"Should this report be routed directly to the Zone C organics crew for return pickup?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["facts", "8"], "text": "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart."}, {"path": ["facts", "1"], "text": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart.", "negative_left": "The cart depicted in the submitted images for Mara Chen's report is labeled as a recycling cart.", "negative_right": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "right": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, "verifier_independent_model": false}, "family": "fast-43-diverse-154-002", "id": "fast-43-diverse-154-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Dispatcher case note for Mara Chen's missed-pickup report at 18 Alder Court, filed 7:22 p.m. Tuesday.", "facts": ["18 Alder Court is confirmed within Zone C service boundaries.", "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "The submitted placement image was captured at 5:58 a.m., before the 6:00 a.m. cutoff.", "The cart placement shown in that image was compliant at capture time: lid closed, curb-set, unobstructed.", "The submitted noncollection image was captured at 7:14 p.m., after the 7:00 p.m. threshold.", "The cart in the noncollection image is the same cart shown in the placement image, identifiable by matching curb position and cart markings.", "That cart remained uncollected and full at the noncollection image's capture time.", "Both submitted images depict a cart located at 18 Alder Court at their respective capture times.", "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart."], "open_question": "Should this report be routed directly to the Zone C organics crew for return pickup?"}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the zone, address, timing and material bindings from the original question and state without adding rules or answers; the counterfactual's swap to a recycling-cart label creates a coherent scheduled-material mismatch rather than a contradiction, and the two focus facts are complete factual sentences with no leaked verdict.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\": \"Dispatcher case note for Mara Chen's missed-pickup report at 18 Alder Court, filed 7:22 p.m. Tuesday.\", \"facts\": [\"18 Alder Court is confirmed within Zone C service boundaries.\", \"Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.\", \"The submitted placement image was captured at 5:58 a.m., before the 6:00 a.m. cutoff.\", \"The cart placement shown in that image was compliant at capture time: lid closed, curb-set, unobstructed.\", \"The submitted noncollection image was captured at 7:14 p.m., after the 7:00 p.m. threshold.\", \"The cart in the noncollection image is the same cart shown in the placement image, identifiable by matching curb position and cart markings.\", \"That cart remained uncollected and full at the noncollection image's capture time.\", \"Both submitted images depict a cart located at 18 Alder Court at their respective capture times.\", \"The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart.\"], \"open_question\": \"Should this report be routed directly to the Zone C organics crew for return pickup?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["facts", "8"], "text": "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart."}, {"path": ["facts", "1"], "text": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart.", "negative_left": "The cart depicted in the submitted images for Mara Chen's report is labeled as a recycling cart.", "negative_right": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "right": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, "verifier_independent_model": false}, "family": "fast-43-diverse-154-002", "id": "fast-43-diverse-154-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Dispatcher case note for Mara Chen's missed-pickup report at 18 Alder Court, filed 7:22 p.m. Tuesday.", "facts": ["18 Alder Court is confirmed within Zone C service boundaries.", "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report.", "The submitted placement image was captured at 5:58 a.m., before the 6:00 a.m. cutoff.", "The cart placement shown in that image was compliant at capture time: lid closed, curb-set, unobstructed.", "The submitted noncollection image was captured at 7:14 p.m., after the 7:00 p.m. threshold.", "The cart in the noncollection image is the same cart shown in the placement image, identifiable by matching curb position and cart markings.", "That cart remained uncollected and full at the noncollection image's capture time.", "Both submitted images depict a cart located at 18 Alder Court at their respective capture times.", "The cart depicted in the submitted images for Mara Chen's report is labeled as a recycling cart."], "open_question": "Should this report be routed directly to the Zone C organics crew for return pickup?"}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original address, zone, timestamps, and schedule facts required by the unchanged question; the counterfactual only swaps the cart's lid label from organics to recycling, creating a plausible material mismatch without contradicting other stated measurements, and neither context includes rule tables, IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Resident reporter Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, in Zone C. A doorbell image timestamped 5:58 a.m. shows a cart at the curb, lid closed, with no visible contamination or obstruction. A second photo timestamped 7:14 p.m. shows the same cart, identifiable by its position and markings, still full at the curb and uncollected. Both images were taken at 18 Alder Court. The dispatcher must determine whether these facts, combined with the address's schedule, support routing the report directly to the Zone C organics crew for a return pickup or instead sending it to the service desk for review.\",\"evidence\":[\"18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected.\",\"The cart depicted in the submitted images for Mara Chen's report is labeled as organics on its lid.\",\"The 5:58 a.m. image shows compliant curb placement of the cart.\",\"The 7:14 p.m. image confirms the same cart remained uncollected past 7:00 p.m.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected."}, {"path": ["evidence", "1"], "text": "The cart depicted in the submitted images for Mara Chen's report is labeled as organics on its lid."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected.", "negative_left": "18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected.", "negative_right": "The cart depicted in the submitted images for Mara Chen's report is labeled as recycling on its lid.", "right": "The cart depicted in the submitted images for Mara Chen's report is labeled as organics on its lid."}, "verifier_independent_model": false}, "family": "fast-43-diverse-154-004", "id": "fast-43-diverse-154-004-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Resident reporter Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, in Zone C. A doorbell image timestamped 5:58 a.m. shows a cart at the curb, lid closed, with no visible contamination or obstruction. A second photo timestamped 7:14 p.m. shows the same cart, identifiable by its position and markings, still full at the curb and uncollected. Both images were taken at 18 Alder Court. The dispatcher must determine whether these facts, combined with the address's schedule, support routing the report directly to the Zone C organics crew for a return pickup or instead sending it to the service desk for review.", "evidence": ["18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected.", "The cart depicted in the submitted images for Mara Chen's report is labeled as organics on its lid.", "The 5:58 a.m. image shows compliant curb placement of the cart.", "The 7:14 p.m. image confirms the same cart remained uncollected past 7:00 p.m."]}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original address, zone, timestamps, and schedule facts required by the unchanged question; the counterfactual only swaps the cart's lid label from organics to recycling, creating a plausible material mismatch without contradicting other stated measurements, and neither context includes rule tables, IDs, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Resident reporter Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, in Zone C. A doorbell image timestamped 5:58 a.m. shows a cart at the curb, lid closed, with no visible contamination or obstruction. A second photo timestamped 7:14 p.m. shows the same cart, identifiable by its position and markings, still full at the curb and uncollected. Both images were taken at 18 Alder Court. The dispatcher must determine whether these facts, combined with the address's schedule, support routing the report directly to the Zone C organics crew for a return pickup or instead sending it to the service desk for review.\",\"evidence\":[\"18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected.\",\"The cart depicted in the submitted images for Mara Chen's report is labeled as organics on its lid.\",\"The 5:58 a.m. image shows compliant curb placement of the cart.\",\"The 7:14 p.m. image confirms the same cart remained uncollected past 7:00 p.m.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected."}, {"path": ["evidence", "1"], "text": "The cart depicted in the submitted images for Mara Chen's report is labeled as organics on its lid."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected.", "negative_left": "18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected.", "negative_right": "The cart depicted in the submitted images for Mara Chen's report is labeled as recycling on its lid.", "right": "The cart depicted in the submitted images for Mara Chen's report is labeled as organics on its lid."}, "verifier_independent_model": false}, "family": "fast-43-diverse-154-004", "id": "fast-43-diverse-154-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Resident reporter Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, in Zone C. A doorbell image timestamped 5:58 a.m. shows a cart at the curb, lid closed, with no visible contamination or obstruction. A second photo timestamped 7:14 p.m. shows the same cart, identifiable by its position and markings, still full at the curb and uncollected. Both images were taken at 18 Alder Court. The dispatcher must determine whether these facts, combined with the address's schedule, support routing the report directly to the Zone C organics crew for a return pickup or instead sending it to the service desk for review.", "evidence": ["18 Alder Court's schedule for the Tuesday of Mara Chen's report lists organics as the material to be collected.", "The cart depicted in the submitted images for Mara Chen's report is labeled as recycling on its lid.", "The 5:58 a.m. image shows compliant curb placement of the cart.", "The 7:14 p.m. image confirms the same cart remained uncollected past 7:00 p.m."]}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question and all governing policy via the original questions object; address, zone, time and path bindings are unchanged; the two focus evidence spans are complete factual sentences; the counterfactual coherently swaps the cart's material label to recycling without contradicting other stated facts; no gold answer, rule table or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Resident reporter Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, requesting a return collection. Dispatch review confirms the following facts: 18 Alder Court is in Zone C. The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection. A placement photo timestamped 5:58 a.m. shows the cart at the curb, lid closed, no contamination or obstruction visible, confirming compliant placement at that time. A second photo timestamped 7:14 p.m. shows the same cart, identifiable by matching curb position and markings, still full and uncollected. Both photos were confirmed to depict a cart at 18 Alder Court at their respective capture times. The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with an organics material designation.\",\"request\":\"Is this report ready to route to the Zone C organics crew for a return pickup?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["context"], "text": "The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection."}, {"path": ["context"], "text": "The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with an organics material designation."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection.", "negative_left": "The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection.", "negative_right": "The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with a recycling material designation.", "right": "The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with an organics material designation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-154-011", "id": "fast-43-diverse-154-011-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Resident reporter Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, requesting a return collection. Dispatch review confirms the following facts: 18 Alder Court is in Zone C. The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection. A placement photo timestamped 5:58 a.m. shows the cart at the curb, lid closed, no contamination or obstruction visible, confirming compliant placement at that time. A second photo timestamped 7:14 p.m. shows the same cart, identifiable by matching curb position and markings, still full and uncollected. Both photos were confirmed to depict a cart at 18 Alder Court at their respective capture times. The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with an organics material designation.", "request": "Is this report ready to route to the Zone C organics crew for a return pickup?"}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question and all governing policy via the original questions object; address, zone, time and path bindings are unchanged; the two focus evidence spans are complete factual sentences; the counterfactual coherently swaps the cart's material label to recycling without contradicting other stated facts; no gold answer, rule table or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Resident reporter Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, requesting a return collection. Dispatch review confirms the following facts: 18 Alder Court is in Zone C. The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection. A placement photo timestamped 5:58 a.m. shows the cart at the curb, lid closed, no contamination or obstruction visible, confirming compliant placement at that time. A second photo timestamped 7:14 p.m. shows the same cart, identifiable by matching curb position and markings, still full and uncollected. Both photos were confirmed to depict a cart at 18 Alder Court at their respective capture times. The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with an organics material designation.\",\"request\":\"Is this report ready to route to the Zone C organics crew for a return pickup?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["context"], "text": "The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection."}, {"path": ["context"], "text": "The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with an organics material designation."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection.", "negative_left": "The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection.", "negative_right": "The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with a recycling material designation.", "right": "The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with an organics material designation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-154-011", "id": "fast-43-diverse-154-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Resident reporter Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, requesting a return collection. Dispatch review confirms the following facts: 18 Alder Court is in Zone C. The Tuesday schedule for 18 Alder Court in Mara Chen's report lists organics as the material for collection. A placement photo timestamped 5:58 a.m. shows the cart at the curb, lid closed, no contamination or obstruction visible, confirming compliant placement at that time. A second photo timestamped 7:14 p.m. shows the same cart, identifiable by matching curb position and markings, still full and uncollected. Both photos were confirmed to depict a cart at 18 Alder Court at their respective capture times. The cart depicted in the submitted placement and noncollection images for Mara Chen's report is labeled with a recycling material designation.", "request": "Is this report ready to route to the Zone C organics crew for a return pickup?"}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same address, zone, timestamps, and compliance facts required by the question; the counterfactual swaps only the cart's labeled material from organics to recycling, creating a coherent single-fact deviation without contradicting other stated facts; evidence entries are plain factual sentences, not policy or rationale, and no gold answer or rule table is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Dispatch case note: Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, requesting a return collection. Records confirm 18 Alder Court sits in Zone C. A doorbell image timestamped 5:58 a.m. Tuesday shows the cart at the curb with the lid closed and no visible contamination or obstruction, establishing compliant placement before 6:00 a.m. A second photo taken at 7:14 p.m. shows the same cart still full at the curb, confirming it remained uncollected after 7:00 p.m. Both images were confirmed to depict the same cart at the same address at their respective capture times. No other zone or address is referenced in the submitted evidence.\",\"evidence\":[\"18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics.\",\"The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics."}, {"path": ["evidence", "1"], "text": "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics.", "negative_left": "18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics.", "negative_right": "The cart depicted in the submitted images for Mara Chen's report is labeled as a recycling cart.", "right": "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart."}, "verifier_independent_model": false}, "family": "fast-43-diverse-154-017", "id": "fast-43-diverse-154-017-base", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Dispatch case note: Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, requesting a return collection. Records confirm 18 Alder Court sits in Zone C. A doorbell image timestamped 5:58 a.m. Tuesday shows the cart at the curb with the lid closed and no visible contamination or obstruction, establishing compliant placement before 6:00 a.m. A second photo taken at 7:14 p.m. shows the same cart still full at the curb, confirming it remained uncollected after 7:00 p.m. Both images were confirmed to depict the same cart at the same address at their respective capture times. No other zone or address is referenced in the submitted evidence.", "evidence": ["18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics.", "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart."]}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same address, zone, timestamps, and compliance facts required by the question; the counterfactual swaps only the cart's labeled material from organics to recycling, creating a coherent single-fact deviation without contradicting other stated facts; evidence entries are plain factual sentences, not policy or rationale, and no gold answer or rule table is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "refuted", "a9": "supported"}, "remove_left": {"a8": "unknown"}, "remove_right": {"a8": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a8": "unknown"}, "negative_pair": {"a8": "refuted"}, "negative_sentence": {"a8": "unknown"}, "positive_pair": {"a8": "supported"}, "right": {"a8": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified image-location atom remains one relationship over an explicit set. The focus atom is factual. Base and counter assignments differ only on the material-match atom and are realizable under the policy. Empty policy_evidence is correct because the governing decision rules are entirely in the retained questions object; the original state adds case observations but no additional substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes Zone C and Tuesday organics identification, ties the submitted images to the address and scheduled material, and establishes compliant placement by 6:00 a.m. plus continued noncollection of the same cart after 7:00 p.m. These conditions are sufficient for a true decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes all timing, placement, address, identity, and noncollection conditions while expressly refuting that the depicted cart material matches the scheduled material. The question requires a no decision when evidence concerns another material.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "18 Alder Court is in Zone C."}, {"id": "a2", "statement": "Organics is the material scheduled at 18 Alder Court on the Tuesday of Mara Chen's report."}, {"id": "a3", "statement": "The submitted placement image for Mara Chen's report was captured no later than 6:00 a.m. on the Tuesday of her report."}, {"id": "a4", "statement": "The cart placement shown in the submitted placement image for Mara Chen's report was compliant at the image's capture time."}, {"id": "a5", "statement": "The submitted noncollection image for Mara Chen's report was captured after 7:00 p.m. on the Tuesday of her report."}, {"id": "a6", "statement": "The cart depicted in the submitted placement image for Mara Chen's report is the same cart depicted in the submitted noncollection image for her report."}, {"id": "a7", "statement": "The cart depicted in the submitted noncollection image for Mara Chen's report was uncollected at the image's capture time."}, {"id": "a8", "statement": "The material designation of the cart depicted in the submitted images for Mara Chen's report matches the material scheduled at 18 Alder Court on the Tuesday of her report."}, {"id": "a9", "statement": "Every submitted cart image for Mara Chen's report depicts a cart at 18 Alder Court at the image's capture time."}], "base_state_json": "{\"context\":\"Dispatch case note: Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, requesting a return collection. Records confirm 18 Alder Court sits in Zone C. A doorbell image timestamped 5:58 a.m. Tuesday shows the cart at the curb with the lid closed and no visible contamination or obstruction, establishing compliant placement before 6:00 a.m. A second photo taken at 7:14 p.m. shows the same cart still full at the curb, confirming it remained uncollected after 7:00 p.m. Both images were confirmed to depict the same cart at the same address at their respective capture times. No other zone or address is referenced in the submitted evidence.\",\"evidence\":[\"18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics.\",\"The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a8", "focus_evidence": [{"path": ["evidence", "0"], "text": "18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics."}, {"path": ["evidence", "1"], "text": "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart."}], "policy_evidence": [], "rules": [{"justification": "The evidence is tied to 18 Alder Court in Zone C and to its Tuesday scheduled organics material, and it establishes compliant placement by 6:00 a.m. and continued noncollection of the same cart after 7:00 p.m.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Even though the evidence is tied to 18 Alder Court and satisfies the placement, timing, identity, and noncollection requirements, the depicted cart's material does not match the material scheduled at that address; the instructions expressly require a no answer when the evidence concerns another material.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "refuted"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics.", "negative_left": "18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics.", "negative_right": "The cart depicted in the submitted images for Mara Chen's report is labeled as a recycling cart.", "right": "The cart depicted in the submitted images for Mara Chen's report is labeled as an organics cart."}, "verifier_independent_model": false}, "family": "fast-43-diverse-154-017", "id": "fast-43-diverse-154-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Do not route directly; send the report to the service desk for review or reject the return-pickup request.", "true": "Route directly to the Zone C organics crew for a return pickup."}, "instructions": "Answer yes when the address and scheduled material identify a crew, evidence shows compliant placement by 6:00 a.m., and the container remained uncollected after 7:00 p.m. Answer no if any requirement is missing, placement was noncompliant, or the evidence concerns another zone or material. Ignore notices and vehicle sightings that do not match the address zone and scheduled material.", "type": "noul"}}, "state": {"context": "Dispatch case note: Mara Chen filed a missed-pickup report at 7:22 p.m. Tuesday for 18 Alder Court, requesting a return collection. Records confirm 18 Alder Court sits in Zone C. A doorbell image timestamped 5:58 a.m. Tuesday shows the cart at the curb with the lid closed and no visible contamination or obstruction, establishing compliant placement before 6:00 a.m. A second photo taken at 7:14 p.m. shows the same cart still full at the curb, confirming it remained uncollected after 7:00 p.m. Both images were confirmed to depict the same cart at the same address at their respective capture times. No other zone or address is referenced in the submitted evidence.", "evidence": ["18 Alder Court's Tuesday scheduled material in Mara Chen's report is organics.", "The cart depicted in the submitted images for Mara Chen's report is labeled as a recycling cart."]}}, "method": "c2d", "provenance": {"source_id": "diverse-154", "source_is_synthetic": true, "source_sha256": "08afe468d75dd9dd572a82e281e20a33f4eec449a16d9053e8e30af9c08df423", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim and all governing policy criteria unchanged, with matching zone/address/time bindings; the two focus sentences are complete factual statements, not policy text; the counterfactual only shifts the report's submission day from Thursday to Friday, which is coherent with the rest of the unchanged context and does not create contradictions or embed any answer, rule table, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[\"18 Alder Mews is listed in the address-zone record as Zone C.\", \"Zone C has scheduled trash collection on Wednesday.\", \"Zone C's scheduled Wednesday trash container is a gray cart.\", \"The cart submitted for Wednesday collection at 18 Alder Mews was a gray cart.\", \"The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday.\", \"Route imagery and crew logs confirm the scheduled crew did not empty the submitted cart at 18 Alder Mews.\", \"The lid on the cart remains fully closed and the waste inside is contained.\", \"No odor, leakage, animals, pests, loose material, obstruction, or acute hazard has been observed at the location.\", \"The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4.\", \"The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Thursday, June 5.\"]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["8"], "text": "The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4."}, {"path": ["9"], "text": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Thursday, June 5."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4.", "negative_left": "The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Friday, June 6.", "right": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Thursday, June 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-155-001", "id": "fast-43-diverse-155-001-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": ["18 Alder Mews is listed in the address-zone record as Zone C.", "Zone C has scheduled trash collection on Wednesday.", "Zone C's scheduled Wednesday trash container is a gray cart.", "The cart submitted for Wednesday collection at 18 Alder Mews was a gray cart.", "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday.", "Route imagery and crew logs confirm the scheduled crew did not empty the submitted cart at 18 Alder Mews.", "The lid on the cart remains fully closed and the waste inside is contained.", "No odor, leakage, animals, pests, loose material, obstruction, or acute hazard has been observed at the location.", "The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4.", "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Thursday, June 5."]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim and all governing policy criteria unchanged, with matching zone/address/time bindings; the two focus sentences are complete factual statements, not policy text; the counterfactual only shifts the report's submission day from Thursday to Friday, which is coherent with the rest of the unchanged context and does not create contradictions or embed any answer, rule table, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[\"18 Alder Mews is listed in the address-zone record as Zone C.\", \"Zone C has scheduled trash collection on Wednesday.\", \"Zone C's scheduled Wednesday trash container is a gray cart.\", \"The cart submitted for Wednesday collection at 18 Alder Mews was a gray cart.\", \"The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday.\", \"Route imagery and crew logs confirm the scheduled crew did not empty the submitted cart at 18 Alder Mews.\", \"The lid on the cart remains fully closed and the waste inside is contained.\", \"No odor, leakage, animals, pests, loose material, obstruction, or acute hazard has been observed at the location.\", \"The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4.\", \"The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Thursday, June 5.\"]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["8"], "text": "The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4."}, {"path": ["9"], "text": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Thursday, June 5."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4.", "negative_left": "The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Friday, June 6.", "right": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Thursday, June 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-155-001", "id": "fast-43-diverse-155-001-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": ["18 Alder Mews is listed in the address-zone record as Zone C.", "Zone C has scheduled trash collection on Wednesday.", "Zone C's scheduled Wednesday trash container is a gray cart.", "The cart submitted for Wednesday collection at 18 Alder Mews was a gray cart.", "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday.", "Route imagery and crew logs confirm the scheduled crew did not empty the submitted cart at 18 Alder Mews.", "The lid on the cart remains fully closed and the waste inside is contained.", "No odor, leakage, animals, pests, loose material, obstruction, or acute hazard has been observed at the location.", "The Zone C Wednesday trash route servicing 18 Alder Mews closed at 3:15 p.m. on Wednesday, June 4.", "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. on Friday, June 6."]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy via the unchanged questions object and only vary the report timestamp (Thursday vs Friday) without altering zone, cart, or route facts, keeping the counterfactual coherent and free of any embedded answer or rule leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Resident reporter\",\"text\":\"18 Alder Mews is on the Zone C route, where Wednesday service uses gray carts. I set the closed gray cart at the curb before 7:00 a.m. Wednesday.\"},{\"speaker\":\"Sanitation dispatcher\",\"text\":\"The address-zone record confirms 18 Alder Mews as Zone C, and route imagery shows the compliant gray cart curbside on time. Supplied evidence establishes the scheduled crew did not empty the cart.\"},{\"speaker\":\"Collection crew supervisor\",\"text\":\"The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday. No exception was logged for the address.\"},{\"speaker\":\"Resident reporter\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Thursday. The lid remains shut, and there is no odor, leakage, animal activity, pests, loose material, obstruction, or dangerous spill; the waste is fully contained.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["2", "text"], "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday."}, {"path": ["3", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Thursday."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.", "negative_left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Friday.", "right": "The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Thursday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-155-005", "id": "fast-43-diverse-155-005-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Resident reporter", "text": "18 Alder Mews is on the Zone C route, where Wednesday service uses gray carts. I set the closed gray cart at the curb before 7:00 a.m. Wednesday."}, {"speaker": "Sanitation dispatcher", "text": "The address-zone record confirms 18 Alder Mews as Zone C, and route imagery shows the compliant gray cart curbside on time. Supplied evidence establishes the scheduled crew did not empty the cart."}, {"speaker": "Collection crew supervisor", "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday. No exception was logged for the address."}, {"speaker": "Resident reporter", "text": "The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Thursday. The lid remains shut, and there is no odor, leakage, animal activity, pests, loose material, obstruction, or dangerous spill; the waste is fully contained."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy via the unchanged questions object and only vary the report timestamp (Thursday vs Friday) without altering zone, cart, or route facts, keeping the counterfactual coherent and free of any embedded answer or rule leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\":\"Resident reporter\",\"text\":\"18 Alder Mews is on the Zone C route, where Wednesday service uses gray carts. I set the closed gray cart at the curb before 7:00 a.m. Wednesday.\"},{\"speaker\":\"Sanitation dispatcher\",\"text\":\"The address-zone record confirms 18 Alder Mews as Zone C, and route imagery shows the compliant gray cart curbside on time. Supplied evidence establishes the scheduled crew did not empty the cart.\"},{\"speaker\":\"Collection crew supervisor\",\"text\":\"The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday. No exception was logged for the address.\"},{\"speaker\":\"Resident reporter\",\"text\":\"The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Thursday. The lid remains shut, and there is no odor, leakage, animal activity, pests, loose material, obstruction, or dangerous spill; the waste is fully contained.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["2", "text"], "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday."}, {"path": ["3", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Thursday."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.", "negative_left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Friday.", "right": "The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Thursday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-155-005", "id": "fast-43-diverse-155-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Resident reporter", "text": "18 Alder Mews is on the Zone C route, where Wednesday service uses gray carts. I set the closed gray cart at the curb before 7:00 a.m. Wednesday."}, {"speaker": "Sanitation dispatcher", "text": "The address-zone record confirms 18 Alder Mews as Zone C, and route imagery shows the compliant gray cart curbside on time. Supplied evidence establishes the scheduled crew did not empty the cart."}, {"speaker": "Collection crew supervisor", "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday. No exception was logged for the address."}, {"speaker": "Resident reporter", "text": "The missed-collection report for 18 Alder Mews was submitted at 9:15 a.m. on Friday. The lid remains shut, and there is no odor, leakage, animal activity, pests, loose material, obstruction, or dangerous spill; the waste is fully contained."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original Zone C/Wednesday/gray-cart bindings and retain the unchanged questions object satisfying the policy contract; the two focus evidence sentences are plain factual statements about route closure and report timing, not policy text; the counterfactual only shifts the report day to Saturday, remaining internally consistent with all other facts; and neither context contains any rule table, label, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\": \"Resident reporter\", \"text\": \"18 Alder Mews is in Zone C, scheduled for Wednesday trash collection with gray carts. I set out the closed gray cart curbside by 6:55 a.m. Wednesday, matching the scheduled material.\"}, {\"speaker\": \"Sanitation dispatcher\", \"text\": \"Photographic evidence and the crew log confirm the scheduled crew did not empty the cart. The cart remains fully contained: no exposed waste, leakage, odor, animals, pests, loose material, obstruction, or dangerous spill was observed.\"}, {\"speaker\": \"Collection crew supervisor\", \"text\": \"The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.\"}, {\"speaker\": \"Resident reporter\", \"text\": \"The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Thursday.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["2", "text"], "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday."}, {"path": ["3", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Thursday."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.", "negative_left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Saturday.", "right": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Thursday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-155-011", "id": "fast-43-diverse-155-011-base", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Resident reporter", "text": "18 Alder Mews is in Zone C, scheduled for Wednesday trash collection with gray carts. I set out the closed gray cart curbside by 6:55 a.m. Wednesday, matching the scheduled material."}, {"speaker": "Sanitation dispatcher", "text": "Photographic evidence and the crew log confirm the scheduled crew did not empty the cart. The cart remains fully contained: no exposed waste, leakage, odor, animals, pests, loose material, obstruction, or dangerous spill was observed."}, {"speaker": "Collection crew supervisor", "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday."}, {"speaker": "Resident reporter", "text": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Thursday."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original Zone C/Wednesday/gray-cart bindings and retain the unchanged questions object satisfying the policy contract; the two focus evidence sentences are plain factual statements about route closure and report timing, not policy text; the counterfactual only shifts the report day to Saturday, remaining internally consistent with all other facts; and neither context contains any rule table, label, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "counterfactual": {"A1": "refuted", "A10": "refuted", "A11": "refuted", "A12": "refuted", "A13": "refuted", "A14": "refuted", "A15": "refuted", "A16": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each describe a factual relationship or condition rather than a final classification. A1 is a factual timing relation, and changing only whether that elapsed time exceeds 24 hours is realizable while holding the other facts fixed. Empty policy_evidence is correct because the governing cutoff, qualification requirements, escalation criteria, priorities, and outcome definitions are all retained automatically in original_input.questions; the state adds no separate substantive policy needed to interpret them.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes a timely report, correct address-zone/day/container, placement by the cutoff, evidence of a genuine miss, contained waste, and explicit absence of every listed sanitation escalation or urgent-hazard condition. It is sufficient for routine return pickup (level 1).", "rule_index": 0, "sound": true}, {"reason": "Refutation of A1 entails that the report exceeded the 24-hour window. The remaining conditions preserve an otherwise compliant miss while expressly excluding sanitation and urgent-hazard outcomes, so the late-report rule is sufficient for no return pickup/service-desk routing (level 0).", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The elapsed time from closure of the relevant Zone C Wednesday trash route to submission of the missed-collection report for 18 Alder Mews was no more than 24 hours."}, {"id": "A2", "statement": "The address-zone record assigns 18 Alder Mews to Zone C."}, {"id": "A3", "statement": "Zone C has scheduled trash collection on Wednesday."}, {"id": "A4", "statement": "Zone C's scheduled Wednesday trash container is a gray cart."}, {"id": "A5", "statement": "The container submitted for Wednesday collection at 18 Alder Mews was a gray cart."}, {"id": "A6", "statement": "The submitted cart at 18 Alder Mews was curbside by 7:00 a.m. Wednesday."}, {"id": "A7", "statement": "The supplied evidence shows that the scheduled crew did not empty the submitted cart at 18 Alder Mews."}, {"id": "A8", "statement": "The waste remaining at 18 Alder Mews is fully contained."}, {"id": "A9", "statement": "The waste remaining at 18 Alder Mews is exposed."}, {"id": "A10", "statement": "The waste or cart at 18 Alder Mews is leaking."}, {"id": "A11", "statement": "The waste or cart at 18 Alder Mews produces an odor."}, {"id": "A12", "statement": "Animals are accessing or disturbing the waste at 18 Alder Mews."}, {"id": "A13", "statement": "Pests are present at the waste at 18 Alder Mews."}, {"id": "A14", "statement": "Loose waste material is present outside the cart at 18 Alder Mews."}, {"id": "A15", "statement": "The waste or cart at 18 Alder Mews creates an obstruction."}, {"id": "A16", "statement": "The waste or cart at 18 Alder Mews creates an acute dangerous spill."}], "base_state_json": "[{\"speaker\": \"Resident reporter\", \"text\": \"18 Alder Mews is in Zone C, scheduled for Wednesday trash collection with gray carts. I set out the closed gray cart curbside by 6:55 a.m. Wednesday, matching the scheduled material.\"}, {\"speaker\": \"Sanitation dispatcher\", \"text\": \"Photographic evidence and the crew log confirm the scheduled crew did not empty the cart. The cart remains fully contained: no exposed waste, leakage, odor, animals, pests, loose material, obstruction, or dangerous spill was observed.\"}, {\"speaker\": \"Collection crew supervisor\", \"text\": \"The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.\"}, {\"speaker\": \"Resident reporter\", \"text\": \"The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Thursday.\"}]", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}], "focus_atom": "A1", "focus_evidence": [{"path": ["2", "text"], "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday."}, {"path": ["3", "text"], "text": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Thursday."}], "policy_evidence": [], "rules": [{"justification": "The report is within 24 hours, the address, collection day, cart type, and placement time are compliant, supplied evidence establishes a miss, and the remaining waste is contained without sanitation or hazard conditions; this is a routine return pickup.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}, {"justification": "The report exceeds the 24-hour qualification window, while the unchanged facts exclude sanitation and urgent-hazard outcomes; the late-report criterion requires no return pickup and service-desk routing.", "target": "0", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}, {"atom_id": "A12", "state": "refuted"}, {"atom_id": "A13", "state": "refuted"}, {"atom_id": "A14", "state": "refuted"}, {"atom_id": "A15", "state": "refuted"}, {"atom_id": "A16", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.", "negative_left": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday.", "negative_right": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Saturday.", "right": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Thursday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-155-011", "id": "fast-43-diverse-155-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No return pickup: route to the service desk for correction or explanation because the report is late, the wrong material/day or address zone is involved, the cart was placed after 7:00 a.m., or evidence shows service occurred.", "1 — Routine return pickup: route to the regular collection crew on its next available return run when a timely, compliant scheduled cart was missed and waste remains contained with no obstruction, leakage, odor, pests, or loose material.", "2 — Expedited sanitation return: route for same-day priority collection when a valid miss also involves exposed waste, leakage, strong odor, animals, pests, or loose material, but does not create an immediate access or traffic hazard.", "3 — Urgent hazard response: immediately dispatch a supervisor or emergency sanitation crew when waste or containers actively block a roadway, sidewalk access, fire route, or create an acute dangerous spill."], "instructions": "Classify the required routing/intervention level. Policy: carts must be curbside by 7:00 a.m.; reports within 24 hours qualify if address-zone data, scheduled material, and supplied evidence support that a compliant cart was missed. Treat an accurate paraphrase as evidence when its details are supported by the original statement and records. A qualifying miss warrants a return pickup. Choose exactly one ordered level.", "type": "score"}}, "state": [{"speaker": "Resident reporter", "text": "18 Alder Mews is in Zone C, scheduled for Wednesday trash collection with gray carts. I set out the closed gray cart curbside by 6:55 a.m. Wednesday, matching the scheduled material."}, {"speaker": "Sanitation dispatcher", "text": "Photographic evidence and the crew log confirm the scheduled crew did not empty the cart. The cart remains fully contained: no exposed waste, leakage, odor, animals, pests, loose material, obstruction, or dangerous spill was observed."}, {"speaker": "Collection crew supervisor", "text": "The Zone C Wednesday trash route serving 18 Alder Mews closed at 2:15 p.m. on Wednesday."}, {"speaker": "Resident reporter", "text": "The missed-collection report for 18 Alder Mews was submitted at 9:00 a.m. the following Saturday."}]}, "method": "c2d", "provenance": {"source_id": "diverse-155", "source_is_synthetic": true, "source_sha256": "f39d8eb452caf19fc5277711789c83ec28f817d816b8aa050b46b52102e2fdb2", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same address, zone, and question object while only altering the crew-completion timestamp, and the counterfactual's crew finishing at 3:30 p.m. after the 2:10 p.m. photo is a coherent, non-contradictory scenario without embedded rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday. The sanitation dispatcher's zone record places the address within municipal collection service in Zone C, where Tuesday is scheduled trash day. Maya selected trash and attached a photo showing the city-issued gray cart upright at the curb, lid closed and unemptied, with no vehicles or other obstructions nearby. Its collection tag is still attached, and GPS metadata matches the address and report. The Zone C crew supervisor confirms the assigned crew was responsible for Alder Lane that day and that the cart was accessible when the crew finished the route. The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m. Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday. The photo remains the only documentation of whether the cart was emptied. No spill, odor, pests, blocked roadway, or pedestrian hazard is reported.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m."}, {"path": [], "text": "Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m.", "negative_left": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 3:30 p.m.", "negative_right": "Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday.", "right": "Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-156-004", "id": "fast-43-diverse-156-004-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday. The sanitation dispatcher's zone record places the address within municipal collection service in Zone C, where Tuesday is scheduled trash day. Maya selected trash and attached a photo showing the city-issued gray cart upright at the curb, lid closed and unemptied, with no vehicles or other obstructions nearby. Its collection tag is still attached, and GPS metadata matches the address and report. The Zone C crew supervisor confirms the assigned crew was responsible for Alder Lane that day and that the cart was accessible when the crew finished the route. The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m. Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday. The photo remains the only documentation of whether the cart was emptied. No spill, odor, pests, blocked roadway, or pedestrian hazard is reported."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same address, zone, and question object while only altering the crew-completion timestamp, and the counterfactual's crew finishing at 3:30 p.m. after the 2:10 p.m. photo is a coherent, non-contradictory scenario without embedded rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday. The sanitation dispatcher's zone record places the address within municipal collection service in Zone C, where Tuesday is scheduled trash day. Maya selected trash and attached a photo showing the city-issued gray cart upright at the curb, lid closed and unemptied, with no vehicles or other obstructions nearby. Its collection tag is still attached, and GPS metadata matches the address and report. The Zone C crew supervisor confirms the assigned crew was responsible for Alder Lane that day and that the cart was accessible when the crew finished the route. The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m. Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday. The photo remains the only documentation of whether the cart was emptied. No spill, odor, pests, blocked roadway, or pedestrian hazard is reported.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m."}, {"path": [], "text": "Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m.", "negative_left": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 3:30 p.m.", "negative_right": "Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday.", "right": "Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-156-004", "id": "fast-43-diverse-156-004-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday. The sanitation dispatcher's zone record places the address within municipal collection service in Zone C, where Tuesday is scheduled trash day. Maya selected trash and attached a photo showing the city-issued gray cart upright at the curb, lid closed and unemptied, with no vehicles or other obstructions nearby. Its collection tag is still attached, and GPS metadata matches the address and report. The Zone C crew supervisor confirms the assigned crew was responsible for Alder Lane that day and that the cart was accessible when the crew finished the route. The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 3:30 p.m. Maya Chen's attached photo was captured at 2:10 p.m. that Tuesday. The photo remains the only documentation of whether the cart was emptied. No spill, odor, pests, blocked roadway, or pedestrian hazard is reported."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric via the original questions object and preserve address/zone/time bindings; the counterfactual only shifts the crew completion time to 3:30 p.m., which stays logically consistent with the unchanged 2:10 p.m. photo capture and unemptied-cart evidence without contradicting other facts; the two focus sentences are complete factual statements, not policy text; neither context reveals a label, code, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday, requesting trash collection. The sanitation dispatcher's zone record places the address in Zone C, where Tuesday is trash day and the Zone C crew was assigned to service the address. Maya attached a photo showing the city-issued gray cart at the curb, lid closed, tag attached, unemptied, with no vehicles or other obstructions nearby, matching GPS and address metadata; this photo is the report's only evidence of whether the cart was emptied. The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m., and the cart remained accessible to the crew at that time. Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday. No spill, odor, pests, blocked roadway, or pedestrian hazard is reported.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m."}, {"path": [], "text": "Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m.", "negative_left": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 3:30 p.m.", "negative_right": "Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday.", "right": "Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-156-007", "id": "fast-43-diverse-156-007-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday, requesting trash collection. The sanitation dispatcher's zone record places the address in Zone C, where Tuesday is trash day and the Zone C crew was assigned to service the address. Maya attached a photo showing the city-issued gray cart at the curb, lid closed, tag attached, unemptied, with no vehicles or other obstructions nearby, matching GPS and address metadata; this photo is the report's only evidence of whether the cart was emptied. The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m., and the cart remained accessible to the crew at that time. Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday. No spill, odor, pests, blocked roadway, or pedestrian hazard is reported."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric via the original questions object and preserve address/zone/time bindings; the counterfactual only shifts the crew completion time to 3:30 p.m., which stays logically consistent with the unchanged 2:10 p.m. photo capture and unemptied-cart evidence without contradicting other facts; the two focus sentences are complete factual statements, not policy text; neither context reveals a label, code, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday, requesting trash collection. The sanitation dispatcher's zone record places the address in Zone C, where Tuesday is trash day and the Zone C crew was assigned to service the address. Maya attached a photo showing the city-issued gray cart at the curb, lid closed, tag attached, unemptied, with no vehicles or other obstructions nearby, matching GPS and address metadata; this photo is the report's only evidence of whether the cart was emptied. The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m., and the cart remained accessible to the crew at that time. Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday. No spill, odor, pests, blocked roadway, or pedestrian hazard is reported.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m."}, {"path": [], "text": "Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 11:45 a.m.", "negative_left": "The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 3:30 p.m.", "negative_right": "Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday.", "right": "Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-156-007", "id": "fast-43-diverse-156-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday, requesting trash collection. The sanitation dispatcher's zone record places the address in Zone C, where Tuesday is trash day and the Zone C crew was assigned to service the address. Maya attached a photo showing the city-issued gray cart at the curb, lid closed, tag attached, unemptied, with no vehicles or other obstructions nearby, matching GPS and address metadata; this photo is the report's only evidence of whether the cart was emptied. The Zone C trash crew completed service at 18 Alder Lane on Tuesday at 3:30 p.m., and the cart remained accessible to the crew at that time. Maya Chen's attached photo shows a recorded capture time of 2:10 p.m. that same Tuesday. No spill, odor, pests, blocked roadway, or pedestrian hazard is reported."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rubric via the unchanged questions object and only vary observed timestamps, keeping zone, path, and time bindings intact; the two focus sentences are plain factual statements; shifting the photo capture time to 9:45 a.m. (before the 11:05 a.m. crew completion) is internally consistent and does not contradict other stated facts; neither context states or implies the classification level or rubric code.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday. Zone C dispatch records confirm Tuesday is trash day for this address, and the report requests trash collection specifically. Maya's attached photo shows the city-issued gray cart, still tagged, sitting curbside with lid closed and no obstructions; the photo is the only documentation in the report bearing on whether the cart was emptied. The photo shows the cart unemptied. GPS metadata confirms the photo was taken at 18 Alder Lane. The Zone C crew supervisor confirms the cart was accessible to the crew when it completed the Alder Lane route. The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday. Maya Chen's attached photo has a recorded capture time of 2:10 p.m. on the report Tuesday. No spill, odor, pests, blocked roadway, or pedestrian hazard is documented anywhere in the report.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday."}, {"path": [], "text": "Maya Chen's attached photo has a recorded capture time of 2:10 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday.", "negative_left": "The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday.", "negative_right": "Maya Chen's attached photo has a recorded capture time of 9:45 a.m. on the report Tuesday.", "right": "Maya Chen's attached photo has a recorded capture time of 2:10 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-156-010", "id": "fast-43-diverse-156-010-base", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday. Zone C dispatch records confirm Tuesday is trash day for this address, and the report requests trash collection specifically. Maya's attached photo shows the city-issued gray cart, still tagged, sitting curbside with lid closed and no obstructions; the photo is the only documentation in the report bearing on whether the cart was emptied. The photo shows the cart unemptied. GPS metadata confirms the photo was taken at 18 Alder Lane. The Zone C crew supervisor confirms the cart was accessible to the crew when it completed the Alder Lane route. The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday. Maya Chen's attached photo has a recorded capture time of 2:10 p.m. on the report Tuesday. No spill, odor, pests, blocked roadway, or pedestrian hazard is documented anywhere in the report."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "public_services-01", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rubric via the unchanged questions object and only vary observed timestamps, keeping zone, path, and time bindings intact; the two focus sentences are plain factual statements; shifting the photo capture time to 9:45 a.m. (before the 11:05 a.m. crew completion) is internally consistent and does not contradict other stated facts; neither context states or implies the classification level or rubric code.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "refuted", "A11": "refuted", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A10": "unknown"}, "remove_right": {"A10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A10": "unknown"}, "negative_pair": {"A10": "refuted"}, "negative_sentence": {"A10": "unknown"}, "positive_pair": {"A10": "supported"}, "right": {"A10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express factual relationships rather than final policy classifications; the focus, A10, is the factual ordering of the photo and crew-completion times. The base and counter assignments can both occur with only that ordering changing: an unemptied-cart photo may be captured either after or before crew completion while the other facts remain fixed. Empty policy_evidence is correct because the governing rubric, ordering, evidence requirements, and exceptions are all retained in the questions object; the original state adds case observations but no separate substantive policy needed to interpret the question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes scheduled, eligible municipal trash service and an accessible reported cart, but the sole evidence that it was unemptied does not postdate crew completion. Thus a post-completion miss cannot be verified. A11 is refuted, excluding level 3, while the missing temporal verification prevents level 2 and supports service-desk review under level 1.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes scheduled municipal trash service, the reported container’s accessibility, assigned-crew completion, and sole status evidence recorded afterward showing the cart still unemptied. A11 is refuted, excluding an urgent hazard, so the conditions sufficiently support a routine return pickup at level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "18 Alder Lane was within municipal collection service on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A2", "statement": "Municipal trash collection was scheduled at 18 Alder Lane on the Tuesday when Maya Chen filed the 2:10 p.m. report."}, {"id": "A3", "statement": "Maya Chen's 2:10 p.m. Tuesday report requested collection of trash."}, {"id": "A4", "statement": "The gray cart depicted in Maya Chen's attached photo was the container identified in her report."}, {"id": "A5", "statement": "The reported gray cart was accessible to the assigned collection crew when that crew completed service at 18 Alder Lane on the report Tuesday."}, {"id": "A6", "statement": "The Zone C trash crew was assigned to service 18 Alder Lane on the report Tuesday."}, {"id": "A7", "statement": "The assigned Zone C trash crew completed service at 18 Alder Lane at its recorded completion time on the report Tuesday."}, {"id": "A8", "statement": "Maya Chen's attached photo depicts the reported gray cart as unemptied."}, {"id": "A9", "statement": "Maya Chen's attached photo is the report's only evidence of whether the reported gray cart was emptied."}, {"id": "A10", "statement": "The recorded capture time of Maya Chen's attached photo is later than the assigned Zone C trash crew's recorded completion time at 18 Alder Lane on the report Tuesday."}, {"id": "A11", "statement": "The report documents an immediate sanitation-or-obstruction hazard caused by the missed collection at 18 Alder Lane."}], "base_state_json": "\"Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday. Zone C dispatch records confirm Tuesday is trash day for this address, and the report requests trash collection specifically. Maya's attached photo shows the city-issued gray cart, still tagged, sitting curbside with lid closed and no obstructions; the photo is the only documentation in the report bearing on whether the cart was emptied. The photo shows the cart unemptied. GPS metadata confirms the photo was taken at 18 Alder Lane. The Zone C crew supervisor confirms the cart was accessible to the crew when it completed the Alder Lane route. The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday. Maya Chen's attached photo has a recorded capture time of 2:10 p.m. on the report Tuesday. No spill, odor, pests, blocked roadway, or pedestrian hazard is documented anywhere in the report.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}], "focus_atom": "A10", "focus_evidence": [{"path": [], "text": "The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday."}, {"path": [], "text": "Maya Chen's attached photo has a recorded capture time of 2:10 p.m. on the report Tuesday."}], "policy_evidence": [], "rules": [{"justification": "The report concerns scheduled municipal trash and identifies an accessible container, but its sole collection-status evidence does not postdate the assigned crew's completion. Consequently, the evidence cannot verify that the container remained uncollected after completion. No immediate sanitation-or-obstruction hazard is documented.", "target": "1", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "refuted"}, {"atom_id": "A11", "state": "refuted"}]}, {"justification": "The evidence verifies that the scheduled municipal trash container was accessible when the assigned crew completed the location and was still unemptied in the sole collection-status evidence recorded afterward. No immediate sanitation-or-obstruction hazard is documented.", "target": "2", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}, {"atom_id": "A11", "state": "refuted"}]}]}, "verified_pair": {"left": "The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday.", "negative_left": "The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday.", "negative_right": "Maya Chen's attached photo has a recorded capture time of 9:45 a.m. on the report Tuesday.", "right": "Maya Chen's attached photo has a recorded capture time of 2:10 p.m. on the report Tuesday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-156-010", "id": "fast-43-diverse-156-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — No collection action: Evidence shows the reported material was not scheduled that day, was already collected, or is outside municipal service; close the report or provide schedule information.", "1 — Service-desk review: Collection eligibility or correct routing cannot be verified because key evidence is missing or conflicting, such as pickup type, zone, placement, timestamp, or crew completion status.", "2 — Standard return pickup: Evidence verifies that the scheduled trash, recycling, or organics container was accessible and remained uncollected after the assigned crew completed the location, with no urgent hazard; route to the matching collection crew for a routine return.", "3 — Urgent field response: The missed collection also creates a documented immediate sanitation or obstruction concern, such as spilled waste, active pests, strong leakage, or a container blocking a road or accessible walkway; dispatch expedited assistance."], "instructions": "Classify the report by the appropriate service response. Apply the ordered rubric and select the single level best supported by the explicit address-zone, schedule, placement, timestamp, photo, and crew-completion evidence.", "type": "score"}}, "state": "Resident reporter Maya Chen submitted a missed-pickup report for 18 Alder Lane at 2:10 p.m. Tuesday. Zone C dispatch records confirm Tuesday is trash day for this address, and the report requests trash collection specifically. Maya's attached photo shows the city-issued gray cart, still tagged, sitting curbside with lid closed and no obstructions; the photo is the only documentation in the report bearing on whether the cart was emptied. The photo shows the cart unemptied. GPS metadata confirms the photo was taken at 18 Alder Lane. The Zone C crew supervisor confirms the cart was accessible to the crew when it completed the Alder Lane route. The Zone C crew completed service at 18 Alder Lane at 11:05 a.m. on the report Tuesday. Maya Chen's attached photo has a recorded capture time of 9:45 a.m. on the report Tuesday. No spill, odor, pests, blocked roadway, or pedestrian hazard is documented anywhere in the report."}, "method": "c2d", "provenance": {"source_id": "diverse-156", "source_is_synthetic": true, "source_sha256": "d5e53cc326ea33c5df7ad6c71fd14d3443d9bfc17f6e9e44b726359770b41de4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "public_services-01", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the capacity policy and form requirement while only altering attendance figures, which stays consistent with the vague attendant remark, and no evidence sentence reveals the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Event organizer\",\"text\":\"I’m requesting the teaching kitchen on October 24, 2026, from 5:30–8:30 p.m. for a dumpling workshop. We need all four induction burners and no other equipment.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The teaching kitchen's maximum attendance capacity is 24 people. Kitchen reservations are ready only when the date, time, attendance, equipment needs, and signed Food Preparation Form are provided.\"},{\"speaker\":\"Event organizer\",\"text\":\"I uploaded the signed sheet confirming we’ll follow the center’s food-handling and cleanup rules—the required paperwork for preparing food onsite. The organizer expects 20 people to attend the October 24, 2026, dumpling workshop.\"},{\"speaker\":\"Community center attendant\",\"text\":\"October 24 is available from 5:30–8:30 p.m., and all four burners are operational. No extra attendant is needed at this attendance level.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The teaching kitchen's maximum attendance capacity is 24 people."}, {"path": ["2", "text"], "text": "The organizer expects 20 people to attend the October 24, 2026, dumpling workshop."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The teaching kitchen's maximum attendance capacity is 24 people.", "negative_left": "The teaching kitchen's maximum attendance capacity is 24 people.", "negative_right": "The organizer expects 30 people to attend the October 24, 2026, dumpling workshop.", "right": "The organizer expects 20 people to attend the October 24, 2026, dumpling workshop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-160-005", "id": "fast-43-diverse-160-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Event organizer", "text": "I’m requesting the teaching kitchen on October 24, 2026, from 5:30–8:30 p.m. for a dumpling workshop. We need all four induction burners and no other equipment."}, {"speaker": "Facility booking coordinator", "text": "The teaching kitchen's maximum attendance capacity is 24 people. Kitchen reservations are ready only when the date, time, attendance, equipment needs, and signed Food Preparation Form are provided."}, {"speaker": "Event organizer", "text": "I uploaded the signed sheet confirming we’ll follow the center’s food-handling and cleanup rules—the required paperwork for preparing food onsite. The organizer expects 20 people to attend the October 24, 2026, dumpling workshop."}, {"speaker": "Community center attendant", "text": "October 24 is available from 5:30–8:30 p.m., and all four burners are operational. No extra attendant is needed at this attendance level."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the capacity policy and form requirement while only altering attendance figures, which stays consistent with the vague attendant remark, and no evidence sentence reveals the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and a3 is a factual capacity relation rather than a policy conclusion. The base and counter assignments can be realized with only whether attendance exceeds capacity changing. Empty policy_evidence is appropriate because the governing readiness requirements are already retained in the questions object; the state adds no separate substantive policy needed to interpret that unchanged question.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction confirms an available date and time, attendance within capacity, fully stated equipment needs, operational requested equipment, and a signed document identified through an acceptable paraphrase as the required Food Preparation Form. These conditions are sufficient for a yes decision.", "rule_index": 0, "sound": true}, {"reason": "Refuting a3 entails that planned attendance exceeds capacity. That capacity violation is sufficient for a no decision regardless of the other confirmed requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The requested date of October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a2", "statement": "The requested time interval from 5:30 p.m. through 8:30 p.m. on October 24, 2026, is available for the teaching-kitchen reservation."}, {"id": "a3", "statement": "The planned attendance for the October 24, 2026, dumpling workshop does not exceed the teaching kitchen's maximum attendance capacity."}, {"id": "a4", "statement": "The organizer has fully stated the equipment needs for the October 24, 2026, dumpling workshop."}, {"id": "a5", "statement": "Every induction burner requested for the October 24, 2026, dumpling workshop is operational."}, {"id": "a6", "statement": "The organizer's uploaded sheet for the October 24, 2026, dumpling workshop is signed."}, {"id": "a7", "statement": "The organizer's uploaded sheet is the required Food Preparation Form for preparing food onsite, as identified by an unambiguous paraphrase of the form's name and purpose."}], "base_state_json": "[{\"speaker\":\"Event organizer\",\"text\":\"I’m requesting the teaching kitchen on October 24, 2026, from 5:30–8:30 p.m. for a dumpling workshop. We need all four induction burners and no other equipment.\"},{\"speaker\":\"Facility booking coordinator\",\"text\":\"The teaching kitchen's maximum attendance capacity is 24 people. Kitchen reservations are ready only when the date, time, attendance, equipment needs, and signed Food Preparation Form are provided.\"},{\"speaker\":\"Event organizer\",\"text\":\"I uploaded the signed sheet confirming we’ll follow the center’s food-handling and cleanup rules—the required paperwork for preparing food onsite. The organizer expects 20 people to attend the October 24, 2026, dumpling workshop.\"},{\"speaker\":\"Community center attendant\",\"text\":\"October 24 is available from 5:30–8:30 p.m., and all four burners are operational. No extra attendant is needed at this attendance level.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The teaching kitchen's maximum attendance capacity is 24 people."}, {"path": ["2", "text"], "text": "The organizer expects 20 people to attend the October 24, 2026, dumpling workshop."}], "policy_evidence": [], "rules": [{"justification": "The requested date and time are available, attendance does not exceed capacity, equipment needs are fully stated and the requested burners are operational, and the uploaded document is the signed required Food Preparation Form.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Even with every other listed readiness requirement confirmed, attendance exceeding the teaching kitchen's capacity is a noncompliant capacity condition and requires a no answer.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The teaching kitchen's maximum attendance capacity is 24 people.", "negative_left": "The teaching kitchen's maximum attendance capacity is 24 people.", "negative_right": "The organizer expects 30 people to attend the October 24, 2026, dumpling workshop.", "right": "The organizer expects 20 people to attend the October 24, 2026, dumpling workshop."}, "verifier_independent_model": false}, "family": "fast-43-diverse-160-005", "id": "fast-43-diverse-160-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one required booking detail, availability check, capacity condition, equipment statement, or form confirmation is missing or noncompliant.", "true": "Yes — every booking-readiness requirement is confirmed, including an acceptable paraphrased confirmation of the required form."}, "instructions": "Determine whether this reservation is booking-ready. Answer yes only if an available date and time, attendance not exceeding capacity, equipment needs, and the signed Food Preparation Form are all confirmed. An unambiguous paraphrase of the form’s name and purpose counts as confirmation. Otherwise answer no.", "type": "noul"}}, "state": [{"speaker": "Event organizer", "text": "I’m requesting the teaching kitchen on October 24, 2026, from 5:30–8:30 p.m. for a dumpling workshop. We need all four induction burners and no other equipment."}, {"speaker": "Facility booking coordinator", "text": "The teaching kitchen's maximum attendance capacity is 24 people. Kitchen reservations are ready only when the date, time, attendance, equipment needs, and signed Food Preparation Form are provided."}, {"speaker": "Event organizer", "text": "I uploaded the signed sheet confirming we’ll follow the center’s food-handling and cleanup rules—the required paperwork for preparing food onsite. The organizer expects 30 people to attend the October 24, 2026, dumpling workshop."}, {"speaker": "Community center attendant", "text": "October 24 is available from 5:30–8:30 p.m., and all four burners are operational. No extra attendant is needed at this attendance level."}]}, "method": "c2d", "provenance": {"source_id": "diverse-160", "source_is_synthetic": true, "source_sha256": "9e480c748b98f3456406e482785aa5767930aceddc01cc9235c887e20a6d0928", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text and question bindings (BR-4471, Oct 18 2026, 2-5pm) with only the attendee count changed (104 vs 78), the two focus evidence sentences are purely factual, and neither context states or implies a specific complexity level or rubric mapping.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\": \"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\", \"evidence\": [\"The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471.\", \"The attendance estimate filed under booking reference BR-4471 lists 104 attendees.\", \"The signed reservation form was submitted.\", \"The organizer explicitly does not request kitchen access and states that no food will be cooked or served.\", \"No specialized space, such as a gymnasium or media lab, is requested; the activity room alone is booked.\", \"No supervised audiovisual or sports equipment is needed.\", \"The community center attendant handles setup, and only the routine closing custodial check is scheduled, with no additional attendant or custodial assignment.\"], \"request\": \"Determine the booking’s operational-complexity level for routing and staffing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471."}, {"path": ["evidence", "1"], "text": "The attendance estimate filed under booking reference BR-4471 lists 104 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471.", "negative_left": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471.", "negative_right": "The attendance estimate filed under booking reference BR-4471 lists 78 attendees.", "right": "The attendance estimate filed under booking reference BR-4471 lists 104 attendees."}, "verifier_independent_model": false}, "family": "fast-43-diverse-162-003", "id": "fast-43-diverse-162-003-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471.", "The attendance estimate filed under booking reference BR-4471 lists 104 attendees.", "The signed reservation form was submitted.", "The organizer explicitly does not request kitchen access and states that no food will be cooked or served.", "No specialized space, such as a gymnasium or media lab, is requested; the activity room alone is booked.", "No supervised audiovisual or sports equipment is needed.", "The community center attendant handles setup, and only the routine closing custodial check is scheduled, with no additional attendant or custodial assignment."], "request": "Determine the booking’s operational-complexity level for routing and staffing."}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy text and question bindings (BR-4471, Oct 18 2026, 2-5pm) with only the attendee count changed (104 vs 78), the two focus evidence sentences are purely factual, and neither context states or implies a specific complexity level or rubric mapping.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\": \"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\", \"evidence\": [\"The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471.\", \"The attendance estimate filed under booking reference BR-4471 lists 104 attendees.\", \"The signed reservation form was submitted.\", \"The organizer explicitly does not request kitchen access and states that no food will be cooked or served.\", \"No specialized space, such as a gymnasium or media lab, is requested; the activity room alone is booked.\", \"No supervised audiovisual or sports equipment is needed.\", \"The community center attendant handles setup, and only the routine closing custodial check is scheduled, with no additional attendant or custodial assignment.\"], \"request\": \"Determine the booking’s operational-complexity level for routing and staffing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471."}, {"path": ["evidence", "1"], "text": "The attendance estimate filed under booking reference BR-4471 lists 104 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471.", "negative_left": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471.", "negative_right": "The attendance estimate filed under booking reference BR-4471 lists 78 attendees.", "right": "The attendance estimate filed under booking reference BR-4471 lists 104 attendees."}, "verifier_independent_model": false}, "family": "fast-43-diverse-162-003", "id": "fast-43-diverse-162-003-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BR-4471.", "The attendance estimate filed under booking reference BR-4471 lists 78 attendees.", "The signed reservation form was submitted.", "The organizer explicitly does not request kitchen access and states that no food will be cooked or served.", "No specialized space, such as a gymnasium or media lab, is requested; the activity room alone is booked.", "No supervised audiovisual or sports equipment is needed.", "The community center attendant handles setup, and only the routine closing custodial check is scheduled, with no additional attendant or custodial assignment."], "request": "Determine the booking’s operational-complexity level for routing and staffing."}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing policy text and RCC-4471/date/time bindings; the focus evidence pair is two factual sentences (booking reference and attendance count) with no policy or instruction language; the counterfactual only swaps the attendee count (74 vs 128) while keeping all other facts consistent, producing no contradictions; neither context states or hints at a rubric level, score, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\":\"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\",\"evidence\":[\"The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471.\",\"The attendance estimate filed under booking reference RCC-4471 lists 128 attendees.\",\"The organizer explicitly does not request kitchen access and states that no food will be cooked or served.\",\"Only the activity room is requested; no additional specialized spaces such as the gymnasium or auditorium are involved.\",\"No supervised audiovisual or sports equipment is requested.\",\"The community center attendant can perform setup, and the custodial scheduler lists only the routine closing check, with no additional staff assignments beyond the standard attendant.\",\"The signed reservation form was submitted and is complete.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471."}, {"path": ["evidence", "1"], "text": "The attendance estimate filed under booking reference RCC-4471 lists 128 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471.", "negative_left": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471.", "negative_right": "The attendance estimate filed under booking reference RCC-4471 lists 74 attendees.", "right": "The attendance estimate filed under booking reference RCC-4471 lists 128 attendees."}, "verifier_independent_model": false}, "family": "fast-43-diverse-162-005", "id": "fast-43-diverse-162-005-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471.", "The attendance estimate filed under booking reference RCC-4471 lists 128 attendees.", "The organizer explicitly does not request kitchen access and states that no food will be cooked or served.", "Only the activity room is requested; no additional specialized spaces such as the gymnasium or auditorium are involved.", "No supervised audiovisual or sports equipment is requested.", "The community center attendant can perform setup, and the custodial scheduler lists only the routine closing check, with no additional staff assignments beyond the standard attendant.", "The signed reservation form was submitted and is complete."]}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing policy text and RCC-4471/date/time bindings; the focus evidence pair is two factual sentences (booking reference and attendance count) with no policy or instruction language; the counterfactual only swaps the attendee count (74 vs 128) while keeping all other facts consistent, producing no contradictions; neither context states or hints at a rubric level, score, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\":\"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\",\"evidence\":[\"The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471.\",\"The attendance estimate filed under booking reference RCC-4471 lists 128 attendees.\",\"The organizer explicitly does not request kitchen access and states that no food will be cooked or served.\",\"Only the activity room is requested; no additional specialized spaces such as the gymnasium or auditorium are involved.\",\"No supervised audiovisual or sports equipment is requested.\",\"The community center attendant can perform setup, and the custodial scheduler lists only the routine closing check, with no additional staff assignments beyond the standard attendant.\",\"The signed reservation form was submitted and is complete.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471."}, {"path": ["evidence", "1"], "text": "The attendance estimate filed under booking reference RCC-4471 lists 128 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471.", "negative_left": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471.", "negative_right": "The attendance estimate filed under booking reference RCC-4471 lists 74 attendees.", "right": "The attendance estimate filed under booking reference RCC-4471 lists 128 attendees."}, "verifier_independent_model": false}, "family": "fast-43-diverse-162-005", "id": "fast-43-diverse-162-005-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. was assigned booking reference RCC-4471.", "The attendance estimate filed under booking reference RCC-4471 lists 74 attendees.", "The organizer explicitly does not request kitchen access and states that no food will be cooked or served.", "Only the activity room is requested; no additional specialized spaces such as the gymnasium or auditorium are involved.", "No supervised audiovisual or sports equipment is requested.", "The community center attendant can perform setup, and the custodial scheduler lists only the routine closing check, with no additional staff assignments beyond the standard attendant.", "The signed reservation form was submitted and is complete."]}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy text and the unchanged question criteria fully define scoring, so policy is preserved. Booking reference, date, time, and room path remain identical across base and counterfactual. The two focus evidence items are complete factual sentences about booking reference and attendance count, not policy or instructions. The counterfactual only changes the attendance figure (128→74) without introducing contradictory duplicate counts, staying coherent with the rest of the unchanged evidence. Neither context states a rule table, level label, or explicit answer, so there is no answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\": \"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\", \"evidence\": [\"The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471.\", \"The attendance estimate filed under booking reference BK-4471 lists 128 attendees.\", \"The organizer explicitly does not request kitchen access and states that no food will be cooked or served.\", \"Only the activity room is reserved; no additional specialized spaces such as the auditorium or gymnasium are requested.\", \"The custodial scheduler lists only the routine closing check, and no additional attendant or custodial staff assignments have been made.\", \"The signed reservation form and attendance estimate were submitted.\"], \"request\": \"Determine the booking\\u2019s operational-complexity level for routing and staffing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471."}, {"path": ["evidence", "1"], "text": "The attendance estimate filed under booking reference BK-4471 lists 128 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471.", "negative_left": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471.", "negative_right": "The attendance estimate filed under booking reference BK-4471 lists 74 attendees.", "right": "The attendance estimate filed under booking reference BK-4471 lists 128 attendees."}, "verifier_independent_model": false}, "family": "fast-43-diverse-162-008", "id": "fast-43-diverse-162-008-base", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471.", "The attendance estimate filed under booking reference BK-4471 lists 128 attendees.", "The organizer explicitly does not request kitchen access and states that no food will be cooked or served.", "Only the activity room is reserved; no additional specialized spaces such as the auditorium or gymnasium are requested.", "The custodial scheduler lists only the routine closing check, and no additional attendant or custodial staff assignments have been made.", "The signed reservation form and attendance estimate were submitted."], "request": "Determine the booking’s operational-complexity level for routing and staffing."}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "public_services-02", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing policy text and the unchanged question criteria fully define scoring, so policy is preserved. Booking reference, date, time, and room path remain identical across base and counterfactual. The two focus evidence items are complete factual sentences about booking reference and attendance count, not policy or instructions. The counterfactual only changes the attendance figure (128→74) without introducing contradictory duplicate counts, staying coherent with the rest of the unchanged evidence. Neither context states a rule table, level label, or explicit answer, so there is no answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "counterfactual": {"a1": "supported", "a2": "refuted", "a3": "refuted", "a4": "refuted", "a5": "refuted"}, "remove_left": {"a2": "unknown"}, "remove_right": {"a2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a2": "unknown"}, "negative_pair": {"a2": "refuted"}, "negative_sentence": {"a2": "unknown"}, "positive_pair": {"a2": "supported"}, "right": {"a2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and the focus atom is a factual attendee-threshold proposition rather than a policy conclusion. The base and counter assignments differ only on the focus and are realizable, representing respectively more than 100 attendees and 61–100 attendees. The policy evidence preserves the substantive operational-complexity rules originating in the original state; criteria in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A filed attendance estimate exceeding 100 establishes the rubric’s more-than-100-attendees trigger, which is independently sufficient for level 4.", "rule_index": 0, "sound": true}, {"reason": "The conditions establish 61–100 attendees, sufficient for level 3, while excluding every level-4 trigger: more than 100 attendees, multiple specialized spaces, at least two additional staff assignments, and the combined kitchen-plus-supervised-equipment trigger through the explicit absence of kitchen access.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The attendee count for the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is at least 61."}, {"id": "a2", "statement": "The attendee count in the attendance estimate filed under the booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. exceeds 100."}, {"id": "a3", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves at least two specialized spaces."}, {"id": "a4", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. involves kitchen access."}, {"id": "a5", "statement": "The Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. has at least two additional staff assignments."}], "base_state_json": "{\"context\": \"Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.\", \"evidence\": [\"The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471.\", \"The attendance estimate filed under booking reference BK-4471 lists 128 attendees.\", \"The organizer explicitly does not request kitchen access and states that no food will be cooked or served.\", \"Only the activity room is reserved; no additional specialized spaces such as the auditorium or gymnasium are requested.\", \"The custodial scheduler lists only the routine closing check, and no additional attendant or custodial staff assignments have been made.\", \"The signed reservation form and attendance estimate were submitted.\"], \"request\": \"Determine the booking\\u2019s operational-complexity level for routing and staffing.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}], "focus_atom": "a2", "focus_evidence": [{"path": ["evidence", "0"], "text": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471."}, {"path": ["evidence", "1"], "text": "The attendance estimate filed under booking reference BK-4471 lists 128 attendees."}], "policy_evidence": [{"path": ["context"], "text": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete."}], "rules": [{"justification": "An attendee count exceeding 100 independently satisfies the very-high criterion.", "target": "4", "when": [{"atom_id": "a2", "state": "supported"}]}, {"justification": "The attendee count is at least 61 and not more than 100, satisfying the high criterion. The request does not involve multiple specialized spaces or at least two additional staff assignments, and the absence of kitchen access makes the combined kitchen-and-supervised-equipment trigger impossible; therefore no very-high trigger applies.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "refuted"}]}]}, "verified_pair": {"left": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471.", "negative_left": "The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471.", "negative_right": "The attendance estimate filed under booking reference BK-4471 lists 74 attendees.", "right": "The attendance estimate filed under booking reference BK-4471 lists 128 attendees."}, "verifier_independent_model": false}, "family": "fast-43-diverse-162-008", "id": "fast-43-diverse-162-008-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Minimal: A complete request for 1–15 attendees needing no room setup, equipment, kitchen access, attendant support, or special custodial work.", "1 — Low: A complete request for 16–30 attendees or a request needing only a routine room opening, with no furniture setup, specialized space, supervised equipment, or extra custodial task.", "2 — Moderate: A complete activity-room request for 31–60 attendees requiring one standard table-and-chair setup and only the routine closing check, with no kitchen use, cooking, food service, specialized equipment, or dedicated supervision.", "3 — High: Any request involving kitchen access, cooking, food service, supervised audiovisual or sports equipment, 61–100 attendees, or an additional attendant or custodial assignment.", "4 — Very high: A request involving more than 100 attendees, multiple specialized spaces, both kitchen and supervised equipment, or two or more additional staff assignments."], "instructions": "Select the single level that matches the stated rubric. Treat explicit statements that a service is not requested as controlling evidence.", "type": "score"}}, "state": {"context": "Riverside Community Center scores reservation requests by operational complexity. Kitchen use, cooking, or dedicated equipment supervision requires level 3 or higher. A standard activity-room booking for 31–60 attendees that needs one furniture setup but no specialized space or equipment is level 2 when the required form is complete.", "evidence": ["The booking reference assigned to the Riverside Community Center reservation request for October 18, 2026, from 2:00–5:00 p.m. is BK-4471.", "The attendance estimate filed under booking reference BK-4471 lists 74 attendees.", "The organizer explicitly does not request kitchen access and states that no food will be cooked or served.", "Only the activity room is reserved; no additional specialized spaces such as the auditorium or gymnasium are requested.", "The custodial scheduler lists only the routine closing check, and no additional attendant or custodial staff assignments have been made.", "The signed reservation form and attendance estimate were submitted."], "request": "Determine the booking’s operational-complexity level for routing and staffing."}}, "method": "c2d", "provenance": {"source_id": "diverse-162", "source_is_synthetic": true, "source_sha256": "5e0e7c3b43cbdf8e25629e4109ec23eee7b104e7a5bd4ec29a69ff3f659c140f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "public_services-02", "split": "train", "variant": "counterfactual"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric and routing policy and the same material type/date/no-packet-number bindings from the original question; the two focus evidence spans are plain factual sentences about the request and P2's cataloged date; the counterfactual only alters P2's cataloged meeting date to 2024-07-09, which is a coherent single-fact change consistent with the rest of the unchanged context (still two packets, one box); neither context states or hints at the final Ready/Not-ready label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\": \"Case note - City Clerk Records Unit, request R: The request specifies the material type as a council meeting packet. Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11. As of 2026-09-17, the archive holds exactly two council meeting packets, P1 and P2, in this office. Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-06-11. Packet P1's cataloged meeting date matches the same 2024-06-11 date. Request R does not include any packet number. Retrieving both P1 and P2 requires pulling from a single archive box. For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["context"], "text": "Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11."}, {"path": ["context"], "text": "Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-06-11."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11.", "negative_left": "Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11.", "negative_right": "Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-07-09.", "right": "Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-06-11."}, "verifier_independent_model": false}, "family": "fast-43-diverse-177-001", "id": "fast-43-diverse-177-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "Case note - City Clerk Records Unit, request R: The request specifies the material type as a council meeting packet. Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11. As of 2026-09-17, the archive holds exactly two council meeting packets, P1 and P2, in this office. Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-06-11. Packet P1's cataloged meeting date matches the same 2024-06-11 date. Request R does not include any packet number. Retrieving both P1 and P2 requires pulling from a single archive box. For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "public_services-05", "split": "train", "variant": "base"} {"domain": "public_services", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric and routing policy and the same material type/date/no-packet-number bindings from the original question; the two focus evidence spans are plain factual sentences about the request and P2's cataloged date; the counterfactual only alters P2's cataloged meeting date to 2024-07-09, which is a coherent single-fact change consistent with the rest of the unchanged context (still two packets, one box); neither context states or hints at the final Ready/Not-ready label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "refuted", "a7": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the exact-set atom is an allowable atomic set relation. The focus atom a5 is a factual catalog-date relation, not a policy conclusion. The base and counter assignments differ only on a5 and are both realizable: P2 can either share P1's matching date or have a different date while all other facts remain fixed. The cited state context preserves the substantive status rubric, scope, and exclusion rule needed to interpret the unchanged question, so policy evidence is complete.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The exact archive set is P1 and P2; P1 matches the specified meeting date and P2 is refuted as matching it. Thus exactly one archived council packet matches. The material type is supplied, and retrieval spans only one box, excluding Ready—complex. The rule therefore sufficiently entails Yes/Ready—standard.", "rule_index": 0, "sound": true}, {"reason": "The exact archive set is P1 and P2, and both packets match the specified meeting date. With no packet number distinguishing them, more than one packet remains possible, so the rule sufficiently entails No/Not ready.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Request R submitted to the City Clerk Records Unit specifies the material type as a council meeting packet."}, {"id": "a2", "statement": "The meeting date specified in request R is 2024-06-11."}, {"id": "a3", "statement": "As of 2026-09-17, the set of archived council meeting packets in the City Clerk Records Unit is exactly {packet P1, packet P2}."}, {"id": "a4", "statement": "The cataloged meeting date of packet P1 is the same as the meeting date specified in request R."}, {"id": "a5", "statement": "The cataloged meeting date of packet P2 is the same as the meeting date specified in request R."}, {"id": "a6", "statement": "Request R supplies a packet number."}, {"id": "a7", "statement": "Retrieval of packets P1 and P2 spans one archive box."}], "base_state_json": "{\"context\": \"Case note - City Clerk Records Unit, request R: The request specifies the material type as a council meeting packet. Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11. As of 2026-09-17, the archive holds exactly two council meeting packets, P1 and P2, in this office. Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-06-11. Packet P1's cataloged meeting date matches the same 2024-06-11 date. Request R does not include any packet number. Retrieving both P1 and P2 requires pulling from a single archive box. For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": ["context"], "text": "Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11."}, {"path": ["context"], "text": "Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-06-11."}], "policy_evidence": [{"path": ["context"], "text": "For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}], "rules": [{"justification": "P1 is the sole archived council meeting packet whose cataloged date matches the supplied meeting date, no competing packet number is supplied, and retrieval spans fewer than three archive boxes. The material type and one uniquely identifying meeting date therefore establish Ready—standard.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Both archived council meeting packets match the supplied meeting date, and no packet number distinguishes them. More than one packet therefore remains possible, so the request is Not ready rather than Ready—standard.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11.", "negative_left": "Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11.", "negative_right": "Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-07-09.", "right": "Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-06-11."}, "verifier_independent_model": false}, "family": "fast-43-diverse-177-001", "id": "fast-43-diverse-177-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — classify the request as Not ready because more than one packet remains possible.", "true": "Yes — classify the request as Ready—standard because exactly one packet is uniquely identified."}, "instructions": "Answer Yes only if the supplied details uniquely identify one packet under the stated status rubric; otherwise answer No.", "type": "noul"}}, "state": {"context": "Case note - City Clerk Records Unit, request R: The request specifies the material type as a council meeting packet. Request R submitted to the City Clerk Records Unit specifies the meeting date as 2024-06-11. As of 2026-09-17, the archive holds exactly two council meeting packets, P1 and P2, in this office. Packet P2's cataloged meeting date in the City Clerk Records Unit is 2024-07-09. Packet P1's cataloged meeting date matches the same 2024-06-11 date. Request R does not include any packet number. Retrieving both P1 and P2 requires pulling from a single archive box. For archived council meeting packets, the municipal information clerk routes requests to the City Clerk Records Unit. Status rubric: Ready—standard requires the material type plus one uniquely identifying meeting date or packet number. Ready—complex applies only after those details are complete and retrieval spans at least three archive boxes. Not ready applies when no unique identifier is supplied. An exclusion does not positively identify a remaining record."}}, "method": "c2d", "provenance": {"source_id": "diverse-177", "source_is_synthetic": true, "source_sha256": "086312f8fe85f69f712c295940c20edfd107070e903933d64f2ce665bb75f1b5", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "public_services-05", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy and entity/time bindings, the two focus evidence sentences are plain factual statements, the counterfactual's changed bottle code (Y-208) is consistent with the stated uniqueness rule and shifts origin logically without contradicting other facts, and neither context reveals any answer code or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Maintenance coordinator Joel reviews the incident from 9:10 p.m., when blue liquid was found trailing from an uncapped detergent bottle on the shelf to the floor near the washer. The only two possible origins considered for this spill are the uncapped detergent bottle on the shelf and the washer itself; no other source is under investigation. Inspector Nia confirms the spill has a single origin. Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114. Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code X-114. Separately, the washer-moisture profile code on file is distinct from the bottle's code, and each profile code uniquely identifies its associated source within the investigated set. The bottle is physically separate from the washer, and its contents are a standard detergent product. Inspection finds no active flooding and no electrical signs near the washer. The washer's logs show no recurring fault, and the cycle notes do not confirm any unbalanced or overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114."}, {"path": [], "text": "Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code X-114."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114.", "negative_left": "Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114.", "negative_right": "Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code Y-208.", "right": "Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code X-114."}, "verifier_independent_model": false}, "family": "fast-43-diverse-181-006", "id": "fast-43-diverse-181-006-base", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Maintenance coordinator Joel reviews the incident from 9:10 p.m., when blue liquid was found trailing from an uncapped detergent bottle on the shelf to the floor near the washer. The only two possible origins considered for this spill are the uncapped detergent bottle on the shelf and the washer itself; no other source is under investigation. Inspector Nia confirms the spill has a single origin. Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114. Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code X-114. Separately, the washer-moisture profile code on file is distinct from the bottle's code, and each profile code uniquely identifies its associated source within the investigated set. The bottle is physically separate from the washer, and its contents are a standard detergent product. Inspection finds no active flooding and no electrical signs near the washer. The washer's logs show no recurring fault, and the cycle notes do not confirm any unbalanced or overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy and entity/time bindings, the two focus evidence sentences are plain factual statements, the counterfactual's changed bottle code (Y-208) is consistent with the stated uniqueness rule and shifts origin logically without contradicting other facts, and neither context reveals any answer code or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Maintenance coordinator Joel reviews the incident from 9:10 p.m., when blue liquid was found trailing from an uncapped detergent bottle on the shelf to the floor near the washer. The only two possible origins considered for this spill are the uncapped detergent bottle on the shelf and the washer itself; no other source is under investigation. Inspector Nia confirms the spill has a single origin. Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114. Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code X-114. Separately, the washer-moisture profile code on file is distinct from the bottle's code, and each profile code uniquely identifies its associated source within the investigated set. The bottle is physically separate from the washer, and its contents are a standard detergent product. Inspection finds no active flooding and no electrical signs near the washer. The washer's logs show no recurring fault, and the cycle notes do not confirm any unbalanced or overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114."}, {"path": [], "text": "Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code X-114."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114.", "negative_left": "Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114.", "negative_right": "Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code Y-208.", "right": "Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code X-114."}, "verifier_independent_model": false}, "family": "fast-43-diverse-181-006", "id": "fast-43-diverse-181-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Maintenance coordinator Joel reviews the incident from 9:10 p.m., when blue liquid was found trailing from an uncapped detergent bottle on the shelf to the floor near the washer. The only two possible origins considered for this spill are the uncapped detergent bottle on the shelf and the washer itself; no other source is under investigation. Inspector Nia confirms the spill has a single origin. Lab technicians recorded the source-profile identifier of the 9:10 p.m. floor-spill sample as profile code X-114. Lab technicians recorded the source-profile identifier of the liquid in the uncapped detergent bottle on the shelf as profile code Y-208. Separately, the washer-moisture profile code on file is distinct from the bottle's code, and each profile code uniquely identifies its associated source within the investigated set. The bottle is physically separate from the washer, and its contents are a standard detergent product. Inspection finds no active flooding and no electrical signs near the washer. The washer's logs show no recurring fault, and the cycle notes do not confirm any unbalanced or overloaded load. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "basic_inspection_hold_high"}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy and unchanged questions, only the bottle-liquid profile code changes from SP-114 to SP-227, which is a coherent single-sentence factual alteration consistent with the stated uniqueness and origin rules, and the focus evidence quotes match exactly with no embedded answer or rule leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Household maintenance coordinator Joel reviews the case: the washer logs show a normal completion, no error code, and no recurring fault. Cycle notes do not confirm an unbalanced or overloaded load. Inspector Nia records dry inlet hoses, drain hose, door seal, and machine base, with no active flooding and no electrical signs present at the washer. The uncapped detergent bottle on the shelf is external to the washer, and its liquid is a detergent product. The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer. Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within that set, the bottle-liquid identifier differs from the washer-moisture identifier, and the spill has exactly one origin. Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114. Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-114. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114."}, {"path": [], "text": "Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-114."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114.", "negative_left": "Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114.", "negative_right": "Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-227.", "right": "Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-114."}, "verifier_independent_model": false}, "family": "fast-43-diverse-181-019", "id": "fast-43-diverse-181-019-base", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Household maintenance coordinator Joel reviews the case: the washer logs show a normal completion, no error code, and no recurring fault. Cycle notes do not confirm an unbalanced or overloaded load. Inspector Nia records dry inlet hoses, drain hose, door seal, and machine base, with no active flooding and no electrical signs present at the washer. The uncapped detergent bottle on the shelf is external to the washer, and its liquid is a detergent product. The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer. Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within that set, the bottle-liquid identifier differs from the washer-moisture identifier, and the spill has exactly one origin. Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114. Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-114. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-01", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy and unchanged questions, only the bottle-liquid profile code changes from SP-114 to SP-227, which is a coherent single-sentence factual alteration consistent with the stated uniqueness and origin rules, and the focus evidence quotes match exactly with no embedded answer or rule leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "refuted", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the quantified uniqueness atom remains a single relationship over the investigation set. The focus atom is a factual identifier-equality claim, not a policy conclusion. Both assignments are jointly realizable: in the base the spill identifier equals the bottle identifier, while in the counter it equals the distinct washer identifier, with all other atoms unchanged. The policy evidence correctly preserves the substantive rules originating in the original state; the remaining criteria and exact-match instructions are already retained in the questions object and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction entails that the spill sample matches the bottle liquid and, through the exact candidate set, unique source-profile mapping, and single-origin condition, that the spill is a confirmed external detergent-product spill. The state policy therefore yields cleanup, reuse permitted, and low urgency. No listed substantive option has that exact combination, while the uncertain-source, recurring-fault, and load-based alternatives are excluded, so none_of_above is sufficient.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a1 means the spill identifier differs from the bottle identifier. Given that the spill identifier is either the bottle or washer identifier and those identifiers differ, it must equal the washer identifier. The uniqueness and single-origin conditions establish washer-origin moisture. With no active flooding or electrical signs, the unchanged question criteria require basic inspection, prohibited reuse, and high urgency. The remaining atoms exclude the competing substantive options.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample is identical to the source-profile identifier measured for the liquid in the uncapped detergent bottle on the shelf."}, {"id": "a2", "statement": "The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer."}, {"id": "a3", "statement": "The source-profile identifier measured for the 9:10 p.m. floor-spill sample belongs to the set containing the bottle-liquid identifier and the washer-moisture identifier."}, {"id": "a4", "statement": "The source-profile identifier measured for the bottle liquid differs from the source-profile identifier measured for washer moisture."}, {"id": "a5", "statement": "Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within the investigated possible-origin set."}, {"id": "a6", "statement": "The 9:10 p.m. floor spill has exactly one origin."}, {"id": "a7", "statement": "The uncapped detergent bottle on the shelf is external to the washer."}, {"id": "a8", "statement": "The liquid in the uncapped detergent bottle on the shelf is a detergent product."}, {"id": "a9", "statement": "No active flooding is present at the washer."}, {"id": "a10", "statement": "No electrical signs are present at the washer."}, {"id": "a11", "statement": "The washer logs show no recurring fault."}, {"id": "a12", "statement": "The cycle notes do not confirm an unbalanced load."}, {"id": "a13", "statement": "The cycle notes do not confirm an overloaded load."}], "base_state_json": "\"Household maintenance coordinator Joel reviews the case: the washer logs show a normal completion, no error code, and no recurring fault. Cycle notes do not confirm an unbalanced or overloaded load. Inspector Nia records dry inlet hoses, drain hose, door seal, and machine base, with no active flooding and no electrical signs present at the washer. The uncapped detergent bottle on the shelf is external to the washer, and its liquid is a detergent product. The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer. Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within that set, the bottle-liquid identifier differs from the washer-moisture identifier, and the spill has exactly one origin. Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114. Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-114. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": [], "text": "Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114."}, {"path": [], "text": "Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-114."}], "policy_evidence": [{"path": [], "text": "Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency."}, {"path": [], "text": "Appliance-origin moisture instead requires basic inspection and no reuse."}], "rules": [{"justification": "The matching source-profile identifiers, exhaustive possible-origin set, unique source identification, and single-origin finding establish that the spill originated from the external detergent-product bottle. Policy therefore requires cleanup, permits reuse, and assigns low urgency. No substantive option has that exact three-field result. The known source excludes the uncertain-source cleanup option, while the absent recurring fault and absent load findings exclude the maintenance and load-correction options.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}, {"justification": "The spill identifier is one of the two candidate-source identifiers but is not the bottle identifier; because the bottle and washer identifiers differ and uniquely identify their associated sources, the single-origin spill is established as washer-origin moisture. With no active flooding or electrical signs, this exactly requires basic inspection, prohibited reuse, and high urgency. Washer origin excludes an external uncertain-source spill, while the absent recurring fault and absent load findings exclude the other substantive routes.", "target": "basic_inspection_hold_high", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}]}]}, "verified_pair": {"left": "Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114.", "negative_left": "Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114.", "negative_right": "Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-227.", "right": "Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-114."}, "verifier_independent_model": false}, "family": "fast-43-diverse-181-019", "id": "fast-43-diverse-181-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"basic_inspection_hold_high": "Route to basic inspection, prohibit reuse, and assign high urgency; applies when evidence shows moisture originating from the appliance but no active flooding or electrical signs.", "cleanup_hold_moderate": "Route to cleanup, do not use the washer again yet, and assign moderate urgency; applies when a spill is external but its source remains uncertain.", "deferred_maintenance_reuse_moderate": "Route to deferred maintenance, permit reuse, and assign moderate urgency; applies when logs show a recurring non-safety fault that does not prevent normal cycles.", "load_correction_reuse_low": "Route to load correction, permit reuse, and assign low urgency; applies when cycle notes confirm an unbalanced or overloaded load without leakage.", "none_of_above": "Select when none of the other options exactly matches the evidence-based route, reuse decision, and urgency."}, "instructions": "Select the one option whose route, reuse decision, and urgency all exactly match the explicit evidence and policy. Urgency is ordered low < moderate < high < emergency. If no substantive option matches all three fields, select none_of_above.", "type": "choice"}}, "state": "Household maintenance coordinator Joel reviews the case: the washer logs show a normal completion, no error code, and no recurring fault. Cycle notes do not confirm an unbalanced or overloaded load. Inspector Nia records dry inlet hoses, drain hose, door seal, and machine base, with no active flooding and no electrical signs present at the washer. The uncapped detergent bottle on the shelf is external to the washer, and its liquid is a detergent product. The investigated possible-origin set for the 9:10 p.m. floor spill consists exactly of the uncapped detergent bottle on the shelf and the washer. Every source-profile identifier used in the investigation uniquely identifies its associated possible origin within that set, the bottle-liquid identifier differs from the washer-moisture identifier, and the spill has exactly one origin. Lab analysis logged the source-profile identifier for the 9:10 p.m. floor-spill sample as profile code SP-114. Lab analysis logged the source-profile identifier for the liquid in the uncapped detergent bottle on the shelf as profile code SP-227. Policy: confirmed external product spills route to cleanup, permit appliance reuse, and receive low urgency. Appliance-origin moisture instead requires basic inspection and no reuse."}, "method": "c2d", "provenance": {"source_id": "diverse-181", "source_is_synthetic": true, "source_sha256": "689ac45cb86fb6472a9efa007a76da8e27fab6bf577de0df1c7eae8f50581ff6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "basic_inspection_hold_high"}, "source_family": "home-01", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy and question bindings, use complete factual evidence sentences, and the counterfactual (45 seconds vs 2:10) coherently alters only the timer measurement without contradicting other evidence or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Rice is fully cooked, and the tofu is browned on every listed side.\",\"Bowls are portioned; two storage containers are labeled and ready.\",\"No ingredients are missing, and nothing is contaminated.\",\"The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving.\",\"The kitchen timer read two minutes and ten seconds when it was stopped.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "3"], "text": "The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving."}, {"path": ["evidence", "4"], "text": "The kitchen timer read two minutes and ten seconds when it was stopped."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving.", "negative_left": "The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving.", "negative_right": "The kitchen timer read forty-five seconds when it was stopped.", "right": "The kitchen timer read two minutes and ten seconds when it was stopped."}, "verifier_independent_model": false}, "family": "fast-43-diverse-187-019", "id": "fast-43-diverse-187-019-base", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Rice is fully cooked, and the tofu is browned on every listed side.", "Bowls are portioned; two storage containers are labeled and ready.", "No ingredients are missing, and nothing is contaminated.", "The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving.", "The kitchen timer read two minutes and ten seconds when it was stopped."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_5_ready"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy and question bindings, use complete factual evidence sentences, and the counterfactual (45 seconds vs 2:10) coherently alters only the timer measurement without contradicting other evidence or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Rice is fully cooked, and the tofu is browned on every listed side.\",\"Bowls are portioned; two storage containers are labeled and ready.\",\"No ingredients are missing, and nothing is contaminated.\",\"The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving.\",\"The kitchen timer read two minutes and ten seconds when it was stopped.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "3"], "text": "The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving."}, {"path": ["evidence", "4"], "text": "The kitchen timer read two minutes and ten seconds when it was stopped."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving.", "negative_left": "The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving.", "negative_right": "The kitchen timer read forty-five seconds when it was stopped.", "right": "The kitchen timer read two minutes and ten seconds when it was stopped."}, "verifier_independent_model": false}, "family": "fast-43-diverse-187-019", "id": "fast-43-diverse-187-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Rice is fully cooked, and the tofu is browned on every listed side.", "Bowls are portioned; two storage containers are labeled and ready.", "No ingredients are missing, and nothing is contaminated.", "The kitchen timer was started when the glaze began bubbling continuously and was stopped when the bubbling ceased before serving.", "The kitchen timer read forty-five seconds when it was stopped."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_2_major_correction"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical recipe/kitchen-rule policy text and unchanged request/decision criteria, only the glaze timer stop reading differs (6:02:15 vs 6:01:10) which is a coherent single-fact change yielding a plausible alternate duration without contradicting other evidence, the focus evidence are two complete factual sentences rather than policy or instructions, and neither context states or hints at the correct readiness level or rule outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Rice is fully cooked, and the tofu is browned on every listed side.\",\"The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased.\",\"The kitchen timer for the glaze read 6:02:15 PM when it was stopped.\",\"Bowls are portioned; two storage containers are labeled and ready.\",\"No ingredients are missing, and nothing is contaminated.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "1"], "text": "The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased."}, {"path": ["evidence", "2"], "text": "The kitchen timer for the glaze read 6:02:15 PM when it was stopped."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased.", "negative_left": "The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased.", "negative_right": "The kitchen timer for the glaze read 6:01:10 PM when it was stopped.", "right": "The kitchen timer for the glaze read 6:02:15 PM when it was stopped."}, "verifier_independent_model": false}, "family": "fast-43-diverse-187-024", "id": "fast-43-diverse-187-024-base", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Rice is fully cooked, and the tofu is browned on every listed side.", "The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased.", "The kitchen timer for the glaze read 6:02:15 PM when it was stopped.", "Bowls are portioned; two storage containers are labeled and ready.", "No ingredients are missing, and nothing is contaminated."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_5_ready"}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical recipe/kitchen-rule policy text and unchanged request/decision criteria, only the glaze timer stop reading differs (6:02:15 vs 6:01:10) which is a coherent single-fact change yielding a plausible alternate duration without contradicting other evidence, the focus evidence are two complete factual sentences rather than policy or instructions, and neither context states or hints at the correct readiness level or rule outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "counterfactual": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relationship; the universally quantified atoms remain atomic because each applies one predicate over an explicit batch or set. A3 is a factual observation about bubbling duration rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A3. Policy evidence preserves the substantive recipe, timed-heating, role, and workflow rules originating in the original state; the unchanged questions automatically preserve the readiness criteria and ordering instructions.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all expressly stated recipe requirements, including the timed heating step, and affirmatively excludes missing ingredients and contamination. No higher-priority indeterminate, discard, or correction outcome remains applicable, so level_5_ready follows.", "rule_index": 0, "sound": true}, {"reason": "Refuting A3 establishes that the required two-minute continuous bubbling step is unmet. The other conditions establish the remaining stated recipe facts and exclude contamination or missing ingredients, so the ordered criteria require level_2_major_correction and return to the home cook and stove.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The rice for the four teriyaki tofu bowls is fully cooked."}, {"id": "A2", "statement": "The tofu for the four teriyaki tofu bowls is browned on every side listed by the recipe."}, {"id": "A3", "statement": "The observed continuous bubbling interval for the glaze before serving lasted at least two minutes."}, {"id": "A4", "statement": "Every ingredient required for the four teriyaki tofu bowls is present."}, {"id": "A5", "statement": "Every food item in the four-teriyaki-tofu-bowl batch is uncontaminated."}, {"id": "A6", "statement": "The four teriyaki tofu bowls are portioned."}, {"id": "A7", "statement": "Every portion designated for refrigerated storage from the four-teriyaki-tofu-bowl batch is labeled."}], "base_state_json": "{\"context\":\"The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.\",\"evidence\":[\"Rice is fully cooked, and the tofu is browned on every listed side.\",\"The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased.\",\"The kitchen timer for the glaze read 6:02:15 PM when it was stopped.\",\"Bowls are portioned; two storage containers are labeled and ready.\",\"No ingredients are missing, and nothing is contaminated.\"],\"request\":\"Rate current readiness and select the required next workflow step.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}], "focus_atom": "A3", "focus_evidence": [{"path": ["evidence", "1"], "text": "The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased."}, {"path": ["evidence", "2"], "text": "The kitchen timer for the glaze read 6:02:15 PM when it was stopped."}], "policy_evidence": [{"path": ["context"], "text": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions."}, {"path": ["request"], "text": "Rate current readiness and select the required next workflow step."}], "rules": [{"justification": "All stated recipe, timed-heating, portioning, labeling, ingredient-presence, and contamination-safety requirements are affirmatively complete.", "target": "level_5_ready", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}, {"justification": "The required two-minute continuous bubbling step is affirmatively unmet, while the remaining required recipe and safety facts are known and do not trigger discard; the rubric therefore requires return to the home cook and stove.", "target": "level_2_major_correction", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}]}]}, "verified_pair": {"left": "The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased.", "negative_left": "The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased.", "negative_right": "The kitchen timer for the glaze read 6:01:10 PM when it was stopped.", "right": "The kitchen timer for the glaze read 6:02:15 PM when it was stopped."}, "verifier_independent_model": false}, "family": "fast-43-diverse-187-024", "id": "fast-43-diverse-187-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_0_indeterminate": "Readiness cannot be rated because evidence about at least one required recipe or safety condition is absent.", "level_1_discard": "Not recoverable: contamination or another stated violation requires discarding the affected food.", "level_2_major_correction": "A required cooking or timed heating step is unmet; route the food back to the home cook and stove before serving.", "level_3_minor_correction": "A recoverable non-heating recipe error remains, such as adding a present ingredient or correcting portions.", "level_4_portion_only": "Cooking is complete, but only portioning, garnishing, labeling, or another non-heating finish remains.", "level_5_ready": "Fully ready: every recipe and kitchen-rule requirement is complete; serve now or seal labeled storage portions."}, "instructions": "Choose exactly one ordered readiness level. Apply the first description that matches the evidence; unmet required heating is a major correction even when all other steps are complete.", "type": "choice"}}, "state": {"context": "The meal planner’s recipe for four teriyaki tofu bowls requires cooked rice, browned tofu, and glaze that bubbles continuously for two minutes before serving. Kitchen rules say any unmet timed cooking step requires return to the stove; garnish or portion adjustments do not. The home cook handles heat, the ingredient prep helper handles garnish, and the cleanup helper labels refrigerated portions.", "evidence": ["Rice is fully cooked, and the tofu is browned on every listed side.", "The kitchen timer for the glaze was started at 6:00:00 PM when bubbling began and stopped when bubbling ceased.", "The kitchen timer for the glaze read 6:01:10 PM when it was stopped.", "Bowls are portioned; two storage containers are labeled and ready.", "No ingredients are missing, and nothing is contaminated."], "request": "Rate current readiness and select the required next workflow step."}}, "method": "c2d", "provenance": {"source_id": "diverse-187", "source_is_synthetic": true, "source_sha256": "ffeda8611cc573ddb164703a3c9aad8bbb15187a1497ef95a6d471cb2c813d75", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_2_major_correction"}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy elements (probe calibration, override rule, four-portion split) fully alongside the unchanged question rubric; entities, request wording, and time framing (6:00 PM cooking finish) are preserved; the two focus evidence items are single factual sentences about cooking completion time and cooling start time; the counterfactual coherently shifts only the cooling start time to 8:45 PM without contradicting any other stated fact; no context contains a gold answer, rule table, or explicit output instruction beyond the narrative kitchen rule already present in the original state.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"Mara is finishing paprika chicken and rice for tonight's dinner.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74\\u00b0C; divide into four equal portions.\",\"Mara simmered the sauce fully before starting the chicken, then cooked it until the probe, freshly ice-water calibrated that morning, read 75\\u00b0C and 74\\u00b0C in the thickest pieces.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"The meal was divided into four equal portions: two plated for serving now and prepared accordingly, and two placed into labeled shallow containers for storage.\",\"Cooking of the chicken finished at 6:00 PM.\",\"The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM.\"],\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "4"], "text": "Cooking of the chicken finished at 6:00 PM."}, {"path": ["evidence", "5"], "text": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Cooking of the chicken finished at 6:00 PM.", "negative_left": "Cooking of the chicken finished at 6:00 PM.", "negative_right": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 8:45 PM.", "right": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-190-015", "id": "fast-43-diverse-190-015-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Mara is finishing paprika chicken and rice for tonight's dinner.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Mara simmered the sauce fully before starting the chicken, then cooked it until the probe, freshly ice-water calibrated that morning, read 75°C and 74°C in the thickest pieces.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "The meal was divided into four equal portions: two plated for serving now and prepared accordingly, and two placed into labeled shallow containers for storage.", "Cooking of the chicken finished at 6:00 PM.", "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."], "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy elements (probe calibration, override rule, four-portion split) fully alongside the unchanged question rubric; entities, request wording, and time framing (6:00 PM cooking finish) are preserved; the two focus evidence items are single factual sentences about cooking completion time and cooling start time; the counterfactual coherently shifts only the cooling start time to 8:45 PM without contradicting any other stated fact; no context contains a gold answer, rule table, or explicit output instruction beyond the narrative kitchen rule already present in the original state.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\":\"Mara is finishing paprika chicken and rice for tonight's dinner.\",\"evidence\":[\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74\\u00b0C; divide into four equal portions.\",\"Mara simmered the sauce fully before starting the chicken, then cooked it until the probe, freshly ice-water calibrated that morning, read 75\\u00b0C and 74\\u00b0C in the thickest pieces.\",\"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\",\"The meal was divided into four equal portions: two plated for serving now and prepared accordingly, and two placed into labeled shallow containers for storage.\",\"Cooking of the chicken finished at 6:00 PM.\",\"The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM.\"],\"request\":\"Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "4"], "text": "Cooking of the chicken finished at 6:00 PM."}, {"path": ["evidence", "5"], "text": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Cooking of the chicken finished at 6:00 PM.", "negative_left": "Cooking of the chicken finished at 6:00 PM.", "negative_right": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 8:45 PM.", "right": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-190-015", "id": "fast-43-diverse-190-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Mara is finishing paprika chicken and rice for tonight's dinner.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Mara simmered the sauce fully before starting the chicken, then cooked it until the probe, freshly ice-water calibrated that morning, read 75°C and 74°C in the thickest pieces.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "The meal was divided into four equal portions: two plated for serving now and prepared accordingly, and two placed into labeled shallow containers for storage.", "Cooking of the chicken finished at 6:00 PM.", "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 8:45 PM."], "request": "Is this meal at Level 3 readiness, allowing serving now and routing the remaining portions to cooling?"}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all original policy elements (recipe checks, probe-override rule) and the question's Level 3 timing rule remains external and unchanged, so no policy is lost or invented; entities, paths, and cooking time (6:00 PM) are preserved across both contexts; the two focus evidence items are complete factual sentences about a finish time and a cooling start time, not policy statements; shifting the cooling start from 7:30 PM to 9:15 PM is a coherent single-fact change that does not contradict any other evidence; neither context states or implies the readiness level or answer, only raw factual details.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\": \"Home cook Mara made paprika chicken and rice. She simmered the sauce fully before starting the chicken, then cooked the chicken until the probe gave two qualifying readings of at least 74°C. Inez, the meal planner, verified the probe's calibration today. Mara divided the finished meal into four equal portions: two set aside for serving now, two designated for storage. Dev, the cleanup helper, prepared the two serving portions for serving and moved the two storage portions toward cooling.\", \"evidence\": [\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\", \"Mara recorded 75°C and 74°C in the thickest pieces.\", \"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\", \"Mara finished cooking the chicken at 6:00 PM.\", \"The two storage-designated portions began compliant cooling at 7:30 PM.\", \"Dev placed two equal portions for serving and prepared them, and set the other two portions in labeled shallow containers.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mara finished cooking the chicken at 6:00 PM."}, {"path": ["evidence", "4"], "text": "The two storage-designated portions began compliant cooling at 7:30 PM."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara finished cooking the chicken at 6:00 PM.", "negative_left": "Mara finished cooking the chicken at 6:00 PM.", "negative_right": "The two storage-designated portions began compliant cooling at 9:15 PM.", "right": "The two storage-designated portions began compliant cooling at 7:30 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-190-024", "id": "fast-43-diverse-190-024-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara made paprika chicken and rice. She simmered the sauce fully before starting the chicken, then cooked the chicken until the probe gave two qualifying readings of at least 74°C. Inez, the meal planner, verified the probe's calibration today. Mara divided the finished meal into four equal portions: two set aside for serving now, two designated for storage. Dev, the cleanup helper, prepared the two serving portions for serving and moved the two storage portions toward cooling.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Mara recorded 75°C and 74°C in the thickest pieces.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara finished cooking the chicken at 6:00 PM.", "The two storage-designated portions began compliant cooling at 7:30 PM.", "Dev placed two equal portions for serving and prepared them, and set the other two portions in labeled shallow containers."]}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all original policy elements (recipe checks, probe-override rule) and the question's Level 3 timing rule remains external and unchanged, so no policy is lost or invented; entities, paths, and cooking time (6:00 PM) are preserved across both contexts; the two focus evidence items are complete factual sentences about a finish time and a cooling start time, not policy statements; shifting the cooling start from 7:30 PM to 9:15 PM is a coherent single-fact change that does not contradict any other evidence; neither context states or implies the readiness level or answer, only raw factual details.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\": \"Home cook Mara made paprika chicken and rice. She simmered the sauce fully before starting the chicken, then cooked the chicken until the probe gave two qualifying readings of at least 74°C. Inez, the meal planner, verified the probe's calibration today. Mara divided the finished meal into four equal portions: two set aside for serving now, two designated for storage. Dev, the cleanup helper, prepared the two serving portions for serving and moved the two storage portions toward cooling.\", \"evidence\": [\"Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.\", \"Mara recorded 75°C and 74°C in the thickest pieces.\", \"Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.\", \"Mara finished cooking the chicken at 6:00 PM.\", \"The two storage-designated portions began compliant cooling at 7:30 PM.\", \"Dev placed two equal portions for serving and prepared them, and set the other two portions in labeled shallow containers.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["evidence", "3"], "text": "Mara finished cooking the chicken at 6:00 PM."}, {"path": ["evidence", "4"], "text": "The two storage-designated portions began compliant cooling at 7:30 PM."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Mara finished cooking the chicken at 6:00 PM.", "negative_left": "Mara finished cooking the chicken at 6:00 PM.", "negative_right": "The two storage-designated portions began compliant cooling at 9:15 PM.", "right": "The two storage-designated portions began compliant cooling at 7:30 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-190-024", "id": "fast-43-diverse-190-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara made paprika chicken and rice. She simmered the sauce fully before starting the chicken, then cooked the chicken until the probe gave two qualifying readings of at least 74°C. Inez, the meal planner, verified the probe's calibration today. Mara divided the finished meal into four equal portions: two set aside for serving now, two designated for storage. Dev, the cleanup helper, prepared the two serving portions for serving and moved the two storage portions toward cooling.", "evidence": ["Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions.", "Mara recorded 75°C and 74°C in the thickest pieces.", "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color.", "Mara finished cooking the chicken at 6:00 PM.", "The two storage-designated portions began compliant cooling at 9:15 PM.", "Dev placed two equal portions for serving and prepared them, and set the other two portions in labeled shallow containers."]}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (probe-over-color rule, calibration, two-hour cooling window remains in the unchanged question) and preserve entity/time bindings except the deliberately varied storage time; the two evidence spans are complete factual sentences with no policy text; the counterfactual only changes the 7:30 PM storage time to 9:15 PM, which is internally consistent and does not contradict other facts; neither context reveals a gold answer or rubric level.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\": \"Home cook Mara made paprika chicken and rice. She first simmered the sauce to completion, then cooked the chicken. Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74\\u00b0C; divide into four equal portions. The probe used gave two readings, each at least 74\\u00b0C. Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color. Mara then divided the finished meal into four equal portions. Two of these portions were designated for serving now and were fully prepared for serving. The other two portions were designated for storage. Cooking of the chicken finished at 6:00 PM. The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["context"], "text": "Cooking of the chicken finished at 6:00 PM."}, {"path": ["context"], "text": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Cooking of the chicken finished at 6:00 PM.", "negative_left": "Cooking of the chicken finished at 6:00 PM.", "negative_right": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 9:15 PM.", "right": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-190-031", "id": "fast-43-diverse-190-031-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara made paprika chicken and rice. She first simmered the sauce to completion, then cooked the chicken. Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions. The probe used gave two readings, each at least 74°C. Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color. Mara then divided the finished meal into four equal portions. Two of these portions were designated for serving now and were fully prepared for serving. The other two portions were designated for storage. Cooking of the chicken finished at 6:00 PM. The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (probe-over-color rule, calibration, two-hour cooling window remains in the unchanged question) and preserve entity/time bindings except the deliberately varied storage time; the two evidence spans are complete factual sentences with no policy text; the counterfactual only changes the 7:30 PM storage time to 9:15 PM, which is internally consistent and does not contradict other facts; neither context reveals a gold answer or rubric level.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a11": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a11": "unknown"}, "remove_right": {"a11": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a11": "unknown"}, "negative_pair": {"a11": "refuted"}, "negative_sentence": {"a11": "unknown"}, "positive_pair": {"a11": "supported"}, "right": {"a11": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship or a permitted universally quantified relationship over an explicit set. The focus atom is a factual timing condition, not a policy classification. The base and counter assignments are jointly realizable: compliant cooling can begin either within two hours or after two hours while all other facts remain fixed. Policy evidence correctly preserves the state-originating recipe requirements and the probe-over-color conflict rule; requirements already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the recipe sequence, calibrated-probe requirement, two qualifying readings, four equal portions, two prepared serving portions, and two storage portions beginning compliant cooling within two hours. These conditions are sufficient for the Level 3 outcome. No exclusion for pink appearance is needed because the preserved kitchen rule makes valid probe readings controlling.", "rule_index": 0, "sound": true}, {"reason": "Refutation of a11 entails that compliant cooling did not begin within the two-hour limit. That leaves the explicit timely-storage requirement unmet and is sufficient for the false/No outcome, regardless of whether the appropriate non-Level-3 classification is Level 1 or Level 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Mara completed simmering the sauce before beginning to cook the chicken."}, {"id": "a2", "statement": "The probe used for the two chicken temperature readings passed its ice-water calibration on the day Mara cooked the meal."}, {"id": "a3", "statement": "The probe used on the chicken produced two temperature readings."}, {"id": "a4", "statement": "Each of the two chicken temperature readings was at least 74°C."}, {"id": "a5", "statement": "Mara divided the cooked meal into exactly four portions."}, {"id": "a6", "statement": "Each of the four meal portions had the same amount of food."}, {"id": "a7", "statement": "Exactly two of the four meal portions were designated for serving now."}, {"id": "a8", "statement": "Exactly two of the four meal portions were designated for storage."}, {"id": "a9", "statement": "Each of the two portions designated for serving now was prepared for serving."}, {"id": "a10", "statement": "Each of the two portions designated for storage began compliant cooling."}, {"id": "a11", "statement": "The elapsed time from the end of cooking to the start of compliant cooling for the two designated storage portions was no more than two hours."}], "base_state_json": "{\"context\": \"Home cook Mara made paprika chicken and rice. She first simmered the sauce to completion, then cooked the chicken. Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74\\u00b0C; divide into four equal portions. The probe used gave two readings, each at least 74\\u00b0C. Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color. Mara then divided the finished meal into four equal portions. Two of these portions were designated for serving now and were fully prepared for serving. The other two portions were designated for storage. Cooking of the chicken finished at 6:00 PM. The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "refuted"}], "focus_atom": "a11", "focus_evidence": [{"path": ["context"], "text": "Cooking of the chicken finished at 6:00 PM."}, {"path": ["context"], "text": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions."}, {"path": ["evidence", "2"], "text": "Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color."}], "rules": [{"justification": "The sauce and chicken were cooked in the required sequence; the calibrated probe supplied two qualifying readings; the meal was divided into four equal portions with two prepared for serving; and the other two portions began compliant cooling within two hours.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}]}, {"justification": "Compliant cooling for the designated storage portions did not begin within two hours, so the explicit timely-storage requirement remains unmet.", "target": "false", "when": [{"atom_id": "a11", "state": "refuted"}]}]}, "verified_pair": {"left": "Cooking of the chicken finished at 6:00 PM.", "negative_left": "Cooking of the chicken finished at 6:00 PM.", "negative_right": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 9:15 PM.", "right": "The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 7:30 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-190-031", "id": "fast-43-diverse-190-031-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 1 or Level 2 because a safety, cooking, portioning, or timely-storage requirement remains unmet.", "true": "Yes — assign Level 3: serve the two designated portions now and continue cooling the two storage portions."}, "instructions": "Answer the yes/no readiness question using this ordered rubric: Level 1 = unsafe or unusable, so stop and discard; Level 2 = not yet ready, so route to a corrective cooking, portioning, or storage step; Level 3 = recipe cooking checks passed, serving portions are prepared, and storage portions have begun compliant cooling within two hours. Apply the stated kitchen rule when visual and thermometer evidence conflict.", "type": "noul"}}, "state": {"context": "Home cook Mara made paprika chicken and rice. She first simmered the sauce to completion, then cooked the chicken. Recipe: simmer sauce, then cook chicken until a calibrated probe gives two readings of at least 74°C; divide into four equal portions. The probe used gave two readings, each at least 74°C. Inez confirmed the probe passed its ice-water calibration today. The kitchen rule says valid probe readings override meat color. Mara then divided the finished meal into four equal portions. Two of these portions were designated for serving now and were fully prepared for serving. The other two portions were designated for storage. Cooking of the chicken finished at 6:00 PM. The two portions designated for storage were placed into the refrigerator's compliant cooling cycle at 9:15 PM."}}, "method": "c2d", "provenance": {"source_id": "diverse-190", "source_is_synthetic": true, "source_sha256": "943473f40b2e0ce2862b95e4b730697a5d68baf82ff35a6bb567ec852a962b92", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (temperature rule, refrigeration window, single-event route) unchanged from the original state and question; only the logged time of event M shifts from 6:42 PM to 6:58 PM, which is a plain factual change not a policy alteration; the focus evidence consists of two complete factual log sentences with no policy language; the counterfactual remains internally consistent since the cook's recognition claim is about mental timing, not log order, so no contradictory duplicate measurement arises; neither context contains a gold answer, rule table, or instruction leaking the intended score.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 6:42 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:42 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:42 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 6:58 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-007", "id": "fast-43-diverse-191-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:42 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (temperature rule, refrigeration window, single-event route) unchanged from the original state and question; only the logged time of event M shifts from 6:42 PM to 6:58 PM, which is a plain factual change not a policy alteration; the focus evidence consists of two complete factual log sentences with no policy language; the counterfactual remains internally consistent since the cook's recognition claim is about mental timing, not log order, so no contradictory duplicate measurement arises; neither context contains a gold answer, rule table, or instruction leaking the intended score.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 6:42 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:42 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:42 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 6:58 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-007", "id": "fast-43-diverse-191-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:58 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:50 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original recipe, temperature rule, and refrigeration policy unchanged alongside the verbatim question; the focus evidence consists of two complete factual log sentences with no policy text; the counterfactual only shifts event M's logged time from 7:05 PM to 7:20 PM, which remains coherent since no other passage fixes the relative order of M and B; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I noticed before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-055", "id": "fast-43-diverse-191-055-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I noticed before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original recipe, temperature rule, and refrigeration policy unchanged alongside the verbatim question; the focus evidence consists of two complete factual log sentences with no policy text; the counterfactual only shifts event M's logged time from 7:05 PM to 7:20 PM, which remains coherent since no other passage fixes the relative order of M and B; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I noticed before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-055", "id": "fast-43-diverse-191-055-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I noticed before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:20 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the recipe, rule, and event definitions unchanged, only altering the logged time of event M from 6:40 PM to 6:55 PM, which is a coherent factual variation; the two evidence sentences are plain factual log entries, not policy text; no answer, code, or rationale is embedded, and all original question entities (chicken, event M, event B) remain bound correctly.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 6:40 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:40 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 6:55 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-058", "id": "fast-43-diverse-191-058-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the recipe, rule, and event definitions unchanged, only altering the logged time of event M from 6:40 PM to 6:55 PM, which is a coherent factual variation; the two evidence sentences are plain factual log entries, not policy text; no answer, code, or rationale is embedded, and all original question entities (chicken, event M, event B) remain bound correctly.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 6:40 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:40 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 6:55 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-058", "id": "fast-43-diverse-191-058-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:55 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing recipe/kitchen rule and the unchanged questions object fully preserves scoring policy; entities, event labels (M, B) and question focus are unchanged. The two focus evidence spans are plain factual log sentences with no rule text or IDs. The counterfactual only shifts M's logged time from 7:14 PM to 7:26 PM, reversing its order relative to B (7:20 PM) without creating duplicate or contradictory measurements, and remains consistent with the unchanged recognition statement about timing relative to the check and portioning. No answer, code, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading for the lemon chicken was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the required reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the internal temperature of the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking of the chicken if event M records below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was recorded in the kitchen log at 7:14 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was recorded in the kitchen log at 7:14 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was recorded in the kitchen log at 7:14 PM.", "negative_left": "Thermometer event M was recorded in the kitchen log at 7:26 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-059", "id": "fast-43-diverse-191-059-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading for the lemon chicken was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the required reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the internal temperature of the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking of the chicken if event M records below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was recorded in the kitchen log at 7:14 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing recipe/kitchen rule and the unchanged questions object fully preserves scoring policy; entities, event labels (M, B) and question focus are unchanged. The two focus evidence spans are plain factual log sentences with no rule text or IDs. The counterfactual only shifts M's logged time from 7:14 PM to 7:26 PM, reversing its order relative to B (7:20 PM) without creating duplicate or contradictory measurements, and remains consistent with the unchanged recognition statement about timing relative to the check and portioning. No answer, code, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading for the lemon chicken was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the required reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the internal temperature of the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking of the chicken if event M records below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was recorded in the kitchen log at 7:14 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was recorded in the kitchen log at 7:14 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was recorded in the kitchen log at 7:14 PM.", "negative_left": "Thermometer event M was recorded in the kitchen log at 7:26 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-059", "id": "fast-43-diverse-191-059-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading for the lemon chicken was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the required reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the internal temperature of the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking of the chicken if event M records below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was recorded in the kitchen log at 7:26 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:20 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the recipe, temperature rule, refrigeration rule, and event descriptions intact, matching the unchanged question; the focus evidence consists of two plain factual timestamp sentences; the counterfactual merely reorders M and B without contradicting any other stated fact, and no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. Before either the thermometer check or the first portioning, I recognized that the required reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was time-stamped by the kitchen log at 7:15 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was time-stamped by the kitchen log at 7:15 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was time-stamped by the kitchen log at 7:15 PM.", "negative_left": "Thermometer event M was time-stamped by the kitchen log at 7:25 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-061", "id": "fast-43-diverse-191-061-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. Before either the thermometer check or the first portioning, I recognized that the required reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was time-stamped by the kitchen log at 7:15 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the recipe, temperature rule, refrigeration rule, and event descriptions intact, matching the unchanged question; the focus evidence consists of two plain factual timestamp sentences; the counterfactual merely reorders M and B without contradicting any other stated fact, and no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. Before either the thermometer check or the first portioning, I recognized that the required reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was time-stamped by the kitchen log at 7:15 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was time-stamped by the kitchen log at 7:15 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was time-stamped by the kitchen log at 7:15 PM.", "negative_left": "Thermometer event M was time-stamped by the kitchen log at 7:25 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-061", "id": "fast-43-diverse-191-061-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. Before either the thermometer check or the first portioning, I recognized that the required reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was time-stamped by the kitchen log at 7:25 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 7:20 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy, entities, and question object; the counterfactual only shifts M's log time to 7:18 PM (after B's 7:12 PM), which reorders logged events without contradicting any other stated fact, and the two focus sentences are plain factual log entries free of rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I realized, before either the thermometer check or the first portioning, that the required reading was still missing.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C, well above the 74°C threshold. The route schedules exactly one thermometer event, M, and requires further cooking only if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:18 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-062", "id": "fast-43-diverse-191-062-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I realized, before either the thermometer check or the first portioning, that the required reading was still missing."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C, well above the 74°C threshold. The route schedules exactly one thermometer event, M, and requires further cooking only if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy, entities, and question object; the counterfactual only shifts M's log time to 7:18 PM (after B's 7:12 PM), which reorders logged events without contradicting any other stated fact, and the two focus sentences are plain factual log entries free of rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I realized, before either the thermometer check or the first portioning, that the required reading was still missing.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C, well above the 74°C threshold. The route schedules exactly one thermometer event, M, and requires further cooking only if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:18 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-062", "id": "fast-43-diverse-191-062-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I realized, before either the thermometer check or the first portioning, that the required reading was still missing."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C, well above the 74°C threshold. The route schedules exactly one thermometer event, M, and requires further cooking only if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:18 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the recipe, temperature rule, and refrigeration policy intact alongside the unchanged questions; the focus evidence sentences are plain factual log entries, not policy or instructions; the counterfactual only changes M's logged time from 6:37 PM to 6:52 PM, a single coherent factual edit that does not contradict any other unchanged statement; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning happened, that the reading was still missing from our records.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The selected route schedules exactly one thermometer event, M, and requires further cooking if the reading is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 6:37 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:37 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:37 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 6:52 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-069", "id": "fast-43-diverse-191-069-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning happened, that the reading was still missing from our records."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The selected route schedules exactly one thermometer event, M, and requires further cooking if the reading is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:37 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the recipe, temperature rule, and refrigeration policy intact alongside the unchanged questions; the focus evidence sentences are plain factual log entries, not policy or instructions; the counterfactual only changes M's logged time from 6:37 PM to 6:52 PM, a single coherent factual edit that does not contradict any other unchanged statement; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning happened, that the reading was still missing from our records.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The selected route schedules exactly one thermometer event, M, and requires further cooking if the reading is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 6:37 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:37 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:37 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 6:52 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-069", "id": "fast-43-diverse-191-069-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning happened, that the reading was still missing from our records."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The selected route schedules exactly one thermometer event, M, and requires further cooking if the reading is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:52 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:45 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only changes the logged time of thermometer event M from 7:05 PM to 7:19 PM, a plausible administrative delay that does not contradict the fixed portioning time, temperature value, or any stated policy, and all governing rules and question bindings remain intact with no leaked answer content.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the required reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was recorded in the kitchen log at 7:05 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was recorded in the kitchen log at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was recorded in the kitchen log at 7:05 PM.", "negative_left": "Thermometer event M was recorded in the kitchen log at 7:19 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-072", "id": "fast-43-diverse-191-072-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the required reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was recorded in the kitchen log at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only changes the logged time of thermometer event M from 7:05 PM to 7:19 PM, a plausible administrative delay that does not contradict the fixed portioning time, temperature value, or any stated policy, and all governing rules and question bindings remain intact with no leaked answer content.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the required reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was recorded in the kitchen log at 7:05 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was recorded in the kitchen log at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was recorded in the kitchen log at 7:05 PM.", "negative_left": "Thermometer event M was recorded in the kitchen log at 7:19 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-072", "id": "fast-43-diverse-191-072-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the required reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was recorded in the kitchen log at 7:19 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was recorded in the kitchen log at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the recipe, temperature rule, and refrigeration policy verbatim while only changing the M log time; the two focus sentences are plain factual log entries, not policy text or answers; the reordering to 7:20 PM after the 7:14 PM portioning is a coherent alternate scenario not contradicted by any other stated fact; entity/time bindings (event M, event B, four dinners) are unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-074", "id": "fast-43-diverse-191-074-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the recipe, temperature rule, and refrigeration policy verbatim while only changing the M log time; the two focus sentences are plain factual log entries, not policy text or answers; the reordering to 7:20 PM after the 7:14 PM portioning is a coherent alternate scenario not contradicted by any other stated fact; entity/time bindings (event M, event B, four dinners) are unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-074", "id": "fast-43-diverse-191-074-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized, before either the thermometer check or the first portioning, that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:20 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:14 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full recipe, temperature, and refrigeration policy unchanged from the original state and question; only the timestamp of event M is altered (7:05 PM to 7:20 PM), which reorders it after event B but does not contradict any explicit stated fact, since the home cook's remark only concerns recognition timing, not measurement order; evidence spans are two complete factual timestamp sentences with no embedded rules, IDs, or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before I checked the thermometer. I already recognized, before either the thermometer check or the first portioning, that the reading was still missing from our log.\"},{\"speaker\":\"Home cook\",\"text\":\"The thermometer check measured the thickest piece of the lemon chicken and read 76°C, above the 74°C threshold. Our route only schedules that one thermometer check, and it calls for further cooking if the reading falls below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. Our route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was time-stamped in the kitchen log at 7:05 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was time-stamped in the kitchen log at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was time-stamped in the kitchen log at 7:05 PM.", "negative_left": "Thermometer event M was time-stamped in the kitchen log at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-080", "id": "fast-43-diverse-191-080-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before I checked the thermometer. I already recognized, before either the thermometer check or the first portioning, that the reading was still missing from our log."}, {"speaker": "Home cook", "text": "The thermometer check measured the thickest piece of the lemon chicken and read 76°C, above the 74°C threshold. Our route only schedules that one thermometer check, and it calls for further cooking if the reading falls below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. Our route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was time-stamped in the kitchen log at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full recipe, temperature, and refrigeration policy unchanged from the original state and question; only the timestamp of event M is altered (7:05 PM to 7:20 PM), which reorders it after event B but does not contradict any explicit stated fact, since the home cook's remark only concerns recognition timing, not measurement order; evidence spans are two complete factual timestamp sentences with no embedded rules, IDs, or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before I checked the thermometer. I already recognized, before either the thermometer check or the first portioning, that the reading was still missing from our log.\"},{\"speaker\":\"Home cook\",\"text\":\"The thermometer check measured the thickest piece of the lemon chicken and read 76°C, above the 74°C threshold. Our route only schedules that one thermometer check, and it calls for further cooking if the reading falls below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. Our route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was time-stamped in the kitchen log at 7:05 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was time-stamped in the kitchen log at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was time-stamped in the kitchen log at 7:05 PM.", "negative_left": "Thermometer event M was time-stamped in the kitchen log at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-080", "id": "fast-43-diverse-191-080-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before I checked the thermometer. I already recognized, before either the thermometer check or the first portioning, that the reading was still missing from our log."}, {"speaker": "Home cook", "text": "The thermometer check measured the thickest piece of the lemon chicken and read 76°C, above the 74°C threshold. Our route only schedules that one thermometer check, and it calls for further cooking if the reading falls below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. Our route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was time-stamped in the kitchen log at 7:20 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped in the kitchen log at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original recipe, temperature rule, and refrigeration policy; the counterfactual only changes the logged time of event M to 7:20 PM, reversing the order relative to event B, which is a coherent single-fact change without contradicting other stated facts, and the two focus evidence lines are plain factual log statements with no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-083", "id": "fast-43-diverse-191-083-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original recipe, temperature rule, and refrigeration policy; the counterfactual only changes the logged time of event M to 7:20 PM, reversing the order relative to event B, which is a coherent single-fact change without contradicting other stated facts, and the two focus evidence lines are plain factual log statements with no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-083", "id": "fast-43-diverse-191-083-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:20 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the recipe/temperature rule and refrigeration policy verbatim alongside the unchanged question; the two focus evidence spans are plain factual timestamp statements, not policy text; swapping M's timestamp from 6:41 PM to 6:53 PM merely reverses the order of M and B without duplicating or contradicting any other measurement, and neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. Before either the thermometer check or the first portioning of the four dinners, I already recognized the reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was time-stamped by the kitchen log at 6:41 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was time-stamped by the kitchen log at 6:41 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was time-stamped by the kitchen log at 6:41 PM.", "negative_left": "Thermometer event M was time-stamped by the kitchen log at 6:53 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-087", "id": "fast-43-diverse-191-087-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. Before either the thermometer check or the first portioning of the four dinners, I already recognized the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was time-stamped by the kitchen log at 6:41 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the recipe/temperature rule and refrigeration policy verbatim alongside the unchanged question; the two focus evidence spans are plain factual timestamp statements, not policy text; swapping M's timestamp from 6:41 PM to 6:53 PM merely reverses the order of M and B without duplicating or contradicting any other measurement, and neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. Before either the thermometer check or the first portioning of the four dinners, I already recognized the reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was time-stamped by the kitchen log at 6:41 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was time-stamped by the kitchen log at 6:41 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was time-stamped by the kitchen log at 6:41 PM.", "negative_left": "Thermometer event M was time-stamped by the kitchen log at 6:53 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-087", "id": "fast-43-diverse-191-087-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. Before either the thermometer check or the first portioning of the four dinners, I already recognized the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was time-stamped by the kitchen log at 6:53 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was time-stamped by the kitchen log at 6:47 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original recipe rule and questions verbatim; the counterfactual only changes the logged time of event M (7:05 PM vs 6:40 PM), reversing the order relative to event B (6:55 PM) in a plausible, non-contradictory way, and both focus_evidence sentences are complete factual log statements with no embedded answer, rule table, or instruction leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading for the chicken has been documented or announced yet. Before either the thermometer check or the first portioning happened, I already realized the reading was still missing.\"},{\"speaker\":\"Home cook\",\"text\":\"I inserted the thermometer into the thickest piece of the lemon chicken for event M, and it read 76°C. Our route only calls for this single thermometer check, and it would require more baking if the reading came in under 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 6:40 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:40 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-097", "id": "fast-43-diverse-191-097-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading for the chicken has been documented or announced yet. Before either the thermometer check or the first portioning happened, I already realized the reading was still missing."}, {"speaker": "Home cook", "text": "I inserted the thermometer into the thickest piece of the lemon chicken for event M, and it read 76°C. Our route only calls for this single thermometer check, and it would require more baking if the reading came in under 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original recipe rule and questions verbatim; the counterfactual only changes the logged time of event M (7:05 PM vs 6:40 PM), reversing the order relative to event B (6:55 PM) in a plausible, non-contradictory way, and both focus_evidence sentences are complete factual log statements with no embedded answer, rule table, or instruction leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading for the chicken has been documented or announced yet. Before either the thermometer check or the first portioning happened, I already realized the reading was still missing.\"},{\"speaker\":\"Home cook\",\"text\":\"I inserted the thermometer into the thickest piece of the lemon chicken for event M, and it read 76°C. Our route only calls for this single thermometer check, and it would require more baking if the reading came in under 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 6:40 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:40 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-097", "id": "fast-43-diverse-191-097-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading for the chicken has been documented or announced yet. Before either the thermometer check or the first portioning happened, I already realized the reading was still missing."}, {"speaker": "Home cook", "text": "I inserted the thermometer into the thickest piece of the lemon chicken for event M, and it read 76°C. Our route only calls for this single thermometer check, and it would require more baking if the reading came in under 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the recipe, temperature and refrigeration rules verbatim and the unchanged questions object carries all scoring criteria; the counterfactual only swaps the thermometer log time from 7:05 PM to 7:20 PM, keeping entity/path/time bindings for events M and B otherwise intact and yielding a plausible alternate ordering (M logged after B) without contradicting any other stated fact; the focus evidence in both contexts consists of two complete factual log sentences with no policy text, rule tables, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I realized before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-098", "id": "fast-43-diverse-191-098-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I realized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the recipe, temperature and refrigeration rules verbatim and the unchanged questions object carries all scoring criteria; the counterfactual only swaps the thermometer log time from 7:05 PM to 7:20 PM, keeping entity/path/time bindings for events M and B otherwise intact and yielding a plausible alternate ordering (M logged after B) without contradicting any other stated fact; the focus evidence in both contexts consists of two complete factual log sentences with no policy text, rule tables, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\": \"Meal planner\", \"text\": \"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"}, {\"speaker\": \"Home cook\", \"text\": \"No internal-temperature reading was documented or announced before the thermometer check. I realized before either the thermometer check or the first portioning that the reading was still undocumented.\"}, {\"speaker\": \"Home cook\", \"text\": \"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"}, {\"speaker\": \"Cleanup helper\", \"text\": \"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Thermometer event M was logged in the kitchen notebook at 7:05 PM.\"}, {\"speaker\": \"Kitchen log\", \"text\": \"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:20 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-098", "id": "fast-43-diverse-191-098-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I realized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:20 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 7:12 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the recipe, temperature rule, and refrigeration policy intact via the unchanged question and identical narrative, only the M-event timestamp changes (6:35→6:53 PM), which is a plausible case-observation shift without contradicting other stated facts, and the evidence sentences are complete factual log entries with no embedded answer or rule text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading has been documented or announced yet. Before either checking the thermometer or portioning the dinners, I already recognized that the required reading was still missing.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 6:35 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:35 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:35 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 6:53 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-099", "id": "fast-43-diverse-191-099-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading has been documented or announced yet. Before either checking the thermometer or portioning the dinners, I already recognized that the required reading was still missing."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:35 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the recipe, temperature rule, and refrigeration policy intact via the unchanged question and identical narrative, only the M-event timestamp changes (6:35→6:53 PM), which is a plausible case-observation shift without contradicting other stated facts, and the evidence sentences are complete factual log entries with no embedded answer or rule text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading has been documented or announced yet. Before either checking the thermometer or portioning the dinners, I already recognized that the required reading was still missing.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 6:35 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:35 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:35 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 6:53 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-099", "id": "fast-43-diverse-191-099-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading has been documented or announced yet. Before either checking the thermometer or portioning the dinners, I already recognized that the required reading was still missing."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:53 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:47 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full recipe rule, refrigeration policy, and single-thermometer-event routing rule from the original state; only the logged time of event M changes (6:40 PM to 7:05 PM), a plain factual edit that coherently reverses the order of M relative to portioning event B without contradicting any other stated fact; both focus_evidence spans are complete factual log-time sentences, not rules or instructions; no answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 6:40 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:40 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-108", "id": "fast-43-diverse-191-108-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full recipe rule, refrigeration policy, and single-thermometer-event routing rule from the original state; only the logged time of event M changes (6:40 PM to 7:05 PM), a plain factual edit that coherently reverses the order of M relative to portioning event B without contradicting any other stated fact; both focus_evidence spans are complete factual log-time sentences, not rules or instructions; no answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was logged in the kitchen notebook at 6:40 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was logged in the kitchen notebook at 6:40 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was logged in the kitchen notebook at 6:40 PM.", "negative_left": "Thermometer event M was logged in the kitchen notebook at 7:05 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-108", "id": "fast-43-diverse-191-108-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I recognized before either the thermometer check or the first portioning that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded at least 74°C. The route schedules exactly one thermometer event, M, and requires further cooking if the temperature is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was logged in the kitchen notebook at 7:05 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was logged in the kitchen notebook at 6:55 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Only the M timestamp changed (7:15→7:25) while B stays at 7:20, a permissible factual variation that still forms a coherent scenario without contradicting any policy, and both contexts keep the governing rules, entities, and evidence sentences intact with no embedded answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I realized, before either the thermometer check or the first portioning happened, that the reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking only if the reading is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was timestamped in the kitchen log at 7:15 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was timestamped in the kitchen log at 7:15 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was timestamped in the kitchen log at 7:15 PM.", "negative_left": "Thermometer event M was timestamped in the kitchen log at 7:25 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-117", "id": "fast-43-diverse-191-117-base", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I realized, before either the thermometer check or the first portioning happened, that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking only if the reading is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was timestamped in the kitchen log at 7:15 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "home-02", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Only the M timestamp changed (7:15→7:25) while B stays at 7:20, a permissible factual variation that still forms a coherent scenario without contradicting any policy, and both contexts keep the governing rules, entities, and evidence sentences intact with no embedded answer or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms express individual factual or route relationships rather than final score classifications. The focus atom is the factual temporal ordering between thermometer event M and handling event B. The base and counter assignments are jointly realizable: M can occur before B in the base and B before M in the counter while all other atom states remain fixed. Policy evidence correctly preserves the temperature/substitution rule and refrigeration deadline originating in the original state; decision criteria in the unchanged questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that the thickest piece is measured, a compliant reading is obtained before the first serving or portioning event, further cooking is required on a low reading, and unserved portions are refrigerated within the required window. These facts are sufficient for score 3.", "rule_index": 0, "sound": true}, {"reason": "The cook recognizes the missing reading before portioning. Refuting that M is earlier than B, together with their distinct recorded times, places the first portioning event before M. M is the route's sole thermometer event and measures the thickest piece, so this entails beginning portioning while planning the later check, which is score 2.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Before thermometer event M, no internal-temperature reading for the lemon chicken is documented or announced."}, {"id": "a2", "statement": "Before the earlier of thermometer event M and chicken-handling event B, the home cook recognizes that the required internal-temperature reading is undocumented."}, {"id": "a3", "statement": "Thermometer event M measures the internal temperature of the thickest piece of the lemon chicken."}, {"id": "a4", "statement": "Thermometer event M records an internal temperature of at least 74°C."}, {"id": "a5", "statement": "Chicken-handling event B is the first serving or portioning of the lemon chicken into the four dinners."}, {"id": "a6", "statement": "The recorded time of thermometer event M is earlier than the recorded time of chicken-handling event B."}, {"id": "a7", "statement": "Thermometer event M and chicken-handling event B have different recorded times."}, {"id": "a8", "statement": "For the actual execution, the selected route schedules exactly one thermometer event for the lemon chicken, event M."}, {"id": "a9", "statement": "The selected route requires further cooking of the lemon chicken if event M records an internal temperature below 74°C."}, {"id": "a10", "statement": "The selected route refrigerates every unserved portion within the required two-hour window."}], "base_state_json": "[{\"speaker\":\"Meal planner\",\"text\":\"Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading.\"},{\"speaker\":\"Home cook\",\"text\":\"No internal-temperature reading was documented or announced before the thermometer check. I realized, before either the thermometer check or the first portioning happened, that the reading was still undocumented.\"},{\"speaker\":\"Home cook\",\"text\":\"Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking only if the reading is below 74°C.\"},{\"speaker\":\"Cleanup helper\",\"text\":\"Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Thermometer event M was timestamped in the kitchen log at 7:15 PM.\"},{\"speaker\":\"Kitchen log\",\"text\":\"Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Thermometer event M was timestamped in the kitchen log at 7:15 PM."}, {"path": ["5", "text"], "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM."}], "policy_evidence": [{"path": ["0", "text"], "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"path": ["3", "text"], "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned."}], "rules": [{"justification": "The cook recognizes the missing evidence, obtains a compliant reading from the thickest piece before the first serving or portioning, retains the required further-cooking contingency for a low result, and routes unserved portions to timely refrigeration.", "target": "3", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}, {"justification": "Because M and B have different times, refuting that M is earlier establishes that the first serving or portioning begins before M. The cook already recognizes that the reading is undocumented, and the selected route schedules M as its thermometer check, so portioning begins while the cook is planning the later check.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}]}]}, "verified_pair": {"left": "Thermometer event M was timestamped in the kitchen log at 7:15 PM.", "negative_left": "Thermometer event M was timestamped in the kitchen log at 7:25 PM.", "negative_right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM.", "right": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM."}, "verifier_independent_model": false}, "family": "fast-43-diverse-191-117", "id": "fast-43-diverse-191-117-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Serve and portion immediately, treating baking time and appearance as sufficient despite the explicit temperature rule.", "1 — Hold the chicken but discard it or restart the entire meal without first obtaining the missing temperature reading.", "2 — Notice that the temperature is undocumented, yet begin serving or portioning while planning to check only a later piece.", "3 — Do not serve or portion the chicken; measure the thickest piece now, cook further if it is below 74°C, and proceed to portioning and timely refrigeration only after a compliant reading."], "instructions": "Rate the best readiness and routing decision under the supplied recipe and kitchen rule. The ordered levels measure decision quality, including whether missing required evidence is handled correctly.", "type": "score"}}, "state": [{"speaker": "Meal planner", "text": "Tonight’s recipe is lemon chicken, rice, and broccoli. Bake chicken 22 minutes at 220°C, then verify at least 74°C in the thickest piece. Kitchen rule: elapsed time, color, and clear juices cannot replace that reading."}, {"speaker": "Home cook", "text": "No internal-temperature reading was documented or announced before the thermometer check. I realized, before either the thermometer check or the first portioning happened, that the reading was still undocumented."}, {"speaker": "Home cook", "text": "Thermometer event M measured the thickest piece of the lemon chicken and recorded 76°C. The route schedules exactly one thermometer event, M, and requires further cooking only if the reading is below 74°C."}, {"speaker": "Cleanup helper", "text": "Four clean containers are ready. Any unserved portions must be refrigerated within two hours, but the meal has not yet been portioned. The route refrigerates every unserved portion within that window."}, {"speaker": "Kitchen log", "text": "Thermometer event M was timestamped in the kitchen log at 7:25 PM."}, {"speaker": "Kitchen log", "text": "Chicken-handling event B, the first portioning of the lemon chicken into the four dinners, was timestamped in the kitchen log at 7:20 PM."}]}, "method": "c2d", "provenance": {"source_id": "diverse-191", "source_is_synthetic": true, "source_sha256": "b2a3412a9c6e76b46178db3172f66748c8a218c2e8c8de04e0dbc8e0a11d5179", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "home-02", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same rule and question, differing only in whether the image shows the panel reversed or not, with no embedded verdicts or rule tables.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"In a living-room workshop, a household assembler completed a fictional Alderline C-4 cabinet build. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.\",\"evidence\":[\"The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly.\",\"The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet in its original manufactured orientation, without any reversal.\",\"At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured, confirming both required straps are installed.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly."}, {"path": ["evidence", "1"], "text": "The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet in its original manufactured orientation, without any reversal."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly.", "negative_left": "The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly.", "negative_right": "The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet flipped from its original manufactured orientation, with the smooth side reversed.", "right": "The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet in its original manufactured orientation, without any reversal."}, "verifier_independent_model": false}, "family": "fast-43-diverse-195-015", "id": "fast-43-diverse-195-015-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "In a living-room workshop, a household assembler completed a fictional Alderline C-4 cabinet build. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.", "evidence": ["The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly.", "The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet in its original manufactured orientation, without any reversal.", "At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured, confirming both required straps are installed."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same rule and question, differing only in whether the image shows the panel reversed or not, with no embedded verdicts or rule tables.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"In a living-room workshop, a household assembler completed a fictional Alderline C-4 cabinet build. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.\",\"evidence\":[\"The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly.\",\"The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet in its original manufactured orientation, without any reversal.\",\"At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured, confirming both required straps are installed.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly."}, {"path": ["evidence", "1"], "text": "The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet in its original manufactured orientation, without any reversal."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly.", "negative_left": "The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly.", "negative_right": "The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet flipped from its original manufactured orientation, with the smooth side reversed.", "right": "The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet in its original manufactured orientation, without any reversal."}, "verifier_independent_model": false}, "family": "fast-43-diverse-195-015", "id": "fast-43-diverse-195-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "In a living-room workshop, a household assembler completed a fictional Alderline C-4 cabinet build. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.", "evidence": ["The back panel used in the Alderline C-4 cabinet is manufactured smooth on one side and textured on the other, with the smooth side designed to face inward when installed correctly.", "The final inspection photo taken at 3:42 PM shows the back panel mounted in the Alderline C-4 cabinet flipped from its original manufactured orientation, with the smooth side reversed.", "At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured, confirming both required straps are installed."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same inspection rule and conflict-resolution policy alongside the unchanged question; the counterfactual only swaps 'smooth' to 'rough' in the manufacturer spec sentence, a single coherent factual change that doesn't contradict the other evidence; both focus sentences are factual statements, not policy text or instructions; no gold answer, rule table, or output directive appears in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\": \"Case note for the fictional Alderline C-4 cabinet, household workshop final inspection. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.\", \"evidence\": [\"At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior.\", \"The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's smooth face.\", \"The 14:30 final-use test log confirms both required anti-tip straps are installed and secured to the wall bracket.\", \"The tester's 14:30 report states the cabinet did not rock when pushed from the front or top during the final-use test.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior."}, {"path": ["evidence", "1"], "text": "The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's smooth face."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior.", "negative_left": "At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior.", "negative_right": "The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's rough face.", "right": "The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's smooth face."}, "verifier_independent_model": false}, "family": "fast-43-diverse-195-017", "id": "fast-43-diverse-195-017-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "Case note for the fictional Alderline C-4 cabinet, household workshop final inspection. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.", "evidence": ["At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior.", "The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's smooth face.", "The 14:30 final-use test log confirms both required anti-tip straps are installed and secured to the wall bracket.", "The tester's 14:30 report states the cabinet did not rock when pushed from the front or top during the final-use test."]}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same inspection rule and conflict-resolution policy alongside the unchanged question; the counterfactual only swaps 'smooth' to 'rough' in the manufacturer spec sentence, a single coherent factual change that doesn't contradict the other evidence; both focus sentences are factual statements, not policy text or instructions; no gold answer, rule table, or output directive appears in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\": \"Case note for the fictional Alderline C-4 cabinet, household workshop final inspection. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.\", \"evidence\": [\"At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior.\", \"The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's smooth face.\", \"The 14:30 final-use test log confirms both required anti-tip straps are installed and secured to the wall bracket.\", \"The tester's 14:30 report states the cabinet did not rock when pushed from the front or top during the final-use test.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior."}, {"path": ["evidence", "1"], "text": "The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's smooth face."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior.", "negative_left": "At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior.", "negative_right": "The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's rough face.", "right": "The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's smooth face."}, "verifier_independent_model": false}, "family": "fast-43-diverse-195-017", "id": "fast-43-diverse-195-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "Case note for the fictional Alderline C-4 cabinet, household workshop final inspection. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls.", "evidence": ["At final inspection, the orientation arrow stamped on the back panel installed in the fictional Alderline C-4 cabinet faces the cabinet's interior.", "The manufacturer's specification for the fictional Alderline C-4 cabinet states that the stamped orientation arrow always points toward the panel's rough face.", "The 14:30 final-use test log confirms both required anti-tip straps are installed and secured to the wall bracket.", "The tester's 14:30 report states the cabinet did not rock when pushed from the front or top during the final-use test."]}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same pass rule and conflict policy from the original question/state, only replacing the conflicting note/image evidence with a direct factual orientation statement, which is a permissible observation change; the counterfactual coherently flips only the orientation fact (inward→outward) without contradicting the straps/no-rock evidence; both focus sentences are plain factual statements, and no answer, code, or rationale is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\": \"In a living-room workshop, a household assembler has finished a fictional Alderline C-4 cabinet. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119. At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces inward. At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured, confirming both required anti-tip straps are properly installed and the cabinet remains stable under load testing.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["context"], "text": "The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119."}, {"path": ["context"], "text": "At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces inward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119.", "negative_left": "The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119.", "negative_right": "At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces outward.", "right": "At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces inward."}, "verifier_independent_model": false}, "family": "fast-43-diverse-195-025", "id": "fast-43-diverse-195-025-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "In a living-room workshop, a household assembler has finished a fictional Alderline C-4 cabinet. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119. At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces inward. At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured, confirming both required anti-tip straps are properly installed and the cabinet remains stable under load testing."}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same pass rule and conflict policy from the original question/state, only replacing the conflicting note/image evidence with a direct factual orientation statement, which is a permissible observation change; the counterfactual coherently flips only the orientation fact (inward→outward) without contradicting the straps/no-rock evidence; both focus sentences are plain factual statements, and no answer, code, or rationale is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\": \"In a living-room workshop, a household assembler has finished a fictional Alderline C-4 cabinet. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119. At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces inward. At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured, confirming both required anti-tip straps are properly installed and the cabinet remains stable under load testing.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["context"], "text": "The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119."}, {"path": ["context"], "text": "At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces inward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119.", "negative_left": "The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119.", "negative_right": "At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces outward.", "right": "At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces inward."}, "verifier_independent_model": false}, "family": "fast-43-diverse-195-025", "id": "fast-43-diverse-195-025-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "In a living-room workshop, a household assembler has finished a fictional Alderline C-4 cabinet. The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing. When notes conflict with timestamped images, the later image controls. The back panel installed in the fictional Alderline C-4 cabinet has its smooth face on the side stamped with serial code B-119. At final inspection, the side stamped with serial code B-119 on the fictional Alderline C-4 cabinet's back panel faces outward. At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured, confirming both required anti-tip straps are properly installed and the cabinet remains stable under load testing."}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the pass rule and conflict policy intact, preserve the same cabinet entity and yes/no request, replace the note/image evidence with a factual sensor reading and its calibration mapping (two factual sentences), the counterfactual coherently swaps 180→0 degrees per the stated mapping without contradicting other evidence, and neither context reveals a gold answer or instructs the output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"In a living-room workshop, a household assembler completed a fictional Alderline C-4 cabinet. Inspection records include sensor logs and physical checks.\",\"evidence\":[\"The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing.\",\"When notes conflict with timestamped images, the later image controls.\",\"At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 180 degrees.\",\"For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward.\",\"At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "2"], "text": "At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 180 degrees."}, {"path": ["evidence", "3"], "text": "For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 180 degrees.", "negative_left": "At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 0 degrees.", "negative_right": "For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward.", "right": "For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward."}, "verifier_independent_model": false}, "family": "fast-43-diverse-195-032", "id": "fast-43-diverse-195-032-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "In a living-room workshop, a household assembler completed a fictional Alderline C-4 cabinet. Inspection records include sensor logs and physical checks.", "evidence": ["The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing.", "When notes conflict with timestamped images, the later image controls.", "At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 180 degrees.", "For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward.", "At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the pass rule and conflict policy intact, preserve the same cabinet entity and yes/no request, replace the note/image evidence with a factual sensor reading and its calibration mapping (two factual sentences), the counterfactual coherently swaps 180→0 degrees per the stated mapping without contradicting other evidence, and neither context reveals a gold answer or instructs the output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "supported", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "supported", "A4": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual condition rather than a policy conclusion or bundled classification. The focus atom A1 is a factual panel-orientation relation. The base and counter assignments are jointly realizable in separate scenarios with only A1 changing. Policy evidence correctly preserves the mandatory-condition list and the timestamped-image conflict rule originating in the original state; rules already contained in the retained questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction verifies every mandatory condition identified by the original state: inward-facing smooth panel, installation of each of the two straps, and no rocking. Under the unchanged question’s pass criterion, this is sufficient for a true outcome.", "rule_index": 0, "sound": true}, {"reason": "A1 is refuted, so the mandatory inward-facing smooth-panel condition fails. The unchanged question explicitly makes any mandatory failure sufficient for a false outcome, regardless of the other supported conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At final inspection, the smooth face of the back panel installed in the fictional Alderline C-4 cabinet points inward."}, {"id": "A2", "statement": "At final inspection, the first of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A3", "statement": "At final inspection, the second of the fictional Alderline C-4 cabinet’s two required anti-tip straps is installed."}, {"id": "A4", "statement": "The fictional Alderline C-4 cabinet does not rock during final-use testing."}], "base_state_json": "{\"context\":\"In a living-room workshop, a household assembler completed a fictional Alderline C-4 cabinet. Inspection records include sensor logs and physical checks.\",\"evidence\":[\"The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing.\",\"When notes conflict with timestamped images, the later image controls.\",\"At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 180 degrees.\",\"For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward.\",\"At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured.\"],\"request\":\"Does the cabinet pass final build inspection?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "2"], "text": "At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 180 degrees."}, {"path": ["evidence", "3"], "text": "For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward."}], "policy_evidence": [{"path": ["context"], "text": "The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing."}, {"path": ["context"], "text": "When notes conflict with timestamped images, the later image controls."}], "rules": [{"justification": "All independently variable mandatory inspection conditions are verified: the back panel has the required orientation, both required straps are installed, and the cabinet does not rock.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}, {"justification": "The required back-panel orientation explicitly fails, so the cabinet cannot pass even though both straps are installed and the no-rock condition is verified.", "target": "false", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 180 degrees.", "negative_left": "At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 0 degrees.", "negative_right": "For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward.", "right": "For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward."}, "verifier_independent_model": false}, "family": "fast-43-diverse-195-032", "id": "fast-43-diverse-195-032-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one mandatory condition fails, so the cabinet must be corrected and reinspected.", "true": "Yes — all mandatory inspection conditions are verified, so the cabinet passes final inspection."}, "instructions": "Answer the yes/no inspection question using the stated pass rule and conflict-resolution policy. If any required condition fails, the item does not pass and should be routed to step correction before reinspection.", "type": "noul"}}, "state": {"context": "In a living-room workshop, a household assembler completed a fictional Alderline C-4 cabinet. Inspection records include sensor logs and physical checks.", "evidence": ["The supplied inspection rule says it passes only if the back panel’s smooth face points inward, both anti-tip straps are installed, and the cabinet does not rock during final-use testing.", "When notes conflict with timestamped images, the later image controls.", "At final inspection, the alignment sensor on the back panel of the Alderline C-4 cabinet reads a rotation value of 0 degrees.", "For the Alderline C-4 cabinet's back panel, a sensor rotation value of 180 degrees corresponds to the smooth face pointing inward, whereas a value of 0 degrees corresponds to the smooth face pointing outward.", "At 14:30, the final-use tester reports no rocking and photographs both anti-tip straps secured."], "request": "Does the cabinet pass final build inspection?"}}, "method": "c2d", "provenance": {"source_id": "diverse-195", "source_is_synthetic": true, "source_sha256": "04a77b76f7a4a5187ecbed5de472aca0e9f7c8b0f35afae4270c2252d72418b1", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full question/rubric context unchanged; the counterfactual only swaps the submitted unit's serial tag to LK-5892, creating a coherent mismatch with session S's LK-4471 unit rather than a contradiction, since the verification file's documentation claim can remain true even if it now refers to the wrong physical unit; the two focus_evidence spans are complete factual sentences about serial tags, not policy or instructions; no gold answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[\"Case file: Larkspur unit, placement decision dated 2026-09-17.\",\"All required safety tests for the Larkspur unit submitted on 2026-09-17 returned passing results as of the placement decision.\",\"Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint has qualifying evidence on file as of 2026-09-17.\",\"No required progress images exist for the step-6 back-panel checkpoint; the count of such images is zero.\",\"During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit's back-panel grooved face oriented inward.\",\"That observation from session S is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\",\"Excluding session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation as of the placement decision.\",\"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471.\",\"The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-4471.\"]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["7"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471."}, {"path": ["8"], "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-4471."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471.", "negative_right": "The physical Larkspur unit submitted for final placement on 2026-09-17 instead bears serial tag LK-5892.", "right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-4471."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-004", "id": "fast-43-diverse-196-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": ["Case file: Larkspur unit, placement decision dated 2026-09-17.", "All required safety tests for the Larkspur unit submitted on 2026-09-17 returned passing results as of the placement decision.", "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint has qualifying evidence on file as of 2026-09-17.", "No required progress images exist for the step-6 back-panel checkpoint; the count of such images is zero.", "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit's back-panel grooved face oriented inward.", "That observation from session S is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.", "Excluding session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation as of the placement decision.", "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471.", "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-4471."]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full question/rubric context unchanged; the counterfactual only swaps the submitted unit's serial tag to LK-5892, creating a coherent mismatch with session S's LK-4471 unit rather than a contradiction, since the verification file's documentation claim can remain true even if it now refers to the wrong physical unit; the two focus_evidence spans are complete factual sentences about serial tags, not policy or instructions; no gold answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[\"Case file: Larkspur unit, placement decision dated 2026-09-17.\",\"All required safety tests for the Larkspur unit submitted on 2026-09-17 returned passing results as of the placement decision.\",\"Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint has qualifying evidence on file as of 2026-09-17.\",\"No required progress images exist for the step-6 back-panel checkpoint; the count of such images is zero.\",\"During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit's back-panel grooved face oriented inward.\",\"That observation from session S is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\",\"Excluding session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation as of the placement decision.\",\"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471.\",\"The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-4471.\"]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["7"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471."}, {"path": ["8"], "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-4471."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471.", "negative_right": "The physical Larkspur unit submitted for final placement on 2026-09-17 instead bears serial tag LK-5892.", "right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-4471."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-004", "id": "fast-43-diverse-196-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": ["Case file: Larkspur unit, placement decision dated 2026-09-17.", "All required safety tests for the Larkspur unit submitted on 2026-09-17 returned passing results as of the placement decision.", "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint has qualifying evidence on file as of 2026-09-17.", "No required progress images exist for the step-6 back-panel checkpoint; the count of such images is zero.", "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit's back-panel grooved face oriented inward.", "That observation from session S is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.", "Excluding session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation as of the placement decision.", "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-4471.", "The physical Larkspur unit submitted for final placement on 2026-09-17 instead bears serial tag LK-5892."]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric intact via the unchanged questions object, preserve the same unit/date/serial bindings, and the two focus_evidence spans are complete factual sentences; the counterfactual only swaps the log auditor's serial tag to LK-9188, creating a plausible single-fact discrepancy that tests whether session S's evidence truly applies to the submitted LK-7042 unit rather than an outright contradiction, and no sentence reveals or hints at the gold decision, rule IDs, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\": \"Household assembler\", \"text\": \"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"}, {\"speaker\": \"Parts checker\", \"text\": \"The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges.\"}, {\"speaker\": \"Final-use tester\", \"text\": \"With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation.\"}, {\"speaker\": \"Inspection clerk\", \"text\": \"All required safety tests for the submitted unit passed as of 2026-09-17, and every concealed-orientation checkpoint other than the step-6 back-panel checkpoint already has qualifying evidence. No required progress image exists for the step-6 checkpoint, and apart from session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation. Session S's direct observation of the back-panel grooved face oriented inward on 2026-09-16 is documented in the step-6 verification file used for this placement decision.\"}, {\"speaker\": \"Log auditor\", \"text\": \"Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-7042 in the inspection log.\"}, {\"speaker\": \"Assembly clerk\", \"text\": \"The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-7042 in the inspection log."}, {"path": ["5", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-7042 in the inspection log.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-9188 in the inspection log.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-011", "id": "fast-43-diverse-196-011-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation."}, {"speaker": "Inspection clerk", "text": "All required safety tests for the submitted unit passed as of 2026-09-17, and every concealed-orientation checkpoint other than the step-6 back-panel checkpoint already has qualifying evidence. No required progress image exists for the step-6 checkpoint, and apart from session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation. Session S's direct observation of the back-panel grooved face oriented inward on 2026-09-16 is documented in the step-6 verification file used for this placement decision."}, {"speaker": "Log auditor", "text": "Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-7042 in the inspection log."}, {"speaker": "Assembly clerk", "text": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric intact via the unchanged questions object, preserve the same unit/date/serial bindings, and the two focus_evidence spans are complete factual sentences; the counterfactual only swaps the log auditor's serial tag to LK-9188, creating a plausible single-fact discrepancy that tests whether session S's evidence truly applies to the submitted LK-7042 unit rather than an outright contradiction, and no sentence reveals or hints at the gold decision, rule IDs, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\": \"Household assembler\", \"text\": \"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"}, {\"speaker\": \"Parts checker\", \"text\": \"The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges.\"}, {\"speaker\": \"Final-use tester\", \"text\": \"With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation.\"}, {\"speaker\": \"Inspection clerk\", \"text\": \"All required safety tests for the submitted unit passed as of 2026-09-17, and every concealed-orientation checkpoint other than the step-6 back-panel checkpoint already has qualifying evidence. No required progress image exists for the step-6 checkpoint, and apart from session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation. Session S's direct observation of the back-panel grooved face oriented inward on 2026-09-16 is documented in the step-6 verification file used for this placement decision.\"}, {\"speaker\": \"Log auditor\", \"text\": \"Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-7042 in the inspection log.\"}, {\"speaker\": \"Assembly clerk\", \"text\": \"The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-7042 in the inspection log."}, {"path": ["5", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-7042 in the inspection log.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-9188 in the inspection log.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-011", "id": "fast-43-diverse-196-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation."}, {"speaker": "Inspection clerk", "text": "All required safety tests for the submitted unit passed as of 2026-09-17, and every concealed-orientation checkpoint other than the step-6 back-panel checkpoint already has qualifying evidence. No required progress image exists for the step-6 checkpoint, and apart from session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation. Session S's direct observation of the back-panel grooved face oriented inward on 2026-09-16 is documented in the step-6 verification file used for this placement decision."}, {"speaker": "Log auditor", "text": "Pre-placement visual-inspection session S on 2026-09-16 was recorded under unit serial tag LK-9188 in the inspection log."}, {"speaker": "Assembly clerk", "text": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-7042 on its assembly ticket."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy language and the unchanged question object; entity (Larkspur unit LK-4471), path (step-6 back-panel orientation checkpoint) and time (2026-09-16/17) bindings are preserved; the counterfactual serial-tag swap (LK-5502 vs LK-4471) is a coherent substantive change showing the verification file documents a different physical unit than the one submitted, without introducing contradictory duplicate measurements; the two focus evidence spans are complete factual sentences, not policy or rationale; neither context reveals a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\": \"Household assembler\", \"text\": \"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"}, {\"speaker\": \"Parts checker\", \"text\": \"The hardware count matches the instructions. All required safety tests passed, and every other concealed-orientation checkpoint has qualifying evidence on file. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges, so zero required progress images exist for that checkpoint.\"}, {\"speaker\": \"Final-use tester\", \"text\": \"With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation myself, and no other current direct visual confirmation of that orientation exists for the submitted unit.\"}, {\"speaker\": \"Inspection records clerk\", \"text\": \"Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-4471. During that session the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and that observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17. The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["3", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-4471."}, {"path": ["3", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-4471.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-5502.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-013", "id": "fast-43-diverse-196-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions. All required safety tests passed, and every other concealed-orientation checkpoint has qualifying evidence on file. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges, so zero required progress images exist for that checkpoint."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation myself, and no other current direct visual confirmation of that orientation exists for the submitted unit."}, {"speaker": "Inspection records clerk", "text": "Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-4471. During that session the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and that observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17. The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy language and the unchanged question object; entity (Larkspur unit LK-4471), path (step-6 back-panel orientation checkpoint) and time (2026-09-16/17) bindings are preserved; the counterfactual serial-tag swap (LK-5502 vs LK-4471) is a coherent substantive change showing the verification file documents a different physical unit than the one submitted, without introducing contradictory duplicate measurements; the two focus evidence spans are complete factual sentences, not policy or rationale; neither context reveals a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\": \"Household assembler\", \"text\": \"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"}, {\"speaker\": \"Parts checker\", \"text\": \"The hardware count matches the instructions. All required safety tests passed, and every other concealed-orientation checkpoint has qualifying evidence on file. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges, so zero required progress images exist for that checkpoint.\"}, {\"speaker\": \"Final-use tester\", \"text\": \"With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation myself, and no other current direct visual confirmation of that orientation exists for the submitted unit.\"}, {\"speaker\": \"Inspection records clerk\", \"text\": \"Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-4471. During that session the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and that observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17. The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["3", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-4471."}, {"path": ["3", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-4471.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-5502.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-013", "id": "fast-43-diverse-196-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions. All required safety tests passed, and every other concealed-orientation checkpoint has qualifying evidence on file. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges, so zero required progress images exist for that checkpoint."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation myself, and no other current direct visual confirmation of that orientation exists for the submitted unit."}, {"speaker": "Inspection records clerk", "text": "Pre-placement visual-inspection session S on 2026-09-16 inspected a physical unit bearing serial tag LK-5502. During that session the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and that observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17. The Larkspur unit submitted for final placement on 2026-09-17 bears serial tag LK-4471."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all QC-log facts, the original question/policy is untouched, and the entity, path and 2026-09-17 timing are preserved; the counterfactual only swaps the session-S unit's serial tag to LKS-9081, which stays logically consistent with the rest of the scene (no duplicated or contradictory measurements) and contains no gold-answer or rule-table leakage, while the two focus_evidence spans are complete factual sentences rather than policy text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Quality control log\",\"text\":\"All required safety tests for the Larkspur unit submitted on 2026-09-17 passed as of the placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Every concealed-orientation checkpoint besides the step-6 back-panel checkpoint has qualifying evidence on file as of the placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Zero required progress images exist for the step-6 back-panel checkpoint as of the placement decision.\"},{\"speaker\":\"Inspector\",\"text\":\"During session S on 2026-09-16 I directly observed the inspected unit's back-panel grooved face oriented inward.\"},{\"speaker\":\"Records clerk\",\"text\":\"That session-S observation is documented in the step-6 verification file used for the September 17 placement decision.\"},{\"speaker\":\"Records clerk\",\"text\":\"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-7742.\"},{\"speaker\":\"Records clerk\",\"text\":\"The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742.\"},{\"speaker\":\"Quality control log\",\"text\":\"Excluding session S, zero current direct visual confirmations exist for the step-6 back-panel orientation as of the placement decision.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["5", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-7742."}, {"path": ["6", "text"], "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-7742.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-9081.", "negative_right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742.", "right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-017", "id": "fast-43-diverse-196-017-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Quality control log", "text": "All required safety tests for the Larkspur unit submitted on 2026-09-17 passed as of the placement decision."}, {"speaker": "Quality control log", "text": "Every concealed-orientation checkpoint besides the step-6 back-panel checkpoint has qualifying evidence on file as of the placement decision."}, {"speaker": "Quality control log", "text": "Zero required progress images exist for the step-6 back-panel checkpoint as of the placement decision."}, {"speaker": "Inspector", "text": "During session S on 2026-09-16 I directly observed the inspected unit's back-panel grooved face oriented inward."}, {"speaker": "Records clerk", "text": "That session-S observation is documented in the step-6 verification file used for the September 17 placement decision."}, {"speaker": "Records clerk", "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-7742."}, {"speaker": "Records clerk", "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742."}, {"speaker": "Quality control log", "text": "Excluding session S, zero current direct visual confirmations exist for the step-6 back-panel orientation as of the placement decision."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all QC-log facts, the original question/policy is untouched, and the entity, path and 2026-09-17 timing are preserved; the counterfactual only swaps the session-S unit's serial tag to LKS-9081, which stays logically consistent with the rest of the scene (no duplicated or contradictory measurements) and contains no gold-answer or rule-table leakage, while the two focus_evidence spans are complete factual sentences rather than policy text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Quality control log\",\"text\":\"All required safety tests for the Larkspur unit submitted on 2026-09-17 passed as of the placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Every concealed-orientation checkpoint besides the step-6 back-panel checkpoint has qualifying evidence on file as of the placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Zero required progress images exist for the step-6 back-panel checkpoint as of the placement decision.\"},{\"speaker\":\"Inspector\",\"text\":\"During session S on 2026-09-16 I directly observed the inspected unit's back-panel grooved face oriented inward.\"},{\"speaker\":\"Records clerk\",\"text\":\"That session-S observation is documented in the step-6 verification file used for the September 17 placement decision.\"},{\"speaker\":\"Records clerk\",\"text\":\"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-7742.\"},{\"speaker\":\"Records clerk\",\"text\":\"The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742.\"},{\"speaker\":\"Quality control log\",\"text\":\"Excluding session S, zero current direct visual confirmations exist for the step-6 back-panel orientation as of the placement decision.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["5", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-7742."}, {"path": ["6", "text"], "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-7742.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-9081.", "negative_right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742.", "right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-017", "id": "fast-43-diverse-196-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Quality control log", "text": "All required safety tests for the Larkspur unit submitted on 2026-09-17 passed as of the placement decision."}, {"speaker": "Quality control log", "text": "Every concealed-orientation checkpoint besides the step-6 back-panel checkpoint has qualifying evidence on file as of the placement decision."}, {"speaker": "Quality control log", "text": "Zero required progress images exist for the step-6 back-panel checkpoint as of the placement decision."}, {"speaker": "Inspector", "text": "During session S on 2026-09-16 I directly observed the inspected unit's back-panel grooved face oriented inward."}, {"speaker": "Records clerk", "text": "That session-S observation is documented in the step-6 verification file used for the September 17 placement decision."}, {"speaker": "Records clerk", "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-9081."}, {"speaker": "Records clerk", "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-7742."}, {"speaker": "Quality control log", "text": "Excluding session S, zero current direct visual confirmations exist for the step-6 back-panel orientation as of the placement decision."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original questions and policy verbatim, keep the same entity/time bindings, and the two focus sentences are complete factual statements rather than policy text; the counterfactual coherently changes only the placement unit's serial tag to LK-5820, showing session S refers to a different unit without contradicting any other measurement, and neither context reveals a gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\": \"Household assembler\", \"text\": \"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"}, {\"speaker\": \"Parts checker\", \"text\": \"The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges.\"}, {\"speaker\": \"Final-use tester\", \"text\": \"With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation.\"}, {\"speaker\": \"Household assembler\", \"text\": \"I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim.\"}, {\"speaker\": \"Inspection log\", \"text\": \"All required safety tests for the submitted Larkspur unit passed as of the 2026-09-17 placement review. Every other concealed-orientation checkpoint besides step-6 back-panel has qualifying evidence on file. No progress image exists for the step-6 back-panel checkpoint. The step-6 verification file references a pre-placement visual-inspection session S conducted on 2026-09-16, during which the inspector directly observed a unit's back-panel grooved face oriented inward, and that observation was entered into the file. Apart from session S, no current direct visual confirmation of the step-6 back-panel orientation exists for the unit under review.\"}, {\"speaker\": \"Serial tag record\", \"text\": \"The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate.\"}, {\"speaker\": \"Placement record\", \"text\": \"The Larkspur unit submitted for final placement on 2026-09-17 also bore serial tag LK-4471 affixed to its base plate.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["5", "text"], "text": "The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate."}, {"path": ["6", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bore serial tag LK-4471 affixed to its base plate."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate.", "negative_left": "The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 instead bore serial tag LK-5820 affixed to its base plate.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 also bore serial tag LK-4471 affixed to its base plate."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-021", "id": "fast-43-diverse-196-021-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation."}, {"speaker": "Household assembler", "text": "I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim."}, {"speaker": "Inspection log", "text": "All required safety tests for the submitted Larkspur unit passed as of the 2026-09-17 placement review. Every other concealed-orientation checkpoint besides step-6 back-panel has qualifying evidence on file. No progress image exists for the step-6 back-panel checkpoint. The step-6 verification file references a pre-placement visual-inspection session S conducted on 2026-09-16, during which the inspector directly observed a unit's back-panel grooved face oriented inward, and that observation was entered into the file. Apart from session S, no current direct visual confirmation of the step-6 back-panel orientation exists for the unit under review."}, {"speaker": "Serial tag record", "text": "The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate."}, {"speaker": "Placement record", "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bore serial tag LK-4471 affixed to its base plate."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original questions and policy verbatim, keep the same entity/time bindings, and the two focus sentences are complete factual statements rather than policy text; the counterfactual coherently changes only the placement unit's serial tag to LK-5820, showing session S refers to a different unit without contradicting any other measurement, and neither context reveals a gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\": \"Household assembler\", \"text\": \"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"}, {\"speaker\": \"Parts checker\", \"text\": \"The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges.\"}, {\"speaker\": \"Final-use tester\", \"text\": \"With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation.\"}, {\"speaker\": \"Household assembler\", \"text\": \"I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim.\"}, {\"speaker\": \"Inspection log\", \"text\": \"All required safety tests for the submitted Larkspur unit passed as of the 2026-09-17 placement review. Every other concealed-orientation checkpoint besides step-6 back-panel has qualifying evidence on file. No progress image exists for the step-6 back-panel checkpoint. The step-6 verification file references a pre-placement visual-inspection session S conducted on 2026-09-16, during which the inspector directly observed a unit's back-panel grooved face oriented inward, and that observation was entered into the file. Apart from session S, no current direct visual confirmation of the step-6 back-panel orientation exists for the unit under review.\"}, {\"speaker\": \"Serial tag record\", \"text\": \"The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate.\"}, {\"speaker\": \"Placement record\", \"text\": \"The Larkspur unit submitted for final placement on 2026-09-17 also bore serial tag LK-4471 affixed to its base plate.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["5", "text"], "text": "The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate."}, {"path": ["6", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bore serial tag LK-4471 affixed to its base plate."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate.", "negative_left": "The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 instead bore serial tag LK-5820 affixed to its base plate.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 also bore serial tag LK-4471 affixed to its base plate."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-021", "id": "fast-43-diverse-196-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation."}, {"speaker": "Household assembler", "text": "I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim."}, {"speaker": "Inspection log", "text": "All required safety tests for the submitted Larkspur unit passed as of the 2026-09-17 placement review. Every other concealed-orientation checkpoint besides step-6 back-panel has qualifying evidence on file. No progress image exists for the step-6 back-panel checkpoint. The step-6 verification file references a pre-placement visual-inspection session S conducted on 2026-09-16, during which the inspector directly observed a unit's back-panel grooved face oriented inward, and that observation was entered into the file. Apart from session S, no current direct visual confirmation of the step-6 back-panel orientation exists for the unit under review."}, {"speaker": "Serial tag record", "text": "The unit inspected during session S on 2026-09-16 bore serial tag LK-4471 affixed to its base plate."}, {"speaker": "Placement record", "text": "The Larkspur unit submitted for final placement on 2026-09-17 instead bore serial tag LK-5820 affixed to its base plate."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric, entity (Larkspur unit), session S, and dates intact, add only observational detail without inventing policy exceptions, and contain no gold-answer or rule-table leakage; the counterfactual's serial-tag change (LK-7742 to LK-9013) is a single coherent factual edit that alters the identity link rather than duplicating or contradicting an existing measurement, and the two focus-evidence spans are complete factual sentences, not instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\": \"Household assembler\", \"text\": \"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"}, {\"speaker\": \"Parts checker\", \"text\": \"The hardware count matches the instructions and every required safety test passed. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges, so there are zero required progress images for that checkpoint.\"}, {\"speaker\": \"Final-use tester\", \"text\": \"With 12 kg distributed across the shelves, it does not wobble or creak. Every other concealed-orientation checkpoint has qualifying evidence on file. Excluding session S, there are zero current direct visual confirmations of the step-6 back-panel orientation.\"}, {\"speaker\": \"Inspection log\", \"text\": \"During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and this observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"}, {\"speaker\": \"Household assembler\", \"text\": \"I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim.\"}, {\"speaker\": \"Placement clerk\", \"text\": \"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7742.\"}, {\"speaker\": \"Placement clerk\", \"text\": \"The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["5", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7742."}, {"path": ["6", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7742.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-9013.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-025", "id": "fast-43-diverse-196-025-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions and every required safety test passed. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges, so there are zero required progress images for that checkpoint."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. Every other concealed-orientation checkpoint has qualifying evidence on file. Excluding session S, there are zero current direct visual confirmations of the step-6 back-panel orientation."}, {"speaker": "Inspection log", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and this observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Household assembler", "text": "I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim."}, {"speaker": "Placement clerk", "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7742."}, {"speaker": "Placement clerk", "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric, entity (Larkspur unit), session S, and dates intact, add only observational detail without inventing policy exceptions, and contain no gold-answer or rule-table leakage; the counterfactual's serial-tag change (LK-7742 to LK-9013) is a single coherent factual edit that alters the identity link rather than duplicating or contradicting an existing measurement, and the two focus-evidence spans are complete factual sentences, not instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\": \"Household assembler\", \"text\": \"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"}, {\"speaker\": \"Parts checker\", \"text\": \"The hardware count matches the instructions and every required safety test passed. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges, so there are zero required progress images for that checkpoint.\"}, {\"speaker\": \"Final-use tester\", \"text\": \"With 12 kg distributed across the shelves, it does not wobble or creak. Every other concealed-orientation checkpoint has qualifying evidence on file. Excluding session S, there are zero current direct visual confirmations of the step-6 back-panel orientation.\"}, {\"speaker\": \"Inspection log\", \"text\": \"During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and this observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17.\"}, {\"speaker\": \"Household assembler\", \"text\": \"I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim.\"}, {\"speaker\": \"Placement clerk\", \"text\": \"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7742.\"}, {\"speaker\": \"Placement clerk\", \"text\": \"The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["5", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7742."}, {"path": ["6", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7742.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-9013.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-025", "id": "fast-43-diverse-196-025-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions and every required safety test passed. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges, so there are zero required progress images for that checkpoint."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. Every other concealed-orientation checkpoint has qualifying evidence on file. Excluding session S, there are zero current direct visual confirmations of the step-6 back-panel orientation."}, {"speaker": "Inspection log", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and this observation is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"speaker": "Household assembler", "text": "I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim."}, {"speaker": "Placement clerk", "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-9013."}, {"speaker": "Placement clerk", "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7742, confirmed by the placement clerk's log."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only swaps the serial tag examined in session S (LK-9981 vs LK-7042), creating a coherent alternate fact where the observed unit differs from the submitted unit, without altering any policy, question bindings, or embedding an answer; both focus_evidence spans are complete factual sentences about serial identity rather than policy text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Quality control log\",\"text\":\"All required safety tests for the Larkspur three-shelf unit passed as of the 2026-09-17 placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Every concealed-orientation checkpoint except the step-6 back-panel checkpoint has qualifying evidence on file.\"},{\"speaker\":\"Quality control log\",\"text\":\"No required progress image exists for the step-6 back-panel checkpoint; count of such images is zero.\"},{\"speaker\":\"Quality control log\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit's back-panel grooved face oriented inward.\"},{\"speaker\":\"Quality control log\",\"text\":\"This observation from session S is documented in the step-6 verification file used for the 2026-09-17 placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Excluding session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation as of the placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7042.\"},{\"speaker\":\"Quality control log\",\"text\":\"The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["6", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7042."}, {"path": ["7", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7042.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-9981.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-028", "id": "fast-43-diverse-196-028-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Quality control log", "text": "All required safety tests for the Larkspur three-shelf unit passed as of the 2026-09-17 placement decision."}, {"speaker": "Quality control log", "text": "Every concealed-orientation checkpoint except the step-6 back-panel checkpoint has qualifying evidence on file."}, {"speaker": "Quality control log", "text": "No required progress image exists for the step-6 back-panel checkpoint; count of such images is zero."}, {"speaker": "Quality control log", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit's back-panel grooved face oriented inward."}, {"speaker": "Quality control log", "text": "This observation from session S is documented in the step-6 verification file used for the 2026-09-17 placement decision."}, {"speaker": "Quality control log", "text": "Excluding session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation as of the placement decision."}, {"speaker": "Quality control log", "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7042."}, {"speaker": "Quality control log", "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only swaps the serial tag examined in session S (LK-9981 vs LK-7042), creating a coherent alternate fact where the observed unit differs from the submitted unit, without altering any policy, question bindings, or embedding an answer; both focus_evidence spans are complete factual sentences about serial identity rather than policy text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Quality control log\",\"text\":\"All required safety tests for the Larkspur three-shelf unit passed as of the 2026-09-17 placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Every concealed-orientation checkpoint except the step-6 back-panel checkpoint has qualifying evidence on file.\"},{\"speaker\":\"Quality control log\",\"text\":\"No required progress image exists for the step-6 back-panel checkpoint; count of such images is zero.\"},{\"speaker\":\"Quality control log\",\"text\":\"During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit's back-panel grooved face oriented inward.\"},{\"speaker\":\"Quality control log\",\"text\":\"This observation from session S is documented in the step-6 verification file used for the 2026-09-17 placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Excluding session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation as of the placement decision.\"},{\"speaker\":\"Quality control log\",\"text\":\"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7042.\"},{\"speaker\":\"Quality control log\",\"text\":\"The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["6", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7042."}, {"path": ["7", "text"], "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-7042.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-9981.", "negative_right": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042.", "right": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-028", "id": "fast-43-diverse-196-028-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Quality control log", "text": "All required safety tests for the Larkspur three-shelf unit passed as of the 2026-09-17 placement decision."}, {"speaker": "Quality control log", "text": "Every concealed-orientation checkpoint except the step-6 back-panel checkpoint has qualifying evidence on file."}, {"speaker": "Quality control log", "text": "No required progress image exists for the step-6 back-panel checkpoint; count of such images is zero."}, {"speaker": "Quality control log", "text": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit's back-panel grooved face oriented inward."}, {"speaker": "Quality control log", "text": "This observation from session S is documented in the step-6 verification file used for the 2026-09-17 placement decision."}, {"speaker": "Quality control log", "text": "Excluding session S, there are zero current direct visual confirmations of the submitted unit's step-6 back-panel orientation as of the placement decision."}, {"speaker": "Quality control log", "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LK-9981."}, {"speaker": "Quality control log", "text": "The Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LK-7042."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only changes the inspected unit's serial tag to LKS-4407 while keeping the placement log's LKS-2291, creating a coherent mismatch that tests whether the inspection applies to the submitted unit without contradicting any other stated facts or policy, and the question's generic Larkspur-unit binding is untouched; no gold answers, rule tables, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Household assembler\",\"text\":\"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"},{\"speaker\":\"Parts checker\",\"text\":\"The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges.\"},{\"speaker\":\"Final-use tester\",\"text\":\"With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation.\"},{\"speaker\":\"Household assembler\",\"text\":\"I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim.\"},{\"speaker\":\"Inspection log\",\"text\":\"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-2291.\"},{\"speaker\":\"Placement log\",\"text\":\"The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291.\"},{\"speaker\":\"Inspection log\",\"text\":\"During session S, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and this observation is documented in the step-6 verification file used for the placement decision.\"},{\"speaker\":\"Safety log\",\"text\":\"All required safety tests for the submitted unit passed as of the placement decision, and every other concealed-orientation checkpoint besides the step-6 back-panel checkpoint has qualifying evidence on file.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-2291."}, {"path": ["5", "text"], "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-2291.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-4407.", "negative_right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291.", "right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-031", "id": "fast-43-diverse-196-031-base", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation."}, {"speaker": "Household assembler", "text": "I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim."}, {"speaker": "Inspection log", "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-2291."}, {"speaker": "Placement log", "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291."}, {"speaker": "Inspection log", "text": "During session S, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and this observation is documented in the step-6 verification file used for the placement decision."}, {"speaker": "Safety log", "text": "All required safety tests for the submitted unit passed as of the placement decision, and every other concealed-orientation checkpoint besides the step-6 back-panel checkpoint has qualifying evidence on file."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-03", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only changes the inspected unit's serial tag to LKS-4407 while keeping the placement log's LKS-2291, creating a coherent mismatch that tests whether the inspection applies to the submitted unit without contradicting any other stated facts or policy, and the question's generic Larkspur-unit binding is untouched; no gold answers, rule tables, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a6": "unknown"}, "remove_right": {"a6": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a6": "unknown"}, "negative_pair": {"a6": "refuted"}, "negative_sentence": {"a6": "unknown"}, "positive_pair": {"a6": "supported"}, "right": {"a6": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation, including permissible universal or count-based relations. The focus atom is the factual identity relation between the inspected and submitted units, not a policy conclusion. Both assignments are realizable while holding the other atoms fixed: the same observation may be documented in the placement file even when it was mistakenly made on a different unit. Empty policy_evidence is correct because all substantive decision rules are contained in the retained questions object; the original state supplies case observations rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions entail that all safety tests pass, all other concealed-orientation checkpoints have qualifying evidence, and session S provides a documented direct observation of the submitted unit’s step-6 orientation. Because a6 establishes unit identity, that observation can serve as the current visual confirmation permitted by the rubric despite the missing progress image. No withholding is required.", "rule_index": 0, "sound": true}, {"reason": "With a6 refuted, session S concerns a different physical unit and cannot confirm the submitted unit’s orientation. Together, a3 and a7 establish that the submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. The remaining conditions establish passing safety tests and qualifying evidence for every other checkpoint, so the evidence-only deficiency requires withholding and reopening and is minor.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "Every safety test required for the physical Larkspur unit submitted for final placement on 2026-09-17 has a passing result as of the placement decision on 2026-09-17."}, {"id": "a2", "statement": "Every concealed-orientation checkpoint other than the step-6 back-panel checkpoint for the physical Larkspur unit submitted for final placement on 2026-09-17 has qualifying evidence as of the placement decision on 2026-09-17."}, {"id": "a3", "statement": "The number of required progress images available for the step-6 back-panel checkpoint of the physical Larkspur unit submitted for final placement on 2026-09-17 is zero as of the placement decision."}, {"id": "a4", "statement": "During pre-placement visual-inspection session S on 2026-09-16, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward."}, {"id": "a5", "statement": "The observation made during pre-placement visual-inspection session S on 2026-09-16 is documented in the step-6 verification file used for the Larkspur placement decision on 2026-09-17."}, {"id": "a6", "statement": "The physical unit inspected during pre-placement visual-inspection session S on 2026-09-16 is the same physical unit as the Larkspur unit submitted for final placement on 2026-09-17."}, {"id": "a7", "statement": "Excluding pre-placement visual-inspection session S, the number of current direct visual confirmations available for the submitted Larkspur unit’s step-6 back-panel orientation is zero as of the placement decision on 2026-09-17."}], "base_state_json": "[{\"speaker\":\"Household assembler\",\"text\":\"The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6.\"},{\"speaker\":\"Parts checker\",\"text\":\"The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges.\"},{\"speaker\":\"Final-use tester\",\"text\":\"With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation.\"},{\"speaker\":\"Household assembler\",\"text\":\"I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim.\"},{\"speaker\":\"Inspection log\",\"text\":\"Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-2291.\"},{\"speaker\":\"Placement log\",\"text\":\"The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291.\"},{\"speaker\":\"Inspection log\",\"text\":\"During session S, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and this observation is documented in the step-6 verification file used for the placement decision.\"},{\"speaker\":\"Safety log\",\"text\":\"All required safety tests for the submitted unit passed as of the placement decision, and every other concealed-orientation checkpoint besides the step-6 back-panel checkpoint has qualifying evidence on file.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a6", "focus_evidence": [{"path": ["4", "text"], "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-2291."}, {"path": ["5", "text"], "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291."}], "policy_evidence": [], "rules": [{"justification": "The documented direct observation from session S is a current visual confirmation of the submitted Larkspur unit’s step-6 back-panel orientation. It therefore satisfies that checkpoint despite the missing progress image. Every required safety test passes and every other concealed-orientation checkpoint has qualifying evidence, so reopening for orientation verification is unnecessary.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "Session S inspected a different physical unit, so its documented observation cannot visually confirm the submitted Larkspur unit’s orientation. The submitted unit has neither the required step-6 progress image nor another current direct visual confirmation. All safety tests pass and every other concealed-orientation checkpoint has qualifying evidence, so only the required evidence is missing; the issue is minor, but the unit must be withheld and reopened for verification.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-2291.", "negative_left": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-4407.", "negative_right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291.", "right": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291."}, "verifier_independent_model": false}, "family": "fast-43-diverse-196-031", "id": "fast-43-diverse-196-031-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — approve the unit for final placement without reopening it, treating the assembler’s note and successful use test as sufficient evidence.", "true": "Yes — withhold the unit from final placement and route it to step correction so the concealed back-panel orientation can be exposed, verified, and documented; rate the issue minor."}, "instructions": "Verification decision: Should the Larkspur unit be withheld from final placement and routed to step correction for reopening and documented orientation verification? Rubric: a build passes only when every safety test passes and each concealed-orientation checkpoint has either a required progress image or current visual confirmation. Notes alone do not satisfy that checkpoint. If only the required evidence is missing, severity is minor but the item must still be withheld and reopened for verification.", "type": "noul"}}, "state": [{"speaker": "Household assembler", "text": "The Larkspur three-shelf unit is square, all eight cam locks are tightened, and my notes say the back panel’s grooved face was installed inward at step 6."}, {"speaker": "Parts checker", "text": "The hardware count matches the instructions. However, the required step-6 photo is missing; the only image was taken after the rear trim concealed the panel edges."}, {"speaker": "Final-use tester", "text": "With 12 kg distributed across the shelves, it does not wobble or creak. The wall strap is fitted, but I cannot visually confirm the concealed back-panel orientation."}, {"speaker": "Household assembler", "text": "I deleted the step-6 photo accidentally. Confirming the orientation now would require removing the rear trim."}, {"speaker": "Inspection log", "text": "Pre-placement visual-inspection session S on 2026-09-16 examined a physical unit bearing serial tag LKS-4407."}, {"speaker": "Placement log", "text": "The physical Larkspur unit submitted for final placement on 2026-09-17 also bears serial tag LKS-2291."}, {"speaker": "Inspection log", "text": "During session S, the inspector directly observed the inspected unit’s back-panel grooved face oriented inward, and this observation is documented in the step-6 verification file used for the placement decision."}, {"speaker": "Safety log", "text": "All required safety tests for the submitted unit passed as of the placement decision, and every other concealed-orientation checkpoint besides the step-6 back-panel checkpoint has qualifying evidence on file."}]}, "method": "c2d", "provenance": {"source_id": "diverse-196", "source_is_synthetic": true, "source_sha256": "c75d24200095863e73b26407be099f340fcdc207a7a64342ad53fe7e5d1b5bc3", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-03", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 3 mm tolerance policy and match the original question's entity/path, the two evidence sentences are factual measurement statements, the counterfactual's changed 12 mm mark yields a 2 mm deviation consistent with the rest of the narrative, and neither context reveals a gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"At the household craft station, sewing hobbyist Mara is finishing Priya's 45 cm cushion cover. The project sheet allows no more than 3 mm seam deviation. At 10:00, Mara marked every panel cut and all dimensions were verified. At 10:50, checker Joel confirmed the zipper and all four seams were stitched, so every required piece and seam is present and all other construction requirements pass. Trimming began at 10:55. At 11:15, Priya test-fitted the cushion and reported a bulging right edge. Joel's 11:25 update recorded measurements for that seam. On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner. On the same inspection, the recorded location of the marked line for that seam is 9 mm from the top corner. Trimming was completed and passed at 11:40, pressing was completed and passed at 11:55, and final inspection was completed and passed at 12:05.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner."}, {"path": [], "text": "On the same inspection, the recorded location of the marked line for that seam is 9 mm from the top corner."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner.", "negative_left": "On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner.", "negative_right": "On the same inspection, the recorded location of the marked line for that seam is 12 mm from the top corner.", "right": "On the same inspection, the recorded location of the marked line for that seam is 9 mm from the top corner."}, "verifier_independent_model": false}, "family": "fast-43-diverse-199-010", "id": "fast-43-diverse-199-010-base", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "At the household craft station, sewing hobbyist Mara is finishing Priya's 45 cm cushion cover. The project sheet allows no more than 3 mm seam deviation. At 10:00, Mara marked every panel cut and all dimensions were verified. At 10:50, checker Joel confirmed the zipper and all four seams were stitched, so every required piece and seam is present and all other construction requirements pass. Trimming began at 10:55. At 11:15, Priya test-fitted the cushion and reported a bulging right edge. Joel's 11:25 update recorded measurements for that seam. On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner. On the same inspection, the recorded location of the marked line for that seam is 9 mm from the top corner. Trimming was completed and passed at 11:40, pressing was completed and passed at 11:55, and final inspection was completed and passed at 12:05."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the 3 mm tolerance policy and match the original question's entity/path, the two evidence sentences are factual measurement statements, the counterfactual's changed 12 mm mark yields a 2 mm deviation consistent with the rest of the narrative, and neither context reveals a gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"At the household craft station, sewing hobbyist Mara is finishing Priya's 45 cm cushion cover. The project sheet allows no more than 3 mm seam deviation. At 10:00, Mara marked every panel cut and all dimensions were verified. At 10:50, checker Joel confirmed the zipper and all four seams were stitched, so every required piece and seam is present and all other construction requirements pass. Trimming began at 10:55. At 11:15, Priya test-fitted the cushion and reported a bulging right edge. Joel's 11:25 update recorded measurements for that seam. On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner. On the same inspection, the recorded location of the marked line for that seam is 9 mm from the top corner. Trimming was completed and passed at 11:40, pressing was completed and passed at 11:55, and final inspection was completed and passed at 12:05.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner."}, {"path": [], "text": "On the same inspection, the recorded location of the marked line for that seam is 9 mm from the top corner."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner.", "negative_left": "On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner.", "negative_right": "On the same inspection, the recorded location of the marked line for that seam is 12 mm from the top corner.", "right": "On the same inspection, the recorded location of the marked line for that seam is 9 mm from the top corner."}, "verifier_independent_model": false}, "family": "fast-43-diverse-199-010", "id": "fast-43-diverse-199-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "At the household craft station, sewing hobbyist Mara is finishing Priya's 45 cm cushion cover. The project sheet allows no more than 3 mm seam deviation. At 10:00, Mara marked every panel cut and all dimensions were verified. At 10:50, checker Joel confirmed the zipper and all four seams were stitched, so every required piece and seam is present and all other construction requirements pass. Trimming began at 10:55. At 11:15, Priya test-fitted the cushion and reported a bulging right edge. Joel's 11:25 update recorded measurements for that seam. On the latest timestamped inspection, the measured location of the existing right-edge seam on Priya's 45 cm cushion cover is 14 mm from the top corner. On the same inspection, the recorded location of the marked line for that seam is 12 mm from the top corner. Trimming was completed and passed at 11:40, pressing was completed and passed at 11:55, and final inspection was completed and passed at 12:05."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_clean_4"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the 3 mm tolerance policy and all construction/trimming/pressing/final-inspection facts, only varying the right-edge seam measurement (14mm vs 10mm) while keeping the marked-line value fixed at 9mm, which is coherent and doesn't duplicate or contradict other seam data; the two focus sentences are plain factual measurements, not policy text; no gold answer, rule table, or output instruction is embedded in either context; question entity, path, and time bindings for Priya's cushion cover remain unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"Inspection log update for Priya's 45 cm cushion cover: All panel dimensions were verified against the pattern, and every required cut piece is accounted for. All four construction seams, including the zipper seam, are present and stitched, and every other construction requirement passes inspection. On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 14 mm from the reference point. On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point. Trimming of loose threads has passed. Pressing has passed. Final inspection has passed. The project sheet allows no more than 3 mm seam deviation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 14 mm from the reference point."}, {"path": [], "text": "On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 14 mm from the reference point.", "negative_left": "On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 10 mm from the reference point.", "negative_right": "On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point.", "right": "On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point."}, "verifier_independent_model": false}, "family": "fast-43-diverse-199-020", "id": "fast-43-diverse-199-020-base", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "Inspection log update for Priya's 45 cm cushion cover: All panel dimensions were verified against the pattern, and every required cut piece is accounted for. All four construction seams, including the zipper seam, are present and stitched, and every other construction requirement passes inspection. On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 14 mm from the reference point. On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point. Trimming of loose threads has passed. Pressing has passed. Final inspection has passed. The project sheet allows no more than 3 mm seam deviation."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "none_of_above"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the 3 mm tolerance policy and all construction/trimming/pressing/final-inspection facts, only varying the right-edge seam measurement (14mm vs 10mm) while keeping the marked-line value fixed at 9mm, which is coherent and doesn't duplicate or contradict other seam data; the two focus sentences are plain factual measurements, not policy text; no gold answer, rule table, or output instruction is embedded in either context; question entity, path, and time bindings for Priya's cushion cover remain unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "refuted", "a6": "supported", "a7": "supported", "a8": "supported"}, "remove_left": {"a5": "unknown"}, "remove_right": {"a5": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a5": "unknown"}, "negative_pair": {"a5": "refuted"}, "negative_sentence": {"a5": "unknown"}, "positive_pair": {"a5": "supported"}, "right": {"a5": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state one factual relation; a4 is a permissible universal claim over the remaining construction requirements rather than a final decision classification. The focus a5 is a factual seam-deviation relation, not policy. The base and counter assignments differ only on a5 and are not prohibited by the explicit policy. Policy evidence correctly preserves the state-originating 3 mm tolerance; all other governing instructions and criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions exclude measuring, cutting, and stitching routes, while establishing an existing seam more than 3 mm from its marked line. The explicit instructions therefore require seam correction. None of the three named outcomes represents seam correction with the corresponding incomplete, repairable-defect rating, so none_of_above is required.", "rule_index": 0, "sound": true}, {"reason": "Refuting a5 establishes that the right-edge seam is not more than 3 mm from its marked line. Together with verified dimensions, present pieces and seams, all other construction requirements passing, and trimming, pressing, and final inspection passing, this satisfies every stated condition for complete_clean_4.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required dimension has been verified."}, {"id": "a2", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required cut piece is present."}, {"id": "a3", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every required construction seam is present."}, {"id": "a4", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, every construction requirement other than required-seam presence and the right-edge seam-deviation limit passes."}, {"id": "a5", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, the absolute difference between the measured location of the existing right-edge seam and the recorded location of its marked line is greater than 3 mm."}, {"id": "a6", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, trimming has passed."}, {"id": "a7", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, pressing has passed."}, {"id": "a8", "statement": "As of the latest timestamped evidence for Priya’s 45 cm cushion cover, final inspection has passed."}], "base_state_json": "\"Inspection log update for Priya's 45 cm cushion cover: All panel dimensions were verified against the pattern, and every required cut piece is accounted for. All four construction seams, including the zipper seam, are present and stitched, and every other construction requirement passes inspection. On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 14 mm from the reference point. On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point. Trimming of loose threads has passed. Pressing has passed. Final inspection has passed. The project sheet allows no more than 3 mm seam deviation.\"", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a5", "focus_evidence": [{"path": [], "text": "On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 14 mm from the reference point."}, {"path": [], "text": "On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point."}], "policy_evidence": [{"path": [], "text": "The project sheet allows no more than 3 mm seam deviation."}], "rules": [{"justification": "The right-edge seam is present, but its location differs from its marked-line location by more than the project sheet’s maximum permitted 3 mm. The instructions therefore require seam correction. Because none of the three listed outcomes specifies seam correction with the corresponding incomplete repair status and rating, none_of_above is required.", "target": "none_of_above", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "All dimensions are verified, all required pieces and seams are present, every other construction requirement passes, and the existing right-edge seam does not deviate from its marked line by more than the permitted 3 mm. Trimming, pressing, and final inspection also pass, so every qualification for complete_clean_4 is satisfied.", "target": "complete_clean_4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "refuted"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 14 mm from the reference point.", "negative_left": "On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 10 mm from the reference point.", "negative_right": "On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point.", "right": "On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point."}, "verifier_independent_model": false}, "family": "fast-43-diverse-199-020", "id": "fast-43-diverse-199-020-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_clean_4": "Do not route further; mark complete; rate 4. Use only when measurements, construction, trimming, pressing, and final inspection all pass.", "finishing_incomplete_3": "Route to finishing; mark incomplete; rate 3. Use only when all construction is within tolerance and only trimming, pressing, or another minor cosmetic task remains.", "none_of_above": "Select when the evidence requires a different route, completion status, or quality rating than all three listed outcomes.", "stitching_incomplete_1": "Route to stitching; mark incomplete; rate 1. Use only when at least one required construction seam is absent and major reconstruction is needed."}, "instructions": "Use the latest timestamped evidence. Route to measuring if dimensions are unverified; cutting if verified pieces are missing; stitching if a required seam is absent; seam correction if an existing seam exceeds tolerance; finishing if construction passes but trimming or pressing remains; otherwise mark complete. Rate finish quality on this ordered scale: 1 = unusable or requiring major reconstruction, 2 = incomplete with a repairable seam or closure defect, 3 = functionally complete with only minor cosmetic finishing needed, 4 = complete and clean. Select the option matching the route, completion decision, and rating.", "type": "choice"}}, "state": "Inspection log update for Priya's 45 cm cushion cover: All panel dimensions were verified against the pattern, and every required cut piece is accounted for. All four construction seams, including the zipper seam, are present and stitched, and every other construction requirement passes inspection. On the latest inspection log for Priya’s 45 cm cushion cover, the measured location of the existing right-edge seam is 10 mm from the reference point. On the same inspection log, the recorded location of the marked line for that right-edge seam is 9 mm from the reference point. Trimming of loose threads has passed. Pressing has passed. Final inspection has passed. The project sheet allows no more than 3 mm seam deviation."}, "method": "c2d", "provenance": {"source_id": "diverse-199", "source_is_synthetic": true, "source_sha256": "b265f00f0f4e8a2417563736156e0a7048590dffeb0cdb8a65a6411601d25284", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_clean_4"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The tag change from OBJ-771 to OBJ-905 on IMG-2 creates a coherent alternative where the puckered-seam photo no longer verifiably depicts Mara's submitted cover (tagged OBJ-771), without contradicting any other evidence; all governing rubric policy remains intact in the unchanged questions object, entity/time bindings (Mara, Lena, project sheet) are unchanged, both focus sentences are complete factual observations, and no answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 \\u00d7 45.0 cm, with \\u00b10.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"evidence\": [\"Ivo measured the laid-flat cover twice as 45.2 x 44.7 cm, both within tolerance.\", \"All four corners of the submitted cover are square, every thread is trimmed, and the cover has been pressed.\", \"Direct inspection of the submitted cover found zero cosmetic defects and zero construction defects outside the zipper-seam area shown in IMG-2.\", \"Every inspected portion of the zipper seam not shown in IMG-2 is even.\", \"Inspection image IMG-2 has an object-tag reading OBJ-771, recorded at intake.\", \"Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log.\", \"IMG-2 shows a puckered zipper seam on the depicted object.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "4"], "text": "Inspection image IMG-2 has an object-tag reading OBJ-771, recorded at intake."}, {"path": ["evidence", "5"], "text": "Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection image IMG-2 has an object-tag reading OBJ-771, recorded at intake.", "negative_left": "Inspection image IMG-2 has an object-tag reading OBJ-905, recorded at intake.", "negative_right": "Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log.", "right": "Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-003", "id": "fast-43-diverse-200-003-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Ivo measured the laid-flat cover twice as 45.2 x 44.7 cm, both within tolerance.", "All four corners of the submitted cover are square, every thread is trimmed, and the cover has been pressed.", "Direct inspection of the submitted cover found zero cosmetic defects and zero construction defects outside the zipper-seam area shown in IMG-2.", "Every inspected portion of the zipper seam not shown in IMG-2 is even.", "Inspection image IMG-2 has an object-tag reading OBJ-771, recorded at intake.", "Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log.", "IMG-2 shows a puckered zipper seam on the depicted object."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The tag change from OBJ-771 to OBJ-905 on IMG-2 creates a coherent alternative where the puckered-seam photo no longer verifiably depicts Mara's submitted cover (tagged OBJ-771), without contradicting any other evidence; all governing rubric policy remains intact in the unchanged questions object, entity/time bindings (Mara, Lena, project sheet) are unchanged, both focus sentences are complete factual observations, and no answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 \\u00d7 45.0 cm, with \\u00b10.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"evidence\": [\"Ivo measured the laid-flat cover twice as 45.2 x 44.7 cm, both within tolerance.\", \"All four corners of the submitted cover are square, every thread is trimmed, and the cover has been pressed.\", \"Direct inspection of the submitted cover found zero cosmetic defects and zero construction defects outside the zipper-seam area shown in IMG-2.\", \"Every inspected portion of the zipper seam not shown in IMG-2 is even.\", \"Inspection image IMG-2 has an object-tag reading OBJ-771, recorded at intake.\", \"Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log.\", \"IMG-2 shows a puckered zipper seam on the depicted object.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "4"], "text": "Inspection image IMG-2 has an object-tag reading OBJ-771, recorded at intake."}, {"path": ["evidence", "5"], "text": "Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection image IMG-2 has an object-tag reading OBJ-771, recorded at intake.", "negative_left": "Inspection image IMG-2 has an object-tag reading OBJ-905, recorded at intake.", "negative_right": "Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log.", "right": "Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-003", "id": "fast-43-diverse-200-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Ivo measured the laid-flat cover twice as 45.2 x 44.7 cm, both within tolerance.", "All four corners of the submitted cover are square, every thread is trimmed, and the cover has been pressed.", "Direct inspection of the submitted cover found zero cosmetic defects and zero construction defects outside the zipper-seam area shown in IMG-2.", "Every inspected portion of the zipper seam not shown in IMG-2 is even.", "Inspection image IMG-2 has an object-tag reading OBJ-905, recorded at intake.", "Mara's submitted zippered cushion cover is tagged OBJ-771 in the submission log.", "IMG-2 shows a puckered zipper seam on the depicted object."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, request scope, and question bindings while only altering the object-identifier fact ('CC-2231' vs 'CC-4407') as a coherent single-sentence counterfactual, and no evidence sentence encodes a rule, rationale, or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Ivo measured the laid-flat cover twice as 45.2 × 44.8 cm, both within tolerance.\",\"The corners are square, threads are trimmed, and the cover has been pressed.\",\"No cosmetic defects were found anywhere on the cover during direct inspection.\",\"Except for the seam portion shown in inspection image IMG-2, the rest of the zipper seam is even, and no construction defects were found elsewhere.\",\"Inspection image IMG-2 shows a puckered zipper seam on the depicted object.\",\"The inspection tag on the object shown in image IMG-2 reads object identifier CC-2231.\",\"The submission log lists Mara's zippered cushion cover under object identifier CC-2231.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "The inspection tag on the object shown in image IMG-2 reads object identifier CC-2231."}, {"path": ["evidence", "6"], "text": "The submission log lists Mara's zippered cushion cover under object identifier CC-2231."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection tag on the object shown in image IMG-2 reads object identifier CC-2231.", "negative_left": "The inspection tag on the object shown in image IMG-2 reads object identifier CC-4407.", "negative_right": "The submission log lists Mara's zippered cushion cover under object identifier CC-2231.", "right": "The submission log lists Mara's zippered cushion cover under object identifier CC-2231."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-007", "id": "fast-43-diverse-200-007-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Ivo measured the laid-flat cover twice as 45.2 × 44.8 cm, both within tolerance.", "The corners are square, threads are trimmed, and the cover has been pressed.", "No cosmetic defects were found anywhere on the cover during direct inspection.", "Except for the seam portion shown in inspection image IMG-2, the rest of the zipper seam is even, and no construction defects were found elsewhere.", "Inspection image IMG-2 shows a puckered zipper seam on the depicted object.", "The inspection tag on the object shown in image IMG-2 reads object identifier CC-2231.", "The submission log lists Mara's zippered cushion cover under object identifier CC-2231."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy, request scope, and question bindings while only altering the object-identifier fact ('CC-2231' vs 'CC-4407') as a coherent single-sentence counterfactual, and no evidence sentence encodes a rule, rationale, or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Ivo measured the laid-flat cover twice as 45.2 × 44.8 cm, both within tolerance.\",\"The corners are square, threads are trimmed, and the cover has been pressed.\",\"No cosmetic defects were found anywhere on the cover during direct inspection.\",\"Except for the seam portion shown in inspection image IMG-2, the rest of the zipper seam is even, and no construction defects were found elsewhere.\",\"Inspection image IMG-2 shows a puckered zipper seam on the depicted object.\",\"The inspection tag on the object shown in image IMG-2 reads object identifier CC-2231.\",\"The submission log lists Mara's zippered cushion cover under object identifier CC-2231.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "The inspection tag on the object shown in image IMG-2 reads object identifier CC-2231."}, {"path": ["evidence", "6"], "text": "The submission log lists Mara's zippered cushion cover under object identifier CC-2231."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection tag on the object shown in image IMG-2 reads object identifier CC-2231.", "negative_left": "The inspection tag on the object shown in image IMG-2 reads object identifier CC-4407.", "negative_right": "The submission log lists Mara's zippered cushion cover under object identifier CC-2231.", "right": "The submission log lists Mara's zippered cushion cover under object identifier CC-2231."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-007", "id": "fast-43-diverse-200-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Ivo measured the laid-flat cover twice as 45.2 × 44.8 cm, both within tolerance.", "The corners are square, threads are trimmed, and the cover has been pressed.", "No cosmetic defects were found anywhere on the cover during direct inspection.", "Except for the seam portion shown in inspection image IMG-2, the rest of the zipper seam is even, and no construction defects were found elsewhere.", "Inspection image IMG-2 shows a puckered zipper seam on the depicted object.", "The inspection tag on the object shown in image IMG-2 reads object identifier CC-4407.", "The submission log lists Mara's zippered cushion cover under object identifier CC-2231."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same entities, dimension tolerances, and construction requirements while the questions object (with its ordered rubric and seam‑defect override rule) is unchanged, so governing policy and bindings are preserved; the two focus evidence spans are complete factual statements, not rules; the counterfactual merely swaps the IMG‑2 metadata ID (CUSH‑7734→CUSH‑9021), a plausible real‑world mismatch rather than a logical contradiction; neither context contains gold labels, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Mara marked the project 'complete' and recorded 45.0 × 45.0 cm.\",\"Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.8 × 45.2 cm.\",\"Corners are square, threads are trimmed, and the cover has been pressed.\",\"Direct inspection found zero cosmetic defects on the cover.\",\"Every inspected portion of the zipper seam not shown in IMG-2 is even.\",\"Inspection image IMG-2 shows a puckered zipper seam on the depicted object.\",\"No construction defects were found outside any zipper-seam portion shown in IMG-2.\",\"The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734.\",\"The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-7734.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "7"], "text": "The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734."}, {"path": ["evidence", "8"], "text": "The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-7734."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734.", "negative_left": "The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734.", "negative_right": "The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-9021.", "right": "The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-7734."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-008", "id": "fast-43-diverse-200-008-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Mara marked the project 'complete' and recorded 45.0 × 45.0 cm.", "Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.8 × 45.2 cm.", "Corners are square, threads are trimmed, and the cover has been pressed.", "Direct inspection found zero cosmetic defects on the cover.", "Every inspected portion of the zipper seam not shown in IMG-2 is even.", "Inspection image IMG-2 shows a puckered zipper seam on the depicted object.", "No construction defects were found outside any zipper-seam portion shown in IMG-2.", "The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734.", "The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-7734."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same entities, dimension tolerances, and construction requirements while the questions object (with its ordered rubric and seam‑defect override rule) is unchanged, so governing policy and bindings are preserved; the two focus evidence spans are complete factual statements, not rules; the counterfactual merely swaps the IMG‑2 metadata ID (CUSH‑7734→CUSH‑9021), a plausible real‑world mismatch rather than a logical contradiction; neither context contains gold labels, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Mara marked the project 'complete' and recorded 45.0 × 45.0 cm.\",\"Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.8 × 45.2 cm.\",\"Corners are square, threads are trimmed, and the cover has been pressed.\",\"Direct inspection found zero cosmetic defects on the cover.\",\"Every inspected portion of the zipper seam not shown in IMG-2 is even.\",\"Inspection image IMG-2 shows a puckered zipper seam on the depicted object.\",\"No construction defects were found outside any zipper-seam portion shown in IMG-2.\",\"The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734.\",\"The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-7734.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "7"], "text": "The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734."}, {"path": ["evidence", "8"], "text": "The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-7734."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734.", "negative_left": "The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734.", "negative_right": "The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-9021.", "right": "The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-7734."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-008", "id": "fast-43-diverse-200-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Mara marked the project 'complete' and recorded 45.0 × 45.0 cm.", "Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.8 × 45.2 cm.", "Corners are square, threads are trimmed, and the cover has been pressed.", "Direct inspection found zero cosmetic defects on the cover.", "Every inspected portion of the zipper seam not shown in IMG-2 is even.", "Inspection image IMG-2 shows a puckered zipper seam on the depicted object.", "No construction defects were found outside any zipper-seam portion shown in IMG-2.", "The inspection log records the object identifier for Mara's submitted zippered cushion cover as CUSH-7734.", "The metadata attached to inspection image IMG-2 lists the depicted object's identifier as CUSH-9021."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions object with all rubric criteria and instructions intact, so governing policy is preserved. The entity (Mara's cushion for Lena), size/tolerance, and request scope remain identical in both contexts. All evidence items are complete factual statements (e.g., 'Checker Ivo measured the laid-flat cover as 44.9 × 45.2 cm.' and 'Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record.') with no policy text or instructions embedded. The counterfactual changes only the IMG-2 lab-log identifier from CUSH-7741 to CUSH-9902, a single coherent factual edit that does not contradict any other measurement or count. Neither context reveals a gold answer, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Checker Ivo measured the laid-flat cover as 44.9 × 45.2 cm.\",\"All corners of the cover were confirmed square, threads were fully trimmed, and the cover has been pressed.\",\"No cosmetic defects were found anywhere on the submitted cover during direct inspection.\",\"Apart from the zipper seam, no construction defects were found outside the region shown in IMG-2.\",\"Every inspected portion of the zipper seam not shown in IMG-2 was found to be even.\",\"Inspection image IMG-2 is labeled with object identifier CUSH-7741 in the lab log.\",\"Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record.\",\"IMG-2 shows a puckered zipper seam on the pictured object.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "Inspection image IMG-2 is labeled with object identifier CUSH-7741 in the lab log."}, {"path": ["evidence", "6"], "text": "Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection image IMG-2 is labeled with object identifier CUSH-7741 in the lab log.", "negative_left": "Inspection image IMG-2 is labeled with object identifier CUSH-9902 in the lab log.", "negative_right": "Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record.", "right": "Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-010", "id": "fast-43-diverse-200-010-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Checker Ivo measured the laid-flat cover as 44.9 × 45.2 cm.", "All corners of the cover were confirmed square, threads were fully trimmed, and the cover has been pressed.", "No cosmetic defects were found anywhere on the submitted cover during direct inspection.", "Apart from the zipper seam, no construction defects were found outside the region shown in IMG-2.", "Every inspected portion of the zipper seam not shown in IMG-2 was found to be even.", "Inspection image IMG-2 is labeled with object identifier CUSH-7741 in the lab log.", "Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record.", "IMG-2 shows a puckered zipper seam on the pictured object."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged questions object with all rubric criteria and instructions intact, so governing policy is preserved. The entity (Mara's cushion for Lena), size/tolerance, and request scope remain identical in both contexts. All evidence items are complete factual statements (e.g., 'Checker Ivo measured the laid-flat cover as 44.9 × 45.2 cm.' and 'Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record.') with no policy text or instructions embedded. The counterfactual changes only the IMG-2 lab-log identifier from CUSH-7741 to CUSH-9902, a single coherent factual edit that does not contradict any other measurement or count. Neither context reveals a gold answer, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Checker Ivo measured the laid-flat cover as 44.9 × 45.2 cm.\",\"All corners of the cover were confirmed square, threads were fully trimmed, and the cover has been pressed.\",\"No cosmetic defects were found anywhere on the submitted cover during direct inspection.\",\"Apart from the zipper seam, no construction defects were found outside the region shown in IMG-2.\",\"Every inspected portion of the zipper seam not shown in IMG-2 was found to be even.\",\"Inspection image IMG-2 is labeled with object identifier CUSH-7741 in the lab log.\",\"Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record.\",\"IMG-2 shows a puckered zipper seam on the pictured object.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "Inspection image IMG-2 is labeled with object identifier CUSH-7741 in the lab log."}, {"path": ["evidence", "6"], "text": "Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection image IMG-2 is labeled with object identifier CUSH-7741 in the lab log.", "negative_left": "Inspection image IMG-2 is labeled with object identifier CUSH-9902 in the lab log.", "negative_right": "Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record.", "right": "Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-010", "id": "fast-43-diverse-200-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Checker Ivo measured the laid-flat cover as 44.9 × 45.2 cm.", "All corners of the cover were confirmed square, threads were fully trimmed, and the cover has been pressed.", "No cosmetic defects were found anywhere on the submitted cover during direct inspection.", "Apart from the zipper seam, no construction defects were found outside the region shown in IMG-2.", "Every inspected portion of the zipper seam not shown in IMG-2 was found to be even.", "Inspection image IMG-2 is labeled with object identifier CUSH-9902 in the lab log.", "Mara's submitted zippered cushion cover is tagged with object identifier CUSH-7741 in the intake record.", "IMG-2 shows a puckered zipper seam on the pictured object."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy, request, and measurements, only varying the IMG-2 tag identifier (CC-7741 vs CC-9028) to change whether the defect image matches Mara's object CC-7741, which is a coherent single-fact change; evidence spans are complete factual sentences with no embedded labels or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"inspection_notes\":\"Ivo measured the laid-flat cover twice at 45.2 × 44.8 cm, both within tolerance. Every corner is square, all threads are trimmed, and the cover has been pressed. No cosmetic defects were found by direct inspection. Every inspected portion of the zipper seam not shown in image IMG-2 is even, and no construction defects were found outside the IMG-2 area. Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image. The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-7741. Mara's submitted zippered cushion cover is recorded under object identifier CC-7741.\",\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["inspection_notes"], "text": "The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-7741."}, {"path": ["inspection_notes"], "text": "Mara's submitted zippered cushion cover is recorded under object identifier CC-7741."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-7741.", "negative_left": "The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-9028.", "negative_right": "Mara's submitted zippered cushion cover is recorded under object identifier CC-7741.", "right": "Mara's submitted zippered cushion cover is recorded under object identifier CC-7741."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-012", "id": "fast-43-diverse-200-012-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "inspection_notes": "Ivo measured the laid-flat cover twice at 45.2 × 44.8 cm, both within tolerance. Every corner is square, all threads are trimmed, and the cover has been pressed. No cosmetic defects were found by direct inspection. Every inspected portion of the zipper seam not shown in image IMG-2 is even, and no construction defects were found outside the IMG-2 area. Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image. The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-7741. Mara's submitted zippered cushion cover is recorded under object identifier CC-7741.", "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy, request, and measurements, only varying the IMG-2 tag identifier (CC-7741 vs CC-9028) to change whether the defect image matches Mara's object CC-7741, which is a coherent single-fact change; evidence spans are complete factual sentences with no embedded labels or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"inspection_notes\":\"Ivo measured the laid-flat cover twice at 45.2 × 44.8 cm, both within tolerance. Every corner is square, all threads are trimmed, and the cover has been pressed. No cosmetic defects were found by direct inspection. Every inspected portion of the zipper seam not shown in image IMG-2 is even, and no construction defects were found outside the IMG-2 area. Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image. The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-7741. Mara's submitted zippered cushion cover is recorded under object identifier CC-7741.\",\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["inspection_notes"], "text": "The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-7741."}, {"path": ["inspection_notes"], "text": "Mara's submitted zippered cushion cover is recorded under object identifier CC-7741."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-7741.", "negative_left": "The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-9028.", "negative_right": "Mara's submitted zippered cushion cover is recorded under object identifier CC-7741.", "right": "Mara's submitted zippered cushion cover is recorded under object identifier CC-7741."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-012", "id": "fast-43-diverse-200-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "inspection_notes": "Ivo measured the laid-flat cover twice at 45.2 × 44.8 cm, both within tolerance. Every corner is square, all threads are trimmed, and the cover has been pressed. No cosmetic defects were found by direct inspection. Every inspected portion of the zipper seam not shown in image IMG-2 is even, and no construction defects were found outside the IMG-2 area. Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image. The inspection tag on the object shown in inspection image IMG-2 reads object identifier CC-9028. Mara's submitted zippered cushion cover is recorded under object identifier CC-7741.", "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only swaps IMG-2's intake-log identifier from CB-7741 to CB-9052, breaking the link between the pictured puckered seam and Mara's actual submission (CB-7741), which is a coherent single-fact change; policy, request, and question bindings remain untouched, and both focus evidence sentences are complete factual statements with no embedded answers or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"evidence\": [\"Ivo measured the laid-flat cover as 45.0 × 45.0 cm, within tolerance on both dimensions.\", \"All four corners are square, all threads are trimmed, and the cover has been pressed.\", \"Direct inspection of the cover found zero cosmetic defects and zero construction defects outside the zipper-seam portion shown in IMG-2.\", \"Every inspected portion of the cover's zipper seam not part of the object shown in IMG-2 is even.\", \"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\", \"Inspection image IMG-2 is labeled with object identifier CB-7741 in the intake log.\", \"Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "Inspection image IMG-2 is labeled with object identifier CB-7741 in the intake log."}, {"path": ["evidence", "6"], "text": "Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection image IMG-2 is labeled with object identifier CB-7741 in the intake log.", "negative_left": "Inspection image IMG-2 is labeled with object identifier CB-9052 in the intake log.", "negative_right": "Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log.", "right": "Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-019", "id": "fast-43-diverse-200-019-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Ivo measured the laid-flat cover as 45.0 × 45.0 cm, within tolerance on both dimensions.", "All four corners are square, all threads are trimmed, and the cover has been pressed.", "Direct inspection of the cover found zero cosmetic defects and zero construction defects outside the zipper-seam portion shown in IMG-2.", "Every inspected portion of the cover's zipper seam not part of the object shown in IMG-2 is even.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "Inspection image IMG-2 is labeled with object identifier CB-7741 in the intake log.", "Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The counterfactual only swaps IMG-2's intake-log identifier from CB-7741 to CB-9052, breaking the link between the pictured puckered seam and Mara's actual submission (CB-7741), which is a coherent single-fact change; policy, request, and question bindings remain untouched, and both focus evidence sentences are complete factual statements with no embedded answers or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"evidence\": [\"Ivo measured the laid-flat cover as 45.0 × 45.0 cm, within tolerance on both dimensions.\", \"All four corners are square, all threads are trimmed, and the cover has been pressed.\", \"Direct inspection of the cover found zero cosmetic defects and zero construction defects outside the zipper-seam portion shown in IMG-2.\", \"Every inspected portion of the cover's zipper seam not part of the object shown in IMG-2 is even.\", \"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\", \"Inspection image IMG-2 is labeled with object identifier CB-7741 in the intake log.\", \"Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "Inspection image IMG-2 is labeled with object identifier CB-7741 in the intake log."}, {"path": ["evidence", "6"], "text": "Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection image IMG-2 is labeled with object identifier CB-7741 in the intake log.", "negative_left": "Inspection image IMG-2 is labeled with object identifier CB-9052 in the intake log.", "negative_right": "Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log.", "right": "Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-019", "id": "fast-43-diverse-200-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Ivo measured the laid-flat cover as 45.0 × 45.0 cm, within tolerance on both dimensions.", "All four corners are square, all threads are trimmed, and the cover has been pressed.", "Direct inspection of the cover found zero cosmetic defects and zero construction defects outside the zipper-seam portion shown in IMG-2.", "Every inspected portion of the cover's zipper seam not part of the object shown in IMG-2 is even.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "Inspection image IMG-2 is labeled with object identifier CB-9052 in the intake log.", "Mara's submitted zippered cushion cover is labeled with object identifier CB-7741 in the intake log."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full rubric/instructions via the unchanged questions object and keep Mara/Lena/cover bindings intact; the focus evidence sentences are plain factual identifier statements, not policy or instructions; the counterfactual changes only the IMG-2 object's stitched-tag code from CVR-7731-B to CVR-4420-C, coherently decoupling the puckered-seam image from Mara's cover without contradicting any other note; no gold labels, rule tables, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"Household craft station: Mara's zippered cushion cover for Lena. The project sheet requires a finished size of 45.0 \\u00d7 45.0 cm, with \\u00b10.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"inspection_notes\": [\"Ivo measured the laid-flat cover twice, recording 44.8 x 45.2 cm, both within the 0.5 cm tolerance of 45.0 cm.\", \"All four corners were checked with a square and found true.\", \"No loose or untrimmed threads remained anywhere on the cover.\", \"The cover has been pressed flat.\", \"Direct inspection found zero cosmetic defects on the cover.\", \"Every inspected portion of the zipper seam not shown in IMG-2 was found even.\", \"Outside any zipper-seam portion shown in IMG-2, zero construction defects were found.\", \"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\", \"The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-7731-B.\", \"The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["inspection_notes", "8"], "text": "The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-7731-B."}, {"path": ["inspection_notes", "9"], "text": "The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-7731-B.", "negative_left": "The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-4420-C.", "negative_right": "The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B.", "right": "The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-020", "id": "fast-43-diverse-200-020-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "Household craft station: Mara's zippered cushion cover for Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "inspection_notes": ["Ivo measured the laid-flat cover twice, recording 44.8 x 45.2 cm, both within the 0.5 cm tolerance of 45.0 cm.", "All four corners were checked with a square and found true.", "No loose or untrimmed threads remained anywhere on the cover.", "The cover has been pressed flat.", "Direct inspection found zero cosmetic defects on the cover.", "Every inspected portion of the zipper seam not shown in IMG-2 was found even.", "Outside any zipper-seam portion shown in IMG-2, zero construction defects were found.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-7731-B.", "The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full rubric/instructions via the unchanged questions object and keep Mara/Lena/cover bindings intact; the focus evidence sentences are plain factual identifier statements, not policy or instructions; the counterfactual changes only the IMG-2 object's stitched-tag code from CVR-7731-B to CVR-4420-C, coherently decoupling the puckered-seam image from Mara's cover without contradicting any other note; no gold labels, rule tables, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"Household craft station: Mara's zippered cushion cover for Lena. The project sheet requires a finished size of 45.0 \\u00d7 45.0 cm, with \\u00b10.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"inspection_notes\": [\"Ivo measured the laid-flat cover twice, recording 44.8 x 45.2 cm, both within the 0.5 cm tolerance of 45.0 cm.\", \"All four corners were checked with a square and found true.\", \"No loose or untrimmed threads remained anywhere on the cover.\", \"The cover has been pressed flat.\", \"Direct inspection found zero cosmetic defects on the cover.\", \"Every inspected portion of the zipper seam not shown in IMG-2 was found even.\", \"Outside any zipper-seam portion shown in IMG-2, zero construction defects were found.\", \"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\", \"The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-7731-B.\", \"The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["inspection_notes", "8"], "text": "The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-7731-B."}, {"path": ["inspection_notes", "9"], "text": "The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-7731-B.", "negative_left": "The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-4420-C.", "negative_right": "The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B.", "right": "The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-020", "id": "fast-43-diverse-200-020-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "Household craft station: Mara's zippered cushion cover for Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "inspection_notes": ["Ivo measured the laid-flat cover twice, recording 44.8 x 45.2 cm, both within the 0.5 cm tolerance of 45.0 cm.", "All four corners were checked with a square and found true.", "No loose or untrimmed threads remained anywhere on the cover.", "The cover has been pressed flat.", "Direct inspection found zero cosmetic defects on the cover.", "Every inspected portion of the zipper seam not shown in IMG-2 was found even.", "Outside any zipper-seam portion shown in IMG-2, zero construction defects were found.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "The object identifier recorded for the object shown in inspection image IMG-2 is stitched-tag code CVR-4420-C.", "The object identifier recorded for Mara's submitted zippered cushion cover is stitched-tag code CVR-7731-B."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original tolerance, defect, and identifier policy language unaltered, with only the IMG-2 metadata ID changed to CUSH-8021, producing a coherent identity mismatch rather than a contradiction; entities, request, and rubric criteria remain bound to the original question; the two focus evidence spans are complete factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"evidence\": [\"Mara marked the project \\u201ccomplete\\u201d and recorded 45.0 × 45.0 cm.\", \"Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.8 × 45.0 cm.\", \"Ivo confirmed the corners are square, all threads are trimmed, and the cover has been pressed.\", \"Ivo found zero cosmetic defects on the cover during direct inspection, and zero construction defects outside any zipper-seam portion shown in inspection image IMG-2.\", \"Every inspected portion of the cover's zipper seam that is not part of the object shown in IMG-2 was found even.\", \"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\", \"The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749.\", \"The object identifier recorded in inspection image IMG-2's file metadata is CUSH-7749.\", \"Lena said the cover looks attractive but has not used it.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "6"], "text": "The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749."}, {"path": ["evidence", "7"], "text": "The object identifier recorded in inspection image IMG-2's file metadata is CUSH-7749."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749.", "negative_left": "The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749.", "negative_right": "The object identifier recorded in inspection image IMG-2's file metadata is CUSH-8021.", "right": "The object identifier recorded in inspection image IMG-2's file metadata is CUSH-7749."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-021", "id": "fast-43-diverse-200-021-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Mara marked the project “complete” and recorded 45.0 × 45.0 cm.", "Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.8 × 45.0 cm.", "Ivo confirmed the corners are square, all threads are trimmed, and the cover has been pressed.", "Ivo found zero cosmetic defects on the cover during direct inspection, and zero construction defects outside any zipper-seam portion shown in inspection image IMG-2.", "Every inspected portion of the cover's zipper seam that is not part of the object shown in IMG-2 was found even.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749.", "The object identifier recorded in inspection image IMG-2's file metadata is CUSH-7749.", "Lena said the cover looks attractive but has not used it."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original tolerance, defect, and identifier policy language unaltered, with only the IMG-2 metadata ID changed to CUSH-8021, producing a coherent identity mismatch rather than a contradiction; entities, request, and rubric criteria remain bound to the original question; the two focus evidence spans are complete factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"evidence\": [\"Mara marked the project \\u201ccomplete\\u201d and recorded 45.0 × 45.0 cm.\", \"Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.8 × 45.0 cm.\", \"Ivo confirmed the corners are square, all threads are trimmed, and the cover has been pressed.\", \"Ivo found zero cosmetic defects on the cover during direct inspection, and zero construction defects outside any zipper-seam portion shown in inspection image IMG-2.\", \"Every inspected portion of the cover's zipper seam that is not part of the object shown in IMG-2 was found even.\", \"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\", \"The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749.\", \"The object identifier recorded in inspection image IMG-2's file metadata is CUSH-7749.\", \"Lena said the cover looks attractive but has not used it.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "6"], "text": "The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749."}, {"path": ["evidence", "7"], "text": "The object identifier recorded in inspection image IMG-2's file metadata is CUSH-7749."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749.", "negative_left": "The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749.", "negative_right": "The object identifier recorded in inspection image IMG-2's file metadata is CUSH-8021.", "right": "The object identifier recorded in inspection image IMG-2's file metadata is CUSH-7749."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-021", "id": "fast-43-diverse-200-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Mara marked the project “complete” and recorded 45.0 × 45.0 cm.", "Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.8 × 45.0 cm.", "Ivo confirmed the corners are square, all threads are trimmed, and the cover has been pressed.", "Ivo found zero cosmetic defects on the cover during direct inspection, and zero construction defects outside any zipper-seam portion shown in inspection image IMG-2.", "Every inspected portion of the cover's zipper seam that is not part of the object shown in IMG-2 was found even.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "The object identifier recorded on Mara's submitted zippered cushion cover's tag is CUSH-7749.", "The object identifier recorded in inspection image IMG-2's file metadata is CUSH-8021.", "Lena said the cover looks attractive but has not used it."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, tolerance and seam-defect policy language verbatim in the question, and the two focus sentences are plain factual statements about tag IDs rather than rules or answers; the counterfactual only swaps the IMG-2 tag ID (CC-2093 vs CC-1147), which coherently changes whether the puckered-seam image matches Mara's logged cover without contradicting any other evidence or leaking the rubric outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"evidence\": [\"Mara marked the project \\u201ccomplete\\u201d and recorded 45.0 × 45.0 cm.\", \"Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.9 × 45.2 cm, both within tolerance.\", \"Ivo confirmed all corners are square, all threads are trimmed, and the cover has been pressed.\", \"Ivo found zero cosmetic defects and zero construction defects outside any zipper-seam portion shown in image IMG-2.\", \"Ivo noted that every inspected portion of the cover's zipper seam not shown in IMG-2 is even.\", \"Image IMG-2 shows a puckered zipper seam on the object it depicts.\", \"The inspection tag on the object shown in image IMG-2 reads object ID CC-1147.\", \"Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet.\", \"Lena said the cover looks attractive but has not used it.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "6"], "text": "The inspection tag on the object shown in image IMG-2 reads object ID CC-1147."}, {"path": ["evidence", "7"], "text": "Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection tag on the object shown in image IMG-2 reads object ID CC-1147.", "negative_left": "The inspection tag on the object shown in image IMG-2 reads object ID CC-2093.", "negative_right": "Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet.", "right": "Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-022", "id": "fast-43-diverse-200-022-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Mara marked the project “complete” and recorded 45.0 × 45.0 cm.", "Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.9 × 45.2 cm, both within tolerance.", "Ivo confirmed all corners are square, all threads are trimmed, and the cover has been pressed.", "Ivo found zero cosmetic defects and zero construction defects outside any zipper-seam portion shown in image IMG-2.", "Ivo noted that every inspected portion of the cover's zipper seam not shown in IMG-2 is even.", "Image IMG-2 shows a puckered zipper seam on the object it depicts.", "The inspection tag on the object shown in image IMG-2 reads object ID CC-1147.", "Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet.", "Lena said the cover looks attractive but has not used it."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same entities, tolerance and seam-defect policy language verbatim in the question, and the two focus sentences are plain factual statements about tag IDs rather than rules or answers; the counterfactual only swaps the IMG-2 tag ID (CC-2093 vs CC-1147), which coherently changes whether the puckered-seam image matches Mara's logged cover without contradicting any other evidence or leaking the rubric outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"evidence\": [\"Mara marked the project \\u201ccomplete\\u201d and recorded 45.0 × 45.0 cm.\", \"Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.9 × 45.2 cm, both within tolerance.\", \"Ivo confirmed all corners are square, all threads are trimmed, and the cover has been pressed.\", \"Ivo found zero cosmetic defects and zero construction defects outside any zipper-seam portion shown in image IMG-2.\", \"Ivo noted that every inspected portion of the cover's zipper seam not shown in IMG-2 is even.\", \"Image IMG-2 shows a puckered zipper seam on the object it depicts.\", \"The inspection tag on the object shown in image IMG-2 reads object ID CC-1147.\", \"Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet.\", \"Lena said the cover looks attractive but has not used it.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "6"], "text": "The inspection tag on the object shown in image IMG-2 reads object ID CC-1147."}, {"path": ["evidence", "7"], "text": "Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection tag on the object shown in image IMG-2 reads object ID CC-1147.", "negative_left": "The inspection tag on the object shown in image IMG-2 reads object ID CC-2093.", "negative_right": "Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet.", "right": "Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-022", "id": "fast-43-diverse-200-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Mara marked the project “complete” and recorded 45.0 × 45.0 cm.", "Pattern and measurement checker Ivo measured the laid-flat cover twice as 44.9 × 45.2 cm, both within tolerance.", "Ivo confirmed all corners are square, all threads are trimmed, and the cover has been pressed.", "Ivo found zero cosmetic defects and zero construction defects outside any zipper-seam portion shown in image IMG-2.", "Ivo noted that every inspected portion of the cover's zipper seam not shown in IMG-2 is even.", "Image IMG-2 shows a puckered zipper seam on the object it depicts.", "The inspection tag on the object shown in image IMG-2 reads object ID CC-2093.", "Mara's submitted zippered cushion cover was logged under object ID CC-1147 on the intake sheet.", "Lena said the cover looks attractive but has not used it."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The single-sentence tag change (OBJ-9915 vs OBJ-7742) plausibly decouples IMG-2's defect from Mara's cover without contradicting other measurements or counts, and both contexts keep the unchanged questions object, entities, and policy intact with no embedded rules or gold labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 \\u00d7 45.0 cm, with \\u00b10.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"inspection_log\": [\"Ivo measured the laid-flat cover twice, confirming 44.9 \\u00d7 45.2 cm, both within tolerance.\", \"All four corners of the cover were checked and found square.\", \"All visible threads on the cover are trimmed, and the cover has been pressed.\", \"No cosmetic defects were found anywhere on the cover during direct inspection.\", \"Every inspected portion of the cover's zipper seam not shown in image IMG-2 was found even.\", \"No construction defects were found outside any zipper-seam portion shown in image IMG-2.\", \"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\", \"Inspection image IMG-2 has the object identifier tag OBJ-7742 recorded in the inspection log.\", \"Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["inspection_log", "7"], "text": "Inspection image IMG-2 has the object identifier tag OBJ-7742 recorded in the inspection log."}, {"path": ["inspection_log", "8"], "text": "Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection image IMG-2 has the object identifier tag OBJ-7742 recorded in the inspection log.", "negative_left": "Inspection image IMG-2 has the object identifier tag OBJ-9915 recorded in the inspection log.", "negative_right": "Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log.", "right": "Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-023", "id": "fast-43-diverse-200-023-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "inspection_log": ["Ivo measured the laid-flat cover twice, confirming 44.9 × 45.2 cm, both within tolerance.", "All four corners of the cover were checked and found square.", "All visible threads on the cover are trimmed, and the cover has been pressed.", "No cosmetic defects were found anywhere on the cover during direct inspection.", "Every inspected portion of the cover's zipper seam not shown in image IMG-2 was found even.", "No construction defects were found outside any zipper-seam portion shown in image IMG-2.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "Inspection image IMG-2 has the object identifier tag OBJ-7742 recorded in the inspection log.", "Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The single-sentence tag change (OBJ-9915 vs OBJ-7742) plausibly decouples IMG-2's defect from Mara's cover without contradicting other measurements or counts, and both contexts keep the unchanged questions object, entities, and policy intact with no embedded rules or gold labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\": \"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 \\u00d7 45.0 cm, with \\u00b10.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\", \"inspection_log\": [\"Ivo measured the laid-flat cover twice, confirming 44.9 \\u00d7 45.2 cm, both within tolerance.\", \"All four corners of the cover were checked and found square.\", \"All visible threads on the cover are trimmed, and the cover has been pressed.\", \"No cosmetic defects were found anywhere on the cover during direct inspection.\", \"Every inspected portion of the cover's zipper seam not shown in image IMG-2 was found even.\", \"No construction defects were found outside any zipper-seam portion shown in image IMG-2.\", \"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\", \"Inspection image IMG-2 has the object identifier tag OBJ-7742 recorded in the inspection log.\", \"Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log.\"], \"request\": \"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["inspection_log", "7"], "text": "Inspection image IMG-2 has the object identifier tag OBJ-7742 recorded in the inspection log."}, {"path": ["inspection_log", "8"], "text": "Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "Inspection image IMG-2 has the object identifier tag OBJ-7742 recorded in the inspection log.", "negative_left": "Inspection image IMG-2 has the object identifier tag OBJ-9915 recorded in the inspection log.", "negative_right": "Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log.", "right": "Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-023", "id": "fast-43-diverse-200-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "inspection_log": ["Ivo measured the laid-flat cover twice, confirming 44.9 × 45.2 cm, both within tolerance.", "All four corners of the cover were checked and found square.", "All visible threads on the cover are trimmed, and the cover has been pressed.", "No cosmetic defects were found anywhere on the cover during direct inspection.", "Every inspected portion of the cover's zipper seam not shown in image IMG-2 was found even.", "No construction defects were found outside any zipper-seam portion shown in image IMG-2.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "Inspection image IMG-2 has the object identifier tag OBJ-9915 recorded in the inspection log.", "Mara's submitted zippered cushion cover has the object identifier tag OBJ-7742 recorded in the inspection log."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question, entities, dimensions, and tolerance policy while using only direct-inspection evidence as instructed; the two focus sentences are plain factual statements, not policy text; the counterfactual coherently reassigns the IMG-2 tag ID to CC-8802 versus the cover's CC-7741, implying the puckered seam belongs to a different object without contradicting any other measurement; no rubric labels, codes, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Ivo measured the laid-flat cover twice, recording 45.2 × 44.8 cm, both within tolerance.\",\"All four corners of the cover are square, all threads are trimmed, and the cover has been pressed.\",\"Direct inspection finds zero cosmetic defects on the cover.\",\"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\",\"Apart from the seam portion shown in IMG-2, every other inspected portion of the cover's zipper seam is even.\",\"No construction defects were found anywhere on the cover outside the seam portion shown in IMG-2.\",\"The inspection tag on the object shown in image IMG-2 reads object ID CC-7741.\",\"Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "6"], "text": "The inspection tag on the object shown in image IMG-2 reads object ID CC-7741."}, {"path": ["evidence", "7"], "text": "Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection tag on the object shown in image IMG-2 reads object ID CC-7741.", "negative_left": "The inspection tag on the object shown in image IMG-2 reads object ID CC-8802.", "negative_right": "Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record.", "right": "Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-026", "id": "fast-43-diverse-200-026-base", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Ivo measured the laid-flat cover twice, recording 45.2 × 44.8 cm, both within tolerance.", "All four corners of the cover are square, all threads are trimmed, and the cover has been pressed.", "Direct inspection finds zero cosmetic defects on the cover.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "Apart from the seam portion shown in IMG-2, every other inspected portion of the cover's zipper seam is even.", "No construction defects were found anywhere on the cover outside the seam portion shown in IMG-2.", "The inspection tag on the object shown in image IMG-2 reads object ID CC-7741.", "Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_seam_correction"}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question, entities, dimensions, and tolerance policy while using only direct-inspection evidence as instructed; the two focus sentences are plain factual statements, not policy text; the counterfactual coherently reassigns the IMG-2 tag ID to CC-8802 versus the cover's CC-7741, implying the puckered seam belongs to a different object without contradicting any other measurement; no rubric labels, codes, or instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universal and zero-count statements remain atomic. A9 is a factual object-identity relation rather than a policy conclusion. The base and counter assignments are jointly realizable: in the base, IMG-2's pucker is a seam-construction defect on Mara's cover while A7 and A10 concern areas outside the depicted seam portion; in the counter, IMG-2 depicts another object, making the defect irrelevant to Mara's otherwise defect-free cover. Only A9 changes. Policy evidence correctly preserves the project specifications and direct-inspection directive originating in the original state; governing classification rules in the unchanged questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A8 establishes a puckered zipper seam in IMG-2, and A9 establishes that the depicted object is Mara's submitted cover. This is a directly inspected zipper/seam-construction defect, which unconditionally requires Needs seam correction under the question's policy.", "rule_index": 0, "sound": true}, {"reason": "A1–A5 establish the dimensional and finishing specifications. With A9 refuted, IMG-2 depicts a different object, so A7 applies to all inspected zipper-seam portions of Mara's cover and A10 excludes construction defects throughout her cover. A6 excludes cosmetic defects. These conditions establish every stated specification and no cosmetic or construction defects, sufficient for Complete—Excellent.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The directly inspected finished width of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required width of 45.0 cm."}, {"id": "A2", "statement": "The directly inspected finished height of Mara's submitted zippered cushion cover is within 0.5 cm of the project-sheet-required height of 45.0 cm."}, {"id": "A3", "statement": "Every corner of Mara's submitted zippered cushion cover is square."}, {"id": "A4", "statement": "Every thread on Mara's submitted zippered cushion cover is trimmed."}, {"id": "A5", "statement": "Mara's submitted zippered cushion cover has been pressed."}, {"id": "A6", "statement": "The number of cosmetic defects found by direct inspection on Mara's submitted zippered cushion cover is zero."}, {"id": "A7", "statement": "Every inspected portion of Mara's submitted cover's zipper seam that is not part of the object shown in inspection image IMG-2 is even."}, {"id": "A8", "statement": "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image."}, {"id": "A9", "statement": "The object identifier recorded for the object shown in inspection image IMG-2 and the object identifier recorded for Mara's submitted zippered cushion cover designate the same physical object."}, {"id": "A10", "statement": "The number of construction defects on Mara's submitted zippered cushion cover outside any zipper-seam portion shown in inspection image IMG-2 is zero."}], "base_state_json": "{\"context\":\"At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.\",\"evidence\":[\"Ivo measured the laid-flat cover twice, recording 45.2 × 44.8 cm, both within tolerance.\",\"All four corners of the cover are square, all threads are trimmed, and the cover has been pressed.\",\"Direct inspection finds zero cosmetic defects on the cover.\",\"Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.\",\"Apart from the seam portion shown in IMG-2, every other inspected portion of the cover's zipper seam is even.\",\"No construction defects were found anywhere on the cover outside the seam portion shown in IMG-2.\",\"The inspection tag on the object shown in image IMG-2 reads object ID CC-7741.\",\"Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record.\"],\"request\":\"Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level.\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "6"], "text": "The inspection tag on the object shown in image IMG-2 reads object ID CC-7741."}, {"path": ["evidence", "7"], "text": "Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record."}], "policy_evidence": [{"path": ["context"], "text": "The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion."}, {"path": ["request"], "text": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}], "rules": [{"justification": "IMG-2 directly shows a puckered zipper seam, and the matching object identifiers establish that the depicted object is Mara's submitted cover. Direct inspection therefore finds a zipper or seam-construction defect, which requires the route_seam_correction choice regardless of a completion mark or recipient preference.", "target": "route_seam_correction", "when": [{"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}]}, {"justification": "Both dimensions are within tolerance, every corner is square, all threads are trimmed, and the cover is pressed. Because IMG-2 depicts a different physical object, every portion of the submitted cover's zipper seam falls within the even-seam observation, and the zero-defect observations cover all cosmetic and construction defects on the submitted cover. Thus every specification is met and direct inspection finds no cosmetic or construction defects.", "target": "complete_excellent", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The inspection tag on the object shown in image IMG-2 reads object ID CC-7741.", "negative_left": "The inspection tag on the object shown in image IMG-2 reads object ID CC-8802.", "negative_right": "Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record.", "right": "Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record."}, "verifier_independent_model": false}, "family": "fast-43-diverse-200-026", "id": "fast-43-diverse-200-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"complete_acceptable": "Complete—Acceptable: every specification is met, with two or more minor cosmetic defects but no dimensional, functional, or seam-construction defect.", "complete_excellent": "Complete—Excellent: every specification is met and direct inspection finds no cosmetic or construction defects.", "complete_good": "Complete—Good: every specification is met, with exactly one minor cosmetic defect that does not involve seam construction or function.", "route_finishing": "Incomplete—Needs finishing: dimensions and construction meet specifications, and the only remaining work is thread trimming or pressing.", "route_seam_correction": "Incomplete—Needs seam correction: at least one checked dimension is outside tolerance or direct inspection finds a zipper, corner, or seam-construction defect."}, "instructions": "Choose exactly one outcome. Ordered levels are Excellent, Good, Acceptable, Needs finishing, and Needs seam correction. Any out-of-tolerance dimension or seam-construction defect requires seam correction regardless of a completion mark or recipient preference.", "type": "choice"}}, "state": {"context": "At the household craft station, sewing hobbyist Mara has made a zippered cushion cover for recipient Lena. The project sheet requires a finished size of 45.0 × 45.0 cm, with ±0.5 cm tolerance, an even zipper seam, square corners, trimmed threads, and pressing before completion.", "evidence": ["Ivo measured the laid-flat cover twice, recording 45.2 × 44.8 cm, both within tolerance.", "All four corners of the cover are square, all threads are trimmed, and the cover has been pressed.", "Direct inspection finds zero cosmetic defects on the cover.", "Inspection image IMG-2 shows a puckered zipper seam on the object depicted in that image.", "Apart from the seam portion shown in IMG-2, every other inspected portion of the cover's zipper seam is even.", "No construction defects were found anywhere on the cover outside the seam portion shown in IMG-2.", "The inspection tag on the object shown in image IMG-2 reads object ID CC-8802.", "Mara's submitted zippered cushion cover is logged under object ID CC-7741 in the intake record."], "request": "Using direct inspection over handwritten status, decide completeness, route the project, and assign the rubric level."}}, "method": "c2d", "provenance": {"source_id": "diverse-200", "source_is_synthetic": true, "source_sha256": "efa40ffb1fc20a77db3bf36783c11b290ef50f7798bb5289f5588a79c835182a", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "complete_excellent"}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the scope, tolerance, exclusion and quality-tier policy language unchanged and the question's entities/instructions intact; the two evidence sentences are plain factual measurements, and the counterfactual only alters the second measurement to 39.2 cm without contradicting any other retained fact or embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"At the household craft station, sewing hobbyist Mara is making two 40 × 40 cm cushion covers for recipient Leo. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Checker Inez confirms both closures work, and photographs trimmed threads and pressed seams. Mara measured the first cushion cover in her stated project scope at 40.3 cm. Mara measured the second cushion cover in her stated project scope at 39.8 cm. One cover has a single 3 mm topstitch wobble; only the excluded inner edges remain raw.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "Mara measured the first cushion cover in her stated project scope at 40.3 cm."}, {"path": [], "text": "Mara measured the second cushion cover in her stated project scope at 39.8 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "Mara measured the first cushion cover in her stated project scope at 40.3 cm.", "negative_left": "Mara measured the first cushion cover in her stated project scope at 40.3 cm.", "negative_right": "Mara measured the second cushion cover in her stated project scope at 39.2 cm.", "right": "Mara measured the second cushion cover in her stated project scope at 39.8 cm."}, "verifier_independent_model": false}, "family": "fast-43-diverse-202-026", "id": "fast-43-diverse-202-026-base", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "At the household craft station, sewing hobbyist Mara is making two 40 × 40 cm cushion covers for recipient Leo. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Checker Inez confirms both closures work, and photographs trimmed threads and pressed seams. Mara measured the first cushion cover in her stated project scope at 40.3 cm. Mara measured the second cushion cover in her stated project scope at 39.8 cm. One cover has a single 3 mm topstitch wobble; only the excluded inner edges remain raw."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "home-04", "split": "train", "variant": "base"} {"domain": "home", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the scope, tolerance, exclusion and quality-tier policy language unchanged and the question's entities/instructions intact; the two evidence sentences are plain factual measurements, and the counterfactual only alters the second measurement to 39.2 cm without contradicting any other retained fact or embedding an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "full_context_fact_states": {"base": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "supported", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "counterfactual": {"a_closures": "supported", "a_cover_count": "supported", "a_dimensions": "refuted", "a_seams": "supported", "a_threads": "supported", "a_wobble_count": "supported", "a_wobble_size": "supported"}, "remove_left": {"a_dimensions": "unknown"}, "remove_right": {"a_dimensions": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_dimensions": "unknown"}, "negative_pair": {"a_dimensions": "refuted"}, "negative_sentence": {"a_dimensions": "unknown"}, "positive_pair": {"a_dimensions": "supported"}, "right": {"a_dimensions": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, including the universally quantified atoms. The focus atom is a factual measurement relation rather than a policy classification. The base and counter assignments are realizable while changing only dimensional compliance, such as changing one recorded cover measurement from within tolerance to 40.6 cm while preserving all other facts. The policy evidence correctly preserves the substantive scope, exception, quality ordering, and Good-limit rules originating in the original state; rules already present in the retained questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a_dimensions entails that at least one of the two covers deviates from 40 cm by more than 0.5 cm, directly violating an in-scope completion requirement. That is sufficient for a false decision regardless of the other facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes both covers, compliant dimensions, working closures, trimmed loose threads, pressed visible seams, and no more than one visible topstitch wobble with every such wobble at most 3 mm. These facts satisfy every stated in-scope requirement and the stated Good-quality limit; the hidden raw edges are explicitly excluded.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_cover_count", "statement": "The number of cushion covers produced within Mara’s stated project scope is two."}, {"id": "a_dimensions", "statement": "For the two cushion covers in Mara’s stated project scope, the larger absolute deviation of their recorded measurements from 40 cm is no greater than 0.5 cm."}, {"id": "a_closures", "statement": "Every cushion cover in Mara’s stated project scope has a working envelope closure."}, {"id": "a_threads", "statement": "Every loose thread on the cushion covers in Mara’s stated project scope is trimmed."}, {"id": "a_seams", "statement": "Every visible seam on the cushion covers in Mara’s stated project scope is pressed."}, {"id": "a_wobble_count", "statement": "The number of visible topstitch wobbles across the two cushion covers in Mara’s stated project scope is no greater than one."}, {"id": "a_wobble_size", "statement": "Every visible topstitch wobble on the two cushion covers in Mara’s stated project scope is no larger than 3 mm."}], "base_state_json": "\"At the household craft station, sewing hobbyist Mara is making two 40 × 40 cm cushion covers for recipient Leo. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Checker Inez confirms both closures work, and photographs trimmed threads and pressed seams. Mara measured the first cushion cover in her stated project scope at 40.3 cm. Mara measured the second cushion cover in her stated project scope at 39.8 cm. One cover has a single 3 mm topstitch wobble; only the excluded inner edges remain raw.\"", "base_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "counter_states": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "refuted"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}], "focus_atom": "a_dimensions", "focus_evidence": [{"path": [], "text": "Mara measured the first cushion cover in her stated project scope at 40.3 cm."}, {"path": [], "text": "Mara measured the second cushion cover in her stated project scope at 39.8 cm."}], "policy_evidence": [{"path": [], "text": "The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams."}, {"path": [], "text": "It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception."}, {"path": [], "text": "Quality is ordered Poor, Fair, Good, Excellent."}, {"path": [], "text": "Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none."}], "rules": [{"justification": "A recorded measurement outside 40 ± 0.5 cm violates an in-scope completion requirement, so the project must be classified false regardless of the unchanged passing facts.", "target": "false", "when": [{"atom_id": "a_dimensions", "state": "refuted"}]}, {"justification": "All in-scope completion requirements are met, the approved hidden-inner-edge exception is outside scope, and at most one visible topstitch wobble of no more than 3 mm falls within the stated Good-quality limit.", "target": "true", "when": [{"atom_id": "a_cover_count", "state": "supported"}, {"atom_id": "a_dimensions", "state": "supported"}, {"atom_id": "a_closures", "state": "supported"}, {"atom_id": "a_threads", "state": "supported"}, {"atom_id": "a_seams", "state": "supported"}, {"atom_id": "a_wobble_count", "state": "supported"}, {"atom_id": "a_wobble_size", "state": "supported"}]}]}, "verified_pair": {"left": "Mara measured the first cushion cover in her stated project scope at 40.3 cm.", "negative_left": "Mara measured the first cushion cover in her stated project scope at 40.3 cm.", "negative_right": "Mara measured the second cushion cover in her stated project scope at 39.2 cm.", "right": "Mara measured the second cushion cover in her stated project scope at 39.8 cm."}, "verifier_independent_model": false}, "family": "fast-43-diverse-202-026", "id": "fast-43-diverse-202-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "The project misses an in-scope completion requirement or has a finish defect exceeding the Good-quality limit.", "true": "The project meets every in-scope completion requirement, and its finish defects fall within the stated Good-quality limit."}, "instructions": "Decide whether the cushion-cover project is complete within its stated scope and qualifies for the finish-quality level Good. Answer yes or no.", "type": "noul"}}, "state": "At the household craft station, sewing hobbyist Mara is making two 40 × 40 cm cushion covers for recipient Leo. The project sheet’s scope requires two covers measuring 40 ± 0.5 cm, working envelope closures, trimmed loose threads, and pressed visible seams. It explicitly excludes finishing the hidden inner raw edges because Leo approved that exception. Quality is ordered Poor, Fair, Good, Excellent. Good permits one visible topstitch wobble no larger than 3 mm; Excellent permits none. Checker Inez confirms both closures work, and photographs trimmed threads and pressed seams. Mara measured the first cushion cover in her stated project scope at 40.3 cm. Mara measured the second cushion cover in her stated project scope at 39.2 cm. One cover has a single 3 mm topstitch wobble; only the excluded inner edges remain raw."}, "method": "c2d", "provenance": {"source_id": "diverse-202", "source_is_synthetic": true, "source_sha256": "e51a5264d14f63f2f1b3d7fe4dbc823aed3b539aed4c4a3011ca733031c663c7", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "home-04", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original incident F entity, encoder-log source, and preshow-countdown timeframe while only altering the end timestamp to shift duration from 2.600s (Level 2) to 1.700s (Level 1), consistently across all statements with no contradictory duplicates; evidence spans are complete factual sentences with no policy text or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"Public encoder log confirms incident F occurred during the preshow countdown and is present in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's end timestamp as 22:14:10.100 during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"No program content was lost following incident F; playback resumed cleanly from the last good frame.\"},{\"speaker\":\"Event producer\",\"text\":\"The stream is stable now and the countdown clock is running normally.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 22:14:10.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 22:14:09.200 during the public preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 22:14:10.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-002", "id": "fast-43-diverse-219-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Public encoder log confirms incident F occurred during the preshow countdown and is present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 22:14:10.100 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "No program content was lost following incident F; playback resumed cleanly from the last good frame."}, {"speaker": "Event producer", "text": "The stream is stable now and the countdown clock is running normally."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original incident F entity, encoder-log source, and preshow-countdown timeframe while only altering the end timestamp to shift duration from 2.600s (Level 2) to 1.700s (Level 1), consistently across all statements with no contradictory duplicates; evidence spans are complete factual sentences with no policy text or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"Public encoder log confirms incident F occurred during the preshow countdown and is present in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's end timestamp as 22:14:10.100 during the public preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"No program content was lost following incident F; playback resumed cleanly from the last good frame.\"},{\"speaker\":\"Event producer\",\"text\":\"The stream is stable now and the countdown clock is running normally.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 22:14:10.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 22:14:09.200 during the public preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 22:14:10.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-002", "id": "fast-43-diverse-219-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Public encoder log confirms incident F occurred during the preshow countdown and is present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 22:14:07.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 22:14:09.200 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "No program content was lost following incident F; playback resumed cleanly from the last good frame."}, {"speaker": "Event producer", "text": "The stream is stable now and the countdown clock is running normally."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions object and rubric untouched, only vary timestamp data for incident F (2.6s vs 1.4s duration) while keeping entity, path, and time-frame bindings intact; the two focus sentences are plain factual log statements, not rule text; the counterfactual's shorter duration remains internally consistent with the unchanged 'no content lost' sentence; no gold label, code, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown, and the defect is present in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100.\"},{\"speaker\":\"Streaming technician\",\"text\":\"After the freeze cleared, playback resumed exactly where it left off, with no program content lost.\"},{\"speaker\":\"Event producer\",\"text\":\"The stream is now stable, and the scheduled opening begins in four minutes.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500."}, {"path": ["2", "text"], "text": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.", "negative_left": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.", "negative_right": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:23.900.", "right": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-003", "id": "fast-43-diverse-219-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F during the public preshow countdown, and the defect is present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500."}, {"speaker": "Streaming technician", "text": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}, {"speaker": "Streaming technician", "text": "After the freeze cleared, playback resumed exactly where it left off, with no program content lost."}, {"speaker": "Event producer", "text": "The stream is now stable, and the scheduled opening begins in four minutes."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions object and rubric untouched, only vary timestamp data for incident F (2.6s vs 1.4s duration) while keeping entity, path, and time-frame bindings intact; the two focus sentences are plain factual log statements, not rule text; the counterfactual's shorter duration remains internally consistent with the unchanged 'no content lost' sentence; no gold label, code, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records video-freeze incident F during the public preshow countdown, and the defect is present in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100.\"},{\"speaker\":\"Streaming technician\",\"text\":\"After the freeze cleared, playback resumed exactly where it left off, with no program content lost.\"},{\"speaker\":\"Event producer\",\"text\":\"The stream is now stable, and the scheduled opening begins in four minutes.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500."}, {"path": ["2", "text"], "text": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.", "negative_left": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.", "negative_right": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:23.900.", "right": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-003", "id": "fast-43-diverse-219-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F during the public preshow countdown, and the defect is present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500."}, {"speaker": "Streaming technician", "text": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:23.900."}, {"speaker": "Streaming technician", "text": "After the freeze cleared, playback resumed exactly where it left off, with no program content lost."}, {"speaker": "Event producer", "text": "The stream is now stable, and the scheduled opening begins in four minutes."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question/rubric and only vary the end timestamp (14:02:10.100 vs 14:02:09.200), shifting duration across the 2.000s threshold without contradicting other facts or revealing the rule outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"The public encoder log records incident F, a video freeze visible in the public encoded output, occurring during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log lists incident F's end timestamp as 14:02:10.100 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"No program content was dropped or lost during incident F.\"}, {\"speaker\": \"Event producer\", \"text\": \"The stream is now stable, and the scheduled opening begins shortly. No viewer complaints have arrived.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 14:02:09.200 during the public preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-004", "id": "fast-43-diverse-219-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log records incident F, a video freeze visible in the public encoded output, occurring during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "No program content was dropped or lost during incident F."}, {"speaker": "Event producer", "text": "The stream is now stable, and the scheduled opening begins shortly. No viewer complaints have arrived."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged question/rubric and only vary the end timestamp (14:02:10.100 vs 14:02:09.200), shifting duration across the 2.000s threshold without contradicting other facts or revealing the rule outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"The public encoder log records incident F, a video freeze visible in the public encoded output, occurring during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log lists incident F's end timestamp as 14:02:10.100 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"No program content was dropped or lost during incident F.\"}, {\"speaker\": \"Event producer\", \"text\": \"The stream is now stable, and the scheduled opening begins shortly. No viewer complaints have arrived.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 14:02:09.200 during the public preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-004", "id": "fast-43-diverse-219-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log records incident F, a video freeze visible in the public encoded output, occurring during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 14:02:09.200 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "No program content was dropped or lost during incident F."}, {"speaker": "Event producer", "text": "The stream is now stable, and the scheduled opening begins shortly. No viewer complaints have arrived."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy via the unchanged question and use two factual encoder-log timestamp sentences as evidence without embedding labels or rules; the counterfactual only alters the end timestamp, changing duration from 2.600s to 1.400s, which is internally consistent and does not contradict other unchanged statements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Reviewing the public encoder log for incident F, I confirm it appears during the public preshow countdown and is present in the public encoded output. The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500. The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100.\"}, {\"speaker\": \"Event producer\", \"text\": \"After the freeze, the program resumed cleanly with no dropped frames or missing segments in the recorded output.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500."}, {"path": ["0", "text"], "text": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.", "negative_left": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.", "negative_right": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:23.900.", "right": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-005", "id": "fast-43-diverse-219-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Reviewing the public encoder log for incident F, I confirm it appears during the public preshow countdown and is present in the public encoded output. The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500. The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}, {"speaker": "Event producer", "text": "After the freeze, the program resumed cleanly with no dropped frames or missing segments in the recorded output."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy via the unchanged question and use two factual encoder-log timestamp sentences as evidence without embedding labels or rules; the counterfactual only alters the end timestamp, changing duration from 2.600s to 1.400s, which is internally consistent and does not contradict other unchanged statements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Reviewing the public encoder log for incident F, I confirm it appears during the public preshow countdown and is present in the public encoded output. The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500. The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100.\"}, {\"speaker\": \"Event producer\", \"text\": \"After the freeze, the program resumed cleanly with no dropped frames or missing segments in the recorded output.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500."}, {"path": ["0", "text"], "text": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.", "negative_left": "The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500.", "negative_right": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:23.900.", "right": "The encoder log lists the end timestamp of video-freeze incident F as 00:14:25.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-005", "id": "fast-43-diverse-219-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Reviewing the public encoder log for incident F, I confirm it appears during the public preshow countdown and is present in the public encoded output. The encoder log lists the start timestamp of video-freeze incident F as 00:14:22.500. The encoder log lists the end timestamp of video-freeze incident F as 00:14:23.900."}, {"speaker": "Event producer", "text": "After the freeze, the program resumed cleanly with no dropped frames or missing segments in the recorded output."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the incident-F freeze scenario with encoder logs and playback confirmation, leaving all governing policy in the original questions untouched, and the counterfactual only alters the end timestamp (21:04:13.100 to 21:04:11.900) to shift duration without introducing contradictions, extra rules, or explicit answer labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"The sponsor slate froze during the public preshow countdown, and the encoder captured the whole incident, which we are calling incident F.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the end timestamp for video-freeze incident F as 21:04:13.100.\"},{\"speaker\":\"Streaming technician\",\"text\":\"Incident F is present in the public encoded output stream, confirmed by playback review.\"},{\"speaker\":\"Event producer\",\"text\":\"Program content resumed cleanly right after the freeze cleared, with nothing dropped or skipped in the broadcast feed.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500."}, {"path": ["2", "text"], "text": "The encoder log lists the end timestamp for video-freeze incident F as 21:04:13.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500.", "negative_left": "The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500.", "negative_right": "The encoder log lists the end timestamp for video-freeze incident F as 21:04:11.900.", "right": "The encoder log lists the end timestamp for video-freeze incident F as 21:04:13.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-006", "id": "fast-43-diverse-219-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "The sponsor slate froze during the public preshow countdown, and the encoder captured the whole incident, which we are calling incident F."}, {"speaker": "Streaming technician", "text": "The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500."}, {"speaker": "Streaming technician", "text": "The encoder log lists the end timestamp for video-freeze incident F as 21:04:13.100."}, {"speaker": "Streaming technician", "text": "Incident F is present in the public encoded output stream, confirmed by playback review."}, {"speaker": "Event producer", "text": "Program content resumed cleanly right after the freeze cleared, with nothing dropped or skipped in the broadcast feed."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the incident-F freeze scenario with encoder logs and playback confirmation, leaving all governing policy in the original questions untouched, and the counterfactual only alters the end timestamp (21:04:13.100 to 21:04:11.900) to shift duration without introducing contradictions, extra rules, or explicit answer labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"The sponsor slate froze during the public preshow countdown, and the encoder captured the whole incident, which we are calling incident F.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the end timestamp for video-freeze incident F as 21:04:13.100.\"},{\"speaker\":\"Streaming technician\",\"text\":\"Incident F is present in the public encoded output stream, confirmed by playback review.\"},{\"speaker\":\"Event producer\",\"text\":\"Program content resumed cleanly right after the freeze cleared, with nothing dropped or skipped in the broadcast feed.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500."}, {"path": ["2", "text"], "text": "The encoder log lists the end timestamp for video-freeze incident F as 21:04:13.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500.", "negative_left": "The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500.", "negative_right": "The encoder log lists the end timestamp for video-freeze incident F as 21:04:11.900.", "right": "The encoder log lists the end timestamp for video-freeze incident F as 21:04:13.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-006", "id": "fast-43-diverse-219-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "The sponsor slate froze during the public preshow countdown, and the encoder captured the whole incident, which we are calling incident F."}, {"speaker": "Streaming technician", "text": "The encoder log lists the start timestamp for video-freeze incident F as 21:04:10.500."}, {"speaker": "Streaming technician", "text": "The encoder log lists the end timestamp for video-freeze incident F as 21:04:11.900."}, {"speaker": "Streaming technician", "text": "Incident F is present in the public encoded output stream, confirmed by playback review."}, {"speaker": "Event producer", "text": "Program content resumed cleanly right after the freeze cleared, with nothing dropped or skipped in the broadcast feed."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the rubric via the unchanged question, keep the same incident/path/time bindings, use two factual timestamp sentences as evidence, alter only the end timestamp coherently to yield a different duration without contradicting other facts, and contain no gold answer or rule leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Incident F, a video freeze, appears in the public encoded output and was captured during the public preshow countdown. The encoder log lists incident F's start timestamp as 20:14:07.500. The encoder log lists incident F's end timestamp as 20:14:10.100.\"}, {\"speaker\": \"Event producer\", \"text\": \"After the freeze resolved, the encoder confirms the program content resumed cleanly frame-for-frame, with no lost segments or dropped frames in the archived output.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The public encoder log is the authoritative source here, not the stage manager's rough estimate of the freeze's length.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log lists incident F's start timestamp as 20:14:07.500."}, {"path": ["0", "text"], "text": "The encoder log lists incident F's end timestamp as 20:14:10.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 20:14:07.500.", "negative_left": "The encoder log lists incident F's start timestamp as 20:14:07.500.", "negative_right": "The encoder log lists incident F's end timestamp as 20:14:09.200.", "right": "The encoder log lists incident F's end timestamp as 20:14:10.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-007", "id": "fast-43-diverse-219-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Incident F, a video freeze, appears in the public encoded output and was captured during the public preshow countdown. The encoder log lists incident F's start timestamp as 20:14:07.500. The encoder log lists incident F's end timestamp as 20:14:10.100."}, {"speaker": "Event producer", "text": "After the freeze resolved, the encoder confirms the program content resumed cleanly frame-for-frame, with no lost segments or dropped frames in the archived output."}, {"speaker": "Streaming technician", "text": "The public encoder log is the authoritative source here, not the stage manager's rough estimate of the freeze's length."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the rubric via the unchanged question, keep the same incident/path/time bindings, use two factual timestamp sentences as evidence, alter only the end timestamp coherently to yield a different duration without contradicting other facts, and contain no gold answer or rule leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Incident F, a video freeze, appears in the public encoded output and was captured during the public preshow countdown. The encoder log lists incident F's start timestamp as 20:14:07.500. The encoder log lists incident F's end timestamp as 20:14:10.100.\"}, {\"speaker\": \"Event producer\", \"text\": \"After the freeze resolved, the encoder confirms the program content resumed cleanly frame-for-frame, with no lost segments or dropped frames in the archived output.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The public encoder log is the authoritative source here, not the stage manager's rough estimate of the freeze's length.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log lists incident F's start timestamp as 20:14:07.500."}, {"path": ["0", "text"], "text": "The encoder log lists incident F's end timestamp as 20:14:10.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 20:14:07.500.", "negative_left": "The encoder log lists incident F's start timestamp as 20:14:07.500.", "negative_right": "The encoder log lists incident F's end timestamp as 20:14:09.200.", "right": "The encoder log lists incident F's end timestamp as 20:14:10.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-007", "id": "fast-43-diverse-219-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Incident F, a video freeze, appears in the public encoded output and was captured during the public preshow countdown. The encoder log lists incident F's start timestamp as 20:14:07.500. The encoder log lists incident F's end timestamp as 20:14:09.200."}, {"speaker": "Event producer", "text": "After the freeze resolved, the encoder confirms the program content resumed cleanly frame-for-frame, with no lost segments or dropped frames in the archived output."}, {"speaker": "Streaming technician", "text": "The public encoder log is the authoritative source here, not the stage manager's rough estimate of the freeze's length."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the video-freeze entity, public-output verification, and no-lost-content clause while only altering the end timestamp, keeping durations calculable without duplicating or contradicting other facts; evidence spans are plain factual log statements with no embedded rule or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log records video-freeze incident F occurring during the public preshow countdown, and it appears in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the end timestamp for video-freeze incident F as 04:12:13.100.\"},{\"speaker\":\"Streaming technician\",\"text\":\"After the freeze resolved, playback resumed at the exact frame where it had stopped, with no program content lost.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500."}, {"path": ["2", "text"], "text": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:13.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500.", "negative_left": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500.", "negative_right": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:12.100.", "right": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:13.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-009", "id": "fast-43-diverse-219-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log records video-freeze incident F occurring during the public preshow countdown, and it appears in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500."}, {"speaker": "Streaming technician", "text": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:13.100."}, {"speaker": "Streaming technician", "text": "After the freeze resolved, playback resumed at the exact frame where it had stopped, with no program content lost."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the video-freeze entity, public-output verification, and no-lost-content clause while only altering the end timestamp, keeping durations calculable without duplicating or contradicting other facts; evidence spans are plain factual log statements with no embedded rule or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log records video-freeze incident F occurring during the public preshow countdown, and it appears in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the end timestamp for video-freeze incident F as 04:12:13.100.\"},{\"speaker\":\"Streaming technician\",\"text\":\"After the freeze resolved, playback resumed at the exact frame where it had stopped, with no program content lost.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500."}, {"path": ["2", "text"], "text": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:13.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500.", "negative_left": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500.", "negative_right": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:12.100.", "right": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:13.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-009", "id": "fast-43-diverse-219-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log records video-freeze incident F occurring during the public preshow countdown, and it appears in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:10.500."}, {"speaker": "Streaming technician", "text": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:12.100."}, {"speaker": "Streaming technician", "text": "After the freeze resolved, playback resumed at the exact frame where it had stopped, with no program content lost."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing Level 1/2 rubric via the unchanged question, keep the same incident/path/time framing, use factual sentences as evidence, and the counterfactual only shifts the end timestamp to 14:02:05.200 (1.7s duration) which stays logically consistent with the unchanged 'no content lost' sentence without leaking any label or rule text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Regarding the preshow countdown freeze: the encoder log confirms incident F appeared in the public encoded output.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log records video-freeze incident F ending at timestamp 14:02:06.100 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"No program content was lost as a result of the freeze; playback resumed cleanly with all frames accounted for after the incident.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log records video-freeze incident F ending at timestamp 14:02:06.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown.", "negative_left": "The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown.", "negative_right": "The encoder log records video-freeze incident F ending at timestamp 14:02:05.200 during the public preshow countdown.", "right": "The encoder log records video-freeze incident F ending at timestamp 14:02:06.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-011", "id": "fast-43-diverse-219-011-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Regarding the preshow countdown freeze: the encoder log confirms incident F appeared in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F ending at timestamp 14:02:06.100 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "No program content was lost as a result of the freeze; playback resumed cleanly with all frames accounted for after the incident."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing Level 1/2 rubric via the unchanged question, keep the same incident/path/time framing, use factual sentences as evidence, and the counterfactual only shifts the end timestamp to 14:02:05.200 (1.7s duration) which stays logically consistent with the unchanged 'no content lost' sentence without leaking any label or rule text.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Regarding the preshow countdown freeze: the encoder log confirms incident F appeared in the public encoded output.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log records video-freeze incident F ending at timestamp 14:02:06.100 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"No program content was lost as a result of the freeze; playback resumed cleanly with all frames accounted for after the incident.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log records video-freeze incident F ending at timestamp 14:02:06.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown.", "negative_left": "The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown.", "negative_right": "The encoder log records video-freeze incident F ending at timestamp 14:02:05.200 during the public preshow countdown.", "right": "The encoder log records video-freeze incident F ending at timestamp 14:02:06.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-011", "id": "fast-43-diverse-219-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Regarding the preshow countdown freeze: the encoder log confirms incident F appeared in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F starting at timestamp 14:02:03.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log records video-freeze incident F ending at timestamp 14:02:05.200 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "No program content was lost as a result of the freeze; playback resumed cleanly with all frames accounted for after the incident."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged Level 2 rubric and routing policy from the original question, only altering the incident's end timestamp, which yields a coherent duration change (2.6s vs 1.4s) without contradicting other stated facts; the focus evidence consists of two complete factual sentences with no embedded rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"During the public preshow countdown I noticed the sponsor slate freeze briefly, then resume normally.\"},{\"speaker\":\"Streaming technician\",\"text\":\"Encoder log confirms incident F is present in the public encoded output. The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown. The encoder log lists incident F's end timestamp as 20:14:06.100 during the public preshow countdown. After the freeze, the encoder log shows the stream resumed normal playback with no gap in the recorded program content.\"},{\"speaker\":\"Event producer\",\"text\":\"No content appears to have been dropped from the broadcast archive; the recording is continuous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown."}, {"path": ["1", "text"], "text": "The encoder log lists incident F's end timestamp as 20:14:06.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 20:14:04.900 during the public preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 20:14:06.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-012", "id": "fast-43-diverse-219-012-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "During the public preshow countdown I noticed the sponsor slate freeze briefly, then resume normally."}, {"speaker": "Streaming technician", "text": "Encoder log confirms incident F is present in the public encoded output. The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown. The encoder log lists incident F's end timestamp as 20:14:06.100 during the public preshow countdown. After the freeze, the encoder log shows the stream resumed normal playback with no gap in the recorded program content."}, {"speaker": "Event producer", "text": "No content appears to have been dropped from the broadcast archive; the recording is continuous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged Level 2 rubric and routing policy from the original question, only altering the incident's end timestamp, which yields a coherent duration change (2.6s vs 1.4s) without contradicting other stated facts; the focus evidence consists of two complete factual sentences with no embedded rules or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Stage manager\",\"text\":\"During the public preshow countdown I noticed the sponsor slate freeze briefly, then resume normally.\"},{\"speaker\":\"Streaming technician\",\"text\":\"Encoder log confirms incident F is present in the public encoded output. The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown. The encoder log lists incident F's end timestamp as 20:14:06.100 during the public preshow countdown. After the freeze, the encoder log shows the stream resumed normal playback with no gap in the recorded program content.\"},{\"speaker\":\"Event producer\",\"text\":\"No content appears to have been dropped from the broadcast archive; the recording is continuous.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown."}, {"path": ["1", "text"], "text": "The encoder log lists incident F's end timestamp as 20:14:06.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 20:14:04.900 during the public preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 20:14:06.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-012", "id": "fast-43-diverse-219-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Stage manager", "text": "During the public preshow countdown I noticed the sponsor slate freeze briefly, then resume normally."}, {"speaker": "Streaming technician", "text": "Encoder log confirms incident F is present in the public encoded output. The encoder log lists incident F's start timestamp as 20:14:03.500 during the public preshow countdown. The encoder log lists incident F's end timestamp as 20:14:04.900 during the public preshow countdown. After the freeze, the encoder log shows the stream resumed normal playback with no gap in the recorded program content."}, {"speaker": "Event producer", "text": "No content appears to have been dropped from the broadcast archive; the recording is continuous."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same incident F identity, preshow-countdown timing and no-lost-frames fact, relying on the unchanged question's rubric for policy; the counterfactual only changes the end timestamp to 21:04:13.700, yielding a shorter but still plausible duration without contradicting other statements; the two focus sentences are plain factual log statements, not policy text; and neither context states or implies a Level classification or decision outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log records one video freeze during the preshow countdown, confirmed present in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's end timestamp as 21:04:14.900 during the preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"After the freeze, playback resumed smoothly and no program frames were dropped or skipped.\"},{\"speaker\":\"Event producer\",\"text\":\"The scheduled opening begins in four minutes and the stream is otherwise stable.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 21:04:14.900 during the preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 21:04:13.700 during the preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 21:04:14.900 during the preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-013", "id": "fast-43-diverse-219-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log records one video freeze during the preshow countdown, confirmed present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 21:04:14.900 during the preshow countdown."}, {"speaker": "Streaming technician", "text": "After the freeze, playback resumed smoothly and no program frames were dropped or skipped."}, {"speaker": "Event producer", "text": "The scheduled opening begins in four minutes and the stream is otherwise stable."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same incident F identity, preshow-countdown timing and no-lost-frames fact, relying on the unchanged question's rubric for policy; the counterfactual only changes the end timestamp to 21:04:13.700, yielding a shorter but still plausible duration without contradicting other statements; the two focus sentences are plain factual log statements, not policy text; and neither context states or implies a Level classification or decision outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log records one video freeze during the preshow countdown, confirmed present in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's end timestamp as 21:04:14.900 during the preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"After the freeze, playback resumed smoothly and no program frames were dropped or skipped.\"},{\"speaker\":\"Event producer\",\"text\":\"The scheduled opening begins in four minutes and the stream is otherwise stable.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 21:04:14.900 during the preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 21:04:13.700 during the preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 21:04:14.900 during the preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-013", "id": "fast-43-diverse-219-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log records one video freeze during the preshow countdown, confirmed present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 21:04:12.150 during the preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 21:04:13.700 during the preshow countdown."}, {"speaker": "Streaming technician", "text": "After the freeze, playback resumed smoothly and no program frames were dropped or skipped."}, {"speaker": "Event producer", "text": "The scheduled opening begins in four minutes and the stream is otherwise stable."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged rubric via the verbatim questions object and only alter the incident F end timestamp, changing duration from 2.600s to 1.400s without contradicting the 'no frames dropped' claim, using two factual, non-policy evidence sentences and no leaked labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Reviewing incident F: the encoder log confirms this video freeze is present in the public encoded output during the preshow countdown. The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown. The encoder log shows video-freeze incident F ending at timestamp 04:12:10.100 during the public preshow countdown.\"}, {\"speaker\": \"Event producer\", \"text\": \"After the freeze cleared, the program resumed normally with no frames dropped and no content lost from the broadcast.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown."}, {"path": ["0", "text"], "text": "The encoder log shows video-freeze incident F ending at timestamp 04:12:10.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown.", "negative_left": "The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown.", "negative_right": "The encoder log shows video-freeze incident F ending at timestamp 04:12:08.900 during the public preshow countdown.", "right": "The encoder log shows video-freeze incident F ending at timestamp 04:12:10.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-014", "id": "fast-43-diverse-219-014-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Reviewing incident F: the encoder log confirms this video freeze is present in the public encoded output during the preshow countdown. The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown. The encoder log shows video-freeze incident F ending at timestamp 04:12:10.100 during the public preshow countdown."}, {"speaker": "Event producer", "text": "After the freeze cleared, the program resumed normally with no frames dropped and no content lost from the broadcast."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the unchanged rubric via the verbatim questions object and only alter the incident F end timestamp, changing duration from 2.600s to 1.400s without contradicting the 'no frames dropped' claim, using two factual, non-policy evidence sentences and no leaked labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Reviewing incident F: the encoder log confirms this video freeze is present in the public encoded output during the preshow countdown. The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown. The encoder log shows video-freeze incident F ending at timestamp 04:12:10.100 during the public preshow countdown.\"}, {\"speaker\": \"Event producer\", \"text\": \"After the freeze cleared, the program resumed normally with no frames dropped and no content lost from the broadcast.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["0", "text"], "text": "The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown."}, {"path": ["0", "text"], "text": "The encoder log shows video-freeze incident F ending at timestamp 04:12:10.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown.", "negative_left": "The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown.", "negative_right": "The encoder log shows video-freeze incident F ending at timestamp 04:12:08.900 during the public preshow countdown.", "right": "The encoder log shows video-freeze incident F ending at timestamp 04:12:10.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-014", "id": "fast-43-diverse-219-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Reviewing incident F: the encoder log confirms this video freeze is present in the public encoded output during the preshow countdown. The encoder log shows video-freeze incident F starting at timestamp 04:12:07.500 during the public preshow countdown. The encoder log shows video-freeze incident F ending at timestamp 04:12:08.900 during the public preshow countdown."}, {"speaker": "Event producer", "text": "After the freeze cleared, the program resumed normally with no frames dropped and no content lost from the broadcast."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rubric via the unchanged question, preserve incident F's identity/location/start time, use two factual timestamp sentences as evidence with a consistent 1.4s vs 2.6s duration change, and contain no explicit level/answer statements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log confirms video-freeze incident F occurred during the public preshow countdown and appears in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's end timestamp as 14:02:10.100 during the preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"No program content was lost after the freeze resolved; playback resumed cleanly from the last valid frame.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 14:02:08.900 during the preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-015", "id": "fast-43-diverse-219-015-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log confirms video-freeze incident F occurred during the public preshow countdown and appears in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the preshow countdown."}, {"speaker": "Streaming technician", "text": "No program content was lost after the freeze resolved; playback resumed cleanly from the last valid frame."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rubric via the unchanged question, preserve incident F's identity/location/start time, use two factual timestamp sentences as evidence with a consistent 1.4s vs 2.6s duration change, and contain no explicit level/answer statements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log confirms video-freeze incident F occurred during the public preshow countdown and appears in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists incident F's end timestamp as 14:02:10.100 during the preshow countdown.\"},{\"speaker\":\"Streaming technician\",\"text\":\"No program content was lost after the freeze resolved; playback resumed cleanly from the last valid frame.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 14:02:08.900 during the preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 14:02:10.100 during the preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-015", "id": "fast-43-diverse-219-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log confirms video-freeze incident F occurred during the public preshow countdown and appears in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 14:02:07.500 during the preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 14:02:08.900 during the preshow countdown."}, {"speaker": "Streaming technician", "text": "No program content was lost after the freeze resolved; playback resumed cleanly from the last valid frame."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same incident F entity, public preshow countdown timing and encoder-log sourcing required by the unchanged questions object; the counterfactual only changes the end timestamp (22.500→23.900), yielding a coherent 1.4s duration consistent with the unchanged 'no content lost' sentence, while the base context's 2.6s duration is also internally consistent; the two focus evidence spans are complete factual sentences about logged timestamps, not policy or instructions; neither context states a Level classification, routing decision, or answer code, so no gold-answer leakage occurs.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"The public encoder log confirms incident F occurred during the public preshow countdown and appears in the public encoded output.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log lists incident F's end timestamp as 00:14:25.100 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"Playback resumed cleanly with all frames accounted for after the freeze, and no program content was lost as a result of incident F.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 00:14:25.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 00:14:23.900 during the public preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 00:14:25.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-016", "id": "fast-43-diverse-219-016-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log confirms incident F occurred during the public preshow countdown and appears in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 00:14:25.100 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "Playback resumed cleanly with all frames accounted for after the freeze, and no program content was lost as a result of incident F."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same incident F entity, public preshow countdown timing and encoder-log sourcing required by the unchanged questions object; the counterfactual only changes the end timestamp (22.500→23.900), yielding a coherent 1.4s duration consistent with the unchanged 'no content lost' sentence, while the base context's 2.6s duration is also internally consistent; the two focus evidence spans are complete factual sentences about logged timestamps, not policy or instructions; neither context states a Level classification, routing decision, or answer code, so no gold-answer leakage occurs.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"The public encoder log confirms incident F occurred during the public preshow countdown and appears in the public encoded output.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log lists incident F's end timestamp as 00:14:25.100 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"Playback resumed cleanly with all frames accounted for after the freeze, and no program content was lost as a result of incident F.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log lists incident F's end timestamp as 00:14:25.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown.", "negative_left": "The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown.", "negative_right": "The encoder log lists incident F's end timestamp as 00:14:23.900 during the public preshow countdown.", "right": "The encoder log lists incident F's end timestamp as 00:14:25.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-016", "id": "fast-43-diverse-219-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log confirms incident F occurred during the public preshow countdown and appears in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's start timestamp as 00:14:22.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log lists incident F's end timestamp as 00:14:23.900 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "Playback resumed cleanly with all frames accounted for after the freeze, and no program content was lost as a result of incident F."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full rubric via the unchanged question, keep the same incident entity F and preshow-countdown timing, use two complete factual timestamp sentences as evidence, and the counterfactual's shorter duration remains internally consistent with the rest of the context without leaking any level/answer labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log confirms one video-freeze incident, designated F, occurring during the public preshow countdown, and it is present in the public encoded output stream.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the end timestamp of video-freeze incident F as 14:02:13.100 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"After the freeze resolved, playback resumed seamlessly and no program content was lost; the countdown graphic simply held on screen before continuing normally.\"},{\"speaker\":\"Event producer\",\"text\":\"No viewer complaints have been received, and the stream is currently stable heading into the scheduled start.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC."}, {"path": ["2", "text"], "text": "The encoder log records the end timestamp of video-freeze incident F as 14:02:13.100 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC.", "negative_left": "The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC.", "negative_right": "The encoder log records the end timestamp of video-freeze incident F as 14:02:11.900 UTC.", "right": "The encoder log records the end timestamp of video-freeze incident F as 14:02:13.100 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-017", "id": "fast-43-diverse-219-017-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log confirms one video-freeze incident, designated F, occurring during the public preshow countdown, and it is present in the public encoded output stream."}, {"speaker": "Streaming technician", "text": "The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC."}, {"speaker": "Streaming technician", "text": "The encoder log records the end timestamp of video-freeze incident F as 14:02:13.100 UTC."}, {"speaker": "Streaming technician", "text": "After the freeze resolved, playback resumed seamlessly and no program content was lost; the countdown graphic simply held on screen before continuing normally."}, {"speaker": "Event producer", "text": "No viewer complaints have been received, and the stream is currently stable heading into the scheduled start."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full rubric via the unchanged question, keep the same incident entity F and preshow-countdown timing, use two complete factual timestamp sentences as evidence, and the counterfactual's shorter duration remains internally consistent with the rest of the context without leaking any level/answer labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log confirms one video-freeze incident, designated F, occurring during the public preshow countdown, and it is present in the public encoded output stream.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log records the end timestamp of video-freeze incident F as 14:02:13.100 UTC.\"},{\"speaker\":\"Streaming technician\",\"text\":\"After the freeze resolved, playback resumed seamlessly and no program content was lost; the countdown graphic simply held on screen before continuing normally.\"},{\"speaker\":\"Event producer\",\"text\":\"No viewer complaints have been received, and the stream is currently stable heading into the scheduled start.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC."}, {"path": ["2", "text"], "text": "The encoder log records the end timestamp of video-freeze incident F as 14:02:13.100 UTC."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC.", "negative_left": "The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC.", "negative_right": "The encoder log records the end timestamp of video-freeze incident F as 14:02:11.900 UTC.", "right": "The encoder log records the end timestamp of video-freeze incident F as 14:02:13.100 UTC."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-017", "id": "fast-43-diverse-219-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log confirms one video-freeze incident, designated F, occurring during the public preshow countdown, and it is present in the public encoded output stream."}, {"speaker": "Streaming technician", "text": "The encoder log records the start timestamp of video-freeze incident F as 14:02:10.500 UTC."}, {"speaker": "Streaming technician", "text": "The encoder log records the end timestamp of video-freeze incident F as 14:02:11.900 UTC."}, {"speaker": "Streaming technician", "text": "After the freeze resolved, playback resumed seamlessly and no program content was lost; the countdown graphic simply held on screen before continuing normally."}, {"speaker": "Event producer", "text": "No viewer complaints have been received, and the stream is currently stable heading into the scheduled start."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original question's Level 1/2 rubric without redefining it, keep the same incident F/video-freeze/public-encoder-log bindings, use two complete factual sentences as evidence, and the counterfactual only changes the end timestamp (yielding a shorter, still-coherent duration) without contradicting other unchanged facts or leaking any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Reviewing the public encoder log for video-freeze incident F during tonight's preshow countdown: the freeze is confirmed present in the public encoded output.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log records the end timestamp for video-freeze incident F as 04:12:21.100 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"No frames of program content were dropped or lost afterward; the feed recovered cleanly and no program content was lost.\"}, {\"speaker\": \"Event producer\", \"text\": \"Noted for the incident report: this is based purely on encoder-recorded timestamps rather than any backstage or approximate observation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log records the end timestamp for video-freeze incident F as 04:12:21.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown.", "negative_left": "The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown.", "negative_right": "The encoder log records the end timestamp for video-freeze incident F as 04:12:19.800 during the public preshow countdown.", "right": "The encoder log records the end timestamp for video-freeze incident F as 04:12:21.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-018", "id": "fast-43-diverse-219-018-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Reviewing the public encoder log for video-freeze incident F during tonight's preshow countdown: the freeze is confirmed present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log records the end timestamp for video-freeze incident F as 04:12:21.100 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "No frames of program content were dropped or lost afterward; the feed recovered cleanly and no program content was lost."}, {"speaker": "Event producer", "text": "Noted for the incident report: this is based purely on encoder-recorded timestamps rather than any backstage or approximate observation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original question's Level 1/2 rubric without redefining it, keep the same incident F/video-freeze/public-encoder-log bindings, use two complete factual sentences as evidence, and the counterfactual only changes the end timestamp (yielding a shorter, still-coherent duration) without contradicting other unchanged facts or leaking any answer/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\": \"Streaming technician\", \"text\": \"Reviewing the public encoder log for video-freeze incident F during tonight's preshow countdown: the freeze is confirmed present in the public encoded output.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"The encoder log records the end timestamp for video-freeze incident F as 04:12:21.100 during the public preshow countdown.\"}, {\"speaker\": \"Streaming technician\", \"text\": \"No frames of program content were dropped or lost afterward; the feed recovered cleanly and no program content was lost.\"}, {\"speaker\": \"Event producer\", \"text\": \"Noted for the incident report: this is based purely on encoder-recorded timestamps rather than any backstage or approximate observation.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown."}, {"path": ["2", "text"], "text": "The encoder log records the end timestamp for video-freeze incident F as 04:12:21.100 during the public preshow countdown."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown.", "negative_left": "The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown.", "negative_right": "The encoder log records the end timestamp for video-freeze incident F as 04:12:19.800 during the public preshow countdown.", "right": "The encoder log records the end timestamp for video-freeze incident F as 04:12:21.100 during the public preshow countdown."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-018", "id": "fast-43-diverse-219-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "Reviewing the public encoder log for video-freeze incident F during tonight's preshow countdown: the freeze is confirmed present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log records the start timestamp for video-freeze incident F as 04:12:18.500 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "The encoder log records the end timestamp for video-freeze incident F as 04:12:19.800 during the public preshow countdown."}, {"speaker": "Streaming technician", "text": "No frames of program content were dropped or lost afterward; the feed recovered cleanly and no program content was lost."}, {"speaker": "Event producer", "text": "Noted for the incident report: this is based purely on encoder-recorded timestamps rather than any backstage or approximate observation."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original decision policy verbatim via the unchanged questions object, keep the same incident/path/time bindings for incident F, use two complete factual timestamp sentences as evidence, and the counterfactual (1.5s duration, no content lost) is internally consistent with no contradictory measurements or leaked answer labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log confirms video-freeze incident F occurred during the public preshow countdown and is present in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the end timestamp for video-freeze incident F as 04:12:10.100.\"},{\"speaker\":\"Event producer\",\"text\":\"Program playback resumed cleanly right after the freeze cleared, with no program content lost.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500."}, {"path": ["2", "text"], "text": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:10.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500.", "negative_left": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500.", "negative_right": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:09.000.", "right": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:10.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-019", "id": "fast-43-diverse-219-019-base", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log confirms video-freeze incident F occurred during the public preshow countdown and is present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500."}, {"speaker": "Streaming technician", "text": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:10.100."}, {"speaker": "Event producer", "text": "Program playback resumed cleanly right after the freeze cleared, with no program content lost."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-02", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original decision policy verbatim via the unchanged questions object, keep the same incident/path/time bindings for incident F, use two complete factual timestamp sentences as evidence, and the counterfactual (1.5s duration, no content lost) is internally consistent with no contradictory measurements or leaked answer labels.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "refuted", "a4": "refuted"}, "remove_left": {"a3": "unknown"}, "remove_right": {"a3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a3": "unknown"}, "negative_pair": {"a3": "refuted"}, "negative_sentence": {"a3": "unknown"}, "positive_pair": {"a3": "supported"}, "right": {"a3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship, and the focus atom is the factual duration-threshold relation rather than a policy classification. The base and counter assignments can be realized with only that threshold fact changing while public presence, encoder verification, and absence of content loss remain fixed. Empty policy evidence is correct because all governing classification rules, thresholds, source priority, and scope are already retained in the questions object; the original state adds no separate substantive policy needed to interpret it.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A verified defect present in public encoded output and lasting more than 2.000 seconds is sufficient for Level 2 under the stated duration criterion; lost-content status need not also be specified.", "rule_index": 0, "sound": true}, {"reason": "A verified defect present in public encoded output that causes lost program content is sufficient for Level 2 regardless of duration.", "rule_index": 1, "sound": true}, {"reason": "Encoder verification and public-output presence establish a verified public defect, while refutation of “more than 2.000 seconds” establishes a duration of 2.000 seconds or less and refutation of content loss supplies the required competing-outcome exclusion. These conditions are sufficient for Level 1 and the false decision.", "rule_index": 2, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The encoder log records video-freeze incident F during the public preshow countdown."}, {"id": "a2", "statement": "Video-freeze incident F is present in the public encoded output."}, {"id": "a3", "statement": "For video-freeze incident F, the encoder-recorded end timestamp minus the encoder-recorded start timestamp is more than 2.000 seconds."}, {"id": "a4", "statement": "Video-freeze incident F caused program content to be lost."}], "base_state_json": "[{\"speaker\":\"Streaming technician\",\"text\":\"The public encoder log confirms video-freeze incident F occurred during the public preshow countdown and is present in the public encoded output.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500.\"},{\"speaker\":\"Streaming technician\",\"text\":\"The encoder log lists the end timestamp for video-freeze incident F as 04:12:10.100.\"},{\"speaker\":\"Event producer\",\"text\":\"Program playback resumed cleanly right after the freeze cleared, with no program content lost.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}], "focus_atom": "a3", "focus_evidence": [{"path": ["1", "text"], "text": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500."}, {"path": ["2", "text"], "text": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:10.100."}], "policy_evidence": [], "rules": [{"justification": "An encoder-verified defect in the public encoded output is Level 2 when its encoder-derived duration is more than 2.000 seconds.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}]}, {"justification": "An encoder-verified defect in the public encoded output is Level 2 when it causes lost program content, regardless of duration.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "An encoder-verified public defect that lasts 2.000 seconds or less and causes no loss of program content remains Level 1.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "refuted"}, {"atom_id": "a4", "state": "refuted"}]}]}, "verified_pair": {"left": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500.", "negative_left": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500.", "negative_right": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:09.000.", "right": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:10.100."}, "verifier_independent_model": false}, "family": "fast-43-diverse-219-019", "id": "fast-43-diverse-219-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No—the incident remains Level 1, so no immediate intervention is required.", "true": "Yes—the incident is Level 2, so immediate intervention is required."}, "instructions": "Decide whether the verified public-output incident meets Level 2 and therefore requires immediate intervention. Use encoder logs over approximate observations. Ignore defects absent from public encoded output. Impact rubric: Level 0 = unverified or non-public defect; Level 1 = verified public defect lasting 2.000 seconds or less, with no program content lost—monitor and route to the relevant production function after the show; Level 2 = defect lasting more than 2.000 seconds or causing lost program content—intervene immediately. Video freezes route to the Streaming technician; public audio defects route to the Audio engineer.", "type": "noul"}}, "state": [{"speaker": "Streaming technician", "text": "The public encoder log confirms video-freeze incident F occurred during the public preshow countdown and is present in the public encoded output."}, {"speaker": "Streaming technician", "text": "The encoder log lists the start timestamp for video-freeze incident F as 04:12:07.500."}, {"speaker": "Streaming technician", "text": "The encoder log lists the end timestamp for video-freeze incident F as 04:12:09.000."}, {"speaker": "Event producer", "text": "Program playback resumed cleanly right after the freeze cleared, with no program content lost."}]}, "method": "c2d", "provenance": {"source_id": "diverse-219", "source_is_synthetic": true, "source_sha256": "ca88ca68f4869867ab861541d2053aa5ca98787d649bd8d57b9e51315f54df25", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-02", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The rubric and questions object are unchanged, preserving policy and bindings; the counterfactual only swaps the checksum tying harbor_mix.prproj to HP-3, breaking the identification chain without contradicting any other assertion, so it remains coherent; the focus evidence are two complete factual sentences; neither context reveals a gold label or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager\\u2019s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 7f2a-19c.\", \"Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 7f2a-19c."}, {"path": ["evidence", "1"], "text": "Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 7f2a-19c.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 9d4e-22b.", "negative_right": "Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-019", "id": "fast-43-diverse-223-019-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 7f2a-19c.", "Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The rubric and questions object are unchanged, preserving policy and bindings; the counterfactual only swaps the checksum tying harbor_mix.prproj to HP-3, breaking the identification chain without contradicting any other assertion, so it remains coherent; the focus evidence are two complete factual sentences; neither context reveals a gold label or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager\\u2019s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 7f2a-19c.\", \"Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 7f2a-19c."}, {"path": ["evidence", "1"], "text": "Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 7f2a-19c.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 9d4e-22b.", "negative_right": "Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-019", "id": "fast-43-diverse-223-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 9d4e-22b.", "Checksum entry 7f2a-19c in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric verbatim and the same request/entity; the counterfactual only swaps the checksum's transfer ID from HP-3 to HP-9, breaking the identity link for harbor_mix.prproj without contradicting any other stated fact, and no evidence sentence states or hints at the final routing decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"Manifest: gulls.wav — 96 kHz field recording; harbor_echo_final.mov — published documentary; harbor_mix.prproj — Adobe project file; pier_still.tif — production photograph.\",\"The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1.\",\"Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-3.\",\"The asset assigned Harbor Echo transfer identifier HP-3 is a project file that is editable and links source media, and it is neither raw captured sound nor a published final export.\",\"Every other asset in the Harbor Echo transfer is a published final export, not raw captured sound, and not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "1"], "text": "The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1."}, {"path": ["evidence", "2"], "text": "Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1.", "negative_left": "The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1.", "negative_right": "Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-9, distinct from HP-3.", "right": "Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-020", "id": "fast-43-diverse-223-020-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["Manifest: gulls.wav — 96 kHz field recording; harbor_echo_final.mov — published documentary; harbor_mix.prproj — Adobe project file; pier_still.tif — production photograph.", "The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1.", "Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is a project file that is editable and links source media, and it is neither raw captured sound nor a published final export.", "Every other asset in the Harbor Echo transfer is a published final export, not raw captured sound, and not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric verbatim and the same request/entity; the counterfactual only swaps the checksum's transfer ID from HP-3 to HP-9, breaking the identity link for harbor_mix.prproj without contradicting any other stated fact, and no evidence sentence states or hints at the final routing decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"Manifest: gulls.wav — 96 kHz field recording; harbor_echo_final.mov — published documentary; harbor_mix.prproj — Adobe project file; pier_still.tif — production photograph.\",\"The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1.\",\"Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-3.\",\"The asset assigned Harbor Echo transfer identifier HP-3 is a project file that is editable and links source media, and it is neither raw captured sound nor a published final export.\",\"Every other asset in the Harbor Echo transfer is a published final export, not raw captured sound, and not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "1"], "text": "The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1."}, {"path": ["evidence", "2"], "text": "Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1.", "negative_left": "The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1.", "negative_right": "Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-9, distinct from HP-3.", "right": "Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-020", "id": "fast-43-diverse-223-020-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["Manifest: gulls.wav — 96 kHz field recording; harbor_echo_final.mov — published documentary; harbor_mix.prproj — Adobe project file; pier_still.tif — production photograph.", "The Harbor Echo transfer manifest logs the checksum for harbor_mix.prproj as 9f3c7a1.", "Manifest entry 9f3c7a1 is recorded under Harbor Echo transfer identifier HP-9, distinct from HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is a project file that is editable and links source media, and it is neither raw captured sound nor a published final export.", "Every other asset in the Harbor Echo transfer is a published final export, not raw captured sound, and not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rubric and identical request/entity; the counterfactual changes only the CX-77 cross-reference from HP-3 to HP-9, coherently reassigning harbor_mix.prproj into the 'other than HP-3' published-export clause without contradicting any other stated fact; evidence spans are plain factual sentences with no rules, IDs, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"Manifest: gulls.wav — 96 kHz field recording; harbor_echo_final.mov — published documentary; harbor_mix.prproj — Adobe project file; pier_still.tif — production photograph.\", \"The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77.\", \"Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-3.\", \"Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export, is not raw captured sound, and is not editable.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is a project file that is editable, links source media, is not raw captured sound, and is not a published final export.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "1"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77."}, {"path": ["evidence", "2"], "text": "Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77.", "negative_right": "Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-9.", "right": "Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-028", "id": "fast-43-diverse-223-028-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["Manifest: gulls.wav — 96 kHz field recording; harbor_echo_final.mov — published documentary; harbor_mix.prproj — Adobe project file; pier_still.tif — production photograph.", "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77.", "Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-3.", "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export, is not raw captured sound, and is not editable.", "The asset assigned Harbor Echo transfer identifier HP-3 is a project file that is editable, links source media, is not raw captured sound, and is not a published final export."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rubric and identical request/entity; the counterfactual changes only the CX-77 cross-reference from HP-3 to HP-9, coherently reassigning harbor_mix.prproj into the 'other than HP-3' published-export clause without contradicting any other stated fact; evidence spans are plain factual sentences with no rules, IDs, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"Manifest: gulls.wav — 96 kHz field recording; harbor_echo_final.mov — published documentary; harbor_mix.prproj — Adobe project file; pier_still.tif — production photograph.\", \"The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77.\", \"Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-3.\", \"Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export, is not raw captured sound, and is not editable.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is a project file that is editable, links source media, is not raw captured sound, and is not a published final export.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "1"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77."}, {"path": ["evidence", "2"], "text": "Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77.", "negative_right": "Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-9.", "right": "Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-028", "id": "fast-43-diverse-223-028-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["Manifest: gulls.wav — 96 kHz field recording; harbor_echo_final.mov — published documentary; harbor_mix.prproj — Adobe project file; pier_still.tif — production photograph.", "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry CX-77.", "Checksum entry CX-77 in the Harbor Echo manifest is cross-referenced to transfer identifier HP-9.", "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export, is not raw captured sound, and is not editable.", "The asset assigned Harbor Echo transfer identifier HP-3 is a project file that is editable, links source media, is not raw captured sound, and is not a published final export."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical rubric and unchanged question, preserving policy and the harbor_mix.prproj entity binding; the two focus evidence spans are complete factual sentences, not rule text; the counterfactual only swaps the transfer id from HP-3 to HP-9 without asserting any contradictory measurement or count, leaving evidentiary linkage merely weakened rather than incoherent; neither context reveals a gold label, rule table, or instruction dictating the output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A.\", \"The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-3.\", \"Metadata specialist notes that the asset tagged HP-3 is an Adobe project file containing a documentary timeline that links to the WAV, MOV, and TIFF source files in the transfer, and remains fully editable rather than a finished export.\", \"All other assets in the Harbor Echo transfer, aside from HP-3, are finished, non-editable exports and are not raw audio captures.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A."}, {"path": ["evidence", "1"], "text": "The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A.", "negative_left": "The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A.", "negative_right": "The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-9.", "right": "The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-033", "id": "fast-43-diverse-223-033-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A.", "The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-3.", "Metadata specialist notes that the asset tagged HP-3 is an Adobe project file containing a documentary timeline that links to the WAV, MOV, and TIFF source files in the transfer, and remains fully editable rather than a finished export.", "All other assets in the Harbor Echo transfer, aside from HP-3, are finished, non-editable exports and are not raw audio captures."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical rubric and unchanged question, preserving policy and the harbor_mix.prproj entity binding; the two focus evidence spans are complete factual sentences, not rule text; the counterfactual only swaps the transfer id from HP-3 to HP-9 without asserting any contradictory measurement or count, leaving evidentiary linkage merely weakened rather than incoherent; neither context reveals a gold label, rule table, or instruction dictating the output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A.\", \"The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-3.\", \"Metadata specialist notes that the asset tagged HP-3 is an Adobe project file containing a documentary timeline that links to the WAV, MOV, and TIFF source files in the transfer, and remains fully editable rather than a finished export.\", \"All other assets in the Harbor Echo transfer, aside from HP-3, are finished, non-editable exports and are not raw audio captures.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A."}, {"path": ["evidence", "1"], "text": "The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A.", "negative_left": "The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A.", "negative_right": "The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-9.", "right": "The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-033", "id": "fast-43-diverse-223-033-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest lists harbor_mix.prproj with checksum 7F2C-91A.", "The manifest entry with checksum 7F2C-91A is tagged with Harbor Echo transfer identifier HP-9.", "Metadata specialist notes that the asset tagged HP-3 is an Adobe project file containing a documentary timeline that links to the WAV, MOV, and TIFF source files in the transfer, and remains fully editable rather than a finished export.", "All other assets in the Harbor Echo transfer, aside from HP-3, are finished, non-editable exports and are not raw audio captures."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The rubric and question criteria are unchanged and carried via the verbatim questions object; the counterfactual only reassigns checksum 7F2C to HP-9 instead of HP-3, which breaks the identification chain linking harbor_mix.prproj to the editable-project description but does not create a direct factual contradiction, since the HP-3 and 'other assets' statements remain internally consistent and now simply fail to resolve to the requested file, plausibly steering toward Metadata Review rather than producing an incoherent state; both focus evidence spans are complete factual manifest statements with no policy language, rule tables, or embedded answer/label leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C.\",\"Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-3.\",\"Manifest entry HP-3: Adobe project file, editable, contains a documentary timeline that links source media (WAV, MOV, TIFF tracks); it is not raw captured sound and not a published final export.\",\"Manifest note: every other asset in the Harbor Echo transfer besides HP-3 is a finished, published export, none of them editable, none of them raw captured sound.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C."}, {"path": ["evidence", "1"], "text": "Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C.", "negative_left": "The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C.", "negative_right": "Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-9.", "right": "Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-048", "id": "fast-43-diverse-223-048-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C.", "Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-3.", "Manifest entry HP-3: Adobe project file, editable, contains a documentary timeline that links source media (WAV, MOV, TIFF tracks); it is not raw captured sound and not a published final export.", "Manifest note: every other asset in the Harbor Echo transfer besides HP-3 is a finished, published export, none of them editable, none of them raw captured sound."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "The rubric and question criteria are unchanged and carried via the verbatim questions object; the counterfactual only reassigns checksum 7F2C to HP-9 instead of HP-3, which breaks the identification chain linking harbor_mix.prproj to the editable-project description but does not create a direct factual contradiction, since the HP-3 and 'other assets' statements remain internally consistent and now simply fail to resolve to the requested file, plausibly steering toward Metadata Review rather than producing an incoherent state; both focus evidence spans are complete factual manifest statements with no policy language, rule tables, or embedded answer/label leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C.\",\"Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-3.\",\"Manifest entry HP-3: Adobe project file, editable, contains a documentary timeline that links source media (WAV, MOV, TIFF tracks); it is not raw captured sound and not a published final export.\",\"Manifest note: every other asset in the Harbor Echo transfer besides HP-3 is a finished, published export, none of them editable, none of them raw captured sound.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C."}, {"path": ["evidence", "1"], "text": "Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C.", "negative_left": "The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C.", "negative_right": "Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-9.", "right": "Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-048", "id": "fast-43-diverse-223-048-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest lists a file named harbor_mix.prproj with checksum tag 7F2C.", "Checksum tag 7F2C is recorded in the manifest as belonging to the asset assigned Harbor Echo transfer identifier HP-9.", "Manifest entry HP-3: Adobe project file, editable, contains a documentary timeline that links source media (WAV, MOV, TIFF tracks); it is not raw captured sound and not a published final export.", "Manifest note: every other asset in the Harbor Echo transfer besides HP-3 is a finished, published export, none of them editable, none of them raw captured sound."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric and criteria intact, and the unchanged questions object supplies the policy scoring rules, so policy is preserved throughout; the request and asset entity (harbor_mix.prproj) are untouched, preserving question bindings; the focus evidence consists of two complete factual sentences about manifest entries, not policy text; the counterfactual only swaps the entry-7 identifier from HP-3 to HP-9, which severs the link between harbor_mix.prproj and the HP-3 description without contradicting any other stated fact, making it a coherent ambiguity-inducing variant; no evidence sentence states or hints at the final routing decision, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager\\u2019s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7.\",\"Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-3.\",\"HP-3 is an Adobe Premiere project file that arranges and links the WAV, MOV, and TIFF source media from the transfer, and remains fully editable by staff.\",\"HP-3 is confirmed to be neither raw captured sound nor a published final export; it is a working project container.\",\"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7."}, {"path": ["evidence", "1"], "text": "Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7.", "negative_left": "The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7.", "negative_right": "Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-9.", "right": "Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-057", "id": "fast-43-diverse-223-057-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7.", "Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-3.", "HP-3 is an Adobe Premiere project file that arranges and links the WAV, MOV, and TIFF source media from the transfer, and remains fully editable by staff.", "HP-3 is confirmed to be neither raw captured sound nor a published final export; it is a working project container.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric and criteria intact, and the unchanged questions object supplies the policy scoring rules, so policy is preserved throughout; the request and asset entity (harbor_mix.prproj) are untouched, preserving question bindings; the focus evidence consists of two complete factual sentences about manifest entries, not policy text; the counterfactual only swaps the entry-7 identifier from HP-3 to HP-9, which severs the link between harbor_mix.prproj and the HP-3 description without contradicting any other stated fact, making it a coherent ambiguity-inducing variant; no evidence sentence states or hints at the final routing decision, rule table, or label rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager\\u2019s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7.\",\"Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-3.\",\"HP-3 is an Adobe Premiere project file that arranges and links the WAV, MOV, and TIFF source media from the transfer, and remains fully editable by staff.\",\"HP-3 is confirmed to be neither raw captured sound nor a published final export; it is a working project container.\",\"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7."}, {"path": ["evidence", "1"], "text": "Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7.", "negative_left": "The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7.", "negative_right": "Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-9.", "right": "Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-057", "id": "fast-43-diverse-223-057-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest lists the file harbor_mix.prproj under transfer entry number 7.", "Transfer entry number 7 in the Harbor Echo manifest is the entry assigned transfer identifier HP-9.", "HP-3 is an Adobe Premiere project file that arranges and links the WAV, MOV, and TIFF source media from the transfer, and remains fully editable by staff.", "HP-3 is confirmed to be neither raw captured sound nor a published final export; it is a working project container.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rubric and the unchanged question object, preserving policy and entity bindings (harbor_mix.prproj, HP-3, Harbor Echo). The counterfactual only swaps the ingest ticket number for harbor_mix.prproj, which is a coherent, non-contradictory case variation (it now fails to link harbor_mix.prproj to HP-3 via IT-4471, a legitimate state rather than a duplicate/contradictory measurement). The two focus evidence sentences are complete factual statements about manifest logging and ticket assignment, not policy or instructions. Neither context reveals a gold answer, rule table, or instructional leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-4471.\", \"Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-4471."}, {"path": ["evidence", "1"], "text": "Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-4471.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-5589.", "negative_right": "Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-066", "id": "fast-43-diverse-223-066-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-4471.", "Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the rubric and the unchanged question object, preserving policy and entity bindings (harbor_mix.prproj, HP-3, Harbor Echo). The counterfactual only swaps the ingest ticket number for harbor_mix.prproj, which is a coherent, non-contradictory case variation (it now fails to link harbor_mix.prproj to HP-3 via IT-4471, a legitimate state rather than a duplicate/contradictory measurement). The two focus evidence sentences are complete factual statements about manifest logging and ticket assignment, not policy or instructions. Neither context reveals a gold answer, rule table, or instructional leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-4471.\", \"Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-4471."}, {"path": ["evidence", "1"], "text": "Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-4471.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-5589.", "negative_right": "Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-066", "id": "fast-43-diverse-223-066-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number IT-5589.", "Ingest ticket number IT-4471 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric and question criteria, so policy is preserved; the entity (harbor_mix.prproj), path, and time bindings from the question remain intact. The focus evidence consists of two complete factual sentences about ticket numbers, not policy text or instructions. The counterfactual changes only the ticket number tying harbor_mix.prproj to HE-6620 instead of HE-5541, which breaks the identity link to HP-3 but does not create a direct contradiction with any other evidence, remaining internally coherent (it merely renders the asset's classification less determinate). Neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-5541.\", \"Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is a project file that remains editable, and it links source media consisting of the gulls.wav recording, the harbor_echo_final.mov documentary, and the pier_still.tif photograph.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is not raw captured sound and is not a published final export.\", \"Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export, is not raw captured sound, and is not editable; this covers gulls.wav, harbor_echo_final.mov, and pier_still.tif as finalized deliverables.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-5541."}, {"path": ["evidence", "1"], "text": "Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-5541.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-6620.", "negative_right": "Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-067", "id": "fast-43-diverse-223-067-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-5541.", "Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is a project file that remains editable, and it links source media consisting of the gulls.wav recording, the harbor_echo_final.mov documentary, and the pier_still.tif photograph.", "The asset assigned Harbor Echo transfer identifier HP-3 is not raw captured sound and is not a published final export.", "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export, is not raw captured sound, and is not editable; this covers gulls.wav, harbor_echo_final.mov, and pier_still.tif as finalized deliverables."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric and question criteria, so policy is preserved; the entity (harbor_mix.prproj), path, and time bindings from the question remain intact. The focus evidence consists of two complete factual sentences about ticket numbers, not policy text or instructions. The counterfactual changes only the ticket number tying harbor_mix.prproj to HE-6620 instead of HE-5541, which breaks the identity link to HP-3 but does not create a direct contradiction with any other evidence, remaining internally coherent (it merely renders the asset's classification less determinate). Neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-5541.\", \"Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is a project file that remains editable, and it links source media consisting of the gulls.wav recording, the harbor_echo_final.mov documentary, and the pier_still.tif photograph.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is not raw captured sound and is not a published final export.\", \"Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export, is not raw captured sound, and is not editable; this covers gulls.wav, harbor_echo_final.mov, and pier_still.tif as finalized deliverables.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-5541."}, {"path": ["evidence", "1"], "text": "Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-5541.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-6620.", "negative_right": "Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-067", "id": "fast-43-diverse-223-067-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number HE-6620.", "Ingest ticket number HE-5541 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is a project file that remains editable, and it links source media consisting of the gulls.wav recording, the harbor_echo_final.mov documentary, and the pier_still.tif photograph.", "The asset assigned Harbor Echo transfer identifier HP-3 is not raw captured sound and is not a published final export.", "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export, is not raw captured sound, and is not editable; this covers gulls.wav, harbor_echo_final.mov, and pier_still.tif as finalized deliverables."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric and question object unchanged, present two complete factual manifest-slot sentences as focus evidence, and the counterfactual’s slot change (M-27 vs M-14) coherently decouples harbor_mix.prproj from HP-3 without contradicting other stated facts; neither context reveals a gold label or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-14.\", \"Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe Premiere project file that remains editable and links source media including gulls.wav, harbor_echo_final.mov, and pier_still.tif.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every asset in the Harbor Echo transfer other than HP-3 is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-14."}, {"path": ["evidence", "1"], "text": "Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-14.", "negative_left": "The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-27.", "negative_right": "Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-068", "id": "fast-43-diverse-223-068-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-14.", "Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe Premiere project file that remains editable and links source media including gulls.wav, harbor_echo_final.mov, and pier_still.tif.", "HP-3 is not raw captured sound and is not a published final export.", "Every asset in the Harbor Echo transfer other than HP-3 is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric and question object unchanged, present two complete factual manifest-slot sentences as focus evidence, and the counterfactual’s slot change (M-27 vs M-14) coherently decouples harbor_mix.prproj from HP-3 without contradicting other stated facts; neither context reveals a gold label or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-14.\", \"Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe Premiere project file that remains editable and links source media including gulls.wav, harbor_echo_final.mov, and pier_still.tif.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every asset in the Harbor Echo transfer other than HP-3 is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-14."}, {"path": ["evidence", "1"], "text": "Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-14.", "negative_left": "The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-27.", "negative_right": "Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-068", "id": "fast-43-diverse-223-068-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer log records harbor_mix.prproj under manifest slot M-27.", "Manifest slot M-14 in the Harbor Echo transfer log is the slot recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe Premiere project file that remains editable and links source media including gulls.wav, harbor_echo_final.mov, and pier_still.tif.", "HP-3 is not raw captured sound and is not a published final export.", "Every asset in the Harbor Echo transfer other than HP-3 is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate the same rubric intact, matching the unchanged question criteria, so policy is preserved. The request and entity (harbor_mix.prproj) are unchanged across both contexts. The two focus evidence spans are plain factual statements, not policy or instructions. The counterfactual only swaps the checksum linking harbor_mix.prproj to entry 6c9d-44e instead of 5b8f-31a, breaking the coreference to HP-3 without contradicting any other stated fact, so it remains internally coherent. Neither context contains a gold label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 5b8f-31a.\", \"Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"HP-3 is an Adobe Premiere project file that remains fully editable and links to gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset cataloged in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 5b8f-31a."}, {"path": ["evidence", "1"], "text": "Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 5b8f-31a.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 6c9d-44e.", "negative_right": "Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-069", "id": "fast-43-diverse-223-069-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 5b8f-31a.", "Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "HP-3 is an Adobe Premiere project file that remains fully editable and links to gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset cataloged in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate the same rubric intact, matching the unchanged question criteria, so policy is preserved. The request and entity (harbor_mix.prproj) are unchanged across both contexts. The two focus evidence spans are plain factual statements, not policy or instructions. The counterfactual only swaps the checksum linking harbor_mix.prproj to entry 6c9d-44e instead of 5b8f-31a, breaking the coreference to HP-3 without contradicting any other stated fact, so it remains internally coherent. Neither context contains a gold label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 5b8f-31a.\", \"Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"HP-3 is an Adobe Premiere project file that remains fully editable and links to gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset cataloged in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 5b8f-31a."}, {"path": ["evidence", "1"], "text": "Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 5b8f-31a.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 6c9d-44e.", "negative_right": "Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-069", "id": "fast-43-diverse-223-069-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 6c9d-44e.", "Checksum entry 5b8f-31a in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "HP-3 is an Adobe Premiere project file that remains fully editable and links to gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset cataloged in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric verbatim and the same request; the counterfactual only changes harbor_mix.prproj's checksum to 6c2e-77a, which logically detaches it from HP-3 rather than contradicting any other stated fact, so it remains internally coherent; both focus evidence spans are complete factual sentences, not rules or instructions; neither context states a routing decision, code, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 4b8f-31d.\", \"Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"HP-3 is described in the manifest as an Adobe project file that remains editable and links to the gulls.wav field recording, the harbor_echo_final.mov documentary export, and the pier_still.tif photograph as its source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is logged as a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 4b8f-31d."}, {"path": ["evidence", "1"], "text": "Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 4b8f-31d.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 6c2e-77a.", "negative_right": "Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-075", "id": "fast-43-diverse-223-075-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 4b8f-31d.", "Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "HP-3 is described in the manifest as an Adobe project file that remains editable and links to the gulls.wav field recording, the harbor_echo_final.mov documentary export, and the pier_still.tif photograph as its source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is logged as a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric verbatim and the same request; the counterfactual only changes harbor_mix.prproj's checksum to 6c2e-77a, which logically detaches it from HP-3 rather than contradicting any other stated fact, so it remains internally coherent; both focus evidence spans are complete factual sentences, not rules or instructions; neither context states a routing decision, code, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 4b8f-31d.\", \"Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"HP-3 is described in the manifest as an Adobe project file that remains editable and links to the gulls.wav field recording, the harbor_echo_final.mov documentary export, and the pier_still.tif photograph as its source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is logged as a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 4b8f-31d."}, {"path": ["evidence", "1"], "text": "Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 4b8f-31d.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 6c2e-77a.", "negative_right": "Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-075", "id": "fast-43-diverse-223-075-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under checksum entry 6c2e-77a.", "Checksum entry 4b8f-31d in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "HP-3 is described in the manifest as an Adobe project file that remains editable and links to the gulls.wav field recording, the harbor_echo_final.mov documentary export, and the pier_still.tif photograph as its source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is logged as a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical rubric and request; the counterfactual only changes the batch tag on harbor_mix.prproj from QX-58 to QX-91, breaking the link to HP-3 without contradicting any other evidence, which is a coherent single-sentence change; the two focus evidence spans are plain factual statements, not policy or rule text; no gold answer, code, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-58.\", \"Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-58."}, {"path": ["evidence", "1"], "text": "Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-58.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-91.", "negative_right": "Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-079", "id": "fast-43-diverse-223-079-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-58.", "Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical rubric and request; the counterfactual only changes the batch tag on harbor_mix.prproj from QX-58 to QX-91, breaking the link to HP-3 without contradicting any other evidence, which is a coherent single-sentence change; the two focus evidence spans are plain factual statements, not policy or rule text; no gold answer, code, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-58.\", \"Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-58."}, {"path": ["evidence", "1"], "text": "Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-58.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-91.", "negative_right": "Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-079", "id": "fast-43-diverse-223-079-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under batch tag QX-91.", "Batch tag QX-58 in the Harbor Echo manifest is the tag recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links to the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric and question, so governing policy is preserved. The request and entity (harbor_mix.prproj) remain identical across contexts. Focus evidence are two complete factual manifest statements, not policy text. The counterfactual only swaps the ticket number for harbor_mix.prproj (T-5521 to T-8890), which breaks the identifying link to HP-3 without creating any direct contradiction or duplicate measurement, so it remains coherent. Neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5521.\", \"Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file cataloged as an editable timeline that links gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media within the transfer.\", \"HP-3 is not raw captured sound and is not a published final export; it remains a working project rather than a finished deliverable.\", \"Every other asset in the Harbor Echo transfer besides HP-3, including the WAV, MOV, and TIFF items, is a published final export, is not raw captured sound, and is not editable, per the manifest’s asset-status column.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5521."}, {"path": ["evidence", "1"], "text": "Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5521.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-8890.", "negative_right": "Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-082", "id": "fast-43-diverse-223-082-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5521.", "Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file cataloged as an editable timeline that links gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media within the transfer.", "HP-3 is not raw captured sound and is not a published final export; it remains a working project rather than a finished deliverable.", "Every other asset in the Harbor Echo transfer besides HP-3, including the WAV, MOV, and TIFF items, is a published final export, is not raw captured sound, and is not editable, per the manifest’s asset-status column."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric and question, so governing policy is preserved. The request and entity (harbor_mix.prproj) remain identical across contexts. Focus evidence are two complete factual manifest statements, not policy text. The counterfactual only swaps the ticket number for harbor_mix.prproj (T-5521 to T-8890), which breaks the identifying link to HP-3 without creating any direct contradiction or duplicate measurement, so it remains coherent. Neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5521.\", \"Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file cataloged as an editable timeline that links gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media within the transfer.\", \"HP-3 is not raw captured sound and is not a published final export; it remains a working project rather than a finished deliverable.\", \"Every other asset in the Harbor Echo transfer besides HP-3, including the WAV, MOV, and TIFF items, is a published final export, is not raw captured sound, and is not editable, per the manifest’s asset-status column.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5521."}, {"path": ["evidence", "1"], "text": "Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5521.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-8890.", "negative_right": "Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-082", "id": "fast-43-diverse-223-082-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-8890.", "Ingest ticket number T-5521 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file cataloged as an editable timeline that links gulls.wav, harbor_echo_final.mov, and pier_still.tif as source media within the transfer.", "HP-3 is not raw captured sound and is not a published final export; it remains a working project rather than a finished deliverable.", "Every other asset in the Harbor Echo transfer besides HP-3, including the WAV, MOV, and TIFF items, is a published final export, is not raw captured sound, and is not editable, per the manifest’s asset-status column."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric and question intact, target the same asset/request, and use plain factual evidence sentences; the counterfactual only swaps the ticket number so harbor_mix.prproj no longer links to HP-3, which coherently shifts the case toward unresolved identity without contradicting other stated facts or leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5591.\",\"Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.\",\"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media into a single documentary timeline.\",\"HP-3 is not raw captured sound and is not a published final export.\",\"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5591."}, {"path": ["evidence", "1"], "text": "Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5591.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-6042.", "negative_right": "Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-083", "id": "fast-43-diverse-223-083-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5591.", "Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media into a single documentary timeline.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the rubric and question intact, target the same asset/request, and use plain factual evidence sentences; the counterfactual only swaps the ticket number so harbor_mix.prproj no longer links to HP-3, which coherently shifts the case toward unresolved identity without contradicting other stated facts or leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5591.\",\"Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.\",\"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media into a single documentary timeline.\",\"HP-3 is not raw captured sound and is not a published final export.\",\"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5591."}, {"path": ["evidence", "1"], "text": "Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-5591.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-6042.", "negative_right": "Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-083", "id": "fast-43-diverse-223-083-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ingest ticket number T-6042.", "Ingest ticket number T-5591 in the Harbor Echo manifest is the ticket recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media into a single documentary timeline.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same rubric, entity, and request untouched; the counterfactual changes only the ledger number logging harbor_mix.prproj from L-58 to L-91, breaking its link to HP-3 without contradicting any other stated fact, count, or measurement, so it remains internally coherent while plausibly shifting the classification toward Metadata Review; the two focus evidence sentences are complete factual statements, not policy or instructions, and neither context states or hints at the intended output label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-58.\", \"Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media into a single timeline.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-58."}, {"path": ["evidence", "1"], "text": "Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-58.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-91.", "negative_right": "Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-084", "id": "fast-43-diverse-223-084-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-58.", "Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media into a single timeline.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same rubric, entity, and request untouched; the counterfactual changes only the ledger number logging harbor_mix.prproj from L-58 to L-91, breaking its link to HP-3 without contradicting any other stated fact, count, or measurement, so it remains internally coherent while plausibly shifting the classification toward Metadata Review; the two focus evidence sentences are complete factual statements, not policy or instructions, and neither context states or hints at the intended output label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\": \"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\", \"evidence\": [\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-58.\", \"Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\", \"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media into a single timeline.\", \"HP-3 is not raw captured sound and is not a published final export.\", \"Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable.\"], \"request\": \"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-58."}, {"path": ["evidence", "1"], "text": "Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-58.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-91.", "negative_right": "Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-084", "id": "fast-43-diverse-223-084-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry L-91.", "Ledger entry L-58 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe project file that remains editable and links the gulls.wav, harbor_echo_final.mov, and pier_still.tif source media into a single timeline.", "HP-3 is not raw captured sound and is not a published final export.", "Every other asset in the Harbor Echo transfer, aside from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric and identical asset/request bindings, so policy and bindings are preserved via the untouched questions object. The focus evidence consists of two complete factual sentences about ledger entries, not policy text or instructions. The counterfactual changes only the ledger number tying harbor_mix.prproj to LX-408, breaking the identity chain to HP-3 without asserting a contradictory duplicate fact, so it remains internally coherent even though it now yields insufficient evidence to link the file to HP-3. Neither context contains a gold label, rule table, or explicit output instruction beyond the rubric already in the question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-408.\",\"Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\",\"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe Premiere project file that remains editable and links source media including gulls.wav, harbor_echo_final.mov, and pier_still.tif.\",\"HP-3 is neither raw captured sound nor a published final export.\",\"Every other asset cataloged in the Harbor Echo transfer, apart from HP-3, is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-408."}, {"path": ["evidence", "1"], "text": "Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-408.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-902.", "negative_right": "Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-085", "id": "fast-43-diverse-223-085-base", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-408.", "Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe Premiere project file that remains editable and links source media including gulls.wav, harbor_echo_final.mov, and pier_still.tif.", "HP-3 is neither raw captured sound nor a published final export.", "Every other asset cataloged in the Harbor Echo transfer, apart from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "production_projects"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged rubric and identical asset/request bindings, so policy and bindings are preserved via the untouched questions object. The focus evidence consists of two complete factual sentences about ledger entries, not policy text or instructions. The counterfactual changes only the ledger number tying harbor_mix.prproj to LX-408, breaking the identity chain to HP-3 without asserting a contradictory duplicate fact, so it remains internally coherent even though it now yields insufficient evidence to link the file to HP-3. Neither context contains a gold label, rule table, or explicit output instruction beyond the rubric already in the question.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "refuted", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "refuted", "A6": "refuted", "A7": "supported", "A8": "supported", "A9": "supported"}, "remove_left": {"A1": "unknown"}, "remove_right": {"A1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A1": "unknown"}, "negative_pair": {"A1": "refuted"}, "negative_sentence": {"A1": "unknown"}, "positive_pair": {"A1": "supported"}, "right": {"A1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universal atoms each apply one property over an explicit set. The focus is the factual identity relation between the requested asset and HP-3. Both assignments are jointly realizable: in the base the requested asset is HP-3, while in the counter it is a different Harbor Echo asset governed by the universal statements; all non-focus atom states can remain unchanged. Policy evidence cites only original-state material, and the unchanged questions already preserve the full routing rubric.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions identify the requested asset with HP-3 and establish that HP-3 is an editable project file linking source media, which is sufficient for Production Projects. They also refute the raw-sound and published-final-export alternatives.", "rule_index": 0, "sound": true}, {"reason": "Refuting identity with HP-3 while placing the requested asset in the Harbor Echo transfer brings it within the stated universal facts about every other transfer asset. Those facts establish that it is a published final export and is neither raw captured sound nor editable, making Finished Works sufficient and excluding the competing substantive collections.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The asset requested as harbor_mix.prproj is the asset assigned Harbor Echo transfer identifier HP-3."}, {"id": "A2", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a project file."}, {"id": "A3", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is editable."}, {"id": "A4", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 links source media."}, {"id": "A5", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is raw captured sound."}, {"id": "A6", "statement": "The asset assigned Harbor Echo transfer identifier HP-3 is a published final export."}, {"id": "A7", "statement": "The asset requested as harbor_mix.prproj belongs to the Harbor Echo transfer."}, {"id": "A8", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is a published final export."}, {"id": "A9", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not raw captured sound."}, {"id": "A10", "statement": "Every asset in the Harbor Echo transfer other than the asset assigned transfer identifier HP-3 is not editable."}], "base_state_json": "{\"context\":\"Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.\",\"evidence\":[\"The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-408.\",\"Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.\",\"The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe Premiere project file that remains editable and links source media including gulls.wav, harbor_echo_final.mov, and pier_still.tif.\",\"HP-3 is neither raw captured sound nor a published final export.\",\"Every other asset cataloged in the Harbor Echo transfer, apart from HP-3, is a published final export, is not raw captured sound, and is not editable.\"],\"request\":\"Where should Lena route harbor_mix.prproj?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A1", "focus_evidence": [{"path": ["evidence", "0"], "text": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-408."}, {"path": ["evidence", "1"], "text": "Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}], "policy_evidence": [{"path": ["context"], "text": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies."}, {"path": ["request"], "text": "Where should Lena route harbor_mix.prproj?"}], "rules": [{"justification": "The requested asset is identified with HP-3, which is an editable project file that links source media. It is explicitly neither raw captured sound nor a published final export, excluding the competing substantive collections.", "target": "production_projects", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "refuted"}]}, {"justification": "The requested asset belongs to the Harbor Echo transfer but is not HP-3, so it falls under the universal facts for the transfer’s other assets: it is a published final export, is not raw captured sound, and is not editable. The supplied evidence therefore distinguishes Finished Works from the competing collections.", "target": "finished_works", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}]}, "verified_pair": {"left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-408.", "negative_left": "The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-902.", "negative_right": "Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "right": "Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-223-085", "id": "fast-43-diverse-223-085-counterfactual", "input": {"questions": {"decision": {"criteria": {"field_recordings": "Route to Field Recordings only if the requested asset is raw captured sound.", "finished_works": "Route to Finished Works only if the requested asset is a published final export.", "metadata_review": "Route to Metadata Review only if the supplied evidence cannot distinguish among the three collections.", "production_projects": "Route to Production Projects if the requested asset is an editable project that arranges or links source media."}, "instructions": "Resolve the note’s coreferences using the manifest and route the requested file under the supplied collection rubric. Select exactly one option.", "type": "choice"}}, "state": {"context": "Digital archivist Lena must route one ambiguous asset from the fictional Harbor Echo transfer. The collection manager’s rubric is: raw captured sound goes to Field Recordings; published final exports go to Finished Works; editable files that arrange source media go to Production Projects; use Metadata Review only when supplied evidence cannot establish which of those applies.", "evidence": ["The Harbor Echo transfer manifest logs harbor_mix.prproj under ledger entry LX-902.", "Ledger entry LX-408 in the Harbor Echo manifest is the entry recorded for the asset assigned Harbor Echo transfer identifier HP-3.", "The asset assigned Harbor Echo transfer identifier HP-3 is an Adobe Premiere project file that remains editable and links source media including gulls.wav, harbor_echo_final.mov, and pier_still.tif.", "HP-3 is neither raw captured sound nor a published final export.", "Every other asset cataloged in the Harbor Echo transfer, apart from HP-3, is a published final export, is not raw captured sound, and is not editable."], "request": "Where should Lena route harbor_mix.prproj?"}}, "method": "c2d", "provenance": {"source_id": "diverse-223", "source_is_synthetic": true, "source_sha256": "830836d20265e2da45a83aa5b05b699cbf6238307b690e6aec00e0b9ac26711f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "finished_works"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the six-field policy and thresholds verbatim via the original questions object; the base and counterfactual contexts each preserve entity/path/time bindings (record #4471, March 3); focus evidence spans are two complete factual sentences; the counterfactual coherently swaps CC-BY-4.0 rights for a blank rights field without contradicting other unchanged facts; neither context contains gold answers, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "full_context_fact_states": {"base": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "counterfactual": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "remove_left": {"rights_present": "unknown"}, "remove_right": {"rights_present": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"rights_present": "unknown"}, "negative_pair": {"rights_present": "refuted"}, "negative_sentence": {"rights_present": "unknown"}, "positive_pair": {"rights_present": "supported"}, "right": {"rights_present": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one metadata-presence relationship. The focus atom, rights presence, is factual rather than a policy classification. The base and counter assignments differ only in rights presence and are both realizable: the base yields all required fields plus one optional field, while the counter yields exactly one missing required field. The cited policy evidence is a valid citation from the original state and accurately preserves the metadata-field bindings and level rules, although the unchanged question already contains the governing criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that exactly one required field—rights—is missing while the other five required fields are present. Level 3 applies regardless of optional-field status, so the target is entailed.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all six required fields are present and exactly one optional field, subject, is present while caption is absent. This is sufficient for level 4 and excludes level 5.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "title_present", "statement": "The supplied record for “Harbor Bells” has title metadata present."}, {"id": "creator_present", "statement": "The supplied record for “Harbor Bells” has creator metadata present."}, {"id": "date_present", "statement": "The supplied record for “Harbor Bells” has date metadata present."}, {"id": "media_type_present", "statement": "The supplied record for “Harbor Bells” has media type metadata present."}, {"id": "duration_present", "statement": "The supplied record for “Harbor Bells” has duration metadata present."}, {"id": "rights_present", "statement": "The supplied record for “Harbor Bells” has rights metadata present."}, {"id": "subject_present", "statement": "The supplied record for “Harbor Bells” has subject metadata present."}, {"id": "caption_present", "statement": "The supplied record for “Harbor Bells” has caption metadata present."}], "base_state_json": "[{\"speaker\":\"Digital archivist\",\"text\":\"The archivist reviewed the metadata fields of the \\\"Harbor Bells\\\" record catalog entry #4471 on March 3.\"},{\"speaker\":\"Digital archivist\",\"text\":\"The record lists title \\\"Harbor Bells,\\\" creator Nia Voss, date 2031-06-14, media type WAV, duration 02:00, and subject \\\"coastal sound.\\\" It has no caption.\"},{\"speaker\":\"Collection manager\",\"text\":\"Catalog entry #4471's rights field lists license code CC-BY-4.0.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing.\"},{\"speaker\":\"Media librarian\",\"text\":\"Which completeness level applies to this record?\"}]", "base_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "counter_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "focus_atom": "rights_present", "focus_evidence": [{"path": ["0", "text"], "text": "The archivist reviewed the metadata fields of the \"Harbor Bells\" record catalog entry #4471 on March 3."}, {"path": ["2", "text"], "text": "Catalog entry #4471's rights field lists license code CC-BY-4.0."}], "policy_evidence": [{"path": ["1", "text"], "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}], "rules": [{"justification": "Rights is the only missing required field; the policy assigns level_3 when exactly one of the six required fields is missing, regardless of optional fields.", "target": "level_3", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}, {"justification": "All six required fields are present and exactly one optional field, subject, is present; the policy assigns level_4.", "target": "level_4", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}]}, "verified_pair": {"left": "The archivist reviewed the metadata fields of the \"Harbor Bells\" record catalog entry #4471 on March 3.", "negative_left": "The archivist reviewed the metadata fields of the \"Harbor Bells\" record catalog entry #4471 on March 3.", "negative_right": "Catalog entry #4471's rights field is left blank with no license code.", "right": "Catalog entry #4471's rights field lists license code CC-BY-4.0."}, "verifier_independent_model": false}, "family": "fast-43-diverse-224-001", "id": "fast-43-diverse-224-001-base", "input": {"questions": {"decision": {"criteria": {"level_1": "Four or more of the six required metadata fields are missing.", "level_2": "Two or three of the six required metadata fields are missing.", "level_3": "Exactly one of the six required metadata fields is missing, regardless of optional fields.", "level_4": "All six required metadata fields are present, with zero or exactly one optional field present.", "level_5": "All six required metadata fields and both optional fields are present."}, "instructions": "Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata.", "type": "choice"}}, "state": [{"speaker": "Digital archivist", "text": "The archivist reviewed the metadata fields of the \"Harbor Bells\" record catalog entry #4471 on March 3."}, {"speaker": "Digital archivist", "text": "The record lists title \"Harbor Bells,\" creator Nia Voss, date 2031-06-14, media type WAV, duration 02:00, and subject \"coastal sound.\" It has no caption."}, {"speaker": "Collection manager", "text": "Catalog entry #4471's rights field lists license code CC-BY-4.0."}, {"speaker": "Metadata specialist", "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}, {"speaker": "Media librarian", "text": "Which completeness level applies to this record?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-224", "source_is_synthetic": true, "source_sha256": "42be2105c51ead25ff90f2762bb2402ddf400ae28dfdeec25ee9b7154cc0ab7b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_4"}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the six-field policy and thresholds verbatim via the original questions object; the base and counterfactual contexts each preserve entity/path/time bindings (record #4471, March 3); focus evidence spans are two complete factual sentences; the counterfactual coherently swaps CC-BY-4.0 rights for a blank rights field without contradicting other unchanged facts; neither context contains gold answers, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "full_context_fact_states": {"base": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "supported", "subject_present": "supported", "title_present": "supported"}, "counterfactual": {"caption_present": "refuted", "creator_present": "supported", "date_present": "supported", "duration_present": "supported", "media_type_present": "supported", "rights_present": "refuted", "subject_present": "supported", "title_present": "supported"}, "remove_left": {"rights_present": "unknown"}, "remove_right": {"rights_present": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"rights_present": "unknown"}, "negative_pair": {"rights_present": "refuted"}, "negative_sentence": {"rights_present": "unknown"}, "positive_pair": {"rights_present": "supported"}, "right": {"rights_present": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one metadata-presence relationship. The focus atom, rights presence, is factual rather than a policy classification. The base and counter assignments differ only in rights presence and are both realizable: the base yields all required fields plus one optional field, while the counter yields exactly one missing required field. The cited policy evidence is a valid citation from the original state and accurately preserves the metadata-field bindings and level rules, although the unchanged question already contains the governing criteria.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes that exactly one required field—rights—is missing while the other five required fields are present. Level 3 applies regardless of optional-field status, so the target is entailed.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that all six required fields are present and exactly one optional field, subject, is present while caption is absent. This is sufficient for level 4 and excludes level 5.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "title_present", "statement": "The supplied record for “Harbor Bells” has title metadata present."}, {"id": "creator_present", "statement": "The supplied record for “Harbor Bells” has creator metadata present."}, {"id": "date_present", "statement": "The supplied record for “Harbor Bells” has date metadata present."}, {"id": "media_type_present", "statement": "The supplied record for “Harbor Bells” has media type metadata present."}, {"id": "duration_present", "statement": "The supplied record for “Harbor Bells” has duration metadata present."}, {"id": "rights_present", "statement": "The supplied record for “Harbor Bells” has rights metadata present."}, {"id": "subject_present", "statement": "The supplied record for “Harbor Bells” has subject metadata present."}, {"id": "caption_present", "statement": "The supplied record for “Harbor Bells” has caption metadata present."}], "base_state_json": "[{\"speaker\":\"Digital archivist\",\"text\":\"The archivist reviewed the metadata fields of the \\\"Harbor Bells\\\" record catalog entry #4471 on March 3.\"},{\"speaker\":\"Digital archivist\",\"text\":\"The record lists title \\\"Harbor Bells,\\\" creator Nia Voss, date 2031-06-14, media type WAV, duration 02:00, and subject \\\"coastal sound.\\\" It has no caption.\"},{\"speaker\":\"Collection manager\",\"text\":\"Catalog entry #4471's rights field lists license code CC-BY-4.0.\"},{\"speaker\":\"Metadata specialist\",\"text\":\"Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing.\"},{\"speaker\":\"Media librarian\",\"text\":\"Which completeness level applies to this record?\"}]", "base_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "counter_states": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}], "focus_atom": "rights_present", "focus_evidence": [{"path": ["0", "text"], "text": "The archivist reviewed the metadata fields of the \"Harbor Bells\" record catalog entry #4471 on March 3."}, {"path": ["2", "text"], "text": "Catalog entry #4471's rights field lists license code CC-BY-4.0."}], "policy_evidence": [{"path": ["1", "text"], "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}], "rules": [{"justification": "Rights is the only missing required field; the policy assigns level_3 when exactly one of the six required fields is missing, regardless of optional fields.", "target": "level_3", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "refuted"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}, {"justification": "All six required fields are present and exactly one optional field, subject, is present; the policy assigns level_4.", "target": "level_4", "when": [{"atom_id": "title_present", "state": "supported"}, {"atom_id": "creator_present", "state": "supported"}, {"atom_id": "date_present", "state": "supported"}, {"atom_id": "media_type_present", "state": "supported"}, {"atom_id": "duration_present", "state": "supported"}, {"atom_id": "rights_present", "state": "supported"}, {"atom_id": "subject_present", "state": "supported"}, {"atom_id": "caption_present", "state": "refuted"}]}]}, "verified_pair": {"left": "The archivist reviewed the metadata fields of the \"Harbor Bells\" record catalog entry #4471 on March 3.", "negative_left": "The archivist reviewed the metadata fields of the \"Harbor Bells\" record catalog entry #4471 on March 3.", "negative_right": "Catalog entry #4471's rights field is left blank with no license code.", "right": "Catalog entry #4471's rights field lists license code CC-BY-4.0."}, "verifier_independent_model": false}, "family": "fast-43-diverse-224-001", "id": "fast-43-diverse-224-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"level_1": "Four or more of the six required metadata fields are missing.", "level_2": "Two or three of the six required metadata fields are missing.", "level_3": "Exactly one of the six required metadata fields is missing, regardless of optional fields.", "level_4": "All six required metadata fields are present, with zero or exactly one optional field present.", "level_5": "All six required metadata fields and both optional fields are present."}, "instructions": "Select the single metadata completeness level dictated by the supplied policy and verified evidence. Do not count unrelated folder items as record metadata.", "type": "choice"}}, "state": [{"speaker": "Digital archivist", "text": "The archivist reviewed the metadata fields of the \"Harbor Bells\" record catalog entry #4471 on March 3."}, {"speaker": "Digital archivist", "text": "The record lists title \"Harbor Bells,\" creator Nia Voss, date 2031-06-14, media type WAV, duration 02:00, and subject \"coastal sound.\" It has no caption."}, {"speaker": "Collection manager", "text": "Catalog entry #4471's rights field is left blank with no license code."}, {"speaker": "Metadata specialist", "text": "Completeness policy has six required fields: title, creator, date, media type, duration, and rights. Optional fields are subject and caption. Rate 5 for six required plus both optional; 4 for six required plus zero or one optional; 3 for exactly one required missing; 2 for two or three required missing; 1 for four or more required missing."}, {"speaker": "Media librarian", "text": "Which completeness level applies to this record?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-224", "source_is_synthetic": true, "source_sha256": "42be2105c51ead25ff90f2762bb2402ddf400ae28dfdeec25ee9b7154cc0ab7b", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "level_3"}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full archive rubric and question unchanged, focus on the same package/date field, and use two factual sentences as evidence; the counterfactual creates a plausible mismatch between recorded and manifest-verified dates (testing an unverified core field) without contradicting other unchanged facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; none is a bundled final classification. The focus atom concerns factual identity between two date values, not archive policy. The base and counter assignments are realizable with only that identity changing: the package date can either match or differ from the verified manifest date while all other atom states remain fixed. Policy evidence preserves the substantive Level 4 requirements and the unverified-core-field cap from the original state. The additional request citation is unnecessary but does not omit or distort policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes verification of title, creator, rights, and collection; establishes that the recorded date matches the verified date supplied by the signed manifest; and supplies an access aid. This is sufficient for every stated Level 4 requirement, so the true target follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the identity atom means the recorded date differs from the verified manifest date. Given that both values are present and the manifest value is verified, the package's date core field is unverified. The explicit cap then limits the rating to Level 3, so the false target follows regardless of the other satisfied requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The title core field of the submitted “River Signal” package is verified."}, {"id": "a2", "statement": "The creator core field of the submitted “River Signal” package is verified."}, {"id": "a3", "statement": "The rights core field of the submitted “River Signal” package is verified."}, {"id": "a4", "statement": "The collection core field of the submitted “River Signal” package is verified."}, {"id": "a5", "statement": "The date core field of the submitted “River Signal” package records a date value."}, {"id": "a6", "statement": "The signed manifest supplies the verified date value for the submitted “River Signal” package."}, {"id": "a7", "statement": "The date value recorded in the submitted “River Signal” package’s date core field is identical to the verified date value supplied by the signed manifest."}, {"id": "a8", "statement": "The submitted “River Signal” package has a supplied access aid."}], "base_state_json": "{\"context\":\"A metadata specialist reviews the fictional \\u201cRiver Signal\\u201d package containing WAV audio, MOV video, TIFF stills, and a Premiere project. The title core field is verified as \\u201cRiver Signal.\\u201d The creator core field is verified as Imani Vale. The rights core field is verified as open educational rights. The collection core field is verified as Experimental Documentary, per a collection-manager routing note. The date core field of the submitted \\u201cRiver Signal\\u201d package records the value 1987-06-14. The signed manifest for the submitted \\u201cRiver Signal\\u201d package specifies the verified date value as 1987-06-14. A timecoded shot list is attached and serves as the package's supplied access aid, though no transcript or subject keywords are present. Archive rubric: Level 4 requires every core field\\u2014title, creator, date, rights, and collection\\u2014to be verified, plus one access aid. Level 5 additionally requires both a transcript and subject keywords. Any unverified core field caps the record at Level 3.\",\"request\":\"Should the metadata completeness rating be Level 4 or higher?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": ["context"], "text": "The date core field of the submitted “River Signal” package records the value 1987-06-14."}, {"path": ["context"], "text": "The signed manifest for the submitted “River Signal” package specifies the verified date value as 1987-06-14."}], "policy_evidence": [{"path": ["request"], "text": "Should the metadata completeness rating be Level 4 or higher?"}, {"path": ["context"], "text": "Archive rubric: Level 4 requires every core field—title, creator, date, rights, and collection—to be verified, plus one access aid. Level 5 additionally requires both a transcript and subject keywords. Any unverified core field caps the record at Level 3."}], "rules": [{"justification": "The title, creator, rights, and collection core fields are verified. The recorded date matches the verified date supplied by the signed manifest, so the date core field is also verified. With a supplied access aid, every Level 4 requirement is met.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The date recorded in the package’s date core field differs from the verified date supplied by the signed manifest. The date core field is therefore unverified, which caps the record at Level 3 even though the other core fields are verified and an access aid is supplied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The date core field of the submitted “River Signal” package records the value 1987-06-14.", "negative_left": "The date core field of the submitted “River Signal” package records the value 1987-06-14.", "negative_right": "The signed manifest for the submitted “River Signal” package specifies the verified date value as 1992-11-03.", "right": "The signed manifest for the submitted “River Signal” package specifies the verified date value as 1987-06-14."}, "verifier_independent_model": false}, "family": "fast-43-diverse-226-012", "id": "fast-43-diverse-226-012-base", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 3 or below because a core field is unverified or no access aid is supplied.", "true": "Yes — assign Level 4 or higher because all Level 4 requirements are met."}, "instructions": "Answer yes if the supplied evidence satisfies the archive’s stated requirements for Level 4 or Level 5; otherwise answer no.", "type": "noul"}}, "state": {"context": "A metadata specialist reviews the fictional “River Signal” package containing WAV audio, MOV video, TIFF stills, and a Premiere project. The title core field is verified as “River Signal.” The creator core field is verified as Imani Vale. The rights core field is verified as open educational rights. The collection core field is verified as Experimental Documentary, per a collection-manager routing note. The date core field of the submitted “River Signal” package records the value 1987-06-14. The signed manifest for the submitted “River Signal” package specifies the verified date value as 1987-06-14. A timecoded shot list is attached and serves as the package's supplied access aid, though no transcript or subject keywords are present. Archive rubric: Level 4 requires every core field—title, creator, date, rights, and collection—to be verified, plus one access aid. Level 5 additionally requires both a transcript and subject keywords. Any unverified core field caps the record at Level 3.", "request": "Should the metadata completeness rating be Level 4 or higher?"}}, "method": "c2d", "provenance": {"source_id": "diverse-226", "source_is_synthetic": true, "source_sha256": "80eaffc080317c92eedee9d75d9492e6904cc7cb6e817c261ad735e9bd4137ed", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full archive rubric and question unchanged, focus on the same package/date field, and use two factual sentences as evidence; the counterfactual creates a plausible mismatch between recorded and manifest-verified dates (testing an unverified core field) without contradicting other unchanged facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "refuted", "a8": "supported"}, "remove_left": {"a7": "unknown"}, "remove_right": {"a7": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a7": "unknown"}, "negative_pair": {"a7": "refuted"}, "negative_sentence": {"a7": "unknown"}, "positive_pair": {"a7": "supported"}, "right": {"a7": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses a single factual relation or status; none is a bundled final classification. The focus atom concerns factual identity between two date values, not archive policy. The base and counter assignments are realizable with only that identity changing: the package date can either match or differ from the verified manifest date while all other atom states remain fixed. Policy evidence preserves the substantive Level 4 requirements and the unverified-core-field cap from the original state. The additional request citation is unnecessary but does not omit or distort policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes verification of title, creator, rights, and collection; establishes that the recorded date matches the verified date supplied by the signed manifest; and supplies an access aid. This is sufficient for every stated Level 4 requirement, so the true target follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the identity atom means the recorded date differs from the verified manifest date. Given that both values are present and the manifest value is verified, the package's date core field is unverified. The explicit cap then limits the rating to Level 3, so the false target follows regardless of the other satisfied requirements.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The title core field of the submitted “River Signal” package is verified."}, {"id": "a2", "statement": "The creator core field of the submitted “River Signal” package is verified."}, {"id": "a3", "statement": "The rights core field of the submitted “River Signal” package is verified."}, {"id": "a4", "statement": "The collection core field of the submitted “River Signal” package is verified."}, {"id": "a5", "statement": "The date core field of the submitted “River Signal” package records a date value."}, {"id": "a6", "statement": "The signed manifest supplies the verified date value for the submitted “River Signal” package."}, {"id": "a7", "statement": "The date value recorded in the submitted “River Signal” package’s date core field is identical to the verified date value supplied by the signed manifest."}, {"id": "a8", "statement": "The submitted “River Signal” package has a supplied access aid."}], "base_state_json": "{\"context\":\"A metadata specialist reviews the fictional \\u201cRiver Signal\\u201d package containing WAV audio, MOV video, TIFF stills, and a Premiere project. The title core field is verified as \\u201cRiver Signal.\\u201d The creator core field is verified as Imani Vale. The rights core field is verified as open educational rights. The collection core field is verified as Experimental Documentary, per a collection-manager routing note. The date core field of the submitted \\u201cRiver Signal\\u201d package records the value 1987-06-14. The signed manifest for the submitted \\u201cRiver Signal\\u201d package specifies the verified date value as 1987-06-14. A timecoded shot list is attached and serves as the package's supplied access aid, though no transcript or subject keywords are present. Archive rubric: Level 4 requires every core field\\u2014title, creator, date, rights, and collection\\u2014to be verified, plus one access aid. Level 5 additionally requires both a transcript and subject keywords. Any unverified core field caps the record at Level 3.\",\"request\":\"Should the metadata completeness rating be Level 4 or higher?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}], "focus_atom": "a7", "focus_evidence": [{"path": ["context"], "text": "The date core field of the submitted “River Signal” package records the value 1987-06-14."}, {"path": ["context"], "text": "The signed manifest for the submitted “River Signal” package specifies the verified date value as 1987-06-14."}], "policy_evidence": [{"path": ["request"], "text": "Should the metadata completeness rating be Level 4 or higher?"}, {"path": ["context"], "text": "Archive rubric: Level 4 requires every core field—title, creator, date, rights, and collection—to be verified, plus one access aid. Level 5 additionally requires both a transcript and subject keywords. Any unverified core field caps the record at Level 3."}], "rules": [{"justification": "The title, creator, rights, and collection core fields are verified. The recorded date matches the verified date supplied by the signed manifest, so the date core field is also verified. With a supplied access aid, every Level 4 requirement is met.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}]}, {"justification": "The date recorded in the package’s date core field differs from the verified date supplied by the signed manifest. The date core field is therefore unverified, which caps the record at Level 3 even though the other core fields are verified and an access aid is supplied.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "refuted"}, {"atom_id": "a8", "state": "supported"}]}]}, "verified_pair": {"left": "The date core field of the submitted “River Signal” package records the value 1987-06-14.", "negative_left": "The date core field of the submitted “River Signal” package records the value 1987-06-14.", "negative_right": "The signed manifest for the submitted “River Signal” package specifies the verified date value as 1992-11-03.", "right": "The signed manifest for the submitted “River Signal” package specifies the verified date value as 1987-06-14."}, "verifier_independent_model": false}, "family": "fast-43-diverse-226-012", "id": "fast-43-diverse-226-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — assign Level 3 or below because a core field is unverified or no access aid is supplied.", "true": "Yes — assign Level 4 or higher because all Level 4 requirements are met."}, "instructions": "Answer yes if the supplied evidence satisfies the archive’s stated requirements for Level 4 or Level 5; otherwise answer no.", "type": "noul"}}, "state": {"context": "A metadata specialist reviews the fictional “River Signal” package containing WAV audio, MOV video, TIFF stills, and a Premiere project. The title core field is verified as “River Signal.” The creator core field is verified as Imani Vale. The rights core field is verified as open educational rights. The collection core field is verified as Experimental Documentary, per a collection-manager routing note. The date core field of the submitted “River Signal” package records the value 1987-06-14. The signed manifest for the submitted “River Signal” package specifies the verified date value as 1992-11-03. A timecoded shot list is attached and serves as the package's supplied access aid, though no transcript or subject keywords are present. Archive rubric: Level 4 requires every core field—title, creator, date, rights, and collection—to be verified, plus one access aid. Level 5 additionally requires both a transcript and subject keywords. Any unverified core field caps the record at Level 3.", "request": "Should the metadata completeness rating be Level 4 or higher?"}}, "method": "c2d", "provenance": {"source_id": "diverse-226", "source_is_synthetic": true, "source_sha256": "80eaffc080317c92eedee9d75d9492e6904cc7cb6e817c261ad735e9bd4137ed", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional instruction and all governing policy elements already covered by the unchanged questions object; only the birth year evidence changes (1990 vs 1985), which is a legitimate counterfactual since it flips the minor-status computation without contradicting other stated facts like blur status or checksums; both focus evidence spans are complete factual sentences, not rules or instructions; no gold answer, label, or rationale is embedded, leaving the age/minor determination to be inferred by the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is June 14, 2005. The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1990.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is June 14, 2005."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1990."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is June 14, 2005.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is June 14, 2005.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1990."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-038", "id": "fast-43-diverse-228-038-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is June 14, 2005. The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1990."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional instruction and all governing policy elements already covered by the unchanged questions object; only the birth year evidence changes (1990 vs 1985), which is a legitimate counterfactual since it flips the minor-status computation without contradicting other stated facts like blur status or checksums; both focus evidence spans are complete factual sentences, not rules or instructions; no gold answer, label, or rationale is embedded, leaving the age/minor determination to be inferred by the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is June 14, 2005. The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1990.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is June 14, 2005."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1990."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is June 14, 2005.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is June 14, 2005.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1990."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-038", "id": "fast-43-diverse-228-038-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is June 14, 2005. The person shown at 00:04:12 in the Harbor Day preservation video was born on June 14, 1985."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full archive policy via the unchanged questions object and only vary the subject's birth date, a factual case detail that shifts age relative to the fixed 2026 event date without altering governing rules; the two evidence spans are complete factual sentences (subject birth date and event date) rather than policy text; the counterfactual coherently reclassifies the individual as an adult rather than a minor without contradicting other unchanged facts; and neither context contains gold labels, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist completes review of the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All four required preservation files are playable and uncorrupted, and each carries a checksum matching the manifest. The authoritative metadata record supplies title, creator, event date, rights holder, and recorded file formats for every file, and the package is reliably identified as belonging to Harbor Day. Rights on file permit archival preservation. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d Reviewing the footage, the archivist notes that every person visible in the video other than the one at 00:04:12 has been blurred, while that individual remains clearly unblurred and identifiable. The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 2010. A usable public-access derivative of the package exists. The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 2010."}, {"path": [], "text": "The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 2010.", "negative_left": "The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 1990.", "negative_right": "The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026.", "right": "The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-052", "id": "fast-43-diverse-228-052-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist completes review of the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All four required preservation files are playable and uncorrupted, and each carries a checksum matching the manifest. The authoritative metadata record supplies title, creator, event date, rights holder, and recorded file formats for every file, and the package is reliably identified as belonging to Harbor Day. Rights on file permit archival preservation. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Reviewing the footage, the archivist notes that every person visible in the video other than the one at 00:04:12 has been blurred, while that individual remains clearly unblurred and identifiable. The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 2010. A usable public-access derivative of the package exists. The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full archive policy via the unchanged questions object and only vary the subject's birth date, a factual case detail that shifts age relative to the fixed 2026 event date without altering governing rules; the two evidence spans are complete factual sentences (subject birth date and event date) rather than policy text; the counterfactual coherently reclassifies the individual as an adult rather than a minor without contradicting other unchanged facts; and neither context contains gold labels, rule tables, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist completes review of the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All four required preservation files are playable and uncorrupted, and each carries a checksum matching the manifest. The authoritative metadata record supplies title, creator, event date, rights holder, and recorded file formats for every file, and the package is reliably identified as belonging to Harbor Day. Rights on file permit archival preservation. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d Reviewing the footage, the archivist notes that every person visible in the video other than the one at 00:04:12 has been blurred, while that individual remains clearly unblurred and identifiable. The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 2010. A usable public-access derivative of the package exists. The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 2010."}, {"path": [], "text": "The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 2010.", "negative_left": "The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 1990.", "negative_right": "The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026.", "right": "The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-052", "id": "fast-43-diverse-228-052-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist completes review of the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All four required preservation files are playable and uncorrupted, and each carries a checksum matching the manifest. The authoritative metadata record supplies title, creator, event date, rights holder, and recorded file formats for every file, and the package is reliably identified as belonging to Harbor Day. Rights on file permit archival preservation. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” Reviewing the footage, the archivist notes that every person visible in the video other than the one at 00:04:12 has been blurred, while that individual remains clearly unblurred and identifiable. The person shown at 00:04:12 in the Harbor Day preservation video is recorded in the archive's subject log as Individual S-7, whose birth date is listed as June 3, 1990. A usable public-access derivative of the package exists. The Harbor Day event date, as documented in the authoritative metadata record, is June 3, 2026."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the donor's conditional policy and all case metadata unchanged, only altering Subject H's derived age from 15 to 22, which is a coherent factual substitution not contradicted elsewhere; the two evidence spans are plain factual sentences about identification and age, and neither context states or implies the final ingestion decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest's checksums match, and title, creator, event date, rights holder, and file formats are recorded and confirmed authoritative. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d At 00:04:12, the video clearly shows one unblurred individual; every other person appearing in the footage is blurred. The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file. Subject H's birth record shows a birth date that makes Subject H 15 years old on the Harbor Day event date. All preservation files are playable, uncorrupted, and reliably identified as belonging to this package, with checksums attached and matching the manifest. Rights documentation permits archival preservation, and the donor's access instruction is on file as authoritative. A usable public-access derivative exists, though it has not yet been screened for blurring compliance.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file."}, {"path": [], "text": "Subject H's birth record shows a birth date that makes Subject H 15 years old on the Harbor Day event date."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file.", "negative_left": "The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file.", "negative_right": "Subject H's birth record shows a birth date that makes Subject H 22 years old on the Harbor Day event date.", "right": "Subject H's birth record shows a birth date that makes Subject H 15 years old on the Harbor Day event date."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-064", "id": "fast-43-diverse-228-064-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest's checksums match, and title, creator, event date, rights holder, and file formats are recorded and confirmed authoritative. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” At 00:04:12, the video clearly shows one unblurred individual; every other person appearing in the footage is blurred. The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file. Subject H's birth record shows a birth date that makes Subject H 15 years old on the Harbor Day event date. All preservation files are playable, uncorrupted, and reliably identified as belonging to this package, with checksums attached and matching the manifest. Rights documentation permits archival preservation, and the donor's access instruction is on file as authoritative. A usable public-access derivative exists, though it has not yet been screened for blurring compliance."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the donor's conditional policy and all case metadata unchanged, only altering Subject H's derived age from 15 to 22, which is a coherent factual substitution not contradicted elsewhere; the two evidence spans are plain factual sentences about identification and age, and neither context states or implies the final ingestion decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest's checksums match, and title, creator, event date, rights holder, and file formats are recorded and confirmed authoritative. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d At 00:04:12, the video clearly shows one unblurred individual; every other person appearing in the footage is blurred. The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file. Subject H's birth record shows a birth date that makes Subject H 15 years old on the Harbor Day event date. All preservation files are playable, uncorrupted, and reliably identified as belonging to this package, with checksums attached and matching the manifest. Rights documentation permits archival preservation, and the donor's access instruction is on file as authoritative. A usable public-access derivative exists, though it has not yet been screened for blurring compliance.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file."}, {"path": [], "text": "Subject H's birth record shows a birth date that makes Subject H 15 years old on the Harbor Day event date."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file.", "negative_left": "The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file.", "negative_right": "Subject H's birth record shows a birth date that makes Subject H 22 years old on the Harbor Day event date.", "right": "Subject H's birth record shows a birth date that makes Subject H 15 years old on the Harbor Day event date."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-064", "id": "fast-43-diverse-228-064-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest's checksums match, and title, creator, event date, rights holder, and file formats are recorded and confirmed authoritative. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” At 00:04:12, the video clearly shows one unblurred individual; every other person appearing in the footage is blurred. The person shown at 00:04:12 in the Harbor Day preservation video is identified in the case file as Subject H, whose birth record is on file. Subject H's birth record shows a birth date that makes Subject H 22 years old on the Harbor Day event date. All preservation files are playable, uncorrupted, and reliably identified as belonging to this package, with checksums attached and matching the manifest. Rights documentation permits archival preservation, and the donor's access instruction is on file as authoritative. A usable public-access derivative exists, though it has not yet been screened for blurring compliance."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional policy verbatim and identical package facts, only the birth year varies, changing the person's age relative to the fixed 2010 event date in a coherent, non-contradictory way; evidence spans are two complete factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable, uncorrupted, and carry checksums matching the manifest. The authoritative metadata record supplies the title, creator, rights holder, and file formats, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder, awaiting final review before release. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1995.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1995."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1995."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-065", "id": "fast-43-diverse-228-065-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable, uncorrupted, and carry checksums matching the manifest. The authoritative metadata record supplies the title, creator, rights holder, and file formats, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder, awaiting final review before release. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1995."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional policy verbatim and identical package facts, only the birth year varies, changing the person's age relative to the fixed 2010 event date in a coherent, non-contradictory way; evidence spans are two complete factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable, uncorrupted, and carry checksums matching the manifest. The authoritative metadata record supplies the title, creator, rights holder, and file formats, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder, awaiting final review before release. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1995.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1995."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1995."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-065", "id": "fast-43-diverse-228-065-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable, uncorrupted, and carry checksums matching the manifest. The authoritative metadata record supplies the title, creator, rights holder, and file formats, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder, awaiting final review before release. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (donor's conditional rule, rights, checksum, derivative facts) unchanged, with only the birth year altered from 1997 to 1985, shifting the person's age relative to the fixed 2012 event date without contradicting any other stated fact; the two evidence sentences are complete factual statements, not policy or rule text; and no gold label, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 9, 2012. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1997.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2012."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1997."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2012.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2012.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1997."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-072", "id": "fast-43-diverse-228-072-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 9, 2012. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1997."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (donor's conditional rule, rights, checksum, derivative facts) unchanged, with only the birth year altered from 1997 to 1985, shifting the person's age relative to the fixed 2012 event date without contradicting any other stated fact; the two evidence sentences are complete factual statements, not policy or rule text; and no gold label, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 9, 2012. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1997.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2012."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1997."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2012.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2012.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1997."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-072", "id": "fast-43-diverse-228-072-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 9, 2012. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 9, 1985."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (donor instruction, rights, checksums) and the identical question object; the two evidence sentences are plain factual statements about the event date and the depicted person's birth date; the counterfactual changes only the birth year (1994→1988), which coherently shifts the person's age relative to the unchanged 2010 event date without contradicting any other stated fact; neither context reveals or hints at a specific decision label, rule ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and the manifest's checksums are present and match. The authoritative metadata record supplies the title, creator, rights holder, and event date, and records the format of every file. The package is reliably identified as Harbor Day material, and rights documentation on file permits archival preservation. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This instruction is treated as authoritative for the package. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person seen elsewhere in the footage is blurred. A usable public-access derivative already sits in the collection folder, ready for release pending the restriction determination. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-075", "id": "fast-43-diverse-228-075-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and the manifest's checksums are present and match. The authoritative metadata record supplies the title, creator, rights holder, and event date, and records the format of every file. The package is reliably identified as Harbor Day material, and rights documentation on file permits archival preservation. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This instruction is treated as authoritative for the package. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person seen elsewhere in the footage is blurred. A usable public-access derivative already sits in the collection folder, ready for release pending the restriction determination. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (donor instruction, rights, checksums) and the identical question object; the two evidence sentences are plain factual statements about the event date and the depicted person's birth date; the counterfactual changes only the birth year (1994→1988), which coherently shifts the person's age relative to the unchanged 2010 event date without contradicting any other stated fact; neither context reveals or hints at a specific decision label, rule ID, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and the manifest's checksums are present and match. The authoritative metadata record supplies the title, creator, rights holder, and event date, and records the format of every file. The package is reliably identified as Harbor Day material, and rights documentation on file permits archival preservation. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This instruction is treated as authoritative for the package. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person seen elsewhere in the footage is blurred. A usable public-access derivative already sits in the collection folder, ready for release pending the restriction determination. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-075", "id": "fast-43-diverse-228-075-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and the manifest's checksums are present and match. The authoritative metadata record supplies the title, creator, rights holder, and event date, and records the format of every file. The package is reliably identified as Harbor Day material, and rights documentation on file permits archival preservation. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This instruction is treated as authoritative for the package. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person seen elsewhere in the footage is blurred. A usable public-access derivative already sits in the collection folder, ready for release pending the restriction determination. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor instruction, checksum/rights/derivative facts, and the two focus sentences as complete factual statements; the counterfactual only changes the birth year (1992→1988), altering the person's age at the 2008 event without contradicting any other stated fact, and neither context reveals a label, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and every checksum is present and matches the manifest. The authoritative metadata record supplies the title, creator, rights holder, and recorded file formats, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This instruction is treated as authoritative for access decisions. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person shown elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already sits in the collection folder awaiting release approval. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-080", "id": "fast-43-diverse-228-080-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and every checksum is present and matches the manifest. The authoritative metadata record supplies the title, creator, rights holder, and recorded file formats, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This instruction is treated as authoritative for access decisions. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person shown elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already sits in the collection folder awaiting release approval. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor instruction, checksum/rights/derivative facts, and the two focus sentences as complete factual statements; the counterfactual only changes the birth year (1992→1988), altering the person's age at the 2008 event without contradicting any other stated fact, and neither context reveals a label, rule table, or instruction to the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and every checksum is present and matches the manifest. The authoritative metadata record supplies the title, creator, rights holder, and recorded file formats, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This instruction is treated as authoritative for access decisions. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person shown elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already sits in the collection folder awaiting release approval. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-080", "id": "fast-43-diverse-228-080-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and every checksum is present and matches the manifest. The authoritative metadata record supplies the title, creator, rights holder, and recorded file formats, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This instruction is treated as authoritative for access decisions. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person shown elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already sits in the collection folder awaiting release approval. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's authoritative access policy and all preservation/rights facts unchanged, only the birth year varies, shifting the person's age from minor (14) to adult (25) without contradicting other stated facts; the two focus evidence spans are complete factual sentences about event date and birth date, not policy text; no gold label, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1996.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1996."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1996."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-082", "id": "fast-43-diverse-228-082-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1996."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's authoritative access policy and all preservation/rights facts unchanged, only the birth year varies, shifting the person's age from minor (14) to adult (25) without contradicting other stated facts; the two focus evidence spans are complete factual sentences about event date and birth date, not policy text; no gold label, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1996.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1996."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1996."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-082", "id": "fast-43-diverse-228-082-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1985."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the donor's conditional access instruction and all preservation/rights facts unchanged, only the birth year of the person at 00:04:12 differs, shifting their age relative to the fixed 2010 event date without altering any other governing fact; both focus_evidence spans are complete factual sentences with no embedded rules, IDs, or output instructions; the counterfactual (1985 birth) remains internally coherent with the unchanged blurred/unblurred and rights facts, merely changing whether the identifiable person is a minor, and no gold answer or rationale is leaked in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable, uncorrupted, and their checksums match the manifest entries. The authoritative metadata record supplies the title, creator, rights holder, and recorded file formats for every item, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the shared drive folder. According to the authoritative metadata record, the Harbor Day event date is September 9, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-088", "id": "fast-43-diverse-228-088-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable, uncorrupted, and their checksums match the manifest entries. The authoritative metadata record supplies the title, creator, rights holder, and recorded file formats for every item, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the shared drive folder. According to the authoritative metadata record, the Harbor Day event date is September 9, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the donor's conditional access instruction and all preservation/rights facts unchanged, only the birth year of the person at 00:04:12 differs, shifting their age relative to the fixed 2010 event date without altering any other governing fact; both focus_evidence spans are complete factual sentences with no embedded rules, IDs, or output instructions; the counterfactual (1985 birth) remains internally coherent with the unchanged blurred/unblurred and rights facts, merely changing whether the identifiable person is a minor, and no gold answer or rationale is leaked in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable, uncorrupted, and their checksums match the manifest entries. The authoritative metadata record supplies the title, creator, rights holder, and recorded file formats for every item, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the shared drive folder. According to the authoritative metadata record, the Harbor Day event date is September 9, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 9, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-088", "id": "fast-43-diverse-228-088-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable, uncorrupted, and their checksums match the manifest entries. The authoritative metadata record supplies the title, creator, rights holder, and recorded file formats for every item, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the shared drive folder. According to the authoritative metadata record, the Harbor Day event date is September 9, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1985."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (checksum, playability, rights, donor's conditional instruction) and the same question bindings; the counterfactual changes only the birth year, shifting the person's age from minor (16) to adult (20) in a coherent, single-fact way; both focus sentences are plain factual statements with no rule tables, IDs, or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and every checksum is present and matches the manifest. The authoritative metadata record supplies the title, creator, rights holder, and file formats for each file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This instruction is treated as authoritative for access decisions. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-089", "id": "fast-43-diverse-228-089-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and every checksum is present and matches the manifest. The authoritative metadata record supplies the title, creator, rights holder, and file formats for each file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This instruction is treated as authoritative for access decisions. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing policy (checksum, playability, rights, donor's conditional instruction) and the same question bindings; the counterfactual changes only the birth year, shifting the person's age from minor (16) to adult (20) in a coherent, single-fact way; both focus sentences are plain factual statements with no rule tables, IDs, or answer leakage.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and every checksum is present and matches the manifest. The authoritative metadata record supplies the title, creator, rights holder, and file formats for each file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This instruction is treated as authoritative for access decisions. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1992."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-089", "id": "fast-43-diverse-228-089-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and every checksum is present and matches the manifest. The authoritative metadata record supplies the title, creator, rights holder, and file formats for each file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This instruction is treated as authoritative for access decisions. The person visible at 00:04:12 is clearly identifiable and appears unblurred, while every other person elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy, checksum, rights, and instruction details while the questions object supplies governing criteria unchanged; the two focus sentences are plain factual statements about event date and birth date; the counterfactual only swaps the birth year (1994→1988), changing minor status coherently without contradicting other facts; no gold answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-094", "id": "fast-43-diverse-228-094-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical policy, checksum, rights, and instruction details while the questions object supplies governing criteria unchanged; the two focus sentences are plain factual statements about event date and birth date; the counterfactual only swaps the birth year (1994→1988), changing minor status coherently without contradicting other facts; no gold answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1994."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-094", "id": "fast-43-diverse-228-094-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event took place on September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1988."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional access policy and all package facts, with only the birthdate changed in the counterfactual, which coherently shifts the minor determination without contradicting other stated facts; the two evidence spans are complete factual sentences with no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 5, 1994.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 5, 1994."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 1, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 5, 1994."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-104", "id": "fast-43-diverse-228-104-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 5, 1994."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional access policy and all package facts, with only the birthdate changed in the counterfactual, which coherently shifts the minor determination without contradicting other stated facts; the two evidence spans are complete factual sentences with no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 5, 1994.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 5, 1994."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 1, 1988.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 5, 1994."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-104", "id": "fast-43-diverse-228-104-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 1, 1988."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional access instruction and all preservation/rights facts, preserving policy and question bindings; the two focus sentences are plain factual statements (event date and birth date) rather than policy text; the counterfactual only changes the birth year (1993→1985), a single coherent factual edit with no contradictory duplicate facts; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All manifest checksums are present and match their entries, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-111", "id": "fast-43-diverse-228-111-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All manifest checksums are present and match their entries, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional access instruction and all preservation/rights facts, preserving policy and question bindings; the two focus sentences are plain factual statements (event date and birth date) rather than policy text; the counterfactual only changes the birth year (1993→1985), a single coherent factual edit with no contradictory duplicate facts; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All manifest checksums are present and match their entries, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1993."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-111", "id": "fast-43-diverse-228-111-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All manifest checksums are present and match their entries, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 10, 1985."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all original policy elements (checksums, playable files, rights documentation, metadata, donor's conditional access statement) alongside the unchanged questions object, so policy and bindings are preserved; the two focus sentences are plain factual statements (event date and birth date) rather than policy or instructions; the counterfactual's single change (birth year 1985 vs 1993) coherently shifts the person's age at the 2008 event from minor to adult without contradicting any other stated fact; neither context contains a gold label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and the manifest checksums are present and match. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1993.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1993."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1993."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-112", "id": "fast-43-diverse-228-112-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and the manifest checksums are present and match. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1993."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all original policy elements (checksums, playable files, rights documentation, metadata, donor's conditional access statement) alongside the unchanged questions object, so policy and bindings are preserved; the two focus sentences are plain factual statements (event date and birth date) rather than policy or instructions; the counterfactual's single change (birth year 1985 vs 1993) coherently shifts the person's age at the 2008 event from minor to adult without contradicting any other stated fact; neither context contains a gold label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and the manifest checksums are present and match. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1993.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1993."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2008.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1993."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-112", "id": "fast-43-diverse-228-112-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. All required preservation files are playable and uncorrupted, and the manifest checksums are present and match. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation, and a usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2008. The person shown at 00:04:12 in the Harbor Day preservation video was born on September 3, 1985."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional access policy and all preservation/rights facts unchanged from the original state, with the question object presumably passed verbatim; the two evidence sentences are plain factual statements (event date, birth date) rather than policy text; changing the birth year from 1995 to 1985 shifts the person's age from a minor to an adult at the 2010 event, a coherent single-fact alteration that does not contradict other unchanged details; neither context contains a rule table, gold answer, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1995.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1995."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1995."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-114", "id": "fast-43-diverse-228-114-base", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1995."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "media-03", "split": "train", "variant": "base"} {"domain": "media", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the donor's conditional access policy and all preservation/rights facts unchanged from the original state, with the question object presumably passed verbatim; the two evidence sentences are plain factual statements (event date, birth date) rather than policy text; changing the birth year from 1995 to 1985 shifts the person's age from a minor to an adult at the 2010 event, a coherent single-fact alteration that does not contradict other unchanged details; neither context contains a rule table, gold answer, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "full_context_fact_states": {"base": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "supported", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "counterfactual": {"a_access_instruction_authoritative": "supported", "a_checksums_match": "supported", "a_checksums_present": "supported", "a_creator_authoritative": "supported", "a_event_date_authoritative": "supported", "a_files_playable": "supported", "a_files_uncorrupted": "supported", "a_focus_under_18": "refuted", "a_formats_recorded": "supported", "a_other_people_blurred": "supported", "a_package_identified": "supported", "a_person_identifiable": "supported", "a_person_unblurred": "supported", "a_public_derivative_usable": "supported", "a_rights_holder_authoritative": "supported", "a_rights_permit_preservation": "supported", "a_title_authoritative": "supported"}, "remove_left": {"a_focus_under_18": "unknown"}, "remove_right": {"a_focus_under_18": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a_focus_under_18": "unknown"}, "negative_pair": {"a_focus_under_18": "refuted"}, "negative_sentence": {"a_focus_under_18": "unknown"}, "positive_pair": {"a_focus_under_18": "supported"}, "right": {"a_focus_under_18": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; universally quantified file/person properties remain atomic and no atom encodes a final readiness classification. The focus is the factual age relationship. The base and counter assignments are jointly realizable with only that age fact changing; an unblurred preservation-video subject can coexist with a separate usable public derivative. Policy evidence correctly preserves the donor’s conditional-access rule from the original state, while the remaining governing criteria are already retained in the questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction satisfies the preservation, checksum, core-metadata, rights, access-instruction, and package-identification requirements, while the supported under-18, identifiable, and unblurred facts trigger the donor’s staff-only restriction. A usable public derivative does not negate that restriction, so level 2 is entailed.", "rule_index": 0, "sound": true}, {"reason": "The focus person is explicitly not under 18, and every other depicted person is blurred, excluding any unblurred identifiable minor. The preservation, checksum, core-metadata, rights, access-instruction, identification, and usable-derivative requirements are satisfied, so no conditional restriction is triggered and level 3 is entailed.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a_focus_under_18", "statement": "The age on the Harbor Day event date of the person shown at 00:04:12 in the Harbor Day preservation video is less than 18 years."}, {"id": "a_person_identifiable", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is identifiable."}, {"id": "a_person_unblurred", "statement": "The person shown at 00:04:12 in the Harbor Day preservation video is unblurred."}, {"id": "a_other_people_blurred", "statement": "Every person other than the person shown at 00:04:12 who appears in the Harbor Day preservation video is blurred."}, {"id": "a_files_playable", "statement": "Every required preservation file in the Harbor Day package is playable."}, {"id": "a_files_uncorrupted", "statement": "Every required preservation file in the Harbor Day package is uncorrupted."}, {"id": "a_checksums_present", "statement": "Every required preservation file in the Harbor Day package has an attached checksum."}, {"id": "a_checksums_match", "statement": "Every required preservation file in the Harbor Day package has a checksum matching its manifest checksum."}, {"id": "a_title_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a title value."}, {"id": "a_creator_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a creator value."}, {"id": "a_event_date_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies an event-date value."}, {"id": "a_rights_holder_authoritative", "statement": "The authoritative metadata record for the Harbor Day package supplies a rights-holder value."}, {"id": "a_access_instruction_authoritative", "statement": "The authoritative donor record for the Harbor Day package supplies the applicable access instruction."}, {"id": "a_formats_recorded", "statement": "Every file in the Harbor Day package has its file format recorded in the authoritative metadata record."}, {"id": "a_rights_permit_preservation", "statement": "The rights governing the Harbor Day package permit archival preservation."}, {"id": "a_package_identified", "statement": "Every required preservation file is reliably identified as belonging to the Harbor Day package."}, {"id": "a_public_derivative_usable", "statement": "The Harbor Day package has a usable public-access derivative."}], "base_state_json": "\"A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, \\u201cIf any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.\\u201d This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1995.\"", "base_states": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "counter_states": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}], "focus_atom": "a_focus_under_18", "focus_evidence": [{"path": [], "text": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010."}, {"path": [], "text": "The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1995."}], "policy_evidence": [{"path": [], "text": "The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.”"}], "rules": [{"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The identifiable, unblurred person at 00:04:12 is under 18, so the donor’s staff-only restriction is triggered; the usable derivative does not remove that restriction.", "target": "2", "when": [{"atom_id": "a_focus_under_18", "state": "supported"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}, {"justification": "All preservation, checksum, authoritative-metadata, rights, access-instruction, and identification requirements are satisfied, excluding levels 0 and 1. The unblurred person at 00:04:12 is not under 18, every other depicted person is blurred, and therefore no identifiable minor is unblurred. The donor’s conditional restriction is not triggered, and a usable public-access derivative is available.", "target": "3", "when": [{"atom_id": "a_focus_under_18", "state": "refuted"}, {"atom_id": "a_person_identifiable", "state": "supported"}, {"atom_id": "a_person_unblurred", "state": "supported"}, {"atom_id": "a_other_people_blurred", "state": "supported"}, {"atom_id": "a_files_playable", "state": "supported"}, {"atom_id": "a_files_uncorrupted", "state": "supported"}, {"atom_id": "a_checksums_present", "state": "supported"}, {"atom_id": "a_checksums_match", "state": "supported"}, {"atom_id": "a_title_authoritative", "state": "supported"}, {"atom_id": "a_creator_authoritative", "state": "supported"}, {"atom_id": "a_event_date_authoritative", "state": "supported"}, {"atom_id": "a_rights_holder_authoritative", "state": "supported"}, {"atom_id": "a_access_instruction_authoritative", "state": "supported"}, {"atom_id": "a_formats_recorded", "state": "supported"}, {"atom_id": "a_rights_permit_preservation", "state": "supported"}, {"atom_id": "a_package_identified", "state": "supported"}, {"atom_id": "a_public_derivative_usable", "state": "supported"}]}]}, "verified_pair": {"left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_left": "According to the authoritative metadata record, the Harbor Day event date is September 3, 2010.", "negative_right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1985.", "right": "The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1995."}, "verifier_independent_model": false}, "family": "fast-43-diverse-228-114", "id": "fast-43-diverse-228-114-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Reject: Do not ingest because rights prohibit preservation, required preservation files are corrupt or unreadable, or the package cannot be reliably identified.", "1 — Hold for remediation: Do not ingest yet because a required preservation file, checksum, title, creator, event date, rights holder, or applicable access instruction is missing or unresolved.", "2 — Ready for restricted ingest: Ingest into staff-only storage when preservation files are playable, checksums and core metadata are complete, rights permit preservation, and a stated restriction is triggered; a public derivative may still be absent.", "3 — Ready for public ingest: Ingest and publish when level-2 preservation and metadata requirements are met, no conditional restriction is triggered, and a usable public-access derivative is available."], "instructions": "Assign the package an ingestion-readiness level using the archive policy below. Apply the donor’s conditional access intent to the supplied file evidence.", "type": "score"}}, "state": "A digital archivist reviews the fictional Harbor Day package: an MP4 preservation video, separate WAV audio, a JPG poster, and a PRPROJ editing file. The manifest checksums are all present and match, and the required preservation files are playable and uncorrupted. The authoritative metadata record supplies the title, creator, rights holder, and file formats for every file, and the package is reliably identified as Harbor Day material. The donor states, “If any identifiable minor is unblurred, keep the package staff-only; otherwise make it public.” This access instruction is treated as authoritative for the package. The person at 00:04:12 is clearly identifiable and appears unblurred, while every other person visible elsewhere in the footage is blurred. Rights documentation on file permits archival preservation. A usable public-access derivative already exists in the collection folder. According to the authoritative metadata record, the Harbor Day event date is September 3, 2010. The person shown at 00:04:12 in the Harbor Day preservation video was born on October 20, 1985."}, "method": "c2d", "provenance": {"source_id": "diverse-228", "source_is_synthetic": true, "source_sha256": "5b3717fc941826afd29ac455957e47e680a85477b13ec8a35f2a1fd4734558f0", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "media-03", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the ecologist's conditional rule and sensor/tray bindings, the counterfactual only changes B-4's reading from 22% (within 18-25% range, passing) to 30% (outside range, failing), which is a coherent single-fact flip without duplicating or contradicting other measurements, and neither context states a pass/fail conclusion or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four sensors governed by this condition were B-1, B-2, B-3, and B-4. At 08:00, technicians logged calibration readings for each sensor. B-1, B-2, and B-3 all passed calibration cleanly. The 08:00 calibration log records sensor B-4's moisture reading at 22% against the required threshold range. Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration. Technician Leon withheld water from tray B for four days and watered tray A daily, and the data quality reviewer confirmed the growth records were complete for both trays.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "The 08:00 calibration log records sensor B-4's moisture reading at 22% against the required threshold range."}, {"path": [], "text": "Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "The 08:00 calibration log records sensor B-4's moisture reading at 22% against the required threshold range.", "negative_left": "The 08:00 calibration log records sensor B-4's moisture reading at 30% against the required threshold range.", "negative_right": "Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration.", "right": "Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-002", "id": "fast-43-diverse-243-002-base", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four sensors governed by this condition were B-1, B-2, B-3, and B-4. At 08:00, technicians logged calibration readings for each sensor. B-1, B-2, and B-3 all passed calibration cleanly. The 08:00 calibration log records sensor B-4's moisture reading at 22% against the required threshold range. Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration. Technician Leon withheld water from tray B for four days and watered tray A daily, and the data quality reviewer confirmed the growth records were complete for both trays."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the ecologist's conditional rule and sensor/tray bindings, the counterfactual only changes B-4's reading from 22% (within 18-25% range, passing) to 30% (outside range, failing), which is a coherent single-fact flip without duplicating or contradicting other measurements, and neither context states a pass/fail conclusion or answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four sensors governed by this condition were B-1, B-2, B-3, and B-4. At 08:00, technicians logged calibration readings for each sensor. B-1, B-2, and B-3 all passed calibration cleanly. The 08:00 calibration log records sensor B-4's moisture reading at 22% against the required threshold range. Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration. Technician Leon withheld water from tray B for four days and watered tray A daily, and the data quality reviewer confirmed the growth records were complete for both trays.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "The 08:00 calibration log records sensor B-4's moisture reading at 22% against the required threshold range."}, {"path": [], "text": "Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "The 08:00 calibration log records sensor B-4's moisture reading at 22% against the required threshold range.", "negative_left": "The 08:00 calibration log records sensor B-4's moisture reading at 30% against the required threshold range.", "negative_right": "Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration.", "right": "Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-002", "id": "fast-43-diverse-243-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four sensors governed by this condition were B-1, B-2, B-3, and B-4. At 08:00, technicians logged calibration readings for each sensor. B-1, B-2, and B-3 all passed calibration cleanly. The 08:00 calibration log records sensor B-4's moisture reading at 30% against the required threshold range. Mira's 08:00 pre-watering calibration protocol states any sensor reading between 18% and 25% moisture passes calibration. Technician Leon withheld water from tray B for four days and watered tray A daily, and the data quality reviewer confirmed the growth records were complete for both trays."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the calibration threshold and conditional policy from the original question intact; the counterfactual changes only B-4's reading to 9 percent, making it fail (below 15 percent), which is internally consistent since B-4 is no longer asserted to pass; evidence spans are two complete factual sentences about the threshold and B-4's reading; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The complete set of soil-moisture sensors governed by this condition was exactly B-1, B-2, B-3, and B-4. Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass. Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration. Sensor B-4 recorded a moisture reading of 22 percent during Mira's 08:00 pre-watering calibration. Greenhouse technician Leon withheld water from tray B for four days and watered tray A daily. The data quality reviewer confirmed that the growth records were complete for all mornings of the trial. The reviewer must now determine whether tray B's drought treatment reflects the documented assignment intent.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass."}, {"path": [], "text": "Sensor B-4 recorded a moisture reading of 22 percent during Mira's 08:00 pre-watering calibration."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass.", "negative_left": "Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass.", "negative_right": "Sensor B-4 recorded a moisture reading of 9 percent during Mira's 08:00 pre-watering calibration.", "right": "Sensor B-4 recorded a moisture reading of 22 percent during Mira's 08:00 pre-watering calibration."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-010", "id": "fast-43-diverse-243-010-base", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The complete set of soil-moisture sensors governed by this condition was exactly B-1, B-2, B-3, and B-4. Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass. Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration. Sensor B-4 recorded a moisture reading of 22 percent during Mira's 08:00 pre-watering calibration. Greenhouse technician Leon withheld water from tray B for four days and watered tray A daily. The data quality reviewer confirmed that the growth records were complete for all mornings of the trial. The reviewer must now determine whether tray B's drought treatment reflects the documented assignment intent."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the calibration threshold and conditional policy from the original question intact; the counterfactual changes only B-4's reading to 9 percent, making it fail (below 15 percent), which is internally consistent since B-4 is no longer asserted to pass; evidence spans are two complete factual sentences about the threshold and B-4's reading; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The complete set of soil-moisture sensors governed by this condition was exactly B-1, B-2, B-3, and B-4. Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass. Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration. Sensor B-4 recorded a moisture reading of 22 percent during Mira's 08:00 pre-watering calibration. Greenhouse technician Leon withheld water from tray B for four days and watered tray A daily. The data quality reviewer confirmed that the growth records were complete for all mornings of the trial. The reviewer must now determine whether tray B's drought treatment reflects the documented assignment intent.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass."}, {"path": [], "text": "Sensor B-4 recorded a moisture reading of 22 percent during Mira's 08:00 pre-watering calibration."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass.", "negative_left": "Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass.", "negative_right": "Sensor B-4 recorded a moisture reading of 9 percent during Mira's 08:00 pre-watering calibration.", "right": "Sensor B-4 recorded a moisture reading of 22 percent during Mira's 08:00 pre-watering calibration."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-010", "id": "fast-43-diverse-243-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The complete set of soil-moisture sensors governed by this condition was exactly B-1, B-2, B-3, and B-4. Mira's 08:00 pre-watering calibration required a moisture reading of at least 15 percent for a sensor to pass. Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration. Sensor B-4 recorded a moisture reading of 9 percent during Mira's 08:00 pre-watering calibration. Greenhouse technician Leon withheld water from tray B for four days and watered tray A daily. The data quality reviewer confirmed that the growth records were complete for all mornings of the trial. The reviewer must now determine whether tray B's drought treatment reflects the documented assignment intent."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's original conditional rule and the question verbatim, only varying the B-4 reading (20% vs 27%) as a changeable observation; the calibration-range statement is a factual description of the protocol rather than a new decision rule, both evidence sentences are complete factual statements, the counterfactual's 27% reading is coherent with B-4 failing calibration, and neither context reveals the intended answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four governed sensors are B-1, B-2, B-3, and B-4. At 08:00, B-1, B-2, and B-3 each passed the calibration. Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass. At 08:00, sensor B-4 recorded a moisture reading of 20%. Greenhouse technician Leon withheld water from tray B for four days and watered tray A daily. The reviewer confirmed the morning growth records were complete and must now determine whether tray B's drought treatment reflects the documented assignment intent.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass."}, {"path": [], "text": "At 08:00, sensor B-4 recorded a moisture reading of 20%."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass.", "negative_left": "Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass.", "negative_right": "At 08:00, sensor B-4 recorded a moisture reading of 27%.", "right": "At 08:00, sensor B-4 recorded a moisture reading of 20%."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-017", "id": "fast-43-diverse-243-017-base", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four governed sensors are B-1, B-2, B-3, and B-4. At 08:00, B-1, B-2, and B-3 each passed the calibration. Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass. At 08:00, sensor B-4 recorded a moisture reading of 20%. Greenhouse technician Leon withheld water from tray B for four days and watered tray A daily. The reviewer confirmed the morning growth records were complete and must now determine whether tray B's drought treatment reflects the documented assignment intent."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep Mira's original conditional rule and the question verbatim, only varying the B-4 reading (20% vs 27%) as a changeable observation; the calibration-range statement is a factual description of the protocol rather than a new decision rule, both evidence sentences are complete factual statements, the counterfactual's 27% reading is coherent with B-4 failing calibration, and neither context reveals the intended answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four governed sensors are B-1, B-2, B-3, and B-4. At 08:00, B-1, B-2, and B-3 each passed the calibration. Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass. At 08:00, sensor B-4 recorded a moisture reading of 20%. Greenhouse technician Leon withheld water from tray B for four days and watered tray A daily. The reviewer confirmed the morning growth records were complete and must now determine whether tray B's drought treatment reflects the documented assignment intent.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass."}, {"path": [], "text": "At 08:00, sensor B-4 recorded a moisture reading of 20%."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass.", "negative_left": "Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass.", "negative_right": "At 08:00, sensor B-4 recorded a moisture reading of 27%.", "right": "At 08:00, sensor B-4 recorded a moisture reading of 20%."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-017", "id": "fast-43-diverse-243-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four governed sensors are B-1, B-2, B-3, and B-4. At 08:00, B-1, B-2, and B-3 each passed the calibration. Mira's 08:00 pre-watering calibration requires a sensor's moisture reading to fall between 18% and 22% to record a pass. At 08:00, sensor B-4 recorded a moisture reading of 27%. Greenhouse technician Leon withheld water from tray B for four days and watered tray A daily. The reviewer confirmed the morning growth records were complete and must now determine whether tray B's drought treatment reflects the documented assignment intent."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the conditional policy and question bindings unchanged except for B-4's deviation value, which coherently shifts from passing (2.1%) to failing (4.5%) against the stated 3.0% threshold; evidence spans are factual sensor/protocol sentences, not answers or rule tables.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” Mira's 08:00 pre-watering calibration governed exactly four sensors: B-1, B-2, B-3, and B-4. At 08:00, sensor B-1 recorded a deviation of 0.8%, B-2 recorded 1.2%, and B-3 recorded 0.5% from the reference standard, so B-1, B-2, and B-3 each passed the calibration. At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 2.1% from the reference standard. Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less. Technician Leon withheld water from tray B for four days and watered tray A daily. The data quality reviewer confirmed the growth records were complete and must now determine whether tray B's drought treatment reflects the documented assignment intent.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 2.1% from the reference standard."}, {"path": [], "text": "Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 2.1% from the reference standard.", "negative_left": "At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 4.5% from the reference standard.", "negative_right": "Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less.", "right": "Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-018", "id": "fast-43-diverse-243-018-base", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” Mira's 08:00 pre-watering calibration governed exactly four sensors: B-1, B-2, B-3, and B-4. At 08:00, sensor B-1 recorded a deviation of 0.8%, B-2 recorded 1.2%, and B-3 recorded 0.5% from the reference standard, so B-1, B-2, and B-3 each passed the calibration. At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 2.1% from the reference standard. Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less. Technician Leon withheld water from tray B for four days and watered tray A daily. The data quality reviewer confirmed the growth records were complete and must now determine whether tray B's drought treatment reflects the documented assignment intent."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the conditional policy and question bindings unchanged except for B-4's deviation value, which coherently shifts from passing (2.1%) to failing (4.5%) against the stated 3.0% threshold; evidence spans are factual sensor/protocol sentences, not answers or rule tables.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” Mira's 08:00 pre-watering calibration governed exactly four sensors: B-1, B-2, B-3, and B-4. At 08:00, sensor B-1 recorded a deviation of 0.8%, B-2 recorded 1.2%, and B-3 recorded 0.5% from the reference standard, so B-1, B-2, and B-3 each passed the calibration. At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 2.1% from the reference standard. Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less. Technician Leon withheld water from tray B for four days and watered tray A daily. The data quality reviewer confirmed the growth records were complete and must now determine whether tray B's drought treatment reflects the documented assignment intent.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 2.1% from the reference standard."}, {"path": [], "text": "Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 2.1% from the reference standard.", "negative_left": "At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 4.5% from the reference standard.", "negative_right": "Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less.", "right": "Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-018", "id": "fast-43-diverse-243-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” Mira's 08:00 pre-watering calibration governed exactly four sensors: B-1, B-2, B-3, and B-4. At 08:00, sensor B-1 recorded a deviation of 0.8%, B-2 recorded 1.2%, and B-3 recorded 0.5% from the reference standard, so B-1, B-2, and B-3 each passed the calibration. At Mira's 08:00 pre-watering calibration, sensor B-4 recorded a moisture-reading deviation of 4.5% from the reference standard. Mira's calibration protocol specifies that a sensor passes the 08:00 pre-watering calibration if its deviation from the reference standard is 3.0% or less. Technician Leon withheld water from tray B for four days and watered tray A daily. The data quality reviewer confirmed the growth records were complete and must now determine whether tray B's drought treatment reflects the documented assignment intent."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original conditional policy and sensor/tray/time bindings, differ only in B-4's calibration reading (1% vs 5%) relative to the stated 2% threshold, and the two evidence sentences are complete factual statements without embedding the answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: \\u201cIf all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.\\u201d The four sensors governed by this condition were B-1, B-2, B-3, and B-4. At 08:00, B-1, B-2, and B-3 each passed calibration. At Mira\\u2019s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass. At 08:00, sensor B-4\\u2019s calibration reading was 1 percent from the reference value. Technician Leon withheld water from tray B for four days and watered tray A daily. Heights were measured each morning, and the data quality reviewer confirmed that the growth records were complete. The reviewer is now reviewing whether the documented calibration condition was met for each sensor before any watering decision was executed.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "At Mira’s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass."}, {"path": [], "text": "At 08:00, sensor B-4’s calibration reading was 1 percent from the reference value."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "At Mira’s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass.", "negative_left": "At Mira’s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass.", "negative_right": "At 08:00, sensor B-4’s calibration reading was 5 percent from the reference value.", "right": "At 08:00, sensor B-4’s calibration reading was 1 percent from the reference value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-026", "id": "fast-43-diverse-243-026-base", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four sensors governed by this condition were B-1, B-2, B-3, and B-4. At 08:00, B-1, B-2, and B-3 each passed calibration. At Mira’s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass. At 08:00, sensor B-4’s calibration reading was 1 percent from the reference value. Technician Leon withheld water from tray B for four days and watered tray A daily. Heights were measured each morning, and the data quality reviewer confirmed that the growth records were complete. The reviewer is now reviewing whether the documented calibration condition was met for each sensor before any watering decision was executed."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-01", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original conditional policy and sensor/tray/time bindings, differ only in B-4's calibration reading (1% vs 5%) relative to the stated 2% threshold, and the two evidence sentences are complete factual statements without embedding the answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "full_context_fact_states": {"base": {"b4_result": "supported", "other_sensor_results": "supported", "sensor_scope": "supported"}, "counterfactual": {"b4_result": "refuted", "other_sensor_results": "supported", "sensor_scope": "supported"}, "remove_left": {"b4_result": "unknown"}, "remove_right": {"b4_result": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"b4_result": "unknown"}, "negative_pair": {"b4_result": "refuted"}, "negative_sentence": {"b4_result": "unknown"}, "positive_pair": {"b4_result": "supported"}, "right": {"b4_result": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the explicit-set universal in other_sensor_results remains atomic. The focus concerns the factual calibration result of B-4, not a policy conclusion. The base and counter assignments are jointly realizable while changing only B-4’s result. The policy evidence correctly preserves the substantive conditional assignment and fallback from the original state; instructions and criteria in questions need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Given that the governed set is exactly B-1 through B-4 and all four are supported as passing the required calibration, Mira’s antecedent is satisfied and tray B is protocol-intended for drought.", "rule_index": 0, "sound": true}, {"reason": "Given that B-4 belongs to the governed four-sensor set and is refuted as having passed, the requirement that all four pass is not satisfied. The explicit otherwise clause therefore assigns both trays to daily watering, making the false target sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "sensor_scope", "statement": "The complete set of soil-moisture sensors governed by Mira’s 08:00 pre-watering calibration condition was exactly B-1, B-2, B-3, and B-4."}, {"id": "other_sensor_results", "statement": "Each of B-1, B-2, and B-3 passed Mira’s 08:00 pre-watering calibration."}, {"id": "b4_result", "statement": "Sensor B-4 passed Mira’s 08:00 pre-watering calibration."}], "base_state_json": "\"Before the drought trial, plant ecologist Mira documented: \\u201cIf all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.\\u201d The four sensors governed by this condition were B-1, B-2, B-3, and B-4. At 08:00, B-1, B-2, and B-3 each passed calibration. At Mira\\u2019s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass. At 08:00, sensor B-4\\u2019s calibration reading was 1 percent from the reference value. Technician Leon withheld water from tray B for four days and watered tray A daily. Heights were measured each morning, and the data quality reviewer confirmed that the growth records were complete. The reviewer is now reviewing whether the documented calibration condition was met for each sensor before any watering decision was executed.\"", "base_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}], "counter_states": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}], "focus_atom": "b4_result", "focus_evidence": [{"path": [], "text": "At Mira’s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass."}, {"path": [], "text": "At 08:00, sensor B-4’s calibration reading was 1 percent from the reference value."}], "policy_evidence": [{"path": [], "text": "“If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.”"}], "rules": [{"justification": "The governed set contains exactly four sensors, and every member of that set passed the required 08:00 pre-watering calibration, so the documented drought-assignment condition for tray B was satisfied.", "target": "true", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "supported"}]}, {"justification": "B-4 is one of the four governed sensors and did not pass the required 08:00 pre-watering calibration, so at least one required condition failed and the daily-watering fallback applied.", "target": "false", "when": [{"atom_id": "sensor_scope", "state": "supported"}, {"atom_id": "other_sensor_results", "state": "supported"}, {"atom_id": "b4_result", "state": "refuted"}]}]}, "verified_pair": {"left": "At Mira’s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass.", "negative_left": "At Mira’s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass.", "negative_right": "At 08:00, sensor B-4’s calibration reading was 5 percent from the reference value.", "right": "At 08:00, sensor B-4’s calibration reading was 1 percent from the reference value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-243-026", "id": "fast-43-diverse-243-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Tray B was not protocol-intended for drought because at least one required condition failed, activating the daily-watering fallback.", "true": "Tray B was protocol-intended for drought because every required condition in the ecologist’s conditional assignment was satisfied."}, "instructions": "Verification task: Decide whether tray B was protocol-intended for drought treatment. Answer yes only if the stated calibration condition for drought assignment was satisfied before watering began. Complete growth measurements do not substitute for satisfying the assignment condition.", "type": "noul"}}, "state": "Before the drought trial, plant ecologist Mira documented: “If all four soil-moisture sensors pass the 08:00 calibration, assign tray B to drought and tray A to daily watering; otherwise water both trays daily as controls.” The four sensors governed by this condition were B-1, B-2, B-3, and B-4. At 08:00, B-1, B-2, and B-3 each passed calibration. At Mira’s 08:00 pre-watering calibration, sensor B-4 must read within 2 percent of the reference value to pass. At 08:00, sensor B-4’s calibration reading was 5 percent from the reference value. Technician Leon withheld water from tray B for four days and watered tray A daily. Heights were measured each morning, and the data quality reviewer confirmed that the growth records were complete. The reviewer is now reviewing whether the documented calibration condition was met for each sensor before any watering decision was executed."}, "method": "c2d", "provenance": {"source_id": "diverse-243", "source_is_synthetic": true, "source_sha256": "b896cf9dcef73bc378eb41f62dcd611d36c9c85c2624f8c85aa57ccdabef410c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-01", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy evidence bullets and question bindings (ST-41, June 3, completeness criteria) while altering only the specimen-total figure to create a coherent mismatch, with no gold answer or rule table embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships, including the universally quantified vial/specimen conditions. The focus a4 is a factual count-equality relation rather than a policy conclusion. The base and counter differ only on a4 and are jointly realizable: identification counts can differ from specimen totals even while every specimen appearing in the specimen record has a resolved identification, for example because the identification sheet contains an extra ST-41 entry. Policy evidence preserves all substantive state-originating rules, including temporal priority, routing, and the confidence rubric, without unnecessarily duplicating instructions from the retained questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a4 entails that the latest controlling ST-41 specimen total and identification count are unequal. The original policy explicitly requires those counts to reconcile, so this failure is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes documented effort, correct labeling and provenance, reconciled specimen and identification counts, resolved identifications, and all four required habitat observations under the latest controlling records. These jointly satisfy every stated completeness requirement and are sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The latest controlling effort record for site ST-41 documents six timed kick samples."}, {"id": "a2", "statement": "Every vial assigned to site ST-41 in the latest controlling submission bears an ST-41 site label."}, {"id": "a3", "statement": "Every vial assigned to site ST-41 in the latest controlling submission was collected at ST-41 according to the latest controlling provenance records."}, {"id": "a4", "statement": "The specimen count assigned to site ST-41 in the latest controlling specimen-total record equals the identification count assigned to ST-41 in the latest controlling identification sheet."}, {"id": "a5", "statement": "Every specimen assigned to site ST-41 in the latest controlling specimen record has a resolved taxonomic identification in the latest controlling identification sheet."}, {"id": "a6", "statement": "The latest controlling habitat record for site ST-41 contains a flow observation."}, {"id": "a7", "statement": "The latest controlling habitat record for site ST-41 contains a substrate observation."}, {"id": "a8", "statement": "The latest controlling habitat record for site ST-41 contains a canopy observation."}, {"id": "a9", "statement": "The latest controlling habitat record for site ST-41 contains a bank-condition observation."}], "base_state_json": "{\"context\":\"Field survey site ST-41 case note: Six timed kick samples were logged, matching the required effort documentation. All vials assigned to ST-41 bear the correct ST-41 site label, and provenance records confirm each was actually collected at ST-41. Every specimen assigned to ST-41 has a resolved taxonomic identification recorded in the identification sheet. Habitat observations for the site include a flow observation, a substrate observation, a canopy observation, and a bank-condition observation, all present in the latest controlling habitat record. The latest controlling specimen-total record dated June 3 assigns a specimen count of 42 to site ST-41. The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41.\",\"evidence\":[\"Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.\",\"Use the latest dated correction over earlier records.\",\"Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification.\"],\"request\":\"Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "The latest controlling specimen-total record dated June 3 assigns a specimen count of 42 to site ST-41."}, {"path": ["context"], "text": "The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations."}, {"path": ["evidence", "1"], "text": "Use the latest dated correction over earlier records."}, {"path": ["evidence", "2"], "text": "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."}, {"path": ["request"], "text": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}], "rules": [{"justification": "Unequal latest controlling ST-41 specimen and identification counts fail the required count reconciliation, making the survey incomplete and subject to data review.", "target": "false", "when": [{"atom_id": "a4", "state": "refuted"}]}, {"justification": "Under the latest controlling information, the documented effort and site labeling are correct, specimen and identification records reconcile, every specimen has a resolved identification, and all four required habitat observations are present.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The latest controlling specimen-total record dated June 3 assigns a specimen count of 42 to site ST-41.", "negative_left": "The latest controlling specimen-total record dated June 3 assigns a specimen count of 47 to site ST-41.", "negative_right": "The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41.", "right": "The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41."}, "verifier_independent_model": false}, "family": "fast-43-diverse-249-001", "id": "fast-43-diverse-249-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one completeness requirement fails under the latest information, so ST-41 is not ready and must be routed according to the stated rule.", "true": "Yes — all required effort, labeling, counts, identifications, and habitat records reconcile under the latest information, so ST-41 is ready for acceptance."}, "instructions": "Answer yes only if every completeness requirement remains satisfied after applying temporal updates. Otherwise answer no. Base confidence on the ordered definitions in the request.", "type": "noul"}}, "state": {"context": "Field survey site ST-41 case note: Six timed kick samples were logged, matching the required effort documentation. All vials assigned to ST-41 bear the correct ST-41 site label, and provenance records confirm each was actually collected at ST-41. Every specimen assigned to ST-41 has a resolved taxonomic identification recorded in the identification sheet. Habitat observations for the site include a flow observation, a substrate observation, a canopy observation, and a bank-condition observation, all present in the latest controlling habitat record. The latest controlling specimen-total record dated June 3 assigns a specimen count of 42 to site ST-41. The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41.", "evidence": ["Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.", "Use the latest dated correction over earlier records.", "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."], "request": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}}, "method": "c2d", "provenance": {"source_id": "diverse-249", "source_is_synthetic": true, "source_sha256": "6e988fdbc9c26cb1580afa5cc854fba73d559d78fe22412ca85c5ea6aaacf1d4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same policy evidence bullets and question bindings (ST-41, June 3, completeness criteria) while altering only the specimen-total figure to create a coherent mismatch, with no gold answer or rule table embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms express individual factual relationships, including the universally quantified vial/specimen conditions. The focus a4 is a factual count-equality relation rather than a policy conclusion. The base and counter differ only on a4 and are jointly realizable: identification counts can differ from specimen totals even while every specimen appearing in the specimen record has a resolved identification, for example because the identification sheet contains an extra ST-41 entry. Policy evidence preserves all substantive state-originating rules, including temporal priority, routing, and the confidence rubric, without unnecessarily duplicating instructions from the retained questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuting a4 entails that the latest controlling ST-41 specimen total and identification count are unequal. The original policy explicitly requires those counts to reconcile, so this failure is sufficient for a false decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes documented effort, correct labeling and provenance, reconciled specimen and identification counts, resolved identifications, and all four required habitat observations under the latest controlling records. These jointly satisfy every stated completeness requirement and are sufficient for a true decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The latest controlling effort record for site ST-41 documents six timed kick samples."}, {"id": "a2", "statement": "Every vial assigned to site ST-41 in the latest controlling submission bears an ST-41 site label."}, {"id": "a3", "statement": "Every vial assigned to site ST-41 in the latest controlling submission was collected at ST-41 according to the latest controlling provenance records."}, {"id": "a4", "statement": "The specimen count assigned to site ST-41 in the latest controlling specimen-total record equals the identification count assigned to ST-41 in the latest controlling identification sheet."}, {"id": "a5", "statement": "Every specimen assigned to site ST-41 in the latest controlling specimen record has a resolved taxonomic identification in the latest controlling identification sheet."}, {"id": "a6", "statement": "The latest controlling habitat record for site ST-41 contains a flow observation."}, {"id": "a7", "statement": "The latest controlling habitat record for site ST-41 contains a substrate observation."}, {"id": "a8", "statement": "The latest controlling habitat record for site ST-41 contains a canopy observation."}, {"id": "a9", "statement": "The latest controlling habitat record for site ST-41 contains a bank-condition observation."}], "base_state_json": "{\"context\":\"Field survey site ST-41 case note: Six timed kick samples were logged, matching the required effort documentation. All vials assigned to ST-41 bear the correct ST-41 site label, and provenance records confirm each was actually collected at ST-41. Every specimen assigned to ST-41 has a resolved taxonomic identification recorded in the identification sheet. Habitat observations for the site include a flow observation, a substrate observation, a canopy observation, and a bank-condition observation, all present in the latest controlling habitat record. The latest controlling specimen-total record dated June 3 assigns a specimen count of 42 to site ST-41. The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41.\",\"evidence\":[\"Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.\",\"Use the latest dated correction over earlier records.\",\"Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification.\"],\"request\":\"Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "The latest controlling specimen-total record dated June 3 assigns a specimen count of 42 to site ST-41."}, {"path": ["context"], "text": "The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations."}, {"path": ["evidence", "1"], "text": "Use the latest dated correction over earlier records."}, {"path": ["evidence", "2"], "text": "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."}, {"path": ["request"], "text": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}], "rules": [{"justification": "Unequal latest controlling ST-41 specimen and identification counts fail the required count reconciliation, making the survey incomplete and subject to data review.", "target": "false", "when": [{"atom_id": "a4", "state": "refuted"}]}, {"justification": "Under the latest controlling information, the documented effort and site labeling are correct, specimen and identification records reconcile, every specimen has a resolved identification, and all four required habitat observations are present.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "The latest controlling specimen-total record dated June 3 assigns a specimen count of 42 to site ST-41.", "negative_left": "The latest controlling specimen-total record dated June 3 assigns a specimen count of 47 to site ST-41.", "negative_right": "The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41.", "right": "The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41."}, "verifier_independent_model": false}, "family": "fast-43-diverse-249-001", "id": "fast-43-diverse-249-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — at least one completeness requirement fails under the latest information, so ST-41 is not ready and must be routed according to the stated rule.", "true": "Yes — all required effort, labeling, counts, identifications, and habitat records reconcile under the latest information, so ST-41 is ready for acceptance."}, "instructions": "Answer yes only if every completeness requirement remains satisfied after applying temporal updates. Otherwise answer no. Base confidence on the ordered definitions in the request.", "type": "noul"}}, "state": {"context": "Field survey site ST-41 case note: Six timed kick samples were logged, matching the required effort documentation. All vials assigned to ST-41 bear the correct ST-41 site label, and provenance records confirm each was actually collected at ST-41. Every specimen assigned to ST-41 has a resolved taxonomic identification recorded in the identification sheet. Habitat observations for the site include a flow observation, a substrate observation, a canopy observation, and a bank-condition observation, all present in the latest controlling habitat record. The latest controlling specimen-total record dated June 3 assigns a specimen count of 47 to site ST-41. The latest controlling identification sheet dated June 3 assigns an identification count of 42 to site ST-41.", "evidence": ["Completeness requires documented effort, correct site labels, reconciled specimen counts and identifications, and all four habitat observations.", "Use the latest dated correction over earlier records.", "Any unresolved site/count mismatch makes the survey incomplete and routes it to data review; missing sampling effort would instead route to sampling, and unresolved species names to identification."], "request": "Is ST-41 complete and ready for acceptance? Also select confidence: High when the rule and decisive record are explicit; Medium when interpretation is needed; Low when required records are absent."}}, "method": "c2d", "provenance": {"source_id": "diverse-249", "source_is_synthetic": true, "source_sha256": "6e988fdbc9c26cb1580afa5cc854fba73d559d78fe22412ca85c5ea6aaacf1d4", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same three sites, route‑sheet scope, and required documentation without altering the rubric's exception logic; the counterfactual only changes PR‑2's count‑sheet tally to 29, creating a coherent single discrepancy rather than a contradictory duplicate; the two focus evidence spans are factual tally statements, not policy text; and neither context states or implies the final completeness score or level.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\", \"evidence\": [\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\", \"PR-1 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.\", \"PR-2 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, and habitat notes present.\", \"PR-2's jar specimen tally from the field survey was recorded as 34 specimens.\", \"PR-2's count-sheet specimen tally submitted with the packet lists 34 specimens.\", \"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\", \"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 34 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 34 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 34 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 34 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 29 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 34 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-011", "id": "fast-43-diverse-252-011-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 34 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 34 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same three sites, route‑sheet scope, and required documentation without altering the rubric's exception logic; the counterfactual only changes PR‑2's count‑sheet tally to 29, creating a coherent single discrepancy rather than a contradictory duplicate; the two focus evidence spans are factual tally statements, not policy text; and neither context states or implies the final completeness score or level.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\", \"evidence\": [\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\", \"PR-1 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.\", \"PR-2 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, and habitat notes present.\", \"PR-2's jar specimen tally from the field survey was recorded as 34 specimens.\", \"PR-2's count-sheet specimen tally submitted with the packet lists 34 specimens.\", \"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\", \"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 34 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 34 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 34 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 34 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 29 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 34 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-011", "id": "fast-43-diverse-252-011-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 34 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 29 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim, preserve site bindings PR-1/PR-2/PR-3 and the dry-site exception language, and the two focus sentences are complete factual claims about specimen counts rather than policy text; the counterfactual changes only the jar count (14→11) creating a coherent single discrepancy without duplicating or contradicting other stated measurements, and neither context reveals a rubric level, code, or instruction dictating the output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"Field packet review note: The route sheet lists exactly three Pine Run sites: PR-1, PR-2, and PR-3. PR-1 is a wet site with recorded 20-minute kick-net sampling effort, a mappable site label, agreeing jar and count-sheet specimen totals, and habitat observations on file. PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations on file. The specimen jar labeled PR-2 was counted and logged at 14 specimens. The PR-2 count sheet submitted with the route records 14 specimens. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and habitat observations are on file for PR-3, satisfying the dry-site exception so no sampling effort, jar, or specimen count is required there. No unresolved required-field deficiency or record discrepancy remains across the three sites beyond what is described above, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.\",\"request\":\"Rate the packet's completeness confidence under the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["context"], "text": "The specimen jar labeled PR-2 was counted and logged at 14 specimens."}, {"path": ["context"], "text": "The PR-2 count sheet submitted with the route records 14 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The specimen jar labeled PR-2 was counted and logged at 14 specimens.", "negative_left": "The specimen jar labeled PR-2 was counted and logged at 11 specimens.", "negative_right": "The PR-2 count sheet submitted with the route records 14 specimens.", "right": "The PR-2 count sheet submitted with the route records 14 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-013", "id": "fast-43-diverse-252-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "Field packet review note: The route sheet lists exactly three Pine Run sites: PR-1, PR-2, and PR-3. PR-1 is a wet site with recorded 20-minute kick-net sampling effort, a mappable site label, agreeing jar and count-sheet specimen totals, and habitat observations on file. PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations on file. The specimen jar labeled PR-2 was counted and logged at 14 specimens. The PR-2 count sheet submitted with the route records 14 specimens. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and habitat observations are on file for PR-3, satisfying the dry-site exception so no sampling effort, jar, or specimen count is required there. No unresolved required-field deficiency or record discrepancy remains across the three sites beyond what is described above, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.", "request": "Rate the packet's completeness confidence under the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original questions verbatim, preserve site bindings PR-1/PR-2/PR-3 and the dry-site exception language, and the two focus sentences are complete factual claims about specimen counts rather than policy text; the counterfactual changes only the jar count (14→11) creating a coherent single discrepancy without duplicating or contradicting other stated measurements, and neither context reveals a rubric level, code, or instruction dictating the output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"Field packet review note: The route sheet lists exactly three Pine Run sites: PR-1, PR-2, and PR-3. PR-1 is a wet site with recorded 20-minute kick-net sampling effort, a mappable site label, agreeing jar and count-sheet specimen totals, and habitat observations on file. PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations on file. The specimen jar labeled PR-2 was counted and logged at 14 specimens. The PR-2 count sheet submitted with the route records 14 specimens. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and habitat observations are on file for PR-3, satisfying the dry-site exception so no sampling effort, jar, or specimen count is required there. No unresolved required-field deficiency or record discrepancy remains across the three sites beyond what is described above, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.\",\"request\":\"Rate the packet's completeness confidence under the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["context"], "text": "The specimen jar labeled PR-2 was counted and logged at 14 specimens."}, {"path": ["context"], "text": "The PR-2 count sheet submitted with the route records 14 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The specimen jar labeled PR-2 was counted and logged at 14 specimens.", "negative_left": "The specimen jar labeled PR-2 was counted and logged at 11 specimens.", "negative_right": "The PR-2 count sheet submitted with the route records 14 specimens.", "right": "The PR-2 count sheet submitted with the route records 14 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-013", "id": "fast-43-diverse-252-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "Field packet review note: The route sheet lists exactly three Pine Run sites: PR-1, PR-2, and PR-3. PR-1 is a wet site with recorded 20-minute kick-net sampling effort, a mappable site label, agreeing jar and count-sheet specimen totals, and habitat observations on file. PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations on file. The specimen jar labeled PR-2 was counted and logged at 11 specimens. The PR-2 count sheet submitted with the route records 14 specimens. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and habitat observations are on file for PR-3, satisfying the dry-site exception so no sampling effort, jar, or specimen count is required there. No unresolved required-field deficiency or record discrepancy remains across the three sites beyond what is described above, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.", "request": "Rate the packet's completeness confidence under the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same site scope, wet/dry classifications, and the rubric's governing exception/discrepancy language, with all policy content already carried by the unchanged questions object; the counterfactual's changed jar count (41 vs 34) is a plausible factual alternative that stays consistent with the context's mention of a 'possible discrepancy' and does not duplicate or contradict other stated measurements; each evidence list contains exactly two complete factual sentences with no policy definitions, rule tables, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager is validating a field survey packet before routing it onward. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with recorded sampling effort, a mappable site label, habitat observations, and its jar specimen count agrees with its count-sheet specimen count. PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and PR-3 has habitat observations. No unresolved required-field deficiency exists across the three sites other than a possible deficiency from the PR-2 specimen-count comparison, no unresolved record discrepancy exists other than a possible discrepancy between PR-2's jar and count-sheet specimen counts, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.\",\"evidence\":[\"The jar for PR-2 was tallied at 34 specimens by the field technician.\",\"The count sheet for PR-2 lists 34 specimens for that same collection.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "0"], "text": "The jar for PR-2 was tallied at 34 specimens by the field technician."}, {"path": ["evidence", "1"], "text": "The count sheet for PR-2 lists 34 specimens for that same collection."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The jar for PR-2 was tallied at 34 specimens by the field technician.", "negative_left": "The jar for PR-2 was tallied at 41 specimens by the field technician.", "negative_right": "The count sheet for PR-2 lists 34 specimens for that same collection.", "right": "The count sheet for PR-2 lists 34 specimens for that same collection."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-017", "id": "fast-43-diverse-252-017-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a field survey packet before routing it onward. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with recorded sampling effort, a mappable site label, habitat observations, and its jar specimen count agrees with its count-sheet specimen count. PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and PR-3 has habitat observations. No unresolved required-field deficiency exists across the three sites other than a possible deficiency from the PR-2 specimen-count comparison, no unresolved record discrepancy exists other than a possible discrepancy between PR-2's jar and count-sheet specimen counts, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.", "evidence": ["The jar for PR-2 was tallied at 34 specimens by the field technician.", "The count sheet for PR-2 lists 34 specimens for that same collection."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same site scope, wet/dry classifications, and the rubric's governing exception/discrepancy language, with all policy content already carried by the unchanged questions object; the counterfactual's changed jar count (41 vs 34) is a plausible factual alternative that stays consistent with the context's mention of a 'possible discrepancy' and does not duplicate or contradict other stated measurements; each evidence list contains exactly two complete factual sentences with no policy definitions, rule tables, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager is validating a field survey packet before routing it onward. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with recorded sampling effort, a mappable site label, habitat observations, and its jar specimen count agrees with its count-sheet specimen count. PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and PR-3 has habitat observations. No unresolved required-field deficiency exists across the three sites other than a possible deficiency from the PR-2 specimen-count comparison, no unresolved record discrepancy exists other than a possible discrepancy between PR-2's jar and count-sheet specimen counts, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.\",\"evidence\":[\"The jar for PR-2 was tallied at 34 specimens by the field technician.\",\"The count sheet for PR-2 lists 34 specimens for that same collection.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "0"], "text": "The jar for PR-2 was tallied at 34 specimens by the field technician."}, {"path": ["evidence", "1"], "text": "The count sheet for PR-2 lists 34 specimens for that same collection."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The jar for PR-2 was tallied at 34 specimens by the field technician.", "negative_left": "The jar for PR-2 was tallied at 41 specimens by the field technician.", "negative_right": "The count sheet for PR-2 lists 34 specimens for that same collection.", "right": "The count sheet for PR-2 lists 34 specimens for that same collection."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-017", "id": "fast-43-diverse-252-017-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a field survey packet before routing it onward. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with recorded sampling effort, a mappable site label, habitat observations, and its jar specimen count agrees with its count-sheet specimen count. PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and PR-3 has habitat observations. No unresolved required-field deficiency exists across the three sites other than a possible deficiency from the PR-2 specimen-count comparison, no unresolved record discrepancy exists other than a possible discrepancy between PR-2's jar and count-sheet specimen counts, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.", "evidence": ["The jar for PR-2 was tallied at 41 specimens by the field technician.", "The count sheet for PR-2 lists 34 specimens for that same collection."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all wet/dry-site policy elements implicitly via the unchanged question, keep PR-1/PR-2/PR-3 bindings intact, use two complete factual sentences as focus evidence, and the counterfactual only swaps the recount figure from 37 to 42 creating a plausible discrepancy without contradicting other fixed facts or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"route_sheet\": \"Lists PR-1, PR-2, PR-3 as the only scoped sites.\", \"PR-1\": \"Wet site with 20-minute kick-net effort, mappable site label, habitat observations recorded, and jar count of 24 specimens matching the count sheet's 24 specimens.\", \"PR-2\": \"Wet site with 20-minute kick-net effort, mappable site label, and habitat observations recorded. The jar labeled PR-2 was recounted on March 4th and contains 37 specimens. The count sheet for PR-2, dated March 4th, lists a total of 37 specimens.\", \"PR-3\": \"Dry site; dated photo documents absence of flowing water, and habitat observations are present, satisfying the dry-site exception with no sampling effort or specimen counts required.\", \"packet_status\": \"No unresolved required-field deficiency remains beyond the possible PR-2 specimen-count comparison, no unresolved record discrepancy remains beyond that same comparison, and no nonrequired formatting or organizational irregularity remains.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["PR-2"], "text": "The jar labeled PR-2 was recounted on March 4th and contains 37 specimens."}, {"path": ["PR-2"], "text": "The count sheet for PR-2, dated March 4th, lists a total of 37 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The jar labeled PR-2 was recounted on March 4th and contains 37 specimens.", "negative_left": "The jar labeled PR-2 was recounted on March 4th and contains 42 specimens.", "negative_right": "The count sheet for PR-2, dated March 4th, lists a total of 37 specimens.", "right": "The count sheet for PR-2, dated March 4th, lists a total of 37 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-022", "id": "fast-43-diverse-252-022-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"PR-1": "Wet site with 20-minute kick-net effort, mappable site label, habitat observations recorded, and jar count of 24 specimens matching the count sheet's 24 specimens.", "PR-2": "Wet site with 20-minute kick-net effort, mappable site label, and habitat observations recorded. The jar labeled PR-2 was recounted on March 4th and contains 37 specimens. The count sheet for PR-2, dated March 4th, lists a total of 37 specimens.", "PR-3": "Dry site; dated photo documents absence of flowing water, and habitat observations are present, satisfying the dry-site exception with no sampling effort or specimen counts required.", "packet_status": "No unresolved required-field deficiency remains beyond the possible PR-2 specimen-count comparison, no unresolved record discrepancy remains beyond that same comparison, and no nonrequired formatting or organizational irregularity remains.", "route_sheet": "Lists PR-1, PR-2, PR-3 as the only scoped sites."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all wet/dry-site policy elements implicitly via the unchanged question, keep PR-1/PR-2/PR-3 bindings intact, use two complete factual sentences as focus evidence, and the counterfactual only swaps the recount figure from 37 to 42 creating a plausible discrepancy without contradicting other fixed facts or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"route_sheet\": \"Lists PR-1, PR-2, PR-3 as the only scoped sites.\", \"PR-1\": \"Wet site with 20-minute kick-net effort, mappable site label, habitat observations recorded, and jar count of 24 specimens matching the count sheet's 24 specimens.\", \"PR-2\": \"Wet site with 20-minute kick-net effort, mappable site label, and habitat observations recorded. The jar labeled PR-2 was recounted on March 4th and contains 37 specimens. The count sheet for PR-2, dated March 4th, lists a total of 37 specimens.\", \"PR-3\": \"Dry site; dated photo documents absence of flowing water, and habitat observations are present, satisfying the dry-site exception with no sampling effort or specimen counts required.\", \"packet_status\": \"No unresolved required-field deficiency remains beyond the possible PR-2 specimen-count comparison, no unresolved record discrepancy remains beyond that same comparison, and no nonrequired formatting or organizational irregularity remains.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["PR-2"], "text": "The jar labeled PR-2 was recounted on March 4th and contains 37 specimens."}, {"path": ["PR-2"], "text": "The count sheet for PR-2, dated March 4th, lists a total of 37 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The jar labeled PR-2 was recounted on March 4th and contains 37 specimens.", "negative_left": "The jar labeled PR-2 was recounted on March 4th and contains 42 specimens.", "negative_right": "The count sheet for PR-2, dated March 4th, lists a total of 37 specimens.", "right": "The count sheet for PR-2, dated March 4th, lists a total of 37 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-022", "id": "fast-43-diverse-252-022-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"PR-1": "Wet site with 20-minute kick-net effort, mappable site label, habitat observations recorded, and jar count of 24 specimens matching the count sheet's 24 specimens.", "PR-2": "Wet site with 20-minute kick-net effort, mappable site label, and habitat observations recorded. The jar labeled PR-2 was recounted on March 4th and contains 42 specimens. The count sheet for PR-2, dated March 4th, lists a total of 37 specimens.", "PR-3": "Dry site; dated photo documents absence of flowing water, and habitat observations are present, satisfying the dry-site exception with no sampling effort or specimen counts required.", "packet_status": "No unresolved required-field deficiency remains beyond the possible PR-2 specimen-count comparison, no unresolved record discrepancy remains beyond that same comparison, and no nonrequired formatting or organizational irregularity remains.", "route_sheet": "Lists PR-1, PR-2, PR-3 as the only scoped sites."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the three-site scope and wet/dry rubric requirements from the original question, so policy is preserved without added exceptions. Focus evidence swaps the jar tally from 34 to 29 while keeping the count-sheet figure at 34, still yielding a single unresolved discrepancy consistent with the rest of the unchanged evidence, so it is coherent. Both focus sentences are complete factual statements, not policy text. Neither context states or implies the rubric score or label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager is validating a packet before routing it to the taxonomic identifier. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\",\"evidence\":[\"PR-1 is a wet site with recorded sampling effort, a mappable site label, agreeing jar and count-sheet specimen counts, and habitat observations.\",\"PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations.\",\"The jar labeled PR-2 was tallied at 34 specimens during the lab count on May 3.\",\"The count sheet for PR-2 records a total of 34 specimens collected.\",\"PR-3 is a dry site. The dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3. PR-3 has habitat observations.\",\"No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison.\",\"No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "2"], "text": "The jar labeled PR-2 was tallied at 34 specimens during the lab count on May 3."}, {"path": ["evidence", "3"], "text": "The count sheet for PR-2 records a total of 34 specimens collected."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The jar labeled PR-2 was tallied at 34 specimens during the lab count on May 3.", "negative_left": "The jar labeled PR-2 was tallied at 29 specimens during the lab count on May 3.", "negative_right": "The count sheet for PR-2 records a total of 34 specimens collected.", "right": "The count sheet for PR-2 records a total of 34 specimens collected."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-026", "id": "fast-43-diverse-252-026-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet before routing it to the taxonomic identifier. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "evidence": ["PR-1 is a wet site with recorded sampling effort, a mappable site label, agreeing jar and count-sheet specimen counts, and habitat observations.", "PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations.", "The jar labeled PR-2 was tallied at 34 specimens during the lab count on May 3.", "The count sheet for PR-2 records a total of 34 specimens collected.", "PR-3 is a dry site. The dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3. PR-3 has habitat observations.", "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the three-site scope and wet/dry rubric requirements from the original question, so policy is preserved without added exceptions. Focus evidence swaps the jar tally from 34 to 29 while keeping the count-sheet figure at 34, still yielding a single unresolved discrepancy consistent with the rest of the unchanged evidence, so it is coherent. Both focus sentences are complete factual statements, not policy text. Neither context states or implies the rubric score or label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager is validating a packet before routing it to the taxonomic identifier. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\",\"evidence\":[\"PR-1 is a wet site with recorded sampling effort, a mappable site label, agreeing jar and count-sheet specimen counts, and habitat observations.\",\"PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations.\",\"The jar labeled PR-2 was tallied at 34 specimens during the lab count on May 3.\",\"The count sheet for PR-2 records a total of 34 specimens collected.\",\"PR-3 is a dry site. The dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3. PR-3 has habitat observations.\",\"No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison.\",\"No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "2"], "text": "The jar labeled PR-2 was tallied at 34 specimens during the lab count on May 3."}, {"path": ["evidence", "3"], "text": "The count sheet for PR-2 records a total of 34 specimens collected."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "The jar labeled PR-2 was tallied at 34 specimens during the lab count on May 3.", "negative_left": "The jar labeled PR-2 was tallied at 29 specimens during the lab count on May 3.", "negative_right": "The count sheet for PR-2 records a total of 34 specimens collected.", "right": "The count sheet for PR-2 records a total of 34 specimens collected."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-026", "id": "fast-43-diverse-252-026-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet before routing it to the taxonomic identifier. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "evidence": ["PR-1 is a wet site with recorded sampling effort, a mappable site label, agreeing jar and count-sheet specimen counts, and habitat observations.", "PR-2 is a wet site with recorded sampling effort, a mappable site label, and habitat observations.", "The jar labeled PR-2 was tallied at 29 specimens during the lab count on May 3.", "The count sheet for PR-2 records a total of 34 specimens collected.", "PR-3 is a dry site. The dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3. PR-3 has habitat observations.", "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the three scoped sites and all governing policy is preserved via the unchanged questions object; the counterfactual changes only the PR-2 count-sheet figure to 29, creating a coherent single discrepancy rather than a contradiction, and neither context embeds rubric labels or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager reviewed the Pine Run packet before routing it onward. The watershed coordinator's route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with a recorded 20-minute kick-net effort, a mappable site label, habitat observations, and its jar specimen count agrees with its count-sheet specimen count. PR-2 is also a wet site with recorded sampling effort, a mappable site label, and habitat observations. PR-2's jar specimen count recorded on the intake jar label is 34 specimens. PR-2's count-sheet specimen count recorded on the field count sheet is 34 specimens. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and PR-3 has habitat observations. No unresolved required-field deficiency exists across the three sites beyond any possible one tied to the PR-2 specimen-count comparison, no unresolved record discrepancy exists beyond any possible one tied to that same comparison, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["context"], "text": "PR-2's jar specimen count recorded on the intake jar label is 34 specimens."}, {"path": ["context"], "text": "PR-2's count-sheet specimen count recorded on the field count sheet is 34 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen count recorded on the intake jar label is 34 specimens.", "negative_left": "PR-2's jar specimen count recorded on the intake jar label is 34 specimens.", "negative_right": "PR-2's count-sheet specimen count recorded on the field count sheet is 29 specimens.", "right": "PR-2's count-sheet specimen count recorded on the field count sheet is 34 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-042", "id": "fast-43-diverse-252-042-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reviewed the Pine Run packet before routing it onward. The watershed coordinator's route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with a recorded 20-minute kick-net effort, a mappable site label, habitat observations, and its jar specimen count agrees with its count-sheet specimen count. PR-2 is also a wet site with recorded sampling effort, a mappable site label, and habitat observations. PR-2's jar specimen count recorded on the intake jar label is 34 specimens. PR-2's count-sheet specimen count recorded on the field count sheet is 34 specimens. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and PR-3 has habitat observations. No unresolved required-field deficiency exists across the three sites beyond any possible one tied to the PR-2 specimen-count comparison, no unresolved record discrepancy exists beyond any possible one tied to that same comparison, and no nonrequired formatting, wording, or organizational irregularity remains in the packet."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the three scoped sites and all governing policy is preserved via the unchanged questions object; the counterfactual changes only the PR-2 count-sheet figure to 29, creating a coherent single discrepancy rather than a contradiction, and neither context embeds rubric labels or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager reviewed the Pine Run packet before routing it onward. The watershed coordinator's route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with a recorded 20-minute kick-net effort, a mappable site label, habitat observations, and its jar specimen count agrees with its count-sheet specimen count. PR-2 is also a wet site with recorded sampling effort, a mappable site label, and habitat observations. PR-2's jar specimen count recorded on the intake jar label is 34 specimens. PR-2's count-sheet specimen count recorded on the field count sheet is 34 specimens. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and PR-3 has habitat observations. No unresolved required-field deficiency exists across the three sites beyond any possible one tied to the PR-2 specimen-count comparison, no unresolved record discrepancy exists beyond any possible one tied to that same comparison, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["context"], "text": "PR-2's jar specimen count recorded on the intake jar label is 34 specimens."}, {"path": ["context"], "text": "PR-2's count-sheet specimen count recorded on the field count sheet is 34 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen count recorded on the intake jar label is 34 specimens.", "negative_left": "PR-2's jar specimen count recorded on the intake jar label is 34 specimens.", "negative_right": "PR-2's count-sheet specimen count recorded on the field count sheet is 29 specimens.", "right": "PR-2's count-sheet specimen count recorded on the field count sheet is 34 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-042", "id": "fast-43-diverse-252-042-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reviewed the Pine Run packet before routing it onward. The watershed coordinator's route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with a recorded 20-minute kick-net effort, a mappable site label, habitat observations, and its jar specimen count agrees with its count-sheet specimen count. PR-2 is also a wet site with recorded sampling effort, a mappable site label, and habitat observations. PR-2's jar specimen count recorded on the intake jar label is 34 specimens. PR-2's count-sheet specimen count recorded on the field count sheet is 29 specimens. PR-3 is a dry site; the dry-site photo submitted for PR-3 bears a date and documents the absence of flowing water at PR-3, and PR-3 has habitat observations. No unresolved required-field deficiency exists across the three sites beyond any possible one tied to the PR-2 specimen-count comparison, no unresolved record discrepancy exists beyond any possible one tied to that same comparison, and no nonrequired formatting, wording, or organizational irregularity remains in the packet."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy (routing rules, wet/dry requirements) unchanged from original state and question, and the three-site scope/entities match; evidence spans are plain factual sentences about specimen counts; the counterfactual changes PR-2's count sheet value from 14 to 11, creating a coherent single discrepancy without contradicting other stated facts; no gold answer, rule table, or instructional leakage appears in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier; unresolved field gaps go to sampling, while record conflicts go to data review. The watershed coordinator's route sheet lists exactly three Pine Run sites: PR-1, PR-2, and PR-3. PR-1 is a wet site with a 20-minute kick-net sampling effort recorded, a clear mappable site label, matching jar and count-sheet specimen totals, and habitat notes present. PR-2 is also a wet site with a 20-minute kick-net effort recorded and a mappable site label, and habitat notes present. PR-2's specimen jar contains 14 specimens. PR-2's count sheet lists a specimen count of 14. PR-3 is marked dry; the submitted dry-site photo bears a date and documents the absence of flowing water at PR-3, and habitat notes are present for PR-3. No sampling effort, jar, or specimen count was recorded for PR-3. Aside from the possible PR-2 specimen comparison, no other required-field deficiency or record discrepancy remains across the three sites, and no nonrequired formatting or organizational irregularity remains in the packet.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["context"], "text": "PR-2's specimen jar contains 14 specimens."}, {"path": ["context"], "text": "PR-2's count sheet lists a specimen count of 14."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's specimen jar contains 14 specimens.", "negative_left": "PR-2's specimen jar contains 14 specimens.", "negative_right": "PR-2's count sheet lists a specimen count of 11.", "right": "PR-2's count sheet lists a specimen count of 14."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-050", "id": "fast-43-diverse-252-050-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier; unresolved field gaps go to sampling, while record conflicts go to data review. The watershed coordinator's route sheet lists exactly three Pine Run sites: PR-1, PR-2, and PR-3. PR-1 is a wet site with a 20-minute kick-net sampling effort recorded, a clear mappable site label, matching jar and count-sheet specimen totals, and habitat notes present. PR-2 is also a wet site with a 20-minute kick-net effort recorded and a mappable site label, and habitat notes present. PR-2's specimen jar contains 14 specimens. PR-2's count sheet lists a specimen count of 14. PR-3 is marked dry; the submitted dry-site photo bears a date and documents the absence of flowing water at PR-3, and habitat notes are present for PR-3. No sampling effort, jar, or specimen count was recorded for PR-3. Aside from the possible PR-2 specimen comparison, no other required-field deficiency or record discrepancy remains across the three sites, and no nonrequired formatting or organizational irregularity remains in the packet."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the governing policy (routing rules, wet/dry requirements) unchanged from original state and question, and the three-site scope/entities match; evidence spans are plain factual sentences about specimen counts; the counterfactual changes PR-2's count sheet value from 14 to 11, creating a coherent single discrepancy without contradicting other stated facts; no gold answer, rule table, or instructional leakage appears in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier; unresolved field gaps go to sampling, while record conflicts go to data review. The watershed coordinator's route sheet lists exactly three Pine Run sites: PR-1, PR-2, and PR-3. PR-1 is a wet site with a 20-minute kick-net sampling effort recorded, a clear mappable site label, matching jar and count-sheet specimen totals, and habitat notes present. PR-2 is also a wet site with a 20-minute kick-net effort recorded and a mappable site label, and habitat notes present. PR-2's specimen jar contains 14 specimens. PR-2's count sheet lists a specimen count of 14. PR-3 is marked dry; the submitted dry-site photo bears a date and documents the absence of flowing water at PR-3, and habitat notes are present for PR-3. No sampling effort, jar, or specimen count was recorded for PR-3. Aside from the possible PR-2 specimen comparison, no other required-field deficiency or record discrepancy remains across the three sites, and no nonrequired formatting or organizational irregularity remains in the packet.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["context"], "text": "PR-2's specimen jar contains 14 specimens."}, {"path": ["context"], "text": "PR-2's count sheet lists a specimen count of 14."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's specimen jar contains 14 specimens.", "negative_left": "PR-2's specimen jar contains 14 specimens.", "negative_right": "PR-2's count sheet lists a specimen count of 11.", "right": "PR-2's count sheet lists a specimen count of 14."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-050", "id": "fast-43-diverse-252-050-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier; unresolved field gaps go to sampling, while record conflicts go to data review. The watershed coordinator's route sheet lists exactly three Pine Run sites: PR-1, PR-2, and PR-3. PR-1 is a wet site with a 20-minute kick-net sampling effort recorded, a clear mappable site label, matching jar and count-sheet specimen totals, and habitat notes present. PR-2 is also a wet site with a 20-minute kick-net effort recorded and a mappable site label, and habitat notes present. PR-2's specimen jar contains 14 specimens. PR-2's count sheet lists a specimen count of 11. PR-3 is marked dry; the submitted dry-site photo bears a date and documents the absence of flowing water at PR-3, and habitat notes are present for PR-3. No sampling effort, jar, or specimen count was recorded for PR-3. Aside from the possible PR-2 specimen comparison, no other required-field deficiency or record discrepancy remains across the three sites, and no nonrequired formatting or organizational irregularity remains in the packet."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the three-site scope and the exception/discrepancy rules already carried in the unchanged questions object; evidence spans are plain factual sentences with no rule tables or gold labels; the counterfactual introduces a genuine jar/count-sheet mismatch (27 vs 31) analogous to the original PR-2 discrepancy, remaining internally coherent rather than duplicating an identical measurement.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager reviewed the route sheet listing Pine Run sites PR-1, PR-2, and PR-3 before routing the packet onward. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with a recorded 20-minute kick-net effort, a clearly mappable site label, matching jar and count-sheet specimen counts, and habitat observations noted in the field log. PR-2 is also a wet site, with recorded sampling effort, a mappable site label, and habitat observations. PR-3 is a dry site; the submitted dry-site photo bears a date and documents the absence of flowing water at PR-3, and habitat observations were recorded for PR-3 as well. No other unresolved required-field deficiency or record discrepancy was found across the three sites, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.\", \"evidence\": [\"PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens.\", \"PR-2's count-sheet specimen count for the collection dated May 14 was also recorded as 27 specimens.\"], \"request\": \"Rate the packet's completeness confidence under the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "0"], "text": "PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens."}, {"path": ["evidence", "1"], "text": "PR-2's count-sheet specimen count for the collection dated May 14 was also recorded as 27 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens.", "negative_left": "PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens.", "negative_right": "PR-2's count-sheet specimen count for the collection dated May 14 was recorded as 31 specimens.", "right": "PR-2's count-sheet specimen count for the collection dated May 14 was also recorded as 27 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-054", "id": "fast-43-diverse-252-054-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reviewed the route sheet listing Pine Run sites PR-1, PR-2, and PR-3 before routing the packet onward. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with a recorded 20-minute kick-net effort, a clearly mappable site label, matching jar and count-sheet specimen counts, and habitat observations noted in the field log. PR-2 is also a wet site, with recorded sampling effort, a mappable site label, and habitat observations. PR-3 is a dry site; the submitted dry-site photo bears a date and documents the absence of flowing water at PR-3, and habitat observations were recorded for PR-3 as well. No other unresolved required-field deficiency or record discrepancy was found across the three sites, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.", "evidence": ["PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens.", "PR-2's count-sheet specimen count for the collection dated May 14 was also recorded as 27 specimens."], "request": "Rate the packet's completeness confidence under the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the three-site scope and the exception/discrepancy rules already carried in the unchanged questions object; evidence spans are plain factual sentences with no rule tables or gold labels; the counterfactual introduces a genuine jar/count-sheet mismatch (27 vs 31) analogous to the original PR-2 discrepancy, remaining internally coherent rather than duplicating an identical measurement.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager reviewed the route sheet listing Pine Run sites PR-1, PR-2, and PR-3 before routing the packet onward. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with a recorded 20-minute kick-net effort, a clearly mappable site label, matching jar and count-sheet specimen counts, and habitat observations noted in the field log. PR-2 is also a wet site, with recorded sampling effort, a mappable site label, and habitat observations. PR-3 is a dry site; the submitted dry-site photo bears a date and documents the absence of flowing water at PR-3, and habitat observations were recorded for PR-3 as well. No other unresolved required-field deficiency or record discrepancy was found across the three sites, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.\", \"evidence\": [\"PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens.\", \"PR-2's count-sheet specimen count for the collection dated May 14 was also recorded as 27 specimens.\"], \"request\": \"Rate the packet's completeness confidence under the supplied rubric.\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "0"], "text": "PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens."}, {"path": ["evidence", "1"], "text": "PR-2's count-sheet specimen count for the collection dated May 14 was also recorded as 27 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens.", "negative_left": "PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens.", "negative_right": "PR-2's count-sheet specimen count for the collection dated May 14 was recorded as 31 specimens.", "right": "PR-2's count-sheet specimen count for the collection dated May 14 was also recorded as 27 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-054", "id": "fast-43-diverse-252-054-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reviewed the route sheet listing Pine Run sites PR-1, PR-2, and PR-3 before routing the packet onward. The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites. PR-1 is a wet site with a recorded 20-minute kick-net effort, a clearly mappable site label, matching jar and count-sheet specimen counts, and habitat observations noted in the field log. PR-2 is also a wet site, with recorded sampling effort, a mappable site label, and habitat observations. PR-3 is a dry site; the submitted dry-site photo bears a date and documents the absence of flowing water at PR-3, and habitat observations were recorded for PR-3 as well. No other unresolved required-field deficiency or record discrepancy was found across the three sites, and no nonrequired formatting, wording, or organizational irregularity remains in the packet.", "evidence": ["PR-2's jar specimen count for the collection dated May 14 was recorded as 27 specimens.", "PR-2's count-sheet specimen count for the collection dated May 14 was recorded as 31 specimens."], "request": "Rate the packet's completeness confidence under the supplied rubric."}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the three-site scope, wet/dry exception language, and the discrepancy framing without adding rules or labels; the counterfactual changes only the count-sheet tally (37 vs 41) creating a genuine discrepancy consistent with the 'possible one tied to PR-2' evidence line, and no gold answer or rule table is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager reviews a packet submitted by a field survey lead ahead of routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\",\"evidence\":[\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\",\"PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.\",\"PR-2 is a wet site with a mappable label, recorded 25-minute kick-net sampling effort, and habitat notes present.\",\"PR-2's jar specimen tally from the field survey was recorded as 41 specimens.\",\"PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.\",\"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\",\"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-079", "id": "fast-43-diverse-252-079-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reviews a packet submitted by a field survey lead ahead of routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 25-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the three-site scope, wet/dry exception language, and the discrepancy framing without adding rules or labels; the counterfactual changes only the count-sheet tally (37 vs 41) creating a genuine discrepancy consistent with the 'possible one tied to PR-2' evidence line, and no gold answer or rule table is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager reviews a packet submitted by a field survey lead ahead of routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\",\"evidence\":[\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\",\"PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.\",\"PR-2 is a wet site with a mappable label, recorded 25-minute kick-net sampling effort, and habitat notes present.\",\"PR-2's jar specimen tally from the field survey was recorded as 41 specimens.\",\"PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.\",\"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\",\"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-079", "id": "fast-43-diverse-252-079-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reviews a packet submitted by a field survey lead ahead of routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 25-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing policy (wet-site requirements, dry-site exception, discrepancy routing rule) as reflected in the unchanged questions object, and preserve the same entities/sites and time scope. The two focus evidence sentences are complete factual statements about jar and count-sheet tallies, not policy text. The counterfactual changes only the count-sheet tally (47→41), which is a coherent single-fact alteration that still creates a specimen-count discrepancy consistent with surrounding 'possible discrepancy' evidence, without contradicting other stated facts. Neither context contains a gold answer, label, rule table, or instruction directing the output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier; unresolved field gaps go to sampling, while record conflicts go to data review. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\",\"evidence\":[\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\",\"PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.\",\"PR-2 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, and habitat notes present.\",\"PR-2's jar specimen tally from the field survey was recorded as 47 specimens.\",\"PR-2's count-sheet specimen tally submitted with the packet lists 47 specimens.\",\"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\",\"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 47 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 47 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 47 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 47 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 47 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-090", "id": "fast-43-diverse-252-090-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier; unresolved field gaps go to sampling, while record conflicts go to data review. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 47 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 47 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing policy (wet-site requirements, dry-site exception, discrepancy routing rule) as reflected in the unchanged questions object, and preserve the same entities/sites and time scope. The two focus evidence sentences are complete factual statements about jar and count-sheet tallies, not policy text. The counterfactual changes only the count-sheet tally (47→41), which is a coherent single-fact alteration that still creates a specimen-count discrepancy consistent with surrounding 'possible discrepancy' evidence, without contradicting other stated facts. Neither context contains a gold answer, label, rule table, or instruction directing the output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier; unresolved field gaps go to sampling, while record conflicts go to data review. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\",\"evidence\":[\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\",\"PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.\",\"PR-2 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, and habitat notes present.\",\"PR-2's jar specimen tally from the field survey was recorded as 47 specimens.\",\"PR-2's count-sheet specimen tally submitted with the packet lists 47 specimens.\",\"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\",\"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 47 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 47 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 47 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 47 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 47 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-090", "id": "fast-43-diverse-252-090-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier; unresolved field gaps go to sampling, while record conflicts go to data review. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 47 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original PR-1/PR-2/PR-3 bindings and rely on the unchanged questions object for governing policy, adding no invented rules; the two focus sentences are plain factual tally statements, not policy text; the counterfactual only changes PR-2's count-sheet tally from 41 to 37, producing a genuine, non-duplicative discrepancy consistent with the rest of the unchanged evidence; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\", \"evidence\": [\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\", \"PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 19 specimens agreeing on both jar and count sheet.\", \"PR-2 is a wet site with a mappable label, recorded 25-minute kick-net sampling effort, and habitat notes present.\", \"PR-2's jar specimen tally from the field survey was recorded as 41 specimens.\", \"PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.\", \"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\", \"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-096", "id": "fast-43-diverse-252-096-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 19 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 25-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original PR-1/PR-2/PR-3 bindings and rely on the unchanged questions object for governing policy, adding no invented rules; the two focus sentences are plain factual tally statements, not policy text; the counterfactual only changes PR-2's count-sheet tally from 41 to 37, producing a genuine, non-duplicative discrepancy consistent with the rest of the unchanged evidence; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\", \"evidence\": [\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\", \"PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 19 specimens agreeing on both jar and count sheet.\", \"PR-2 is a wet site with a mappable label, recorded 25-minute kick-net sampling effort, and habitat notes present.\", \"PR-2's jar specimen tally from the field survey was recorded as 41 specimens.\", \"PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.\", \"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\", \"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-096", "id": "fast-43-diverse-252-096-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, habitat notes present, and 19 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 25-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the three-site scope, wet/dry exception details, and the possible-discrepancy hedge required by the unchanged rubric; the counterfactual only changes PR-2's count-sheet tally from 41 to 37, a coherent factual edit that still matches the 'possible one discrepancy' language, and neither context states a score or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\", \"evidence\": [\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\", \"PR-1 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.\", \"PR-2 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, and habitat notes present.\", \"PR-2's jar specimen tally from the field survey was recorded as 41 specimens.\", \"PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.\", \"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\", \"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-099", "id": "fast-43-diverse-252-099-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the three-site scope, wet/dry exception details, and the possible-discrepancy hedge required by the unchanged rubric; the counterfactual only changes PR-2's count-sheet tally from 41 to 37, a coherent factual edit that still matches the 'possible one discrepancy' language, and neither context states a score or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\": \"An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\", \"evidence\": [\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\", \"PR-1 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.\", \"PR-2 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, and habitat notes present.\", \"PR-2's jar specimen tally from the field survey was recorded as 41 specimens.\", \"PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.\", \"PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\", \"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\", \"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_left": "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-099", "id": "fast-43-diverse-252-099-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager is validating a packet from a field survey lead before routing it to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, recorded 20-minute kick-net sampling effort, habitat notes present, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, recorded 15-minute kick-net sampling effort, and habitat notes present.", "PR-2's jar specimen tally from the field survey was recorded as 41 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "PR-3 is marked dry; the dated photo and habitat notes document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the three-site scope, wet/dry policy language stays only in the unchanged question, and the focus evidence pair are complete factual sentences about PR-2's tallies; the counterfactual changes the count-sheet figure from 41 to 37, creating the discrepancy already anticipated by the surrounding 'possible one' caveats, so it remains internally consistent without contradicting other measurements or leaking any rule/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager reviews a packet before routing to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\",\"evidence\":[\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\",\"PR-1 is a wet site with a mappable label, a recorded 15-minute dip-net sampling effort, habitat observations noting cobble substrate, and 24 specimens agreeing on both jar and count sheet.\",\"PR-2 is a wet site with a mappable label, a recorded 15-minute dip-net sampling effort, and habitat observations noting silty banks.\",\"PR-2's jar specimen tally recorded during the field survey lists 41 specimens.\",\"PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.\",\"PR-3 is marked dry; the dated photo and habitat observations document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\",\"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally recorded during the field survey lists 41 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally recorded during the field survey lists 41 specimens.", "negative_left": "PR-2's jar specimen tally recorded during the field survey lists 41 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-101", "id": "fast-43-diverse-252-101-base", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reviews a packet before routing to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, a recorded 15-minute dip-net sampling effort, habitat observations noting cobble substrate, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, a recorded 15-minute dip-net sampling effort, and habitat observations noting silty banks.", "PR-2's jar specimen tally recorded during the field survey lists 41 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.", "PR-3 is marked dry; the dated photo and habitat observations document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 4}, "source_family": "science-02", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the three-site scope, wet/dry policy language stays only in the unchanged question, and the focus evidence pair are complete factual sentences about PR-2's tallies; the counterfactual changes the count-sheet figure from 41 to 37, creating the discrepancy already anticipated by the surrounding 'possible one' caveats, so it remains internally consistent without contradicting other measurements or leaking any rule/label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a10": "supported", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a10": "refuted", "a11": "supported", "a12": "supported", "a13": "supported", "a14": "supported", "a15": "supported", "a16": "supported", "a17": "supported", "a18": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a10": "unknown"}, "remove_right": {"a10": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a10": "unknown"}, "negative_pair": {"a10": "refuted"}, "negative_sentence": {"a10": "unknown"}, "positive_pair": {"a10": "supported"}, "right": {"a10": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a factual relationship; the aggregate exclusion atoms quantify a single deficiency/discrepancy relationship over the explicit three-site scope rather than imposing final classifications. The focus is the factual agreement of two PR-2 specimen counts. Base and counter assignments differ only on that agreement and are both realizable: matching versus mismatching counts leave all other facts unchanged. Empty policy evidence is appropriate because the governing criteria, scope rule, dry-site exception, discrepancy counting, and routing instruction are already retained in the questions object; the state contributes case observations and incidental site identifiers rather than additional governing policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the three-site scope, satisfies every listed wet-site requirement for PR-1 and PR-2, documents PR-3's dry-site exception, excludes all other required-field deficiencies and record discrepancies, and excludes nonrequired irregularities. With PR-2 count agreement supported, level 4 follows.", "rule_index": 0, "sound": true}, {"reason": "Refutation of PR-2 count agreement establishes one specimen-count disagreement, which the question explicitly treats as one unresolved record discrepancy. The remaining atoms satisfy all other applicable requirements and exclude any additional deficiency or discrepancy, so exactly one unresolved issue remains and level 2 follows.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites."}, {"id": "a2", "statement": "PR-1 is a wet site."}, {"id": "a3", "statement": "PR-1 has recorded sampling effort."}, {"id": "a4", "statement": "PR-1 has a mappable site label."}, {"id": "a5", "statement": "PR-1's jar specimen count agrees with PR-1's count-sheet specimen count."}, {"id": "a6", "statement": "PR-1 has habitat observations."}, {"id": "a7", "statement": "PR-2 is a wet site."}, {"id": "a8", "statement": "PR-2 has recorded sampling effort."}, {"id": "a9", "statement": "PR-2 has a mappable site label."}, {"id": "a10", "statement": "PR-2's jar specimen count agrees with PR-2's count-sheet specimen count."}, {"id": "a11", "statement": "PR-2 has habitat observations."}, {"id": "a12", "statement": "PR-3 is a dry site."}, {"id": "a13", "statement": "The dry-site photo submitted for PR-3 bears a date."}, {"id": "a14", "statement": "The dry-site photo submitted for PR-3 documents the absence of flowing water at PR-3."}, {"id": "a15", "statement": "PR-3 has habitat observations."}, {"id": "a16", "statement": "No unresolved required-field deficiency exists across PR-1, PR-2, and PR-3 other than a possible deficiency represented by the PR-2 specimen-count comparison."}, {"id": "a17", "statement": "No unresolved record discrepancy exists across PR-1, PR-2, and PR-3 other than a possible discrepancy between PR-2's jar specimen count and count-sheet specimen count."}, {"id": "a18", "statement": "No nonrequired formatting, wording, or organizational irregularity remains in the packet."}], "base_state_json": "{\"context\":\"An ecology data manager reviews a packet before routing to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.\",\"evidence\":[\"The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.\",\"PR-1 is a wet site with a mappable label, a recorded 15-minute dip-net sampling effort, habitat observations noting cobble substrate, and 24 specimens agreeing on both jar and count sheet.\",\"PR-2 is a wet site with a mappable label, a recorded 15-minute dip-net sampling effort, and habitat observations noting silty banks.\",\"PR-2's jar specimen tally recorded during the field survey lists 41 specimens.\",\"PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens.\",\"PR-3 is marked dry; the dated photo and habitat observations document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.\",\"No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.\",\"No nonrequired formatting, wording, or organizational irregularity remains in the packet.\"]}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}], "focus_atom": "a10", "focus_evidence": [{"path": ["evidence", "3"], "text": "PR-2's jar specimen tally recorded during the field survey lists 41 specimens."}, {"path": ["evidence", "4"], "text": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}], "policy_evidence": [], "rules": [{"justification": "All three scoped sites satisfy their applicable requirements; both wet sites have agreeing specimen counts; the PR-3 dry-site exception is documented by a dated dry-site photo and habitat observations; and no required, record, or nonrequired irregularity remains.", "target": "4", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "supported"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}, {"justification": "The explicitly refuted PR-2 count agreement creates exactly one unresolved record discrepancy. All other applicable site requirements and the PR-3 dry-site exception are satisfied, and the exclusion atoms establish that no competing deficiency or discrepancy remains.", "target": "2", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}, {"atom_id": "a10", "state": "refuted"}, {"atom_id": "a11", "state": "supported"}, {"atom_id": "a12", "state": "supported"}, {"atom_id": "a13", "state": "supported"}, {"atom_id": "a14", "state": "supported"}, {"atom_id": "a15", "state": "supported"}, {"atom_id": "a16", "state": "supported"}, {"atom_id": "a17", "state": "supported"}, {"atom_id": "a18", "state": "supported"}]}]}, "verified_pair": {"left": "PR-2's jar specimen tally recorded during the field survey lists 41 specimens.", "negative_left": "PR-2's jar specimen tally recorded during the field survey lists 41 specimens.", "negative_right": "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "right": "PR-2's count-sheet specimen tally submitted with the packet lists 41 specimens."}, "verifier_independent_model": false}, "family": "fast-43-diverse-252-101", "id": "fast-43-diverse-252-101-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Unassessable: the route sheet is missing, no listed site can be matched to evidence, or no site has enough documentation to apply the wet-site requirements or dry-site exception.", "1 — Substantially incomplete: two or more unresolved required-field deficiencies or record discrepancies remain across the scoped sites, including any wet site with an unmappable label or missing sampling evidence.", "2 — Partially complete: exactly one unresolved required-field deficiency or record discrepancy remains after applying documented exceptions; every other scoped site satisfies its applicable requirements.", "3 — Complete with minor nonrequired irregularity: all required fields agree and every exception is documented, but one nonrequired formatting, wording, or organizational issue remains that does not obscure site identity or evidence.", "4 — Fully complete and corroborated: all scoped sites satisfy their applicable requirements, all wet-site records agree, every exception is fully documented, and no irregularities remain."], "instructions": "Assess only the three route-sheet sites. Every wet site requires sampling effort, a mappable site label, specimen counts agreeing across jar and count sheet, and habitat observations. Exception: a dry site needs none of the sampling or specimen fields if both a dated dry-site photo and habitat observations are supplied. Apply the exception before counting deficiencies. Use the ordered completeness levels below. A specimen-count disagreement is one unresolved record discrepancy and routes the packet to data review rather than sampling.", "type": "score"}}, "state": {"context": "An ecology data manager reviews a packet before routing to the taxonomic identifier. The watershed coordinator's route sheet lists Pine Run sites PR-1, PR-2, and PR-3.", "evidence": ["The route sheet lists exactly PR-1, PR-2, and PR-3 as its scoped sites.", "PR-1 is a wet site with a mappable label, a recorded 15-minute dip-net sampling effort, habitat observations noting cobble substrate, and 24 specimens agreeing on both jar and count sheet.", "PR-2 is a wet site with a mappable label, a recorded 15-minute dip-net sampling effort, and habitat observations noting silty banks.", "PR-2's jar specimen tally recorded during the field survey lists 41 specimens.", "PR-2's count-sheet specimen tally submitted with the packet lists 37 specimens.", "PR-3 is marked dry; the dated photo and habitat observations document no flowing water at PR-3, and no sampling effort, jar, or specimen count is recorded for PR-3.", "No unresolved required-field deficiency exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No unresolved record discrepancy exists across the three sites other than a possible one tied to the PR-2 specimen-count comparison.", "No nonrequired formatting, wording, or organizational irregularity remains in the packet."]}}, "method": "c2d", "provenance": {"source_id": "diverse-252", "source_is_synthetic": true, "source_sha256": "c2255d696ca5fc3c64ad917e828d430c3efe238a9cb409a52349f67d50114d99", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 2}, "source_family": "science-02", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy text and original question object with unchanged thresholds and entity/time bindings, evidence items are complete factual sentences without embedded rules or answers, and the counterfactual only alters the R6 count value (4050→4120) without duplicating or contradicting other measurements, remaining internally coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"09:10 sample overlay suggested a 1.6 px offset.\",\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward.\",\"Transmitted-light review found folds in 2 of 30 fields.\",\"Technical QC passed and replicate CV was measured at 8%.\",\"Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6.\",\"Comparison run R6 on slide batch 22 counted 4,050 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6."}, {"path": ["evidence", "5"], "text": "Comparison run R6 on slide batch 22 counted 4,050 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6.", "negative_left": "Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6.", "negative_right": "Comparison run R6 on slide batch 22 counted 4,120 cells.", "right": "Comparison run R6 on slide batch 22 counted 4,050 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-001", "id": "fast-43-diverse-259-001-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward.", "Transmitted-light review found folds in 2 of 30 fields.", "Technical QC passed and replicate CV was measured at 8%.", "Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6.", "Comparison run R6 on slide batch 22 counted 4,050 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy text and original question object with unchanged thresholds and entity/time bindings, evidence items are complete factual sentences without embedded rules or answers, and the counterfactual only alters the R6 count value (4050→4120) without duplicating or contradicting other measurements, remaining internally coherent.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"09:10 sample overlay suggested a 1.6 px offset.\",\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward.\",\"Transmitted-light review found folds in 2 of 30 fields.\",\"Technical QC passed and replicate CV was measured at 8%.\",\"Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6.\",\"Comparison run R6 on slide batch 22 counted 4,050 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6."}, {"path": ["evidence", "5"], "text": "Comparison run R6 on slide batch 22 counted 4,050 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6.", "negative_left": "Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6.", "negative_right": "Comparison run R6 on slide batch 22 counted 4,120 cells.", "right": "Comparison run R6 on slide batch 22 counted 4,050 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-001", "id": "fast-43-diverse-259-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward.", "Transmitted-light review found folds in 2 of 30 fields.", "Technical QC passed and replicate CV was measured at 8%.", "Segmentation rerun R7 on slide batch 22 counted 4,300 cells, using the same raw image dataset as comparison run R6.", "Comparison run R6 on slide batch 22 counted 4,120 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and original questions, use factual evidence sentences, and the counterfactual (5,120 vs 5,000, a 2.4% change) coherently shifts the outcome from analysis routing to acceptance without contradicting other measurements or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and calibration passed.\",\"Transmitted-light review found folds in 2 of 30 fields.\",\"Technical QC for the current run passed, and replicate CV was measured at 8%.\",\"On March 4th, the segmentation rerun labeled Run-B7 counted 5,320 cells using the same raw image dataset as comparison Run-B6.\",\"Comparison Run-B6, from the same raw image dataset, counted 5,000 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "3"], "text": "On March 4th, the segmentation rerun labeled Run-B7 counted 5,320 cells using the same raw image dataset as comparison Run-B6."}, {"path": ["evidence", "4"], "text": "Comparison Run-B6, from the same raw image dataset, counted 5,000 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "On March 4th, the segmentation rerun labeled Run-B7 counted 5,320 cells using the same raw image dataset as comparison Run-B6.", "negative_left": "On March 4th, the segmentation rerun labeled Run-B7 counted 5,120 cells using the same raw image dataset as comparison Run-B6.", "negative_right": "Comparison Run-B6, from the same raw image dataset, counted 5,000 cells.", "right": "Comparison Run-B6, from the same raw image dataset, counted 5,000 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-002", "id": "fast-43-diverse-259-002-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and calibration passed.", "Transmitted-light review found folds in 2 of 30 fields.", "Technical QC for the current run passed, and replicate CV was measured at 8%.", "On March 4th, the segmentation rerun labeled Run-B7 counted 5,320 cells using the same raw image dataset as comparison Run-B6.", "Comparison Run-B6, from the same raw image dataset, counted 5,000 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and original questions, use factual evidence sentences, and the counterfactual (5,120 vs 5,000, a 2.4% change) coherently shifts the outcome from analysis routing to acceptance without contradicting other measurements or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and calibration passed.\",\"Transmitted-light review found folds in 2 of 30 fields.\",\"Technical QC for the current run passed, and replicate CV was measured at 8%.\",\"On March 4th, the segmentation rerun labeled Run-B7 counted 5,320 cells using the same raw image dataset as comparison Run-B6.\",\"Comparison Run-B6, from the same raw image dataset, counted 5,000 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "3"], "text": "On March 4th, the segmentation rerun labeled Run-B7 counted 5,320 cells using the same raw image dataset as comparison Run-B6."}, {"path": ["evidence", "4"], "text": "Comparison Run-B6, from the same raw image dataset, counted 5,000 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "On March 4th, the segmentation rerun labeled Run-B7 counted 5,320 cells using the same raw image dataset as comparison Run-B6.", "negative_left": "On March 4th, the segmentation rerun labeled Run-B7 counted 5,120 cells using the same raw image dataset as comparison Run-B6.", "negative_right": "Comparison Run-B6, from the same raw image dataset, counted 5,000 cells.", "right": "Comparison Run-B6, from the same raw image dataset, counted 5,000 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-002", "id": "fast-43-diverse-259-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and calibration passed.", "Transmitted-light review found folds in 2 of 30 fields.", "Technical QC for the current run passed, and replicate CV was measured at 8%.", "On March 4th, the segmentation rerun labeled Run-B7 counted 5,120 cells using the same raw image dataset as comparison Run-B6.", "Comparison Run-B6, from the same raw image dataset, counted 5,000 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all four routing thresholds and only vary the R7 count (4120 vs 3960) against the fixed C7=3900 baseline, producing a coherent >5% vs <5% change without contradicting other measurements; the two evidence spans are plain factual count statements, not policy text; no gold label, rule table, or output instruction is embedded, and question entity/path/time bindings are unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"Case note: Bead test at 11:30 measured 0.4 px offset after recalibration, well under the 1.0 px threshold, and logs confirm the optics were unchanged afterward, so the offset finding is refuted. Transmitted-light review found folds in only 2 of 30 fields (about 6.7%), refuting the fixed slide defect threshold. Current calibration passes technical validation, and technical QC also passes. Replicate CV was measured at 8%, refuting the exceeds-15% condition. Segmentation rerun R7 on slide batch 22 counted 4120 cells using the same raw image dataset as comparison run C7. Comparison run C7 counted 3900 cells on that same raw image dataset. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies. Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "Segmentation rerun R7 on slide batch 22 counted 4120 cells using the same raw image dataset as comparison run C7."}, {"path": ["context"], "text": "Comparison run C7 counted 3900 cells on that same raw image dataset."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch 22 counted 4120 cells using the same raw image dataset as comparison run C7.", "negative_left": "Segmentation rerun R7 on slide batch 22 counted 3960 cells using the same raw image dataset as comparison run C7.", "negative_right": "Comparison run C7 counted 3900 cells on that same raw image dataset.", "right": "Comparison run C7 counted 3900 cells on that same raw image dataset."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-004", "id": "fast-43-diverse-259-004-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Case note: Bead test at 11:30 measured 0.4 px offset after recalibration, well under the 1.0 px threshold, and logs confirm the optics were unchanged afterward, so the offset finding is refuted. Transmitted-light review found folds in only 2 of 30 fields (about 6.7%), refuting the fixed slide defect threshold. Current calibration passes technical validation, and technical QC also passes. Replicate CV was measured at 8%, refuting the exceeds-15% condition. Segmentation rerun R7 on slide batch 22 counted 4120 cells using the same raw image dataset as comparison run C7. Comparison run C7 counted 3900 cells on that same raw image dataset. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies. Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all four routing thresholds and only vary the R7 count (4120 vs 3960) against the fixed C7=3900 baseline, producing a coherent >5% vs <5% change without contradicting other measurements; the two evidence spans are plain factual count statements, not policy text; no gold label, rule table, or output instruction is embedded, and question entity/path/time bindings are unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"Case note: Bead test at 11:30 measured 0.4 px offset after recalibration, well under the 1.0 px threshold, and logs confirm the optics were unchanged afterward, so the offset finding is refuted. Transmitted-light review found folds in only 2 of 30 fields (about 6.7%), refuting the fixed slide defect threshold. Current calibration passes technical validation, and technical QC also passes. Replicate CV was measured at 8%, refuting the exceeds-15% condition. Segmentation rerun R7 on slide batch 22 counted 4120 cells using the same raw image dataset as comparison run C7. Comparison run C7 counted 3900 cells on that same raw image dataset. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies. Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "Segmentation rerun R7 on slide batch 22 counted 4120 cells using the same raw image dataset as comparison run C7."}, {"path": ["context"], "text": "Comparison run C7 counted 3900 cells on that same raw image dataset."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch 22 counted 4120 cells using the same raw image dataset as comparison run C7.", "negative_left": "Segmentation rerun R7 on slide batch 22 counted 3960 cells using the same raw image dataset as comparison run C7.", "negative_right": "Comparison run C7 counted 3900 cells on that same raw image dataset.", "right": "Comparison run C7 counted 3900 cells on that same raw image dataset."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-004", "id": "fast-43-diverse-259-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Case note: Bead test at 11:30 measured 0.4 px offset after recalibration, well under the 1.0 px threshold, and logs confirm the optics were unchanged afterward, so the offset finding is refuted. Transmitted-light review found folds in only 2 of 30 fields (about 6.7%), refuting the fixed slide defect threshold. Current calibration passes technical validation, and technical QC also passes. Replicate CV was measured at 8%, refuting the exceeds-15% condition. Segmentation rerun R7 on slide batch 22 counted 3960 cells using the same raw image dataset as comparison run C7. Comparison run C7 counted 3900 cells on that same raw image dataset. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies. Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and question bindings, split the original combined measurement into clear factual sentences, and the counterfactual merely changes the rerun count (842→812) without contradicting other evidence or embedding any answer cues.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"evidence\": [\"09:10 sample overlay suggested a 1.6 px offset, superseded by later calibration.\", \"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.\", \"Transmitted-light review found folds in 2 of 30 fields, well under the 10% threshold.\", \"Technical QC for the current run passed all checks.\", \"Replicate coefficient of variation for the most recent relevant replicate set was 8%, below 15%.\", \"Segmentation rerun R7 on slide batch S12 counted 842 cells, using the same raw image dataset as comparison run R6.\", \"Comparison run R6 on slide batch S12 counted 790 cells.\"], \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "5"], "text": "Segmentation rerun R7 on slide batch S12 counted 842 cells, using the same raw image dataset as comparison run R6."}, {"path": ["evidence", "6"], "text": "Comparison run R6 on slide batch S12 counted 790 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch S12 counted 842 cells, using the same raw image dataset as comparison run R6.", "negative_left": "Segmentation rerun R7 on slide batch S12 counted 812 cells, using the same raw image dataset as comparison run R6.", "negative_right": "Comparison run R6 on slide batch S12 counted 790 cells.", "right": "Comparison run R6 on slide batch S12 counted 790 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-006", "id": "fast-43-diverse-259-006-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset, superseded by later calibration.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, well under the 10% threshold.", "Technical QC for the current run passed all checks.", "Replicate coefficient of variation for the most recent relevant replicate set was 8%, below 15%.", "Segmentation rerun R7 on slide batch S12 counted 842 cells, using the same raw image dataset as comparison run R6.", "Comparison run R6 on slide batch S12 counted 790 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and question bindings, split the original combined measurement into clear factual sentences, and the counterfactual merely changes the rerun count (842→812) without contradicting other evidence or embedding any answer cues.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"evidence\": [\"09:10 sample overlay suggested a 1.6 px offset, superseded by later calibration.\", \"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.\", \"Transmitted-light review found folds in 2 of 30 fields, well under the 10% threshold.\", \"Technical QC for the current run passed all checks.\", \"Replicate coefficient of variation for the most recent relevant replicate set was 8%, below 15%.\", \"Segmentation rerun R7 on slide batch S12 counted 842 cells, using the same raw image dataset as comparison run R6.\", \"Comparison run R6 on slide batch S12 counted 790 cells.\"], \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "5"], "text": "Segmentation rerun R7 on slide batch S12 counted 842 cells, using the same raw image dataset as comparison run R6."}, {"path": ["evidence", "6"], "text": "Comparison run R6 on slide batch S12 counted 790 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch S12 counted 842 cells, using the same raw image dataset as comparison run R6.", "negative_left": "Segmentation rerun R7 on slide batch S12 counted 812 cells, using the same raw image dataset as comparison run R6.", "negative_right": "Comparison run R6 on slide batch S12 counted 790 cells.", "right": "Comparison run R6 on slide batch S12 counted 790 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-006", "id": "fast-43-diverse-259-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset, superseded by later calibration.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, well under the 10% threshold.", "Technical QC for the current run passed all checks.", "Replicate coefficient of variation for the most recent relevant replicate set was 8%, below 15%.", "Segmentation rerun R7 on slide batch S12 counted 812 cells, using the same raw image dataset as comparison run R6.", "Comparison run R6 on slide batch S12 counted 790 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy verbatim alongside the unchanged question object, preserve the S-204 entity and 'latest evidence' framing, and the two evidence sentences are plain factual count statements ('...counted 8,420 cells...' and 'Comparison run R-11...counted 8,000 cells.') with only the first altered to 8,190 in the counterfactual, a single coherent numeric change (2.4% vs 5.25% delta) that does not contradict any other stated fact or embed a label, rule table, or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"Case note: Slide batch S-204 assay under review. 09:10 sample overlay suggested a 1.6 px offset. 11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward. Transmitted-light review found folds in 2 of 30 fields. Current technical QC passed. Replicate CV was measured at 8%, well under threshold. The most recent segmentation rerun on slide batch S-204 counted 8,420 cells using the same raw image dataset as comparison run R-11. Comparison run R-11 on slide batch S-204 counted 8,000 cells. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "The most recent segmentation rerun on slide batch S-204 counted 8,420 cells using the same raw image dataset as comparison run R-11."}, {"path": ["context"], "text": "Comparison run R-11 on slide batch S-204 counted 8,000 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The most recent segmentation rerun on slide batch S-204 counted 8,420 cells using the same raw image dataset as comparison run R-11.", "negative_left": "The most recent segmentation rerun on slide batch S-204 counted 8,190 cells using the same raw image dataset as comparison run R-11.", "negative_right": "Comparison run R-11 on slide batch S-204 counted 8,000 cells.", "right": "Comparison run R-11 on slide batch S-204 counted 8,000 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-009", "id": "fast-43-diverse-259-009-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Case note: Slide batch S-204 assay under review. 09:10 sample overlay suggested a 1.6 px offset. 11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward. Transmitted-light review found folds in 2 of 30 fields. Current technical QC passed. Replicate CV was measured at 8%, well under threshold. The most recent segmentation rerun on slide batch S-204 counted 8,420 cells using the same raw image dataset as comparison run R-11. Comparison run R-11 on slide batch S-204 counted 8,000 cells. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy verbatim alongside the unchanged question object, preserve the S-204 entity and 'latest evidence' framing, and the two evidence sentences are plain factual count statements ('...counted 8,420 cells...' and 'Comparison run R-11...counted 8,000 cells.') with only the first altered to 8,190 in the counterfactual, a single coherent numeric change (2.4% vs 5.25% delta) that does not contradict any other stated fact or embed a label, rule table, or instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"Case note: Slide batch S-204 assay under review. 09:10 sample overlay suggested a 1.6 px offset. 11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward. Transmitted-light review found folds in 2 of 30 fields. Current technical QC passed. Replicate CV was measured at 8%, well under threshold. The most recent segmentation rerun on slide batch S-204 counted 8,420 cells using the same raw image dataset as comparison run R-11. Comparison run R-11 on slide batch S-204 counted 8,000 cells. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "The most recent segmentation rerun on slide batch S-204 counted 8,420 cells using the same raw image dataset as comparison run R-11."}, {"path": ["context"], "text": "Comparison run R-11 on slide batch S-204 counted 8,000 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The most recent segmentation rerun on slide batch S-204 counted 8,420 cells using the same raw image dataset as comparison run R-11.", "negative_left": "The most recent segmentation rerun on slide batch S-204 counted 8,190 cells using the same raw image dataset as comparison run R-11.", "negative_right": "Comparison run R-11 on slide batch S-204 counted 8,000 cells.", "right": "Comparison run R-11 on slide batch S-204 counted 8,000 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-009", "id": "fast-43-diverse-259-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Case note: Slide batch S-204 assay under review. 09:10 sample overlay suggested a 1.6 px offset. 11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward. Transmitted-light review found folds in 2 of 30 fields. Current technical QC passed. Replicate CV was measured at 8%, well under threshold. The most recent segmentation rerun on slide batch S-204 counted 8,190 cells using the same raw image dataset as comparison run R-11. Comparison run R-11 on slide batch S-204 counted 8,000 cells. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and question, preserve the same entities/paths/times, use two complete factual sentences as evidence, and the counterfactual's 905 vs 890 count is a coherent modification that still respects the stated policy without contradicting other measurements or leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"Case note (assay QC review): Microscopy technician confirmed the 11:30 bead recalibration measured 0.4 px offset, within tolerance, and logs confirm unchanged optics afterward. Transmitted-light review found folds in 2 of 30 fields, well under 10%. Current calibration passes and technical QC passes. Segmentation rerun R-118 on slide batch B-22 counted 942 cells using the identical raw image dataset as its comparison run C-118. Comparison run C-118 on slide batch B-22 counted 890 cells. The assay scientist noted replicate CV of 8%, below the 15% threshold. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "Segmentation rerun R-118 on slide batch B-22 counted 942 cells using the identical raw image dataset as its comparison run C-118."}, {"path": ["context"], "text": "Comparison run C-118 on slide batch B-22 counted 890 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R-118 on slide batch B-22 counted 942 cells using the identical raw image dataset as its comparison run C-118.", "negative_left": "Segmentation rerun R-118 on slide batch B-22 counted 905 cells using the identical raw image dataset as its comparison run C-118.", "negative_right": "Comparison run C-118 on slide batch B-22 counted 890 cells.", "right": "Comparison run C-118 on slide batch B-22 counted 890 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-017", "id": "fast-43-diverse-259-017-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Case note (assay QC review): Microscopy technician confirmed the 11:30 bead recalibration measured 0.4 px offset, within tolerance, and logs confirm unchanged optics afterward. Transmitted-light review found folds in 2 of 30 fields, well under 10%. Current calibration passes and technical QC passes. Segmentation rerun R-118 on slide batch B-22 counted 942 cells using the identical raw image dataset as its comparison run C-118. Comparison run C-118 on slide batch B-22 counted 890 cells. The assay scientist noted replicate CV of 8%, below the 15% threshold. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and question, preserve the same entities/paths/times, use two complete factual sentences as evidence, and the counterfactual's 905 vs 890 count is a coherent modification that still respects the stated policy without contradicting other measurements or leaking an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"Case note (assay QC review): Microscopy technician confirmed the 11:30 bead recalibration measured 0.4 px offset, within tolerance, and logs confirm unchanged optics afterward. Transmitted-light review found folds in 2 of 30 fields, well under 10%. Current calibration passes and technical QC passes. Segmentation rerun R-118 on slide batch B-22 counted 942 cells using the identical raw image dataset as its comparison run C-118. Comparison run C-118 on slide batch B-22 counted 890 cells. The assay scientist noted replicate CV of 8%, below the 15% threshold. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["context"], "text": "Segmentation rerun R-118 on slide batch B-22 counted 942 cells using the identical raw image dataset as its comparison run C-118."}, {"path": ["context"], "text": "Comparison run C-118 on slide batch B-22 counted 890 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R-118 on slide batch B-22 counted 942 cells using the identical raw image dataset as its comparison run C-118.", "negative_left": "Segmentation rerun R-118 on slide batch B-22 counted 905 cells using the identical raw image dataset as its comparison run C-118.", "negative_right": "Comparison run C-118 on slide batch B-22 counted 890 cells.", "right": "Comparison run C-118 on slide batch B-22 counted 890 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-017", "id": "fast-43-diverse-259-017-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Case note (assay QC review): Microscopy technician confirmed the 11:30 bead recalibration measured 0.4 px offset, within tolerance, and logs confirm unchanged optics afterward. Transmitted-light review found folds in 2 of 30 fields, well under 10%. Current calibration passes and technical QC passes. Segmentation rerun R-118 on slide batch B-22 counted 905 cells using the identical raw image dataset as its comparison run C-118. Comparison run C-118 on slide batch B-22 counted 890 cells. The assay scientist noted replicate CV of 8%, below the 15% threshold. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy text alongside the unchanged questions object, preserving all thresholds and entities/paths/times; the two focus evidence items are complete factual sentences about R-88 and C-88 counts with no policy language; the counterfactual only changes the R-88 count (9540→9030) while keeping C-88 at 9000, remaining internally consistent with no duplicate or contradictory measurements; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"evidence\": [\"09:10 sample overlay suggested a 1.6 px offset.\", \"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward.\", \"Transmitted-light review found folds in 2 of 30 fields.\", \"Technical QC passed and current calibration passed.\", \"Replicate CV for the most recent relevant replicate set was 8%, within the 15% threshold.\", \"In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,540 cells.\", \"The comparison run C-88 on that same raw image dataset reported 9,000 cells.\"], \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "5"], "text": "In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,540 cells."}, {"path": ["evidence", "6"], "text": "The comparison run C-88 on that same raw image dataset reported 9,000 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,540 cells.", "negative_left": "In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,030 cells.", "negative_right": "The comparison run C-88 on that same raw image dataset reported 9,000 cells.", "right": "The comparison run C-88 on that same raw image dataset reported 9,000 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-019", "id": "fast-43-diverse-259-019-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward.", "Transmitted-light review found folds in 2 of 30 fields.", "Technical QC passed and current calibration passed.", "Replicate CV for the most recent relevant replicate set was 8%, within the 15% threshold.", "In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,540 cells.", "The comparison run C-88 on that same raw image dataset reported 9,000 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy text alongside the unchanged questions object, preserving all thresholds and entities/paths/times; the two focus evidence items are complete factual sentences about R-88 and C-88 counts with no policy language; the counterfactual only changes the R-88 count (9540→9030) while keeping C-88 at 9000, remaining internally consistent with no duplicate or contradictory measurements; neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"evidence\": [\"09:10 sample overlay suggested a 1.6 px offset.\", \"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward.\", \"Transmitted-light review found folds in 2 of 30 fields.\", \"Technical QC passed and current calibration passed.\", \"Replicate CV for the most recent relevant replicate set was 8%, within the 15% threshold.\", \"In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,540 cells.\", \"The comparison run C-88 on that same raw image dataset reported 9,000 cells.\"], \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "5"], "text": "In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,540 cells."}, {"path": ["evidence", "6"], "text": "The comparison run C-88 on that same raw image dataset reported 9,000 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,540 cells.", "negative_left": "In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,030 cells.", "negative_right": "The comparison run C-88 on that same raw image dataset reported 9,000 cells.", "right": "The comparison run C-88 on that same raw image dataset reported 9,000 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-019", "id": "fast-43-diverse-259-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward.", "Transmitted-light review found folds in 2 of 30 fields.", "Technical QC passed and current calibration passed.", "Replicate CV for the most recent relevant replicate set was 8%, within the 15% threshold.", "In segmentation rerun R-88, run on the same raw image dataset as its comparison run C-88, the automated cell counter reported 9,030 cells.", "The comparison run C-88 on that same raw image dataset reported 9,000 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing thresholds and evidence structure, differing only in the rerun cell count (8,400 vs 8,150) which changes the percent-change outcome without introducing contradictions, extra rules, or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"Case note: Slide batch S-22 underwent standard three-channel acquisition, preparation, segmentation, and replicate scoring review. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"09:10 sample overlay suggested a 1.6 px offset.\",\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and current calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields, well under the 10% threshold.\",\"Technical QC for the run passed, and replicate CV was measured at 8%, below the 15% threshold.\",\"The most recent segmentation rerun for slide batch S-22 counted 8,400 cells, using the same raw image dataset as its comparison run.\",\"The comparison run for slide batch S-22 counted 7,900 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "The most recent segmentation rerun for slide batch S-22 counted 8,400 cells, using the same raw image dataset as its comparison run."}, {"path": ["evidence", "5"], "text": "The comparison run for slide batch S-22 counted 7,900 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The most recent segmentation rerun for slide batch S-22 counted 8,400 cells, using the same raw image dataset as its comparison run.", "negative_left": "The most recent segmentation rerun for slide batch S-22 counted 8,150 cells, using the same raw image dataset as its comparison run.", "negative_right": "The comparison run for slide batch S-22 counted 7,900 cells.", "right": "The comparison run for slide batch S-22 counted 7,900 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-021", "id": "fast-43-diverse-259-021-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Case note: Slide batch S-22 underwent standard three-channel acquisition, preparation, segmentation, and replicate scoring review. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, well under the 10% threshold.", "Technical QC for the run passed, and replicate CV was measured at 8%, below the 15% threshold.", "The most recent segmentation rerun for slide batch S-22 counted 8,400 cells, using the same raw image dataset as its comparison run.", "The comparison run for slide batch S-22 counted 7,900 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original routing thresholds and evidence structure, differing only in the rerun cell count (8,400 vs 8,150) which changes the percent-change outcome without introducing contradictions, extra rules, or answer hints.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"Case note: Slide batch S-22 underwent standard three-channel acquisition, preparation, segmentation, and replicate scoring review. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"09:10 sample overlay suggested a 1.6 px offset.\",\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and current calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields, well under the 10% threshold.\",\"Technical QC for the run passed, and replicate CV was measured at 8%, below the 15% threshold.\",\"The most recent segmentation rerun for slide batch S-22 counted 8,400 cells, using the same raw image dataset as its comparison run.\",\"The comparison run for slide batch S-22 counted 7,900 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "The most recent segmentation rerun for slide batch S-22 counted 8,400 cells, using the same raw image dataset as its comparison run."}, {"path": ["evidence", "5"], "text": "The comparison run for slide batch S-22 counted 7,900 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "The most recent segmentation rerun for slide batch S-22 counted 8,400 cells, using the same raw image dataset as its comparison run.", "negative_left": "The most recent segmentation rerun for slide batch S-22 counted 8,150 cells, using the same raw image dataset as its comparison run.", "negative_right": "The comparison run for slide batch S-22 counted 7,900 cells.", "right": "The comparison run for slide batch S-22 counted 7,900 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-021", "id": "fast-43-diverse-259-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "Case note: Slide batch S-22 underwent standard three-channel acquisition, preparation, segmentation, and replicate scoring review. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, well under the 10% threshold.", "Technical QC for the run passed, and replicate CV was measured at 8%, below the 15% threshold.", "The most recent segmentation rerun for slide batch S-22 counted 8,150 cells, using the same raw image dataset as its comparison run.", "The comparison run for slide batch S-22 counted 7,900 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy and unchanged question bindings, use complete factual evidence sentences without embedding policy tables or answer labels, and the counterfactual only changes the C7 count (750→785) coherently without contradicting other retained measurements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px, confirming channel offset is well under 1.0 px; logs confirm unchanged optics afterward, so current calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields, a proportion well under 10%.\",\"Technical QC for the current run passed all checks.\",\"Replicate coefficient of variation for the most recent relevant replicate set was 8%, under the 15% threshold.\",\"Segmentation rerun R7 recorded 812 cells on Slide-114.\",\"Comparison run C7, using the same raw image dataset, recorded 750 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun R7 recorded 812 cells on Slide-114."}, {"path": ["evidence", "5"], "text": "Comparison run C7, using the same raw image dataset, recorded 750 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 recorded 812 cells on Slide-114.", "negative_left": "Segmentation rerun R7 recorded 812 cells on Slide-114.", "negative_right": "Comparison run C7, using the same raw image dataset, recorded 785 cells.", "right": "Comparison run C7, using the same raw image dataset, recorded 750 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-022", "id": "fast-43-diverse-259-022-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px, confirming channel offset is well under 1.0 px; logs confirm unchanged optics afterward, so current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, a proportion well under 10%.", "Technical QC for the current run passed all checks.", "Replicate coefficient of variation for the most recent relevant replicate set was 8%, under the 15% threshold.", "Segmentation rerun R7 recorded 812 cells on Slide-114.", "Comparison run C7, using the same raw image dataset, recorded 750 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy and unchanged question bindings, use complete factual evidence sentences without embedding policy tables or answer labels, and the counterfactual only changes the C7 count (750→785) coherently without contradicting other retained measurements.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px, confirming channel offset is well under 1.0 px; logs confirm unchanged optics afterward, so current calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields, a proportion well under 10%.\",\"Technical QC for the current run passed all checks.\",\"Replicate coefficient of variation for the most recent relevant replicate set was 8%, under the 15% threshold.\",\"Segmentation rerun R7 recorded 812 cells on Slide-114.\",\"Comparison run C7, using the same raw image dataset, recorded 750 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun R7 recorded 812 cells on Slide-114."}, {"path": ["evidence", "5"], "text": "Comparison run C7, using the same raw image dataset, recorded 750 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 recorded 812 cells on Slide-114.", "negative_left": "Segmentation rerun R7 recorded 812 cells on Slide-114.", "negative_right": "Comparison run C7, using the same raw image dataset, recorded 785 cells.", "right": "Comparison run C7, using the same raw image dataset, recorded 750 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-022", "id": "fast-43-diverse-259-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px, confirming channel offset is well under 1.0 px; logs confirm unchanged optics afterward, so current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, a proportion well under 10%.", "Technical QC for the current run passed all checks.", "Replicate coefficient of variation for the most recent relevant replicate set was 8%, under the 15% threshold.", "Segmentation rerun R7 recorded 812 cells on Slide-114.", "Comparison run C7, using the same raw image dataset, recorded 785 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy and original entity/path bindings while supplying consistent factual evidence; the base shows a 5.76% count change triggering analysis, the counterfactual shows a 1.7% change staying under threshold, and neither reveals a label or instruction, so all criteria are satisfied.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay for slide batch SB-14; an image analyst segmented nuclei, and the assay scientist reviewed replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields, well under the 10% preparation threshold.\",\"Technical QC for the session passed all checks.\",\"Replicate CV for the latest relevant run was measured at 8%, below the 15% threshold.\",\"Segmentation rerun R-207 on slide batch SB-14 counted 3,120 cells using the same raw image dataset as comparison run R-206.\",\"Comparison run R-206 on slide batch SB-14 counted 2,950 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun R-207 on slide batch SB-14 counted 3,120 cells using the same raw image dataset as comparison run R-206."}, {"path": ["evidence", "5"], "text": "Comparison run R-206 on slide batch SB-14 counted 2,950 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R-207 on slide batch SB-14 counted 3,120 cells using the same raw image dataset as comparison run R-206.", "negative_left": "Segmentation rerun R-207 on slide batch SB-14 counted 3,000 cells using the same raw image dataset as comparison run R-206.", "negative_right": "Comparison run R-206 on slide batch SB-14 counted 2,950 cells.", "right": "Comparison run R-206 on slide batch SB-14 counted 2,950 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-023", "id": "fast-43-diverse-259-023-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay for slide batch SB-14; an image analyst segmented nuclei, and the assay scientist reviewed replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, well under the 10% preparation threshold.", "Technical QC for the session passed all checks.", "Replicate CV for the latest relevant run was measured at 8%, below the 15% threshold.", "Segmentation rerun R-207 on slide batch SB-14 counted 3,120 cells using the same raw image dataset as comparison run R-206.", "Comparison run R-206 on slide batch SB-14 counted 2,950 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing policy and original entity/path bindings while supplying consistent factual evidence; the base shows a 5.76% count change triggering analysis, the counterfactual shows a 1.7% change staying under threshold, and neither reveals a label or instruction, so all criteria are satisfied.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay for slide batch SB-14; an image analyst segmented nuclei, and the assay scientist reviewed replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields, well under the 10% preparation threshold.\",\"Technical QC for the session passed all checks.\",\"Replicate CV for the latest relevant run was measured at 8%, below the 15% threshold.\",\"Segmentation rerun R-207 on slide batch SB-14 counted 3,120 cells using the same raw image dataset as comparison run R-206.\",\"Comparison run R-206 on slide batch SB-14 counted 2,950 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun R-207 on slide batch SB-14 counted 3,120 cells using the same raw image dataset as comparison run R-206."}, {"path": ["evidence", "5"], "text": "Comparison run R-206 on slide batch SB-14 counted 2,950 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R-207 on slide batch SB-14 counted 3,120 cells using the same raw image dataset as comparison run R-206.", "negative_left": "Segmentation rerun R-207 on slide batch SB-14 counted 3,000 cells using the same raw image dataset as comparison run R-206.", "negative_right": "Comparison run R-206 on slide batch SB-14 counted 2,950 cells.", "right": "Comparison run R-206 on slide batch SB-14 counted 2,950 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-023", "id": "fast-43-diverse-259-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay for slide batch SB-14; an image analyst segmented nuclei, and the assay scientist reviewed replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, well under the 10% preparation threshold.", "Technical QC for the session passed all checks.", "Replicate CV for the latest relevant run was measured at 8%, below the 15% threshold.", "Segmentation rerun R-207 on slide batch SB-14 counted 3,000 cells using the same raw image dataset as comparison run R-206.", "Comparison run R-206 on slide batch SB-14 counted 2,950 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the full routing policy and question bindings verbatim; each evidence item is a complete factual statement, not an instruction; the counterfactual changes only the R7 count (905 vs 940) yielding a different but internally consistent percentage change without contradicting other measurements; neither context states or implies the final disposition.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"09:10 sample overlay suggested a 1.6 px offset.\",\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward; calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields.\",\"Technical QC passed at 13:00.\",\"Replicate CV for the latest relevant assay was 8%, well under 15%.\",\"Segmentation rerun R7 on slide batch 22 counted 940 cells using the same raw image dataset as comparison run R6.\",\"Comparison run R6 counted 880 cells on that same raw image dataset.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "5"], "text": "Segmentation rerun R7 on slide batch 22 counted 940 cells using the same raw image dataset as comparison run R6."}, {"path": ["evidence", "6"], "text": "Comparison run R6 counted 880 cells on that same raw image dataset."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch 22 counted 940 cells using the same raw image dataset as comparison run R6.", "negative_left": "Segmentation rerun R7 on slide batch 22 counted 905 cells using the same raw image dataset as comparison run R6.", "negative_right": "Comparison run R6 counted 880 cells on that same raw image dataset.", "right": "Comparison run R6 counted 880 cells on that same raw image dataset."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-024", "id": "fast-43-diverse-259-024-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward; calibration passes.", "Transmitted-light review found folds in 2 of 30 fields.", "Technical QC passed at 13:00.", "Replicate CV for the latest relevant assay was 8%, well under 15%.", "Segmentation rerun R7 on slide batch 22 counted 940 cells using the same raw image dataset as comparison run R6.", "Comparison run R6 counted 880 cells on that same raw image dataset."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the full routing policy and question bindings verbatim; each evidence item is a complete factual statement, not an instruction; the counterfactual changes only the R7 count (905 vs 940) yielding a different but internally consistent percentage change without contradicting other measurements; neither context states or implies the final disposition.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"09:10 sample overlay suggested a 1.6 px offset.\",\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward; calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields.\",\"Technical QC passed at 13:00.\",\"Replicate CV for the latest relevant assay was 8%, well under 15%.\",\"Segmentation rerun R7 on slide batch 22 counted 940 cells using the same raw image dataset as comparison run R6.\",\"Comparison run R6 counted 880 cells on that same raw image dataset.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "5"], "text": "Segmentation rerun R7 on slide batch 22 counted 940 cells using the same raw image dataset as comparison run R6."}, {"path": ["evidence", "6"], "text": "Comparison run R6 counted 880 cells on that same raw image dataset."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch 22 counted 940 cells using the same raw image dataset as comparison run R6.", "negative_left": "Segmentation rerun R7 on slide batch 22 counted 905 cells using the same raw image dataset as comparison run R6.", "negative_right": "Comparison run R6 counted 880 cells on that same raw image dataset.", "right": "Comparison run R6 counted 880 cells on that same raw image dataset."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-024", "id": "fast-43-diverse-259-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:10 sample overlay suggested a 1.6 px offset.", "11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward; calibration passes.", "Transmitted-light review found folds in 2 of 30 fields.", "Technical QC passed at 13:00.", "Replicate CV for the latest relevant assay was 8%, well under 15%.", "Segmentation rerun R7 on slide batch 22 counted 905 cells using the same raw image dataset as comparison run R6.", "Comparison run R6 counted 880 cells on that same raw image dataset."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy text and unchanged question object, preserving thresholds and entity/time bindings; only the R7 segmentation count differs (4,180 vs 4,050) while R6 stays at 3,950, a single coherent factual edit without contradicting other measurements; both focus sentences are plain factual statements, not policy or rule text; no evidence reveals or hints at the intended disposition.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields, well under 10%.\",\"Technical QC for the current run passes all internal checks.\",\"Replicate CV for the most recent relevant set was 8%, below the 15% threshold.\",\"Segmentation rerun R7 on slide batch 22 counted 4,180 cells using the same raw image dataset as comparison run R6.\",\"Comparison run R6 on slide batch 22 counted 3,950 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun R7 on slide batch 22 counted 4,180 cells using the same raw image dataset as comparison run R6."}, {"path": ["evidence", "5"], "text": "Comparison run R6 on slide batch 22 counted 3,950 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch 22 counted 4,180 cells using the same raw image dataset as comparison run R6.", "negative_left": "Segmentation rerun R7 on slide batch 22 counted 4,050 cells using the same raw image dataset as comparison run R6.", "negative_right": "Comparison run R6 on slide batch 22 counted 3,950 cells.", "right": "Comparison run R6 on slide batch 22 counted 3,950 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-031", "id": "fast-43-diverse-259-031-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, well under 10%.", "Technical QC for the current run passes all internal checks.", "Replicate CV for the most recent relevant set was 8%, below the 15% threshold.", "Segmentation rerun R7 on slide batch 22 counted 4,180 cells using the same raw image dataset as comparison run R6.", "Comparison run R6 on slide batch 22 counted 3,950 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy text and unchanged question object, preserving thresholds and entity/time bindings; only the R7 segmentation count differs (4,180 vs 4,050) while R6 stays at 3,950, a single coherent factual edit without contradicting other measurements; both focus sentences are plain factual statements, not policy or rule text; no evidence reveals or hints at the intended disposition.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.\",\"Transmitted-light review found folds in 2 of 30 fields, well under 10%.\",\"Technical QC for the current run passes all internal checks.\",\"Replicate CV for the most recent relevant set was 8%, below the 15% threshold.\",\"Segmentation rerun R7 on slide batch 22 counted 4,180 cells using the same raw image dataset as comparison run R6.\",\"Comparison run R6 on slide batch 22 counted 3,950 cells.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun R7 on slide batch 22 counted 4,180 cells using the same raw image dataset as comparison run R6."}, {"path": ["evidence", "5"], "text": "Comparison run R6 on slide batch 22 counted 3,950 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R7 on slide batch 22 counted 4,180 cells using the same raw image dataset as comparison run R6.", "negative_left": "Segmentation rerun R7 on slide batch 22 counted 4,050 cells using the same raw image dataset as comparison run R6.", "negative_right": "Comparison run R6 on slide batch 22 counted 3,950 cells.", "right": "Comparison run R6 on slide batch 22 counted 3,950 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-031", "id": "fast-43-diverse-259-031-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, so current calibration passes.", "Transmitted-light review found folds in 2 of 30 fields, well under 10%.", "Technical QC for the current run passes all internal checks.", "Replicate CV for the most recent relevant set was 8%, below the 15% threshold.", "Segmentation rerun R7 on slide batch 22 counted 4,050 cells using the same raw image dataset as comparison run R6.", "Comparison run R6 on slide batch 22 counted 3,950 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and unchanged question bindings, present two factual measurement sentences as evidence without policy language or answer leakage, and the counterfactual's 4,080 count (vs 4,000, ~2% change) is internally consistent with the unchanged bead offset, defect, and CV evidence, differing coherently from the base's 4,300 (~7.5% change).", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a three-channel cell assay for slide batch S-12; the image analyst reran segmentation as R-58 and compared it to run C-58; the assay scientist logged replicate variability. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and calibration passed review.\",\"Transmitted-light review found folds in 2 of 30 fields, well under the defect threshold.\",\"Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,300, using the same raw image dataset as its comparison run C-58.\",\"Comparison run C-58 for slide batch S-12 reported a cell count of 4,000.\",\"Replicate CV for the latest assay run was 9%, and technical QC passed all standard checks.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,300, using the same raw image dataset as its comparison run C-58."}, {"path": ["evidence", "3"], "text": "Comparison run C-58 for slide batch S-12 reported a cell count of 4,000."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,300, using the same raw image dataset as its comparison run C-58.", "negative_left": "Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,080, using the same raw image dataset as its comparison run C-58.", "negative_right": "Comparison run C-58 for slide batch S-12 reported a cell count of 4,000.", "right": "Comparison run C-58 for slide batch S-12 reported a cell count of 4,000."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-121", "id": "fast-43-diverse-259-121-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a three-channel cell assay for slide batch S-12; the image analyst reran segmentation as R-58 and compared it to run C-58; the assay scientist logged replicate variability. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and calibration passed review.", "Transmitted-light review found folds in 2 of 30 fields, well under the defect threshold.", "Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,300, using the same raw image dataset as its comparison run C-58.", "Comparison run C-58 for slide batch S-12 reported a cell count of 4,000.", "Replicate CV for the latest assay run was 9%, and technical QC passed all standard checks."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and unchanged question bindings, present two factual measurement sentences as evidence without policy language or answer leakage, and the counterfactual's 4,080 count (vs 4,000, ~2% change) is internally consistent with the unchanged bead offset, defect, and CV evidence, differing coherently from the base's 4,300 (~7.5% change).", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\":\"A microscopy technician captured a three-channel cell assay for slide batch S-12; the image analyst reran segmentation as R-58 and compared it to run C-58; the assay scientist logged replicate variability. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\",\"evidence\":[\"11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and calibration passed review.\",\"Transmitted-light review found folds in 2 of 30 fields, well under the defect threshold.\",\"Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,300, using the same raw image dataset as its comparison run C-58.\",\"Comparison run C-58 for slide batch S-12 reported a cell count of 4,000.\",\"Replicate CV for the latest assay run was 9%, and technical QC passed all standard checks.\"],\"request\":\"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "2"], "text": "Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,300, using the same raw image dataset as its comparison run C-58."}, {"path": ["evidence", "3"], "text": "Comparison run C-58 for slide batch S-12 reported a cell count of 4,000."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,300, using the same raw image dataset as its comparison run C-58.", "negative_left": "Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,080, using the same raw image dataset as its comparison run C-58.", "negative_right": "Comparison run C-58 for slide batch S-12 reported a cell count of 4,000.", "right": "Comparison run C-58 for slide batch S-12 reported a cell count of 4,000."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-121", "id": "fast-43-diverse-259-121-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a three-channel cell assay for slide batch S-12; the image analyst reran segmentation as R-58 and compared it to run C-58; the assay scientist logged replicate variability. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["11:30 bead recalibration measured 0.4 px; logs confirm unchanged optics afterward, and calibration passed review.", "Transmitted-light review found folds in 2 of 30 fields, well under the defect threshold.", "Segmentation rerun R-58 on slide batch S-12 reported a cell count of 4,080, using the same raw image dataset as its comparison run C-58.", "Comparison run C-58 for slide batch S-12 reported a cell count of 4,000.", "Replicate CV for the latest assay run was 9%, and technical QC passed all standard checks."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and question object, use complete factual evidence sentences, and the counterfactual only changes the S8 count (4050→4200), shifting the percent change from ~6.2% to ~2.4% without contradicting other measurements or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"evidence\": [\"09:40 sample overlay suggested a 1.7 px offset.\", \"12:05 bead recalibration measured 0.3 px; logs confirm unchanged optics afterward.\", \"Transmitted-light review found folds in 2 of 34 fields.\", \"Technical QC passed and replicate CV was measured at 7%.\", \"Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8.\", \"Comparison run S8 on slide set 14 counted 4,050 cells.\"], \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8."}, {"path": ["evidence", "5"], "text": "Comparison run S8 on slide set 14 counted 4,050 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8.", "negative_left": "Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8.", "negative_right": "Comparison run S8 on slide set 14 counted 4,200 cells.", "right": "Comparison run S8 on slide set 14 counted 4,050 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-124", "id": "fast-43-diverse-259-124-base", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:40 sample overlay suggested a 1.7 px offset.", "12:05 bead recalibration measured 0.3 px; logs confirm unchanged optics afterward.", "Transmitted-light review found folds in 2 of 34 fields.", "Technical QC passed and replicate CV was measured at 7%.", "Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8.", "Comparison run S8 on slide set 14 counted 4,050 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing policy and question object, use complete factual evidence sentences, and the counterfactual only changes the S8 count (4050→4200), shifting the percent change from ~6.2% to ~2.4% without contradicting other measurements or leaking any answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "full_context_fact_states": {"base": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "refuted", "a7": "supported"}, "counterfactual": {"a1": "refuted", "a2": "refuted", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "refuted", "a7": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual threshold, status, or dataset relationship; none is a bundled policy classification. The focus atom a4 is factual. The base and counter assignments can be realized with only the segmentation count-change threshold changing while all other facts remain fixed. The state-derived policy evidence preserves the routing rules needed from the original state; additional governing instructions and criteria in the questions object are automatically retained and need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes the analysis trigger: current calibration passes, the rerun used unchanged raw images, and the count change exceeds 5%. It also excludes acquisition, preparation, and assay-review triggers, making route_analysis sufficient under the exactly-one policy.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes that acquisition and preparation thresholds are not exceeded, unchanged-image segmentation is stable within 5%, and replicate CV does not exceed 15%. Calibration and technical QC pass, and no specialist-routing trigger remains, so accept_workflow is sufficient.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The measured channel offset in the latest same-session bead test is greater than 1.0 px."}, {"id": "a2", "statement": "The proportion of reviewed fields affected by a fixed slide defect in the most recent relevant slide review is greater than 10%."}, {"id": "a3", "statement": "The current calibration passes."}, {"id": "a4", "statement": "The measured cell-count change associated with the most recent relevant segmentation rerun, relative to its comparison run, is greater than 5%."}, {"id": "a5", "statement": "The most recent relevant segmentation rerun and its comparison run used the same raw image dataset."}, {"id": "a6", "statement": "The most recent relevant replicate coefficient of variation exceeds 15%."}, {"id": "a7", "statement": "The current technical QC passes."}], "base_state_json": "{\"context\": \"A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.\", \"evidence\": [\"09:40 sample overlay suggested a 1.7 px offset.\", \"12:05 bead recalibration measured 0.3 px; logs confirm unchanged optics afterward.\", \"Transmitted-light review found folds in 2 of 34 fields.\", \"Technical QC passed and replicate CV was measured at 7%.\", \"Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8.\", \"Comparison run S8 on slide set 14 counted 4,050 cells.\"], \"request\": \"Using the latest evidence, choose the single correct disposition.\"}", "base_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["evidence", "4"], "text": "Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8."}, {"path": ["evidence", "5"], "text": "Comparison run S8 on slide set 14 counted 4,050 cells."}], "policy_evidence": [{"path": ["context"], "text": "Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies."}, {"path": ["request"], "text": "Using the latest evidence, choose the single correct disposition."}], "rules": [{"justification": "The current calibration passes, and the most recent relevant segmentation rerun on unchanged raw images changes cell counts by more than 5%. The latest acquisition and preparation thresholds are not exceeded, and replicate CV does not exceed 15%, excluding the competing specialist routes.", "target": "route_analysis", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}, {"justification": "The latest acquisition and preparation thresholds are not exceeded, current calibration and technical QC pass, the most recent relevant unchanged-image segmentation rerun does not alter counts by more than 5%, and the most recent relevant replicate CV does not exceed 15%; therefore all stated acceptance criteria pass.", "target": "accept_workflow", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "refuted"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "refuted"}, {"atom_id": "a7", "state": "supported"}]}]}, "verified_pair": {"left": "Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8.", "negative_left": "Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8.", "negative_right": "Comparison run S8 on slide set 14 counted 4,200 cells.", "right": "Comparison run S8 on slide set 14 counted 4,050 cells."}, "verifier_independent_model": false}, "family": "fast-43-diverse-259-124", "id": "fast-43-diverse-259-124-counterfactual", "input": {"questions": {"decision": {"criteria": {"accept_workflow": "Accept without specialist routing only when acquisition, preparation, segmentation stability, and replicate-CV criteria all pass.", "none_of_above": "Use only when the supplied evidence is insufficient or conflicting under the stated routing policy.", "route_acquisition": "Route to the microscopy technician for acquisition correction when the latest same-session bead test shows channel offset above 1.0 px.", "route_analysis": "Route to the image analyst when current calibration passes but segmentation reruns on unchanged raw images alter counts by more than 5%.", "route_assay_review": "Route to the assay scientist for method review when technical QC passes but replicate CV exceeds 15%.", "route_preparation": "Route to a preparation specialist when a fixed slide defect affects more than 10% of reviewed fields."}, "instructions": "Apply the stated thresholds to the most recent relevant measurements. Later validated calibration evidence supersedes the earlier sample-based offset estimate. Select exactly one mutually exclusive disposition.", "type": "choice"}}, "state": {"context": "A microscopy technician captured a fictional three-channel cell assay; an image analyst segmented nuclei, and the assay scientist is reviewing replicate scores. Routing policy: acquisition if the latest same-session bead test shows channel offset >1.0 px; preparation if a fixed slide defect affects >10% of fields; analysis if calibration passes but rerunning segmentation changes counts >5%; assay review if technical QC passes yet replicate CV exceeds 15%; accept only if none applies.", "evidence": ["09:40 sample overlay suggested a 1.7 px offset.", "12:05 bead recalibration measured 0.3 px; logs confirm unchanged optics afterward.", "Transmitted-light review found folds in 2 of 34 fields.", "Technical QC passed and replicate CV was measured at 7%.", "Segmentation rerun S9 on slide set 14 counted 4,300 cells, using the same raw image dataset as comparison run S8.", "Comparison run S8 on slide set 14 counted 4,200 cells."], "request": "Using the latest evidence, choose the single correct disposition."}}, "method": "c2d", "provenance": {"source_id": "diverse-259", "source_is_synthetic": true, "source_sha256": "f3dd4a26b2a6b1f571cd0da2f4f24460fa54bed3a53df14bee67a271c599bf9f", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "accept_workflow"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric verbatim and the same case ID/question; the single-sentence edit changes the recalculation correction from 0.9 to 0.3 points, shifting total drift from 2.3% (exceeds 2.0%) to 1.7% (does not exceed), which coherently changes the controlling criterion without contradicting other measurements; the two focus evidence spans are factual statements about drift values, not policy or rule text; neither context states or hints at the final routing decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"image analyst\",\"text\":\"Checking the supplied raw images for case CW-4471, I measure exactly 3.0% bubble coverage and 1.0% saturation, so neither exceeds its threshold. Replicate segmentation disagreement is 5.2%, which does exceed the 5.0% threshold.\"},{\"speaker\":\"microscopy technician\",\"text\":\"At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation.\"},{\"speaker\":\"microscopy technician\",\"text\":\"The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.9 percentage points to the initial drift value.\"},{\"speaker\":\"assay scientist\",\"text\":\"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation."}, {"path": ["3", "text"], "text": "The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.9 percentage points to the initial drift value."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation.", "negative_left": "At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation.", "negative_right": "The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.3 percentage points to the initial drift value.", "right": "The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.9 percentage points to the initial drift value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-260-014", "id": "fast-43-diverse-260-014-base", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "Checking the supplied raw images for case CW-4471, I measure exactly 3.0% bubble coverage and 1.0% saturation, so neither exceeds its threshold. Replicate segmentation disagreement is 5.2%, which does exceed the 5.0% threshold."}, {"speaker": "microscopy technician", "text": "At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation."}, {"speaker": "microscopy technician", "text": "The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.9 percentage points to the initial drift value."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_acquisition"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric verbatim and the same case ID/question; the single-sentence edit changes the recalculation correction from 0.9 to 0.3 points, shifting total drift from 2.3% (exceeds 2.0%) to 1.7% (does not exceed), which coherently changes the controlling criterion without contradicting other measurements; the two focus evidence spans are factual statements about drift values, not policy or rule text; neither context states or hints at the final routing decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\":\"assay scientist\",\"text\":\"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"},{\"speaker\":\"image analyst\",\"text\":\"Checking the supplied raw images for case CW-4471, I measure exactly 3.0% bubble coverage and 1.0% saturation, so neither exceeds its threshold. Replicate segmentation disagreement is 5.2%, which does exceed the 5.0% threshold.\"},{\"speaker\":\"microscopy technician\",\"text\":\"At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation.\"},{\"speaker\":\"microscopy technician\",\"text\":\"The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.9 percentage points to the initial drift value.\"},{\"speaker\":\"assay scientist\",\"text\":\"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation."}, {"path": ["3", "text"], "text": "The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.9 percentage points to the initial drift value."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation.", "negative_left": "At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation.", "negative_right": "The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.3 percentage points to the initial drift value.", "right": "The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.9 percentage points to the initial drift value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-260-014", "id": "fast-43-diverse-260-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "Checking the supplied raw images for case CW-4471, I measure exactly 3.0% bubble coverage and 1.0% saturation, so neither exceeds its threshold. Replicate segmentation disagreement is 5.2%, which does exceed the 5.0% threshold."}, {"speaker": "microscopy technician", "text": "At routing time, the calibration log for the case workflow, case CW-4471, records an initial drift value of 1.4% before recalculation."}, {"speaker": "microscopy technician", "text": "The recalculation applied to case CW-4471's calibration log at routing time adds a correction of 0.3 percentage points to the initial drift value."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric and unchanged question, use the same two factual technician sentences as evidence, and the counterfactual coherently lowers the raw drift value (1.4%+0.1%=1.5%) without contradicting any other measurement, while no gold answer or rule table is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\": \"assay scientist\", \"text\": \"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"}, {\"speaker\": \"image analyst\", \"text\": \"Checking the supplied raw images, I measure exactly 2.5% bubble coverage and 0.6% saturation, so neither preparation nor saturation-based acquisition criteria are met. Replicate segmentation disagreement is 5.2%, above the analysis threshold.\"}, {\"speaker\": \"microscopy technician\", \"text\": \"At routing time, the calibration log for the case workflow recorded a raw drift value of 3.4% before recalculation.\"}, {\"speaker\": \"microscopy technician\", \"text\": \"The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value.\"}, {\"speaker\": \"assay scientist\", \"text\": \"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At routing time, the calibration log for the case workflow recorded a raw drift value of 3.4% before recalculation."}, {"path": ["3", "text"], "text": "The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, the calibration log for the case workflow recorded a raw drift value of 3.4% before recalculation.", "negative_left": "At routing time, the calibration log for the case workflow recorded a raw drift value of 1.4% before recalculation.", "negative_right": "The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value.", "right": "The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-260-015", "id": "fast-43-diverse-260-015-base", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "Checking the supplied raw images, I measure exactly 2.5% bubble coverage and 0.6% saturation, so neither preparation nor saturation-based acquisition criteria are met. Replicate segmentation disagreement is 5.2%, above the analysis threshold."}, {"speaker": "microscopy technician", "text": "At routing time, the calibration log for the case workflow recorded a raw drift value of 3.4% before recalculation."}, {"speaker": "microscopy technician", "text": "The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_acquisition"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric and unchanged question, use the same two factual technician sentences as evidence, and the counterfactual coherently lowers the raw drift value (1.4%+0.1%=1.5%) without contradicting any other measurement, while no gold answer or rule table is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\": \"assay scientist\", \"text\": \"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"}, {\"speaker\": \"image analyst\", \"text\": \"Checking the supplied raw images, I measure exactly 2.5% bubble coverage and 0.6% saturation, so neither preparation nor saturation-based acquisition criteria are met. Replicate segmentation disagreement is 5.2%, above the analysis threshold.\"}, {\"speaker\": \"microscopy technician\", \"text\": \"At routing time, the calibration log for the case workflow recorded a raw drift value of 3.4% before recalculation.\"}, {\"speaker\": \"microscopy technician\", \"text\": \"The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value.\"}, {\"speaker\": \"assay scientist\", \"text\": \"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "At routing time, the calibration log for the case workflow recorded a raw drift value of 3.4% before recalculation."}, {"path": ["3", "text"], "text": "The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "At routing time, the calibration log for the case workflow recorded a raw drift value of 3.4% before recalculation.", "negative_left": "At routing time, the calibration log for the case workflow recorded a raw drift value of 1.4% before recalculation.", "negative_right": "The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value.", "right": "The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value."}, "verifier_independent_model": false}, "family": "fast-43-diverse-260-015", "id": "fast-43-diverse-260-015-counterfactual", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "Checking the supplied raw images, I measure exactly 2.5% bubble coverage and 0.6% saturation, so neither preparation nor saturation-based acquisition criteria are met. Replicate segmentation disagreement is 5.2%, above the analysis threshold."}, {"speaker": "microscopy technician", "text": "At routing time, the calibration log for the case workflow recorded a raw drift value of 1.4% before recalculation."}, {"speaker": "microscopy technician", "text": "The recalculation formula applied to the case workflow's calibration log adds a fixed correction of 0.1% to the raw drift value."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric, case ID, and question verbatim; the counterfactual only changes the sensor reading (97.4→98.5), yielding a coherent alternate drift value without contradicting other measurements; the two evidence spans are exact factual sentences from the technician's turn with no policy text, rule tables, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\": \"assay scientist\", \"text\": \"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"}, {\"speaker\": \"image analyst\", \"text\": \"Checking the supplied raw images for case workflow C-482, bubble coverage measures exactly 3.0%, which does not exceed the 3.0% preparation threshold. Saturated pixels measure exactly 1.0%, which does not exceed the 1.0% threshold. Replicate segmentation disagreement is 5.2%, exceeding the 5.0% threshold.\"}, {\"speaker\": \"microscopy technician\", \"text\": \"The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time. Recalculating using the current sensor reading for case workflow C-482, which measures 97.4 units, produces a calibration drift equal to the percentage difference between the two values.\"}, {\"speaker\": \"assay scientist\", \"text\": \"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time."}, {"path": ["2", "text"], "text": "Recalculating using the current sensor reading for case workflow C-482, which measures 97.4 units, produces a calibration drift equal to the percentage difference between the two values."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time.", "negative_left": "The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time.", "negative_right": "Recalculating using the current sensor reading for case workflow C-482, which measures 98.5 units, produces a calibration drift equal to the percentage difference between the two values.", "right": "Recalculating using the current sensor reading for case workflow C-482, which measures 97.4 units, produces a calibration drift equal to the percentage difference between the two values."}, "verifier_independent_model": false}, "family": "fast-43-diverse-260-023", "id": "fast-43-diverse-260-023-base", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "Checking the supplied raw images for case workflow C-482, bubble coverage measures exactly 3.0%, which does not exceed the 3.0% preparation threshold. Saturated pixels measure exactly 1.0%, which does not exceed the 1.0% threshold. Replicate segmentation disagreement is 5.2%, exceeding the 5.0% threshold."}, {"speaker": "microscopy technician", "text": "The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time. Recalculating using the current sensor reading for case workflow C-482, which measures 97.4 units, produces a calibration drift equal to the percentage difference between the two values."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_acquisition"}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original rubric, case ID, and question verbatim; the counterfactual only changes the sensor reading (97.4→98.5), yielding a coherent alternate drift value without contradicting other measurements; the two evidence spans are exact factual sentences from the technician's turn with no policy text, rule tables, or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "refuted", "A4": "supported"}, "counterfactual": {"A1": "refuted", "A2": "refuted", "A3": "refuted", "A4": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "All atoms state single factual threshold relationships, and A2 is a factual measurement relation rather than a policy conclusion. The base and counter assignments differ only on A2 and are jointly realizable as separate synthetic scenarios. The cited state evidence preserves the routing order, thresholds, exact-threshold treatment, and source-precedence rules needed from the original state. The rule table is partial but valid; omitted preparation and saturation-only acquisition cases may abstain.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A1 refuted excludes preparation, and A2 supported triggers acquisition through recalculated drift exceeding 2.0%. Saturation and segmentation values cannot change that outcome under the stated precedence.", "rule_index": 0, "sound": true}, {"reason": "A1 refuted excludes preparation; A2 and A3 refuted exclude both acquisition triggers; and A4 supported triggers analysis. Thus every assignment satisfying the conjunction routes to analysis.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows image-confirmed bubble coverage exceeding 3.0%."}, {"id": "A2", "statement": "At routing time, recalculation from the calibration log associated with the case workflow shows calibration drift exceeding 2.0%."}, {"id": "A3", "statement": "At routing time, the direct measurement from the supplied raw images for the case workflow shows saturated pixels exceeding 1.0%."}, {"id": "A4", "statement": "At routing time, replicate segmentation for the case workflow shows disagreement exceeding 5.0%."}], "base_state_json": "[{\"speaker\": \"assay scientist\", \"text\": \"Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries.\"}, {\"speaker\": \"image analyst\", \"text\": \"Checking the supplied raw images for case workflow C-482, bubble coverage measures exactly 3.0%, which does not exceed the 3.0% preparation threshold. Saturated pixels measure exactly 1.0%, which does not exceed the 1.0% threshold. Replicate segmentation disagreement is 5.2%, exceeding the 5.0% threshold.\"}, {\"speaker\": \"microscopy technician\", \"text\": \"The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time. Recalculating using the current sensor reading for case workflow C-482, which measures 97.4 units, produces a calibration drift equal to the percentage difference between the two values.\"}, {\"speaker\": \"assay scientist\", \"text\": \"Which specialist should receive the case?\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": ["2", "text"], "text": "The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time."}, {"path": ["2", "text"], "text": "Recalculating using the current sensor reading for case workflow C-482, which measures 97.4 units, produces a calibration drift equal to the percentage difference between the two values."}], "policy_evidence": [{"path": ["0", "text"], "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}], "rules": [{"justification": "Preparation does not trigger because controlling image-confirmed bubble coverage does not exceed 3.0%. Acquisition then triggers because controlling recalculated calibration drift exceeds 2.0%.", "target": "route_acquisition", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}]}, {"justification": "Preparation does not trigger, neither acquisition threshold is exceeded, and replicate segmentation disagreement exceeds 5.0%, so the case routes to analysis.", "target": "route_analysis", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "supported"}]}]}, "verified_pair": {"left": "The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time.", "negative_left": "The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time.", "negative_right": "Recalculating using the current sensor reading for case workflow C-482, which measures 98.5 units, produces a calibration drift equal to the percentage difference between the two values.", "right": "Recalculating using the current sensor reading for case workflow C-482, which measures 97.4 units, produces a calibration drift equal to the percentage difference between the two values."}, "verifier_independent_model": false}, "family": "fast-43-diverse-260-023", "id": "fast-43-diverse-260-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"route_acquisition": "Route to the microscopy technician for acquisition review because controlling recalculated drift exceeds 2.0% or controlling saturation exceeds 1.0%.", "route_analysis": "Route to the image analyst because preparation and acquisition criteria do not trigger, while replicate segmentation disagreement exceeds 5.0%.", "route_preparation": "Route to the microscopy technician for preparation review because controlling image-confirmed bubble coverage exceeds 3.0%."}, "instructions": "Route the workflow to exactly one specialist by applying the stated evidence-precedence rules and thresholds.", "type": "choice"}}, "state": [{"speaker": "assay scientist", "text": "Use this routing rubric: preparation if image-confirmed bubble coverage exceeds 3.0%; otherwise acquisition if recalculated calibration drift exceeds 2.0% or saturated pixels exceed 1.0%; otherwise analysis if replicate segmentation disagreement exceeds 5.0%. Exact thresholds pass. Direct image measurements override technician notes, and recalculations override software summaries."}, {"speaker": "image analyst", "text": "Checking the supplied raw images for case workflow C-482, bubble coverage measures exactly 3.0%, which does not exceed the 3.0% preparation threshold. Saturated pixels measure exactly 1.0%, which does not exceed the 1.0% threshold. Replicate segmentation disagreement is 5.2%, exceeding the 5.0% threshold."}, {"speaker": "microscopy technician", "text": "The calibration log for case workflow C-482 records a baseline calibration value of 100.0 units at routing time. Recalculating using the current sensor reading for case workflow C-482, which measures 98.5 units, produces a calibration drift equal to the percentage difference between the two values."}, {"speaker": "assay scientist", "text": "Which specialist should receive the case?"}]}, "method": "c2d", "provenance": {"source_id": "diverse-260", "source_is_synthetic": true, "source_sha256": "66e1be58d132b40118774833f59ee8be60dd2efe99917e0e4f8c6442b571a5bf", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "route_analysis"}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration gate policy and all quality indicators, matching the original question's rubric; the counterfactual only changes the measured calibration value (112.0→105.0), yielding a coherent single-sentence change (12%→5% error) without duplicating or contradicting other measurements; both evidence spans are complete factual sentences about calibration values, not policy statements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"A microscopy technician prepared a fictional fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields. The image analyst reports segmentation IoU of 0.94, meeting validation criteria, and replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms that the positive-control intensity falls within its expected range, and artifacts across all fields are negligible. No major preparation, capture, or segmentation defect was identified in the review. The capture session includes a mandatory calibration record. The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers. During the capture session, the calibration record's measured value for that same reference standard is 112.0 micrometers. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers."}, {"path": [], "text": "During the capture session, the calibration record's measured value for that same reference standard is 112.0 micrometers."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers.", "negative_left": "The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers.", "negative_right": "During the capture session, the calibration record's measured value for that same reference standard is 105.0 micrometers.", "right": "During the capture session, the calibration record's measured value for that same reference standard is 112.0 micrometers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-264-009", "id": "fast-43-diverse-264-009-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "A microscopy technician prepared a fictional fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields. The image analyst reports segmentation IoU of 0.94, meeting validation criteria, and replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms that the positive-control intensity falls within its expected range, and artifacts across all fields are negligible. No major preparation, capture, or segmentation defect was identified in the review. The capture session includes a mandatory calibration record. The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers. During the capture session, the calibration record's measured value for that same reference standard is 112.0 micrometers. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration gate policy and all quality indicators, matching the original question's rubric; the counterfactual only changes the measured calibration value (112.0→105.0), yielding a coherent single-sentence change (12%→5% error) without duplicating or contradicting other measurements; both evidence spans are complete factual sentences about calibration values, not policy statements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"A microscopy technician prepared a fictional fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields. The image analyst reports segmentation IoU of 0.94, meeting validation criteria, and replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms that the positive-control intensity falls within its expected range, and artifacts across all fields are negligible. No major preparation, capture, or segmentation defect was identified in the review. The capture session includes a mandatory calibration record. The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers. During the capture session, the calibration record's measured value for that same reference standard is 112.0 micrometers. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers."}, {"path": [], "text": "During the capture session, the calibration record's measured value for that same reference standard is 112.0 micrometers."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers.", "negative_left": "The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers.", "negative_right": "During the capture session, the calibration record's measured value for that same reference standard is 105.0 micrometers.", "right": "During the capture session, the calibration record's measured value for that same reference standard is 112.0 micrometers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-264-009", "id": "fast-43-diverse-264-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "A microscopy technician prepared a fictional fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields. The image analyst reports segmentation IoU of 0.94, meeting validation criteria, and replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms that the positive-control intensity falls within its expected range, and artifacts across all fields are negligible. No major preparation, capture, or segmentation defect was identified in the review. The capture session includes a mandatory calibration record. The microscopy workflow's mandatory calibration record has a certified reference standard measuring exactly 100.0 micrometers. During the capture session, the calibration record's measured value for that same reference standard is 105.0 micrometers. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration-gate policy verbatim and keep the same specimen batch M-77 entity binding; the evidence spans are two complete factual sentences about the calibration record and its measured value, not policy text; the counterfactual only changes the measured value (505 vs 560 microns) which is coherent with all other unchanged facts and does not contradict any other measurement; neither context states a reliability level, code, or instructs an output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"A microscopy technician prepared specimen batch M-77 with even staining, no folds, and sharp focus across 12 fields. The image analyst reports a segmentation IoU of 0.94, meeting validation criteria, and replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms the positive-control intensity falls within its expected range, and no artifacts were observed in the captured images. No major preparation, capture, or segmentation defects were noted anywhere in the workflow. The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns. That calibration record's measured value against the certified standard was 560 microns. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns."}, {"path": [], "text": "That calibration record's measured value against the certified standard was 560 microns."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns.", "negative_left": "The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns.", "negative_right": "That calibration record's measured value against the certified standard was 505 microns.", "right": "That calibration record's measured value against the certified standard was 560 microns."}, "verifier_independent_model": false}, "family": "fast-43-diverse-264-010", "id": "fast-43-diverse-264-010-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "A microscopy technician prepared specimen batch M-77 with even staining, no folds, and sharp focus across 12 fields. The image analyst reports a segmentation IoU of 0.94, meeting validation criteria, and replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms the positive-control intensity falls within its expected range, and no artifacts were observed in the captured images. No major preparation, capture, or segmentation defects were noted anywhere in the workflow. The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns. That calibration record's measured value against the certified standard was 560 microns. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration-gate policy verbatim and keep the same specimen batch M-77 entity binding; the evidence spans are two complete factual sentences about the calibration record and its measured value, not policy text; the counterfactual only changes the measured value (505 vs 560 microns) which is coherent with all other unchanged facts and does not contradict any other measurement; neither context states a reliability level, code, or instructs an output.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"A microscopy technician prepared specimen batch M-77 with even staining, no folds, and sharp focus across 12 fields. The image analyst reports a segmentation IoU of 0.94, meeting validation criteria, and replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms the positive-control intensity falls within its expected range, and no artifacts were observed in the captured images. No major preparation, capture, or segmentation defects were noted anywhere in the workflow. The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns. That calibration record's measured value against the certified standard was 560 microns. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns."}, {"path": [], "text": "That calibration record's measured value against the certified standard was 560 microns."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns.", "negative_left": "The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns.", "negative_right": "That calibration record's measured value against the certified standard was 505 microns.", "right": "That calibration record's measured value against the certified standard was 560 microns."}, "verifier_independent_model": false}, "family": "fast-43-diverse-264-010", "id": "fast-43-diverse-264-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "A microscopy technician prepared specimen batch M-77 with even staining, no folds, and sharp focus across 12 fields. The image analyst reports a segmentation IoU of 0.94, meeting validation criteria, and replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms the positive-control intensity falls within its expected range, and no artifacts were observed in the captured images. No major preparation, capture, or segmentation defects were noted anywhere in the workflow. The capture session for specimen batch M-77's microscopy run includes a calibration record used for quantitative size and density measurements, referencing a certified stage micrometer standard of 500 microns. That calibration record's measured value against the certified standard was 505 microns. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration gate policy verbatim, preserve the same workflow entity and quality facts, differ only in the imaged bead diameter (5.60 vs 5.03 µm) which stays a coherent single-sentence factual change, and neither context includes any label, rule table, or explicit answer instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"A microscopy technician prepared a fictional fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields. The image analyst reports segmentation IoU of 0.94 and replicate object-count CV of 4%. The assay scientist confirms that the positive-control intensity falls within its expected range. A mandatory calibration record was generated during the capture session, referencing a certified reference bead of known diameter. The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers. The same calibration record measures that reference bead's imaged diameter as 5.60 micrometers. No major preparation, capture, or segmentation defects were noted, and artifacts across the imaged fields are negligible. Segmentation meets validation criteria and replicate measurements are consistent across all fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers."}, {"path": [], "text": "The same calibration record measures that reference bead's imaged diameter as 5.60 micrometers."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers.", "negative_left": "The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers.", "negative_right": "The same calibration record measures that reference bead's imaged diameter as 5.03 micrometers.", "right": "The same calibration record measures that reference bead's imaged diameter as 5.60 micrometers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-264-028", "id": "fast-43-diverse-264-028-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "A microscopy technician prepared a fictional fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields. The image analyst reports segmentation IoU of 0.94 and replicate object-count CV of 4%. The assay scientist confirms that the positive-control intensity falls within its expected range. A mandatory calibration record was generated during the capture session, referencing a certified reference bead of known diameter. The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers. The same calibration record measures that reference bead's imaged diameter as 5.60 micrometers. No major preparation, capture, or segmentation defects were noted, and artifacts across the imaged fields are negligible. Segmentation meets validation criteria and replicate measurements are consistent across all fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration gate policy verbatim, preserve the same workflow entity and quality facts, differ only in the imaged bead diameter (5.60 vs 5.03 µm) which stays a coherent single-sentence factual change, and neither context includes any label, rule table, or explicit answer instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"A microscopy technician prepared a fictional fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields. The image analyst reports segmentation IoU of 0.94 and replicate object-count CV of 4%. The assay scientist confirms that the positive-control intensity falls within its expected range. A mandatory calibration record was generated during the capture session, referencing a certified reference bead of known diameter. The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers. The same calibration record measures that reference bead's imaged diameter as 5.60 micrometers. No major preparation, capture, or segmentation defects were noted, and artifacts across the imaged fields are negligible. Segmentation meets validation criteria and replicate measurements are consistent across all fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers."}, {"path": [], "text": "The same calibration record measures that reference bead's imaged diameter as 5.60 micrometers."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers.", "negative_left": "The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers.", "negative_right": "The same calibration record measures that reference bead's imaged diameter as 5.03 micrometers.", "right": "The same calibration record measures that reference bead's imaged diameter as 5.60 micrometers."}, "verifier_independent_model": false}, "family": "fast-43-diverse-264-028", "id": "fast-43-diverse-264-028-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "A microscopy technician prepared a fictional fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields. The image analyst reports segmentation IoU of 0.94 and replicate object-count CV of 4%. The assay scientist confirms that the positive-control intensity falls within its expected range. A mandatory calibration record was generated during the capture session, referencing a certified reference bead of known diameter. The capture session's calibration record for quantitative size and density measurements reports a certified reference bead diameter of 5.00 micrometers. The same calibration record measures that reference bead's imaged diameter as 5.03 micrometers. No major preparation, capture, or segmentation defects were noted, and artifacts across the imaged fields are negligible. Segmentation meets validation criteria and replicate measurements are consistent across all fields. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration gate policy and CS-217 entity/path bindings unchanged from the question; the two evidence spans are complete factual measurement sentences, not policy text; the counterfactual only alters the re-measured value (0.285→0.255) without contradicting other unchanged facts, and neither context reveals a gold label or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"A microscopy technician prepared a fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields, part of capture session CS-217, which has a mandatory calibration record. The image analyst reports segmentation IoU of 0.94, meeting validation criteria, with replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms the positive-control intensity falls within its expected range, and artifacts in the captured images are negligible. No major preparation, capture, or segmentation defect was identified that would bias the measurements. The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel. When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.285 µm/pixel. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel."}, {"path": [], "text": "When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.285 µm/pixel."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel.", "negative_left": "The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel.", "negative_right": "When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.255 µm/pixel.", "right": "When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.285 µm/pixel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-264-031", "id": "fast-43-diverse-264-031-base", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "A microscopy technician prepared a fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields, part of capture session CS-217, which has a mandatory calibration record. The image analyst reports segmentation IoU of 0.94, meeting validation criteria, with replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms the positive-control intensity falls within its expected range, and artifacts in the captured images are negligible. No major preparation, capture, or segmentation defect was identified that would bias the measurements. The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel. When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.285 µm/pixel. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "science-04", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the mandatory calibration gate policy and CS-217 entity/path bindings unchanged from the question; the two evidence spans are complete factual measurement sentences, not policy text; the counterfactual only alters the re-measured value (0.285→0.255) without contradicting other unchanged facts, and neither context reveals a gold label or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A2": "supported", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "counterfactual": {"A1": "supported", "A2": "refuted", "A3": "refuted", "A4": "refuted", "A5": "refuted", "A6": "supported", "A7": "supported", "A8": "supported"}, "remove_left": {"A2": "unknown"}, "remove_right": {"A2": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A2": "unknown"}, "negative_pair": {"A2": "refuted"}, "negative_sentence": {"A2": "unknown"}, "positive_pair": {"A2": "supported"}, "right": {"A2": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual relationship, and A2 is a factual calibration condition rather than a policy conclusion. The base and counter assignments are realizable with only A2 changing: a present calibration can respectively have an error above 10% or not above 10%, while the other quality facts remain favorable. The policy evidence correctly preserves the substantive mandatory-gate rule originating in the original state; the full rubric remains automatically available through the unchanged questions object.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "A supported A2 means the calibration scale error exceeds 10%, which mandates Level 0 regardless of favorable preparation, segmentation, replicate, or control facts.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes a present calibration with no error above 10%, excludes all listed major defects, and requires negligible artifacts, validated segmentation, and consistent replicates. These conditions are sufficient for Level 3 and exclude the lower-level conditions.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The microscopy workflow’s capture session has a mandatory calibration record."}, {"id": "A2", "statement": "The scale error of the mandatory calibration record used for quantitative size and density measurements in the microscopy workflow’s capture session exceeds 10%."}, {"id": "A3", "statement": "A major preparation defect materially biases the microscopy workflow’s measurements."}, {"id": "A4", "statement": "A major capture defect materially biases the microscopy workflow’s measurements."}, {"id": "A5", "statement": "A major segmentation defect materially biases the microscopy workflow’s measurements."}, {"id": "A6", "statement": "Artifacts in the microscopy workflow’s images are negligible."}, {"id": "A7", "statement": "The microscopy workflow’s segmentation meets validation criteria."}, {"id": "A8", "statement": "The microscopy workflow’s replicate measurements are consistent."}], "base_state_json": "\"A microscopy technician prepared a fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields, part of capture session CS-217, which has a mandatory calibration record. The image analyst reports segmentation IoU of 0.94, meeting validation criteria, with replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms the positive-control intensity falls within its expected range, and artifacts in the captured images are negligible. No major preparation, capture, or segmentation defect was identified that would bias the measurements. The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel. When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.285 µm/pixel. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators.\"", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}], "focus_atom": "A2", "focus_evidence": [{"path": [], "text": "The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel."}, {"path": [], "text": "When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.285 µm/pixel."}], "policy_evidence": [{"path": [], "text": "Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}], "rules": [{"justification": "A scale error exceeding 10% triggers Level 0 regardless of all favorable workflow facts.", "target": "0", "when": [{"atom_id": "A2", "state": "supported"}]}, {"justification": "The calibration is present and does not exceed the failure threshold, no major preparation, capture, or segmentation defect materially biases measurements, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent.", "target": "3", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "refuted"}, {"atom_id": "A3", "state": "refuted"}, {"atom_id": "A4", "state": "refuted"}, {"atom_id": "A5", "state": "refuted"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}]}]}, "verified_pair": {"left": "The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel.", "negative_left": "The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel.", "negative_right": "When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.255 µm/pixel.", "right": "When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.285 µm/pixel."}, "verifier_independent_model": false}, "family": "fast-43-diverse-264-031", "id": "fast-43-diverse-264-031-counterfactual", "input": {"questions": {"decision": {"criteria": ["Level 0 — Unusable: A mandatory calibration is missing or its scale error exceeds 10%; quantitative size or density results must be rejected regardless of other workflow quality.", "Level 1 — Low reliability: Calibration passes, but major preparation, capture, or segmentation defects materially bias measurements, so results require reacquisition or reanalysis before use.", "Level 2 — Moderate reliability: Calibration passes and no major defect exists, but minor artifacts or replicate variability limit precision; results may be used only with documented qualification.", "Level 3 — High reliability: Calibration passes, artifacts are negligible, segmentation meets validation criteria, and replicate measurements are consistent; results are accepted without qualification."], "instructions": "Assign the workflow’s measurement-reliability level using the four-level rubric. Treat favorable preparation, segmentation, replicate, or control facts as distractors when a mandatory calibration gate fails.", "type": "score"}}, "state": "A microscopy technician prepared a fluorescent tissue slide with even staining, no folds, and sharp focus across 12 fields, part of capture session CS-217, which has a mandatory calibration record. The image analyst reports segmentation IoU of 0.94, meeting validation criteria, with replicate object-count CV of 4%, indicating consistent replicate measurements. The assay scientist confirms the positive-control intensity falls within its expected range, and artifacts in the captured images are negligible. No major preparation, capture, or segmentation defect was identified that would bias the measurements. The capture session labeled CS-217 used a calibration slide with a certified pixel-to-micron conversion of 0.250 µm/pixel. When re-measured against the NIST-traceable standard, the calibration slide used in CS-217 actually corresponds to 0.255 µm/pixel. Policy makes calibration a mandatory gate: missing calibration or scale error above 10% renders all size and density measurements unusable, regardless of other quality indicators."}, "method": "c2d", "provenance": {"source_id": "diverse-264", "source_is_synthetic": true, "source_sha256": "abf725258890969aa3c8d2ab8aa05aece67f209eedcabf6d54d91d48d6ab75d6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 3}, "source_family": "science-04", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol policy verbatim and preserve the WHC, temperature range, and CO2 measurement bindings; the two log-note sentences are purely factual timestamps with no rule tables or answer hints; the counterfactual simply extends the excursion end time to 12:30, yielding a 195-minute excursion that coherently exceeds the 90-minute exception without contradicting any other stated fact.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"A soil scientist reviews a technician's 14-day moisture microcosm replication chamber log. All jars began at 61% water-holding capacity, and CO₂ was measured on days 3, 7, and 14. The incubation temperature stayed between 19°C and 21°C at all times except for one shared excursion affecting every treatment and paired control alike.\",\"policy\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\"],\"log_notes\":[\"The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation.\",\"The temperature logger recorded the end of that same excursion at 10:30 on day 6 of the incubation.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["log_notes", "0"], "text": "The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation."}, {"path": ["log_notes", "1"], "text": "The temperature logger recorded the end of that same excursion at 10:30 on day 6 of the incubation."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation.", "negative_left": "The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation.", "negative_right": "The temperature logger recorded the end of that same excursion at 12:30 on day 6 of the incubation.", "right": "The temperature logger recorded the end of that same excursion at 10:30 on day 6 of the incubation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-002", "id": "fast-43-diverse-267-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist reviews a technician's 14-day moisture microcosm replication chamber log. All jars began at 61% water-holding capacity, and CO₂ was measured on days 3, 7, and 14. The incubation temperature stayed between 19°C and 21°C at all times except for one shared excursion affecting every treatment and paired control alike.", "log_notes": ["The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation.", "The temperature logger recorded the end of that same excursion at 10:30 on day 6 of the incubation."], "policy": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol policy verbatim and preserve the WHC, temperature range, and CO2 measurement bindings; the two log-note sentences are purely factual timestamps with no rule tables or answer hints; the counterfactual simply extends the excursion end time to 12:30, yielding a 195-minute excursion that coherently exceeds the 90-minute exception without contradicting any other stated fact.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"A soil scientist reviews a technician's 14-day moisture microcosm replication chamber log. All jars began at 61% water-holding capacity, and CO₂ was measured on days 3, 7, and 14. The incubation temperature stayed between 19°C and 21°C at all times except for one shared excursion affecting every treatment and paired control alike.\",\"policy\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\"],\"log_notes\":[\"The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation.\",\"The temperature logger recorded the end of that same excursion at 10:30 on day 6 of the incubation.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["log_notes", "0"], "text": "The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation."}, {"path": ["log_notes", "1"], "text": "The temperature logger recorded the end of that same excursion at 10:30 on day 6 of the incubation."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation.", "negative_left": "The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation.", "negative_right": "The temperature logger recorded the end of that same excursion at 12:30 on day 6 of the incubation.", "right": "The temperature logger recorded the end of that same excursion at 10:30 on day 6 of the incubation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-002", "id": "fast-43-diverse-267-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist reviews a technician's 14-day moisture microcosm replication chamber log. All jars began at 61% water-holding capacity, and CO₂ was measured on days 3, 7, and 14. The incubation temperature stayed between 19°C and 21°C at all times except for one shared excursion affecting every treatment and paired control alike.", "log_notes": ["The temperature logger recorded the start of the single excursion at 09:15 on day 6 of the incubation.", "The temperature logger recorded the end of that same excursion at 12:30 on day 6 of the incubation."], "policy": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol requirements and 90-minute shared-excursion exception verbatim, preserve the same entities/timeframe/request, and change only the excursion end time (10:30→11:10), yielding a coherent shift from a 75-minute to a 115-minute excursion without contradicting other evidence or embedding any rule table, label, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\": \"A soil scientist is reviewing a technician's 14-day moisture microcosm replication. An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.\", \"evidence\": [\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\", \"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\", \"All jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14.\", \"Outside the single excursion, the incubator held steady between 19°C and 21°C for the full 14-day run.\", \"The single excursion affected every treatment jar and its paired control simultaneously.\", \"The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6.\", \"The logbook for jar batch M lists the temperature excursion as ending at 10:30 on day 6.\"], \"request\": \"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6."}, {"path": ["evidence", "6"], "text": "The logbook for jar batch M lists the temperature excursion as ending at 10:30 on day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6.", "negative_left": "The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6.", "negative_right": "The logbook for jar batch M lists the temperature excursion as ending at 11:10 on day 6.", "right": "The logbook for jar batch M lists the temperature excursion as ending at 10:30 on day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-003", "id": "fast-43-diverse-267-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a technician's 14-day moisture microcosm replication. An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "All jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14.", "Outside the single excursion, the incubator held steady between 19°C and 21°C for the full 14-day run.", "The single excursion affected every treatment jar and its paired control simultaneously.", "The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6.", "The logbook for jar batch M lists the temperature excursion as ending at 10:30 on day 6."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol requirements and 90-minute shared-excursion exception verbatim, preserve the same entities/timeframe/request, and change only the excursion end time (10:30→11:10), yielding a coherent shift from a 75-minute to a 115-minute excursion without contradicting other evidence or embedding any rule table, label, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\": \"A soil scientist is reviewing a technician's 14-day moisture microcosm replication. An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.\", \"evidence\": [\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\", \"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\", \"All jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14.\", \"Outside the single excursion, the incubator held steady between 19°C and 21°C for the full 14-day run.\", \"The single excursion affected every treatment jar and its paired control simultaneously.\", \"The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6.\", \"The logbook for jar batch M lists the temperature excursion as ending at 10:30 on day 6.\"], \"request\": \"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6."}, {"path": ["evidence", "6"], "text": "The logbook for jar batch M lists the temperature excursion as ending at 10:30 on day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6.", "negative_left": "The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6.", "negative_right": "The logbook for jar batch M lists the temperature excursion as ending at 11:10 on day 6.", "right": "The logbook for jar batch M lists the temperature excursion as ending at 10:30 on day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-003", "id": "fast-43-diverse-267-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a technician's 14-day moisture microcosm replication. An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "All jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14.", "Outside the single excursion, the incubator held steady between 19°C and 21°C for the full 14-day run.", "The single excursion affected every treatment jar and its paired control simultaneously.", "The logbook for jar batch M lists the temperature excursion as beginning at 09:15 on day 6.", "The logbook for jar batch M lists the temperature excursion as ending at 11:10 on day 6."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged protocol requirements and 90-minute exception policy, keep the same jar-set-Delta entity and 14-day timeframe, and the two focus evidence items are plain factual timestamp statements rather than policy text; the counterfactual merely extends the excursion end time to 10:10, yielding a coherent 115-minute duration without contradicting other stated facts, and no gold answer or rule table is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\": \"Case note: A soil scientist is reviewing a technician's 14-day moisture microcosm replication. The chamber log confirms all jars began at 61% water-holding capacity, within the 58-62% range, and all three required CO₂ measurements were recorded on days 3, 7, and 14. Temperature stayed between 19°C and 21°C at all times except for a single excursion, which the operator's log shows affected every treatment and paired control in jar set Delta simultaneously. No other deviations were noted during incubation.\", \"evidence\": [\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\", \"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\", \"The logged start time of the single temperature excursion in jar set Delta was 08:15.\", \"The logged end time of the single temperature excursion in jar set Delta was 09:30.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "2"], "text": "The logged start time of the single temperature excursion in jar set Delta was 08:15."}, {"path": ["evidence", "3"], "text": "The logged end time of the single temperature excursion in jar set Delta was 09:30."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logged start time of the single temperature excursion in jar set Delta was 08:15.", "negative_left": "The logged start time of the single temperature excursion in jar set Delta was 08:15.", "negative_right": "The logged end time of the single temperature excursion in jar set Delta was 10:10.", "right": "The logged end time of the single temperature excursion in jar set Delta was 09:30."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-004", "id": "fast-43-diverse-267-004-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "Case note: A soil scientist is reviewing a technician's 14-day moisture microcosm replication. The chamber log confirms all jars began at 61% water-holding capacity, within the 58-62% range, and all three required CO₂ measurements were recorded on days 3, 7, and 14. Temperature stayed between 19°C and 21°C at all times except for a single excursion, which the operator's log shows affected every treatment and paired control in jar set Delta simultaneously. No other deviations were noted during incubation.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "The logged start time of the single temperature excursion in jar set Delta was 08:15.", "The logged end time of the single temperature excursion in jar set Delta was 09:30."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the unchanged protocol requirements and 90-minute exception policy, keep the same jar-set-Delta entity and 14-day timeframe, and the two focus evidence items are plain factual timestamp statements rather than policy text; the counterfactual merely extends the excursion end time to 10:10, yielding a coherent 115-minute duration without contradicting other stated facts, and no gold answer or rule table is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\": \"Case note: A soil scientist is reviewing a technician's 14-day moisture microcosm replication. The chamber log confirms all jars began at 61% water-holding capacity, within the 58-62% range, and all three required CO₂ measurements were recorded on days 3, 7, and 14. Temperature stayed between 19°C and 21°C at all times except for a single excursion, which the operator's log shows affected every treatment and paired control in jar set Delta simultaneously. No other deviations were noted during incubation.\", \"evidence\": [\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\", \"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\", \"The logged start time of the single temperature excursion in jar set Delta was 08:15.\", \"The logged end time of the single temperature excursion in jar set Delta was 09:30.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "2"], "text": "The logged start time of the single temperature excursion in jar set Delta was 08:15."}, {"path": ["evidence", "3"], "text": "The logged end time of the single temperature excursion in jar set Delta was 09:30."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logged start time of the single temperature excursion in jar set Delta was 08:15.", "negative_left": "The logged start time of the single temperature excursion in jar set Delta was 08:15.", "negative_right": "The logged end time of the single temperature excursion in jar set Delta was 10:10.", "right": "The logged end time of the single temperature excursion in jar set Delta was 09:30."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-004", "id": "fast-43-diverse-267-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "Case note: A soil scientist is reviewing a technician's 14-day moisture microcosm replication. The chamber log confirms all jars began at 61% water-holding capacity, within the 58-62% range, and all three required CO₂ measurements were recorded on days 3, 7, and 14. Temperature stayed between 19°C and 21°C at all times except for a single excursion, which the operator's log shows affected every treatment and paired control in jar set Delta simultaneously. No other deviations were noted during incubation.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "The logged start time of the single temperature excursion in jar set Delta was 08:15.", "The logged end time of the single temperature excursion in jar set Delta was 10:10."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the protocol requirement and 90-minute exception verbatim, keep the same jar-set/14-day/day-6 bindings, the focus evidence is two complete factual sentences about excursion start/end times, the counterfactual only changes the end time to 11:30 (135 min) which is internally consistent with the single shared excursion narrative, and no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"Case note: Reviewer audit of a 14-day moisture microcosm replication for jar set B. All jars began between 58% and 62% water-holding capacity, and CO₂ readings were logged on days 3, 7, and 14 as required. The chamber log shows temperatures held between 19°C and 21°C throughout incubation except for a single shared excursion affecting every treatment and paired control on incubation day 6.\",\"log_entries\":[\"All jars in the set began at 61% water-holding capacity, within the 58–62% range.\",\"CO₂ measurements were recorded on days 3, 7, and 14 as scheduled.\",\"Temperature stayed within 19–21°C at all other times outside the single excursion.\",\"The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6.\",\"The logbook for jar set B records the single temperature excursion as ending at 10:30 on incubation day 6.\",\"Only one temperature excursion occurred during the full incubation period, and it affected every treatment and paired control simultaneously.\",\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["log_entries", "3"], "text": "The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6."}, {"path": ["log_entries", "4"], "text": "The logbook for jar set B records the single temperature excursion as ending at 10:30 on incubation day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6.", "negative_left": "The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6.", "negative_right": "The logbook for jar set B records the single temperature excursion as ending at 11:30 on incubation day 6.", "right": "The logbook for jar set B records the single temperature excursion as ending at 10:30 on incubation day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-018", "id": "fast-43-diverse-267-018-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "Case note: Reviewer audit of a 14-day moisture microcosm replication for jar set B. All jars began between 58% and 62% water-holding capacity, and CO₂ readings were logged on days 3, 7, and 14 as required. The chamber log shows temperatures held between 19°C and 21°C throughout incubation except for a single shared excursion affecting every treatment and paired control on incubation day 6.", "log_entries": ["All jars in the set began at 61% water-holding capacity, within the 58–62% range.", "CO₂ measurements were recorded on days 3, 7, and 14 as scheduled.", "Temperature stayed within 19–21°C at all other times outside the single excursion.", "The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6.", "The logbook for jar set B records the single temperature excursion as ending at 10:30 on incubation day 6.", "Only one temperature excursion occurred during the full incubation period, and it affected every treatment and paired control simultaneously.", "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the protocol requirement and 90-minute exception verbatim, keep the same jar-set/14-day/day-6 bindings, the focus evidence is two complete factual sentences about excursion start/end times, the counterfactual only changes the end time to 11:30 (135 min) which is internally consistent with the single shared excursion narrative, and no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"Case note: Reviewer audit of a 14-day moisture microcosm replication for jar set B. All jars began between 58% and 62% water-holding capacity, and CO₂ readings were logged on days 3, 7, and 14 as required. The chamber log shows temperatures held between 19°C and 21°C throughout incubation except for a single shared excursion affecting every treatment and paired control on incubation day 6.\",\"log_entries\":[\"All jars in the set began at 61% water-holding capacity, within the 58–62% range.\",\"CO₂ measurements were recorded on days 3, 7, and 14 as scheduled.\",\"Temperature stayed within 19–21°C at all other times outside the single excursion.\",\"The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6.\",\"The logbook for jar set B records the single temperature excursion as ending at 10:30 on incubation day 6.\",\"Only one temperature excursion occurred during the full incubation period, and it affected every treatment and paired control simultaneously.\",\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["log_entries", "3"], "text": "The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6."}, {"path": ["log_entries", "4"], "text": "The logbook for jar set B records the single temperature excursion as ending at 10:30 on incubation day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6.", "negative_left": "The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6.", "negative_right": "The logbook for jar set B records the single temperature excursion as ending at 11:30 on incubation day 6.", "right": "The logbook for jar set B records the single temperature excursion as ending at 10:30 on incubation day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-018", "id": "fast-43-diverse-267-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "Case note: Reviewer audit of a 14-day moisture microcosm replication for jar set B. All jars began between 58% and 62% water-holding capacity, and CO₂ readings were logged on days 3, 7, and 14 as required. The chamber log shows temperatures held between 19°C and 21°C throughout incubation except for a single shared excursion affecting every treatment and paired control on incubation day 6.", "log_entries": ["All jars in the set began at 61% water-holding capacity, within the 58–62% range.", "CO₂ measurements were recorded on days 3, 7, and 14 as scheduled.", "Temperature stayed within 19–21°C at all other times outside the single excursion.", "The logbook for jar set B records the single temperature excursion as beginning at 09:15 on incubation day 6.", "The logbook for jar set B records the single temperature excursion as ending at 11:30 on incubation day 6.", "Only one temperature excursion occurred during the full incubation period, and it affected every treatment and paired control simultaneously.", "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol thresholds, exception rule, and measurement facts without alteration; the only change is the excursion end time (10:32 -> 11:32), shifting duration from 78 to 138 minutes, which is a coherent factual variation, not a contradiction; no answer, rule table, or rationale is embedded, and question entity/path/time bindings are unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"A soil scientist is reviewing a technician’s 14-day moisture microcosm replication. An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"All jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14.\",\"Aside from one shared excursion affecting every treatment and paired control, the incubation temperature stayed between 19°C and 21°C throughout the 14-day run.\",\"The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6.\",\"The logbook for the incubation run records the single temperature excursion ending at 10:32 on day 6.\"],\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "4"], "text": "The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6."}, {"path": ["evidence", "5"], "text": "The logbook for the incubation run records the single temperature excursion ending at 10:32 on day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6.", "negative_left": "The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6.", "negative_right": "The logbook for the incubation run records the single temperature excursion ending at 11:32 on day 6.", "right": "The logbook for the incubation run records the single temperature excursion ending at 10:32 on day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-021", "id": "fast-43-diverse-267-021-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a technician’s 14-day moisture microcosm replication. An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "All jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14.", "Aside from one shared excursion affecting every treatment and paired control, the incubation temperature stayed between 19°C and 21°C throughout the 14-day run.", "The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6.", "The logbook for the incubation run records the single temperature excursion ending at 10:32 on day 6."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol thresholds, exception rule, and measurement facts without alteration; the only change is the excursion end time (10:32 -> 11:32), shifting duration from 78 to 138 minutes, which is a coherent factual variation, not a contradiction; no answer, rule table, or rationale is embedded, and question entity/path/time bindings are unchanged.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"A soil scientist is reviewing a technician’s 14-day moisture microcosm replication. An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"All jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14.\",\"Aside from one shared excursion affecting every treatment and paired control, the incubation temperature stayed between 19°C and 21°C throughout the 14-day run.\",\"The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6.\",\"The logbook for the incubation run records the single temperature excursion ending at 10:32 on day 6.\"],\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "4"], "text": "The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6."}, {"path": ["evidence", "5"], "text": "The logbook for the incubation run records the single temperature excursion ending at 10:32 on day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6.", "negative_left": "The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6.", "negative_right": "The logbook for the incubation run records the single temperature excursion ending at 11:32 on day 6.", "right": "The logbook for the incubation run records the single temperature excursion ending at 10:32 on day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-021", "id": "fast-43-diverse-267-021-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a technician’s 14-day moisture microcosm replication. An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "All jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14.", "Aside from one shared excursion affecting every treatment and paired control, the incubation temperature stayed between 19°C and 21°C throughout the 14-day run.", "The logbook for the incubation run records the single temperature excursion beginning at 09:14 on day 6.", "The logbook for the incubation run records the single temperature excursion ending at 11:32 on day 6."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the two governing policy statements verbatim and the original question is unchanged, preserving entity/path/time bindings for jar B12 day 6; the focus evidence consists of two complete factual sentences about excursion start/end times; the counterfactual coherently shifts the end time to 11:29, making the excursion 135 minutes (vs 75 minutes in base) without contradicting any other logged fact; neither context states or implies the final yes/no verdict.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"experiment\":\"14-day moisture microcosm replication, jar B12 series\",\"jars\":\"all began between 58% and 62% water-holding capacity\",\"co2_measurements\":\"recorded on days 3, 7, and 14\",\"temperature_log\":\"maintained between 19°C and 21°C except one shared excursion affecting all treatments and paired controls\",\"policy\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\"],\"excursion_log\":{\"start\":\"The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6.\",\"end\":\"The logged end time of the single temperature excursion in jar B12 was 10:29 on incubation day 6.\"}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["excursion_log", "start"], "text": "The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6."}, {"path": ["excursion_log", "end"], "text": "The logged end time of the single temperature excursion in jar B12 was 10:29 on incubation day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6.", "negative_left": "The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6.", "negative_right": "The logged end time of the single temperature excursion in jar B12 was 11:29 on incubation day 6.", "right": "The logged end time of the single temperature excursion in jar B12 was 10:29 on incubation day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-022", "id": "fast-43-diverse-267-022-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"co2_measurements": "recorded on days 3, 7, and 14", "excursion_log": {"end": "The logged end time of the single temperature excursion in jar B12 was 10:29 on incubation day 6.", "start": "The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6."}, "experiment": "14-day moisture microcosm replication, jar B12 series", "jars": "all began between 58% and 62% water-holding capacity", "policy": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."], "temperature_log": "maintained between 19°C and 21°C except one shared excursion affecting all treatments and paired controls"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the two governing policy statements verbatim and the original question is unchanged, preserving entity/path/time bindings for jar B12 day 6; the focus evidence consists of two complete factual sentences about excursion start/end times; the counterfactual coherently shifts the end time to 11:29, making the excursion 135 minutes (vs 75 minutes in base) without contradicting any other logged fact; neither context states or implies the final yes/no verdict.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"experiment\":\"14-day moisture microcosm replication, jar B12 series\",\"jars\":\"all began between 58% and 62% water-holding capacity\",\"co2_measurements\":\"recorded on days 3, 7, and 14\",\"temperature_log\":\"maintained between 19°C and 21°C except one shared excursion affecting all treatments and paired controls\",\"policy\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\"],\"excursion_log\":{\"start\":\"The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6.\",\"end\":\"The logged end time of the single temperature excursion in jar B12 was 10:29 on incubation day 6.\"}}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["excursion_log", "start"], "text": "The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6."}, {"path": ["excursion_log", "end"], "text": "The logged end time of the single temperature excursion in jar B12 was 10:29 on incubation day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6.", "negative_left": "The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6.", "negative_right": "The logged end time of the single temperature excursion in jar B12 was 11:29 on incubation day 6.", "right": "The logged end time of the single temperature excursion in jar B12 was 10:29 on incubation day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-022", "id": "fast-43-diverse-267-022-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"co2_measurements": "recorded on days 3, 7, and 14", "excursion_log": {"end": "The logged end time of the single temperature excursion in jar B12 was 11:29 on incubation day 6.", "start": "The logged start time of the single temperature excursion in jar B12 was 09:14 on incubation day 6."}, "experiment": "14-day moisture microcosm replication, jar B12 series", "jars": "all began between 58% and 62% water-holding capacity", "policy": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."], "temperature_log": "maintained between 19°C and 21°C except one shared excursion affecting all treatments and paired controls"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol and exception policy verbatim, keep the same Jar Set 7 entity and day-5 time binding, use two factual timestamp sentences as evidence, and the counterfactual coherently changes only the end time (09:12–11:30, 138 min) without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\": \"A soil scientist is reviewing a technician's 14-day moisture microcosm replication (Jar Set 7). The chamber log shows all jars began at 61% water-holding capacity, and all required CO\\u2082 measurements were recorded on days 3, 7, and 14. A single shared temperature excursion affected every treatment and paired control on day 5, with temperature remaining between 19\\u00b0C and 21\\u00b0C at all other times. The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation. The single temperature excursion in Jar Set 7 was logged as ending at 10:30 on day 5 of incubation. The reviewer must determine whether this excursion satisfies the incubation exception before confirming protocol compliance.\", \"facts\": {\"A1\": \"supported\", \"A2\": \"supported\", \"A3\": \"supported\", \"A4\": \"supported\", \"A5\": \"supported\", \"A6\": \"supported\", \"A7\": \"supported\", \"A8\": \"supported\", \"A10\": \"supported\"}, \"policy\": [\"The protocol covers 14-day moisture microcosms and requires 60\\u00b12% water-holding capacity, 20\\u00b11\\u00b0C incubation, and CO\\u2082 measurements on days 3, 7, and 14.\", \"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation."}, {"path": ["context"], "text": "The single temperature excursion in Jar Set 7 was logged as ending at 10:30 on day 5 of incubation."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation.", "negative_left": "The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation.", "negative_right": "The single temperature excursion in Jar Set 7 was logged as ending at 11:30 on day 5 of incubation.", "right": "The single temperature excursion in Jar Set 7 was logged as ending at 10:30 on day 5 of incubation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-023", "id": "fast-43-diverse-267-023-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a technician's 14-day moisture microcosm replication (Jar Set 7). The chamber log shows all jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14. A single shared temperature excursion affected every treatment and paired control on day 5, with temperature remaining between 19°C and 21°C at all other times. The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation. The single temperature excursion in Jar Set 7 was logged as ending at 10:30 on day 5 of incubation. The reviewer must determine whether this excursion satisfies the incubation exception before confirming protocol compliance.", "facts": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "policy": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original protocol and exception policy verbatim, keep the same Jar Set 7 entity and day-5 time binding, use two factual timestamp sentences as evidence, and the counterfactual coherently changes only the end time (09:12–11:30, 138 min) without contradicting other facts or leaking the answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\": \"A soil scientist is reviewing a technician's 14-day moisture microcosm replication (Jar Set 7). The chamber log shows all jars began at 61% water-holding capacity, and all required CO\\u2082 measurements were recorded on days 3, 7, and 14. A single shared temperature excursion affected every treatment and paired control on day 5, with temperature remaining between 19\\u00b0C and 21\\u00b0C at all other times. The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation. The single temperature excursion in Jar Set 7 was logged as ending at 10:30 on day 5 of incubation. The reviewer must determine whether this excursion satisfies the incubation exception before confirming protocol compliance.\", \"facts\": {\"A1\": \"supported\", \"A2\": \"supported\", \"A3\": \"supported\", \"A4\": \"supported\", \"A5\": \"supported\", \"A6\": \"supported\", \"A7\": \"supported\", \"A8\": \"supported\", \"A10\": \"supported\"}, \"policy\": [\"The protocol covers 14-day moisture microcosms and requires 60\\u00b12% water-holding capacity, 20\\u00b11\\u00b0C incubation, and CO\\u2082 measurements on days 3, 7, and 14.\", \"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\"]}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["context"], "text": "The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation."}, {"path": ["context"], "text": "The single temperature excursion in Jar Set 7 was logged as ending at 10:30 on day 5 of incubation."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation.", "negative_left": "The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation.", "negative_right": "The single temperature excursion in Jar Set 7 was logged as ending at 11:30 on day 5 of incubation.", "right": "The single temperature excursion in Jar Set 7 was logged as ending at 10:30 on day 5 of incubation."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-023", "id": "fast-43-diverse-267-023-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a technician's 14-day moisture microcosm replication (Jar Set 7). The chamber log shows all jars began at 61% water-holding capacity, and all required CO₂ measurements were recorded on days 3, 7, and 14. A single shared temperature excursion affected every treatment and paired control on day 5, with temperature remaining between 19°C and 21°C at all other times. The single temperature excursion in Jar Set 7 was logged as starting at 09:12 on day 5 of incubation. The single temperature excursion in Jar Set 7 was logged as ending at 11:30 on day 5 of incubation. The reviewer must determine whether this excursion satisfies the incubation exception before confirming protocol compliance.", "facts": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported"}, "policy": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."]}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full protocol and exception text verbatim, preserve the same jar set B entity, 14-day timeframe, and request, and the focus evidence pair are complete factual timestamp sentences; the counterfactual only changes the excursion end time (10:30→12:30), shifting duration from 78 to 198 minutes without contradicting any other stated fact, and neither context states or implies the compliance verdict.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"A soil scientist is reviewing a technician's 14-day moisture microcosm replication (jar set B). An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"All jars began at 61% water-holding capacity, and CO₂ was measured on days 3, 7, and 14 as required.\",\"Outside the single excursion window, the logger recorded temperatures continuously between 19°C and 21°C for the full 14-day run.\",\"The single excursion was recorded as affecting every treatment jar and its paired control jar simultaneously.\",\"The temperature logger for jar set B recorded the excursion start at 09:12 on day 6.\",\"The temperature logger for jar set B recorded the excursion end at 10:30 on day 6.\"],\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "The temperature logger for jar set B recorded the excursion start at 09:12 on day 6."}, {"path": ["evidence", "6"], "text": "The temperature logger for jar set B recorded the excursion end at 10:30 on day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The temperature logger for jar set B recorded the excursion start at 09:12 on day 6.", "negative_left": "The temperature logger for jar set B recorded the excursion start at 09:12 on day 6.", "negative_right": "The temperature logger for jar set B recorded the excursion end at 12:30 on day 6.", "right": "The temperature logger for jar set B recorded the excursion end at 10:30 on day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-026", "id": "fast-43-diverse-267-026-base", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a technician's 14-day moisture microcosm replication (jar set B). An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "All jars began at 61% water-holding capacity, and CO₂ was measured on days 3, 7, and 14 as required.", "Outside the single excursion window, the logger recorded temperatures continuously between 19°C and 21°C for the full 14-day run.", "The single excursion was recorded as affecting every treatment jar and its paired control jar simultaneously.", "The temperature logger for jar set B recorded the excursion start at 09:12 on day 6.", "The temperature logger for jar set B recorded the excursion end at 10:30 on day 6."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full protocol and exception text verbatim, preserve the same jar set B entity, 14-day timeframe, and request, and the focus evidence pair are complete factual timestamp sentences; the counterfactual only changes the excursion end time (10:30→12:30), shifting duration from 78 to 198 minutes without contradicting any other stated fact, and neither context states or implies the compliance verdict.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "full_context_fact_states": {"base": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "supported"}, "counterfactual": {"A1": "supported", "A10": "supported", "A2": "supported", "A3": "supported", "A4": "supported", "A5": "supported", "A6": "supported", "A7": "supported", "A8": "supported", "A9": "refuted"}, "remove_left": {"A9": "unknown"}, "remove_right": {"A9": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A9": "unknown"}, "negative_pair": {"A9": "refuted"}, "negative_sentence": {"A9": "unknown"}, "positive_pair": {"A9": "supported"}, "right": {"A9": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; the universally quantified jar condition remains atomic. A9 is a factual duration-threshold proposition rather than a policy conclusion. The base and counter assignments are jointly realizable while changing only A9: the same single shared excursion can last at most 90 minutes in the base and more than 90 minutes in the counter, with all other facts unchanged. The policy evidence correctly preserves the substantive scope, protocol requirements, and temperature-excursion exception originating in the original state.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes protocol scope, compliant starting water-holding capacity, all required CO₂ measurement days, compliant temperature outside exactly one excursion, and every condition needed for that excursion to be covered by the stated exception. It is sufficient for qualification.", "rule_index": 0, "sound": true}, {"reason": "A9 being refuted entails that the recorded excursion lasted more than 90 consecutive minutes. For an in-scope 14-day moisture microcosm with such an excursion, the explicit policy says any longer excursion fails replication regardless of shared exposure, so the conjunction is sufficient for a false decision.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "The experiment under review is a moisture microcosm."}, {"id": "A2", "statement": "The experiment under review has an incubation duration of 14 days."}, {"id": "A3", "statement": "Every jar in the experiment under review began between 58% and 62% water-holding capacity, inclusive."}, {"id": "A4", "statement": "A CO₂ measurement for the experiment under review was recorded on day 3."}, {"id": "A5", "statement": "A CO₂ measurement for the experiment under review was recorded on day 7."}, {"id": "A6", "statement": "A CO₂ measurement for the experiment under review was recorded on day 14."}, {"id": "A7", "statement": "Exactly one consecutive temperature excursion outside 19°C to 21°C occurred during the incubation of the experiment under review."}, {"id": "A8", "statement": "The single temperature excursion during the experiment under review affected every treatment and paired control."}, {"id": "A9", "statement": "The elapsed time from the recorded start of the single temperature excursion to its recorded end was no more than 90 consecutive minutes."}, {"id": "A10", "statement": "At every time outside the single temperature excursion, the incubation temperature of the experiment under review was between 19°C and 21°C, inclusive."}], "base_state_json": "{\"context\":\"A soil scientist is reviewing a technician's 14-day moisture microcosm replication (jar set B). An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.\",\"evidence\":[\"The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.\",\"Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.\",\"All jars began at 61% water-holding capacity, and CO₂ was measured on days 3, 7, and 14 as required.\",\"Outside the single excursion window, the logger recorded temperatures continuously between 19°C and 21°C for the full 14-day run.\",\"The single excursion was recorded as affecting every treatment jar and its paired control jar simultaneously.\",\"The temperature logger for jar set B recorded the excursion start at 09:12 on day 6.\",\"The temperature logger for jar set B recorded the excursion end at 10:30 on day 6.\"],\"request\":\"Does this experiment qualify as an acceptable replication under the stated protocol?\"}", "base_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}, {"atom_id": "A10", "state": "supported"}], "focus_atom": "A9", "focus_evidence": [{"path": ["evidence", "5"], "text": "The temperature logger for jar set B recorded the excursion start at 09:12 on day 6."}, {"path": ["evidence", "6"], "text": "The temperature logger for jar set B recorded the excursion end at 10:30 on day 6."}], "policy_evidence": [{"path": ["evidence", "0"], "text": "The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14."}, {"path": ["evidence", "1"], "text": "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure."}], "rules": [{"justification": "The experiment is within the protocol's 14-day moisture-microcosm scope; its water-holding capacity and all three required CO₂ measurements comply; temperatures comply outside exactly one shared excursion; and that excursion satisfies both the 90-minute limit and the all-treatment-and-control exposure condition.", "target": "true", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}, {"atom_id": "A4", "state": "supported"}, {"atom_id": "A5", "state": "supported"}, {"atom_id": "A6", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A8", "state": "supported"}, {"atom_id": "A9", "state": "supported"}, {"atom_id": "A10", "state": "supported"}]}, {"justification": "For an in-scope 14-day moisture microcosm with a temperature excursion, refutation of the proposition that the excursion lasted no more than 90 consecutive minutes entails a longer excursion. The policy states that any longer excursion fails replication regardless of shared exposure.", "target": "false", "when": [{"atom_id": "A1", "state": "supported"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A7", "state": "supported"}, {"atom_id": "A9", "state": "refuted"}]}]}, "verified_pair": {"left": "The temperature logger for jar set B recorded the excursion start at 09:12 on day 6.", "negative_left": "The temperature logger for jar set B recorded the excursion start at 09:12 on day 6.", "negative_right": "The temperature logger for jar set B recorded the excursion end at 12:30 on day 6.", "right": "The temperature logger for jar set B recorded the excursion end at 10:30 on day 6."}, "verifier_independent_model": false}, "family": "fast-43-diverse-267-026", "id": "fast-43-diverse-267-026-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — the experiment does not qualify because at least one applicable requirement is unmet and not covered by an exception.", "true": "Yes — the experiment qualifies as an acceptable protocol-compliant replication."}, "instructions": "Perform a verification decision using only the supplied scope, exception, and observations. Return yes only if every applicable protocol requirement is met or covered by an exception; otherwise return no.", "type": "noul"}}, "state": {"context": "A soil scientist is reviewing a technician's 14-day moisture microcosm replication (jar set B). An instrument operator supplied the chamber log, and a replication reviewer must verify protocol compliance.", "evidence": ["The protocol covers 14-day moisture microcosms and requires 60±2% water-holding capacity, 20±1°C incubation, and CO₂ measurements on days 3, 7, and 14.", "Incubation exception: one shared temperature excursion may be accepted only if it lasts no more than 90 consecutive minutes and affects every treatment and paired control. Any longer excursion fails replication, regardless of shared exposure.", "All jars began at 61% water-holding capacity, and CO₂ was measured on days 3, 7, and 14 as required.", "Outside the single excursion window, the logger recorded temperatures continuously between 19°C and 21°C for the full 14-day run.", "The single excursion was recorded as affecting every treatment jar and its paired control jar simultaneously.", "The temperature logger for jar set B recorded the excursion start at 09:12 on day 6.", "The temperature logger for jar set B recorded the excursion end at 12:30 on day 6."], "request": "Does this experiment qualify as an acceptable replication under the stated protocol?"}}, "method": "c2d", "provenance": {"source_id": "diverse-267", "source_is_synthetic": true, "source_sha256": "63c9d39d573853cd106ca2e1e4840d83adda5003872dd716ba661a469b0bd61c", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing rubric and jar roles unchanged, only the correction day (Friday vs Thursday) differs, which coherently shifts the outage duration from exactly 72 hours to 48 hours without contradicting other stated facts, and neither context states or implies the routing outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"soil scientist\",\"text\":\"In the decomposition trial, the set of jars is exactly Cedar and Elm. Cedar is a sealed test jar, and Elm is excluded from replication judgments.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar was corrected at 9:00 AM on the following Friday.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar had exactly one continuous episode outside 55–65% water-holding capacity, beginning with that first low reading and ending at correction.\"},{\"speaker\":\"instrument operator\",\"text\":\"Every sealing check for Cedar during the trial passed, and every instrument check relevant to Cedar's moisture readings passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday."}, {"path": ["2", "text"], "text": "Cedar was corrected at 9:00 AM on the following Friday."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday.", "negative_left": "Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday.", "negative_right": "Cedar was corrected at 9:00 AM on the following Thursday.", "right": "Cedar was corrected at 9:00 AM on the following Friday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-268-018", "id": "fast-43-diverse-268-018-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "In the decomposition trial, the set of jars is exactly Cedar and Elm. Cedar is a sealed test jar, and Elm is excluded from replication judgments."}, {"speaker": "microcosm technician", "text": "Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday."}, {"speaker": "microcosm technician", "text": "Cedar was corrected at 9:00 AM on the following Friday."}, {"speaker": "microcosm technician", "text": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity, beginning with that first low reading and ending at correction."}, {"speaker": "instrument operator", "text": "Every sealing check for Cedar during the trial passed, and every instrument check relevant to Cedar's moisture readings passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing rubric and jar roles unchanged, only the correction day (Friday vs Thursday) differs, which coherently shifts the outage duration from exactly 72 hours to 48 hours without contradicting other stated facts, and neither context states or implies the routing outcome.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\":\"soil scientist\",\"text\":\"In the decomposition trial, the set of jars is exactly Cedar and Elm. Cedar is a sealed test jar, and Elm is excluded from replication judgments.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar was corrected at 9:00 AM on the following Friday.\"},{\"speaker\":\"microcosm technician\",\"text\":\"Cedar had exactly one continuous episode outside 55–65% water-holding capacity, beginning with that first low reading and ending at correction.\"},{\"speaker\":\"instrument operator\",\"text\":\"Every sealing check for Cedar during the trial passed, and every instrument check relevant to Cedar's moisture readings passed.\"},{\"speaker\":\"replication reviewer\",\"text\":\"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday."}, {"path": ["2", "text"], "text": "Cedar was corrected at 9:00 AM on the following Friday."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday.", "negative_left": "Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday.", "negative_right": "Cedar was corrected at 9:00 AM on the following Thursday.", "right": "Cedar was corrected at 9:00 AM on the following Friday."}, "verifier_independent_model": false}, "family": "fast-43-diverse-268-018", "id": "fast-43-diverse-268-018-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "In the decomposition trial, the set of jars is exactly Cedar and Elm. Cedar is a sealed test jar, and Elm is excluded from replication judgments."}, {"speaker": "microcosm technician", "text": "Cedar's first low moisture reading on day 2 was logged at 9:00 AM on Tuesday."}, {"speaker": "microcosm technician", "text": "Cedar was corrected at 9:00 AM on the following Thursday."}, {"speaker": "microcosm technician", "text": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity, beginning with that first low reading and ending at correction."}, {"speaker": "instrument operator", "text": "Every sealing check for Cedar during the trial passed, and every instrument check relevant to Cedar's moisture readings passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing rubric and entity/time bindings unchanged from the original; the counterfactual only shifts the correction day from 5 to 4, yielding a coherent 48-hour episode instead of 72 hours without contradicting other facts; the focus evidence consists of two complete factual sentences about readings, not policy text; neither context states or implies the routing decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\": \"soil scientist\", \"text\": \"In the decomposition trial, the jars are exactly Cedar and Elm. Cedar is a sealed test jar and Elm is a balance-check jar, excluded from replication judgments.\"}, {\"speaker\": \"microcosm technician\", \"text\": \"Cedar's first low moisture reading on day 2 was logged at 8:00 AM.\"}, {\"speaker\": \"microcosm technician\", \"text\": \"Cedar's correction was logged at 8:00 AM on day 5. This was Cedar's only continuous episode outside 55–65% water-holding capacity, beginning with the first low reading and ending at correction.\"}, {\"speaker\": \"instrument operator\", \"text\": \"Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed.\"}, {\"speaker\": \"replication reviewer\", \"text\": \"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 was logged at 8:00 AM."}, {"path": ["2", "text"], "text": "Cedar's correction was logged at 8:00 AM on day 5."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was logged at 8:00 AM.", "negative_left": "Cedar's first low moisture reading on day 2 was logged at 8:00 AM.", "negative_right": "Cedar's correction was logged at 8:00 AM on day 4.", "right": "Cedar's correction was logged at 8:00 AM on day 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-268-024", "id": "fast-43-diverse-268-024-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "In the decomposition trial, the jars are exactly Cedar and Elm. Cedar is a sealed test jar and Elm is a balance-check jar, excluded from replication judgments."}, {"speaker": "microcosm technician", "text": "Cedar's first low moisture reading on day 2 was logged at 8:00 AM."}, {"speaker": "microcosm technician", "text": "Cedar's correction was logged at 8:00 AM on day 5. This was Cedar's only continuous episode outside 55–65% water-holding capacity, beginning with the first low reading and ending at correction."}, {"speaker": "instrument operator", "text": "Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "science-05", "split": "train", "variant": "base"} {"domain": "science", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing rubric and entity/time bindings unchanged from the original; the counterfactual only shifts the correction day from 5 to 4, yielding a coherent 48-hour episode instead of 72 hours without contradicting other facts; the focus evidence consists of two complete factual sentences about readings, not policy text; neither context states or implies the routing decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "counterfactual": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "refuted", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "supported"}, "remove_left": {"a4": "unknown"}, "remove_right": {"a4": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a4": "unknown"}, "negative_pair": {"a4": "refuted"}, "negative_sentence": {"a4": "unknown"}, "positive_pair": {"a4": "supported"}, "right": {"a4": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship; the universally quantified check atoms remain atomic. The focus atom is a factual elapsed-time relation, not a policy conclusion. The base and counter assignments are jointly realizable with only the correction timing—and thus a4—changing while the same unique episode and endpoints remain supported. Policy evidence correctly preserves the substantive routing rubric originating in the original state; instructions already retained in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conditions establish that Cedar is the only relevant non-excluded sealed test jar, that its sole continuous out-of-range episode ran from the first low reading through correction for at least 72 hours, and that no relevant sealing or instrument check failed. This is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting “at least 72 hours” entails a duration below 72 hours. Together with the episode’s stated endpoints, its uniqueness, the exhaustive jar set, and Elm’s exclusion, the conditions exclude every qualifying non-excluded jar and support the false outcome.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The set of jars in the decomposition trial is exactly {Cedar, Elm}."}, {"id": "a2", "statement": "Cedar is a sealed test jar in the decomposition trial."}, {"id": "a3", "statement": "Elm is excluded from replication judgments in the decomposition trial."}, {"id": "a4", "statement": "At least 72 hours elapsed between Cedar's first low moisture reading on day 2 and Cedar's correction."}, {"id": "a5", "statement": "Cedar had exactly one continuous episode outside 55–65% water-holding capacity during the incubation."}, {"id": "a6", "statement": "Cedar's continuous out-of-range episode began with Cedar's first low moisture reading on day 2."}, {"id": "a7", "statement": "Cedar's continuous out-of-range episode ended when Cedar was corrected."}, {"id": "a8", "statement": "Every sealing check for Cedar during the decomposition trial passed."}, {"id": "a9", "statement": "Every instrument check relevant to Cedar's moisture readings during the decomposition trial passed."}], "base_state_json": "[{\"speaker\": \"soil scientist\", \"text\": \"In the decomposition trial, the jars are exactly Cedar and Elm. Cedar is a sealed test jar and Elm is a balance-check jar, excluded from replication judgments.\"}, {\"speaker\": \"microcosm technician\", \"text\": \"Cedar's first low moisture reading on day 2 was logged at 8:00 AM.\"}, {\"speaker\": \"microcosm technician\", \"text\": \"Cedar's correction was logged at 8:00 AM on day 5. This was Cedar's only continuous episode outside 55–65% water-holding capacity, beginning with the first low reading and ending at correction.\"}, {\"speaker\": \"instrument operator\", \"text\": \"Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed.\"}, {\"speaker\": \"replication reviewer\", \"text\": \"Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}], "focus_atom": "a4", "focus_evidence": [{"path": ["1", "text"], "text": "Cedar's first low moisture reading on day 2 was logged at 8:00 AM."}, {"path": ["2", "text"], "text": "Cedar's correction was logged at 8:00 AM on day 5."}], "policy_evidence": [{"path": ["3", "text"], "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}], "rules": [{"justification": "Cedar is the only non-excluded sealed test jar, its sole continuous out-of-range episode lasted from its first low reading until correction for at least 72 hours, and no failed sealing or instrument check triggers a competing issue.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}, {"justification": "Cedar's sole continuous out-of-range episode ended at correction less than 72 hours after it began, and Elm is the only other jar and is excluded, so no non-excluded sealed test jar satisfies the duration threshold.", "target": "false", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "refuted"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "supported"}]}]}, "verified_pair": {"left": "Cedar's first low moisture reading on day 2 was logged at 8:00 AM.", "negative_left": "Cedar's first low moisture reading on day 2 was logged at 8:00 AM.", "negative_right": "Cedar's correction was logged at 8:00 AM on day 4.", "right": "Cedar's correction was logged at 8:00 AM on day 5."}, "verifier_independent_model": false}, "family": "fast-43-diverse-268-024", "id": "fast-43-diverse-268-024-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route the case as an incubation issue.", "true": "Yes — route the case as an incubation issue."}, "instructions": "Using the stated routing rubric and the dialogue’s coreferences, decide whether this case should be routed as an incubation issue.", "type": "noul"}}, "state": [{"speaker": "soil scientist", "text": "In the decomposition trial, the jars are exactly Cedar and Elm. Cedar is a sealed test jar and Elm is a balance-check jar, excluded from replication judgments."}, {"speaker": "microcosm technician", "text": "Cedar's first low moisture reading on day 2 was logged at 8:00 AM."}, {"speaker": "microcosm technician", "text": "Cedar's correction was logged at 8:00 AM on day 4. This was Cedar's only continuous episode outside 55–65% water-holding capacity, beginning with the first low reading and ending at correction."}, {"speaker": "instrument operator", "text": "Every sealing check for Cedar passed, and every instrument check relevant to Cedar's moisture readings passed."}, {"speaker": "replication reviewer", "text": "Route as an incubation issue when a sealed test jar remains outside 55–65% for at least 72 hours; exactly 72 qualifies. Ignore excluded jars. Setup and measurement issues take precedence only when sealing or instrument checks fail."}]}, "method": "c2d", "provenance": {"source_id": "diverse-268", "source_is_synthetic": true, "source_sha256": "cf535c50ff2ac7ae12e870568504167ce104bf56399825d7e261d9af3fc4cc97", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "science-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate the same routing/threshold policy verbatim, matching the original state's rubric, and the PO identifier binding (PO-731) is preserved while only the observed substitution count changes as a case fact; the two evidence sentences are plain factual statements about PO size and substitution count, not policy text; the counterfactual (90/500=18%) is internally consistent with the unchanged no-damage and no-other-discrepancy statements and merely shifts the substitution rate past the 10% threshold without contradicting other measurements; neither context states or implies the final queue/severity label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered 500 units of item Z from the supplier.\"}, {\"speaker\": \"Inspection team\", \"text\": \"All units received under PO-731 were physically intact upon inspection, with no damage found anywhere in the shipment.\"}, {\"speaker\": \"Analyst\", \"text\": \"An inspection confirmed that 40 units received under PO-731 were unauthorized substitutions.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"No other discrepancies were noted in the shipment records.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of item Z from the supplier."}, {"path": ["2", "text"], "text": "An inspection confirmed that 40 units received under PO-731 were unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of item Z from the supplier.", "negative_left": "PO-731 ordered 500 units of item Z from the supplier.", "negative_right": "An inspection confirmed that 90 units received under PO-731 were unauthorized substitutions.", "right": "An inspection confirmed that 40 units received under PO-731 were unauthorized substitutions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-002", "id": "fast-43-diverse-278-002-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of item Z from the supplier."}, {"speaker": "Inspection team", "text": "All units received under PO-731 were physically intact upon inspection, with no damage found anywhere in the shipment."}, {"speaker": "Analyst", "text": "An inspection confirmed that 40 units received under PO-731 were unauthorized substitutions."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "No other discrepancies were noted in the shipment records."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts restate the same routing/threshold policy verbatim, matching the original state's rubric, and the PO identifier binding (PO-731) is preserved while only the observed substitution count changes as a case fact; the two evidence sentences are plain factual statements about PO size and substitution count, not policy text; the counterfactual (90/500=18%) is internally consistent with the unchanged no-damage and no-other-discrepancy statements and merely shifts the substitution rate past the 10% threshold without contradicting other measurements; neither context states or implies the final queue/severity label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered 500 units of item Z from the supplier.\"}, {\"speaker\": \"Inspection team\", \"text\": \"All units received under PO-731 were physically intact upon inspection, with no damage found anywhere in the shipment.\"}, {\"speaker\": \"Analyst\", \"text\": \"An inspection confirmed that 40 units received under PO-731 were unauthorized substitutions.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"No other discrepancies were noted in the shipment records.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of item Z from the supplier."}, {"path": ["2", "text"], "text": "An inspection confirmed that 40 units received under PO-731 were unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of item Z from the supplier.", "negative_left": "PO-731 ordered 500 units of item Z from the supplier.", "negative_right": "An inspection confirmed that 90 units received under PO-731 were unauthorized substitutions.", "right": "An inspection confirmed that 40 units received under PO-731 were unauthorized substitutions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-002", "id": "fast-43-diverse-278-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of item Z from the supplier."}, {"speaker": "Inspection team", "text": "All units received under PO-731 were physically intact upon inspection, with no damage found anywhere in the shipment."}, {"speaker": "Analyst", "text": "An inspection confirmed that 90 units received under PO-731 were unauthorized substitutions."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "No other discrepancies were noted in the shipment records."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the PO-731 quantity and the dock supervisor's unchanged routing/threshold policy, matching the original question's criteria; the counterfactual only changes the affected-unit count (45→75), which is internally consistent with the fixed 500-unit PO and no other reported discrepancies; evidence spans are plain factual statements about order size and substitution quantity, not policy text; no queue names, severity labels, or rationale are asserted as answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 500 units of the component.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Carton C5 was labeled with a substitute part number that Procurement never approved; those units arrived intact with no sign of physical damage.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"The unauthorized substitution under PO-731 affected 45 units of the component.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"No other discrepancy—no shortage, overage, or damage—has been recorded on this shipment. Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of the component."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 45 units of the component."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of the component.", "negative_left": "PO-731 ordered 500 units of the component.", "negative_right": "The unauthorized substitution under PO-731 affected 75 units of the component.", "right": "The unauthorized substitution under PO-731 affected 45 units of the component."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-004", "id": "fast-43-diverse-278-004-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of the component."}, {"speaker": "Inventory control analyst", "text": "Carton C5 was labeled with a substitute part number that Procurement never approved; those units arrived intact with no sign of physical damage."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 45 units of the component."}, {"speaker": "Dock supervisor", "text": "No other discrepancy—no shortage, overage, or damage—has been recorded on this shipment. Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the PO-731 quantity and the dock supervisor's unchanged routing/threshold policy, matching the original question's criteria; the counterfactual only changes the affected-unit count (45→75), which is internally consistent with the fixed 500-unit PO and no other reported discrepancies; evidence spans are plain factual statements about order size and substitution quantity, not policy text; no queue names, severity labels, or rationale are asserted as answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 500 units of the component.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Carton C5 was labeled with a substitute part number that Procurement never approved; those units arrived intact with no sign of physical damage.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"The unauthorized substitution under PO-731 affected 45 units of the component.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"No other discrepancy—no shortage, overage, or damage—has been recorded on this shipment. Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of the component."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 45 units of the component."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of the component.", "negative_left": "PO-731 ordered 500 units of the component.", "negative_right": "The unauthorized substitution under PO-731 affected 75 units of the component.", "right": "The unauthorized substitution under PO-731 affected 45 units of the component."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-004", "id": "fast-43-diverse-278-004-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of the component."}, {"speaker": "Inventory control analyst", "text": "Carton C5 was labeled with a substitute part number that Procurement never approved; those units arrived intact with no sign of physical damage."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 75 units of the component."}, {"speaker": "Dock supervisor", "text": "No other discrepancy—no shortage, overage, or damage—has been recorded on this shipment. Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the PO-731 substitution scenario and full routing/threshold rubric from the original policy, only the affected-unit count varies (40 vs 60) which is an allowed observation change, evidence spans are two complete factual sentences without policy definitions, and no gold label, rule table, or output instruction is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered 500 units in total. Carton C5 contained units labeled Q-9, which were proposed as substitutes for the ordered P-9 seals.\"}, {\"speaker\": \"Inventory control analyst\", \"text\": \"Procurement confirms the Q-9 substitution was never approved. All received units, including those in C5, were inspected and found intact with no dents, cracks, or moisture damage.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"The unauthorized substitution under PO-731 affected 40 units.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units in total."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 40 units."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units in total.", "negative_left": "PO-731 ordered 500 units in total.", "negative_right": "The unauthorized substitution under PO-731 affected 60 units.", "right": "The unauthorized substitution under PO-731 affected 40 units."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-005", "id": "fast-43-diverse-278-005-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units in total. Carton C5 contained units labeled Q-9, which were proposed as substitutes for the ordered P-9 seals."}, {"speaker": "Inventory control analyst", "text": "Procurement confirms the Q-9 substitution was never approved. All received units, including those in C5, were inspected and found intact with no dents, cracks, or moisture damage."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 40 units."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the PO-731 substitution scenario and full routing/threshold rubric from the original policy, only the affected-unit count varies (40 vs 60) which is an allowed observation change, evidence spans are two complete factual sentences without policy definitions, and no gold label, rule table, or output instruction is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered 500 units in total. Carton C5 contained units labeled Q-9, which were proposed as substitutes for the ordered P-9 seals.\"}, {\"speaker\": \"Inventory control analyst\", \"text\": \"Procurement confirms the Q-9 substitution was never approved. All received units, including those in C5, were inspected and found intact with no dents, cracks, or moisture damage.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"The unauthorized substitution under PO-731 affected 40 units.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units in total."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 40 units."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units in total.", "negative_left": "PO-731 ordered 500 units in total.", "negative_right": "The unauthorized substitution under PO-731 affected 60 units.", "right": "The unauthorized substitution under PO-731 affected 40 units."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-005", "id": "fast-43-diverse-278-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units in total. Carton C5 contained units labeled Q-9, which were proposed as substitutes for the ordered P-9 seals."}, {"speaker": "Inventory control analyst", "text": "Procurement confirms the Q-9 substitution was never approved. All received units, including those in C5, were inspected and found intact with no dents, cracks, or moisture damage."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 60 units."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing rubric and PO-731 bindings from the original question/state, the two focus sentences are plain factual claims about quantity and affected units, and the counterfactual's 75-unit change is internally consistent (15% of 500) without contradicting other facts or leaking the gold label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"Inspection of the PO-731 carton confirmed all units arrived intact; no physical damage was found anywhere in the shipment.\"},{\"speaker\":\"Procurement analyst\",\"text\":\"The carton labeled for PO-731 carried a different part number than ordered, and those units had not been approved as substitutes, confirming an unauthorized substitution.\"},{\"speaker\":\"Auditor\",\"text\":\"PO-731 has a total ordered quantity of 500 units.\"},{\"speaker\":\"Auditor\",\"text\":\"An audit of PO-731 found that 50 units were affected by the unauthorized substitution.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "PO-731 has a total ordered quantity of 500 units."}, {"path": ["3", "text"], "text": "An audit of PO-731 found that 50 units were affected by the unauthorized substitution."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 has a total ordered quantity of 500 units.", "negative_left": "PO-731 has a total ordered quantity of 500 units.", "negative_right": "An audit of PO-731 found that 75 units were affected by the unauthorized substitution.", "right": "An audit of PO-731 found that 50 units were affected by the unauthorized substitution."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-006", "id": "fast-43-diverse-278-006-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "Inspection of the PO-731 carton confirmed all units arrived intact; no physical damage was found anywhere in the shipment."}, {"speaker": "Procurement analyst", "text": "The carton labeled for PO-731 carried a different part number than ordered, and those units had not been approved as substitutes, confirming an unauthorized substitution."}, {"speaker": "Auditor", "text": "PO-731 has a total ordered quantity of 500 units."}, {"speaker": "Auditor", "text": "An audit of PO-731 found that 50 units were affected by the unauthorized substitution."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing rubric and PO-731 bindings from the original question/state, the two focus sentences are plain factual claims about quantity and affected units, and the counterfactual's 75-unit change is internally consistent (15% of 500) without contradicting other facts or leaking the gold label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"Inspection of the PO-731 carton confirmed all units arrived intact; no physical damage was found anywhere in the shipment.\"},{\"speaker\":\"Procurement analyst\",\"text\":\"The carton labeled for PO-731 carried a different part number than ordered, and those units had not been approved as substitutes, confirming an unauthorized substitution.\"},{\"speaker\":\"Auditor\",\"text\":\"PO-731 has a total ordered quantity of 500 units.\"},{\"speaker\":\"Auditor\",\"text\":\"An audit of PO-731 found that 50 units were affected by the unauthorized substitution.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["2", "text"], "text": "PO-731 has a total ordered quantity of 500 units."}, {"path": ["3", "text"], "text": "An audit of PO-731 found that 50 units were affected by the unauthorized substitution."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 has a total ordered quantity of 500 units.", "negative_left": "PO-731 has a total ordered quantity of 500 units.", "negative_right": "An audit of PO-731 found that 75 units were affected by the unauthorized substitution.", "right": "An audit of PO-731 found that 50 units were affected by the unauthorized substitution."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-006", "id": "fast-43-diverse-278-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "Inspection of the PO-731 carton confirmed all units arrived intact; no physical damage was found anywhere in the shipment."}, {"speaker": "Procurement analyst", "text": "The carton labeled for PO-731 carried a different part number than ordered, and those units had not been approved as substitutes, confirming an unauthorized substitution."}, {"speaker": "Auditor", "text": "PO-731 has a total ordered quantity of 500 units."}, {"speaker": "Auditor", "text": "An audit of PO-731 found that 75 units were affected by the unauthorized substitution."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing priority and threshold rubric matching the original policy, preserve PO-731/C5 bindings, and only vary the substitution quantity (40 vs 90) without contradicting other stated facts; the two focus evidence spans are complete factual sentences, not instructions; no answer, label, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"Discrepancy review for PO-731 is underway. Inspection confirmed all received units were intact, with no physical damage found on any carton or item.\"}, {\"speaker\": \"Procurement analyst\", \"text\": \"PO-731 ordered 500 units of the component. Carton C5 was labeled with an alternate part number and its units featured blue caps; these had been proposed as substitutes but were never approved, establishing an unauthorized substitution among the units received under PO-731.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"The unauthorized substitution under PO-731 affected 40 units of the component. No other discrepancies, shortages, or overages were recorded for this shipment.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "PO-731 ordered 500 units of the component."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 40 units of the component."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of the component.", "negative_left": "PO-731 ordered 500 units of the component.", "negative_right": "The unauthorized substitution under PO-731 affected 90 units of the component.", "right": "The unauthorized substitution under PO-731 affected 40 units of the component."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-009", "id": "fast-43-diverse-278-009-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "Discrepancy review for PO-731 is underway. Inspection confirmed all received units were intact, with no physical damage found on any carton or item."}, {"speaker": "Procurement analyst", "text": "PO-731 ordered 500 units of the component. Carton C5 was labeled with an alternate part number and its units featured blue caps; these had been proposed as substitutes but were never approved, establishing an unauthorized substitution among the units received under PO-731."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 40 units of the component. No other discrepancies, shortages, or overages were recorded for this shipment."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing priority and threshold rubric matching the original policy, preserve PO-731/C5 bindings, and only vary the substitution quantity (40 vs 90) without contradicting other stated facts; the two focus evidence spans are complete factual sentences, not instructions; no answer, label, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"Discrepancy review for PO-731 is underway. Inspection confirmed all received units were intact, with no physical damage found on any carton or item.\"}, {\"speaker\": \"Procurement analyst\", \"text\": \"PO-731 ordered 500 units of the component. Carton C5 was labeled with an alternate part number and its units featured blue caps; these had been proposed as substitutes but were never approved, establishing an unauthorized substitution among the units received under PO-731.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"The unauthorized substitution under PO-731 affected 40 units of the component. No other discrepancies, shortages, or overages were recorded for this shipment.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["1", "text"], "text": "PO-731 ordered 500 units of the component."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 40 units of the component."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of the component.", "negative_left": "PO-731 ordered 500 units of the component.", "negative_right": "The unauthorized substitution under PO-731 affected 90 units of the component.", "right": "The unauthorized substitution under PO-731 affected 40 units of the component."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-009", "id": "fast-43-diverse-278-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "Discrepancy review for PO-731 is underway. Inspection confirmed all received units were intact, with no physical damage found on any carton or item."}, {"speaker": "Procurement analyst", "text": "PO-731 ordered 500 units of the component. Carton C5 was labeled with an alternate part number and its units featured blue caps; these had been proposed as substitutes but were never approved, establishing an unauthorized substitution among the units received under PO-731."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 90 units of the component. No other discrepancies, shortages, or overages were recorded for this shipment."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep PO-731 entity and the full routing rubric verbatim, only the affected-quantity figure changes (40 vs 60 of 500), which is a permissible case-observation edit; the two focus sentences are plain factual statements with no rule table or answer embedded, and the counterfactual's altered count does not conflict with any other retained figure.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered a total quantity of 500 units.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Carton C5 was labeled Q-9 with blue-capped units proposed as substitutes for the ordered P-9 seals, but these were not approved. Inspection found all units intact, so no physical damage applies.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"The unauthorized substitution under PO-731 affected 40 of the units received.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered a total quantity of 500 units."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 40 of the units received."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered a total quantity of 500 units.", "negative_left": "PO-731 ordered a total quantity of 500 units.", "negative_right": "The unauthorized substitution under PO-731 affected 60 of the units received.", "right": "The unauthorized substitution under PO-731 affected 40 of the units received."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-010", "id": "fast-43-diverse-278-010-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered a total quantity of 500 units."}, {"speaker": "Inventory control analyst", "text": "Carton C5 was labeled Q-9 with blue-capped units proposed as substitutes for the ordered P-9 seals, but these were not approved. Inspection found all units intact, so no physical damage applies."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 40 of the units received."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep PO-731 entity and the full routing rubric verbatim, only the affected-quantity figure changes (40 vs 60 of 500), which is a permissible case-observation edit; the two focus sentences are plain factual statements with no rule table or answer embedded, and the counterfactual's altered count does not conflict with any other retained figure.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered a total quantity of 500 units.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Carton C5 was labeled Q-9 with blue-capped units proposed as substitutes for the ordered P-9 seals, but these were not approved. Inspection found all units intact, so no physical damage applies.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"The unauthorized substitution under PO-731 affected 40 of the units received.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered a total quantity of 500 units."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 40 of the units received."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered a total quantity of 500 units.", "negative_left": "PO-731 ordered a total quantity of 500 units.", "negative_right": "The unauthorized substitution under PO-731 affected 60 of the units received.", "right": "The unauthorized substitution under PO-731 affected 40 of the units received."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-010", "id": "fast-43-diverse-278-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered a total quantity of 500 units."}, {"speaker": "Inventory control analyst", "text": "Carton C5 was labeled Q-9 with blue-capped units proposed as substitutes for the ordered P-9 seals, but these were not approved. Inspection found all units intact, so no physical damage applies."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 60 of the units received."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing priority and threshold rubric as the original policy, keep PO-731 entity and quantity bindings intact, and only vary the substitution quantity (40 vs 90) which shifts percentage from 8% to 18% without contradicting other facts; evidence spans are two complete factual sentences with no rule tables or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered a quantity of 500 units. ASN-731 declares receipt matching that order. My count sheet shows most units as ordered, but one carton labeled differently.\"}, {\"speaker\": \"Inventory control analyst\", \"text\": \"Procurement's note confirms the carton's label lists an alternate part number; those units were proposed as substitutes, but were not approved. Inspection found all units intact, so physical damage does not apply.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"The unauthorized substitution under PO-731 affected 40 units. No other discrepancy is recorded.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered a quantity of 500 units."}, {"path": ["3", "text"], "text": "The unauthorized substitution under PO-731 affected 40 units."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered a quantity of 500 units.", "negative_left": "PO-731 ordered a quantity of 500 units.", "negative_right": "The unauthorized substitution under PO-731 affected 90 units.", "right": "The unauthorized substitution under PO-731 affected 40 units."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-011", "id": "fast-43-diverse-278-011-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered a quantity of 500 units. ASN-731 declares receipt matching that order. My count sheet shows most units as ordered, but one carton labeled differently."}, {"speaker": "Inventory control analyst", "text": "Procurement's note confirms the carton's label lists an alternate part number; those units were proposed as substitutes, but were not approved. Inspection found all units intact, so physical damage does not apply."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 40 units. No other discrepancy is recorded."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing priority and threshold rubric as the original policy, keep PO-731 entity and quantity bindings intact, and only vary the substitution quantity (40 vs 90) which shifts percentage from 8% to 18% without contradicting other facts; evidence spans are two complete factual sentences with no rule tables or answer hints embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered a quantity of 500 units. ASN-731 declares receipt matching that order. My count sheet shows most units as ordered, but one carton labeled differently.\"}, {\"speaker\": \"Inventory control analyst\", \"text\": \"Procurement's note confirms the carton's label lists an alternate part number; those units were proposed as substitutes, but were not approved. Inspection found all units intact, so physical damage does not apply.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"The unauthorized substitution under PO-731 affected 40 units. No other discrepancy is recorded.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered a quantity of 500 units."}, {"path": ["3", "text"], "text": "The unauthorized substitution under PO-731 affected 40 units."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered a quantity of 500 units.", "negative_left": "PO-731 ordered a quantity of 500 units.", "negative_right": "The unauthorized substitution under PO-731 affected 90 units.", "right": "The unauthorized substitution under PO-731 affected 40 units."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-011", "id": "fast-43-diverse-278-011-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered a quantity of 500 units. ASN-731 declares receipt matching that order. My count sheet shows most units as ordered, but one carton labeled differently."}, {"speaker": "Inventory control analyst", "text": "Procurement's note confirms the carton's label lists an alternate part number; those units were proposed as substitutes, but were not approved. Inspection found all units intact, so physical damage does not apply."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 90 units. No other discrepancy is recorded."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep PO-731 identity and the full routing/threshold rubric verbatim, matching the original question's criteria; the two focus evidence spans are complete factual statements, not policy or rationale; the counterfactual swaps 40→90 units cleanly without contradicting the 'no damage' fact or duplicating the original count; neither context states or implies the final routing/severity label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 500 units of connector housings.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Inspectors confirmed that 40 units received under PO-731 were unauthorized substitutions.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"Inspection found all units intact, with no signs of physical damage.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of connector housings."}, {"path": ["1", "text"], "text": "Inspectors confirmed that 40 units received under PO-731 were unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of connector housings.", "negative_left": "PO-731 ordered 500 units of connector housings.", "negative_right": "Inspectors confirmed that 90 units received under PO-731 were unauthorized substitutions.", "right": "Inspectors confirmed that 40 units received under PO-731 were unauthorized substitutions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-013", "id": "fast-43-diverse-278-013-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of connector housings."}, {"speaker": "Inventory control analyst", "text": "Inspectors confirmed that 40 units received under PO-731 were unauthorized substitutions."}, {"speaker": "Receiving clerk", "text": "Inspection found all units intact, with no signs of physical damage."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep PO-731 identity and the full routing/threshold rubric verbatim, matching the original question's criteria; the two focus evidence spans are complete factual statements, not policy or rationale; the counterfactual swaps 40→90 units cleanly without contradicting the 'no damage' fact or duplicating the original count; neither context states or implies the final routing/severity label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 500 units of connector housings.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Inspectors confirmed that 40 units received under PO-731 were unauthorized substitutions.\"},{\"speaker\":\"Receiving clerk\",\"text\":\"Inspection found all units intact, with no signs of physical damage.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of connector housings."}, {"path": ["1", "text"], "text": "Inspectors confirmed that 40 units received under PO-731 were unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of connector housings.", "negative_left": "PO-731 ordered 500 units of connector housings.", "negative_right": "Inspectors confirmed that 90 units received under PO-731 were unauthorized substitutions.", "right": "Inspectors confirmed that 40 units received under PO-731 were unauthorized substitutions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-013", "id": "fast-43-diverse-278-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of connector housings."}, {"speaker": "Inventory control analyst", "text": "Inspectors confirmed that 90 units received under PO-731 were unauthorized substitutions."}, {"speaker": "Receiving clerk", "text": "Inspection found all units intact, with no signs of physical damage."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing priority and threshold rubric via the dock supervisor's speech, keep PO-731 as the bound entity, use two factual sentences as evidence, and the counterfactual's changed substitution count (55 vs 35) is internally consistent with the unchanged 400-unit total and no-damage finding, without leaking any gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 400 units in total.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"C5's label lists Q-9, proposed as substitutes but not approved. Inspection found all units intact, so no physical damage applies.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"An audit found that 35 units received under PO-731 were substituted with an unauthorized alternate part.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 400 units in total."}, {"path": ["2", "text"], "text": "An audit found that 35 units received under PO-731 were substituted with an unauthorized alternate part."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 400 units in total.", "negative_left": "PO-731 ordered 400 units in total.", "negative_right": "An audit found that 55 units received under PO-731 were substituted with an unauthorized alternate part.", "right": "An audit found that 35 units received under PO-731 were substituted with an unauthorized alternate part."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-016", "id": "fast-43-diverse-278-016-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 400 units in total."}, {"speaker": "Inventory control analyst", "text": "C5's label lists Q-9, proposed as substitutes but not approved. Inspection found all units intact, so no physical damage applies."}, {"speaker": "Supplier claims coordinator", "text": "An audit found that 35 units received under PO-731 were substituted with an unauthorized alternate part."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full routing priority and threshold rubric via the dock supervisor's speech, keep PO-731 as the bound entity, use two factual sentences as evidence, and the counterfactual's changed substitution count (55 vs 35) is internally consistent with the unchanged 400-unit total and no-damage finding, without leaking any gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 400 units in total.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"C5's label lists Q-9, proposed as substitutes but not approved. Inspection found all units intact, so no physical damage applies.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"An audit found that 35 units received under PO-731 were substituted with an unauthorized alternate part.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 400 units in total."}, {"path": ["2", "text"], "text": "An audit found that 35 units received under PO-731 were substituted with an unauthorized alternate part."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 400 units in total.", "negative_left": "PO-731 ordered 400 units in total.", "negative_right": "An audit found that 55 units received under PO-731 were substituted with an unauthorized alternate part.", "right": "An audit found that 35 units received under PO-731 were substituted with an unauthorized alternate part."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-016", "id": "fast-43-diverse-278-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 400 units in total."}, {"speaker": "Inventory control analyst", "text": "C5's label lists Q-9, proposed as substitutes but not approved. Inspection found all units intact, so no physical damage applies."}, {"speaker": "Supplier claims coordinator", "text": "An audit found that 55 units received under PO-731 were substituted with an unauthorized alternate part."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the dock supervisor's full routing/threshold policy verbatim, keep PO-731/connector-housing bindings intact, and the counterfactual only changes the substitution quantity (40→90 units) without contradicting other facts, yielding a coherent 8%→18% shift with no embedded answer or rule table beyond the original policy statement.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered 500 units of connector housings. Inspection found all units intact, so no physical damage was recorded on this shipment.\"}, {\"speaker\": \"Inventory control analyst\", \"text\": \"Procurement confirmed that a batch was mislabeled and shipped without prior authorization from the supplier.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"An inspection confirmed that 40 units of connector housings were unauthorized substitutions.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of connector housings."}, {"path": ["3", "text"], "text": "An inspection confirmed that 40 units of connector housings were unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of connector housings.", "negative_left": "PO-731 ordered 500 units of connector housings.", "negative_right": "An inspection confirmed that 90 units of connector housings were unauthorized substitutions.", "right": "An inspection confirmed that 40 units of connector housings were unauthorized substitutions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-019", "id": "fast-43-diverse-278-019-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of connector housings. Inspection found all units intact, so no physical damage was recorded on this shipment."}, {"speaker": "Inventory control analyst", "text": "Procurement confirmed that a batch was mislabeled and shipped without prior authorization from the supplier."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "An inspection confirmed that 40 units of connector housings were unauthorized substitutions."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the dock supervisor's full routing/threshold policy verbatim, keep PO-731/connector-housing bindings intact, and the counterfactual only changes the substitution quantity (40→90 units) without contradicting other facts, yielding a coherent 8%→18% shift with no embedded answer or rule table beyond the original policy statement.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered 500 units of connector housings. Inspection found all units intact, so no physical damage was recorded on this shipment.\"}, {\"speaker\": \"Inventory control analyst\", \"text\": \"Procurement confirmed that a batch was mislabeled and shipped without prior authorization from the supplier.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"An inspection confirmed that 40 units of connector housings were unauthorized substitutions.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of connector housings."}, {"path": ["3", "text"], "text": "An inspection confirmed that 40 units of connector housings were unauthorized substitutions."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of connector housings.", "negative_left": "PO-731 ordered 500 units of connector housings.", "negative_right": "An inspection confirmed that 90 units of connector housings were unauthorized substitutions.", "right": "An inspection confirmed that 40 units of connector housings were unauthorized substitutions."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-019", "id": "fast-43-diverse-278-019-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of connector housings. Inspection found all units intact, so no physical damage was recorded on this shipment."}, {"speaker": "Inventory control analyst", "text": "Procurement confirmed that a batch was mislabeled and shipped without prior authorization from the supplier."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "An inspection confirmed that 90 units of connector housings were unauthorized substitutions."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing rubric and PO-731 binding, only varying the substitution quantity (40 vs 90 of 500), which keeps the scenario internally consistent without duplicate or conflicting counts; the focus evidence consists of two complete factual sentences about the PO size and audit finding, and neither context states or implies the final queue/severity label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 500 units of item Z.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Inspection found all units on the dock intact, with no signs of physical damage to any carton or item.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"The label on the affected carton listed a different part number than the PO called for, and the substituted units were never approved by procurement, establishing an unauthorized substitution.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"An audit found 40 units of item Z under PO-731 were affected by the unauthorized substitution.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of item Z."}, {"path": ["4", "text"], "text": "An audit found 40 units of item Z under PO-731 were affected by the unauthorized substitution."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of item Z.", "negative_left": "PO-731 ordered 500 units of item Z.", "negative_right": "An audit found 90 units of item Z under PO-731 were affected by the unauthorized substitution.", "right": "An audit found 40 units of item Z under PO-731 were affected by the unauthorized substitution."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-031", "id": "fast-43-diverse-278-031-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of item Z."}, {"speaker": "Inventory control analyst", "text": "Inspection found all units on the dock intact, with no signs of physical damage to any carton or item."}, {"speaker": "Inventory control analyst", "text": "The label on the affected carton listed a different part number than the PO called for, and the substituted units were never approved by procurement, establishing an unauthorized substitution."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "An audit found 40 units of item Z under PO-731 were affected by the unauthorized substitution."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the identical routing rubric and PO-731 binding, only varying the substitution quantity (40 vs 90 of 500), which keeps the scenario internally consistent without duplicate or conflicting counts; the focus evidence consists of two complete factual sentences about the PO size and audit finding, and neither context states or implies the final queue/severity label.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 500 units of item Z.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Inspection found all units on the dock intact, with no signs of physical damage to any carton or item.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"The label on the affected carton listed a different part number than the PO called for, and the substituted units were never approved by procurement, establishing an unauthorized substitution.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"An audit found 40 units of item Z under PO-731 were affected by the unauthorized substitution.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of item Z."}, {"path": ["4", "text"], "text": "An audit found 40 units of item Z under PO-731 were affected by the unauthorized substitution."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of item Z.", "negative_left": "PO-731 ordered 500 units of item Z.", "negative_right": "An audit found 90 units of item Z under PO-731 were affected by the unauthorized substitution.", "right": "An audit found 40 units of item Z under PO-731 were affected by the unauthorized substitution."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-031", "id": "fast-43-diverse-278-031-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of item Z."}, {"speaker": "Inventory control analyst", "text": "Inspection found all units on the dock intact, with no signs of physical damage to any carton or item."}, {"speaker": "Inventory control analyst", "text": "The label on the affected carton listed a different part number than the PO called for, and the substituted units were never approved by procurement, establishing an unauthorized substitution."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "An audit found 90 units of item Z under PO-731 were affected by the unauthorized substitution."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy text and PO/item bindings, only the substituted quantity changes (45→145) which cleanly shifts severity without contradicting other stated facts, evidence spans are plain factual sentences, and no answer or rule table is leaked.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 500 units of item Z-14.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Inspection confirmed no physical damage to any of the received units under PO-731; all cartons were intact and sealed.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"Receiving inspection found 45 units of item Z-14 substituted with an unauthorized alternate part under PO-731.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"Procurement confirmed the alternate part was never authorized as a substitute for Z-14, and no other discrepancy is recorded on this PO.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of item Z-14."}, {"path": ["3", "text"], "text": "Receiving inspection found 45 units of item Z-14 substituted with an unauthorized alternate part under PO-731."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of item Z-14.", "negative_left": "PO-731 ordered 500 units of item Z-14.", "negative_right": "Receiving inspection found 145 units of item Z-14 substituted with an unauthorized alternate part under PO-731.", "right": "Receiving inspection found 45 units of item Z-14 substituted with an unauthorized alternate part under PO-731."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-032", "id": "fast-43-diverse-278-032-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of item Z-14."}, {"speaker": "Inventory control analyst", "text": "Inspection confirmed no physical damage to any of the received units under PO-731; all cartons were intact and sealed."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "Receiving inspection found 45 units of item Z-14 substituted with an unauthorized alternate part under PO-731."}, {"speaker": "Supplier claims coordinator", "text": "Procurement confirmed the alternate part was never authorized as a substitute for Z-14, and no other discrepancy is recorded on this PO."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same routing policy text and PO/item bindings, only the substituted quantity changes (45→145) which cleanly shifts severity without contradicting other stated facts, evidence spans are plain factual sentences, and no answer or rule table is leaked.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\":\"Receiving clerk\",\"text\":\"PO-731 ordered 500 units of item Z-14.\"},{\"speaker\":\"Inventory control analyst\",\"text\":\"Inspection confirmed no physical damage to any of the received units under PO-731; all cartons were intact and sealed.\"},{\"speaker\":\"Dock supervisor\",\"text\":\"Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"Receiving inspection found 45 units of item Z-14 substituted with an unauthorized alternate part under PO-731.\"},{\"speaker\":\"Supplier claims coordinator\",\"text\":\"Procurement confirmed the alternate part was never authorized as a substitute for Z-14, and no other discrepancy is recorded on this PO.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of item Z-14."}, {"path": ["3", "text"], "text": "Receiving inspection found 45 units of item Z-14 substituted with an unauthorized alternate part under PO-731."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of item Z-14.", "negative_left": "PO-731 ordered 500 units of item Z-14.", "negative_right": "Receiving inspection found 145 units of item Z-14 substituted with an unauthorized alternate part under PO-731.", "right": "Receiving inspection found 45 units of item Z-14 substituted with an unauthorized alternate part under PO-731."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-032", "id": "fast-43-diverse-278-032-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of item Z-14."}, {"speaker": "Inventory control analyst", "text": "Inspection confirmed no physical damage to any of the received units under PO-731; all cartons were intact and sealed."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}, {"speaker": "Supplier claims coordinator", "text": "Receiving inspection found 145 units of item Z-14 substituted with an unauthorized alternate part under PO-731."}, {"speaker": "Supplier claims coordinator", "text": "Procurement confirmed the alternate part was never authorized as a substitute for Z-14, and no other discrepancy is recorded on this PO."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the dock supervisor's routing priority and threshold rubric verbatim, keep PO-731 as the bound entity, and the counterfactual only changes the affected unit count (40→90) without contradicting any other stated fact; evidence spans are plain factual sentences with no embedded labels or rule tables.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered 500 units of the component. Our count sheet and ASN both reference the same shipment on the dock today.\"}, {\"speaker\": \"Inventory control analyst\", \"text\": \"Inspection confirmed all units on the dock are intact, so physical damage does not apply to the shipment received under PO-731. Procurement confirmed that carton C5 contained units labeled with an alternate part number that had not been approved as substitutes, establishing an unauthorized substitution among the units received under PO-731.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"The unauthorized substitution under PO-731 affected 40 units of the component. No other discrepancy is recorded for this shipment.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of the component."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 40 units of the component."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of the component.", "negative_left": "PO-731 ordered 500 units of the component.", "negative_right": "The unauthorized substitution under PO-731 affected 90 units of the component.", "right": "The unauthorized substitution under PO-731 affected 40 units of the component."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-033", "id": "fast-43-diverse-278-033-base", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of the component. Our count sheet and ASN both reference the same shipment on the dock today."}, {"speaker": "Inventory control analyst", "text": "Inspection confirmed all units on the dock are intact, so physical damage does not apply to the shipment received under PO-731. Procurement confirmed that carton C5 contained units labeled with an alternate part number that had not been approved as substitutes, establishing an unauthorized substitution among the units received under PO-731."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 40 units of the component. No other discrepancy is recorded for this shipment."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_medium_exact_threshold"}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the dock supervisor's routing priority and threshold rubric verbatim, keep PO-731 as the bound entity, and the counterfactual only changes the affected unit count (40→90) without contradicting any other stated fact; evidence spans are plain factual sentences with no embedded labels or rule tables.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "full_context_fact_states": {"base": {"A1": "refuted", "A2": "supported", "A3": "supported"}, "counterfactual": {"A1": "refuted", "A2": "supported", "A3": "refuted"}, "remove_left": {"A3": "unknown"}, "remove_right": {"A3": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"A3": "unknown"}, "negative_pair": {"A3": "refuted"}, "negative_sentence": {"A3": "unknown"}, "positive_pair": {"A3": "supported"}, "right": {"A3": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states one factual relationship, and A3 is a factual threshold relation rather than a policy classification. Base and counter assignments are realizable while changing only A3, for example by using an affected substitution quantity of 10% versus more than 10%, with no damage in either case. The policy evidence accurately preserves the routing priority and substitution thresholds originating in the original state; rules already present in the retained questions need not be duplicated. Both rules include the necessary exclusion of the competing higher-priority damage outcome.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction excludes the higher-priority physical-damage route, establishes unauthorized substitution, and establishes that the affected quantity is at most 10% of the PO. This is sufficient for Supplier Claims with Medium severity, including exactly 10%.", "rule_index": 0, "sound": true}, {"reason": "The conjunction excludes physical damage, establishes unauthorized substitution, and refutes that the affected quantity is at most 10%. For a numeric affected proportion, this entails more than 10%, which is sufficient for Supplier Claims with High severity.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "A1", "statement": "Physical damage applies to the units received under PO-731."}, {"id": "A2", "statement": "An unauthorized substitution is established among the units received under PO-731."}, {"id": "A3", "statement": "The quantity affected by the unauthorized substitution under PO-731 is at most 10% of the quantity ordered under PO-731."}], "base_state_json": "[{\"speaker\": \"Receiving clerk\", \"text\": \"PO-731 ordered 500 units of the component. Our count sheet and ASN both reference the same shipment on the dock today.\"}, {\"speaker\": \"Inventory control analyst\", \"text\": \"Inspection confirmed all units on the dock are intact, so physical damage does not apply to the shipment received under PO-731. Procurement confirmed that carton C5 contained units labeled with an alternate part number that had not been approved as substitutes, establishing an unauthorized substitution among the units received under PO-731.\"}, {\"speaker\": \"Supplier claims coordinator\", \"text\": \"The unauthorized substitution under PO-731 affected 40 units of the component. No other discrepancy is recorded for this shipment.\"}, {\"speaker\": \"Dock supervisor\", \"text\": \"Apply one route by priority: physical damage \\u2192 Dock Hold; unauthorized substitution \\u2192 Supplier Claims; confirmed shortage or overage \\u2192 Inventory Control; documentation-only ambiguity \\u2192 Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium.\"}]", "base_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}], "counter_states": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}], "focus_atom": "A3", "focus_evidence": [{"path": ["0", "text"], "text": "PO-731 ordered 500 units of the component."}, {"path": ["2", "text"], "text": "The unauthorized substitution under PO-731 affected 40 units of the component."}], "policy_evidence": [{"path": ["2", "text"], "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}], "rules": [{"justification": "With physical damage excluded, unauthorized substitution is the first applicable discrepancy; an affected quantity at or below 10% routes to Supplier Claims with Medium severity, including exactly 10%.", "target": "supplier_claims_medium_exact_threshold", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "supported"}]}, {"justification": "With physical damage excluded and unauthorized substitution established, explicit refutation of an at-most-10% affected quantity entails that more than 10% is affected, which routes to Supplier Claims with High severity.", "target": "supplier_claims_high", "when": [{"atom_id": "A1", "state": "refuted"}, {"atom_id": "A2", "state": "supported"}, {"atom_id": "A3", "state": "refuted"}]}]}, "verified_pair": {"left": "PO-731 ordered 500 units of the component.", "negative_left": "PO-731 ordered 500 units of the component.", "negative_right": "The unauthorized substitution under PO-731 affected 90 units of the component.", "right": "The unauthorized substitution under PO-731 affected 40 units of the component."}, "verifier_independent_model": false}, "family": "fast-43-diverse-278-033", "id": "fast-43-diverse-278-033-counterfactual", "input": {"questions": {"decision": {"criteria": {"dock_hold_medium": "Choose when physical damage is the first applicable discrepancy and the affected quantity is at most 10% of the PO.", "inventory_control_medium": "Choose when neither damage nor unauthorized substitution applies, but a confirmed shortage or overage affects at most 10% of the PO.", "none_of_above": "Choose only when the evidence and stated routing rubric require a queue or severity combination not represented by the other options.", "receiving_review_low": "Choose when there is only a documentation ambiguity, with no established damage, unauthorized substitution, shortage, or overage.", "supplier_claims_high": "Choose when no physical damage applies, an unauthorized substitution is established, and more than 10% of the PO is affected.", "supplier_claims_medium_exact_threshold": "Choose when no physical damage applies, an unauthorized substitution is established, and at most 10% of the PO is affected, including exactly 10%."}, "instructions": "Resolve the references in the procurement note, apply the routing priority, and select the single correct queue and severity under the stated threshold rubric.", "type": "choice"}}, "state": [{"speaker": "Receiving clerk", "text": "PO-731 ordered 500 units of the component. Our count sheet and ASN both reference the same shipment on the dock today."}, {"speaker": "Inventory control analyst", "text": "Inspection confirmed all units on the dock are intact, so physical damage does not apply to the shipment received under PO-731. Procurement confirmed that carton C5 contained units labeled with an alternate part number that had not been approved as substitutes, establishing an unauthorized substitution among the units received under PO-731."}, {"speaker": "Supplier claims coordinator", "text": "The unauthorized substitution under PO-731 affected 90 units of the component. No other discrepancy is recorded for this shipment."}, {"speaker": "Dock supervisor", "text": "Apply one route by priority: physical damage → Dock Hold; unauthorized substitution → Supplier Claims; confirmed shortage or overage → Inventory Control; documentation-only ambiguity → Receiving Review. For substitutions, Medium is at most 10% of the PO; High is more than 10%. Exactly 10% is Medium."}]}, "method": "c2d", "provenance": {"source_id": "diverse-278", "source_is_synthetic": true, "source_sha256": "8403d1a7c9f41e9ac9e7755a7dd57b454a0b30ecf4c08fd1ee57328deccf53e6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": "supplier_claims_high"}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing routing policy and question object; the counterfactual only changes the PO's SKU, which is a consistent, non-contradictory factual edit; evidence spans are plain factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\": \"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\", \"evidence\": [\"Purchase order PO 7714 specifies SKU-4021 for the ordered goods.\", \"ASN 7714-A likewise lists SKU-4021 as the shipped item.\", \"The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098.\", \"The Dock Supervisor found no change notice or substitution approval on file.\", \"The inspection note records the alternate goods as fully intact with no damage.\"], \"request\": \"Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Purchase order PO 7714 specifies SKU-4021 for the ordered goods."}, {"path": ["evidence", "2"], "text": "The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "Purchase order PO 7714 specifies SKU-4021 for the ordered goods.", "negative_left": "Purchase order PO 7714 specifies SKU-4098 for the ordered goods.", "negative_right": "The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098.", "right": "The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098."}, "verifier_independent_model": false}, "family": "fast-43-diverse-280-016", "id": "fast-43-diverse-280-016-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["Purchase order PO 7714 specifies SKU-4021 for the ordered goods.", "ASN 7714-A likewise lists SKU-4021 as the shipped item.", "The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098.", "The Dock Supervisor found no change notice or substitution approval on file.", "The inspection note records the alternate goods as fully intact with no damage."], "request": "Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?"}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-02", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing routing policy and question object; the counterfactual only changes the PO's SKU, which is a consistent, non-contradictory factual edit; evidence spans are plain factual sentences with no embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom expresses one factual relationship, and the focus atom is a factual SKU comparison rather than a policy conclusion. The policy evidence correctly preserves the governing routing definition and sequence from the original state; question-level instructions need not be repeated. The base and counter assignments are jointly realizable with only a1 changing: in the counter, the received SKU can match the purchase order while differing from the ASN, with authorization absent and the goods undamaged.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction states every policy condition: the received SKU differs from both the purchase order and ASN, no written authorization exists, and the alternate goods are undamaged. This is sufficient for the true routing outcome.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 means the received SKU does not differ from the purchase order. Therefore the required condition that it differ from both documents fails, which is sufficient for the false outcome under the stated rubric.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in PO 7714."}, {"id": "a2", "statement": "The SKU of the goods received in the shipment associated with PO 7714 and ASN 7714-A differs from the SKU specified in ASN 7714-A."}, {"id": "a3", "statement": "No written authorization exists for substituting the goods received in the shipment associated with PO 7714 and ASN 7714-A."}, {"id": "a4", "statement": "The alternate goods received in the shipment associated with PO 7714 and ASN 7714-A are undamaged."}], "base_state_json": "{\"context\": \"At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.\", \"evidence\": [\"Purchase order PO 7714 specifies SKU-4021 for the ordered goods.\", \"ASN 7714-A likewise lists SKU-4021 as the shipped item.\", \"The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098.\", \"The Dock Supervisor found no change notice or substitution approval on file.\", \"The inspection note records the alternate goods as fully intact with no damage.\"], \"request\": \"Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?\"}", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}], "focus_atom": "a1", "focus_evidence": [{"path": ["evidence", "0"], "text": "Purchase order PO 7714 specifies SKU-4021 for the ordered goods."}, {"path": ["evidence", "2"], "text": "The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098."}], "policy_evidence": [{"path": ["context"], "text": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review."}], "rules": [{"justification": "The received SKU differs from both governing shipment documents, no written authorization exists, and the alternate goods are undamaged, so the shipment meets every stated condition for an undocumented substitution and must go first to the Inventory Control Analyst.", "target": "true", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}, {"justification": "The received SKU does not differ from the purchase order, so the policy requirement that it differ from both the purchase order and advance shipping notice is not met; it must not be routed as an undocumented substitution.", "target": "false", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}]}]}, "verified_pair": {"left": "Purchase order PO 7714 specifies SKU-4021 for the ordered goods.", "negative_left": "Purchase order PO 7714 specifies SKU-4098 for the ordered goods.", "negative_right": "The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098.", "right": "The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098."}, "verifier_independent_model": false}, "family": "fast-43-diverse-280-016", "id": "fast-43-diverse-280-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not route it to Inventory Control as an undocumented substitution.", "true": "Yes — route it first to the Inventory Control Analyst as an undocumented substitution."}, "instructions": "Answer yes or no using the stated routing rubric. Treat wording that conveys an unapproved alternate version as evidence of substitution even if it does not use that exact term.", "type": "noul"}}, "state": {"context": "At North Quay Warehouse, routing policy says a shipment is an undocumented substitution when the received SKU differs from both the purchase order and advance shipping notice, no written authorization exists, and the alternate goods are undamaged. Such cases go first to the Inventory Control Analyst; the Supplier Claims Coordinator acts only after that review.", "evidence": ["Purchase order PO 7714 specifies SKU-4098 for the ordered goods.", "ASN 7714-A likewise lists SKU-4021 as the shipped item.", "The shipment received against PO 7714 and ASN 7714-A contains goods labeled SKU-4098.", "The Dock Supervisor found no change notice or substitution approval on file.", "The inspection note records the alternate goods as fully intact with no damage."], "request": "Should this shipment be routed to the Inventory Control Analyst as an undocumented substitution?"}}, "method": "c2d", "provenance": {"source_id": "diverse-280", "source_is_synthetic": true, "source_sha256": "fcefd55951e4e2dcf38c18111ea39bf9adeac4f08c9f3a9ba3bf553dfa8e2ca6", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-02", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing policy clause and all readiness conditions, with only the staffing count varied as a factual observation; the two focus sentences are plain factual statements, not policy text; the counterfactual's reduced operator count (4 vs required 6) is internally coherent and creates a genuine shortfall without contradicting other facts; neither context states or implies the release decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support confirmed all required white resin and labels are staged before release. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions."}, {"path": [], "text": "7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_right": "4 trained operators are currently assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-001", "id": "fast-43-diverse-292-001-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support confirmed all required white resin and labels are staged before release. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the same governing policy clause and all readiness conditions, with only the staffing count varied as a factual observation; the two focus sentences are plain factual statements, not policy text; the counterfactual's reduced operator count (4 vs required 6) is internally coherent and creates a genuine shortfall without contradicting other facts; neither context states or implies the release decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support confirmed all required white resin and labels are staged before release. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions."}, {"path": [], "text": "7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_right": "4 trained operators are currently assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-001", "id": "fast-43-diverse-292-001-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support confirmed all required white resin and labels are staged before release. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 4 trained operators are currently assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing readiness conditions from the original policy (materials, staffing, cleaning, tooling, maintenance, first-piece approval) alongside the unchanged question; the staffing requirement is a case observation that may legitimately change and does not alter the question's entity/time bindings (Line 4, 14:00 change). The two evidence sentences are complete factual statements, not policy definitions or instructions. The counterfactual coherently reduces the assigned operator count from 6 to 4 while keeping the 6-operator requirement fixed, creating a genuine shortfall without contradicting other unchanged facts. Neither context states a gold answer, rule table, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white valve caps at 14:00. All required white resin and labels are staged before release. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. No maintenance work remains open, and yesterday's temperature alarm did not recur during today's test and is not an unresolved defect. The quality inspector approved first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:50, 6 trained operators have been assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "As of 13:50, 6 trained operators have been assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "As of 13:50, 4 trained operators have been assigned to Line 4 for the scheduled 14:00 change.", "right": "As of 13:50, 6 trained operators have been assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-002", "id": "fast-43-diverse-292-002-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white valve caps at 14:00. All required white resin and labels are staged before release. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. No maintenance work remains open, and yesterday's temperature alarm did not recur during today's test and is not an unresolved defect. The quality inspector approved first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:50, 6 trained operators have been assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all governing readiness conditions from the original policy (materials, staffing, cleaning, tooling, maintenance, first-piece approval) alongside the unchanged question; the staffing requirement is a case observation that may legitimately change and does not alter the question's entity/time bindings (Line 4, 14:00 change). The two evidence sentences are complete factual statements, not policy definitions or instructions. The counterfactual coherently reduces the assigned operator count from 6 to 4 while keeping the 6-operator requirement fixed, creating a genuine shortfall without contradicting other unchanged facts. Neither context states a gold answer, rule table, or explicit output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white valve caps at 14:00. All required white resin and labels are staged before release. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. No maintenance work remains open, and yesterday's temperature alarm did not recur during today's test and is not an unresolved defect. The quality inspector approved first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:50, 6 trained operators have been assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "As of 13:50, 6 trained operators have been assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "As of 13:50, 4 trained operators have been assigned to Line 4 for the scheduled 14:00 change.", "right": "As of 13:50, 6 trained operators have been assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-002", "id": "fast-43-diverse-292-002-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white valve caps at 14:00. All required white resin and labels are staged before release. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. No maintenance work remains open, and yesterday's temperature alarm did not recur during today's test and is not an unresolved defect. The quality inspector approved first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:50, 4 trained operators have been assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing release policy verbatim and only vary the staffing observation, which is permitted; the two evidence spans are complete factual sentences with no policy definitions; the counterfactual (4 vs required 6) is internally consistent with no duplicate or contradictory measurements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At Northstar Components, Line 4 is scheduled for its 14:00 changeover from blue to white valve caps. All required white resin and labels are staged for the change. The cleaning record confirms completion at 13:20. The correct mold is installed and all tooling checks passed. No maintenance work remains open on Line 4, and no maintenance defect affecting the line remains unresolved. The quality inspector approved the first piece for the 14:00 production of white valve caps. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-003", "id": "fast-43-diverse-292-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At Northstar Components, Line 4 is scheduled for its 14:00 changeover from blue to white valve caps. All required white resin and labels are staged for the change. The cleaning record confirms completion at 13:20. The correct mold is installed and all tooling checks passed. No maintenance work remains open on Line 4, and no maintenance defect affecting the line remains unresolved. The quality inspector approved the first piece for the 14:00 production of white valve caps. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the governing release policy verbatim and only vary the staffing observation, which is permitted; the two evidence spans are complete factual sentences with no policy definitions; the counterfactual (4 vs required 6) is internally consistent with no duplicate or contradictory measurements; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At Northstar Components, Line 4 is scheduled for its 14:00 changeover from blue to white valve caps. All required white resin and labels are staged for the change. The cleaning record confirms completion at 13:20. The correct mold is installed and all tooling checks passed. No maintenance work remains open on Line 4, and no maintenance defect affecting the line remains unresolved. The quality inspector approved the first piece for the 14:00 production of white valve caps. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-003", "id": "fast-43-diverse-292-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At Northstar Components, Line 4 is scheduled for its 14:00 changeover from blue to white valve caps. All required white resin and labels are staged for the change. The cleaning record confirms completion at 13:20. The correct mold is installed and all tooling checks passed. No maintenance work remains open on Line 4, and no maintenance defect affecting the line remains unresolved. The quality inspector approved the first piece for the 14:00 production of white valve caps. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 4 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the release policy and question bindings, the focus evidence consists of two factual sentences about staffing, and the counterfactual coherently varies the operator count without contradicting other unchanged facts or leaking answer-determining language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At Northstar Components, Line 4 is scheduled to switch from blue to white valve caps at 14:00. Materials support confirmed all required white resin and labels are staged. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks have passed. A maintenance technician confirmed no maintenance work remains open and no defect is unresolved. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:45, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "As of 13:45, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "As of 13:45, 4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "As of 13:45, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-005", "id": "fast-43-diverse-292-005-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At Northstar Components, Line 4 is scheduled to switch from blue to white valve caps at 14:00. Materials support confirmed all required white resin and labels are staged. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks have passed. A maintenance technician confirmed no maintenance work remains open and no defect is unresolved. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:45, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the release policy and question bindings, the focus evidence consists of two factual sentences about staffing, and the counterfactual coherently varies the operator count without contradicting other unchanged facts or leaking answer-determining language.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At Northstar Components, Line 4 is scheduled to switch from blue to white valve caps at 14:00. Materials support confirmed all required white resin and labels are staged. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks have passed. A maintenance technician confirmed no maintenance work remains open and no defect is unresolved. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:45, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "As of 13:45, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "As of 13:45, 4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "As of 13:45, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-005", "id": "fast-43-diverse-292-005-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At Northstar Components, Line 4 is scheduled to switch from blue to white valve caps at 14:00. Materials support confirmed all required white resin and labels are staged. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks have passed. A maintenance technician confirmed no maintenance work remains open and no defect is unresolved. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:45, 4 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical governing policy and entity/time bindings for Line 4's 14:00 change, only the assigned-operator count varies (7 vs 4) which is a permissible observation change; the two focus sentences are plain factual statements, not policy text; the counterfactual (4<6) is internally coherent and does not contradict other unchanged facts; neither context reveals the yes/no answer, rationale, or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials staff confirmed all required white resin and labels are staged for the change. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-006", "id": "fast-43-diverse-292-006-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials staff confirmed all required white resin and labels are staged for the change. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical governing policy and entity/time bindings for Line 4's 14:00 change, only the assigned-operator count varies (7 vs 4) which is a permissible observation change; the two focus sentences are plain factual statements, not policy text; the counterfactual (4<6) is internally coherent and does not contradict other unchanged facts; neither context reveals the yes/no answer, rationale, or instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials staff confirmed all required white resin and labels are staged for the change. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-006", "id": "fast-43-diverse-292-006-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials staff confirmed all required white resin and labels are staged for the change. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 4 trained operators are assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy statement and Line 4/14:00 bindings; the base context restates staffing as 6 required/7 assigned (meeting spec) while the counterfactual changes only the assigned count to 4 (failing spec), a coherent single-fact change; the two focus evidence spans are complete factual sentences about staffing requirement and assignment, not policy or instructions; no gold answer, label, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support confirms all required white resin and labels are staged for the change. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. A maintenance technician confirmed yesterday's temperature alarm did not recur during today's test and is not an unresolved defect, and no maintenance work remains open. The quality inspector approved the first-piece dimensions and color before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_right": "4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-009", "id": "fast-43-diverse-292-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support confirms all required white resin and labels are staged for the change. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. A maintenance technician confirmed yesterday's temperature alarm did not recur during today's test and is not an unresolved defect, and no maintenance work remains open. The quality inspector approved the first-piece dimensions and color before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve the original policy statement and Line 4/14:00 bindings; the base context restates staffing as 6 required/7 assigned (meeting spec) while the counterfactual changes only the assigned count to 4 (failing spec), a coherent single-fact change; the two focus evidence spans are complete factual sentences about staffing requirement and assignment, not policy or instructions; no gold answer, label, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support confirms all required white resin and labels are staged for the change. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. A maintenance technician confirmed yesterday's temperature alarm did not recur during today's test and is not an unresolved defect, and no maintenance work remains open. The quality inspector approved the first-piece dimensions and color before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_right": "4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-009", "id": "fast-43-diverse-292-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support confirms all required white resin and labels are staged for the change. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. A maintenance technician confirmed yesterday's temperature alarm did not recur during today's test and is not an unresolved defect, and no maintenance work remains open. The quality inspector approved the first-piece dimensions and color before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 4 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy sentence and Line 4/14:00 bindings; the evidence consists of two factual sentences (staffing requirement and assigned operators) rather than policy text; the counterfactual simply lowers the assigned operator count to 5, which is internally consistent and does not contradict other facts; neither context states or implies the final yes/no decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. All required white resin and labels are staged for the change. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. No maintenance work remains open, and no maintenance defect affecting Line 4 remains unresolved at the release decision. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "5 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-010", "id": "fast-43-diverse-292-010-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. All required white resin and labels are staged for the change. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. No maintenance work remains open, and no maintenance defect affecting Line 4 remains unresolved at the release decision. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original policy sentence and Line 4/14:00 bindings; the evidence consists of two factual sentences (staffing requirement and assigned operators) rather than policy text; the counterfactual simply lowers the assigned operator count to 5, which is internally consistent and does not contradict other facts; neither context states or implies the final yes/no decision.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. All required white resin and labels are staged for the change. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. No maintenance work remains open, and no maintenance defect affecting Line 4 remains unresolved at the release decision. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "5 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-010", "id": "fast-43-diverse-292-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. All required white resin and labels are staged for the change. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 5 trained operators are assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. No maintenance work remains open, and no maintenance defect affecting Line 4 remains unresolved at the release decision. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release policy and question bindings for Line 4's 14:00 changeover; the counterfactual coherently reduces staffing from 6 to 4 without contradicting other facts, and the two evidence sentences are exact factual quotes with no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support staged all required white resin and labels. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:50, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "As of 13:50, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "As of 13:50, 4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "As of 13:50, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-012", "id": "fast-43-diverse-292-012-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support staged all required white resin and labels. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:50, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release policy and question bindings for Line 4's 14:00 changeover; the counterfactual coherently reduces staffing from 6 to 4 without contradicting other facts, and the two evidence sentences are exact factual quotes with no embedded answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support staged all required white resin and labels. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:50, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "As of 13:50, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "As of 13:50, 4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "As of 13:50, 6 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-012", "id": "fast-43-diverse-292-012-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support staged all required white resin and labels. The signed cleaning record shows completion at 13:20. The correct mold is installed and tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. As of 13:50, 4 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release policy and all readiness categories; the staffing numbers are permissible observation changes, the focus evidence is two factual sentences (not policy text), the counterfactual coherently drops assigned operators to 4 without contradicting other unchanged facts, and neither context reveals a gold answer or instructs the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At Northstar Components, Line 4 is scheduled for a 14:00 changeover from blue valve caps to white valve caps. Materials support staged all required white resin and labels before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed. Tooling checks passed. No maintenance work remains open; yesterday's temperature alarm did not recur during today's test and is not an unresolved defect. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "4 trained operators are currently assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-013", "id": "fast-43-diverse-292-013-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At Northstar Components, Line 4 is scheduled for a 14:00 changeover from blue valve caps to white valve caps. Materials support staged all required white resin and labels before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed. Tooling checks passed. No maintenance work remains open; yesterday's temperature alarm did not recur during today's test and is not an unresolved defect. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original release policy and all readiness categories; the staffing numbers are permissible observation changes, the focus evidence is two factual sentences (not policy text), the counterfactual coherently drops assigned operators to 4 without contradicting other unchanged facts, and neither context reveals a gold answer or instructs the classifier.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At Northstar Components, Line 4 is scheduled for a 14:00 changeover from blue valve caps to white valve caps. Materials support staged all required white resin and labels before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed. Tooling checks passed. No maintenance work remains open; yesterday's temperature alarm did not recur during today's test and is not an unresolved defect. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "4 trained operators are currently assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are currently assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-013", "id": "fast-43-diverse-292-013-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At Northstar Components, Line 4 is scheduled for a 14:00 changeover from blue valve caps to white valve caps. Materials support staged all required white resin and labels before release. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. 4 trained operators are currently assigned to Line 4 for the scheduled 14:00 change. The signed cleaning record shows completion at 13:20. The correct mold is installed. Tooling checks passed. No maintenance work remains open; yesterday's temperature alarm did not recur during today's test and is not an unresolved defect. The quality inspector approved the first-piece dimensions and color. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical governing policy and all readiness domains from the original question, only altering the case-specific staffing count (6 required vs 4 assigned) which is a permissible observation change; the two evidence sentences are plain factual statements, not policy definitions; the counterfactual staffing shortfall is internally consistent with the unchanged requirement and does not contradict other facts; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support staged all required white resin and labels for the change. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. There are 6 trained operators currently assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "There are 6 trained operators currently assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "There are 4 trained operators currently assigned to Line 4 for the scheduled 14:00 change.", "right": "There are 6 trained operators currently assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-014", "id": "fast-43-diverse-292-014-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support staged all required white resin and labels for the change. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. There are 6 trained operators currently assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain identical governing policy and all readiness domains from the original question, only altering the case-specific staffing count (6 required vs 4 assigned) which is a permissible observation change; the two evidence sentences are plain factual statements, not policy definitions; the counterfactual staffing shortfall is internally consistent with the unchanged requirement and does not contradict other facts; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support staged all required white resin and labels for the change. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. There are 6 trained operators currently assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators."}, {"path": [], "text": "There are 6 trained operators currently assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators.", "negative_right": "There are 4 trained operators currently assigned to Line 4 for the scheduled 14:00 change.", "right": "There are 6 trained operators currently assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-014", "id": "fast-43-diverse-292-014-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support staged all required white resin and labels for the change. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. A maintenance technician recorded that yesterday's temperature alarm did not recur during today's test and is not an unresolved defect; no maintenance work remains open. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operators. There are 4 trained operators currently assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy clause and all readiness conditions from the question; the focus evidence consists of two complete factual sentences about staffing requirement and assignment, not policy text; the counterfactual changes only the assigned operator count (7→4) creating a staffing shortfall without contradicting other unchanged facts; entity (Line 4), time (14:00) and path bindings are preserved; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support has staged all required white resin and labels. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. No maintenance work remains open, and the technician confirmed no unresolved maintenance defect affects the line. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_right": "4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-016", "id": "fast-43-diverse-292-016-base", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support has staged all required white resin and labels. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. No maintenance work remains open, and the technician confirmed no unresolved maintenance defect affects the line. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the full policy clause and all readiness conditions from the question; the focus evidence consists of two complete factual sentences about staffing requirement and assignment, not policy text; the counterfactual changes only the assigned operator count (7→4) creating a staffing shortfall without contradicting other unchanged facts; entity (Line 4), time (14:00) and path bindings are preserved; no gold answer, rule table, or output instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "full_context_fact_states": {"base": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "supported", "tooling_checks_passed": "supported"}, "counterfactual": {"cleaning_complete": "supported", "correct_mold_installed": "supported", "first_piece_approved": "supported", "maintenance_work_complete": "supported", "materials_staged": "supported", "no_unresolved_maintenance_defect": "supported", "staffing_sufficient": "refuted", "tooling_checks_passed": "supported"}, "remove_left": {"staffing_sufficient": "unknown"}, "remove_right": {"staffing_sufficient": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"staffing_sufficient": "unknown"}, "negative_pair": {"staffing_sufficient": "refuted"}, "negative_sentence": {"staffing_sufficient": "unknown"}, "positive_pair": {"staffing_sufficient": "supported"}, "right": {"staffing_sufficient": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "The atoms each state a single factual readiness relationship; splitting tooling and maintenance into separate facts does not create improper bundles. The focus atom is factual, and the base and counter assignments can differ only in staffing while all other readiness facts remain satisfied. Policy evidence correctly preserves the substantive release-only rule from the original state; requirements appearing in the retained questions need not be duplicated there.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Refuted staffing sufficiency entails that a required readiness condition is incomplete. Under the unchanged false criterion, any incomplete required condition is sufficient for a no-release decision.", "rule_index": 0, "sound": true}, {"reason": "The conjunction establishes completion of materials, staffing, cleaning, mold/tooling, maintenance work, and first-piece approval, while also expressly excluding any unresolved maintenance defect. It is therefore sufficient for release under the stated criteria and policy.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "materials_staged", "statement": "Every material required for Line 4's scheduled 14:00 change from blue valve caps to white valve caps is staged for that change before release."}, {"id": "staffing_sufficient", "statement": "The number of trained operators assigned to Line 4 for the scheduled 14:00 change is at least the number of positions required by the signed staffing specification for that change."}, {"id": "cleaning_complete", "statement": "The required cleaning of Line 4 for the scheduled 14:00 change is complete before release."}, {"id": "correct_mold_installed", "statement": "The mold specified for Line 4's scheduled 14:00 change to white valve caps is installed before release."}, {"id": "tooling_checks_passed", "statement": "Every required tooling check for Line 4's scheduled 14:00 change has passed before release."}, {"id": "maintenance_work_complete", "statement": "No required maintenance work for Line 4 remains open at the release decision."}, {"id": "no_unresolved_maintenance_defect", "statement": "No maintenance defect affecting Line 4 remains unresolved at the release decision."}, {"id": "first_piece_approved", "statement": "The quality inspector has approved the first piece for Line 4's scheduled 14:00 production of white valve caps before release."}], "base_state_json": "\"At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support has staged all required white resin and labels. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. No maintenance work remains open, and the technician confirmed no unresolved maintenance defect affects the line. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 7 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete.\"", "base_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "counter_states": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "refuted"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}], "focus_atom": "staffing_sufficient", "focus_evidence": [{"path": [], "text": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions."}, {"path": [], "text": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}], "policy_evidence": [{"path": [], "text": "Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}], "rules": [{"justification": "A refuted staffing-sufficiency comparison establishes that the trained-operator count is below the signed required-position count, so a required readiness condition is incomplete.", "target": "false", "when": [{"atom_id": "staffing_sufficient", "state": "refuted"}]}, {"justification": "All required materials, staffing, cleaning, tooling, maintenance, and first-piece readiness conditions are complete, and no unresolved maintenance defect remains.", "target": "true", "when": [{"atom_id": "materials_staged", "state": "supported"}, {"atom_id": "staffing_sufficient", "state": "supported"}, {"atom_id": "cleaning_complete", "state": "supported"}, {"atom_id": "correct_mold_installed", "state": "supported"}, {"atom_id": "tooling_checks_passed", "state": "supported"}, {"atom_id": "maintenance_work_complete", "state": "supported"}, {"atom_id": "no_unresolved_maintenance_defect", "state": "supported"}, {"atom_id": "first_piece_approved", "state": "supported"}]}]}, "verified_pair": {"left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_left": "The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions.", "negative_right": "4 trained operators are assigned to Line 4 for the scheduled 14:00 change.", "right": "7 trained operators are assigned to Line 4 for the scheduled 14:00 change."}, "verifier_independent_model": false}, "family": "fast-43-diverse-292-016", "id": "fast-43-diverse-292-016-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No — do not release Line 4 because at least one required readiness condition is incomplete or an unresolved maintenance defect remains.", "true": "Yes — release Line 4 because every required readiness condition is complete and no unresolved maintenance defect remains."}, "instructions": "Decide whether Line 4 should be released to production under the stated policy. Answer yes or no.", "type": "noul"}}, "state": "At fictional Northstar Components, Line 4 is scheduled to change from blue valve caps to white caps at 14:00. Materials support has staged all required white resin and labels. The signed cleaning record shows completion at 13:20. The correct mold is installed and all tooling checks passed. No maintenance work remains open, and the technician confirmed no unresolved maintenance defect affects the line. The quality inspector approved the first-piece dimensions and color. The signed staffing specification for Line 4's scheduled 14:00 change requires 6 trained operator positions. 4 trained operators are assigned to Line 4 for the scheduled 14:00 change. Policy permits release only when materials, staffing, cleaning, tooling, maintenance, and first-piece approval are all complete."}, "method": "c2d", "provenance": {"source_id": "diverse-292", "source_is_synthetic": true, "source_sha256": "7fb1f72997e555e40c61d3f22fd3bddb22f1c9e548f2b6355b95ff79dc2f0fcb", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all six mandatory gates (mold, cleaning record, operators, quality approval, cleared stop-work hold, resin sufficiency) and the L4/06:52 bindings, with the counterfactual altering only the single staged-resin figure (480→380 kg) coherently against the unchanged 450 kg requirement, and the two evidence sentences are plain factual statements with no embedded labels or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"Changeover L4 (Jar-A to Jar-B) is being reviewed as of 06:52.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The Jar-B mold is installed, cleaning record C-441 is signed, and all four required operators are present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"The earlier 06:28 stop-work hold was cleared at 06:46 after switch replacement and successful dry cycles, so no current stop-work hold applies.\"},{\"speaker\":\"Quality inspector\",\"text\":\"I verified C-441 and approved the setup sample for line release. The delayed reserve resin delivery is documented as an issue that could disrupt continued production, though it does not block line release.\"},{\"speaker\":\"Scale technician\",\"text\":\"At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The first lot for changeover L4 requires 450 kilograms of resin per the current work order.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["4", "text"], "text": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover."}, {"path": ["5", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin per the current work order."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.", "negative_left": "At 06:52, the scale log shows 380 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin per the current work order.", "right": "The first lot for changeover L4 requires 450 kilograms of resin per the current work order."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-007", "id": "fast-43-diverse-294-007-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "Changeover L4 (Jar-A to Jar-B) is being reviewed as of 06:52."}, {"speaker": "Line supervisor", "text": "The Jar-B mold is installed, cleaning record C-441 is signed, and all four required operators are present."}, {"speaker": "Maintenance technician", "text": "The earlier 06:28 stop-work hold was cleared at 06:46 after switch replacement and successful dry cycles, so no current stop-work hold applies."}, {"speaker": "Quality inspector", "text": "I verified C-441 and approved the setup sample for line release. The delayed reserve resin delivery is documented as an issue that could disrupt continued production, though it does not block line release."}, {"speaker": "Scale technician", "text": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover."}, {"speaker": "Production scheduler", "text": "The first lot for changeover L4 requires 450 kilograms of resin per the current work order."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all six mandatory gates (mold, cleaning record, operators, quality approval, cleared stop-work hold, resin sufficiency) and the L4/06:52 bindings, with the counterfactual altering only the single staged-resin figure (480→380 kg) coherently against the unchanged 450 kg requirement, and the two evidence sentences are plain factual statements with no embedded labels or rules.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"Changeover L4 (Jar-A to Jar-B) is being reviewed as of 06:52.\"},{\"speaker\":\"Line supervisor\",\"text\":\"The Jar-B mold is installed, cleaning record C-441 is signed, and all four required operators are present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"The earlier 06:28 stop-work hold was cleared at 06:46 after switch replacement and successful dry cycles, so no current stop-work hold applies.\"},{\"speaker\":\"Quality inspector\",\"text\":\"I verified C-441 and approved the setup sample for line release. The delayed reserve resin delivery is documented as an issue that could disrupt continued production, though it does not block line release.\"},{\"speaker\":\"Scale technician\",\"text\":\"At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The first lot for changeover L4 requires 450 kilograms of resin per the current work order.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["4", "text"], "text": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover."}, {"path": ["5", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin per the current work order."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.", "negative_left": "At 06:52, the scale log shows 380 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin per the current work order.", "right": "The first lot for changeover L4 requires 450 kilograms of resin per the current work order."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-007", "id": "fast-43-diverse-294-007-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "Changeover L4 (Jar-A to Jar-B) is being reviewed as of 06:52."}, {"speaker": "Line supervisor", "text": "The Jar-B mold is installed, cleaning record C-441 is signed, and all four required operators are present."}, {"speaker": "Maintenance technician", "text": "The earlier 06:28 stop-work hold was cleared at 06:46 after switch replacement and successful dry cycles, so no current stop-work hold applies."}, {"speaker": "Quality inspector", "text": "I verified C-441 and approved the setup sample for line release. The delayed reserve resin delivery is documented as an issue that could disrupt continued production, though it does not block line release."}, {"speaker": "Scale technician", "text": "At 06:52, the scale log shows 380 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover."}, {"speaker": "Production scheduler", "text": "The first lot for changeover L4 requires 450 kilograms of resin per the current work order."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six release gates and the 06:52 timeframe/L4 Jar-A→Jar-B binding from the untouched question; the focus evidence is two complete factual sentences with no policy text; the counterfactual only changes staged resin to 380kg, consistently creating a shortfall against the untouched 450kg requirement without contradicting other facts; no answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:46, the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, C-441 was verified and the setup sample was approved for line release. The reserve-resin delivery remains delayed and is documented as a nonblocking issue that could disrupt continued production, though it does not block release of this changeover.\"},{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The first lot for changeover L4 requires 450 kilograms of resin.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms."}, {"path": ["4", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms.", "negative_left": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 380 kilograms.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin.", "right": "The first lot for changeover L4 requires 450 kilograms of resin."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-009", "id": "fast-43-diverse-294-009-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:46, the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared."}, {"speaker": "Quality inspector", "text": "At 06:52, C-441 was verified and the setup sample was approved for line release. The reserve-resin delivery remains delayed and is documented as a nonblocking issue that could disrupt continued production, though it does not block release of this changeover."}, {"speaker": "Production scheduler", "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms."}, {"speaker": "Production scheduler", "text": "The first lot for changeover L4 requires 450 kilograms of resin."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six release gates and the 06:52 timeframe/L4 Jar-A→Jar-B binding from the untouched question; the focus evidence is two complete factual sentences with no policy text; the counterfactual only changes staged resin to 380kg, consistently creating a shortfall against the untouched 450kg requirement without contradicting other facts; no answer, rule table, or rationale is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:46, the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, C-441 was verified and the setup sample was approved for line release. The reserve-resin delivery remains delayed and is documented as a nonblocking issue that could disrupt continued production, though it does not block release of this changeover.\"},{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The first lot for changeover L4 requires 450 kilograms of resin.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms."}, {"path": ["4", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms.", "negative_left": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 380 kilograms.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin.", "right": "The first lot for changeover L4 requires 450 kilograms of resin."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-009", "id": "fast-43-diverse-294-009-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:46, the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared."}, {"speaker": "Quality inspector", "text": "At 06:52, C-441 was verified and the setup sample was approved for line release. The reserve-resin delivery remains delayed and is documented as a nonblocking issue that could disrupt continued production, though it does not block release of this changeover."}, {"speaker": "Production scheduler", "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 380 kilograms."}, {"speaker": "Production scheduler", "text": "The first lot for changeover L4 requires 450 kilograms of resin."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all mandatory release gates and the original question verbatim, focus evidence consists of two complete factual sentences about lot requirement and staged resin, the counterfactual only changes the staged resin amount (200→120 kg) without contradicting other timestamped facts, and no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored as a documented nonblocking concern that could disrupt continued production, but it does not block line release.\"},{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the first lot for changeover L4 requires 180 kilograms of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 200 kilograms.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["5", "text"], "text": "At 06:52, the first lot for changeover L4 requires 180 kilograms of resin."}, {"path": ["6", "text"], "text": "At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 200 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first lot for changeover L4 requires 180 kilograms of resin.", "negative_left": "At 06:52, the first lot for changeover L4 requires 180 kilograms of resin.", "negative_right": "At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 120 kilograms.", "right": "At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 200 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-010", "id": "fast-43-diverse-294-010-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored as a documented nonblocking concern that could disrupt continued production, but it does not block line release."}, {"speaker": "Production scheduler", "text": "At 06:52, the first lot for changeover L4 requires 180 kilograms of resin."}, {"speaker": "Line supervisor", "text": "At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 200 kilograms."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all mandatory release gates and the original question verbatim, focus evidence consists of two complete factual sentences about lot requirement and staged resin, the counterfactual only changes the staged resin amount (200→120 kg) without contradicting other timestamped facts, and no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored as a documented nonblocking concern that could disrupt continued production, but it does not block line release.\"},{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the first lot for changeover L4 requires 180 kilograms of resin.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 200 kilograms.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["5", "text"], "text": "At 06:52, the first lot for changeover L4 requires 180 kilograms of resin."}, {"path": ["6", "text"], "text": "At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 200 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first lot for changeover L4 requires 180 kilograms of resin.", "negative_left": "At 06:52, the first lot for changeover L4 requires 180 kilograms of resin.", "negative_right": "At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 120 kilograms.", "right": "At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 200 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-010", "id": "fast-43-diverse-294-010-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored as a documented nonblocking concern that could disrupt continued production, but it does not block line release."}, {"speaker": "Production scheduler", "text": "At 06:52, the first lot for changeover L4 requires 180 kilograms of resin."}, {"speaker": "Line supervisor", "text": "At 06:52, the resin staged at line L4 for the Jar-A to Jar-B changeover totals 120 kilograms."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep all six mandatory gate facts and only vary the staged-resin quantity as a case observation, with the original question’s criteria supplying the governing policy; the two evidence sentences are plain factual statements, the counterfactual’s 380 kg (< 450 kg required) is internally consistent with the rest of the unchanged timeline, and no gold answer, rule table, or instructive rationale is embedded—the inspector's 'does not itself block line release' merely echoes generic criteria language, not a case-specific verdict.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The first lot for changeover L4 requires 450 kilograms of resin. The delayed reserve resin should be monitored as a documented issue that could disrupt continued production, though it does not itself block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms."}, {"path": ["4", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms.", "negative_left": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 380 kilograms.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin.", "right": "The first lot for changeover L4 requires 450 kilograms of resin."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-012", "id": "fast-43-diverse-294-012-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The first lot for changeover L4 requires 450 kilograms of resin. The delayed reserve resin should be monitored as a documented issue that could disrupt continued production, though it does not itself block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep all six mandatory gate facts and only vary the staged-resin quantity as a case observation, with the original question’s criteria supplying the governing policy; the two evidence sentences are plain factual statements, the counterfactual’s 380 kg (< 450 kg required) is internally consistent with the rest of the unchanged timeline, and no gold answer, rule table, or instructive rationale is embedded—the inspector's 'does not itself block line release' merely echoes generic criteria language, not a case-specific verdict.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The first lot for changeover L4 requires 450 kilograms of resin. The delayed reserve resin should be monitored as a documented issue that could disrupt continued production, though it does not itself block line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms."}, {"path": ["4", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 480 kilograms.", "negative_left": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 380 kilograms.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin.", "right": "The first lot for changeover L4 requires 450 kilograms of resin."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-012", "id": "fast-43-diverse-294-012-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B weighs 380 kilograms."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The first lot for changeover L4 requires 450 kilograms of resin. The delayed reserve resin should be monitored as a documented issue that could disrupt continued production, though it does not itself block line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates and the same question object; only the staged resin quantity changes (480 kg vs 380 kg) against the fixed 450 kg lot requirement, giving a coherent, non-contradictory counterfactual; evidence spans are plain factual sentences with no gold answer, rule table, or output instructions embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored as a nonblocking issue that could disrupt continued production.\"},{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, the scale log shows 480 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap.\"},{\"speaker\":\"Materials clerk\",\"text\":\"The first lot for changeover L4 requires 450 kilograms of resin.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["5", "text"], "text": "At 06:52, the scale log shows 480 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap."}, {"path": ["6", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the scale log shows 480 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap.", "negative_left": "At 06:52, the scale log shows 380 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin.", "right": "The first lot for changeover L4 requires 450 kilograms of resin."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-013", "id": "fast-43-diverse-294-013-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored as a nonblocking issue that could disrupt continued production."}, {"speaker": "Materials clerk", "text": "At 06:52, the scale log shows 480 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap."}, {"speaker": "Materials clerk", "text": "The first lot for changeover L4 requires 450 kilograms of resin."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates and the same question object; only the staged resin quantity changes (480 kg vs 380 kg) against the fixed 450 kg lot requirement, giving a coherent, non-contradictory counterfactual; evidence spans are plain factual sentences with no gold answer, rule table, or output instructions embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored as a nonblocking issue that could disrupt continued production.\"},{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, the scale log shows 480 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap.\"},{\"speaker\":\"Materials clerk\",\"text\":\"The first lot for changeover L4 requires 450 kilograms of resin.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["5", "text"], "text": "At 06:52, the scale log shows 480 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap."}, {"path": ["6", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the scale log shows 480 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap.", "negative_left": "At 06:52, the scale log shows 380 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin.", "right": "The first lot for changeover L4 requires 450 kilograms of resin."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-013", "id": "fast-43-diverse-294-013-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored as a nonblocking issue that could disrupt continued production."}, {"speaker": "Materials clerk", "text": "At 06:52, the scale log shows 380 kilograms of resin staged at changeover L4 for the Jar-A to Jar-B swap."}, {"speaker": "Materials clerk", "text": "The first lot for changeover L4 requires 450 kilograms of resin."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates from the original question/policy, keep entity 'changeover L4 Jar-A to Jar-B' and timestamp 06:52 intact, use two factual quantity sentences as evidence ('requires 180 kilograms' and 'measures 200/150 kilograms'), the counterfactual coherently alters only the staged resin amount without contradicting other unchanged facts, and neither context reveals a gold label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\": \"Production scheduler\", \"text\": \"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00.\"}, {\"speaker\": \"Line supervisor\", \"text\": \"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"}, {\"speaker\": \"Maintenance technician\", \"text\": \"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"}, {\"speaker\": \"Maintenance technician\", \"text\": \"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"}, {\"speaker\": \"Quality inspector\", \"text\": \"At 06:52, I verified C-441 and approved the setup sample for line release. No current stop-work hold applies. The delayed reserve resin is documented as an issue that could disrupt continued production, though it does not block release.\"}, {\"speaker\": \"Production scheduler\", \"text\": \"At 06:52, the first lot of changeover L4 requires 180 kilograms of resin.\"}, {\"speaker\": \"Production scheduler\", \"text\": \"At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 200 kilograms.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["5", "text"], "text": "At 06:52, the first lot of changeover L4 requires 180 kilograms of resin."}, {"path": ["6", "text"], "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 200 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first lot of changeover L4 requires 180 kilograms of resin.", "negative_left": "At 06:52, the first lot of changeover L4 requires 180 kilograms of resin.", "negative_right": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 150 kilograms.", "right": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 200 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-016", "id": "fast-43-diverse-294-016-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. No current stop-work hold applies. The delayed reserve resin is documented as an issue that could disrupt continued production, though it does not block release."}, {"speaker": "Production scheduler", "text": "At 06:52, the first lot of changeover L4 requires 180 kilograms of resin."}, {"speaker": "Production scheduler", "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 200 kilograms."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates from the original question/policy, keep entity 'changeover L4 Jar-A to Jar-B' and timestamp 06:52 intact, use two factual quantity sentences as evidence ('requires 180 kilograms' and 'measures 200/150 kilograms'), the counterfactual coherently alters only the staged resin amount without contradicting other unchanged facts, and neither context reveals a gold label, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\": \"Production scheduler\", \"text\": \"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00.\"}, {\"speaker\": \"Line supervisor\", \"text\": \"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"}, {\"speaker\": \"Maintenance technician\", \"text\": \"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"}, {\"speaker\": \"Maintenance technician\", \"text\": \"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"}, {\"speaker\": \"Quality inspector\", \"text\": \"At 06:52, I verified C-441 and approved the setup sample for line release. No current stop-work hold applies. The delayed reserve resin is documented as an issue that could disrupt continued production, though it does not block release.\"}, {\"speaker\": \"Production scheduler\", \"text\": \"At 06:52, the first lot of changeover L4 requires 180 kilograms of resin.\"}, {\"speaker\": \"Production scheduler\", \"text\": \"At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 200 kilograms.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["5", "text"], "text": "At 06:52, the first lot of changeover L4 requires 180 kilograms of resin."}, {"path": ["6", "text"], "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 200 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first lot of changeover L4 requires 180 kilograms of resin.", "negative_left": "At 06:52, the first lot of changeover L4 requires 180 kilograms of resin.", "negative_right": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 150 kilograms.", "right": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 200 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-016", "id": "fast-43-diverse-294-016-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. No current stop-work hold applies. The delayed reserve resin is documented as an issue that could disrupt continued production, though it does not block release."}, {"speaker": "Production scheduler", "text": "At 06:52, the first lot of changeover L4 requires 180 kilograms of resin."}, {"speaker": "Production scheduler", "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B measures 150 kilograms."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all mandatory gate evidence (mold, cleaning, operators, hold clearance, quality approval, resin need) while only the staged resin weight changes between 180kg (met) and 120kg (unmet), a coherent single-fact counterfactual with no embedded answers or rule tables.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 180 kilograms. The first lot for changeover L4 requires 150 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin, now expected at 10:00 instead of 08:00, is documented as an issue that could disrupt continued production, though it does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 180 kilograms."}, {"path": ["0", "text"], "text": "The first lot for changeover L4 requires 150 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 180 kilograms.", "negative_left": "At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 120 kilograms.", "negative_right": "The first lot for changeover L4 requires 150 kilograms of resin.", "right": "The first lot for changeover L4 requires 150 kilograms of resin."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-017", "id": "fast-43-diverse-294-017-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 180 kilograms. The first lot for changeover L4 requires 150 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin, now expected at 10:00 instead of 08:00, is documented as an issue that could disrupt continued production, though it does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all mandatory gate evidence (mold, cleaning, operators, hold clearance, quality approval, resin need) while only the staged resin weight changes between 180kg (met) and 120kg (unmet), a coherent single-fact counterfactual with no embedded answers or rule tables.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 180 kilograms. The first lot for changeover L4 requires 150 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin, now expected at 10:00 instead of 08:00, is documented as an issue that could disrupt continued production, though it does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 180 kilograms."}, {"path": ["0", "text"], "text": "The first lot for changeover L4 requires 150 kilograms of resin."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 180 kilograms.", "negative_left": "At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 120 kilograms.", "negative_right": "The first lot for changeover L4 requires 150 kilograms of resin.", "right": "The first lot for changeover L4 requires 150 kilograms of resin."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-017", "id": "fast-43-diverse-294-017-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the resin staged at station L4 for the Jar-A to Jar-B changeover weighs 120 kilograms. The first lot for changeover L4 requires 150 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin, now expected at 10:00 instead of 08:00, is documented as an issue that could disrupt continued production, though it does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the governing release‑gate policy intact via the unchanged question, only altering case‑specific staged/required resin quantities, which is a permissible observation change; the focus evidence consists of two verbatim factual sentences with no policy text or gold‑answer leakage, and the counterfactual (120kg vs 150kg required) is internally consistent with the rest of the unchanged context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release.\"},{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, the scale log shows 180 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover.\"},{\"speaker\":\"Materials clerk\",\"text\":\"The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The reserve-resin delivery for changeover L4 remains delayed and is documented as an issue that could disrupt continued production, though it does not block current line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "At 06:52, the scale log shows 180 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover."}, {"path": ["4", "text"], "text": "The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the scale log shows 180 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover.", "negative_left": "At 06:52, the scale log shows 120 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover.", "negative_right": "The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket.", "right": "The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-019", "id": "fast-43-diverse-294-019-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release."}, {"speaker": "Materials clerk", "text": "At 06:52, the scale log shows 180 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover."}, {"speaker": "Materials clerk", "text": "The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket."}, {"speaker": "Production scheduler", "text": "The reserve-resin delivery for changeover L4 remains delayed and is documented as an issue that could disrupt continued production, though it does not block current line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the governing release‑gate policy intact via the unchanged question, only altering case‑specific staged/required resin quantities, which is a permissible observation change; the focus evidence consists of two verbatim factual sentences with no policy text or gold‑answer leakage, and the counterfactual (120kg vs 150kg required) is internally consistent with the rest of the unchanged context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release.\"},{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, the scale log shows 180 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover.\"},{\"speaker\":\"Materials clerk\",\"text\":\"The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The reserve-resin delivery for changeover L4 remains delayed and is documented as an issue that could disrupt continued production, though it does not block current line release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["3", "text"], "text": "At 06:52, the scale log shows 180 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover."}, {"path": ["4", "text"], "text": "The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the scale log shows 180 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover.", "negative_left": "At 06:52, the scale log shows 120 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover.", "negative_right": "The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket.", "right": "The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-019", "id": "fast-43-diverse-294-019-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release."}, {"speaker": "Materials clerk", "text": "At 06:52, the scale log shows 120 kilograms of resin staged at line L4 for the Jar-A to Jar-B changeover."}, {"speaker": "Materials clerk", "text": "The first lot for changeover L4 requires 150 kilograms of resin per the batch ticket."}, {"speaker": "Production scheduler", "text": "The reserve-resin delivery for changeover L4 remains delayed and is documented as an issue that could disrupt continued production, though it does not block current line release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question object and all mandatory gate criteria while only altering the staged-resin quantity (480kg vs 420kg) against a required 450kg, a coherent single-fact change with no contradictions, no embedded gold answers or rule tables, and evidence spans are single factual sentences about staged/required resin quantities.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\": \"Production scheduler\", \"text\": \"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, the resin staged for changeover L4's first lot totals 480 kilograms.\"}, {\"speaker\": \"Line supervisor\", \"text\": \"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"}, {\"speaker\": \"Maintenance technician\", \"text\": \"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test. Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"}, {\"speaker\": \"Quality inspector\", \"text\": \"At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms. The reserve delivery has slipped from 08:00 to 10:00 and should be monitored as a documented issue that could disrupt continued production, though it does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the resin staged for changeover L4's first lot totals 480 kilograms."}, {"path": ["3", "text"], "text": "As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the resin staged for changeover L4's first lot totals 480 kilograms.", "negative_left": "As of 06:52, the resin staged for changeover L4's first lot totals 420 kilograms.", "negative_right": "As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms.", "right": "As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-023", "id": "fast-43-diverse-294-023-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, the resin staged for changeover L4's first lot totals 480 kilograms."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test. Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms. The reserve delivery has slipped from 08:00 to 10:00 and should be monitored as a documented issue that could disrupt continued production, though it does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original question object and all mandatory gate criteria while only altering the staged-resin quantity (480kg vs 420kg) against a required 450kg, a coherent single-fact change with no contradictions, no embedded gold answers or rule tables, and evidence spans are single factual sentences about staged/required resin quantities.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\": \"Production scheduler\", \"text\": \"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, the resin staged for changeover L4's first lot totals 480 kilograms.\"}, {\"speaker\": \"Line supervisor\", \"text\": \"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"}, {\"speaker\": \"Maintenance technician\", \"text\": \"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test. Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"}, {\"speaker\": \"Quality inspector\", \"text\": \"At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms. The reserve delivery has slipped from 08:00 to 10:00 and should be monitored as a documented issue that could disrupt continued production, though it does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the resin staged for changeover L4's first lot totals 480 kilograms."}, {"path": ["3", "text"], "text": "As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the resin staged for changeover L4's first lot totals 480 kilograms.", "negative_left": "As of 06:52, the resin staged for changeover L4's first lot totals 420 kilograms.", "negative_right": "As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms.", "right": "As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-023", "id": "fast-43-diverse-294-023-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, the resin staged for changeover L4's first lot totals 420 kilograms."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test. Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, the resin quantity required for changeover L4's first lot is 450 kilograms. The reserve delivery has slipped from 08:00 to 10:00 and should be monitored as a documented issue that could disrupt continued production, though it does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original gates and timestamps while only changing the staged resin figure (240→180 kg) in a single factual sentence, with no policy, exception, or answer leakage introduced; the two focus sentences are plain factual statements ('requires 220 kilograms' and 'measures 240/180 kilograms') not policy text, and the counterfactual does not create any duplicate or contradictory measurement elsewhere in the context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00, an issue that could disrupt continued production, though it does not block line release.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored.\"},{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin.\"},{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the resin staged for changeover L4 measures 240 kilograms.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["5", "text"], "text": "At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin."}, {"path": ["6", "text"], "text": "At 06:52, the resin staged for changeover L4 measures 240 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin.", "negative_left": "At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin.", "negative_right": "At 06:52, the resin staged for changeover L4 measures 180 kilograms.", "right": "At 06:52, the resin staged for changeover L4 measures 240 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-024", "id": "fast-43-diverse-294-024-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00, an issue that could disrupt continued production, though it does not block line release."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored."}, {"speaker": "Production scheduler", "text": "At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin."}, {"speaker": "Production scheduler", "text": "At 06:52, the resin staged for changeover L4 measures 240 kilograms."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original gates and timestamps while only changing the staged resin figure (240→180 kg) in a single factual sentence, with no policy, exception, or answer leakage introduced; the two focus sentences are plain factual statements ('requires 220 kilograms' and 'measures 240/180 kilograms') not policy text, and the counterfactual does not create any duplicate or contradictory measurement elsewhere in the context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00, an issue that could disrupt continued production, though it does not block line release.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored.\"},{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin.\"},{\"speaker\":\"Production scheduler\",\"text\":\"At 06:52, the resin staged for changeover L4 measures 240 kilograms.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["5", "text"], "text": "At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin."}, {"path": ["6", "text"], "text": "At 06:52, the resin staged for changeover L4 measures 240 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin.", "negative_left": "At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin.", "negative_right": "At 06:52, the resin staged for changeover L4 measures 180 kilograms.", "right": "At 06:52, the resin staged for changeover L4 measures 240 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-024", "id": "fast-43-diverse-294-024-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00, an issue that could disrupt continued production, though it does not block line release."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored."}, {"speaker": "Production scheduler", "text": "At 06:52, the first lot for changeover L4 from Jar-A to Jar-B requires 220 kilograms of resin."}, {"speaker": "Production scheduler", "text": "At 06:52, the resin staged for changeover L4 measures 180 kilograms."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all mandatory gate facts and the unchanged question preserves policy; only the staged-resin figure changes (480→380 kg) which is a permissible case observation, not a policy element; the two focus evidence spans are complete factual sentences ('At 06:52, the scale log shows 480 kilograms of resin staged...' and 'The first lot for changeover L4 requires 450 kilograms of resin...'); no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Case log\",\"text\":\"Changeover L4 status as of 06:52. Jar-B mold installed and verified. Cleaning record C-441 signed and verified by quality. Four required operators confirmed present at the line. Quality inspector approved the setup sample for release. No stop-work hold currently applies; the earlier proximity-switch hold was cleared at 06:46 after successful dry cycles. The reserve-resin delivery remains delayed and is documented as an open issue that could disrupt continued production once the first lot is running, though it does not block the immediate release decision.\"},{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The first lot for changeover L4 requires 450 kilograms of resin per the production order.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover."}, {"path": ["2", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin per the production order."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.", "negative_left": "At 06:52, the scale log shows 380 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin per the production order.", "right": "The first lot for changeover L4 requires 450 kilograms of resin per the production order."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-026", "id": "fast-43-diverse-294-026-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Case log", "text": "Changeover L4 status as of 06:52. Jar-B mold installed and verified. Cleaning record C-441 signed and verified by quality. Four required operators confirmed present at the line. Quality inspector approved the setup sample for release. No stop-work hold currently applies; the earlier proximity-switch hold was cleared at 06:46 after successful dry cycles. The reserve-resin delivery remains delayed and is documented as an open issue that could disrupt continued production once the first lot is running, though it does not block the immediate release decision."}, {"speaker": "Materials clerk", "text": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover."}, {"speaker": "Production scheduler", "text": "The first lot for changeover L4 requires 450 kilograms of resin per the production order."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all mandatory gate facts and the unchanged question preserves policy; only the staged-resin figure changes (480→380 kg) which is a permissible case observation, not a policy element; the two focus evidence spans are complete factual sentences ('At 06:52, the scale log shows 480 kilograms of resin staged...' and 'The first lot for changeover L4 requires 450 kilograms of resin...'); no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Case log\",\"text\":\"Changeover L4 status as of 06:52. Jar-B mold installed and verified. Cleaning record C-441 signed and verified by quality. Four required operators confirmed present at the line. Quality inspector approved the setup sample for release. No stop-work hold currently applies; the earlier proximity-switch hold was cleared at 06:46 after successful dry cycles. The reserve-resin delivery remains delayed and is documented as an open issue that could disrupt continued production once the first lot is running, though it does not block the immediate release decision.\"},{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.\"},{\"speaker\":\"Production scheduler\",\"text\":\"The first lot for changeover L4 requires 450 kilograms of resin per the production order.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["1", "text"], "text": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover."}, {"path": ["2", "text"], "text": "The first lot for changeover L4 requires 450 kilograms of resin per the production order."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the scale log shows 480 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.", "negative_left": "At 06:52, the scale log shows 380 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover.", "negative_right": "The first lot for changeover L4 requires 450 kilograms of resin per the production order.", "right": "The first lot for changeover L4 requires 450 kilograms of resin per the production order."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-026", "id": "fast-43-diverse-294-026-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Case log", "text": "Changeover L4 status as of 06:52. Jar-B mold installed and verified. Cleaning record C-441 signed and verified by quality. Four required operators confirmed present at the line. Quality inspector approved the setup sample for release. No stop-work hold currently applies; the earlier proximity-switch hold was cleared at 06:46 after successful dry cycles. The reserve-resin delivery remains delayed and is documented as an open issue that could disrupt continued production once the first lot is running, though it does not block the immediate release decision."}, {"speaker": "Materials clerk", "text": "At 06:52, the scale log shows 380 kilograms of resin staged at press L4 for the Jar-A to Jar-B changeover."}, {"speaker": "Production scheduler", "text": "The first lot for changeover L4 requires 450 kilograms of resin per the production order."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates and the original question verbatim, keep the same L4/06:52 bindings, use two complete factual sentences as evidence, and the counterfactual coherently swaps staged resin to 140kg (below the 180kg requirement) without contradicting other unchanged facts or leaking any gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742. The reserve delivery has slipped from 08:00 to 10:00, and this delay is logged as an open issue that could disrupt continued production if not monitored.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present at station L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared. No hold currently applies.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, 210 kilograms of resin are staged at station L4 per staging log SL-9903. The delayed reserve resin should be monitored.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742."}, {"path": ["3", "text"], "text": "As of 06:52, 210 kilograms of resin are staged at station L4 per staging log SL-9903."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742.", "negative_left": "As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742.", "negative_right": "As of 06:52, 140 kilograms of resin are staged at station L4 per staging log SL-9903.", "right": "As of 06:52, 210 kilograms of resin are staged at station L4 per staging log SL-9903."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-034", "id": "fast-43-diverse-294-034-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742. The reserve delivery has slipped from 08:00 to 10:00, and this delay is logged as an open issue that could disrupt continued production if not monitored."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present at station L4."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared. No hold currently applies."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, 210 kilograms of resin are staged at station L4 per staging log SL-9903. The delayed reserve resin should be monitored."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates and the original question verbatim, keep the same L4/06:52 bindings, use two complete factual sentences as evidence, and the counterfactual coherently swaps staged resin to 140kg (below the 180kg requirement) without contradicting other unchanged facts or leaking any gold answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742. The reserve delivery has slipped from 08:00 to 10:00, and this delay is logged as an open issue that could disrupt continued production if not monitored.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present at station L4.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared. No hold currently applies.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, 210 kilograms of resin are staged at station L4 per staging log SL-9903. The delayed reserve resin should be monitored.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742."}, {"path": ["3", "text"], "text": "As of 06:52, 210 kilograms of resin are staged at station L4 per staging log SL-9903."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742.", "negative_left": "As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742.", "negative_right": "As of 06:52, 140 kilograms of resin are staged at station L4 per staging log SL-9903.", "right": "As of 06:52, 210 kilograms of resin are staged at station L4 per staging log SL-9903."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-034", "id": "fast-43-diverse-294-034-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, changeover L4's first lot requires 180 kilograms of resin according to work order WO-7742. The reserve delivery has slipped from 08:00 to 10:00, and this delay is logged as an open issue that could disrupt continued production if not monitored."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present at station L4."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: the mold-clamp proximity switch was replaced, three successful dry cycles were completed, and the stop-work hold was formally cleared. No hold currently applies."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, 140 kilograms of resin are staged at station L4 per staging log SL-9903. The delayed reserve resin should be monitored."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates and routing rules via the unchanged question, only altering staged resin quantity (200kg vs 150kg) while keeping the 180kg requirement fixed, which is a coherent case-level change since it plausibly fails the material gate without contradicting other reported facts; the evidence spans are two complete factual sentences reporting resin requirement and staged amount, and neither context contains rule tables, gold labels, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00 and remains a documented concern for continued production.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, 200 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch. The delayed reserve resin should be monitored but does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin."}, {"path": ["4", "text"], "text": "As of 06:52, 200 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin.", "negative_left": "As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin.", "negative_right": "As of 06:52, 150 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch.", "right": "As of 06:52, 200 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-035", "id": "fast-43-diverse-294-035-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00 and remains a documented concern for continued production."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, 200 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch. The delayed reserve resin should be monitored but does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates and routing rules via the unchanged question, only altering staged resin quantity (200kg vs 150kg) while keeping the 180kg requirement fixed, which is a coherent case-level change since it plausibly fails the material gate without contradicting other reported facts; the evidence spans are two complete factual sentences reporting resin requirement and staged amount, and neither context contains rule tables, gold labels, or output instructions.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00 and remains a documented concern for continued production.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, 200 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch. The delayed reserve resin should be monitored but does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin."}, {"path": ["4", "text"], "text": "As of 06:52, 200 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin.", "negative_left": "As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin.", "negative_right": "As of 06:52, 150 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch.", "right": "As of 06:52, 200 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-035", "id": "fast-43-diverse-294-035-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. As of 06:52, the first lot for changeover L4 requires 180 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00 and remains a documented concern for continued production."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. As of 06:52, 150 kilograms of resin are staged at changeover L4 for the Jar-A to Jar-B switch. The delayed reserve resin should be monitored but does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all mandatory gate criteria and timestamps from the unchanged question while altering only the staged resin quantity, and the evidence spans are two complete factual sentences without embedded rules or gold answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\": \"Production scheduler\", \"text\": \"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms. The reserve delivery has slipped from 08:00 to 10:00 and is being tracked as a documented issue that could disrupt continued production.\"}, {\"speaker\": \"Line supervisor\", \"text\": \"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"}, {\"speaker\": \"Maintenance technician\", \"text\": \"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test. Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"}, {\"speaker\": \"Materials clerk\", \"text\": \"At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 210 kilograms.\"}, {\"speaker\": \"Quality inspector\", \"text\": \"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored but does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms."}, {"path": ["3", "text"], "text": "At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 210 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms.", "negative_left": "At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms.", "negative_right": "At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 140 kilograms.", "right": "At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 210 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-040", "id": "fast-43-diverse-294-040-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms. The reserve delivery has slipped from 08:00 to 10:00 and is being tracked as a documented issue that could disrupt continued production."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test. Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Materials clerk", "text": "At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 210 kilograms."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored but does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all mandatory gate criteria and timestamps from the unchanged question while altering only the staged resin quantity, and the evidence spans are two complete factual sentences without embedded rules or gold answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\": \"Production scheduler\", \"text\": \"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms. The reserve delivery has slipped from 08:00 to 10:00 and is being tracked as a documented issue that could disrupt continued production.\"}, {\"speaker\": \"Line supervisor\", \"text\": \"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"}, {\"speaker\": \"Maintenance technician\", \"text\": \"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test. Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"}, {\"speaker\": \"Materials clerk\", \"text\": \"At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 210 kilograms.\"}, {\"speaker\": \"Quality inspector\", \"text\": \"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored but does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms."}, {"path": ["3", "text"], "text": "At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 210 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms.", "negative_left": "At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms.", "negative_right": "At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 140 kilograms.", "right": "At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 210 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-040", "id": "fast-43-diverse-294-040-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the first-lot resin requirement for changeover L4 (Jar-A to Jar-B) per work order WO-7719 is 180 kilograms. The reserve delivery has slipped from 08:00 to 10:00 and is being tracked as a documented issue that could disrupt continued production."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test. Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Materials clerk", "text": "At 06:52, the resin staged at station L4 under batch tag B-7719 weighs 140 kilograms."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored but does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates and the quality inspector's nonblocking-concern language exactly as required by the unchanged question, with only the staged-resin quantity altered (260kg to 190kg) to flip that gate below the 240kg requirement, which is a coherent single-fact change with no contradictory duplicate measurements; the two focus evidence spans are complete factual sentences reporting the work-order requirement and staged amount, not policy text; and neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00, but the shortfall does not affect tooling or staffing plans.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms. At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 260 kilograms.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin remains a documented, nonblocking concern that could disrupt continued production, but it does not itself block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["4", "text"], "text": "At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms."}, {"path": ["4", "text"], "text": "At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 260 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms.", "negative_left": "At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms.", "negative_right": "At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 190 kilograms.", "right": "At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 260 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-041", "id": "fast-43-diverse-294-041-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00, but the shortfall does not affect tooling or staffing plans."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Materials clerk", "text": "At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms. At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 260 kilograms."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin remains a documented, nonblocking concern that could disrupt continued production, but it does not itself block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all six mandatory gates and the quality inspector's nonblocking-concern language exactly as required by the unchanged question, with only the staged-resin quantity altered (260kg to 190kg) to flip that gate below the 240kg requirement, which is a coherent single-fact change with no contradictory duplicate measurements; the two focus evidence spans are complete factual sentences reporting the work-order requirement and staged amount, not policy text; and neither context contains a gold answer, rule table, or output instruction.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00, but the shortfall does not affect tooling or staffing plans.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold.\"},{\"speaker\":\"Materials clerk\",\"text\":\"At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms. At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 260 kilograms.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin remains a documented, nonblocking concern that could disrupt continued production, but it does not itself block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["4", "text"], "text": "At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms."}, {"path": ["4", "text"], "text": "At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 260 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms.", "negative_left": "At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms.", "negative_right": "At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 190 kilograms.", "right": "At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 260 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-041", "id": "fast-43-diverse-294-041-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. The reserve delivery has slipped from 08:00 to 10:00, but the shortfall does not affect tooling or staffing plans."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present."}, {"speaker": "Maintenance technician", "text": "At 06:28, I placed a stop-work hold because the mold-clamp proximity switch failed its test."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold."}, {"speaker": "Materials clerk", "text": "At 06:52, changeover L4's first-lot resin requirement is recorded on work order W-118 as 240 kilograms. At 06:52, the resin staged at the L4 line for the Jar-A to Jar-B changeover measures 190 kilograms."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin remains a documented, nonblocking concern that could disrupt continued production, but it does not itself block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all mandatory gate facts (tooling, cleaning record, operators, stop-work clearance, quality approval) alongside the unchanged question criteria, with only the material figures altered as permitted case observations; the counterfactual changes staged resin from 200kg to 150kg, a single coherent factual edit that flips the material-sufficiency gate without contradicting other timestamped facts; both focus evidence spans are complete factual sentences with exact quotes ('requires 180 kilograms of resin.' and 'logged at 200 kilograms.'/'logged at 150 kilograms.'); no rule tables, gold labels, or output instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00, and this delay is noted in the shift log as a documented concern that could disrupt continued production if not resolved before the reserve stock is needed.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present and remain on shift at 06:52.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold. No stop-work hold currently applies to changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. At 06:52, the resin staged for changeover L4 is logged at 200 kilograms. The delayed reserve resin should be monitored as it does not block release but remains a documented open issue.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin."}, {"path": ["3", "text"], "text": "At 06:52, the resin staged for changeover L4 is logged at 200 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin.", "negative_left": "At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin.", "negative_right": "At 06:52, the resin staged for changeover L4 is logged at 150 kilograms.", "right": "At 06:52, the resin staged for changeover L4 is logged at 200 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-042", "id": "fast-43-diverse-294-042-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00, and this delay is noted in the shift log as a documented concern that could disrupt continued production if not resolved before the reserve stock is needed."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present and remain on shift at 06:52."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold. No stop-work hold currently applies to changeover L4."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. At 06:52, the resin staged for changeover L4 is logged at 200 kilograms. The delayed reserve resin should be monitored as it does not block release but remains a documented open issue."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts preserve all mandatory gate facts (tooling, cleaning record, operators, stop-work clearance, quality approval) alongside the unchanged question criteria, with only the material figures altered as permitted case observations; the counterfactual changes staged resin from 200kg to 150kg, a single coherent factual edit that flips the material-sufficiency gate without contradicting other timestamped facts; both focus evidence spans are complete factual sentences with exact quotes ('requires 180 kilograms of resin.' and 'logged at 200 kilograms.'/'logged at 150 kilograms.'); no rule tables, gold labels, or output instructions are embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00, and this delay is noted in the shift log as a documented concern that could disrupt continued production if not resolved before the reserve stock is needed.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present and remain on shift at 06:52.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold. No stop-work hold currently applies to changeover L4.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. At 06:52, the resin staged for changeover L4 is logged at 200 kilograms. The delayed reserve resin should be monitored as it does not block release but remains a documented open issue.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin."}, {"path": ["3", "text"], "text": "At 06:52, the resin staged for changeover L4 is logged at 200 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin.", "negative_left": "At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin.", "negative_right": "At 06:52, the resin staged for changeover L4 is logged at 150 kilograms.", "right": "At 06:52, the resin staged for changeover L4 is logged at 200 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-042", "id": "fast-43-diverse-294-042-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, the first lot for changeover L4 (Jar-A to Jar-B) requires 180 kilograms of resin. The reserve delivery has slipped from 08:00 to 10:00, and this delay is noted in the shift log as a documented concern that could disrupt continued production if not resolved before the reserve stock is needed."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present and remain on shift at 06:52."}, {"speaker": "Maintenance technician", "text": "Update at 06:46: I replaced the switch, completed three successful dry cycles, and formally cleared the stop-work hold. No stop-work hold currently applies to changeover L4."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. At 06:52, the resin staged for changeover L4 is logged at 150 kilograms. The delayed reserve resin should be monitored as it does not block release but remains a documented open issue."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all mandatory gate facts and question bindings (L4, Jar-A→Jar-B, BT-772) while only the staged resin quantity changes from 200kg to 140kg, which is a permissible case observation change; the two focus sentences are complete factual statements, not policy text; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772. The reserve delivery has slipped from 08:00 to 10:00, which is documented as a monitoring concern that could disrupt continued production.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present. At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 200 kilograms.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:46, I replaced the faulty proximity switch, completed three successful dry cycles, and formally cleared the earlier stop-work hold; no stop-work hold currently applies.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored but does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772."}, {"path": ["1", "text"], "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 200 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772.", "negative_left": "At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772.", "negative_right": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 140 kilograms.", "right": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 200 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-058", "id": "fast-43-diverse-294-058-base", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772. The reserve delivery has slipped from 08:00 to 10:00, which is documented as a monitoring concern that could disrupt continued production."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present. At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 200 kilograms."}, {"speaker": "Maintenance technician", "text": "At 06:46, I replaced the faulty proximity switch, completed three successful dry cycles, and formally cleared the earlier stop-work hold; no stop-work hold currently applies."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored but does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 1}, "source_family": "supply_chain-04", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain all mandatory gate facts and question bindings (L4, Jar-A→Jar-B, BT-772) while only the staged resin quantity changes from 200kg to 140kg, which is a permissible case observation change; the two focus sentences are complete factual statements, not policy text; no gold answer, rule table, or instruction is embedded in either context.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "full_context_fact_states": {"base": {"a1": "supported", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "counterfactual": {"a1": "refuted", "a2": "supported", "a3": "supported", "a4": "supported", "a5": "supported", "a6": "supported", "a7": "supported", "a8": "supported", "a9": "refuted"}, "remove_left": {"a1": "unknown"}, "remove_right": {"a1": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"a1": "unknown"}, "negative_pair": {"a1": "refuted"}, "negative_sentence": {"a1": "unknown"}, "positive_pair": {"a1": "supported"}, "right": {"a1": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Each atom states a single factual relationship; a9 concerns the factual release-blocking status of the delayed delivery rather than assigning a final readiness level. The focus a1 is a factual quantity comparison. The base and counter assignments differ only on a1 and are realizable: first-lot material can become insufficient for a reason independent of the documented reserve-delivery concern, while the remaining facts stay fixed. Empty policy_evidence is correct because all governing criteria, timing rules, gates, and routing instructions are already retained in original_input.questions; the state adds case observations rather than additional policy.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The conjunction establishes all six mandatory gates: sufficient first-lot material, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. It also establishes a documented issue that could disrupt continued production and explicitly refutes that the issue blocks release, which is sufficient for level 1.", "rule_index": 0, "sound": true}, {"reason": "Refuting a1 establishes that the mandatory first-lot-material gate is unmet. One unmet mandatory gate is sufficient for level 0 regardless of the status of the other gates or the nonblocking concern.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "a1", "statement": "As of 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is at least the resin quantity required for its first lot."}, {"id": "a2", "statement": "As of 06:52, the Jar-B mold is installed for changeover L4."}, {"id": "a3", "statement": "As of 06:52, cleaning record C-441 for changeover L4 is signed."}, {"id": "a4", "statement": "As of 06:52, four required operators are present for changeover L4."}, {"id": "a5", "statement": "As of 06:52, quality has approved the setup sample for release of changeover L4."}, {"id": "a6", "statement": "As of 06:52, no current stop-work hold applies to changeover L4."}, {"id": "a7", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 is documented as an issue."}, {"id": "a8", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 could disrupt continued production."}, {"id": "a9", "statement": "As of 06:52, the delayed reserve-resin delivery for changeover L4 blocks line release."}], "base_state_json": "[{\"speaker\":\"Production scheduler\",\"text\":\"At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772. The reserve delivery has slipped from 08:00 to 10:00, which is documented as a monitoring concern that could disrupt continued production.\"},{\"speaker\":\"Line supervisor\",\"text\":\"At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present. At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 200 kilograms.\"},{\"speaker\":\"Maintenance technician\",\"text\":\"At 06:46, I replaced the faulty proximity switch, completed three successful dry cycles, and formally cleared the earlier stop-work hold; no stop-work hold currently applies.\"},{\"speaker\":\"Quality inspector\",\"text\":\"At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored but does not block release.\"}]", "base_states": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "counter_states": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}], "focus_atom": "a1", "focus_evidence": [{"path": ["0", "text"], "text": "At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772."}, {"path": ["1", "text"], "text": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 200 kilograms."}], "policy_evidence": [], "rules": [{"justification": "Every mandatory gate is met, while the documented reserve-resin issue is nonblocking but could disrupt continued production.", "target": "1", "when": [{"atom_id": "a1", "state": "supported"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}, {"justification": "The first-lot material gate is unmet, so at least one mandatory release gate remains unmet even though the other gates are met.", "target": "0", "when": [{"atom_id": "a1", "state": "refuted"}, {"atom_id": "a2", "state": "supported"}, {"atom_id": "a3", "state": "supported"}, {"atom_id": "a4", "state": "supported"}, {"atom_id": "a5", "state": "supported"}, {"atom_id": "a6", "state": "supported"}, {"atom_id": "a7", "state": "supported"}, {"atom_id": "a8", "state": "supported"}, {"atom_id": "a9", "state": "refuted"}]}]}, "verified_pair": {"left": "At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772.", "negative_left": "At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772.", "negative_right": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 140 kilograms.", "right": "At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 200 kilograms."}, "verifier_independent_model": false}, "family": "fast-43-diverse-294-058", "id": "fast-43-diverse-294-058-counterfactual", "input": {"questions": {"decision": {"criteria": ["0 — Not ready: At least one mandatory release gate remains unmet under the latest evidence. Do not release the line; route to the function responsible for the unresolved gate.", "1 — Ready with controlled delay risk: Every mandatory release gate is met, but a documented nonblocking issue could disrupt continued production. Release the line and route the concern to materials support for monitoring.", "2 — Fully ready: Every mandatory release gate is met and no documented nonblocking concern remains. Release the line with no support routing required."], "instructions": "Determine changeover readiness as of 06:52. Later timestamped updates supersede earlier reports about the same issue. Mandatory release gates are: enough material for the first lot, installed tooling, signed cleaning record, four operators, quality approval, and no current stop-work hold. Select the single matching level; its stated routing and release decision apply.", "type": "score"}}, "state": [{"speaker": "Production scheduler", "text": "At 06:10, changeover L4 from Jar-A to Jar-B remained scheduled for 07:00. At 06:52, changeover L4's first-lot resin requirement is logged as 180 kilograms in batch ticket BT-772. The reserve delivery has slipped from 08:00 to 10:00, which is documented as a monitoring concern that could disrupt continued production."}, {"speaker": "Line supervisor", "text": "At 06:20, the Jar-B mold was installed, cleaning record C-441 was signed, and all four required operators were present. At 06:52, the resin staged for changeover L4 from Jar-A to Jar-B is measured at 140 kilograms."}, {"speaker": "Maintenance technician", "text": "At 06:46, I replaced the faulty proximity switch, completed three successful dry cycles, and formally cleared the earlier stop-work hold; no stop-work hold currently applies."}, {"speaker": "Quality inspector", "text": "At 06:52, I verified C-441 and approved the setup sample for line release. The delayed reserve resin should be monitored but does not block release."}]}, "method": "c2d", "provenance": {"source_id": "diverse-294", "source_is_synthetic": true, "source_sha256": "a68dbe6da3ce152523d40abc63a93d15256dd93a89c08ce35fd95e4554f32f81", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": 0}, "source_family": "supply_chain-04", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the stop instruction and all governing policy, only vary the alternate-delivery request timing relative to the 16:10 decision, keep entity/path bindings intact, use two complete factual sentences as evidence, and contain no gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present at Stop 18 for FM-297. Depot operations lead Priya confirms secure hold space is available if needed. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10. At 15:50, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10."}, {"path": [], "text": "At 15:50, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10.", "negative_left": "Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10.", "negative_right": "At 16:30, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system.", "right": "At 15:50, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system."}, "verifier_independent_model": false}, "family": "fast-43-diverse-297-003", "id": "fast-43-diverse-297-003-base", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present at Stop 18 for FM-297. Depot operations lead Priya confirms secure hold space is available if needed. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10. At 15:50, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the stop instruction and all governing policy, only vary the alternate-delivery request timing relative to the 16:10 decision, keep entity/path bindings intact, use two complete factual sentences as evidence, and contain no gold answer or rule table.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present at Stop 18 for FM-297. Depot operations lead Priya confirms secure hold space is available if needed. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10. At 15:50, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10."}, {"path": [], "text": "At 15:50, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10.", "negative_left": "Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10.", "negative_right": "At 16:30, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system.", "right": "At 15:50, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system."}, "verifier_independent_model": false}, "family": "fast-43-diverse-297-003", "id": "fast-43-diverse-297-003-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present at Stop 18 for FM-297. Depot operations lead Priya confirms secure hold space is available if needed. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Dispatcher Mina's routing decision for parcel FM-297 was made at 16:10. At 16:30, the customer submitted an alternate-delivery request for parcel FM-297, logged in the tracking system."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the stop instruction and question criteria intact, only varying the timing of the alternate-delivery request, which is a legitimate factual variable already referenced in the question's exception clause, and the evidence spans are two complete factual sentences without embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day. The customer submitted an alternate-delivery request for parcel FM-297 at 15:50 on the delivery day.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day."}, {"path": [], "text": "The customer submitted an alternate-delivery request for parcel FM-297 at 15:50 on the delivery day."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day.", "negative_left": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day.", "negative_right": "The customer submitted an alternate-delivery request for parcel FM-297 at 16:20 on the delivery day.", "right": "The customer submitted an alternate-delivery request for parcel FM-297 at 15:50 on the delivery day."}, "verifier_independent_model": false}, "family": "fast-43-diverse-297-007", "id": "fast-43-diverse-297-007-base", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day. The customer submitted an alternate-delivery request for parcel FM-297 at 15:50 on the delivery day."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the stop instruction and question criteria intact, only varying the timing of the alternate-delivery request, which is a legitimate factual variable already referenced in the question's exception clause, and the evidence spans are two complete factual sentences without embedded rules or answers.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day. The customer submitted an alternate-delivery request for parcel FM-297 at 15:50 on the delivery day.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day."}, {"path": [], "text": "The customer submitted an alternate-delivery request for parcel FM-297 at 15:50 on the delivery day."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day.", "negative_left": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day.", "negative_right": "The customer submitted an alternate-delivery request for parcel FM-297 at 16:20 on the delivery day.", "right": "The customer submitted an alternate-delivery request for parcel FM-297 at 15:50 on the delivery day."}, "verifier_independent_model": false}, "family": "fast-43-diverse-297-007", "id": "fast-43-diverse-297-007-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 16:05 on the delivery day. The customer submitted an alternate-delivery request for parcel FM-297 at 16:20 on the delivery day."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original stop instruction and all other policy language via the unchanged questions object, only the factual timing of the alternate-delivery request differs (09:10 vs 13:00) relative to the fixed 11:00 decision, which is a coherent single-fact counterfactual change; the two focus evidence sentences are plain factual statements about decision and request timing, not policy text; and neither context reveals a gold answer, code, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: \\u201cIf the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.\\u201d The 15:42 scan reads \\u2018attempted\\u2014recipient unavailable.\\u2019 Joel notes that nobody answered and no authorized neighbor was present. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3. An alternate-delivery request for parcel FM-297 was logged at 09:10 on May 3.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3."}, {"path": [], "text": "An alternate-delivery request for parcel FM-297 was logged at 09:10 on May 3."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3.", "negative_left": "Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3.", "negative_right": "An alternate-delivery request for parcel FM-297 was logged at 13:00 on May 3.", "right": "An alternate-delivery request for parcel FM-297 was logged at 09:10 on May 3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-297-008", "id": "fast-43-diverse-297-008-base", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads ‘attempted—recipient unavailable.’ Joel notes that nobody answered and no authorized neighbor was present. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3. An alternate-delivery request for parcel FM-297 was logged at 09:10 on May 3."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the original stop instruction and all other policy language via the unchanged questions object, only the factual timing of the alternate-delivery request differs (09:10 vs 13:00) relative to the fixed 11:00 decision, which is a coherent single-fact counterfactual change; the two focus evidence sentences are plain factual statements about decision and request timing, not policy text; and neither context reveals a gold answer, code, or rationale.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: \\u201cIf the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.\\u201d The 15:42 scan reads \\u2018attempted\\u2014recipient unavailable.\\u2019 Joel notes that nobody answered and no authorized neighbor was present. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3. An alternate-delivery request for parcel FM-297 was logged at 09:10 on May 3.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3."}, {"path": [], "text": "An alternate-delivery request for parcel FM-297 was logged at 09:10 on May 3."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3.", "negative_left": "Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3.", "negative_right": "An alternate-delivery request for parcel FM-297 was logged at 13:00 on May 3.", "right": "An alternate-delivery request for parcel FM-297 was logged at 09:10 on May 3."}, "verifier_independent_model": false}, "family": "fast-43-diverse-297-008", "id": "fast-43-diverse-297-008-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads ‘attempted—recipient unavailable.’ Joel notes that nobody answered and no authorized neighbor was present. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 11:00 on May 3. An alternate-delivery request for parcel FM-297 was logged at 13:00 on May 3."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the stop instruction and all governing policy via the unchanged question; only the alternate-delivery request timing (15:50 vs 16:20 relative to the 16:05 decision) changes as a case fact, which is coherent and not contradictory; evidence spans are two complete factual sentences with no policy text; no gold answer, rule table, or output instruction is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: \\u201cIf the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.\\u201d The 15:42 scan reads \\u201cattempted\\u2014recipient unavailable.\\u201d Joel notes that nobody answered and no authorized neighbor was present, confirming recipient unavailability. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 16:05. Customer service agent Leon logs incoming requests as they arrive and checks the queue before each routing decision. An alternate-delivery request for parcel FM-297 was logged at 15:50.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05."}, {"path": [], "text": "An alternate-delivery request for parcel FM-297 was logged at 15:50."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05.", "negative_left": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05.", "negative_right": "An alternate-delivery request for parcel FM-297 was logged at 16:20.", "right": "An alternate-delivery request for parcel FM-297 was logged at 15:50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-297-009", "id": "fast-43-diverse-297-009-base", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present, confirming recipient unavailability. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 16:05. Customer service agent Leon logs incoming requests as they arrive and checks the queue before each routing decision. An alternate-delivery request for parcel FM-297 was logged at 15:50."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts retain the stop instruction and all governing policy via the unchanged question; only the alternate-delivery request timing (15:50 vs 16:20 relative to the 16:05 decision) changes as a case fact, which is coherent and not contradictory; evidence spans are two complete factual sentences with no policy text; no gold answer, rule table, or output instruction is embedded.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "full_context_fact_states": {"base": {"alternate_request_exists": "supported", "recipient_unavailable": "supported"}, "counterfactual": {"alternate_request_exists": "refuted", "recipient_unavailable": "supported"}, "remove_left": {"alternate_request_exists": "unknown"}, "remove_right": {"alternate_request_exists": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"alternate_request_exists": "unknown"}, "negative_pair": {"alternate_request_exists": "refuted"}, "negative_sentence": {"alternate_request_exists": "unknown"}, "positive_pair": {"alternate_request_exists": "supported"}, "right": {"alternate_request_exists": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms state single factual relationships, and the focus is the factual existence of an alternate-delivery request rather than a policy conclusion. The base and counter assignments are jointly realizable and differ only in the focus atom. The cited state evidence preserves the substantive stop instruction originating in the original state; the exception and decision criteria already remain available verbatim in the questions object and therefore need not be repeated in policy_evidence. The two-rule partial table correctly covers the relevant exception contrast without treating unknown as false.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "Recipient unavailability together with an existing alternate-delivery request satisfies the question’s explicit exception to the otherwise mandatory depot-hold/no-redelivery instruction. No competing outcome condition stated in the policy remains unexcluded.", "rule_index": 0, "sound": true}, {"reason": "Recipient unavailability with the alternate-delivery request refuted directly triggers the stated prohibition on redelivery and requirement for depot hold.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "recipient_unavailable", "statement": "At the 15:42 Stop 18 delivery attempt for parcel FM-297, the named recipient was unavailable."}, {"id": "alternate_request_exists", "statement": "Before dispatcher Mina's routing decision for parcel FM-297, an alternate-delivery request for parcel FM-297 existed."}], "base_state_json": "\"Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: \\u201cIf the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.\\u201d The 15:42 scan reads \\u201cattempted\\u2014recipient unavailable.\\u201d Joel notes that nobody answered and no authorized neighbor was present, confirming recipient unavailability. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 16:05. Customer service agent Leon logs incoming requests as they arrive and checks the queue before each routing decision. An alternate-delivery request for parcel FM-297 was logged at 15:50.\"", "base_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}], "counter_states": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}], "focus_atom": "alternate_request_exists", "focus_evidence": [{"path": [], "text": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05."}, {"path": [], "text": "An alternate-delivery request for parcel FM-297 was logged at 15:50."}], "policy_evidence": [{"path": [], "text": "The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.”"}], "rules": [{"justification": "For recipient unavailability, the instruction prohibits redelivery unless an alternate-delivery request exists. With that request established for FM-297, redelivery is the stated exception route.", "target": "true", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "supported"}]}, {"justification": "When the named recipient was unavailable and no alternate-delivery request existed for FM-297, the stop instruction requires depot hold and prohibits redelivery.", "target": "false", "when": [{"atom_id": "recipient_unavailable", "state": "supported"}, {"atom_id": "alternate_request_exists", "state": "refuted"}]}]}, "verified_pair": {"left": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05.", "negative_left": "Dispatcher Mina made her routing decision for parcel FM-297 at 16:05.", "negative_right": "An alternate-delivery request for parcel FM-297 was logged at 16:20.", "right": "An alternate-delivery request for parcel FM-297 was logged at 15:50."}, "verifier_independent_model": false}, "family": "fast-43-diverse-297-009", "id": "fast-43-diverse-297-009-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "No means redelivery is prohibited or unsupported and another route, here depot hold, is required.", "true": "Yes means the record and stated policy make redelivery the correct exception route."}, "instructions": "Decide whether this exception should be routed for redelivery. Apply the stop instruction exactly: recipient unavailability requires depot hold and explicitly prohibits redelivery unless an alternate-delivery request exists. Answer yes or no.", "type": "noul"}}, "state": "Route dispatcher Mina reviews parcel FM-297 after driver Joel returns from Stop 18. The stop instruction says: “If the named recipient is unavailable, do not schedule redelivery; place the parcel on depot hold for customer pickup.” The 15:42 scan reads “attempted—recipient unavailable.” Joel notes that nobody answered and no authorized neighbor was present, confirming recipient unavailability. The address is complete, the van had sufficient capacity, and the sealed carton is not damaged. Depot operations lead Priya confirms secure hold space is available. Dispatcher Mina made her routing decision for parcel FM-297 at 16:05. Customer service agent Leon logs incoming requests as they arrive and checks the queue before each routing decision. An alternate-delivery request for parcel FM-297 was logged at 16:20."}, "method": "c2d", "provenance": {"source_id": "diverse-297", "source_is_synthetic": true, "source_sha256": "f492d8656565e6612cc41594f1aac2292ee33c68c56c733aeb4596c4dccce082", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same policy, entities, route, and address; the focus evidence sentences are complete factual statements about a conversation and its content; the counterfactual coherently changes Alvarez's statement to an unconfirmed status without contradicting other unchanged facts; no evidence sentence states the final classification or rule outcome, only that the dispatcher's instruction 'applies to this stop', which is a relevance statement not an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"building_staff_confirmed_move": "supported", "instruction_applies": "supported"}, "full_context_fact_states": {"base": {"building_staff_confirmed_move": "supported", "instruction_applies": "supported"}, "counterfactual": {"building_staff_confirmed_move": "refuted", "instruction_applies": "supported"}, "remove_left": {"building_staff_confirmed_move": "unknown"}, "remove_right": {"building_staff_confirmed_move": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"building_staff_confirmed_move": "unknown"}, "negative_pair": {"building_staff_confirmed_move": "refuted"}, "negative_sentence": {"building_staff_confirmed_move": "unknown"}, "positive_pair": {"building_staff_confirmed_move": "supported"}, "right": {"building_staff_confirmed_move": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express individual factual relationships. The focus concerns whether building staff confirmed the move, not a policy conclusion. The base and counter assignments can be realized while keeping the instruction applicable and changing only whether the required confirmation occurred. Policy evidence preserves the substantive state-originating depot rule and the dispatcher’s conditional instruction; instructions and criteria in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The rule requires both that the dispatcher’s conditional instruction applies and that building staff confirmed the specified move from Unit 4B. Under the preserved depot policy and instruction, this is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the building-staff confirmation entails that the dispatcher’s stated condition was not confirmed. The unchanged question explicitly requires the false outcome when that condition was not confirmed, so no additional ordinary-redelivery facts are necessary.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "instruction_applies", "statement": "For package ZX-418 on Route 12 at the unsuccessful-delivery incident for 8 Harbor Lane, Unit 4B, the route dispatcher’s conditional instruction applies."}, {"id": "building_staff_confirmed_move", "statement": "At the unsuccessful-delivery incident for package ZX-418, a building staff member confirmed to the delivery driver that the recipient had moved from Unit 4B."}], "base_state_json": "{\"context\":\"Package ZX-418, Route 12, destined for 8 Harbor Lane, Unit 4B, was flagged for exception handling per dispatcher instruction. Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery.\",\"evidence\":[\"Route dispatcher: \\u201cIf building staff confirms the recipient moved from 4B, do not retry; request address clarification.\\u201d\",\"The dispatcher's conditional instruction applies to this stop.\",\"During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM.\",\"Mr. Alvarez told the driver that the recipient had moved out of Unit 4B.\",\"Scan events show arrival and unsuccessful delivery, with no completion scan.\",\"Depot operations lead recorded intact packaging and no vehicle-capacity constraint.\"],\"request\":\"Under depot policy, should this exception be routed to address clarification rather than redelivery?\"}", "base_states": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "supported"}], "counter_states": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "refuted"}], "focus_atom": "building_staff_confirmed_move", "focus_evidence": [{"path": ["evidence", "2"], "text": "During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM."}, {"path": ["evidence", "3"], "text": "Mr. Alvarez told the driver that the recipient had moved out of Unit 4B."}], "policy_evidence": [{"path": ["request"], "text": "Under depot policy, should this exception be routed to address clarification rather than redelivery?"}, {"path": ["context"], "text": "Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery."}, {"path": ["evidence", "0"], "text": "Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”"}], "rules": [{"justification": "The applicable dispatcher instruction’s stated address-conflict condition is confirmed, so depot policy routes the exception to address clarification.", "target": "true", "when": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "supported"}]}, {"justification": "The applicable dispatcher instruction requires confirmation by building staff that the recipient moved from Unit 4B; explicit refutation of that confirmation means the condition was not confirmed.", "target": "false", "when": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "refuted"}]}]}, "verified_pair": {"left": "During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM.", "negative_left": "During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM.", "negative_right": "Mr. Alvarez told the driver that he did not know whether the recipient still lived in Unit 4B.", "right": "Mr. Alvarez told the driver that the recipient had moved out of Unit 4B."}, "verifier_independent_model": false}, "family": "fast-43-diverse-298-010", "id": "fast-43-diverse-298-010-base", "input": {"questions": {"decision": {"criteria": {"false": "Route the incident for redelivery because the dispatcher’s address-conflict condition was not confirmed.", "true": "Route the incident to address clarification because the condition in the dispatcher’s instruction was confirmed."}, "instructions": "Answer yes if the dispatcher’s conditional address-conflict instruction was activated by confirmed evidence. Answer no if that condition was not confirmed or if the record instead supports ordinary redelivery.", "type": "noul"}}, "state": {"context": "Package ZX-418, Route 12, destined for 8 Harbor Lane, Unit 4B, was flagged for exception handling per dispatcher instruction. Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery.", "evidence": ["Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”", "The dispatcher's conditional instruction applies to this stop.", "During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM.", "Mr. Alvarez told the driver that the recipient had moved out of Unit 4B.", "Scan events show arrival and unsuccessful delivery, with no completion scan.", "Depot operations lead recorded intact packaging and no vehicle-capacity constraint."], "request": "Under depot policy, should this exception be routed to address clarification rather than redelivery?"}}, "method": "c2d", "provenance": {"source_id": "diverse-298", "source_is_synthetic": true, "source_sha256": "d5c7c6aa5962d6762278e5ae977fe87f27888823aa0a70204e56364a3bea30d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": true}, "source_family": "supply_chain-05", "split": "train", "variant": "base"} {"domain": "supply_chain", "evidence_certificate": {"context_audit": {"counterfactual_is_coherent": true, "evidence_is_two_factual_sentences": true, "explanation": "Both contexts keep the same policy, entities, route, and address; the focus evidence sentences are complete factual statements about a conversation and its content; the counterfactual coherently changes Alvarez's statement to an unconfirmed status without contradicting other unchanged facts; no evidence sentence states the final classification or rule outcome, only that the dispatcher's instruction 'applies to this stop', which is a relevance statement not an answer.", "no_answer_leakage": true, "policy_preserved": true, "question_bindings_preserved": true}, "fact_states": {"building_staff_confirmed_move": "refuted", "instruction_applies": "supported"}, "full_context_fact_states": {"base": {"building_staff_confirmed_move": "supported", "instruction_applies": "supported"}, "counterfactual": {"building_staff_confirmed_move": "refuted", "instruction_applies": "supported"}, "remove_left": {"building_staff_confirmed_move": "unknown"}, "remove_right": {"building_staff_confirmed_move": "unknown"}}, "necessity_checks_passed": true, "pair_fact_states": {"left": {"building_staff_confirmed_move": "unknown"}, "negative_pair": {"building_staff_confirmed_move": "refuted"}, "negative_sentence": {"building_staff_confirmed_move": "unknown"}, "positive_pair": {"building_staff_confirmed_move": "supported"}, "right": {"building_staff_confirmed_move": "unknown"}}, "pipeline_version": "evidence-curation-v3", "rule_audit": {"assignments_are_realizable": true, "atoms_are_atomic": true, "explanation": "Both atoms express individual factual relationships. The focus concerns whether building staff confirmed the move, not a policy conclusion. The base and counter assignments can be realized while keeping the instruction applicable and changing only whether the required confirmation occurred. Policy evidence preserves the substantive state-originating depot rule and the dispatcher’s conditional instruction; instructions and criteria in the questions object need not be duplicated.", "focus_is_factual": true, "policy_complete": true, "rule_checks": [{"reason": "The rule requires both that the dispatcher’s conditional instruction applies and that building staff confirmed the specified move from Unit 4B. Under the preserved depot policy and instruction, this is sufficient for the true outcome.", "rule_index": 0, "sound": true}, {"reason": "Refutation of the building-staff confirmation entails that the dispatcher’s stated condition was not confirmed. The unchanged question explicitly requires the false outcome when that condition was not confirmed, so no additional ordinary-redelivery facts are necessary.", "rule_index": 1, "sound": true}]}, "spec": {"atoms": [{"id": "instruction_applies", "statement": "For package ZX-418 on Route 12 at the unsuccessful-delivery incident for 8 Harbor Lane, Unit 4B, the route dispatcher’s conditional instruction applies."}, {"id": "building_staff_confirmed_move", "statement": "At the unsuccessful-delivery incident for package ZX-418, a building staff member confirmed to the delivery driver that the recipient had moved from Unit 4B."}], "base_state_json": "{\"context\":\"Package ZX-418, Route 12, destined for 8 Harbor Lane, Unit 4B, was flagged for exception handling per dispatcher instruction. Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery.\",\"evidence\":[\"Route dispatcher: \\u201cIf building staff confirms the recipient moved from 4B, do not retry; request address clarification.\\u201d\",\"The dispatcher's conditional instruction applies to this stop.\",\"During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM.\",\"Mr. Alvarez told the driver that the recipient had moved out of Unit 4B.\",\"Scan events show arrival and unsuccessful delivery, with no completion scan.\",\"Depot operations lead recorded intact packaging and no vehicle-capacity constraint.\"],\"request\":\"Under depot policy, should this exception be routed to address clarification rather than redelivery?\"}", "base_states": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "supported"}], "counter_states": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "refuted"}], "focus_atom": "building_staff_confirmed_move", "focus_evidence": [{"path": ["evidence", "2"], "text": "During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM."}, {"path": ["evidence", "3"], "text": "Mr. Alvarez told the driver that the recipient had moved out of Unit 4B."}], "policy_evidence": [{"path": ["request"], "text": "Under depot policy, should this exception be routed to address clarification rather than redelivery?"}, {"path": ["context"], "text": "Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery."}, {"path": ["evidence", "0"], "text": "Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”"}], "rules": [{"justification": "The applicable dispatcher instruction’s stated address-conflict condition is confirmed, so depot policy routes the exception to address clarification.", "target": "true", "when": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "supported"}]}, {"justification": "The applicable dispatcher instruction requires confirmation by building staff that the recipient moved from Unit 4B; explicit refutation of that confirmation means the condition was not confirmed.", "target": "false", "when": [{"atom_id": "instruction_applies", "state": "supported"}, {"atom_id": "building_staff_confirmed_move", "state": "refuted"}]}]}, "verified_pair": {"left": "During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM.", "negative_left": "During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM.", "negative_right": "Mr. Alvarez told the driver that he did not know whether the recipient still lived in Unit 4B.", "right": "Mr. Alvarez told the driver that the recipient had moved out of Unit 4B."}, "verifier_independent_model": false}, "family": "fast-43-diverse-298-010", "id": "fast-43-diverse-298-010-counterfactual", "input": {"questions": {"decision": {"criteria": {"false": "Route the incident for redelivery because the dispatcher’s address-conflict condition was not confirmed.", "true": "Route the incident to address clarification because the condition in the dispatcher’s instruction was confirmed."}, "instructions": "Answer yes if the dispatcher’s conditional address-conflict instruction was activated by confirmed evidence. Answer no if that condition was not confirmed or if the record instead supports ordinary redelivery.", "type": "noul"}}, "state": {"context": "Package ZX-418, Route 12, destined for 8 Harbor Lane, Unit 4B, was flagged for exception handling per dispatcher instruction. Depot policy routes an exception to address clarification when a dispatcher’s conditional instruction applies and the stated address-conflict condition is confirmed; otherwise, an intact package with an unavailable recipient is routed for redelivery.", "evidence": ["Route dispatcher: “If building staff confirms the recipient moved from 4B, do not retry; request address clarification.”", "The dispatcher's conditional instruction applies to this stop.", "During the ZX-418 delivery attempt at 8 Harbor Lane, Unit 4B, the driver spoke with building staff member Mr. Alvarez at 3:15 PM.", "Mr. Alvarez told the driver that he did not know whether the recipient still lived in Unit 4B.", "Scan events show arrival and unsuccessful delivery, with no completion scan.", "Depot operations lead recorded intact packaging and no vehicle-capacity constraint."], "request": "Under depot policy, should this exception be routed to address clarification rather than redelivery?"}}, "method": "c2d", "provenance": {"source_id": "diverse-298", "source_is_synthetic": true, "source_sha256": "d5c7c6aa5962d6762278e5ae977fe87f27888823aa0a70204e56364a3bea30d8", "source_split": "train"}, "quality_status": "evidence_model_checked", "reference": {"human_reviewed": false, "source": "audited_rule_over_verified_facts", "target": false}, "source_family": "supply_chain-05", "split": "train", "variant": "counterfactual"}